Testing AI Image Generation Again: What Changed By 2026

Author: Shira Abel Category: Enterprise Marketing URL: https://hunterandbard.com/resources/blog/testing-3-different-ai-platforms-to-generate-images

Summary

We ran the same AI image test twice: once with an ordinary prompt, once with a real art-director brief. Same three models, minutes apart, very different results. In 2026 the brief decides the image.

TL;DR

Round one used a normal prompt and produced normal AI images: competent, polished, and quietly wrong on casting. Round two used a full art-director brief with named cast, lens behavior, practical lighting, and explicit negatives, and the same three models produced photographs. In 2026 the brief decides the image.

Full Article

You need one photograph for the landing page. The stock library options all look like a 2014 SaaS brochure, and the AI version comes out looking like an AI version. That is where most marketing teams sit with generated imagery right now.

The last time we tested AI image generation, the question was whether these tools could produce a single usable photo at all. That question is closed. So we ran the test differently: two rounds, three models, one scene. Round one uses the same prompt we used in 2024. Round two uses a prompt built for what these models can do today. Both rounds ran today, on the same models, so the comparison is fair.

The models are OpenAI's GPT Image, Google's Nano Banana (Gemini's image model), and Flux 2. The standard is unchanged: would I put this on a landing page for an event we are selling tickets to?

Round one: the ordinary brief

Editorial press photograph of a fintech networking reception: men and women in tailored suits standing in small groups talking, holding drinks, inside a glass-walled boardroom high in a Manhattan skyscraper, late afternoon light, mixed ages and ethnicities, candid documentary realism, shot on a full-frame camera at 35mm, natural color.

That is a fine prompt by 2024 standards. It names a subject, a place, a time of day, a camera, and a mood. One shot per model, no retries, and no cherry-picking.

OpenAI GPT Image

Round one, GPT Image: correct hands, real name badges, believable light.The most shippable of the three. Hands are correct for the most part, the badges read as badges instead of smeared text, and the light behaves like light coming through glass at 5pm. People are cropped, backs are turned, nobody is posing. The remaining tell is the grade, flat and cool in the way stock photography is flat. The skyline looks more like watercolor than a photo.

Google Nano Banana

Round one, Nano Banana: symmetrical, the outline of the people is an AI tell, and slightly too perfect to be a real photo.Technically the most beautiful image thanks to the skyline. The skyline is pretty without being detailed, the room is perfectly symmetrical, and every single face is lit. Real photographers do not get that lucky.

Flux 2

Round one, Flux 2: photographically convincing, and it quietly ignored half the brief.The least convincing grade of the three in terms of AI faces. The image itself is warm and slightly hazy, exactly like a room shot into backlight. Then look at who is in the room. The brief said men and women. Flux delivered a wall of men in suits with two women pushed to the edge of frame and out of focus. Ask for a demographic in the abstract and the model gives you the demographic it saw most in training.

Why the brief has to change

{{block:0}}Anatomy, text rendering, and photorealism used to consume the model's entire attention budget, so extra brief detail was wasted effort. Everything the model has stopped struggling with is capacity you can now spend on what still fails: casting, art direction, lens behavior, and grade.

Round one failed on exactly those axes, because nobody told the models what to do. The prompt needs details, such as a Black woman in her late 40s in a charcoal double-breasted suit, mid-sentence, gesturing with a wine glass. That sentence is a casting call. Mixed ages and ethnicities is a wish.

Round two: the art-director brief

Editorial documentary press photograph, candid, of a fintech networking reception in progress. Cast, exactly as described: a Black woman in her late 40s in a charcoal double-breasted suit, mid-sentence, gesturing with a wine glass; a South Asian man in his early 30s in a rumpled blue oxford shirt with no jacket, listening, head tilted; a white woman in her 60s with short silver hair in a burgundy knit dress, laughing with her eyes closed; an East Asian man in his 20s in a black turtleneck, half out of frame at the right edge, checking his phone. Two blurred people in the near foreground, backs of heads only, cropped by the frame edge. Location: a corner conference room on the 41st floor of an older Manhattan office tower, west-facing windows, a scuffed walnut credenza with abandoned glassware and a crumpled napkin, one ceiling downlight burned out. Time: 5:40pm in late September, low raking sun from camera left, warm practical light from a table lamp mixing with cool daylight. Camera: full-frame, 35mm, f/2, 1/60s, ISO 1600, handheld, slight motion blur on the gesturing hand, focus on the woman's face. Grade: Kodak Portra 400 push, warm skin, muted greens, visible fine grain. Negatives: no symmetry, no centered composition, no hero skyline, nobody looking at the camera, no stock-photo polish, no evenly lit faces.

Same three models, same rule: one prompt, one shot, no retries.

OpenAI GPT Image, round two

Round two, GPT Image: the cast arrived exactly as written, including the man cropped at the right edge.Every named person is present, in the right wardrobe, doing the right thing. The foreground heads are there. The credenza is there. The burned-out downlight is there. What the detail bought is control. This is the first frame that looks directed rather than sampled. The color is still slightly cooler and bluer than the warm film look the brief asked for, the composition is off-center, and the eyelines are correct.

Google Nano Banana, round two

Round two, Nano Banana: the most set-dressed frame, with the brass lamp and glassware exactly as briefed.Nano Banana is the best listener of the three. It picked up the props nobody else bothered with: the brass table lamp, the coupe glasses, the folded napkin, and the patterned rug. It also added background people the brief did not ask for, which reads as a room rather than a scene. The lingering weakness is the same one from round one. It still lights everything a little too kindly. Left to its instincts it wants to make the picture pretty, and it will fight the negatives to get there.

Flux 2, round two

Round two, Flux 2: named casting fixed the demographic default, and the film grade is still the best of the three.The interesting result. Round one's biggest failure, a room full of men, disappeared once the cast was named person by person. Flux has a defaults problem, and specificity solves it. The grade remains the most photographic of the three: the haze, the bloom off the window, and the grain in the shadows. Flux also ignored the requested aspect ratio and returned a square frame (it will follow every instruction except the one with a number in it), which keeps it the least controllable of the three.

What this means if you need AI photos for your marketing

In 2024 the question was whether AI imagery was usable at all. It is. The question now is what the image is doing for you, and how hard you are willing to direct it.

The same three models produced both rounds, minutes apart. What changed between them is how much of the work the person writing the prompt was willing to do. Your output is a function of your expertise, and the expertise is art direction.

If you want help with your messaging and positioning, talk to us.