Most AI-generated product shots have a tell within about two seconds: skin too smooth, hands slightly wrong, lighting that doesn't match the product, a pose no real photographer would ever call. A brand tries it once, gets something that reads as obviously synthetic, and goes back to booking real shoots.
Most of that gap traces back to the prompting process, not the image model. A single detailed prompt, even a good one, produces one lucky image and nine that miss. Consistent, natural-looking output takes a real production process behind it, the same prep a photographer runs before a shoot.
Brand name below is anonymized ("a founder-led apparel brand"). Every step, prompt structure, and image below is from real, paid client work.
The situation
- A founder-led apparel brand needed lifestyle and editorial photography for a product launch, on a timeline and budget that didn't fit a studio shoot
- An initial attempt at AI-generated images had already been tried and shelved, the output read as synthetic within a glance
- The founder had a specific editorial vision in mind (Vogue and Elle-adjacent), not a generic catalogue look
What they needed: campaign-ready photography without a shoot.
What I delivered: a 7-step generation process and a set of images that read as a real shoot on first look.
How I approached it
- Understood the ICP. Age, demographic, gender, ethnicity, colour palette, matched to the brand's actual customer, not a generic stock-model default.
- Built a model document. A single reference file locking the model's identity (face shape, body type, skin tone, styling) so every image pointed back to the same source.
- Documented the product. Dimensions, colours, SKUs, fits, cuts, angles, the same brief a real photographer would need before a shoot.
- Aligned on shot expectations. Pulled reference imagery from Vogue, Elle, and comparable editorial sources, then sat with the founder directly to confirm the vision before generating anything.
- Locked poses, props, lighting, and grading. Naturality lives in these details more than in the model's face.
- Generated with full context loaded. Every prompt carried the ICP, model document, product spec, and shot reference together, not in isolation.
- Iterated 3 to 4 rounds. Fine-tuned against founder feedback each round, usually on a detail (hand position, fabric fall, a light source), not the concept.

Skipping the model document was the costliest shortcut to test: without it, every image in a set looked like a different person wearing similar clothes, worse for usable output than one clearly synthetic image, since it couldn't be assembled into a consistent-feeling set at all. Skipping the shot-expectations step produced images that were technically well-lit and well-composed but didn't match what the founder had in mind, which meant the iteration rounds got spent on a concept disagreement instead of a detail fix.
What shipped
Generated, not photographed. No studio, no camera, no model on a call sheet.

Hard directional light, real shadow falloff on the stairs, a pose a photographer would actually direct rather than a default catalogue stance. The brief called for editorial, not catalogue, and step 5 is where that got enforced.

A genuine-feeling candid moment between two people is a harder ask than a single posed shot, expressions have to read as reactive, not performed. This is where iteration earned its keep: the first pass read a beat too posed.
What this doesn't claim
This covers the generation process and the images that shipped from it, not a claim that AI photography replaces a real shoot in every case. Motion-heavy campaign video and physical in-hand product detail still need a camera. Within the scope this process targets, catalogue and lifestyle stills, the output matched what a shoot would have delivered, without booking a studio, a model, or a location.
