This article was originally released in The Interline’s AI Report 2026.
Including profiles and exclusive interviews with 17 AI companies, stories and opinions from 12 different industry perspectives, and real survey feedback from around 100 voices from every level of fashion, bottled and analysed, The AI Report 2026 is essential reading for anyone interested in AI for fashion.

Image generation has been AI’s fastest and flashiest success story. 

I could mount an argument that the frontier language models released in the last six months represent the most fundamental change to what agentic AI is capable of, when it’s properly grounded, used, and reinforced with appropriate context. But the qualitative difference between a previous-generation LLM and the current state of the art requires someone to really sit down, then personalise, connect, and use it for deep work to perceive. At a glance, a textbox is still a textbox, even if the model behind it can now take on complex, multi-turn projects rather than just answering staccato questions.

In just over four years, though, image models have transformed at both the output level, in a way that’s immediately visible and arresting, and in how people interact with, experience, and use them. 

In early 2022, generating an image was a matter of invoking Midjourney Version 1 through Discord slash commands and then experimenting with the new concept of prompting – a series of steps that were alien to most end users.

In The Interline’s infancy, before we committed to assigning creative budget to new artistic  talent every year, the cover for our DPC Report 2022 was actually created this way. And while I struggle to remember as far back as this year’s DPC Report (two weekly podcasts take their toll on the working memory!), I do know that no part of that process felt as though it was a near-term threat to real artists or real photographers, or a complete overhaul of the what it means to put pixels in front of people.

It felt like an experiment in something different, rather than a replacement for a well-worn set of tools and workflows we already had. Which meant that the audience for those first image-gen models was made up of existing photographers, digital artists, and Photoshop professionals to some extent, but it was primarily people who’d never touched an image creation workflow before who were thrilled at what seemed like a crazy new set of possibilities.

But outside of a few savvy self-promoters (and some deluded ones as well) very few people using Midjourney Version 1 in a public channel thought they were about to become working photographers or creative fashion designers.

Times have changed.

That same 2022 report also contains the earliest analysis I personally wrote about image generation. That article, available here, has aged ok in some respects, but it’s accompanied by images that really capture that moment in time. They’re very crude, by today’s standards. It also includes, in the opening sting, a line of questioning I want to pick up on… because it seems both kind of prescient and very naive at once.

I asked: “is generative AI going to change anything fundamental about how digital workflows operate in fashion?”

You can laugh. With the benefit of hindsight, I did, too. That’s probably one of the biggest under-estimations I’ve ever committed to writing. Because, as everyone reading this knows, in the intervening four years fashion, along with every other visual-first medium, has seen the work of creating visual content, for either in-house teams or external parties, completely upended.

And, to be very clear: someone using an image generation model can, in a very direct sense, do the work of a designer or a photographer. They might not do it very well, which is an edge I’ll talk about in just a minute, but they can sit down, pay a fraction of a dollar, and end up with something that looks like it came out of a designer’s 2D CAD application, or a photographer’s pipeline.

Midjourney was a technological and cultural curio. What we have today is a fundamental change in how visual outputs are made, with a commensurate impact on our collective conception of who does the work, and why.

At the application level, generation and editing now lives as a native modality in popular chatbots and directly in photo apps, putting it on everyone’s device. 

There’s no better benchmark for how ubiquitous computational pipelines have become than the fact that you now need to go to the App Store and download a separate application, from a third party developer, if you want to see pictures that are assembled from just the lens and the sensor in your phone.

If you want to change the outfit someone was wearing in a candid capture, though? That’s baked into the first-party apps that ship with the latest versions of Android and iOS.

That’s a profound reversal of people’s relationship with photography and photo editing.

For fashion’s purposes, image generation now scarcely even resembles the public chatbot prompt, turn, re-prompt paradigm. It’s expressed in private multi-model, node-based professional workspaces that put a full complement of generative tools into applications and workflows that feel at once familiar and endlessly powerful. The developers of those tools are now working to re-integrate capabilities from traditional pixel editing applications into their platforms, to create a vision for fully end-to-end AI image workspaces, and there’s now a buoyant ecosystem of platform-makers across both industry-agnostic applications and more heavily verticalised, fashion-specific or footwear-specific usage.

But as potent as the shift in environment has been, the rapid maturation of the underlying models – which are used in both consumer and enterprise pipelines – is staggering to stop and consider. Four years ago, as you can see in the images we extracted from the 2022 Report that appear earlier in this article, we struggled to generate a running jacket that had the normal number of sleeves, and a shoe that didn’t have pieces of outsole sticking out at bizarre angles. These images are quaint today, but they took time to produce because the hit ratio of prompt to successful output was wildly out of proportion.

Again: times have changed. I can pause typing, open up one of the several different canvases available to me, and generate a fully photoreal-looking example of either product category, and it would cost me fifty cents or less, per image, at 4K output resolution, to do so. 

(To avoid any accusations of bias, The Interline does not accept free workspace seats or tokens for generative tools, and we pay directly for all licenses and usage billing.)

Perhaps even more importantly, I can then edit those generations (a misnomer to some extent, since every generation is still exactly that, a fresh, steerable generation rather than a true ‘edit’) with similar ease and at similarly low cost. If I wanted to replace a material or swap a colour on that running jacket, I could do it in under two minutes, with no technical skills, no traditional software or hardware, and no need to even watch the work happening; I can lodge the request, walk away, and come back to find the work done. 

And I can apply the same hands-off logic to batch processing, codified workflows, and other steps if I decide I’m willing to transfer more creative responsibility over to an orchestrating agent – either one that’s built into the generative workspace I’m already in, or one that exists in my everyday chatbot of choice, through MCP tool-calling.

If I want to take the products we could only loosely approximate using an early-stage image generation model in 2022, and build out a quick campaign using them as inputs, and a couple of modern image models as the execution layer, I could do that too.

So, as a fun exercise, I did. Separately, I also tasked other members of The Interline team with coming up with their own fake brands, their own synthetic models, and their own campaign shoots, and you’ll see the outputs of some of these accompanying this article – as well as being used to illustrate other stories [in The AI Report 2026].

In a very meaningful sense, it is now absolutely trivial to generate ‘photography’ of products that don’t exist, people who aren’t real, and locations that would be expensive or impractical to shoot in. 

Now, I’m not suggesting that any of what you see on this page is good enough to stand in for real creative direction. We’re just a publication playing around with the tools. Writing a prompt, or letting an agent do it for you, does not make you a photographer or a stylist or a location scout or anything else… but it still lets you create output that resembles photography. And if you’re good enough at it, that resemblance can now be very close.

There are, today, forward-deployed creatives working for major generative workspace providers who do fit that brief: they are formally-trained artists, stylists, photographers and people with groundwork in similar disciplines, who now act as pace-setters for the frontier of generative AI usage, and that frontier is pushing ever closer to direct substitution for traditional workflows by dint of being faster, cheaper, and available around the clock at unmatched scale.

And there are, too, retailers and brands who recognise this and have made generative AI a cornerstone of their content creation, from product detail page to full-blown campaigns. Some of these companies are shy about the prominence they’re placing on generative workflows, fearing consumer backlash, but the data contained in the survey portion of this report suggests that fear is dwindling. Other companies are more candid about it.

I’ve been pretty effusive about what I’m calling the “post-image era” here. And I think that’s reasonable as a pure technology analyst and observer. We can easily fall into cynicism when we’re talking about AI, but as someone who likes tech and understands the work that goes into exponentially improving it, I want to underline just how damned impressive image generation and generative editing is today.

So what’s the risk? As amazing as it seems to be able to just spin out visual outputs that are getting incredibly close to being indistinguishable from real photography, why do I consider this a cautionary moment as much as an intoxicating one?

For the first proviso, I’ll borrow Midjourney’s own language. The earliest versions of that model, as used in our 2022 generations, exhibited what the team there refer to as “low coherence”. This is part of the same AI lexicon as adherence, but while that word means how close the generation was to the prompt (how obedient the model was), coherence refers to the internal logic exhibited in the final generation, and in multi-prompt workflows it was, four years ago, effectively zero.

Both of those metrics have now improved massively. Image generators have found a good balance of prompt-interpretation (supplemented by internal prompt enhancement that the user never sees) between literally executing on what the user wants, and on using some measure of proxy judgment drawn from training data to avoid taking things too literally and betraying real-world constants. And coherence, from image to image, is better… but it’s not perfect.

For the user, what was once a completely blind exercise in prompting and re-prompting, with no consistency and little predictability, is now, ironically, a more complicated task. Because generative models do such a comparatively good job of outputting believable results that deliver what the user expected, fashion professionals need a much more discerning eye to see where coherence is breaking down.

I have seen, as I’m sure many readers have, countless instances of people promoting their AI workflows on LinkedIn and other platforms, and extolling the results they’ve been able to achieve in making sure a garment or accessory looks the same from shot to shot. I remember an especially proud one where an AI artist was convinced they’d completely solved coherence in a series of campaign-style images, but they’d failed to notice that the glasses the synthetic model was wearing had different arms in every angle.

As another example: the running jacket in our original 2022 image had a hood. The version the model recreated for 2026 doesn’t. And although construction details, trim placements etc. look consistent across generations, there are still subtle variations.

None of that (except for the hood, which is pretty egregious but could have been solved with a more comprehensive prompt) is worth burning it all down over, but that’s because we’re doing low-stakes work here. I’m illustrating an opinion piece, in a report that’s going in front of people who work with AI already.

The challenge would be if I was doing much higher-stakes work, such as creating campaign or PDP images of real products that I was aiming to sell to people – products that had to be made through codified sewing operations, and that consumers were going to wear and then inevitably compare to the marketing material.

These quirks of image generation are becoming much harder to spot, and much more fine-grained, but in a way this represents a much harder problem to solve if the goal is to turn generated images into real products.

Later [in The AI Report 2026] you’ll hear from a panel of different technology vendors, but there are two quotes I want to pull forward to here, because I think they capture the challenges of the post-image era quite neatly.

The first is from Jaden Oh, Founder of CLO Virtual Fashion:

“General-purpose image models lack precise control, so designers end up in this loop of tweaking the text prompt over and over, hoping the AI tool eventually guesses right. That unpredictability is a barrier to trust. Our approach is grounded in the customer’s actual libraries: their patterns, their fabrics, their trims. Anchoring it in real, manufacturable data that they own raises the visual quality, but the more important part is that what you see on screen is something that can actually be produced.”

The second is from Kirti Poonia, CEO and Co-Founder of Caimera:

“The question today is not whether AI agents can help fashion companies create sketches, renders, tech packs or campaigns. The question is how intuitive the tools are for seamless adoption and improved workflow. Taking away redundant tasks and sparing more time for human creativity to create bestsellers.”

This articulates what I see as the second big concern: adoption and intelligent, effective usage. As I’ve already spent many paragraphs saying, it’s now incredibly easy to spit out images that look like fashion photography, or interim stages of the visual pipeline, but the upshot of this scale and speed needs to be greater intentionality, not just a blithe acceptance that we can now do something more efficiently and on an effectively infinite canvas.

This is a challenge for brands, retailers, and tool-creators to address, by pairing that intentionality with the frameworks and the tools to actually execute on it, and by having the judgment and the sensitivity not to just go after the easy target of ubiquity.

The reason I’m referring to this moment as the beginning of the post-image era, then, is that the very definition and promise of an image has changed. It used to be the case that a picture was a scarce resource: expensive to produce, important to the recipient, and, as a consequence, a high-effort artifact that was implicitly positioned as a carrier of truth.

Now the opposite is true: an image is a commodity that’s everywhere, that requires literally no effort to create, is effectively free, and risks becoming throwaway and untrustworthy to the intended audience.

If we, as a society and as an industry, are not careful, the upshot is that the bar for images as a useful medium will lower until it hits the floor. And the trust we want to transfer with them will vanish unless it’s cultivated.

I don’t know what the image generations of 2030 will look like. I’ve learnt the hard way not to try and predict these things. But I do know that we’ll need to spend considerably more time thinking about what they mean than we will creating them.