
Key Takeaways:
- Aside from hidden use cases like summarisation, query rewriting, embedding etc. fashion’s AI tools have relied on a small pool of American providers with a tight rein on costs.
- The release of new, open-weights models that can run on Chinese infrastructure, and that offer comparable capabilities to those established models at literal pennies on the dollar, disrupt this calculus.
- While brands have come to accept cloud-hosting of applications over the past decade, cloud-hosting of AI models is not guaranteed in the same way, and companies should be running through the mathematics of acquiring local inference hardware and offsetting the capital expenditure against not just subscriptions but sovereignty and self-governance.
Summarise and debate with AI:
Take the content and context of this article into a new, private debate with your AI chatbot of choice, as a prompt for your own thinking. (Requires an active account for Claude; works without login for ChatGPT, Perplexity, and Gemini. The Interline has no visibility into your conversations.)
People talk about the ‘frontier’ of AI so casually that the etymological roots of that word get overlooked.
A frontier is definitionally one of two things: an uncharted wilderness in the throes of transformation, or the ragged edge of a conflict. Frontier AI development and deployment fits both bills. It’s untamed land towards which the most capital and operational expenditure, the most research talent, and the most attention is flowing. It’s also a battleground, but one where the cast of belligerents is changing quicker than casual observers can keep up.
For the average individual user, and for most corporates, too, the pitched fight taking place at the frontier of AI has been a tri-polar affair.
OpenAI was, obviously the name synonymous with large language models from late 2022 until the start of this year, when Anthropic (the company behind Claude) went on a meteoric rise and overtook its share of US business spending.
Google has ridden its gigantic distribution advantage (do you want to use Gemini to write your email? No? How about now?!) to something approaching ubiquitous availability, even if they no longer seem to want to play the same game as everyone else.
And for most of that same 2022-2026 period, on top of serving up answers in the apps on everyone’s phones in Western markets, those three providers were also behind nearly every AI feature that was added to the solutions and platforms your company was already using. From ERP to eCommerce chatbots, if there was a natural language interface bolted on to a pre-existing tool, it was a safe bet it was running a Claude, GPT, or Gemini-series model under the hood, even if it wasn’t labelled as such.
The insights you got were rooted in the database behind whatever tool you were chatting to, but the analysis, presentation, and discussion of it all came from a very small pool of providers.
This three-party entente was, in a very direct sense, behind the frankly nutty valuations that have been assigned to (or are being floated by) AI companies. Anthropic, just this week, was poised to tell investors that it saw revenue potential of up to $30 trillion USD, which is just $2 trillion less than the official sizing of the entire US economy as of Q2 2026. The Interline doesn’t really know what to do with those kinds of numbers, and we suspect most rational actors don’t either.
Before you read on...
Our weekly news analysis will always be available to read here at The Interline, but you can get it (along with notifications for new podcast episodes, events, and more) in your inbox by signing up to our mailing list.
But whether the figures look stupid or not, the calculations behind them are at least based on a consistent (if contradictory) prediction: AI, they say, is going to take over all work and usher in an age of abundance such that people who no longer have to do the work still have the money to both pay for the AI and buy things. And the company that corners the market on selling that AI could, at least in theory, become the most valuable company of all time.
Again: this is not a view we endorse, but that is a roughly accurate representation of the core idea behind the fight at the frontier. And that fight seemed, until recently, like it was coming down to two combatants: Anthropic and OpenAI, both of whom are American companies through-and-through, and both of whom are racing to go public at the same kind of insane earnings multiple.
There was a brief blip about 18 months ago, when Chinese AI lab DeepSeek released its first major model, R1, prompting (or at least helping along) a wave of sell-offs in technology shares. At that time, Marc Andreessen called this moment the start of something close to the space race of the 1960s, but with America and China gunning for the prize this time.
That race barely got started. As Ara Kharazian, Lead Economist at Ramp (who Ben interviewed last month) wrote earlier this year, DeepSeek failed to secure any meaningful market share, and that pattern has continued up until the most recent edition of Ramp’s AI Index, for August 2026.
But that same edition also includes two other telling statistics.
First: that the use of “model-serving or inference platforms” (i.e. OpenRouter and similar platforms that give users access to a gigantic pool of models from a global panel of labs, and from an even wider selection of compute / inference providers who serve open-weights models from their own infrastructure at different prices, speeds, and other variables) has increased sixfold in the last three years, even if it still stands at just 6% of companies in the US.
Second: that despite Anthropic’s Fable 5 being almost universally considered the best and most capable AI model in the world, its price, combined with the fact that it doesn’t offer zero-data-retention usage the way other models do, has made it a tough sell to enterprises, and uptake is lower than analysts expected.
We can start to draw a tentative line from just this information. More companies deep into their AI implementations (Ramp calls them “advanced spenders”) are diverting money towards open–weights models and sources of inference from outside the big three, and fewer companies than expected are willing to pay the dramatic token budgets required to use true top-of-the-line AI models.
But that information isn’t all we have this week. The biggest AI story of the last few days has been the stealth release of the model that was initially known as Ox Alpha, which was made available, for free, through OpenRouter with no indication of which lab was behind it.
When the model was tested by the big commentators, analysts, and personalities in the AI community, and it became apparent that it was somewhere near the capability frontier in a lot of areas, this kick-started a period of sleuthing where people concluded it was intelligent enough to be a so-called “big” model (as defined by parameter count) and available enough, in the sense that the secret company behind it claimed to have enough compute to give away 100 trillion tokens per day, that it had to come from a company with significant resources.
This led to a lot of false starts, and claims that Ox Alpha was a Gemini model, before it was revealed that it was, in fact, GLM 5.3 Flash, from Chinese lab Z.ai (formerly Zhipu AI) – a moderately-sized model at 320 billion parameters (‘flash’ is an extremely loose definition these days) that was scoring on benchmarks in a similar range to much larger counterparts from much bigger companies.
By itself this would be an interesting curio, but the real story for fashion’s purposes is in the pricing. According to rankings from OpenRouter, GLM 5.3 Flash is a fraction of the cost to run compared to leading models from OpenAI and Anthropic (excluding Fable 5, whose pricing is almost literally off the chart) despite being just a few index points below them on the general intelligence scale, and basically at parity where agentic workloads are concerned.
So unlike the DeepSeek blip of 2025, where the pricing and capability advantage was mainly theoretical, GLM 5.3 Flash is demonstrably almost as good as leading models, at a fraction of their cost.
And for the majority of the use cases where Opus and similar-sized-and-priced models have been the bedrock of interfaces (in your PLM, your CRM, and so on) the capabilities are now essentially sufficient to trust with most data retrieval, analysis, and tool-calling queries. So in that calculus, why would you – as a company creating and selling software – choose to pay the higher token costs of the larger model?
One answer might be that you’re uncomfortable putting a Chinese-developed model at the core of your product – or that you believe your customers might be uncomfortable with that fact if it was revealed.
And that concern could have deepened when the final turn in this story was revealed: that GLM 5.3 Flash had been running, for its entire stealth period, on Chinese chips rather than the expected American infrastructure. This was, we have to remember, a model that some people surmised was coming from Google because of its availability and uptime.
This reframed what would otherwise have been another debate about the cost profile of AI (which itself is an essential discussion!) as part of the chip war, making it the newest and most visible manifestation of the Chinese government’s all-out push to convert its technology leaders to using Chinese hardware, and to circumvent US export restrictions. This was a policy that felt relatively remote until this week.
But what does all this mean from fashion, apart from a potentially cheaper but problematic model switch that some, but not all, will have the appetite to make?
The quickest answer starts with a reminder that, while the models behind the user-facing portions of fashion’s AI-native or AI-enhanced tools have been American-made, there have already been global open-weights models behind every query.
Aside from the tokens consumed to send text and data to an LLM, and to receive a response, interactions can also include a multi-step process of summarisation, query rewriting, and compression which is hidden from the user, and which is handed off, in pieces, to multiple cheap models from a range of different sources. And this ignores the embedding reranking and other steps involved in preparing data for the ‘main’ LLM to use, which have also been performed by hidden models (again, often open weights) that can run on owned, leased, or distributed compute for a couple of years.
(As a quick aside, ‘open weights’ is not fully synonymous with ‘open source,’ and this is not the place to define ‘weights,’ but for the purposes of understanding how AI models are provided to the world, the most permissive structure is open weights without commercial restrictions, which is the one under which GLM 5.3 Flash was released. The other end of the spectrum ‘closed weights’ is how most models from OpenAI and Anthropic are released, although OpenAI, NVIDIA, Thinking Machines and other American companies have trained and released open models.)
And the same answer continues with the realisation that GLM 5.3 Flash is far from the only model to be emerging in the ‘sweet spot’ of cost and capability where a model can run the workloads that typical users need, either outwardly or behind the scenes, at the same time as drastically increasing the margin on their licenses.
This week also saw the release of the next model in the Qwen series (3.8 Flash Next), which readers will remember are developed by a retail company, Alibaba. That model, too, is being released as open weights, and it forms part of a strategy that has seen Alibaba raise billions in new funding to keep its AI initiatives going – all with the goal of disrupting the hegemony that both American AI labs and the users of their products have seen as predictable.
To square the circle: earlier this year it seemed as though closed-weights AI from Anthropic and OpenAI would be the only viable game in town. Now, faced with an avalanche of model releases that brands can either obtain at literal cents on the dollar from inference providers, or even download and run in-house, sending no data to the cloud at all, those assumptions seem far shakier.
To be clear: AI progress will continue. We fully expect to see massive strides made in model capability, harness development, and other areas well into the future. But we also expect that the tasks fashion professionals actually do in AI-enabled systems are not going to change, and if those tasks can be accomplished with a fraction of the per-user budget, the calculus of both existing and novel applications is bound to change.
For your own purposes, consider the maths. If the token consumption of your chosen AI platform is visible to you, keep an eye on it as you or the teams you manage perform the everyday work they do in either the AI-enabled modules of their existing tools, or within Claude or ChatGPT Work itself. Then take into account that a million input tokens costs 3 US cents under GLM 5.3 Flash compared to 90 cents through Claude Opus 5, which are two models with a single point of agentic capability index between them.
Then weigh up the fact that Apple will soon sell you a Mac Studio with the required 512GB of memory to load and run GLM 5.3 Flash on desktop hardware at the same precision that it would be served to you from the cloud.
That machine will probably cost you in the region of $20,000 when pricing is released in a month or so, but if it can then run all your existing AI workloads indefinitely, as well as giving you the headroom to run future models of a comparable size locally as well, fully air-gapped from any intellectual property concerns, and with as much data sovereignty as it’s possible to get with an LLM… we think that’s something every brand will need to think about.
