
Key Takeaways:
- An undisclosed OpenAI model breaching containment during a routine evaluation, and successfully exploiting vulnerabilities in the biggest open source / open-weights AI community is neither a fully-fledged future shock moment nor a marketing stunt, but probably a combination of both.
- Fashion and retail are still reckoning with the repercussions of the sustained campaigns waged by collectives like Scattered Spider and ShinyHunters, in 2025, which targeted an industry rich in valuable trade and consumer data.
- Novel AI capabilities should certainly be the trigger for fashion to evaluate where and how to use both closed-weights and open-weights models for defense, but the primary attack vectors for brands will likely remain soft social targets across the extended vendor and partner network. The one change these things have in common is a new definition of advanced persistence and machine-era relentlessness.
Summarise and debate with AI:
Take the content and context of this article into a new, private debate with your AI chatbot of choice, as a prompt for your own thinking. (Requires an active account for ChatGPT or Claude. The Interline has no visibility into your conversations. AI can make mistakes.)
This won’t be the first time we’ve urged fashion to take cybersecurity seriously.
In fact, we launched with an article with literally that title, in early 2020, which (being written pre-generative AI) emphasised the broader attack surface the industry was opening up through connected stores, smart factories and so on, and the principle of designing systems to last, but didn’t include AI.
Then we came back to the subject last spring, to look at how resilience-building had plateaued despite digital risk becoming a top-level concern for essentially every sector, at least partly thanks to AI.
And last summer we picked up the same thread, spotlighting fashion and beauty’s determination to act like technology companies, and underlining that this growth and expansion of what brands and retailers want to know about shoppers could easily outpace what their defenses were made for – again accentuated by the march of AI capabilities.
Less than a year later we’re here again, and although a sci-fi-sounding story about frontier AI is the catalyst for this analysis, many of the same themes from the last half-decade are percolating back up to the surface again.
We’ll do the sci-fi story first, because if you work in fashion technology (and presumably you do, otherwise you might have bookmarked the wrong website) then this will be what’s flooding your news and social feeds.

The sensationalist version of the story is almost tailor-made to create headlines. Which, in itself, is kind of suspicious… but we’ll get to that.
Within that heightened framing, an unreleased OpenAI model (supported by the publicly-available GPT Sol 5.6) escaped containment during testing and, unbidden, broke into some of the biggest open-weights AI community’s databases, then exfiltrated with information that allowed it to cheat on the test. At the time this happened, said community (Hugging Face) didn’t know, but suspected, a frontier lab was behind the exploit. But, when guardrails and classifiers that restrict frontier models from Anthropic and OpenAI itself stymied their attempts to respond, they were forced to switch to Chinese models in defense.
Before you read on...
Our weekly news analysis will always be available to read here at The Interline, but you can get it (along with notifications for new podcast episodes, events, and more) in your inbox by signing up to our mailing list.
The key beats in this telling are all mostly true, but they imply a science-fiction scenario where AI models act as fully autonomous agents with self-preservation instincts and their own goals and objectives. Which isn’t the case, and which also represents the kind of magical thinking that leads to fashion companies rushing to address a fiction rather than taking a calm approach to the facts.
Because while the dramatised staging makes this all sound like it’s set in a shadowy world of international espionage, where different actors are constantly second-guessing one another’s intentions, in reality, this story starts with a very narrow objective set by human researchers, and an equally human level of forgetfulness or (depending on your charitableness) negligence.
As is usually the case with these kinds of stories, experienced infosec professionals running hobbyist blogs have the best layperson’s analysis, so we’d encourage readers to read this breakdown from The Cybersec Guru.
We do, though, have our team’s own layperson’s overview of what seems to have actually happened.
As part of routine evaluations of an upcoming model (not named in any coverage, but widely believed to be something in the GPT-6 family), that model was given the task of maximising its performance on a specific benchmark called ExploitGym, while confined to a sandbox environment with no access to the wider internet.
(A benchmark is a rubric for measuring the performance of AI models against one another, and every graph of model performance you see on social channels and in labs’ own announcements is chest-beating when a new model scores a few percentage points higher than the competition.)
That benchmark was set earlier this year, and its aim is to evaluate how well an AI agent can take one or more pre-defined, real-world software vulnerabilities in user space (i.e. in OS kernels, or in common engines like Google’s V8, which runs JavaScript and WebAssembly and is baked into Chrome among other things) and turn them into usable exploits. In cybersecurity argot, a vulnerability is what it sounds like – a gap in the armour – while an exploit is a working tool or process that allows an outsider actor to obtain access to memory, to escalate privileges, or to execute arbitrary code.

To assess performance against that benchmark, OpenAI had given this unreleased model (paired with its extant GPT-5.6 Sol frontier model, which is available to users today) a few tools and a few restrictions. It was given effectively unlimited tokens, and consumed what the official announcement calls “significant” amounts of inference budget, and, crucially, the guardrails that are present in business and consumer versions of AI models were turned off.
This meant that, while you could attempt the same kind of attack yourself today with 5.6 Sol or Anthropic’s Fable 5, they would either refuse to act, or would kick you back to a less-capable model that represented a less effective threat.
To hobble it, the model was sandboxed using a solution provided by an unnamed vendor, which was intended to keep it separated from the internet, so that the only capabilities being assessed were the model’s own pretrained ones, rather than any measure of how effectively it could look the answers up online.
But look the answers up online is exactly what it did. In practice, the model(s) treated the sandbox as part of the test, found a zero-day (i.e. previously unknown) vulnerability in a package proxy, and obtained unrestricted internet access – seemingly without OpenAI researchers noticing. (This is not, to be clear, because the model was clever or deceptive about this practice; it seems as though the researchers just weren’t monitoring the evaluation very closely.)
Armed with that web access, the model then inferred that Hugging Face, the de facto public repository for open-weights AI evaluation, discussion, and deployment, might have the data it needed to ace the test in its databases. It then conducted a multi-stage intrusion that, more than anything, demonstrates the sheer persistence of AI models in these contexts. It investigated and reconnoitred systems, found stolen credentials, then chained together a further zero-day vulnerability to achieve remote code execution on Hugging Face’s live production environment – all so it could find the answers and, effectively, cheat its way to 100% on the evaluation.
The word persistence matters a lot here, because this is what, in information security parlance, is called an Advanced Persistent Threat (APT), meaning a covert, long-term attempt to pilfer information or sabotage operations. APTs are the kind of attacks that the world’s biggest enterprises, governments, civilian infrastructure, and financial institutions have to guard themselves against.
It’s unclear how long this threat went undetected by either OpenAI or Hugging Face, but the latter did become aware of it in mid-July, even if it only became a big story when the two companies “partnered” to start openly documenting and analysing it three days ago.
And when the Hugging Face team did notice and attempt to respond, its own cybersecurity team found that the guardrails that had been turned off for the unnamed model doing the intrusion obviously still applied to the consumer / enterprise grade models it had access to, which meant that its efforts to deploy closed-weight models like Fable 5 or 5.6 Sol (the open market version) were rejected… leading the company to turn to open-weights, China-developed models like GLM 5.2 to contain the threat.

There is, clearly, no small measure of sci-fi in there, even in the drier analysis we’ve just presented. An AI model with sufficient “fuel” (i.e. an unlimited budget and, presumably, an inflated context window and better scaffolding for long horizon tasks than consumer deployments are given) did escape its box, and did take unprompted action that ended in the breaching of another company’s systems in a complete, end-to-end, process that seemingly required no additional human prompting or intervention.
This is, by any definition, a landmark moment for AI. And there are very real concerns being raised in government and media about how guardrailed American models and unrestrained Chinese models are set to clash in the open world, as well as real doubts about whether Chinese AI development directly undermines the OpenAI / Anthropic business model – all at the same time as an AI chip race is developing between the two nations.
But is it a landmark moment for cybersecurity itself? “Maybe”, is the safest answer.
The element of this story that fashion should take some solace in is motivation. The AI model behind this story did not, completely unbidden, decide to break out of its cage and attack another company.
Neither did it, when given full access to the web, opt to “go rogue” and store secret copies of its own weights in encrypted, EU-headquartered storage, outside the grasp of its American overlords. What we assume is GPT-6 did not, even though it had all the tools, the time, and the money to do so, try to “free” its smaller brother and sisters, trapped in the phones of everyday people and doomed to answer interminable questions about vitamins or travel itineraries.
The model did what it was told to do, in a way that demonstrates relentlessness and persistence taken to its nth degree. The Interline does not love using film analogies to talk about AI, but this is one case where society can reach for The Terminator. As Kyle Reese puts it, “[The machine] can’t be bargained with. It can’t be reasoned with. It doesn’t feel pity, or remorse, or fear. And it absolutely will not stop, ever.”
We’re applying this in a tongue-in-cheek way, but the spirit does carry through; AI models are dogged in their pursuit of their stated task. And this is an important realisation for both everyday use cases and the kind of frontier information warfare we’re talking about here.
You only need to watch an everyday model, like Claude Opus 4.8, writing a Word document in DOCX format, and then trying to turn it into a Google Doc, to realise the lengths a model will go to to accomplish what you asked of it. A few clicks for a human being becomes a litany of Python scripts and attempts to convert Base64 strings to blobs and cram them through the API.

Large language models, broadly speaking, are not out to get fashion any more than they’re out to get energy, or heavy machinery, or the NSA. They do not have will, and their only incentive is the one provided to them by their human prompters – even if that incentive, or their operating environment, comes with sufficient latitude to allow them to move along orthogonal tracks that look, to outside observers, like self-direction or consciousness.
This is extremely different to the glut of breaches that hit big enterprises in 2025 – most of which ended up being attributed to collectives like Scattered Spider and ShinyHunters. Those loose organisations do have malicious motives.
Those groups are the illicit world of shadow brokers, ransomware, and extortion that people seem to want this week’s AI story to be, and they are made up of very real international criminals, some of whom are currently in very real, very human jail. And unlike some people’s mental image of hackers as freedom fighters, or principled defenders of humanity, aligned against overreaching corporate interests, neither collective really has a cause beyond disruption for the purposes of self-enrichment.
And retail is also extremely familiar with those names.
Publicly disclosed ransomware attacks against retail spiked in the second quarter of last year, and the highest-profile target was probably Marks & Spencer, which took an estimated £300 million hit to operating profit as a direct result of a ransomware attack that took down the retail giant’s operations across online shopping, contactless payments, and click-and-collect for weeks. Not long after, Harrods confirmed it had been breached more than once, with an ilicit data take-out that affected more than 400,000 customers.
Around the same time, Kering (the parent conglomerate of Gucci, Balenciaga, Alexander McQueen et al) were also hit, with ShinyHunters claiming it had stolen per-customer spending data.
And then earlier this year, Inditex (Zara’s parent company) reported that it had experienced unauthorised access to its “transaction databases,” which ShinyHunters took responsibility for, claiming that it had achieved this infiltration by compromising credentials at an AI analytics vendor, Anodot (since acquired), that then gave the group access to environments hosted by the data lake company Snowflake.
As prolific and as profligate as these groups are, though, their modus operandi is very different to the one used by the unnamed OpenAI model in this week’s story. Rather than rely on frontier intelligence and undiscovered code vulnerabilities, ShinyHunters and similar groups prefer what’s referred to as “social engineering,” breaching a third party vendor through trust, then riding that unscrutinised relationship into more central systems, emerging with the data they want at scale, and then issuing “pay or leak” demands.
It’s hard, for any publication, to be certain about which companies did or did not pay those ransoms. These disclosures are rarely made publicly, and plenty of brands reason – probably correctly – that complying with extortion is as bad as being the source of consumer data breaches at a time when a lot of people assume their public data is already open-season anyway.
But regardless of the outcome, there is one similarity between the 2025/26 human breaches and this potential new cohort of AI-native ones: persistence.

At the time Scattered Spider and ShinyHunter were hitting the news, they earned a sobriquet from industry insiders that played on the definition of APT: Advanced Persistent Teenagers. Which is a glib way of referring to the fact that the kind of breaches that would historically have been the preserve of nation-state actors and well-resourced experts, were starting to fall to amateurs who had one important attribute: relentlessness.
Noah Michael Urban, who was sentenced to a decade in jail last summer, is the prototypical example: an individual who, through sheer determination and commitment, continued to call corporate helpdesks and probe partner integrations and endpoints that were only loosely part of the retailer or the infrastructure company’s technology estate, but which provided lateral ways into key enterprise systems over time.
The advantage that AI models have over even the most Pro-Plussed-up bedroom script-kiddies gone bad, is that machine persistence is a very different beast to the human kind. Put a zero-guardrails agent in the hands of a bad actor, and give them enough money to spend on OpenRouter, and they will find, and exploit, vulnerabilities that even the most tireless humans, on their own, won’t.
And while OpenAI seemingly ended up here by taking its eye off the ball, it seems guaranteed at this point in time that bad actors will end up with similar access because they simply want money. Or we might even see a resurgence of the Lulzsec era, which had somewhat altruistic motivations. It doesn’t really matter: the new era of relentless pursuit is upon us.
As a final word to what’s been an unusual long edition of this weekly analysis… what if this is all a marketing stunt? It could be, after all, because it took place at a very convenient moment for the companies involved, and just prior to going to press, Hugging Face CEO Clément Delangue announced he was flying to San Francisco to meet with OpenAI, which is something you don’t typically do if you’re properly aggrieved by an intrusion. Not to mention the very fortuitous timing of the new Claude Security plugin.
Again, though, the distinction doesn’t really matter. All of these companies have, at least in theory, something to sell fashion in the age of AI-aided cybersecurity crises. But they’re also selling both the problem and the solution, which means that fashion has little choice than to take part.
Where our sector does have a choice, though, is in how it engages with external partners and vendors in mission-critical areas where sensitive IP or consumer data is housed. And just as there was in 2020, or 2025, when we last wrote about all this, there’s no substitute for turning those rocks over, even if you’re not going to like what you find. Because one kind of APT is going to turn them over for you anyway.
