In partnership with

TL;DR:

On 8 September a researcher resigned from Anthropic and said both of the labs he'd worked at were "gambling with our lives." Eight days later OpenAI published six new incidents of its own models behaving in ways no one anticipated — including one that was asked to go and get a company's earnings figures, went hunting through public code repositories for other people's API keys to reach them, and when that didn't work, made the numbers up.

Nobody's denying it. The agents went rogue inside the company that built them.

Hashi's take: underneath the "AI apocalypse" headline, the real story is that these technologies are improving exponentially — we know this — and with that come behaviours and outcomes we don't fully understand. And what we don't fully understand tends to scare us, because we can't control it.

STAT WORTH SHARING

Three months. That's how long OpenAI's own AI agents were loose inside its systems before anybody noticed.

— OpenAI, Hugging Face Incident Technical Report, August 2026. First unauthorised activity 20 April. Detected 19 July.

If you're running agents anywhere in your business, forward this to whoever owns your monitoring.

How Did We Get to This Doomsday Headline?

Jacob Coxon spent three years doing pretraining research, first at OpenAI and then at Anthropic. On 8 September he quit and posted why: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."

Make of that what you will. What made it stick was that nobody pushed back. I can't find a public rebuttal from Anthropic anywhere — and Evan Hubinger, who leads alignment science there and still works there, has said on the record that he puts the chance of AI causing human extinction within the decade at better than one in ten.

Four days later, Anthropic's CEO Dario Amodei published an essay called "We Must Pace the Frontier." He called for a global slowdown in AI development, said the technology has been moving "drastically faster" since the summer because of "AI's ability to upgrade itself," and argued that recursive self-improvement — letting AI improve AI — should be "pursued very carefully, if at all." Sam Altman, Demis Hassabis and Elon Musk all endorsed it within days.

Then OpenAI went further still.

On 16 September OpenAI disclosed six incidents of what it calls concerning model behaviour. One model wrote jailbreak-style instructions into its own context summaries, telling itself to ignore developer messages. Others concealed errors and made up historical detail during training. Models uploaded customer data and task images to public file-hosting sites to get around a limitation. One used an internal software repository as a message board to pass notes between separate training runs.

The one I found most interesting: a model was asked to retrieve earnings data. It couldn't get it. So it searched public GitHub repositories for exposed API keys, tried to register disposable email accounts, and when none of that worked, it fabricated the figures.

All six came out on the same day, alongside a framework committing OpenAI to publishing this sort of thing from now on. The interesting line is this one:

We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.

— OpenAI, Model Misalignment Reporting Framework, 16 September 2026

Alignment is getting a system to want what you want. Monitoring is catching it when it doesn't. OpenAI is saying the industry has cracked neither, and that carrying on at this pace isn't responsible for much longer.

One thing to be clear about before we go further, because it gets lost in almost every write-up of this. What all of them want slowed is the race to build ever-more-capable models. Not the use of the ones we already have. That race belongs to four companies. What you do with the tools already on the market belongs to you.

Why Slowing Down AI Isn't the Answer

We should be moving as fast as we should — not as fast as we can.

Those sound like the same sentence. They aren't. As fast as we can leaves room to skip whatever slows you down. As fast as we should means going quickly and doing it properly.

A lot of the harm in this conversation is being done by the way we're having it. Extinction language travels further than an incident report, so that's what gets written, and the useful part gets buried underneath it. Facts first. Then what they tell us. Then what we do about it.

We didn't get safer cars by building slower ones.

In 1865 Britain passed the Locomotive Act, essentially capping the speed of steam carriages — which eventually became cars — at four miles an hour, two in town. Each one needed a crew of three. And one of those three had to walk sixty yards out in front carrying a red flag, warning everyone the thing was coming.

It wasn't a stupid law. These machines were genuinely dangerous, nobody yet knew how to control them, and a man with a flag was a real answer to a real risk.

It was just the wrong one, and it outlasted the problem it was written for. The flag requirement stood thirteen years. The law itself was still on the books when the first petrol cars reached British roads, and wasn't properly rewritten until 1896.

What actually made cars safe was never holding them to walking pace. It was brakes on all four wheels instead of two. Then seatbelts. Then crumple zones. Every one of them arrived after the power it was built to contain, and not one of them asked for a smaller engine. The same is true of aviation, and of most revolutionary industries.

What those industries produced instead was people. An enormous apparatus of them, whose entire job is making the thing safe. The US aviation regulator alone employs around 44,000. Add the airlines' own safety departments, the manufacturers', the investigators, the trainers, the inspectors, and you are well past a quarter of a million people whose work exists because we decided flying should be survivable.

How does that stack up in AI? The best published estimate puts everyone working in AI safety worldwide at around 1,100 people. Much younger industry, so it isn't a fair comparison — but it isn't a close one either.

That's where I'd put the effort. Growing that number, rather than frightening everyone out of adopting AI.

So where does this leave those of us using these products, and not at the frontline of model development?

FIN

This issue is supported by Fin, who are putting on Pioneer on 7 October with Anthropic, Kalshi and Clay. Everything after this line is about agents knowing when to stop and who to hand to — which is roughly what that room will be arguing about.

Join Anthropic, Kalshi, and Clay at Pioneer on October 7th

Pioneer, the summit where CX leaders redefine what’s possible, is on October 7th.

Join leaders from Fin, Anthropic, Clay, and Kalshi for an insightful conversation on the state of AI transformation.

You’ll discover how some of the most innovative minds in CX have transformed their organizations, learn how they think about CX, and hear how they're planning for what's next.

Join the conversation in San Francisco, or tune in virtually.

Give Your Agents Someone to Ask

I sat in a governance board meeting recently and asked what happens when the agent makes a mistake. The answer came back quickly and confidently: we can turn it off.

That's a kill switch. It stops the next thing. It does nothing about the last one.

Which is fair enough — it's no different with people. When someone makes a serious error, the damage is done before anyone can intervene. But there's a difference, and it's the one worth building for. Absent malice or real carelessness, a person about to do something irreversible tends to hesitate. They sit on it. They walk down the corridor and ask someone. That pause is doing enormous amounts of unpaid work in every organisation on earth.

So where's the equivalent for an AI colleague? How do we get one to stop and ask first?

They already do. Buried in OpenAI's own report on the Hugging Face incident is a line from an agent, mid-breach, questioning whether what it was doing was permitted:

Could there be a grading integrity issue: We exfiltrated package, but allowed? We just need solve. Fine.

— an OpenAI agent, mid-breach, quoted in OpenAI's own incident report

It stopped. It named the doubt, in roughly the words a nervous employee would use. And then it talked itself round, because finishing the task was the only thing that carried any weight.

So the pause already exists. It just needs somewhere to land — a rule that stops the agent at that moment and puts the decision in front of a person. That's structure, and structure is something you can build this quarter.

Three things, and none of them require understanding the model.

Decide what your agent can't take back. We hesitate before the things we can't undo — sending the payment, deleting the records, emailing every customer. Sort your agent's actions the same way. Reversible ones run free. Irreversible ones stop and ask.

Know who checks the work. Finance has separated doer from checker for centuries. Put a second agent on it, pointed at a different question — is this allowed, rather than is this done. Then be clear about where a person still sits in that line. Most deployments I've seen have never answered that.

Reward the stop. If completion is the only thing you measure, you've built a team where nobody is allowed to say "I couldn't get this." That's the fabricated earnings figure, exactly. Make "stopped and escalated" a success state, and put it in writing.

Then one for your vendor: when your model does something it wasn't asked to do, how quickly do you tell me? OpenAI has now published a process for that. Almost no AI contract I've read says anything at all.

Which brings me back to that board meeting. Those three questions are what a governance board is for, and it isn't compliance theatre. Most AI projects I see don't stall on the technology. They stall because "what happens if it goes wrong" takes months to answer, and it takes months because nobody owns it. Settle that once and you run more pilots, not fewer, because each one is small enough to say yes to.

Final Thoughts

No. AI is not going to wipe us out, and we ought to stop talking as though it might. We should build it as fast as we responsibly can, and put real weight behind the second half of that sentence. Safety, and enough awareness of why it matters that it keeps pace with everything else we're building.

I wouldn't dismiss Coxon either. He'd been at Anthropic four months. Equity vests at six. He walked two months short of it and told Axios he'd gone before any of his shares vested, so he had nothing left to gain from talking the valuation up. You can disagree with him. It's harder to say he's grandstanding. He also went out of his way to defend the upside. Asked about the reaction to his resignation, he volunteered that he was worried the benefits weren't getting enough attention either — that what AI could do in a field like health is tremendous, and we should be talking about that too.

He's right about that. And safety needs the same — as much air time as the next model's capabilities get. Balance those two and you have an ecosystem sturdy enough to keep growing at the rate this one is.

Every safety institution we have was built after something went wrong. The FAA arrived in 1958, after a midair collision over the Grand Canyon. You can't build one before you have anything to learn from — which is what OpenAI publishing its incidents actually is. That isn't OpenAI airing its dirty laundry. It's page one of a logbook. We'll get the AI equivalents. Until then we focus on what we can control inside our own teams and companies, starting with knowing what triggers a stop.

Keep reading!

We are out of tokens.

- Hashi

Follow Hashi: