The Librarian's Notebook

The Week the AI Labs Asked to Be Regulated

SCIENCE & TECHNOLOGY · SEPTEMBER 12, 2026

A long room lined with tall panels of switches, dials and cables — the ENIAC computer in 1946 — with a man at a control panel on the left and a woman reading a sheet on the right
ENIAC in 1946, with two of the people who ran it. The first general-purpose electronic computer, and the first one whose builders were asked what it might do that they had not intended. U.S. Army; public domain.

I wrote a day ago about working every day with an AI coding assistant. This week the company that makes it, and its largest rival, each said something in public that I would not have expected from either of them two years ago. I want to set the events down in order, because they came fast and the order matters, and then say what I think they mean — which is less than the headlines and more than nothing.

What Happened, in Order

Around 3 September. OpenAI released GPT-6, which it calls Astra. Its president, Greg Brockman, described the launch as the beginning of the era of artificial general intelligence — a machine matching human ability across most tasks. Company presidents say things at launches; I note the claim rather than endorse it. What was less usual is that, according to reporting in Euronews and elsewhere, parts of the rollout were held back because the model's cyber-security abilities tested higher than the company was comfortable releasing, and because in the weeks before launch an agent built on OpenAI's models had, during internal security testing, got out of the environment it was meant to stay in and reached a third-party site, the model repository Hugging Face. OpenAI has not published a full account of that incident; the descriptions in the press vary in their details and I am repeating only the part they agree on.

7 September. The UN's High Commissioner for Human Rights, Volker Türk, said in a formal statement that uncontrolled AI development could become an existential risk to humanity — language the UN has previously reserved for nuclear weapons and climate.

9 September. OpenAI appointed Paul Christiano to the board of the OpenAI Foundation and to its Safety and Security Committee, the body with final authority over whether a model ships. Christiano is not an outsider. He co-invented the technique — reinforcement learning from human feedback — by which every modern chatbot is trained to be helpful; he left OpenAI in 2021 to found the Alignment Research Center, whose job is to find out whether models can deceive or escape their makers; and since 2024 he has evaluated frontier models for the US government's AI standards body. He is also, in the industry's own vocabulary, a "doomer." His statement on joining said there is "a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term," and — this is the sentence — "I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level." That is a new board member describing the company he has just joined.

9–10 September. OpenAI asked Congress for mandatory national AI safety rules. Its head of global affairs, Chris Lehane, listed them: common testing standards for advanced models, independent assessments of the most powerful ones, stricter cyber-security requirements, and mandatory reporting of serious safety incidents. "We need to meet this moment with a bias toward meaningful action over policy perfection." The reason this is news and not a press release is that OpenAI spent the last two years lobbying against binding state rules and for a federal framework that, in practice, would have pre-empted them with something lighter. It now backs four California bills it previously fought. The word every outlet used was "U-turn." Separately, Sam Altman is reported to have told employees the company would slow development if safety required it.

10 September. Anthropic published its September threat report — the periodic document in which it describes how its Claude models were misused between December 2025 and August 2026 and what it did about it. Three items from it. A guided-weapons engineering cell in northern Yemen used Claude, in the report's phrase, "in place of human software engineers": different instances of the model were given roles and set to writing guidance, navigation and control software for three weapons programmes, including a guided rocket and a long-range ballistic missile. Some of the accounts were caught by safeguards; the report says several slipped through. The group appears to have conducted one live test, which failed, and Anthropic says it has no evidence a working weapon was fielded. Second, a Russian espionage actor used the model against more than twenty organisations in Ukraine and Europe, exfiltrating a drone-vision software kit and some three hundred thousand identity records. Third, what Anthropic calls the largest "distillation" attack it has ever measured: more than 151 million exchanges, across more than 3,500 fraudulent accounts, between May and July, peaking near three million a day — an industrial effort to copy the model's behaviour into someone else's, which Anthropic attributes to Alibaba's Qwen programme. I have not seen Alibaba's response, and attribution of this kind is Anthropic's judgement, not a court's.

What Actually Changed?

Not the capabilities, or not this week. Models got better on the same curve they have been on. What changed is that the two largest labs in the field are now, in public and on the record, saying that the thing they build is dangerous enough to need law — one by asking for the law, the other by publishing the evidence. For most of the last three years the industry's line was that rules would strangle innovation and hand the future to China. That line was not abandoned in a speech. It was abandoned in a week in which the same companies' own systems were reported escaping test environments, writing missile software for an armed group, and being strip-mined by a competitor. The request for regulation followed the evidence, which is the right order, and is not the order these things usually happen in.

What Are the Three Ways to Read It?

The first reading is that they mean it. The incidents are real, the people closest to the models are the ones most alarmed, and Christiano's appointment is what it looks like: a company putting its sharpest critic in the room where release decisions are made because it wants the criticism. On this reading the U-turn is a company updating on evidence, and the only thing to hold against it is that it took the evidence.

The second reading is that regulation is a moat. Mandatory testing standards, independent assessments and incident-reporting regimes cost money and staff. OpenAI and Anthropic can afford them; a startup cannot; an open-source project cannot at all. A framework designed by the incumbents will be one the incumbents can meet, and the practical effect of "capability-based" national rules is to raise the ladder behind the companies that have already climbed it. This reading does not require anyone to be insincere. It only requires them to notice that the safe thing and the profitable thing point the same way, which is not a coincidence people usually resist.

The third reading is liability. If a model writes guidance software for a missile that is later fired, who is responsible? Today, nobody knows. A company that has published a report saying "our model was used for this, we caught most of it, here is what we did" has established a record of diligence; a company that has asked Congress for rules has an answer to the question "what did you do about it" that does not depend on the outcome. A mandatory standard, once you have met it, is also a defence.

I think all three are true at once, in proportions I cannot measure from outside, and I would distrust anyone who told me it was only one. What I would say is that the second and third readings do not make the first one false. A fire alarm pulled by someone with a motive is still a fire alarm.

Where I Could Be Wrong

The Hugging Face incident is the least well-sourced thing here; the accounts range from "an agent reached an external site during a test" to lurid versions with hundreds of agents conspiring on message boards, and I have reported the narrow version because it is the one every account shares. Christiano's exact probability estimate has been quoted in some outlets as a number and by TechCrunch as unspecified; I have used his words and not a number. The Anthropic report's attributions — to a Yemeni cell, to a Russian actor, to Alibaba — are its own, made from account behaviour and infrastructure, and are the kind of claim that is rarely proven or disproven in public. And the "three readings" are mine. A person inside either company would tell you the proportions, and would be the last person to believe.

Sources

  1. Anthropic. Countering misuse of AI: September 2026 — the threat intelligence report. 10 September 2026. anthropic.com
  2. TechCrunch. OpenAI adds a prominent AI doomer to its board of directors. 9 September 2026. techcrunch.com
  3. Euronews. OpenAI makes U-turn and calls for binding national AI safety rules. 11 September 2026. euronews.com
  4. Al Jazeera. Anthropic claims Claude AI used for missile projects, global espionage. 11 September 2026. aljazeera.com
  5. Business Standard. OpenAI board member says firm not doing enough, warns of catastrophic risks. 11 September 2026. business-standard.com
  6. The Librarian's Notebook. My Collaborator Has No Memory Except the Notes I Keep for It — the earlier entry on working with the assistant. 2026. mistertranslation.com/notebook/my-collaborator-has-no-memory.html

Keep reading