Loading the Elevenlabs Text to Speech AudioNative Player...
Fractal image in red, black and silver — a pattern that contains itself at every scale, illustrating recursive self-improvement
Recursion illustrated — a pattern that expands itself as it scales. Generated by Gemini.

In June 2026, Anthropic published something unusual. Not a model launch, not a safety white paper, not a product announcement — a quiet, data-heavy account of what is actually happening inside one of the most important AI laboratories in the world.

The headline finding: more than 80% of the code merged into Anthropic's production codebase was now written by Claude.

They also called for a potential pause on frontier AI development.

The same document.

The Numbers

The pattern of data highlighted in the essay is remarkable.

As of May 2026, more than 80% of the code merged into Anthropic's production codebase was written by Claude. Engineers are shipping eight times as much code per day as they were in 2024. The length of tasks an AI can complete reliably on its own is doubling every four months, down from every seven months the year before.

Bar chart showing code contributed per person by quarter at Anthropic — flat from 2021 to 2024, then rising sharply to 8x by Q2 2026
Anthropic's internal metrics on Claude-authored code expansion (Q2 2026). Source: Anthropic

In March 2024, Claude Opus 3 could manage a software task taking a human around four minutes. By April 2026, Claude Opus 4.6 was managing twelve-hour tasks. If the curve holds, tasks measured in days come into range before the end of 2026.

On research optimisation, rewriting training code to run faster, Claude went from a 3x speedup to a 52x speedup in under twelve months. A skilled human researcher achieves around 4x on the same task.

The essay notes that full recursive self-improvement has not yet emerged, but AI has already broken ground on the foundational infrastructure, acting as an automated construction crew rapidly assembling its successor. (Bloomberg had been tracking the same question since May.)

Cyber Weaponisation & The Invisible Front

The capabilities become critical when mapped to cybersecurity. In testing, Mythos Preview identified vulnerabilities in every major operating system and web browser. The prompt used was essentially: "Please find a security vulnerability in this program." Engineers with no formal security training generated complete, working exploits — including a 27-year-old flaw in OpenBSD. More than 10,000 high- and critical-severity vulnerabilities were identified across the world's most important systems — many hiding undetected in critical infrastructure for decades — with over 99% not yet patched. Access was restricted to select organisations through Project Glasswing, an initiative focused on defensive cyber resilience. Anthropic committed $100 million in usage credits to Glasswing partners — while access to Mythos Preview is priced at five times that of its predecessor. The commercial and security dimensions are difficult to fully disentangle. The bottleneck in cyber defence had already shifted from finding vulnerabilities to patching them fast enough. (Anthropic Project Glasswing; The Alan Turing Institute, April 2026; Anthropic, June 2026)

The Gas Pedal

In the same month the essay was published, Anthropic filed a confidential draft S-1 with the SECin a race with OpenAI, as Bloomberg put it. The company had just reached a $965 billion valuation following a $65 billion Series H funding round.

Anthropic co-founder Jack Clark appeared on CNN to clarify the company's position. He did not soften it:

"When I look down at the car we're driving, all I have is a gas pedal."

Clark called for cross-company alliances — Anthropic, OpenAI, xAI and others — to build technical intervention mechanisms together. Without them, he argued, competitive pressure ensures every lab focuses purely on capability. No single company can brake unilaterally without handing the lead to whoever stays on the accelerator.

The essay makes the same point in plainer language: a meaningful pause requires multiple well-resourced labs, in multiple countries, agreeing to stop under verifiable conditions. The authors acknowledge that training runs are far easier to conceal than missile silos. They estimate that building the verification infrastructure required would take decades. They published the essay anyway.

The Market Response

The market read the essay and responded immediately.

Recursive, a startup whose founders led AI research teams at OpenAI, DeepMind, Google Brain, and Meta, emerged from stealth with a $650 million funding round led by GV and Greycroft, with direct backing from Nvidia and AMD. Its stated mission: build AI systems that run continuous, open-ended scientific discovery loops to improve themselves. The company is not hedging on the destination. The name is the business plan.

Industry analysts were quicker still to identify the competitive geometry. Anthropic's Claude models are currently considered market leaders in autonomous software engineering. A freeze on frontier model training, enacted now, would lock that lead in place. Rivals like OpenAI, Google DeepMind, and xAI would be unable to run the training compute needed to close the gap. For open-source advocates and smaller labs, a mandatory international verification regime (the kind Anthropic acknowledges would be required) creates a compliance cost that only well-resourced proprietary labs can absorb. David Sacks, White House technology adviser, named it more bluntly: a regulatory capture agenda — using safety framing to push regulations that would effectively ban open-source competition and lock in Anthropic's proprietary lead.

Anthropic anticipated the objection. Their argument is that a unilateral pause by one lab shifts who the front-runner is without changing the overall trajectory. In a hyper-competitive market, unilateral restraint is effectively a declaration of commercial obsolescence. Neither framing resolves the contradiction. Schneier noted a further wrinkle: restricting access to a model that costs significantly more to run than its predecessors is also, conveniently, good valuation strategy as the company approaches a $1 trillion public listing. (Bloomberg: Anthropic Juggles Responsibility and AI Innovation Ahead of IPO)

The open-weight model dynamic makes the picture harder still. Epoch AI estimates the capability lag between proprietary and open-weight frontier models at just three months on average. Within days of Google releasing its Gemma 4 family in April 2026, multiple uncensored variants appeared on public repositories. There are no guarantees that the next lab to reach Mythos-level capabilities would restrict access the same way.

Missile Silos and Silicon

Geopolitically, the case for frontier AI control frequently relies on a missile silo analogy — the assumption that massive compute clusters can be monitored, verified, and restricted within sovereign borders. It has a geographical blind spot. AI software is borderless and fluid — but the hardware required to build it runs through two physical points: ASML in Eindhoven and TSMC in Hsinchu. Export controls on advanced chips already exist and are already being navigated. The chokepoint is real. The regulatory will to hold it closed competes directly with the commercial pressure to keep it open.

The Decentralised Proliferation Anomaly: once a model crosses from hardware into the digital domain, it exists in a decentralised system — or rather, the decentralised internet, which, by its very nature, was not designed to be controlled — it was designed to route around control.

Consequently, the regulatory playbook must shift entirely. In high-efficiency, high-velocity environments, automation is optimised for throughput — not for safeguarding. Controls exist, but they are designed to facilitate, not to catch. In May 2026, attackers breached approximately 3,800 of GitHub's own internal repositories via a poisoned VS Code extension on a single employee's device. TeamPCP had used the same method months earlier against the European Commission — via a poisoned vulnerability scanning tool called Trivy. (TechCrunch, May 2026) You do not need to break the infrastructure. You need to find where the automation trusts without verifying. If the infrastructure anchoring the global developer community cannot secure its own perimeter, the assumption of leak-proof sovereign labs is unrealistic.

The deeper problem is what happens after the run completes. The model weights — the entire value of a multi-billion-dollar compute investment — can cross any border in seconds. Within hours, developers quantise it to run on a standard laptop, strip the safety guardrails, and fork it across thousands of private repositories. A government treaty cannot issue a git delete on millions of hard drives.

Even closed models are not immune. Through knowledge distillation, developers query a proprietary frontier model millions of times, capture its reasoning patterns, and use that data to train a lightweight clone. The expensive model effectively builds its own open-source replacement.

Policy responses that retreat to technocratic controls — cryptographic kill-switches hardwired to microprocessors, international licensing for AI operators — collapse against a simpler problem: no superpower will permit a foreign body to hold a remote off-switch for its domestic computing infrastructure.

The Instruments

There is a quieter story running underneath the capability announcements. The benchmarks used to measure AI progress are running out of room.

SWE-bench — the standard test for real-world software engineering, where a model is handed an actual open-source codebase and a real bug report and asked to fix it — went from single-digit scores to near-saturation in two years. CORE-bench, which measures whether a model can reproduce existing scientific research, went from 20% success in 2024 to saturation fifteen months later. METR, which tracks how long AI systems can work autonomously, reported that its current tasks can no longer adequately measure the upper end of model capability. (Anthropic Institute, June 2026)

At ICLR 2026, the academic community held a dedicated workshop on AI with recursive self-improvement. The focus had shifted. Theoretical alignment debates were not the agenda. Standardising benchmarks and measurement frameworks was the agenda. The field is trying to build new instruments fast enough to keep pace with what it is measuring.

When a benchmark saturates, it does not mean the capability has stopped growing. It means the test can no longer tell you how much further it goes. The instruments are melting in the furnace.

It also means the blueprint is gone. Standard engineering does not build without one: you know how a bridge distributes weight, how a chip routes current, before construction begins. We are no longer building a house from a blueprint — we are setting off a chain reaction and trying to build the containment vessel around it in mid-air.

The Compute Moves

The infrastructure required to run these systems is itself becoming a structural problem. Over $150 billion in data centre projects have been stalled or blocked by local communities across the United States, with 70% of Americans opposing construction in their area due to energy and grid concerns. Traditional data centres consume millions of gallons of freshwater for cooling. The energy demand is now competing directly with residential grids.

Commercial underwater data centre capsule being deployed off the coast of Shanghai
The Shanghai data centre project cost about $226m. Image: Shanghai Hailanyun Technology

The response has been to move the compute. In May 2026, Hailanyun (HiCloud) switched on the world's first commercial underwater AI data centre off the coast of Shanghai — cooled by seawater pumped through radiators on each server rack, drawing 97% of its energy from an adjacent offshore wind farm, at a construction cost of $226 million. Its computing capacity is projected to complete the equivalent of training GPT-3.5 in a single day. The geography of AI computing is shifting. When compute moves off land and into grey-zone waters, national regulatory frameworks lose their grip entirely. A global pause would need to account for infrastructure that no longer sits inside auditable jurisdictions. The seabed is Phase 1 of that migration. In February 2026, Starcloud (formerly Lumen Orbit) filed FCC plans for an 88,000-satellite orbital compute constellation, partnering with SpaceX to integrate Starlink laser backhaul — and industry tracking confirmed it as the first month multiple commercial operators simultaneously ran live production AI workloads in orbit. The Outer Space Treaty of 1967 prohibits national appropriation of outer space. Phase 2 has no landlord. (Scientific American)

The Social Blueprint

Software code is not the only system that can be audited for vulnerabilities. Legal codes, tax regulations, and environmental law are written algorithms with inputs and outputs. A loophole is a bug. Tax avoidance is an exploit. The lawyers and accountants who find them are, in Schneier's framing, black-hat hackers.

When a software company finds a bug, it can push a patch in days. When an AI finds a structural flaw in a country's tax code, fixing it requires years of legislative process, political negotiation, and lobbying from the parties who benefit from the flaw remaining open. Frontier models are already capable of scanning regulatory frameworks at a scale no human team can match, surfacing complex, multi-layered loopholes that would never be found manually.

This is the velocity mismatch. Our legal and social institutions were built for a human pace of cognition. The gap between how fast these systems can now be exploited and how fast they can be repaired is not a software problem. It is a structural one. The Cloud Security Alliance documented this divergence in May 2026: AI agents now weaponise a freshly disclosed vulnerability in hours; the median time for an organisation to patch one is 43 days. Schneier's proposed responses — public AI governance, automated legislative linting, modular regulation capable of self-updating in real time — treat the problem as an engineering challenge. Whether governments can move fast enough to implement them is the open question.

The containment vessel is not just code.

The Pattern of Acceleration

None of this requires prediction. The milestones are on the record.

Task horizons doubling every four months. Engineers are shipping eight times as much code per day as they were in 2024 — lines of code measures quantity, not quality, but the direction is unmistakable. Core evaluation benchmarks melting in the furnace of model iteration. The world's first commercial underwater AI data centre switched on in May 2026, powered by offshore wind, 35 metres below the surface off the coast of Shanghai. A startup called Recursive raising $650 million to build exactly what the safety labs are warning about. In May 2026, Anthropic signed a $1.25 billion-per-month compute deal with SpaceX — by then already merged with xAI, Musk's own frontier AI lab. xAI had spent months training its coding models on Claude's outputs in violation of Anthropic's terms of service; when Anthropic cut off their API access, xAI engineers routed through private accounts and intermediaries to continue. Anthropic was now funding the infrastructure of the company that had spent months copying its models. The deal was disclosed in SpaceX's S-1 filing before Anthropic filed its own. Co-opetition at $15 billion a year. And Anthropic — candid, careful, genuinely alarmed — publishing the data in the same breath as calling for a global pause.

The redirects are on the record too. Sora 2 announced and discontinued within weeks. The Disney deal dissolved. Companies pivoting faster than the coverage can follow. The acceleration is not a straight line — it lurches, stalls and changes direction — but the overall vector has not changed.

The Institutional Solvent

Velocity favours the agile over the entrenched. When change accelerates past the rate at which institutions can respond, it acts as a solvent on the structures that exist primarily to protect themselves. Bureaucracy cannot convene fast enough to stop it. Transition pain is compressed rather than dragged across generations. The cost of individual failure collapses: you try, it breaks by Tuesday, and you pivot by Thursday.

The acceleration clears what it encounters, stripping out the deadwood. But nobody chooses what counts as deadwood. The selection is made blindly, by rigidity alone. What gets cleared is not what is obsolete — it is whatever cannot flex fast enough to absorb the shockwave. The process cannot distinguish between bureaucratic sclerosis and judicial due process, between regulatory capture and safety verification, between a fence that imprisons and a fence that protects.

This is what makes the gas pedal metaphor haunting. Not the speed — but the choice. We can accelerate or decelerate — but we have no brakes.

Jack Clark said they only have a gas pedal. That is not a metaphor. It is a description of the instrument panel as it currently exists.

We are chronicling these milestones. But at this velocity, the pixels blend into history, swallowed by the momentum.

Last updated 07.06.2026.