Startup Bets

Anthropic asks for testing instead of bans

 ·  By Sophronia Wentworth
Anthropic asks for testing instead of bans - anthropic safety testing
Anthropic asks for testing instead of bans

Anthropic CEO Dario Amodei is pushing back against calls for a ban on open-weight AI models, instead proposing a system of mandatory safety testing for highly capable systems. The proposal comes after a week of criticism and directly addresses a letter signed by OpenAI and Google, which seeks stricter restrictions on the distribution of these models.

Amodei answered that letter today with his own counterproposal. He argues that models without dangerous capabilities are a public good that should remain available. “Anthropic has never advocated for a ban on open-weight models,” Amodei writes, calling models without dangerous capabilities “a public good” that cost nothing beyond the compute to run them.

The letter outlines three main components. First, it calls for tighter export controls on advanced chips and chipmaking equipment. Second, it seeks to crack down on industrial-scale model distillation. Finally, it proposes requiring safety evaluations for every sufficiently capable model before release.

Amodei suggests this testing idea is close to consensus. He notes that the Trump administration has moved in this direction and that recent industry proposals would apply testing regardless of a model’s origin or license. The exemption for less capable models is critical, though. Once some models are exempt and others are not, the policy effectively becomes a capability threshold by another name.

That is the central ambiguity of the plan. Nobody has defined where “sufficiently capable” begins, and nobody has said who would set the number or who would run the tests. Model capability could be measured by benchmark results, training compute, or scores evaluated against defined risk categories. And the chosen metric would dictate the model type, effectively setting release timing for every lab that publishes weights.

Related: IBM Disappoints in Second Quarter Earnings Report

On Monday, Nvidia launched the Open Secure AI Alliance with more than 35 founding members. The group includes Microsoft, IBM, Hugging Face, Palantir, Red Hat, CrowdStrike, and the Linux Foundation. Nvidia argues that defenders need open frontier models to investigate incidents. The company cited the recent case in which closed tooling blocked forensic work on a Hugging Face incident and Hugging Face ran the analysis on an open-weight model instead.

Amodei concedes the testing “would need to be global, which means even the CCP would need to be on board,” and figures that may be possible. This is the largest unanswered question around the story. Export controls and anti-distillation measures are policies Washington can impose on its own, but a global capability-testing regime depends on broad international agreements that do not yet exist (and may never exist).

The letter so far counts 50 signatories, with OpenAI and Google joining Microsoft, Nvidia, Meta, AMD, Cisco, Cloudflare, GitHub and Ollama. Amodei and Amazon remain notable holdouts. To his credit, Amodei addressed the obvious criticism directly. A U.S. ban, he said, would mostly shield domestic AI companies from competition while doing little to stop bad actors. “It would protect U.S. AI companies from competition, but that has never been my goal.” That’s fair. It’s also true that Amodei does not release open-weight models, so the statement feels a touch empty.

Amodei argues that a testing regime is the only way to balance safety and openness. The idea of a global capability-testing regime depends on broad international agreements that do not yet exist. [1]Pilot Protocol has launched a platform to facilitate this agent economy, but the central ambiguity of the plan remains: nobody has defined where “sufficiently capable” begins.

Leave a Comment

Your email address will not be published.