ESSAY

Does AI Safety Testing Change What Gets Built?

The evidence is stronger for release than for construction

Niraj Shukla 9 min read

Eight ideas in this essay
Begin reading
01

AI safety tests can change how models are built and released, but that does not prove they make AI safer.

01 The Argument

AI safety testing already shapes how advanced models are developed, secured and released. It can redirect engineering, strengthen safeguards and alter the conditions under which a model reaches the public. Yet when a commitment becomes competitively costly, a company may revise the rule rather than let the rule determine its conduct. Even where testing changes development or deployment, the public evidence tells us much less about whether those changes reduce harm after release.

The purpose of this article is to identify the evidence that would distinguish meaningful constraint from organized compliance. The question is no longer whether safety frameworks have effects but whether they remain binding when their consequences become expensive, and whether the decisions they produce make deployed systems safer.

03

Anthropic added stronger protections to Claude Opus 4 because its tests could not rule out more dangerous capabilities.

03 The May Release

Major developers have adopted voluntary frameworks built on a similar principle: stronger capabilities should bring stronger safeguards (Williams & Freund, 2026). Anthropic offers the clearest public example of both their influence and their limits. Companies that disclose less about their policies provide less evidence to examine, so the public record may not represent the industry as a whole.

In May 2025, Anthropic released Claude Opus 4 under its AI Safety Level 3 (ASL-3) protections. The company had not concluded that the model crossed the relevant capability threshold, but said it could no longer confidently rule that out. It applied stronger deployment and security measures while continuing its evaluation (Anthropic, 2025). The policy materially affected how the model reached the public.

04

Anthropic later weakened some future commitments after arguing that acting alone could put it at a competitive disadvantage.

04 The Rewrite

In February 2026, Anthropic substantially rewrote the policy. Earlier versions committed the company to pause development or deployment if it could not implement the safeguards required before reaching the next capability level. Version 3.0 replaced some of those commitments with company plans, public goals and recommendations for the industry. Anthropic argued that higher safeguards might be impossible or strategically damaging for one company to implement while competitors continued without them (Anthropic, 2026a; Williams & Freund, 2026).

The earlier policy already allowed Anthropic to lower safeguards if competitors failed to adopt comparable standards, and the company did not reduce its existing ASL-3 protections. What changed was the strength of its commitment to future safeguards (Williams & Freund, 2026). A later revision clarified that Anthropic remained free to pause development even when the policy did not require one (Anthropic, 2026b).

In 2025, the framework changed a deployment and in 2026, anticipated competitive pressure changed the framework. The rule acted on the organization, and the organization acted on the rule.

In 2025, the framework changed a deployment and in 2026, anticipated competitive pressure changed the framework.

05

Safety rules can influence real decisions, but their force depends on how companies respond when following them becomes costly.

05 Engine or Costume

Donald MacKenzie’s history of the Black-Scholes formula offers one interpretation, and his title states the claim: the formula was an engine, not a camera. It began as an account of how options ought to be priced and became a force shaping how they were priced (MacKenzie, 2006). MacKenzie called the strong form of this Barnesian performativity: using a model makes the world more like the model. A safety threshold can work the same way. Once written down, it enters the environment in which technical and commercial decisions are made rather than describing that environment from outside.

John Meyer and Brian Rowan offer the skeptical account. Organizations often adopt formal structures as ceremony, gaining legitimacy while insulating core activity from their demands (Meyer & Rowan, 1977). On this view, a safety framework is a costume: visible and socially useful, but not necessarily connected to the work underneath.

Neither account is sufficient. Anthropic’s framework was not an empty costume because it changed a release. Nor was it an engine running on fixed settings because Anthropic weakened its strongest forward commitment after concluding that unilateral compliance could conflict with competitive pressure. Matthew Kraatz and Edward Zajac found a similar limit to institutional conformity among American colleges, which departed from accepted templates when market and technical pressures favoured change without suffering the expected penalties (Kraatz & Zajac, 1996).

06

Evidence that testing changes a model or its release conditions is not the same as evidence that it prevents harm.

06 Three Outcomes

This suggests there is a better way to assess AI safety rules. Three outcomes must be separated: whether testing changes the model, whether it changes deployment, and whether either change reduces harm after release.

Evidence about the first remains thin. One possible signal would be clustering below the AI Act’s compute threshold, since economists use bunching around income tax boundaries to detect when a legal line changes behaviour (Saez, 2010). Repeatedly capping or restructuring training runs below 1025 operations would show the threshold entering development decisions. For now, too few runs near the line are publicly documented with enough precision for a credible analysis. Epoch AI identified more than 30 models estimated to exceed the threshold as of June 2025, but many estimates were derived indirectly from model performance rather than company disclosures (Rahman et al., 2025). The more tractable question is qualitative: has any developer acknowledged redesigning a training run because crossing the line would trigger additional obligations?

Evidence about deployment is easier to find. Anthropic’s ASL-3 decision shows testing producing stronger access controls and release conditions. Comparable evidence would include delayed releases, restricted tools, withheld model weights or cancelled features, provided the company can identify the evaluation result that caused the decision and explain what would otherwise have happened.

07

Tests can record what happened before release, but they struggle to predict what users may discover afterwards.

07 The Open World

The third outcome is the hardest and most important. The apparatus creates detailed records of evaluations and organizational decisions, but much weaker evidence that those decisions reduce harm. The European Code of Practice addresses this through continuing risk management, incident reporting, post-market monitoring and external evaluation (European Commission, 2025). These measures may reveal failures. They cannot readily establish how much worse an event would have been without a safeguard, or whether a model that passed a bounded test will remain safe against new prompts, tools, fine-tuning methods and attacks.

The operational layer therefore deserves scrutiny. Independent experts prepared the aforementioned Code of Practice through a multi-stakeholder process, but providers hold much of the knowledge needed to decide how hard a model has been tested, which methods count as adequate elicitation and what remaining risk is acceptable. Anthropic’s own disclosures show how much room that leaves. Reviewing its first year under the framework, the company reported that some evaluations had lacked basic elicitation techniques such as chain-of-thought prompting, and it responded in part by extending the evaluation interval from three months to six so that future assessments would not be rushed. In February 2026 it judged that a model did not cross its automated AI research threshold while noting that confidently ruling this out was becoming difficult and required assessments more subjective than it would like. Providers can also revise the thresholds themselves: versions 3.3 and 3.4 changed those covering chemical and biological weapons production and automated AI research to better reflect the company’s stated threat models (Anthropic, 2026b). As Philip Mirowski and Edward Nik-Khah argued in another context, understanding what a measure does is incomplete without examining who selects and operationalizes it (Mirowski & Nik-Khah, 2007).

08

A safety system can be serious, expensive and influential while still steering the industry toward an imperfect measure of safety.

08 The Wrong Target

Patricia Bromley and Walter Powell describe organizations that implement a policy in full and never establish whether it achieves anything. They call this means-ends decoupling: the policy is implemented, the money is spent and the procedures are followed, while the connection between the activity and its intended result goes untested (Bromley & Powell, 2012). The regime may be sincere, technically demanding and powerful enough to reshape organizations, while still being a poor guide to what happens after release.

The reassuring and cynical accounts are both incomplete. Safety testing is not theatre, but neither is it a fixed constraint that companies obey unchanged. It is a contested instrument. It shapes models and releases while firms, markets and regulators reshape it in return.

A costume wastes cloth. An engine aimed slightly wrong can reorganize an industry around the wrong target.

A costume wastes cloth. An engine aimed slightly wrong can reorganize an industry around the wrong target.

Sources

References

  1. Anthropic. (2025, May 22). Activating AI Safety Level 3 protections.
  2. Anthropic. (2026a, February 24). Responsible Scaling Policy Version 3.0.
  3. Anthropic. (2026b, July 8). Anthropic’s Responsible Scaling Policy: Version 3.4 and version history.
  4. Bromley, P., & Powell, W. W. (2012). From smoke and mirrors to walking the talk: Decoupling in the contemporary world. Academy of Management Annals, 6(1), 483–530. doi:10.5465/19416520.2012.684462
  5. European Commission. (2025, July 10). General-Purpose AI Code of Practice: Safety and Security.
  6. European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. Official Journal of the European Union, L 2024/1689, July 12, 2024.
  7. Kraatz, M. S., & Zajac, E. J. (1996). Exploring the limits of the new institutionalism: The causes and consequences of illegitimate organizational change. American Sociological Review, 61(5), 812–836. doi:10.2307/2096455
  8. MacKenzie, D. (2006). An engine, not a camera: How financial models shape markets. MIT Press.
  9. Meyer, J. W., & Rowan, B. (1977). Institutionalized organizations: Formal structure as myth and ceremony. American Journal of Sociology, 83(2), 340–363. doi:10.1086/226550
  10. Mirowski, P., & Nik-Khah, E. (2007). Markets made flesh: Performativity, and a problem in science studies, augmented with consideration of the FCC auctions. In D. MacKenzie, F. Muniesa, & L. Siu (Eds.), Do economists make markets? On the performativity of economics (pp. 190–224). Princeton University Press.
  11. Rahman, R., Heindrich, L., Owen, D., & Emberson, L. (2025, June 6). Over 30 AI models have been trained at the scale of GPT-4. Epoch AI.
  12. Saez, E. (2010). Do taxpayers bunch at kink points? American Economic Journal: Economic Policy, 2(3), 180–212. doi:10.1257/pol.2.3.180
  13. Williams, S., & Freund, J. (2026, March 5). Anthropic’s RSP v3.0: How it works, what’s changed, and some reflections. Centre for the Governance of AI.