AI policy is a rapidly emerging field facing a core problem: some of its most consequential decisions may need to be made before reliable evidence exists about the systems being governed.
For example, suppose frontier AI developers were required to report very large training runs. The immediate effect seems straightforward because reporting gives regulators more visibility. The harder questions concern what happens next. Would firms comply, reorganize training to avoid thresholds, invest in algorithmic efficiency, lobby for exemptions, or relocate activity? Would rival countries interpret the restriction as reassuring restraint or as an opportunity to accelerate?
Simulation can help study these interactions, but its purpose should not simply be to predict exactly what AI will look like at a future date. A more defensible goal is to ask when a policy works, how it fails, and which changes in the world should cause policymakers to reconsider it.
Exploring uncertain futures
RAND researcher Steve Bankes distinguished exploratory modeling from the more conventional use of models to consolidate existing knowledge into a representation of a real system. When a system is sufficiently well understood, such a model can be used as a reasonable surrogate for reality. Exploratory modeling is more appropriate when important mechanisms remain uncertain. Instead of relying on one best representation of the world, researchers run computational experiments across many plausible assumptions and examine how the conclusions change (Bankes, 1993).
AI governance fits this second category. We do not know with confidence how quickly capabilities will improve, how firms will respond to regulation, how international competition will evolve, or which technical bottlenecks will matter.
Robust Decision Making, or RDM, is a framework for making policy choices when the future is too uncertain to support a single reliable forecast. Developed by Robert Lempert, Steven Popper, Steve Bankes and colleagues at RAND, RDM reverses the usual modeling approach. Instead of first choosing the most likely future and then finding the policy that performs best in it, researchers begin with possible policy options and test them across many plausible futures. The goal is to identify strategies that continue to meet important policy objectives across a wide range of conditions, while also revealing the circumstances in which they fail (Lempert, Popper & Bankes, 2003).
In AI governance, researchers might compare compute reporting, licensing, mandatory safety evaluations and no new intervention. Each option could be tested under different assumptions about capability growth, regulatory enforcement, corporate compliance and international competition. The objective would be to identify which strategies continue to meet specified goals across many plausible conditions and where each strategy begins to break down.
Running hundreds of simulations, however, is not enough. Researchers also need to understand why particular policies fail. Scenario discovery provides methods for identifying combinations of uncertain conditions associated with important outcomes. Bryant and Lempert (2010), for example, show how analysts can identify regions of the uncertainty space where a policy becomes vulnerable.
A regulation might work under most tested conditions but fail when enforcement is weak, competitive pressure is high and algorithmic efficiency improves rapidly. This produces a more useful finding than a precise-looking risk estimate:
“This policy performs well under many tested conditions, but becomes fragile under weak enforcement and high competitive pressure.”
A strategic multi-agent simulation
Agent-based modeling provides one useful foundation for studying the effects of AI policy in a dynamic system of interacting actors. Governments, companies and regulators can each be represented with different objectives, resources and information, allowing the simulation to examine how their individual decisions combine to produce broader outcomes (Bonabeau, 2002).
AI governance may depend heavily on a relatively small number of powerful actors, such as frontier AI labs, major governments and key compute providers. The central problem is therefore less about mass behavior and more about small-N strategic interaction. In other words, a small number of actors make decisions partly in anticipation of how the others will respond.
For that reason, AI policy simulations should borrow not only from agent-based modeling but also from game theory and wargaming. Armstrong, Bostrom and Shulman’s Racing to the Precipice, for example, models how competition between AI developers can create incentives to reduce safety precautions (Armstrong, Bostrom & Shulman, 2016).
A complementary approach appears in Intelligence Rising, developed by Shahar Avin, Ross Gruetzemacher and James Fox. In this role-play methodology, human participants represent governments, AI companies and other actors navigating possible AI futures. The purpose is to explore how strategic interactions and unexpected dynamics can emerge as technology develops (Avin, Gruetzemacher & Fox, 2020).
A useful design may therefore be a strategic multi-agent simulation that combines an explicit model of the world with adaptive actors.
Turning policy into mechanisms
One of the hardest parts of building this kind of simulation is translating real-world policy into rules and variables that the model can actually use.
A compute-reporting policy cannot simply be entered into the simulation as “more regulation.” The model needs to specify who imposes the rule, which companies it applies to, what level of compute triggers reporting, how violations are detected, what penalties apply, and how companies might respond.
Sastry, Heim, Belfield and colleagues provide a useful framework. They argue that compute is particularly promising for governance because it is relatively detectable, excludable and quantifiable, while important parts of the advanced semiconductor supply chain are highly concentrated. They group compute governance around visibility, allocation and enforcement (Sastry et al., 2024). In practical terms, this means compute is something governments may be able to observe, control access to, and use as a basis for enforcing rules on advanced AI development.
A simulation could therefore vary factors such as detection probability, enforcement strength, compliance costs, willingness to evade and the speed of algorithmic improvement.
Policy itself should also be dynamic. Elections, courts, lobbying and changes of government can weaken, strengthen or repeal rules. Policy durability should therefore be modeled as an uncertainty rather than assumed to remain constant.
Using LLM agents carefully
Large language models could allow simulated actors to reason about competition, regulation and negotiation rather than simply selecting from a fixed list of actions.
But realistic-sounding behavior is not necessarily valid behavior.
LLMs introduce several methodological problems. First, because they are trained on large bodies of public text, simulated actors may reproduce familiar narratives about AI races, regulation or geopolitics instead of responding only to the incentives specified by the researcher.
Second, LLM behavior can contain systematic biases. Chuang and colleagues, for example, found that LLM agents in an opinion-dynamics simulation tended toward factual consensus, making some forms of persistent disagreement difficult to reproduce (Chuang et al., 2024). In an AI governance simulation, this kind of bias could distort results if agents appear more willing to converge than real governments or companies would be.
Third, reproducibility is a challenge. Chen, Zaharia and Zou found that the behavior of the same commercial LLM service can change substantially across model versions (Chen, Zaharia & Zou, 2024). A simulation may therefore produce different results after the underlying model is updated, even if the code and prompts remain unchanged.
A stronger design is likely to be hybrid. Explicit models can determine measurable quantities such as compute, budgets and training costs, while LLMs handle genuinely qualitative decisions such as negotiation or strategic response. Researchers should record prompts, model versions and decision traces and test whether important conclusions remain stable across different models.
From simulation to adaptive policy
Return to the compute-reporting example. Suppose repeated simulations show that reporting consistently improves government visibility but does little to reduce risk when enforcement is weak. Researchers could then test stronger audits, international information sharing or additional enforcement mechanisms and identify which combinations remain effective across uncertain futures.
RDM and scenario discovery can therefore reveal where a policy becomes vulnerable. The next question is what policymakers should do when the world begins moving toward one of those vulnerable conditions.
Dynamic Adaptive Policy Pathways, or DAPP, is a planning approach designed for situations where conditions are expected to change and no single policy is likely to remain appropriate indefinitely. Rather than selecting one fixed policy for the future, DAPP sets out a sequence of possible actions, together with signposts that policymakers monitor and triggers that indicate when a change in policy may be necessary (Haasnoot et al., 2013).
In AI governance, a government might begin with reporting requirements, introduce audits if evidence of evasion increases, and move toward licensing if model capabilities or training scale cross predetermined thresholds.
Simulation can therefore help design not only a policy, but a policy pathway.
Finally, the model itself must remain open to scrutiny. Researchers decide which actors matter, which outcomes count as harmful, which assumptions are considered plausible and which futures are included. Running more scenarios does not eliminate these judgments.
The purpose should therefore not be to build an oracle. It should be to build a policy stress-testing laboratory that exposes failure modes, strategic reactions and critical assumptions and helps policymakers understand when their strategy should change.