Technology Speed Limits
Abstract: We study optimal technology regulation when private learning occurs both through doing (scaling up the technology) and through waiting (as time passes). We show that an adaptive speed limit -- a cap on the rate at which the technology can increase per-unit time -- delivers optimal worst-case guarantees over all learning processes and/or preferences, and is the only time-consistent mechanism that does so.
- Policy Learning with Adaptively Collected Data (2021)
- Inducing Social Optimality in Games via Adaptive Incentive Design (2022)
- Monitoring with Rich Data (2023)
- Robust Monopoly Regulation (2019)
- Dynamic Evidence Disclosure: Delay the Good to Accelerate the Bad (2024)
- Robust Technology Regulation (2024)
- Pricing Novel Goods (2022)
- The speed of sequential asymptotic learning (2017)
- From Best Responses to Learning: Investment Efficiency in Dynamic Environment (2025)
- Continuous time asymptotic representations for adaptive experiments (2026)
Summary
- The paper demonstrates that adaptive speed limits uniquely maximize worst-case regulatory welfare under uncertain agent learning.
- It employs a multiparameter filtration framework distinguishing learning-by-doing from learning-by-waiting, enabling robust control analysis.
- The approach highlights the option value of cautious scaling and ensures time-consistent policy implementation for high-risk technologies.
Technology Speed Limits: A Formal Analysis of Robust Mechanisms for Regulating Risky Technological Scale-up
Introduction
This paper studies the design of optimal regulatory mechanisms for steering the scale-up of technologies—especially those with unknown, potentially catastrophic externalities—when learning occurs not only from direct experimentation (learning-by-doing) but also simply from the passage of time (learning-by-waiting). The agent (e.g., a technology developer with possibly misaligned incentives) faces uncertainty about the technology's value and impact, and updates beliefs through both modes of learning. The regulatory objective is to maximize (or robustly guarantee) aggregate welfare under the worst-case possible agent learning process and/or preferences. The primary result is the characterization of "adaptive speed limits" as the uniquely time-consistent, robustly optimal regulatory mechanism.
Model Structure
The technology level l evolves over discrete time periods t=1,…,T, and both agent and principal have partial knowledge about an underlying binary state θ (representing "safe" or "dangerous" technology). Their payoffs depend on both state and technology scale. Importantly, information accrues along a two-dimensional random field indexed by (l,t), following multiparameter filtration theory. The agent chooses an irreversible, non-decreasing technology path, and a regulatory mechanism imposes transfers (including hard constraints) as a function of the trajectory realized so far.



Figure 1: Schematic illustration of possible regulatory mechanisms, including level caps, adaptive time-varying caps, constant and increasing marginal taxes.
Central to the framework is the contrast between learning-by-doing (increasing l at fixed t) and learning-by-waiting (increasing t at fixed l). Unlike most bandit or optimal experimentation models, these learning channels are not assumed to commute; information about safety can arrive in a way that revises previous inferences or that is only interpretable after enough time. The associated stochastic control problem thus generalizes classic bandit learning, encompassing settings where the regulator cannot predict learning rates or their interaction.
Policy Mechanism Space
The regulatory mechanism is modeled as a (potentially path-dependent) transfer function. The most generic instruments considered are sequences of contingent transfers—encompassing both "hard" constraints (infinite penalties) and soft taxes. The mechanism can be made adaptive over time and observed history.
A key object is the "adaptive speed limit"—a time-indexed, possibly information-adaptive cap on the allowable increment in technology scale between periods. The regulator specifies, at each t and conditional on all observed data up to t, a maximal allowable scale. The agent cannot exceed this cap without incurring an infinite penalty; below the cap, there is no transfer.
Main Results: Robustness and Time-Consistency
Robust Worst-case Optimality
The central theorem asserts that adaptive speed limits uniquely maximize the regulator’s worst-case expected payoff over all admissible learning processes and agent preferences. That is, for any possible agent (with any convex, misaligned utility), and for any process specifying how much information becomes available (to the agent or principal), the agent cannot generate a path that would systematically undermine the regulator’s value guarantee given an optimally-chosen, path-adaptive speed limit.
In technical terms, solving the principal's direct-control problem under her information filtration yields the optimal cap sequence. Any deviation by the agent towards more cautious paths (relative to the cap) provides additional "option value"—the opportunity to halt or slow scale-up if adverse news arrives in the future, reflecting the irreversibility of the technology.


Figure 2: Illustration of speed limits and feasible technology trajectories (red: speed limit, blue: agent strategy within the limit).
This result is highly nontrivial, given the generality of the stochastic control problem. The agent's information structure and preferences are (from the regulator’s perspective) unknown and possibly adversarial; yet, simple, non-punitive rules based only on current technology and time are proven sufficient (and necessary) for robust control.
Option Value of Slowing Down
An important comparative statics insight is that more conservative scaling (whether chosen by the agent out of precaution or imposed by stricter caps) increases the value of the option to halt—precisely because the technology is irreversible and (in expectation) new information may reveal catastrophic downside. Thus, the regulatory logic does not incentivize maximal speed except where and when the information justifies it.

Figure 3: Illustration of the option value—relative payoffs when the agent slows the pace of scale-up relative to the speed limit.
Time-Consistent Implementation
Distinctively, the adaptive speed limit is not only optimal ex ante, but is also the unique mechanism that remains optimal at all interim histories—i.e., is time-consistent without requiring regulatory commitment or revising the set of admissible learning processes ex post. Formally, this is developed through a careful recursive analysis of continuation games after arbitrary partial history.
Any alternative mechanism attempting to deviate (e.g., by using taxes, penalties, or more complex path contingencies) would either lose robustness or create incentives for premature stopping or waiting that strictly worsen regulatory performance in some informational realization.
Figure 4: Illustration of mechanism “ironing”: Flattening continuation transfers enforces time consistency. If transfers slope downward, this creates incentives for premature stopping; ironing them to be flat below the speed limit is required for optimality.
Practical and Theoretical Implications
Application to AI and High-risk Technologies
This framework directly motivates the case for regulatory speed limits for scaling up AI systems whose real-world impacts—including catastrophic downside risks—are not known ex ante and for which we learn through both testing and the passage of time. The paper's abstract model gives a normative foundation to the emerging practice of "responsible scaling policies" now advocated by AI labs and some policy organizations.
Broad Regulatory Principle
A significant normative implication is that, absent perfect knowledge of both the process and preferences of agents at the technological frontier, regulators should default to speed limits whose magnitude is dynamically adjusted by accumulating public information. Taxes or other incentive-based mechanisms are strictly suboptimal if the regulator cannot commit ex ante or anticipate all possible learning dynamics.
Theoretical Developments
The multiparameter learning formalism advances the methodology for analyzing experimentation environments where two sources of learning interact and may not commute. This is notably stronger than classical sequential experimentation or bandit frameworks, and opens avenues for general regulation under deep uncertainty.
Furthermore, the robust/time-consistent characterization of speed limits suggests a modular approach for regulation under similar "learning in the face of irreversibility" scenarios beyond AI—such as biotechnology, resource extraction, or financial innovation.
Extensions: Mitigation, Endogenous Risk Reduction
The paper also considers (in a stylized way) the scenario in which slow scaling buys time not just for learning but for mitigation—such as institutional adaptation or safety R&D. In this extension, the arrival rate of transition from "dangerous" to "safe" technological states is endogenous and decreases with scale. Here, the same core conclusion holds: the regulator’s optimal mechanism is an information-adaptive speed limit, and its time-consistency/robustness is retained.
Conclusion
The paper provides a rigorous foundation for the regulation of risky, poorly-understood technologies, demonstrating that adaptive speed limits are the single robust policy instrument that achieves optimal worst-case performance across all possible agent learning processes and preferences, and that are time-consistent. This result directly informs debates on AI and other high-stakes domains, where both irreversibility and hard-to-anticipate emergent risks are central.
The methods generalize prior learning-by-doing and learning-by-waiting frameworks, and the formal application of multiparameter filtrations to policy design solves a previously open question in the design of simple, interpretable alignment mechanisms under extreme uncertainty. Future work may explore richer information structures, strategic agent learning, or multi-agent environments, but the foundational character of the speed limit mechanism appears robust to such extensions.
References
See (2606.01424) for full bibliography and proofs.
Paper to Video (Beta)
No one has generated a video about this paper yet.
Whiteboard
No one has generated a whiteboard explanation for this paper yet.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Open Problems
We haven't generated a list of open problems mentioned in this paper yet.
Continue Learning
- How do adaptive speed limits compare with traditional regulatory instruments in controlling technological risks?
- What advantages does the multiparameter filtration framework offer over classical sequential experimentation models?
- How might regulatory policymakers implement adaptive speed limits in real-world scenarios, particularly for AI technologies?
- What are the implications of agent misaligned incentives on the effectiveness of the proposed regulatory mechanism?
- Find recent papers about adaptive regulatory mechanisms.
Collections
Sign up for free to add this paper to one or more collections.
Tweets
Sign up for free to view the 1 tweet with 4 likes about this paper.