Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how AI agents can solve difficult, long-running tasks more effectively.
An AI agent is a LLM that does more than give one answer. It may work through many steps, ask other model “workers” to help, check its progress, use tools, and decide when to stop.
The authors introduce a method called agentic meta-reasoning. In simple terms, this means that the AI does not only think about the problem—it also thinks about how it is solving the problem.
For example, if an AI is trying to prove a difficult mathematics problem, it might ask itself:
- Should I continue my current proof?
- Should I check whether one step is correct?
- Should I try a completely different approach?
- Which previous attempts are worth reusing?
- Have I already found a good answer, and should I stop?
The paper argues that these decisions should be handled by a special controller, rather than being mixed together with the ordinary problem-solving work.
2. What questions are the researchers asking?
The researchers mainly want to know whether separating “thinking about the task” from “thinking about how to solve the task” makes AI agents better.
Their key questions are:
- Does meta-reasoning improve performance? Can an agent solve more problems correctly if it has a separate controller?
- Does it help more on longer and harder tasks? The researchers want to see whether the method becomes especially useful when an agent has many steps and a large computing budget.
- How does the agent use its time and computing power? Does it create new attempts, check old ones, or combine several earlier ideas?
- Does the agent find correct answers more often? Sometimes an agent creates a correct answer but fails to recognize it.
- Can the agent choose the correct answer from several possibilities? Finding a good answer and selecting it as the final answer are not always the same thing.
3. How did the researchers test the idea?
The controller and the workers
The proposed system has two main parts:
- Workers do the actual task, such as writing a proof or creating computer code.
- A controller decides what the workers should do next.
Imagine a school project. The workers are students doing research, writing drafts, and checking facts. The controller is like a project leader who decides:
- Which draft should be improved?
- Which idea needs checking?
- Should someone start a new approach?
- Which information should each student receive?
The controller does not need to remember every word from the entire project. Instead, it keeps a short summary of what has happened and stores the full work in a kind of memory.
Four stages of control
Each decision cycle has four stages:
- Assess The controller reviews what has been learned so far and updates its summary.
- Propose It suggests several possible next steps.
- Evaluate It decides which option is most useful given the remaining time or computing budget.
- Dispatch It sends the chosen task, along with the relevant earlier work, to one or more workers.
The agent can also decide to stop and submit an answer.
Artifact memory and graphs
Every piece of work is saved as an artifact. An artifact could be:
- A first attempt at a proof
- A criticism of that proof
- A revised solution
- A controller’s note
- A piece of computer code
The researchers record which artifacts were used to create later artifacts. This produces an artifact graph.
For example:
1 2 3 |
First proof attempt → Critique → Revised proof
↘
Second check |
This graph helps researchers see whether the AI is building on earlier work or simply starting many unrelated attempts.
Comparison systems
The main comparison was between:
- Meta-Reasoning Agent: uses the separate controller and the four stages.
- Direct Control Agent: uses the same workers and tools but chooses each next action in one step, based on its accumulated history.
This comparison is useful because the two systems have similar abilities. The main difference is how they control the work.
Tasks and budgets
The systems were tested on four types of problems:
| Benchmark | What it tests |
|---|---|
| IMO ProofBench-Advanced | Writing difficult mathematical proofs |
| ARC-AGI-2 | Discovering hidden rules in colored-grid puzzles |
| LongCoT-mini | Long reasoning problems in areas such as science, chess, and mathematics |
| ProgramBench | Rebuilding computer programs from descriptions and testing them |
The researchers used three advanced LLMs: Gemini 3.1 Pro, GPT-5.5, and Opus 4.8.
They gave the agents limited numbers of model calls. A model call is like one turn in which the AI thinks or performs an action. The reasoning tasks allowed up to 100 calls, while the programming tasks allowed up to 1,200 calls.
4. What did the researchers find?
Meta-reasoning usually improved results
At the main budgets, the Meta-Reasoning Agent performed better than the Direct Control Agent in all 12 matched comparisons: four benchmarks tested with three models.
The average improvements over direct control were:
- About 4.0 percentage points on the proof benchmark
- About 4.2 percentage points on ARC-AGI-2
- About 3.6 percentage points on LongCoT-mini
- Clear improvements on ProgramBench as well
For example, on ProgramBench:
- GPT-5.5 with Direct Control: 63.7%
- GPT-5.5 with Meta-Reasoning: 71.5%
The meta-reasoning system also scored higher than some existing coding agents. With GPT-5.5, it achieved 71.5%, compared with 58.0% for Codex.
It benefited more from larger budgets
When the agents were allowed more model calls, meta-reasoning generally kept improving.
For example, on ProgramBench with GPT-5.5, increasing the budget from 400 to 1,200 calls raised performance from 64.1% to 71.5%.
The Direct Control Agent often stopped improving. With the larger budget, it sometimes used more calls but did not produce much better answers.
This suggests that simply giving an AI more time is not enough. The AI also needs a good plan for deciding how to spend that time.
It reused earlier work more often
The Meta-Reasoning Agent created more links between earlier and later work. Instead of treating every attempt as completely new, it was more likely to:
- Check an earlier attempt
- Repair a mistake
- Combine several ideas
- Ask a worker to improve a promising answer
This is similar to writing an essay. A weak approach is to write ten completely separate essays. A better approach may be to write a draft, find its problems, revise it, and then ask someone else to check the revision.
It found correct answers more often
The researchers separated two different abilities:
- Coverage: Did the agent create a correct answer at some point?
- Selection: Did the agent actually submit that correct answer?
Meta-reasoning improved coverage in most settings. This means that more runs contained at least one correct candidate answer.
However, finding a correct answer did not always mean that the agent submitted it.
Choosing the best answer was still difficult
Sometimes the agent had both a correct and an incorrect answer in its memory. It then had to decide which one to submit.
The controller’s judgments were often better than the workers’ own confidence ratings. For example, on one proof task, worker confidence was only slightly better than guessing, while the controller’s judgments were much more accurate at ranking good and bad answers.
However, the improvement in final answer selection was not consistent in every test. Meta-reasoning helped in some settings but was tied with or slightly worse than direct control in others.
The method can be expensive at small budgets
The controller itself uses model calls. These calls do not directly solve the problem; they are used to plan and evaluate the work.
At small budgets, this extra planning can take away too many calls from the workers. As a result, meta-reasoning sometimes performed worse than direct control when the budget was very limited.
This is an important trade-off:
- With little time, extra planning may be a burden.
- With lots of time, better planning may help the agent use its resources wisely.
5. Why are these results important?
The paper suggests that future AI systems should not only become better at producing individual answers. They should also become better at managing long problem-solving processes.
A powerful agent may need to:
- Remember useful earlier work
- Detect when an idea is failing
- Decide whether to explore or improve
- Spend more effort on uncertain parts
- Recognize when it has already found the right answer
- Avoid throwing away good work
This is especially important for tasks such as:
- Writing large software systems
- Proving difficult mathematics
- Conducting scientific research
- Planning complicated projects
- Solving problems that require hundreds or thousands of steps
The artifact graphs also give researchers a way to study how an AI reached an answer, rather than looking only at whether the final answer was right or wrong.
Conclusion
The main message of the paper is that AI agents can improve when they spend some of their computing power thinking about what to do next.
The Meta-Reasoning Agent separates planning and self-checking from ordinary task-solving. It keeps a short summary of progress, stores detailed work in memory, and chooses new actions through a structured process.
The experiments show that this approach usually leads to better results, especially on long and complicated tasks with larger budgets. It helps agents reuse earlier work and find correct answers more often. However, it also costs extra time and computing power, so it may not help when the budget is very small.
In the future, AI systems may work less like a person answering one question at a time and more like a whole team: some parts doing the work, and another part organizing, checking, and deciding how the team should proceed.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The paper does not isolate which component of the four-stage controller—
Assess,Propose,Evaluate, orDispatch—produces the observed gains; the Direct Control comparison changes staging, state compression, memory access, and deliberation simultaneously. - The contribution of compact state compression is not separately compared with alternatives such as full-history control, structured summaries, learned memory, or programmatic retrieval.
- The paper does not establish whether explicit meta-reasoning is superior to a simpler controller with equivalent additional calls, such as extra worker samples, verifier calls, critique rounds, or self-consistency selection.
- The qualitative
Evaluatestage is not compared with a calibrated or learned value-of-computation estimator, leaving unclear whether the gains arise from meta-reasoning itself or from improved budget allocation. - The controller's proposal quality is not evaluated independently from its ability to select among proposals; failed runs are not attributed separately to missing useful actions versus poor evaluation of available actions.
- The paper does not report ablations on controller-stage prompts, stage order, number of controller calls, or whether all four stages are necessary.
- The controller is not trained or fine-tuned for orchestration, so it remains unclear whether the approach scales with specialized controller training, reinforcement learning, preference optimization, or task-specific orchestration policies.
- The experiments use only three frontier models and do not establish whether the method benefits weaker models, open-weight models, multimodal models, or models with different instruction-following and self-evaluation capabilities.
- The benchmarks are limited in number and scope, with especially small samples for IMO ProofBench-Advanced (30 problems) and ProgramBench (200 problems); the statistical significance and confidence intervals of the reported gains are not provided.
- The paper reports point estimates but does not adequately quantify run-to-run variance, random-seed sensitivity, or uncertainty across repeated executions of the same problem.
- The external-baseline comparisons are not fully matched: production coding agents have different native interfaces, prompts, stopping behavior, and tool-use procedures, so their results cannot cleanly identify the effect of meta-reasoning.
- The nominal model-call budget is only a coarse resource measure; the systems are not compared using token consumption, FLOPs, monetary cost, wall-clock latency, energy use, or parallelism-adjusted compute.
- Budget enforcement allows workers dispatched before exhaustion to finish, potentially causing overspending; the effect of this accounting rule on comparisons is not analyzed.
- The paper does not determine the minimum budget at which meta-reasoning becomes cost-effective, nor whether its overhead can be adaptively reduced for easy problems.
- The reported scaling trends cover only budgets up to 100 calls for reasoning tasks and 1,200 calls for ProgramBench; it remains unknown whether the gains persist, saturate, or reverse at substantially larger budgets.
- The method's behavior under strict latency, concurrency, API-rate, or monetary constraints is not evaluated, despite the additional sequential controller stages likely increasing response time.
- The artifact graph records only controller-selected context dependencies and therefore omits information transferred through environment state, files, tools, hidden model context, shared prompts, and unrecorded worker interactions.
- Because the artifact graph is incomplete for coding agents, the reported topology differences may not accurately represent the full computation performed on ProgramBench.
- The paper does not test whether alternative memory-retrieval strategies—embedding search, relevance ranking, recency policies, hierarchical summaries, or learned retrieval—improve on controller-selected artifact access.
- The controller's self-authored notes are repeatedly reread, but the paper does not establish whether these notes contain reliable information, introduce confirmation bias, or amplify early mistakes.
- The impact of erroneous or lossy
Assesssummaries is not quantified; a mistaken compressed state could systematically discard useful evidence from the persistent artifact memory. - The system lacks an explicit mechanism for correcting corrupted state or detecting when its own summary conflicts with the underlying artifacts.
- Worker context selection is treated as a controller decision, but the paper does not measure how often relevant artifacts are omitted, irrelevant artifacts are included, or context selection causes failures.
- The analysis of coverage and selection is limited to reasoning benchmarks; no equivalent intermediate-artifact grading protocol is provided for ProgramBench, where the largest practical gains are reported.
- Intermediate solution correctness is not always comparable to final benchmark correctness: IMO ProofBench uses partial proof scores, whereas the coverage analysis binarizes artifacts, potentially discarding meaningful degrees of correctness.
- The reliability of intermediate-artifact labels and per-domain verifiers is not examined, including grader error, evaluator-model bias, or disagreement between automated and human assessment.
- Worker and controller confidence are evaluated using Type-2 AUC only; calibration, absolute confidence quality, selective prediction, and decision-theoretic usefulness are left unresolved.
- The analysis does not determine whether controller verdicts improve because they access more evidence, because they receive better prompts, or because they are given a distinct role from workers.
- The frontier-selection baseline is structurally defined and may not represent the strongest available selection strategy; comparisons against learned verifiers, majority voting, pairwise comparison, or independent judge ensembles are absent.
- Selection gain can reflect choosing a shallower artifact rather than improving terminal-candidate discrimination, but these effects are not separated.
- The paper does not investigate failure modes in which meta-reasoning overcommits to an initially plausible but incorrect approach, repeatedly reuses flawed artifacts, or wastes budget on increasingly elaborate repairs.
- The experiments do not evaluate adversarial, deceptive, or distribution-shifted intermediate artifacts designed to mislead the controller.
- The method's robustness to incorrect worker outputs, unreliable tools, execution failures, malformed artifacts, or inconsistent environmental state is not characterized.
- ProgramBench evaluation uses hidden tests, but the paper does not analyze whether meta-reasoning improves generalization to substantially different program interfaces, languages, repositories, or execution environments.
- The reasoning benchmarks do not establish transfer to interactive, open-ended, safety-critical, or real-world tasks where correctness is ambiguous and no deterministic verifier is available.
- The paper does not compare meta-reasoning with human-designed hierarchical planning, beam search, Monte Carlo tree search, debate, verifier-guided search, or adaptive test-time scaling under equalized compute.
- The controller's proposed actions are not analyzed for diversity; it remains unclear whether meta-reasoning expands the search over genuinely different approaches or mainly produces correlated variations of the same line of work.
- The relationship between graph reuse and actual causal usefulness is unresolved: more edges and deeper graphs may indicate productive composition, redundant context accumulation, or merely more controller activity.
- The paper does not test whether deliberately constraining graph depth, fan-in, branching, or memory reads can preserve performance while reducing cost and latency.
- The method's stopping behavior is underexplored; there is no systematic analysis of premature stopping, unnecessary continuation, forced submission after budget exhaustion, or optimal stopping criteria.
- The paper does not establish whether the controller can reliably recognize that no available artifact is correct and should initiate a genuinely fresh search rather than continue repairing existing work.
- Reproducibility is limited by missing implementation details in the provided text, including complete prompts, model sampling parameters, retrieval policies, stage-specific call limits, exact baseline modifications, and handling of malformed or failed calls.
- The paper excerpt ends before the complete memory/context results, appendices, and presumably detailed statistical analyses, leaving unresolved whether the reported qualitative conclusions are fully supported by the omitted evidence.
Practical Applications
Immediate Applications
The paper’s main practical contribution is an inference-time orchestration pattern: separate task execution from control of execution, retain intermediate artifacts in persistent memory, and use staged assessment, proposal, evaluation, and dispatch to decide what to compute next.
- Long-horizon software engineering agents — software development
- Integrate the framework into coding agents that inspect repositories, edit files, run tests, and revise implementations.
- A controller could:
- summarize the current implementation state;
- identify unresolved failures or missing functionality;
- propose alternatives such as debugging a failing test, inspecting a dependency, reverting to an earlier design, or starting a fresh implementation;
- allocate additional calls to the highest-value option;
- pass only relevant files, test logs, and prior patches to each coding worker.
- This could produce tools for repository reconstruction, legacy-system migration, automated bug fixing, and large refactoring workflows.
- Evidence: On ProgramBench, the meta-reasoning system achieved a higher test-pass rate than the matched direct-control agent and the evaluated coding-agent baselines, particularly at larger budgets.
- Dependencies and assumptions: The repository must support reliable execution and testing; hidden-test performance must correlate with real-world correctness; controller calls must be affordable; and access controls are needed before agents can modify production systems.
- Adaptive debugging and test-generation workflows — software quality assurance
- Use artifact graphs to preserve relationships among source files, patches, test outputs, critiques, and failed approaches.
- When a patch fails, the controller can decide whether to:
- repair the current patch;
- generate an independent solution;
- request targeted test generation;
- inspect an earlier artifact;
- compare competing implementations.
- A practical product could be an IDE or CI plugin that maintains a persistent “debugging graph” rather than a linear chat transcript.
- Dependencies and assumptions: Tests must be sufficiently informative; artifact provenance must be captured accurately; and the system should avoid repeatedly exploring equivalent fixes.
- Research-problem solving assistants — academia
- Apply the architecture to mathematical proofs, formal derivations, literature-grounded analysis, and experimental planning.
- Workers could independently propose proofs or hypotheses, while the controller tracks which lemmas, assumptions, experiments, or citations remain unverified.
- The system could assign specialist workers to proof checking, counterexample search, statistical review, literature retrieval, or result synthesis.
- The coverage/selection distinction is especially useful: a research assistant can separately report whether a valid candidate was found and whether it was selected as the final result.
- Dependencies and assumptions: Intermediate outputs need domain-appropriate verifiers. Human experts remain necessary for claims that lack deterministic validation, and the system must preserve citations, provenance, and uncertainty rather than treating controller judgments as proof.
- Automated verification and review pipelines — formal methods, compliance, and cybersecurity
- Use explicit controller assessment to review competing artifacts such as proofs, security patches, threat analyses, configuration changes, or compliance reports.
- The controller can prioritize targeted verification of uncertain components instead of rerunning an entire analysis.
- Artifact graphs can provide an audit trail showing which evidence supported each conclusion and which workers generated or reviewed it.
- Dependencies and assumptions: The review process requires trusted validators or independent checkers. Language-model confidence alone is insufficient, as the paper shows that worker confidence can be close to chance in some settings.
- Cost-aware orchestration of model calls — AI infrastructure
- Implement the four-stage control cycle as a general-purpose runtime for multi-agent systems.
- The runtime can dynamically choose between independent sampling, critique, repair, synthesis, memory retrieval, and termination according to a remaining call budget.
- This is applicable to customer-support automation, document processing, data analysis, and internal knowledge assistants where tasks vary substantially in difficulty.
- Compact state can reduce the need to replay the full interaction history, potentially lowering context length and improving run manageability.
- Dependencies and assumptions: The current evaluation uses model-call budgets rather than a complete cost model. Real deployments must account for token volume, latency, parallelism, API pricing, memory retrieval costs, and tool execution time.
- Agent observability and failure diagnosis — MLOps and enterprise automation
- Record each worker output as an artifact and each supplied context as a directed dependency edge.
- Operators can then distinguish:
- failures where no correct candidate was produced;
- failures where a correct candidate existed but was not selected;
- excessive branching or redundant work;
- overreliance on a flawed artifact;
- premature stopping.
- A monitoring dashboard could expose graph depth, branching, reuse, budget utilization, candidate coverage, and selection quality.
- Dependencies and assumptions: The artifact graph captures only controller-recorded dependencies and may omit information obtained through unlogged tools or environmental state. Production systems therefore need comprehensive event logging and privacy controls.
- Adaptive educational tutoring — education
- A tutoring agent could maintain artifacts representing a student’s attempted solution, hints, misconceptions, corrections, and alternative explanations.
- Instead of always generating the next hint, the controller could decide whether to:
- ask the student to retry;
- provide a targeted explanation;
- verify a specific step;
- present a counterexample;
- switch instructional strategies.
- The compact state could preserve the student’s learning trajectory without repeatedly exposing the full transcript to every worker.
- Dependencies and assumptions: Educational quality requires pedagogical evaluation, not only answer accuracy. The system must avoid revealing solutions too early, protect student data, and adapt to individual learning needs.
- Multi-stage document and data analysis — legal, finance, and business operations
- Use workers for extraction, independent analysis, critique, fact checking, and synthesis, with the controller selecting which stage to execute next.
- For example, a financial-analysis agent could preserve candidate forecasts, assumptions, sensitivity analyses, and validation results, then direct additional work toward the most consequential uncertainty.
- In legal workflows, the artifact graph could connect claims, source passages, counterarguments, and review notes.
- Dependencies and assumptions: Domain-specific retrieval and verification are essential. The system must not infer that a candidate is reliable merely because it has undergone more agentic processing.
- Daily-life planning assistants — personal productivity
- Apply the framework to complex planning tasks such as travel, household projects, job searches, event planning, or budgeting.
- The assistant could maintain alternatives, constraints, reservations, costs, and user preferences as artifacts, then decide whether to search for more options, verify a constraint, revise a plan, or stop.
- Dependencies and assumptions: External information must be current; actions involving purchases, bookings, or messages require explicit user confirmation; and the system must handle changing constraints and sensitive personal data.
Long-Term Applications
The reported results suggest broader applications, but these require stronger validation, improved cost models, better monitoring, and integration with reliable external tools.
- Autonomous scientific discovery systems — science and research
- A future system could coordinate literature search, hypothesis generation, simulation, experiment design, statistical analysis, and replication.
- The artifact graph would represent dependencies among hypotheses, datasets, experimental results, code, and interpretations.
- The controller could spend additional computation on the uncertainty most likely to change the scientific conclusion rather than simply producing more independent hypotheses.
- Dependencies: This requires trustworthy scientific simulators, laboratory interfaces, statistical safeguards, reproducibility protocols, and human approval for consequential experiments. The paper demonstrates improved reasoning benchmarks, not end-to-end scientific discovery.
- Formal theorem-proving and mathematical research agents — mathematics and verification
- Workers could generate proofs, search for counterexamples, invoke proof assistants, repair failed lemmas, and translate informal arguments into formal code.
- The controller could prioritize unresolved bottlenecks and select a proof artifact only after formal verification.
- Dependencies: Integration with proof assistants and dependable proof checkers is necessary. Qualitative controller judgments should guide search, but formal verification must determine correctness.
- Robotic task planning and recovery — robotics
- In embodied agents, workers could propose action sequences, inspect sensor data, simulate alternatives, diagnose failures, or recover from unexpected states.
- Persistent artifacts could include maps, trajectory plans, object detections, failed actions, and safety constraints.
- The controller could choose between continuing a plan, replanning, gathering another observation, or returning to a safe state.
- Dependencies: Real-time latency, partial observability, actuator reliability, simulation-to-reality transfer, and safety certification are major barriers. The paper’s model-call experiments do not establish suitability for physical-time control.
- Healthcare decision-support and clinical workflow orchestration — healthcare
- A system could coordinate medical-record summarization, guideline retrieval, differential-diagnosis generation, evidence checking, and patient-specific risk analysis.
- Its artifact graph could make the provenance of recommendations explicit and identify whether a decision was limited by missing evidence or by poor selection among available options.
- Dependencies: Clinical validation, privacy protection, calibrated uncertainty, bias assessment, regulatory approval, and clinician oversight are mandatory. The system should support—not replace—licensed clinical judgment.
- Energy-grid and industrial operations planning — energy and manufacturing
- Controllers could coordinate forecasts, simulations, maintenance plans, fault diagnoses, and contingency strategies.
- Additional computation could be directed toward the constraint or scenario with the greatest operational value, such as demand uncertainty or equipment failure.
- Artifact graphs could provide traceability for operational recommendations.
- Dependencies: Real-time systems require deterministic latency, validated simulators, secure interfaces, and fail-safe behavior. The qualitative value assessment described in the paper would need to be replaced or supplemented with calibrated operational objectives.
- Financial research, portfolio analysis, and risk management — finance
- Workers could generate investment theses, stress tests, scenario analyses, and independent risk reviews, while a controller allocates effort toward unresolved assumptions.
- Persistent artifacts would support auditability of forecasts and enable separation of “no viable analysis was found” from “a viable analysis was found but not selected.”
- Dependencies: Market nonstationarity, data quality, regulatory requirements, model risk, and adversarial incentives make uncalibrated self-assessment dangerous. Deployment would require strict human approval and independent quantitative validation.
- Self-optimizing agent runtimes — AI systems research
- Artifact-graph data could support learning better controllers, estimating the value of computation, and automatically adapting worker roles, context selection, and stopping policies.
- Future runtimes might learn when to use deep staged control, when to use a cheaper direct strategy, and how to predict whether a run is coverage-bound or selection-bound.
- Dependencies: The paper’s controller uses prompted qualitative evaluation rather than a learned value-of-computation model. Large, representative execution traces and reliable outcome labels would be needed to train and validate such systems.
- General-purpose autonomous organizations and workflow systems — enterprise automation
- Multiple specialized agents could maintain shared artifact memory for product development, procurement, incident response, policy analysis, or large-scale operations.
- A supervisory controller could allocate work across teams of agents, preserve decisions and evidence, and escalate uncertain or high-impact actions to humans.
- Dependencies: This requires robust identity and permission management, conflict resolution, shared data schemas, accountability mechanisms, and safeguards against error propagation through the artifact graph.
- Personalized lifelong AI assistants — daily life
- A long-term assistant could maintain structured memories of projects, preferences, commitments, and prior decisions while selectively retrieving only relevant artifacts.
- It could distinguish between planning, verification, execution, and reflection rather than treating every interaction as a linear conversation.
- Dependencies: Long-term memory raises substantial privacy, consent, deletion, security, and ownership issues. The assistant must also avoid reinforcing outdated or incorrect notes through repeated retrieval.
- Policy simulation and public-sector decision support — government and public policy
- Workers could model policy alternatives, summarize stakeholder evidence, identify implementation risks, and test outcomes under different assumptions.
- The controller could allocate computation to contested assumptions and preserve an auditable graph connecting recommendations to evidence.
- Dependencies: Policy decisions involve normative judgments that cannot be reduced to model confidence or benchmark accuracy. Public-sector use would require transparency, democratic accountability, independent review, and explicit separation between factual analysis and value judgments.
Overall, the most deployable near-term use is budget-aware orchestration of long-running software and reasoning agents with persistent artifact memory and execution tracing. The strongest long-term opportunities involve domains where intermediate artifacts can be independently verified; domains without reliable verification should treat the controller’s assessments as prioritization signals rather than authoritative judgments.
Glossary
- Abstract reasoning: Problem solving that requires manipulation of concepts or patterns rather than direct procedural execution. “ARC-AGI-2, spanning proof search, abstract reasoning, long-horizon reasoning, and program reconstruction”
- Agentic harness: A runtime system that coordinates model calls, tools, memory, and task execution. “Agentic harnesses instead use the model's own judgments to direct next computations”
- Agentic inference: Inference in which an agent adaptively performs actions and computations toward a goal. “We call this approach agentic meta-reasoning”
- Artifact graph: A directed graph representing artifacts and the dependencies between them. “We record the resulting work as an artifact graph.”
- Artifact memory: Persistent storage containing intermediate outputs and notes produced during an agent run. “Artifact memory holds worker outputs and the controller's own notes side by side.”
- AUC: Area under the receiver operating characteristic curve, a measure of ranking or classification quality. “We call this Type-2 AUC because the system is judging its own answers”
- Budget utilization: The fraction of an available computational allowance that an agent actually consumes. “Budget utilization measures the fraction of the allowance actually used.”
- Candidate solution: An intermediate artifact that may serve as a possible final answer. “Let $V_{\mathrm{sol}$ be the artifacts that can be graded as candidate solutions.”
- Chain of thought: A model-generated sequence of intermediate reasoning steps. “In a single model call, the object of that control is the chain of thought.”
- Convergence frontier: The deepest terminal artifacts in an artifact graph. “We call the deepest terminal artifacts the convergence frontier.”
- Coverage: The probability that a run produces at least one correct candidate solution. “Coverage counts a run as successful once it contains a correct solution”
- Direct Control Agent: A comparison agent that chooses actions directly from its accumulated history without explicit staged control. “We also introduce a Direct Control Agent, a matched variant with no explicit separation of control”
- Directed acyclic graph: A directed graph containing no cycles, so dependencies proceed in one direction. “Because the input artifacts already exist when the worker starts, these edges form a directed acyclic graph.”
- Dispatch: The process of converting a selected proposal into an executable worker action. “Dispatch turns the selected proposal into an executable action.”
- Epistemic action: An action performed to organize, acquire, or modify an agent’s knowledge. “These reads and writes are epistemic actions”
- Fan-in: The number of incoming edges to a node in a graph. “Fan-in counts a node's incoming edges”
- FLOPs: Floating-point operations, a common measure of computational work. “Note that equal call allowances do not imply equal token cost, latency, or FLOPs.”
- Frontier selection gain: Improvement over randomly selecting a candidate from the deepest terminal artifacts. “Positive gain means the agent does better than choosing uniformly from its frontier.”
- Inference-time computation: Computation performed while generating an answer rather than during model training. “These established inference-time computation as a major scaling lever”
- Long-horizon reasoning: Reasoning over extended sequences of interdependent steps. “LongCoT-mini, which tests long-horizon agentic capability”
- Meta-cognitive control: Monitoring and regulating one’s own reasoning or computational process. “This is a problem of metacognitive control”
- Meta-reasoning: Reasoning about and acting on an agent’s own inference process. “We call this approach agentic meta-reasoning”
- Monitoring: Assessing the quality or correctness of intermediate work during execution. “We also ask the controller for its own judgment”
- Nominal budget: The stated maximum computational allowance assigned to a run. “The reasoning benchmarks use nominal budgets of 25, 50, and 100 model calls per problem.”
- Object-level computation: Computation directly aimed at solving the target task rather than managing the solving process. “separate from the object-level computation it directs”
- Online artifact-graph construction: Building a dependency graph incrementally as an agent performs a task. “This motivates our view of agentic inference as online artifact-graph construction”
- Orchestration: Coordination of multiple models, agents, tools, or computational steps. “A complementary line of work learns, searches, or optimizes the system that coordinates model calls.”
- Persistent memory: Storage that remains available across multiple decisions or execution cycles. “Full worker outputs remain in persistent memory”
- Program reconstruction: Recreating a program’s behavior from documentation and an executable reference. “the solver must construct a self-contained codebase whose compiled executable matches the reference”
- Proof search: The process of exploring possible mathematical arguments to find a valid proof. “spanning proof search, abstract reasoning, long-horizon reasoning, and program reconstruction”
- Recursive LLM: A model framework that uses code to invoke sub-models recursively over parts of a task. “The Recursive LLM (RLM) holds the context in a variable the model edits with code”
- Selection: Choosing a correct candidate after one has been generated. “Success therefore factors into finding a correct answer and choosing it once one exists”
- State compression: Representing a large execution history using a smaller summary. “The comparison thus isolates the combined control design, but it does not separately identify the effects of staging, state compression, or memory access.”
- Stochastic output: An output influenced by randomness and therefore not fully determined by the same inputs. “A coding worker may read files not represented in , modify the environment, and produce stochastic outputs.”
- Test-time scaling: Improving performance by allocating more computation during inference or execution. “Whether such automation improves on simpler test-time scaling is under active investigation”
- Type-2 AUC: A ranking measure for how well a system evaluates the correctness of its own outputs. “We call this Type-2 AUC because the system is judging its own answers”
- Value of computation: The expected usefulness of spending additional computational resources on a particular action. “This is a prompted, qualitative assessment of computational value rather than an exact optimization”
- Worker dispatch: Assignment of a task, instructions, and selected context to a worker agent. “Let be the resulting worker-dispatch or stopping action.”






