Papers
Topics
Authors
Recent
Search
2000 character limit reached

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Published 29 Sep 2026 in cs.AI | (2609.38147v1)

Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

Summary

  • The paper presents a Meta-Reasoning Agent framework that separates control from computation, improving problem-solving performance by 4.0%, 4.2%, and 3.6% on average for three benchmarks.
  • The framework uses an explicit four-stage control cycle: Assess, Propose, Evaluate, and Dispatch, each with access to persistent memory and internal deliberation, with a method that does not ensure uniform improvements but rather benefits substantial computational horizons.
  • The controller’s compact state remains significantly smaller than direct-control accumulated history, limiting retrieval noise but requiring effective information- retrieval strategies.

Problem formulation and central contribution

“Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning” (2609.38147) addresses a control problem that becomes increasingly consequential as agents execute longer inference runs. An agent must repeatedly decide which partial result to trust, whether to refine or abandon an approach, how much context to expose to a worker, whether to initiate independent exploration, and when to stop. In conventional agent architectures, these decisions are interleaved with object-level problem solving and are often made in a single control step conditioned on an expanding transcript. The paper argues that control itself should be treated as an explicit agentic computation.

The proposed Meta-Reasoning Agent separates task-level computation from inference control. Workers generate proofs, analyses, or code, while a controller reasons over the evolving state of the run. The controller maintains a compact textual state and stores full worker outputs in persistent artifact memory. It therefore need not replay the entire history at every decision point. The framework represents a run as an artifact graph: nodes are worker or controller outputs, and directed edges record which prior artifacts were selected as context for subsequent work.

The core design is an explicit four-stage control cycle:

  1. Assess updates the controller’s account of what has been established, which approaches have failed, and which candidates remain promising.
  2. Propose generates possible next computations, including fresh attempts, targeted verification, repair, synthesis, and stopping.
  3. Evaluate estimates which proposal is worth executing under the remaining model-call budget.
  4. Dispatch converts the selected proposal into a worker instruction and selects the artifact context that worker receives.

The controller stages are themselves agentic processes with access to memory and, where appropriate, additional internal deliberation. Their calls are charged against the same nominal budget as worker calls. This constraint is important: the method does not assume that additional control is free. Its claim is instead that structured control can produce enough improvement in downstream computation to compensate for its overhead.

Figure 1

Figure 1: Meta-reasoning separates explicit control over compact state and persistent memory from object-level worker computation.

The principal baseline is a Direct Control Agent using the same underlying models, workers, interfaces, delegation mechanism, artifact-selection capability, stopping mechanism, and nominal call budgets. Direct control chooses actions in a single turn over the accumulated history. This comparison tests the complete control design, but not the independent causal contribution of staged control, state compression, or persistent memory access.

Experimental scope and evaluation methodology

The evaluation spans four substantially different tasks: IMO ProofBench-Advanced, ARC-AGI-2, LongCoT-mini, and ProgramBench. The first three use model-call workers that produce textual artifacts. ProgramBench uses coding-agent workers capable of inspecting files, executing commands, testing implementations, and reconstructing a program from an execute-only reference binary.

The underlying models are Gemini 3.1 Pro, GPT-5.5, and Opus 4.8. Reasoning benchmarks use nominal budgets of 25, 50, and 100 model calls per problem; ProgramBench uses 400, 800, and 1200 calls because coding workers require many calls within a single assignment. Controller and worker calls are both counted, and parallel calls are summed rather than treated as one batch. Since calls can differ in token length, the paper explicitly limits its resource claim to common nominal call budgets rather than equal token cost, latency, or FLOPs.

The benchmark suite probes distinct failure modes. IMO ProofBench-Advanced evaluates proof construction using scores in {0,1,6,7}\{0,1,6,7\}, ARC-AGI-2 requires exact grid transformation, LongCoT-mini tests long-horizon reasoning across several domains, and ProgramBench measures the fraction of hidden behavioral tests passed by reconstructed code. The external comparisons on ProgramBench include mini-SWE Agent, Claude Code, and Codex; Recursive LLM is evaluated on the reasoning tasks with modifications intended to impose a comparable model-call accounting and restrict free computation inside its workspace.

The paper supplements final accuracy with trajectory-level diagnostics. For reasoning benchmarks, intermediate solution artifacts are graded so that success can be factored into coverage—whether a correct candidate appeared anywhere in the run—and selection—whether the agent submitted a correct candidate once one existed. It also measures Type-2 AUC for worker confidence and controller verdicts, artifact-graph topology, memory operations, controller-state size, budget utilization, and frontier selection gain. This diagnostic decomposition is methodologically significant because identical final scores can conceal materially different failure modes: one run may never generate a correct answer, whereas another may generate one but discard it.

Performance across benchmarks and models

At the principal budgets—100 calls on the reasoning benchmarks and 1200 on ProgramBench—the Meta-Reasoning Agent obtains the highest reported point estimate in all 12 matched comparisons against Direct Control. Averaged across models, gains over direct control are 4.0 points on IMO ProofBench-Advanced, 4.2 points on ARC-AGI-2, and 3.6 points on LongCoT-mini. The effects are heterogeneous. Gemini 3.1 Pro receives the largest gain on IMO ProofBench-Advanced, 8.6 points, and the largest gain on LongCoT-mini, 9.2 points. ARC-AGI-2 is more consistent, with gains ranging from 2.5 to 6.7 points across models. GPT-5.5 improves only 0.4 points on LongCoT-mini, showing that the framework’s benefit is not uniform across model-task combinations.

The detailed matched results are:

Model Benchmark Direct control Meta-reasoning Difference
Gemini 3.1 Pro IMO ProofBench-Advanced 82.7 91.3 +8.6
Gemini 3.1 Pro ARC-AGI-2 77.5 84.2 +6.7
Gemini 3.1 Pro LongCoT-mini 53.5 62.7 +9.2
GPT-5.5 IMO ProofBench-Advanced 93.3 94.6 +1.3
GPT-5.5 ARC-AGI-2 75.8 79.2 +3.4
GPT-5.5 LongCoT-mini 64.7 65.1 +0.4
Opus 4.8 IMO ProofBench-Advanced 78.0 80.2 +2.2
Opus 4.8 ARC-AGI-2 77.5 80.0 +2.5
Opus 4.8 LongCoT-mini 65.3 66.5 +1.2

The strongest ProgramBench result is obtained with GPT-5.5: meta-reasoning reaches 71.5% mean per-problem test-pass rate, compared with 63.7% for matched direct control and 58.0% for Codex. With Opus 4.8, it reaches 67.2%, compared with 65.3% for direct control and 65.5% for Claude Code. It exceeds mini-SWE Agent for all three models by 2.5 to 13.9 points. These external comparisons establish useful end-to-end reference points, although the matched Direct Control comparison is the more informative test of the control mechanism because the external systems differ in prompts, implementation, and native orchestration.

Figure 2

Figure 2: Performance across four benchmarks and three models at the principal nominal budgets.

The paper’s central empirical claim is therefore specific rather than universal: under the tested configurations, explicit control produces consistent point-estimate improvements at sufficiently large budgets, even when the underlying workers and models are held fixed. The results do not establish statistical significance for every individual difference, particularly on the 30-problem proof benchmark, but the direction of the matched comparisons is uniform.

Scaling behavior and the cost of explicit control

The budget sweeps provide the most important evidence for the paper’s scaling argument. Meta-reasoning generally consumes more of the available allowance as the nominal budget increases, whereas direct control often stops early. On ProgramBench with GPT-5.5, direct control uses approximately 18% of the 1200-call allowance, while meta-reasoning uses roughly 89% to 101% across models. The slight excess above the nominal allowance can occur because workers dispatched before the cap are allowed to complete.

This additional computation is associated with improved performance. For GPT-5.5 on ProgramBench, increasing the allowance from 400 to 1200 calls raises meta-reasoning performance from 64.1% to 71.5%; direct control remains near 64% across the same sweep. With Opus 4.8, direct control increases its actual usage from 376 to 768 calls but improves only from 62.7% to 65.3%, peaking at the intermediate budget. Thus, the result is not merely that meta-reasoning spends more calls. Its control policy determines how those calls are allocated and connected.

The method has a clear low-budget disadvantage. With Opus 4.8 at 400 ProgramBench calls, meta-reasoning scores 56.6%, below direct control’s 62.7%; GPT-5.5 also trails direct control at the smallest ProgramBench allowance. The four-stage controller incurs a fixed and variable overhead that must be amortized over a sufficiently long run. This result qualifies the paper’s scaling claim: explicit meta-reasoning is not uniformly preferable, and the relevant regime is one in which the task provides enough computational horizon for improved control to repay its cost.

Artifact topology and reuse of prior work

The artifact graphs show that meta-reasoning changes not only the quantity of computation but also its structure. On the reasoning benchmarks, it generates more worker artifacts than direct control at the main budget, while recorded dependencies increase more rapidly. For example, on IMO ProofBench-Advanced with Gemini 3.1 Pro, worker artifacts approximately double under meta-reasoning, whereas recorded dependencies increase by roughly an order of magnitude.

The same pattern appears on ProgramBench. Raising GPT-5.5’s allowance from 400 to 1200 calls produces more than 2.5 times as many worker-artifact nodes and nearly four times as many dependency edges under meta-reasoning. Direct control does not exhibit comparable graph growth. The method therefore uses additional inference budget to construct increasingly interconnected computation rather than simply generating a larger collection of independent attempts.

Figure 3

Figure 3: Worker-artifact topology on ProgramBench as nominal budgets increase; meta-reasoning produces larger and more interconnected computation graphs.

The graphs distinguish exploration, verification, repair, and synthesis. A root represents work launched without prior artifact context; descendants represent follow-up work based on selected earlier outputs; fan-in indicates synthesis or joint analysis of multiple artifacts. On ARC-AGI-2, for example, a direct-control run may produce several independent attempts and a few one-hop follow-ups, whereas meta-reasoning explores more attempts, links later computation to earlier results, and reaches a frontier containing multiple candidates. This structure is consistent with a controller that explicitly decides when to branch and when to reuse.

Figure 4

Figure 4: Example artifact graphs comparing independent attempts and dependency-rich follow-up computation under the two control policies.

The interpretation should remain bounded by the graph definition. Edges represent only controller-selected artifact context. They do not capture all information channels, all within-worker interactions, or environmental state. In ProgramBench, for instance, a coding worker may inspect and modify repository state that is not fully represented by the selected artifact context. Consequently, the graph is a useful operational trace of orchestration, not a complete causal description of computation.

Coverage, monitoring, and final selection

The paper separates the appearance of correct work from the decision to submit it. If C\mathcal{C} denotes the event that a run contains a correct candidate and S\mathcal{S} the event that its submission is correct, then S⊆C\mathcal{S}\subseteq\mathcal{C} and final success can be written as Pr⁡(S)=Pr⁡(C)Pr⁡(S∣C)\Pr(\mathcal{S})=\Pr(\mathcal{C})\Pr(\mathcal{S}\mid\mathcal{C}). This decomposition provides a principled account of whether an agent is coverage-bound or selection-bound.

Meta-reasoning improves correct-candidate coverage in most evaluated settings. The largest reported increases occur with Gemini 3.1 Pro: 20 percentage points on IMO ProofBench-Advanced and 10 points on LongCoT-mini. ARC-AGI-2 improves across models, while Opus 4.8 on LongCoT-mini is approximately tied with direct control. These results imply that a substantial portion of the final performance improvement comes from generating correct work more often, not solely from selecting among a fixed set of candidates.

Monitoring quality also changes. Worker confidence can be nearly uninformative: Gemini 3.1 Pro’s worker-confidence Type-2 AUC on IMO ProofBench-Advanced is approximately 0.55, close to chance. The controller’s Assess-stage verdict reaches 0.88 on the same setting. A similar separation occurs on ARC-AGI-2. The result supports the claim that a dedicated controller can evaluate candidate quality more effectively than workers’ self-reported confidence in several settings. It does not establish calibration, however; Type-2 AUC measures ranking discrimination, not whether confidence values correspond to empirical probabilities.

Figure 5

Figure 5: Meta-reasoning improves correct-candidate coverage in most settings, while monitoring and final-selection gains vary by model and task.

Improved monitoring does not uniformly translate into improved final selection. On ARC-AGI-2 with GPT-5.5, meta-reasoning obtains an 11-percentage-point frontier selection gain relative to uniform choice from the run’s deepest terminal artifacts, compared with 3 points for direct control. It also improves the corresponding measure for ARC-AGI-2 with Opus 4.8 and LongCoT-mini with GPT-5.5. On LongCoT-mini with Opus 4.8, the methods are approximately tied, and with Gemini 3.1 Pro direct control is slightly better.

This nonuniformity is an important qualification. The controller can produce better assessment signals without consistently making better final commitments. In approximately 83% of ARC-AGI-2 and LongCoT-mini runs, the submitted artifact comes from the deepest terminal frontier, and roughly three-quarters of those runs contain several frontier candidates. Selection remains a difficult decision even after correct work has been found.

Persistent memory and compact controller state

The memory system performs two distinct functions. It preserves full worker artifacts for later retrieval and stores controller-authored notes that compress the evolving state of the run. The controller’s state is rewritten at each cycle, while artifacts retain stable identifiers and remain available through explicit reads.

The controller frequently revisits its own notes. Among entries retrieved at least once, controller notes receive between 4.4 and 9.8 reads on ARC-AGI-2 and LongCoT-mini, compared with 1.9 to 3.3 reads for worker artifacts. Reads and writes are concentrated in Assess and Propose, with writes outnumbering reads for every model. Assess receives new artifacts directly and therefore performs relatively few explicit reads of older work; Propose and Evaluate retrieve prior artifacts when generating and comparing candidate actions.

Figure 6

Figure 6: Memory operations concentrate in assessment and proposal, with repeated reads of selected controller-authored notes.

The strongest systems-level effect concerns the representation persisted between control decisions. On the reasoning benchmarks, Direct Control’s accumulated history exceeds one million characters on the longest Opus 4.8 IMO ProofBench-Advanced runs. The Meta-Reasoning Agent’s compact state remains on the order of thousands to tens of thousands of characters and often shrinks as the run progresses. On ProgramBench, the gap narrows to a factor of a few because the controller must retain increasingly detailed repository information, and its state grows over time.

Figure 7

Figure 7: The persisted meta-reasoning state remains far smaller than direct control’s accumulated message history, especially on long reasoning runs.

The implication is not simply reduced context length. Compact state changes the information-access policy: important details are removed from the immediate control representation but can be recovered through memory reads. This introduces a trade-off. Compression can prevent context saturation and reduce retrieval noise, but it is lossy if the controller fails to preserve a fact or fails to retrieve it when it later becomes relevant. The paper suggests that direct control’s plateau may partly reflect long-context degradation, but it does not directly isolate this mechanism through a state-length ablation.

Limitations and open questions

The principal causal limitation is that the matched comparison changes several components simultaneously: staged Assess–Propose–Evaluate–Dispatch control, compact state, persistent memory, and the associated retrieval policy. The experiments establish the value of the evaluated design as a whole, but cannot determine whether the gains arise primarily from state compression, explicit proposal generation, qualitative budget evaluation, controller memory, or their interaction. Stage-removal, state-only, memory-only, and matched-context ablations are needed for component-level attribution.

Resource matching is also incomplete. Equal model-call allowances do not imply equal token consumption, wall-clock time, or FLOPs. The systems choose to stop at different points, and the meta-reasoning controller may use substantially more controller tokens even when call counts are equal. The results therefore support a comparison under nominal call budgets, not an efficiency claim at matched monetary or latency cost.

The benchmark evidence is limited in breadth and statistical resolution. The proof benchmark contains only 30 problems, and the positive point estimates do not establish significance for each comparison. Generalization to weaker models, other task distributions, or substantially longer horizons remains open. External-agent results are configuration-specific rather than controlled architectural comparisons.

Finally, the paper reports a concrete failure mode on the chess subset of LongCoT-mini. Meta-reasoning sometimes engages in unproductive reconsideration and destabilizes an answer that was already correct. This demonstrates that additional verification is not intrinsically beneficial: a controller can over-deliberate, introduce regressions, or replace a valid solution with a less reliable revision. The paper leaves open how to detect when further checking has negative expected value and how to calibrate stopping against the risk of destabilizing correct work.

Conclusion

The paper presents agentic meta-reasoning as an inference-time control architecture in which deciding what to compute next becomes an explicit structured process. Across four benchmarks and three models, it improves all 12 matched point estimates at the principal budgets, achieves 71.5% on GPT-5.5 ProgramBench versus 63.7% for direct control, and continues to benefit from larger allowances in settings where direct control plateaus. Its runs contain more reused computation, more correct candidates in most settings, and substantially smaller persistent controller state than full-history control.

The evidence supports a narrower but important conclusion: for long-horizon agents, object-level competence is only one determinant of performance. The allocation, assessment, reuse, and termination of inference computation constitute a separate capability whose value becomes more apparent as the available run length increases.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper studies how AI agents can solve difficult, long-running tasks more effectively.

An AI agent is a LLM that does more than give one answer. It may work through many steps, ask other model “workers” to help, check its progress, use tools, and decide when to stop.

The authors introduce a method called agentic meta-reasoning. In simple terms, this means that the AI does not only think about the problem—it also thinks about how it is solving the problem.

For example, if an AI is trying to prove a difficult mathematics problem, it might ask itself:

  • Should I continue my current proof?
  • Should I check whether one step is correct?
  • Should I try a completely different approach?
  • Which previous attempts are worth reusing?
  • Have I already found a good answer, and should I stop?

The paper argues that these decisions should be handled by a special controller, rather than being mixed together with the ordinary problem-solving work.

2. What questions are the researchers asking?

The researchers mainly want to know whether separating “thinking about the task” from “thinking about how to solve the task” makes AI agents better.

Their key questions are:

  1. Does meta-reasoning improve performance? Can an agent solve more problems correctly if it has a separate controller?
  2. Does it help more on longer and harder tasks? The researchers want to see whether the method becomes especially useful when an agent has many steps and a large computing budget.
  3. How does the agent use its time and computing power? Does it create new attempts, check old ones, or combine several earlier ideas?
  4. Does the agent find correct answers more often? Sometimes an agent creates a correct answer but fails to recognize it.
  5. Can the agent choose the correct answer from several possibilities? Finding a good answer and selecting it as the final answer are not always the same thing.

3. How did the researchers test the idea?

The controller and the workers

The proposed system has two main parts:

  • Workers do the actual task, such as writing a proof or creating computer code.
  • A controller decides what the workers should do next.

Imagine a school project. The workers are students doing research, writing drafts, and checking facts. The controller is like a project leader who decides:

  • Which draft should be improved?
  • Which idea needs checking?
  • Should someone start a new approach?
  • Which information should each student receive?

The controller does not need to remember every word from the entire project. Instead, it keeps a short summary of what has happened and stores the full work in a kind of memory.

Four stages of control

Each decision cycle has four stages:

  1. Assess The controller reviews what has been learned so far and updates its summary.
  2. Propose It suggests several possible next steps.
  3. Evaluate It decides which option is most useful given the remaining time or computing budget.
  4. Dispatch It sends the chosen task, along with the relevant earlier work, to one or more workers.

The agent can also decide to stop and submit an answer.

Artifact memory and graphs

Every piece of work is saved as an artifact. An artifact could be:

  • A first attempt at a proof
  • A criticism of that proof
  • A revised solution
  • A controller’s note
  • A piece of computer code

The researchers record which artifacts were used to create later artifacts. This produces an artifact graph.

For example:

1
2
3
First proof attempt → Critique → Revised proof
                  ↘
                   Second check

This graph helps researchers see whether the AI is building on earlier work or simply starting many unrelated attempts.

Comparison systems

The main comparison was between:

  • Meta-Reasoning Agent: uses the separate controller and the four stages.
  • Direct Control Agent: uses the same workers and tools but chooses each next action in one step, based on its accumulated history.

This comparison is useful because the two systems have similar abilities. The main difference is how they control the work.

Tasks and budgets

The systems were tested on four types of problems:

Benchmark What it tests
IMO ProofBench-Advanced Writing difficult mathematical proofs
ARC-AGI-2 Discovering hidden rules in colored-grid puzzles
LongCoT-mini Long reasoning problems in areas such as science, chess, and mathematics
ProgramBench Rebuilding computer programs from descriptions and testing them

The researchers used three advanced LLMs: Gemini 3.1 Pro, GPT-5.5, and Opus 4.8.

They gave the agents limited numbers of model calls. A model call is like one turn in which the AI thinks or performs an action. The reasoning tasks allowed up to 100 calls, while the programming tasks allowed up to 1,200 calls.

4. What did the researchers find?

Meta-reasoning usually improved results

At the main budgets, the Meta-Reasoning Agent performed better than the Direct Control Agent in all 12 matched comparisons: four benchmarks tested with three models.

The average improvements over direct control were:

  • About 4.0 percentage points on the proof benchmark
  • About 4.2 percentage points on ARC-AGI-2
  • About 3.6 percentage points on LongCoT-mini
  • Clear improvements on ProgramBench as well

For example, on ProgramBench:

  • GPT-5.5 with Direct Control: 63.7%
  • GPT-5.5 with Meta-Reasoning: 71.5%

The meta-reasoning system also scored higher than some existing coding agents. With GPT-5.5, it achieved 71.5%, compared with 58.0% for Codex.

It benefited more from larger budgets

When the agents were allowed more model calls, meta-reasoning generally kept improving.

For example, on ProgramBench with GPT-5.5, increasing the budget from 400 to 1,200 calls raised performance from 64.1% to 71.5%.

The Direct Control Agent often stopped improving. With the larger budget, it sometimes used more calls but did not produce much better answers.

This suggests that simply giving an AI more time is not enough. The AI also needs a good plan for deciding how to spend that time.

It reused earlier work more often

The Meta-Reasoning Agent created more links between earlier and later work. Instead of treating every attempt as completely new, it was more likely to:

  • Check an earlier attempt
  • Repair a mistake
  • Combine several ideas
  • Ask a worker to improve a promising answer

This is similar to writing an essay. A weak approach is to write ten completely separate essays. A better approach may be to write a draft, find its problems, revise it, and then ask someone else to check the revision.

It found correct answers more often

The researchers separated two different abilities:

  1. Coverage: Did the agent create a correct answer at some point?
  2. Selection: Did the agent actually submit that correct answer?

Meta-reasoning improved coverage in most settings. This means that more runs contained at least one correct candidate answer.

However, finding a correct answer did not always mean that the agent submitted it.

Choosing the best answer was still difficult

Sometimes the agent had both a correct and an incorrect answer in its memory. It then had to decide which one to submit.

The controller’s judgments were often better than the workers’ own confidence ratings. For example, on one proof task, worker confidence was only slightly better than guessing, while the controller’s judgments were much more accurate at ranking good and bad answers.

However, the improvement in final answer selection was not consistent in every test. Meta-reasoning helped in some settings but was tied with or slightly worse than direct control in others.

The method can be expensive at small budgets

The controller itself uses model calls. These calls do not directly solve the problem; they are used to plan and evaluate the work.

At small budgets, this extra planning can take away too many calls from the workers. As a result, meta-reasoning sometimes performed worse than direct control when the budget was very limited.

This is an important trade-off:

  • With little time, extra planning may be a burden.
  • With lots of time, better planning may help the agent use its resources wisely.

5. Why are these results important?

The paper suggests that future AI systems should not only become better at producing individual answers. They should also become better at managing long problem-solving processes.

A powerful agent may need to:

  • Remember useful earlier work
  • Detect when an idea is failing
  • Decide whether to explore or improve
  • Spend more effort on uncertain parts
  • Recognize when it has already found the right answer
  • Avoid throwing away good work

This is especially important for tasks such as:

  • Writing large software systems
  • Proving difficult mathematics
  • Conducting scientific research
  • Planning complicated projects
  • Solving problems that require hundreds or thousands of steps

The artifact graphs also give researchers a way to study how an AI reached an answer, rather than looking only at whether the final answer was right or wrong.

Conclusion

The main message of the paper is that AI agents can improve when they spend some of their computing power thinking about what to do next.

The Meta-Reasoning Agent separates planning and self-checking from ordinary task-solving. It keeps a short summary of progress, stores detailed work in memory, and chooses new actions through a structured process.

The experiments show that this approach usually leads to better results, especially on long and complicated tasks with larger budgets. It helps agents reuse earlier work and find correct answers more often. However, it also costs extra time and computing power, so it may not help when the budget is very small.

In the future, AI systems may work less like a person answering one question at a time and more like a whole team: some parts doing the work, and another part organizing, checking, and deciding how the team should proceed.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • The paper does not isolate which component of the four-stage controller—Assess, Propose, Evaluate, or Dispatch—produces the observed gains; the Direct Control comparison changes staging, state compression, memory access, and deliberation simultaneously.
  • The contribution of compact state compression is not separately compared with alternatives such as full-history control, structured summaries, learned memory, or programmatic retrieval.
  • The paper does not establish whether explicit meta-reasoning is superior to a simpler controller with equivalent additional calls, such as extra worker samples, verifier calls, critique rounds, or self-consistency selection.
  • The qualitative Evaluate stage is not compared with a calibrated or learned value-of-computation estimator, leaving unclear whether the gains arise from meta-reasoning itself or from improved budget allocation.
  • The controller's proposal quality is not evaluated independently from its ability to select among proposals; failed runs are not attributed separately to missing useful actions versus poor evaluation of available actions.
  • The paper does not report ablations on controller-stage prompts, stage order, number of controller calls, or whether all four stages are necessary.
  • The controller is not trained or fine-tuned for orchestration, so it remains unclear whether the approach scales with specialized controller training, reinforcement learning, preference optimization, or task-specific orchestration policies.
  • The experiments use only three frontier models and do not establish whether the method benefits weaker models, open-weight models, multimodal models, or models with different instruction-following and self-evaluation capabilities.
  • The benchmarks are limited in number and scope, with especially small samples for IMO ProofBench-Advanced (30 problems) and ProgramBench (200 problems); the statistical significance and confidence intervals of the reported gains are not provided.
  • The paper reports point estimates but does not adequately quantify run-to-run variance, random-seed sensitivity, or uncertainty across repeated executions of the same problem.
  • The external-baseline comparisons are not fully matched: production coding agents have different native interfaces, prompts, stopping behavior, and tool-use procedures, so their results cannot cleanly identify the effect of meta-reasoning.
  • The nominal model-call budget is only a coarse resource measure; the systems are not compared using token consumption, FLOPs, monetary cost, wall-clock latency, energy use, or parallelism-adjusted compute.
  • Budget enforcement allows workers dispatched before exhaustion to finish, potentially causing overspending; the effect of this accounting rule on comparisons is not analyzed.
  • The paper does not determine the minimum budget at which meta-reasoning becomes cost-effective, nor whether its overhead can be adaptively reduced for easy problems.
  • The reported scaling trends cover only budgets up to 100 calls for reasoning tasks and 1,200 calls for ProgramBench; it remains unknown whether the gains persist, saturate, or reverse at substantially larger budgets.
  • The method's behavior under strict latency, concurrency, API-rate, or monetary constraints is not evaluated, despite the additional sequential controller stages likely increasing response time.
  • The artifact graph records only controller-selected context dependencies and therefore omits information transferred through environment state, files, tools, hidden model context, shared prompts, and unrecorded worker interactions.
  • Because the artifact graph is incomplete for coding agents, the reported topology differences may not accurately represent the full computation performed on ProgramBench.
  • The paper does not test whether alternative memory-retrieval strategies—embedding search, relevance ranking, recency policies, hierarchical summaries, or learned retrieval—improve on controller-selected artifact access.
  • The controller's self-authored notes are repeatedly reread, but the paper does not establish whether these notes contain reliable information, introduce confirmation bias, or amplify early mistakes.
  • The impact of erroneous or lossy Assess summaries is not quantified; a mistaken compressed state could systematically discard useful evidence from the persistent artifact memory.
  • The system lacks an explicit mechanism for correcting corrupted state or detecting when its own summary conflicts with the underlying artifacts.
  • Worker context selection is treated as a controller decision, but the paper does not measure how often relevant artifacts are omitted, irrelevant artifacts are included, or context selection causes failures.
  • The analysis of coverage and selection is limited to reasoning benchmarks; no equivalent intermediate-artifact grading protocol is provided for ProgramBench, where the largest practical gains are reported.
  • Intermediate solution correctness is not always comparable to final benchmark correctness: IMO ProofBench uses partial proof scores, whereas the coverage analysis binarizes artifacts, potentially discarding meaningful degrees of correctness.
  • The reliability of intermediate-artifact labels and per-domain verifiers is not examined, including grader error, evaluator-model bias, or disagreement between automated and human assessment.
  • Worker and controller confidence are evaluated using Type-2 AUC only; calibration, absolute confidence quality, selective prediction, and decision-theoretic usefulness are left unresolved.
  • The analysis does not determine whether controller verdicts improve because they access more evidence, because they receive better prompts, or because they are given a distinct role from workers.
  • The frontier-selection baseline is structurally defined and may not represent the strongest available selection strategy; comparisons against learned verifiers, majority voting, pairwise comparison, or independent judge ensembles are absent.
  • Selection gain can reflect choosing a shallower artifact rather than improving terminal-candidate discrimination, but these effects are not separated.
  • The paper does not investigate failure modes in which meta-reasoning overcommits to an initially plausible but incorrect approach, repeatedly reuses flawed artifacts, or wastes budget on increasingly elaborate repairs.
  • The experiments do not evaluate adversarial, deceptive, or distribution-shifted intermediate artifacts designed to mislead the controller.
  • The method's robustness to incorrect worker outputs, unreliable tools, execution failures, malformed artifacts, or inconsistent environmental state is not characterized.
  • ProgramBench evaluation uses hidden tests, but the paper does not analyze whether meta-reasoning improves generalization to substantially different program interfaces, languages, repositories, or execution environments.
  • The reasoning benchmarks do not establish transfer to interactive, open-ended, safety-critical, or real-world tasks where correctness is ambiguous and no deterministic verifier is available.
  • The paper does not compare meta-reasoning with human-designed hierarchical planning, beam search, Monte Carlo tree search, debate, verifier-guided search, or adaptive test-time scaling under equalized compute.
  • The controller's proposed actions are not analyzed for diversity; it remains unclear whether meta-reasoning expands the search over genuinely different approaches or mainly produces correlated variations of the same line of work.
  • The relationship between graph reuse and actual causal usefulness is unresolved: more edges and deeper graphs may indicate productive composition, redundant context accumulation, or merely more controller activity.
  • The paper does not test whether deliberately constraining graph depth, fan-in, branching, or memory reads can preserve performance while reducing cost and latency.
  • The method's stopping behavior is underexplored; there is no systematic analysis of premature stopping, unnecessary continuation, forced submission after budget exhaustion, or optimal stopping criteria.
  • The paper does not establish whether the controller can reliably recognize that no available artifact is correct and should initiate a genuinely fresh search rather than continue repairing existing work.
  • Reproducibility is limited by missing implementation details in the provided text, including complete prompts, model sampling parameters, retrieval policies, stage-specific call limits, exact baseline modifications, and handling of malformed or failed calls.
  • The paper excerpt ends before the complete memory/context results, appendices, and presumably detailed statistical analyses, leaving unresolved whether the reported qualitative conclusions are fully supported by the omitted evidence.

Practical Applications

Immediate Applications

The paper’s main practical contribution is an inference-time orchestration pattern: separate task execution from control of execution, retain intermediate artifacts in persistent memory, and use staged assessment, proposal, evaluation, and dispatch to decide what to compute next.

  • Long-horizon software engineering agents — software development
    • Integrate the framework into coding agents that inspect repositories, edit files, run tests, and revise implementations.
    • A controller could:
    • summarize the current implementation state;
    • identify unresolved failures or missing functionality;
    • propose alternatives such as debugging a failing test, inspecting a dependency, reverting to an earlier design, or starting a fresh implementation;
    • allocate additional calls to the highest-value option;
    • pass only relevant files, test logs, and prior patches to each coding worker.
    • This could produce tools for repository reconstruction, legacy-system migration, automated bug fixing, and large refactoring workflows.
    • Evidence: On ProgramBench, the meta-reasoning system achieved a higher test-pass rate than the matched direct-control agent and the evaluated coding-agent baselines, particularly at larger budgets.
    • Dependencies and assumptions: The repository must support reliable execution and testing; hidden-test performance must correlate with real-world correctness; controller calls must be affordable; and access controls are needed before agents can modify production systems.
  • Adaptive debugging and test-generation workflows — software quality assurance
    • Use artifact graphs to preserve relationships among source files, patches, test outputs, critiques, and failed approaches.
    • When a patch fails, the controller can decide whether to:
    • repair the current patch;
    • generate an independent solution;
    • request targeted test generation;
    • inspect an earlier artifact;
    • compare competing implementations.
    • A practical product could be an IDE or CI plugin that maintains a persistent “debugging graph” rather than a linear chat transcript.
    • Dependencies and assumptions: Tests must be sufficiently informative; artifact provenance must be captured accurately; and the system should avoid repeatedly exploring equivalent fixes.
  • Research-problem solving assistants — academia
    • Apply the architecture to mathematical proofs, formal derivations, literature-grounded analysis, and experimental planning.
    • Workers could independently propose proofs or hypotheses, while the controller tracks which lemmas, assumptions, experiments, or citations remain unverified.
    • The system could assign specialist workers to proof checking, counterexample search, statistical review, literature retrieval, or result synthesis.
    • The coverage/selection distinction is especially useful: a research assistant can separately report whether a valid candidate was found and whether it was selected as the final result.
    • Dependencies and assumptions: Intermediate outputs need domain-appropriate verifiers. Human experts remain necessary for claims that lack deterministic validation, and the system must preserve citations, provenance, and uncertainty rather than treating controller judgments as proof.
  • Automated verification and review pipelines — formal methods, compliance, and cybersecurity
    • Use explicit controller assessment to review competing artifacts such as proofs, security patches, threat analyses, configuration changes, or compliance reports.
    • The controller can prioritize targeted verification of uncertain components instead of rerunning an entire analysis.
    • Artifact graphs can provide an audit trail showing which evidence supported each conclusion and which workers generated or reviewed it.
    • Dependencies and assumptions: The review process requires trusted validators or independent checkers. Language-model confidence alone is insufficient, as the paper shows that worker confidence can be close to chance in some settings.
  • Cost-aware orchestration of model calls — AI infrastructure
    • Implement the four-stage control cycle as a general-purpose runtime for multi-agent systems.
    • The runtime can dynamically choose between independent sampling, critique, repair, synthesis, memory retrieval, and termination according to a remaining call budget.
    • This is applicable to customer-support automation, document processing, data analysis, and internal knowledge assistants where tasks vary substantially in difficulty.
    • Compact state can reduce the need to replay the full interaction history, potentially lowering context length and improving run manageability.
    • Dependencies and assumptions: The current evaluation uses model-call budgets rather than a complete cost model. Real deployments must account for token volume, latency, parallelism, API pricing, memory retrieval costs, and tool execution time.
  • Agent observability and failure diagnosis — MLOps and enterprise automation
    • Record each worker output as an artifact and each supplied context as a directed dependency edge.
    • Operators can then distinguish:
    • failures where no correct candidate was produced;
    • failures where a correct candidate existed but was not selected;
    • excessive branching or redundant work;
    • overreliance on a flawed artifact;
    • premature stopping.
    • A monitoring dashboard could expose graph depth, branching, reuse, budget utilization, candidate coverage, and selection quality.
    • Dependencies and assumptions: The artifact graph captures only controller-recorded dependencies and may omit information obtained through unlogged tools or environmental state. Production systems therefore need comprehensive event logging and privacy controls.
  • Adaptive educational tutoring — education
    • A tutoring agent could maintain artifacts representing a student’s attempted solution, hints, misconceptions, corrections, and alternative explanations.
    • Instead of always generating the next hint, the controller could decide whether to:
    • ask the student to retry;
    • provide a targeted explanation;
    • verify a specific step;
    • present a counterexample;
    • switch instructional strategies.
    • The compact state could preserve the student’s learning trajectory without repeatedly exposing the full transcript to every worker.
    • Dependencies and assumptions: Educational quality requires pedagogical evaluation, not only answer accuracy. The system must avoid revealing solutions too early, protect student data, and adapt to individual learning needs.
  • Multi-stage document and data analysis — legal, finance, and business operations
    • Use workers for extraction, independent analysis, critique, fact checking, and synthesis, with the controller selecting which stage to execute next.
    • For example, a financial-analysis agent could preserve candidate forecasts, assumptions, sensitivity analyses, and validation results, then direct additional work toward the most consequential uncertainty.
    • In legal workflows, the artifact graph could connect claims, source passages, counterarguments, and review notes.
    • Dependencies and assumptions: Domain-specific retrieval and verification are essential. The system must not infer that a candidate is reliable merely because it has undergone more agentic processing.
  • Daily-life planning assistants — personal productivity
    • Apply the framework to complex planning tasks such as travel, household projects, job searches, event planning, or budgeting.
    • The assistant could maintain alternatives, constraints, reservations, costs, and user preferences as artifacts, then decide whether to search for more options, verify a constraint, revise a plan, or stop.
    • Dependencies and assumptions: External information must be current; actions involving purchases, bookings, or messages require explicit user confirmation; and the system must handle changing constraints and sensitive personal data.

Long-Term Applications

The reported results suggest broader applications, but these require stronger validation, improved cost models, better monitoring, and integration with reliable external tools.

  • Autonomous scientific discovery systems — science and research
    • A future system could coordinate literature search, hypothesis generation, simulation, experiment design, statistical analysis, and replication.
    • The artifact graph would represent dependencies among hypotheses, datasets, experimental results, code, and interpretations.
    • The controller could spend additional computation on the uncertainty most likely to change the scientific conclusion rather than simply producing more independent hypotheses.
    • Dependencies: This requires trustworthy scientific simulators, laboratory interfaces, statistical safeguards, reproducibility protocols, and human approval for consequential experiments. The paper demonstrates improved reasoning benchmarks, not end-to-end scientific discovery.
  • Formal theorem-proving and mathematical research agents — mathematics and verification
    • Workers could generate proofs, search for counterexamples, invoke proof assistants, repair failed lemmas, and translate informal arguments into formal code.
    • The controller could prioritize unresolved bottlenecks and select a proof artifact only after formal verification.
    • Dependencies: Integration with proof assistants and dependable proof checkers is necessary. Qualitative controller judgments should guide search, but formal verification must determine correctness.
  • Robotic task planning and recovery — robotics
    • In embodied agents, workers could propose action sequences, inspect sensor data, simulate alternatives, diagnose failures, or recover from unexpected states.
    • Persistent artifacts could include maps, trajectory plans, object detections, failed actions, and safety constraints.
    • The controller could choose between continuing a plan, replanning, gathering another observation, or returning to a safe state.
    • Dependencies: Real-time latency, partial observability, actuator reliability, simulation-to-reality transfer, and safety certification are major barriers. The paper’s model-call experiments do not establish suitability for physical-time control.
  • Healthcare decision-support and clinical workflow orchestration — healthcare
    • A system could coordinate medical-record summarization, guideline retrieval, differential-diagnosis generation, evidence checking, and patient-specific risk analysis.
    • Its artifact graph could make the provenance of recommendations explicit and identify whether a decision was limited by missing evidence or by poor selection among available options.
    • Dependencies: Clinical validation, privacy protection, calibrated uncertainty, bias assessment, regulatory approval, and clinician oversight are mandatory. The system should support—not replace—licensed clinical judgment.
  • Energy-grid and industrial operations planning — energy and manufacturing
    • Controllers could coordinate forecasts, simulations, maintenance plans, fault diagnoses, and contingency strategies.
    • Additional computation could be directed toward the constraint or scenario with the greatest operational value, such as demand uncertainty or equipment failure.
    • Artifact graphs could provide traceability for operational recommendations.
    • Dependencies: Real-time systems require deterministic latency, validated simulators, secure interfaces, and fail-safe behavior. The qualitative value assessment described in the paper would need to be replaced or supplemented with calibrated operational objectives.
  • Financial research, portfolio analysis, and risk management — finance
    • Workers could generate investment theses, stress tests, scenario analyses, and independent risk reviews, while a controller allocates effort toward unresolved assumptions.
    • Persistent artifacts would support auditability of forecasts and enable separation of “no viable analysis was found” from “a viable analysis was found but not selected.”
    • Dependencies: Market nonstationarity, data quality, regulatory requirements, model risk, and adversarial incentives make uncalibrated self-assessment dangerous. Deployment would require strict human approval and independent quantitative validation.
  • Self-optimizing agent runtimes — AI systems research
    • Artifact-graph data could support learning better controllers, estimating the value of computation, and automatically adapting worker roles, context selection, and stopping policies.
    • Future runtimes might learn when to use deep staged control, when to use a cheaper direct strategy, and how to predict whether a run is coverage-bound or selection-bound.
    • Dependencies: The paper’s controller uses prompted qualitative evaluation rather than a learned value-of-computation model. Large, representative execution traces and reliable outcome labels would be needed to train and validate such systems.
  • General-purpose autonomous organizations and workflow systems — enterprise automation
    • Multiple specialized agents could maintain shared artifact memory for product development, procurement, incident response, policy analysis, or large-scale operations.
    • A supervisory controller could allocate work across teams of agents, preserve decisions and evidence, and escalate uncertain or high-impact actions to humans.
    • Dependencies: This requires robust identity and permission management, conflict resolution, shared data schemas, accountability mechanisms, and safeguards against error propagation through the artifact graph.
  • Personalized lifelong AI assistants — daily life
    • A long-term assistant could maintain structured memories of projects, preferences, commitments, and prior decisions while selectively retrieving only relevant artifacts.
    • It could distinguish between planning, verification, execution, and reflection rather than treating every interaction as a linear conversation.
    • Dependencies: Long-term memory raises substantial privacy, consent, deletion, security, and ownership issues. The assistant must also avoid reinforcing outdated or incorrect notes through repeated retrieval.
  • Policy simulation and public-sector decision support — government and public policy
    • Workers could model policy alternatives, summarize stakeholder evidence, identify implementation risks, and test outcomes under different assumptions.
    • The controller could allocate computation to contested assumptions and preserve an auditable graph connecting recommendations to evidence.
    • Dependencies: Policy decisions involve normative judgments that cannot be reduced to model confidence or benchmark accuracy. Public-sector use would require transparency, democratic accountability, independent review, and explicit separation between factual analysis and value judgments.

Overall, the most deployable near-term use is budget-aware orchestration of long-running software and reasoning agents with persistent artifact memory and execution tracing. The strongest long-term opportunities involve domains where intermediate artifacts can be independently verified; domains without reliable verification should treat the controller’s assessments as prioritization signals rather than authoritative judgments.

Glossary

  • Abstract reasoning: Problem solving that requires manipulation of concepts or patterns rather than direct procedural execution. “ARC-AGI-2, spanning proof search, abstract reasoning, long-horizon reasoning, and program reconstruction”
  • Agentic harness: A runtime system that coordinates model calls, tools, memory, and task execution. “Agentic harnesses instead use the model's own judgments to direct next computations”
  • Agentic inference: Inference in which an agent adaptively performs actions and computations toward a goal. “We call this approach agentic meta-reasoning”
  • Artifact graph: A directed graph representing artifacts and the dependencies between them. “We record the resulting work as an artifact graph.”
  • Artifact memory: Persistent storage containing intermediate outputs and notes produced during an agent run. “Artifact memory holds worker outputs and the controller's own notes side by side.”
  • AUC: Area under the receiver operating characteristic curve, a measure of ranking or classification quality. “We call this Type-2 AUC because the system is judging its own answers”
  • Budget utilization: The fraction of an available computational allowance that an agent actually consumes. “Budget utilization measures the fraction of the allowance actually used.”
  • Candidate solution: An intermediate artifact that may serve as a possible final answer. “Let $V_{\mathrm{sol}$ be the artifacts that can be graded as candidate solutions.”
  • Chain of thought: A model-generated sequence of intermediate reasoning steps. “In a single model call, the object of that control is the chain of thought.”
  • Convergence frontier: The deepest terminal artifacts in an artifact graph. “We call the deepest terminal artifacts the convergence frontier.”
  • Coverage: The probability that a run produces at least one correct candidate solution. “Coverage counts a run as successful once it contains a correct solution”
  • Direct Control Agent: A comparison agent that chooses actions directly from its accumulated history without explicit staged control. “We also introduce a Direct Control Agent, a matched variant with no explicit separation of control”
  • Directed acyclic graph: A directed graph containing no cycles, so dependencies proceed in one direction. “Because the input artifacts already exist when the worker starts, these edges form a directed acyclic graph.”
  • Dispatch: The process of converting a selected proposal into an executable worker action. “Dispatch turns the selected proposal into an executable action.”
  • Epistemic action: An action performed to organize, acquire, or modify an agent’s knowledge. “These reads and writes are epistemic actions”
  • Fan-in: The number of incoming edges to a node in a graph. “Fan-in counts a node's incoming edges”
  • FLOPs: Floating-point operations, a common measure of computational work. “Note that equal call allowances do not imply equal token cost, latency, or FLOPs.”
  • Frontier selection gain: Improvement over randomly selecting a candidate from the deepest terminal artifacts. “Positive gain means the agent does better than choosing uniformly from its frontier.”
  • Inference-time computation: Computation performed while generating an answer rather than during model training. “These established inference-time computation as a major scaling lever”
  • Long-horizon reasoning: Reasoning over extended sequences of interdependent steps. “LongCoT-mini, which tests long-horizon agentic capability”
  • Meta-cognitive control: Monitoring and regulating one’s own reasoning or computational process. “This is a problem of metacognitive control”
  • Meta-reasoning: Reasoning about and acting on an agent’s own inference process. “We call this approach agentic meta-reasoning”
  • Monitoring: Assessing the quality or correctness of intermediate work during execution. “We also ask the controller for its own judgment”
  • Nominal budget: The stated maximum computational allowance assigned to a run. “The reasoning benchmarks use nominal budgets of 25, 50, and 100 model calls per problem.”
  • Object-level computation: Computation directly aimed at solving the target task rather than managing the solving process. “separate from the object-level computation it directs”
  • Online artifact-graph construction: Building a dependency graph incrementally as an agent performs a task. “This motivates our view of agentic inference as online artifact-graph construction”
  • Orchestration: Coordination of multiple models, agents, tools, or computational steps. “A complementary line of work learns, searches, or optimizes the system that coordinates model calls.”
  • Persistent memory: Storage that remains available across multiple decisions or execution cycles. “Full worker outputs remain in persistent memory”
  • Program reconstruction: Recreating a program’s behavior from documentation and an executable reference. “the solver must construct a self-contained codebase whose compiled executable matches the reference”
  • Proof search: The process of exploring possible mathematical arguments to find a valid proof. “spanning proof search, abstract reasoning, long-horizon reasoning, and program reconstruction”
  • Recursive LLM: A model framework that uses code to invoke sub-models recursively over parts of a task. “The Recursive LLM (RLM) holds the context in a variable the model edits with code”
  • Selection: Choosing a correct candidate after one has been generated. “Success therefore factors into finding a correct answer and choosing it once one exists”
  • State compression: Representing a large execution history using a smaller summary. “The comparison thus isolates the combined control design, but it does not separately identify the effects of staging, state compression, or memory access.”
  • Stochastic output: An output influenced by randomness and therefore not fully determined by the same inputs. “A coding worker may read files not represented in CC, modify the environment, and produce stochastic outputs.”
  • Test-time scaling: Improving performance by allocating more computation during inference or execution. “Whether such automation improves on simpler test-time scaling is under active investigation”
  • Type-2 AUC: A ranking measure for how well a system evaluates the correctness of its own outputs. “We call this Type-2 AUC because the system is judging its own answers”
  • Value of computation: The expected usefulness of spending additional computational resources on a particular action. “This is a prompted, qualitative assessment of computational value rather than an exact optimization”
  • Worker dispatch: Assignment of a task, instructions, and selected context to a worker agent. “Let ata_t be the resulting worker-dispatch or stopping action.”

Tweets

Sign up for free to view the 13 tweets with 311 likes about this paper.