---
title: Agent-Native Mid-Training Paradigm
url: https://www.emergentmind.com/topics/agent-native-mid-training-paradigm
type: topic
---

# Agent-Native Mid-Training Paradigm

Agent-Native Mid-Training Paradigm refers to a spectrum of training algorithms and data regimes for instilling agentic behaviors—such as planning, tool-use, reflection, or multi-agent coordination—directly within model or agent parameters, at an intermediate stage between pretraining and downstream post-training. Unlike pipeline-based systems, where agentic capabilities are modularized and orchestrated by external scripts or policies, agent-native mid-training emphasizes the internalization and direct optimization of agent workflows, feedback loops, and interactive environment signals, either within fixed model weights or via structured adaptive context. This paradigm encompasses a wide array of methodologies: trajectory-centric next-token modeling on agentic data, in-place reinforcement learning from agent execution, meta tool learning and self-reflection without weight updates, multi-agent mid-training with interaction-centric objectives, and function-library adaptation with model weights frozen. Across domains—software engineering, STEM, multimodal reasoning, tool-based research, and multi-agent environments—agent-native mid-training marks a methodological pivot towards agents that learn, adapt, and specialize through authentic, dynamic experience rather than static demonstrations or scripted scaffolds.

## 1. Defining Agent-Native Mid-Training: Foundations and Motivation

Agent-native mid-training (ANMT) is the intermediate training regime that bridges large-scale unsupervised pretraining and data-intensive supervised or reinforcement learning (RL) post-training. Its key distinction is the use of agent-native data: supervision that encodes the full action-observation-feedback sequences native to autonomous agents operating in authentic environments. In formal terms, given a policy $\pi_\theta$, observations $\mathrm{Obs}$, and an agent-native trajectory corpus $\mathcal{D}_{\text{MT}}$, the objective is:

\[
\mathcal{L}_{\rm MT}(\theta) = -\sum_{(x,y) \in \mathcal{D}_{\rm MT}} \sum_{t=1}^{|y|} \log p_\theta(y_t|x_{<t})
\]

while ensuring that each sample $(x,y)$ encodes multi-step sequences $\{(a_t, o_t)\}_{t=1}^T$ as encountered by a deployed agent [2601.18418].

ANMT is motivated by the observed inefficiencies and domain gaps of both static pretraining—which lacks dynamic, interactive state-action feedback—and pure RL, which is computationally prohibitive and constrained by the base model’s representational limits. By leveraging supervised learning on authentic agentic rollouts at scale, ANMT offers a scalable, data-efficient, and task-general paradigm for injecting agentic priors and behaviors [2601.18418, 2512.24618].

## 2. Categories of Agent-Native Data and Trajectories

Across ANMT implementations, the foundation lies in the careful construction or synthesis of agent-native data, which falls broadly into the following categories:

| Type                        | Description                                                       | Example Domains                  |
|-----------------------------|-------------------------------------------------------------------|----------------------------------|
| Contextually-native         | Full action-observation history, navigation context, edit traces  | Code PR workflows [2601.18418]   |
| Environmentally-native      | Real environment interactive feedback (tool output, errors)       | Executable code edits [2601.18418]|
| Agentic–Chain-of-Thought    | Structured multi-phase (plan, act, reflect) generator trajectories| STEM, math, research [2512.24618]|
| Multi-agent interactive     | Self-play or multi-agent dialogs; role-specific decision logs     | Coordination, theory-of-mind [2512.08743]|
| Task-oriented machine tokens| LLM-learned embedding-based message trajectories                  | Agent communication [2507.21454] |
| Self-Reflection & Meta-Tool | Episodic reflection, tool-use logs, context augmentation          | Knowledge agents [2508.00271]    |
| Offline function editing    | Incremental library synthesis/edit traces with failure feedback   | Symbolic reasoning, QA [2402.11359]|

For example, daVinci-Dev synthesizes both contextually-native (from pull requests and repo navigation) and environmentally-native (from in-Docker agent execution and tool feedback) code trajectories, enabling large-scale exposure to authentic agentic feedback loops [2601.18418]. In Youtu-LLM, agentic mid-training leverages over 200B tokens of structured math, code, research, and tool-use trajectories, each annotated with explicit phase tags (<Analysis>, <Plan>, <Action>, etc.), branching at critical action or failure points [2512.24618].

## 3. Training Algorithms and Optimization Objectives

The ANMT paradigm encompasses varying learning modalities and objectives, including:

- **Trajectory-centric next-token modeling:** Standard cross-entropy loss on assistant- or agent-output tokens in agentic trajectories, often with segment masking to focus learning on agent behaviors [2512.24618, 2601.18418].
- **Reinforcement learning from native execution (Agent Lightning):** Treats agent executions as Markov Decision Processes (MDPs), logging tuples $(s_t, a_t, r_t)$ at runtime and applying hierarchical RL (e.g., GRPO) on transitions without altering agent orchestration logic [2508.03680].

\[
\mathcal{L}(\theta) = -\,\mathbb{E}_{x\sim\mathcal{X}}\,\mathbb{E}_{(s_t, a_t)\sim\tau}
\left[\,\sum_{j=1}^{N_t}\log\pi_\theta(y_{t,j}\mid s_t,y_{t,<j})\,(G_t - b_x) \right]
\]

- **Meta tool learning and context synthesis (MetaAgent):** Incorporates self-reflection and verified reflection into dynamic context banks and in-house knowledge bases to shape agent behavior without parameter updates [2508.00271].
- **Early experience and self-reflection (Agent Learning via Early Experience):** Generates agent rollouts from current policies, applying auxiliary objectives for implicit world modeling and natural language rationale generation to improve policy grounding and reasoning [2510.08558].
- **Function library optimization with frozen models:** Treats function sets $F$ as “agent parameters”, using an LLM-based optimizer to incrementally edit $F$ with roll-back and early-stop mechanisms, while the core model weights remain unchanged [2402.11359].
- **Multi-agent losses (native multi-agent mid-training):** Combines understanding (theory-of-mind), joint planning, communication efficiency, and adaptation, each with specific loss functions, interleaved within minibatches according to a multi-task mixture [2512.08743].

## 4. Architectural and System Considerations

Agent-native mid-training unifies models, data, and environment interfaces:

- **Execution-agent/trainer decoupling (Agent Lightning):** Direct observability frameworks allow off-policy RL algorithms to consume agent-generated transitions with negligible code overhead; native agent workflows remain untouched [2508.03680].
- **Multi-stage curricula:** Progressive exposure to agentic trajectories is embedded as the final phase in multi-stage pretraining (e.g., commonsense → STEM → agentic) to facilitate internalization of planning and reflection [2512.24618].
- **Long-context support and specialized vocabularies:** Architectures such as Multi-Latent Attention (MLA) and STEM-optimized vocabularies are deployed to handle long-horizon agentic sequences and reduce compression overhead [2512.24618].
- **Unified communication and representation (machine language tokens):** Agents are trained to generate and interpret specialized token embeddings as an efficient channel for agent-to-agent or multi-modal communication, jointly optimized for task loss and embedding robustness [2507.21454].

## 5. Empirical Evidence and Scaling Laws

Experimental studies across benchmarks and domains validate the efficacy of ANMT:

- On SWE-Bench Verified, daVinci-Dev Qwen2.5-72B with full agent-native mid-training achieves 58.5% resolution rate, outperforming the agentless Kimi-Dev MT (48.6%) with less than half the training tokens [2601.18418].
- Youtu-LLM’s agentic mid-training yields a +14.4% average relative improvement across six downstream agentic benchmarks, with ~42.7% relative lift in code Pass@1 at k = 1 [2512.24618].
- Agent Lightning demonstrates stable performance improvement when deploying RL training in situ across text-to-SQL, retrieval-augmented generation, and math tool-use agents, with no agent code modification [2508.03680].
- MetaAgent’s meta tool learning elevates a frozen LLM from novice to expert-level tool reasoning without any weight updates, with ablations showing up to ~8-point drop in EM if the in-house knowledge base is omitted [2508.00271].
- Agent Learning via Early Experience records +9.6 points in in-domain and +9.4 in out-of-domain generalization over imitation-only baselines, with further gains in RL-ready settings [2510.08558].
- Scaling analysis in Youtu-LLM indicates logarithmic agentic performance scaling with mid-training token size, and unsaturated learning curves with increasing scale [2512.24618, 2601.18418].

## 6. Limitations, Challenges, and Future Research Directions

Challenges and open areas in agent-native mid-training include:

- **Data authenticity and coverage:** Ensuring the breadth (contextual diversity) and depth (authentic feedback) of agentic trajectories remains a bottleneck for non-code domains and underexplored environments [2601.18418].
- **Privacy and reproducibility:** Persisting developer identifiers, reliance on patched test harnesses, and single-model evaluations restrict generalizability and reproducibility in code domains [2601.18418].
- **Multi-agent scaling:** Pure single-agent scaling does not spontaneously yield robust multi-agent intelligence, as shown by plateaus on ToMBench and CoordinationQA without targeted multi-agent mid-training [2512.08743].
- **Sample efficiency and credit assignment:** Advanced techniques for hierarchical credit assignment and value learning promise further improvement in RL-based agent optimization, especially in distributed or tool-rich settings [2508.03680].
- **Beyond RL and LLMs:** Adaptive prompt optimization and modalities outside text (audio, vision, tactile) are conceptual extensions of the paradigm [2507.21454, 2512.24618].
- **Curricular and architectural synergy:** The synergy of contextually- and environmentally-native data, progressive curricula, and scalable architectures underlies robust agentic skill acquisition [2512.24618, 2601.18418].

A plausible implication is that as mid-training scales in both trajectory diversity and size, the marginal gains in agentic behavior are maintained, providing a pathway to increasingly general, robust, and intrinsically agentic models.

## 7. Comparative Table: Key Agent-Native Mid-Training Paradigms

| Approach     | Core Mechanism                         | Domain         | Weight Update | Empirical Gains                                                           |
|--------------|----------------------------------------|---------------|--------------|---------------------------------------------------------------------------|
| daVinci-Dev  | Trajectory-centric next-token modeling | Code          | Yes          | Pass@1 up to 58.5% on SWE-Bench Verified [2601.18418]                     |
| Youtu-LLM    | Scalable agentic data curriculum       | General/STEM  | Yes          | +14.4% avg. lift, log-linear scaling with data size [2512.24618]          |
| MetaAgent    | Meta tool learning, context enrichment | Web research  | No           | EM up to 52.1%; ablation drops up to 8 pts on EM [2508.00271]             |
| Agent Lightning | Agent execution → RL on MDP         | Any           | Yes          | Continuous improvement in SQL, RAG, math tool-use [2508.03680]            |
| Early Experience | Self-reflective rollout modeling   | General       | Yes          | +9.6 points in-domain, +9.4 OOD, improved post-RL ceiling [2510.08558]    |
| Function Editing | LLM-driven library optimization    | Symbolic/math | No           | +3–11% on MATH, TabMWP, GAIA [2402.11359]                                 |
| Machine Language Tokens | Embedding-based comm.       | Multi-modal   | Yes          | Compression ratio ≈0.01, <5% accuracy loss at low SNR [2507.21454]         |
| Native Multi-Agent | Multi-agent loss interleaving    | Multi-agent   | Yes          | Blueprint only; required to surpass coordination accuracy plateaus [2512.08743] |

Source: https://www.emergentmind.com/topics/agent-native-mid-training-paradigm