Miles v0.1: Production-Level Post-Training
Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
่ฎบๆๆฆ่ฟฐ
่ฟ็ฏ่ฎบๆไป็ปไบ Miles v0.1๏ผ่ฟๆฏไธไธชๅธฎๅฉ็ ็ฉถไบบๅ่ฎญ็ปๅคงๅไบบๅทฅๆบ่ฝๆจกๅ็็ณป็ปใ
็ฎๅๆฅ่ฏด๏ผMiles ็็ฎๆ ๆฏ่ฎฉๅคงๅ่ฏญ่จๆจกๅๅญฆไผๆดๅฅฝๅฐๅฎๆไปปๅก๏ผๅฐคๅ ถๆฏ้ฃไบ้่ฆๅคๆญฅ่กๅจ็ไปปๅกใไพๅฆ๏ผๆจกๅๅฏ่ฝ้่ฆ๏ผ
- ้ ่ฏปไธไธช็ผ็จ้ฎ้ข๏ผ
- ๆๅผ็ป็ซฏ๏ผ
- ็ผๅๆไฟฎๆนไปฃ็ ๏ผ
- ่ฟ่กๆต่ฏ๏ผ
- ๆ นๆฎ็ปๆ็ปง็ปญไฟฎๆนไปฃ็ ๏ผ
- ๆๅๅพๅฐไธไธชๅๆฐใ
่ฟ็ฑป่ฎญ็ปๆฏๆฎ้็โ่พๅ ฅ้ฎ้ขโ่พๅบ็ญๆกโๆดๅคๆ๏ผๅ ไธบๆจกๅ้่ฆไธๆญ่กๅจใไฝฟ็จๅทฅๅ ท๏ผๅนถไป็ฏๅขไธญ่ทๅพๅ้ฆใ
็ ็ฉถ็ฎๆ
่ฟ็ฏ่ฎบๆไธป่ฆๆณๅ็ญไปฅไธ้ฎ้ข๏ผ
- ๆๆ ทๆดๅฟซๅฐ่ฎญ็ปๅคงๅ่ฏญ่จๆจกๅ๏ผ
- ๆๆ ท่ฎฉ่ฎญ็ป่ฟ็จๆดๅ ๅ็กฎ๏ผ
- ๆๆ ท้ฟๅ ๆจกๅๅจ่ฎญ็ปๆถไฝฟ็จไบไธๅฎ้ ็ๆ่ฟ็จไธๅ็ๆฐๆฎ๏ผ
- ๆๆ ท่ฎฉ่ฎญ็ป็ณป็ปๆฏๆๆดๅคง็ๆจกๅๅๆดๅค็ GPU๏ผ
- ๆๆ ท่ฎฉ็ณป็ป้ๅบไธๅ็ๆจกๅใ็กฌไปถๅ่ฎญ็ปๆนๆณ๏ผ
ไฝ่ ็นๅซๅ ณๆณจไธไธช้ฎ้ข๏ผๆจกๅ็ๆ็ญๆก็็ณป็ปๅ็ๆญฃ่ฟ่ก่ฎญ็ป็็ณป็ป๏ผๅฏ่ฝไผ็จ็ฅๆไธๅ็ๆนๅผ่ฎก็ฎ็ปๆใๅณไฝฟๅทฎๅซๅพๅฐ๏ผ็ป่ฟๅพๅค่ฝฎ่ฎญ็ปๅ๏ผไนๅฏ่ฝๅฏผ่ดๆจกๅ่กจ็ฐๅๅทฎ๏ผ็่ณ่ฎญ็ปๅคฑ่ดฅใ
Miles ๆฏๆๆ ทๅทฅไฝ็๏ผ
Miles ็่ฎญ็ป่ฟ็จๅฏไปฅๆณ่ฑกๆไธไธชไธๆญๅพช็ฏ็ๅญฆไน ๆธธๆใๅฎไธป่ฆๅไธบไธไธช้ถๆฎตใ
1. ็ๆไปปๅกๅฐ่ฏ
้ฆๅ ๏ผๆจกๅๅฐ่ฏๅฎๆ่ฎธๅคไปปๅกใ่ฟไบๅฐ่ฏๅซไฝ ่ฝจ่ฟนใ
ๅฏนไบๆฎ้้ฎ้ข๏ผไธๆก่ฝจ่ฟนๅฏ่ฝๅชๆฏๆจกๅๅๅบไธๆฎต็ญๆกใๅฏนไบ็ผ็จไปปๅก๏ผไธๆก่ฝจ่ฟนๅฏ่ฝๅ ๅซ๏ผ
- ๆจกๅ่ฏดไบไปไน๏ผ
- ๆจกๅ่ฐ็จไบๅชไบๅทฅๅ ท๏ผ
- ๅทฅๅ ท่ฟๅไบไปไน๏ผ
- ๆจกๅๆฅไธๆฅ้ๅไบไปไน่กๅจ๏ผ
- ๆๅไปปๅกๅพๅฐไบๅคๅฐๅใ
ๆจกๅ้ๅธธไผๅฏนๅไธไธช้ฎ้ขๅฐ่ฏๅคๆฌกใๆฅ่ชๅไธไธช้ฎ้ข็ๅคๆกๅฐ่ฏ่ขซ็งฐไธบไธไธช ่ฝจ่ฟน็ปใๆๅฎไปฌๆพๅจไธ่ตทๅพ้่ฆ๏ผๅ ไธบ็ณป็ปๅฏไปฅๆฏ่พ่ฟไบๅฐ่ฏ๏ผๅคๆญๅชไบๅๅพๆดๅฅฝใ
Miles ไฝฟ็จไธไธชๅซไฝ SGLang ็็ณป็ปๆฅๅฟซ้็ๆ่ฟไบ่ฝจ่ฟนใๅฎ่ฟไผๅฐฝ้่ฎฉๅไธไธชๅคๆญฅ้ชคไปปๅกไธ็ด็ฑๅไธไธช GPU ๆๅกใ่ฟๆ ท๏ผไนๅ็ๅฏน่ฏๅ ๅฎนๅฏไปฅไฟๅญๅจ้ซ้็ผๅญไธญ๏ผไธๅฟ ๆฏๆฌก้ฝ้ๆฐๅค็ใ
่ฟๅฐฑๅไธไธชๅญฆ็ไธ็ดไฝฟ็จๅไธๆฌๆๅผ็็ฌ่ฎฐ๏ผ่ไธๆฏๆฏๆฌกๅ้ข้ฝ้ๆฐๆๅๅ้ข็ๅ ๅฎนใ
2. ๆ นๆฎ็ปๆ่ฎญ็ปๆจกๅ
ๆฅไธๆฅ๏ผ่ฎญ็ปๅจไผๆฅ็ๆจกๅ็ๅฐ่ฏๅๅพๅฐ็ๅๆฐ๏ผๅนถๆดๆฐๆจกๅ็ๅๆฐใ
ๆจกๅๅๆฐๅฏไปฅ็่งฃไธบๆจกๅๅ ้จ็ๅคง้โ่ฎฐๅฟๆ้ฎโใ่ฎญ็ป็่ฟ็จๅฐฑๆฏๆ นๆฎ็ปๆ่ฐๆด่ฟไบๆ้ฎ๏ผ
- ๅๅพๅฅฝ็่กไธบๅบ่ฏฅๆดๅฎนๆๅๆฌกๅบ็ฐ๏ผ
- ๅๅพไธๅฅฝ็่กไธบๅบ่ฏฅๅๅฐใ
Miles ๆฏๆไธค็งไธป่ฆ็่ฎญ็ปๅทฅๅ ท๏ผ
- Megatron-LM
- PyTorch FSDP
่ฟไบๅทฅๅ ท่ฝๅคๆ้ๅธธๅคง็ๆจกๅๅๆฃๅฐ่ฎธๅค GPU ไธๅ ฑๅ่ฎญ็ปใ
3. ๆๆฐๆจกๅ้ๅ็ๆ็ณป็ป
่ฎญ็ปๅฎๆๅ๏ผMiles ไผๆๆดๆฐๅ็ๆจกๅๅๆฐไผ ๅ็ๆ็ณป็ปใ่ฟๆ ท๏ผๆจกๅไธไธๆฌกๅฐ่ฏไปปๅกๆถ๏ผๅฐฑไผไฝฟ็จๅๅๅญฆๅฐ็ๆฐ็ฅ่ฏใ
่ฟไธไธช้ถๆฎตไผไธๆญ้ๅค๏ผ
็ๆๅฐ่ฏ โ ๆ นๆฎ็ปๆ่ฎญ็ป โ ๆดๆฐๆจกๅ โ ๅ็ๆๅฐ่ฏ
ๅๆญฅ่ฎญ็ปๅๅผๆญฅ่ฎญ็ป
ไผ ็ป่ฎญ็ป้ๅธธๅๆ้ไธๆ ท่ฟ่ก๏ผ
- ็ญๆๆๆจกๅๅฐ่ฏๅฎๆ๏ผ
- ่ฎญ็ปๆจกๅ๏ผ
- ็ญ่ฎญ็ป็ปๆ๏ผ
- ๅๅผๅงไธไธ่ฝฎๅฐ่ฏใ
่ฟ็งๆนๆณ็้ฎ้ขๆฏ๏ผ้ๅบฆ่พๆ ข็ไปปๅกไผๆไฝๆๆไบบใๅฐฑๅไธ็พคๅญฆ็ไธ่ตทๅฎๆไฝไธ๏ผๅฟ ้กป็ญๆๅไธไธชๅญฆ็ไบคไฝไธๅ๏ผ่ๅธๆ่ฝๆนๆนใ
Miles ๆฏๆ ๅฎๅ จๅผๆญฅ่ฎญ็ปใๅจ่ฟ็งๆนๅผไธ๏ผ
- ไธไบ GPU ๆญฃๅจ็ๆๆฐ็ไปปๅกๅฐ่ฏ๏ผ
- ๅฆไธไบ GPU ๅๆถ่ฎญ็ปๆจกๅ๏ผ
- ็ๆๅ่ฎญ็ปๅฏไปฅๅๆถ่ฟ่กใ
่ฟๆ ทๅฏไปฅๅๅฐ GPU ็ญๅพ ็ๆถ้ด๏ผๆ้ซๆดไฝ้ๅบฆใ
Miles ่ฟ่ฎพ็ฝฎไบไธไธช็ฑปไผผโ็ญๅพ ๅบโ็ ๆฐๆฎ็ผๅฒๅบใๅทฒ็ปๅฎๆ็่ฝจ่ฟนไผๅ ๆพๅจ้ฃ้๏ผ่ฎญ็ปๅจ้่ฆๆฐๆฎๆถๅๅ่ตฐใ
็ณป็ปไผๆฃๆฅ่ฟไบๆฐๆฎๆฏๅฆไป็ถๆ็จใๅฆๆๆๆก่ฝจ่ฟน็ญๅพ ๅคชไน ใไฝฟ็จ็ๆฏ่ฟๆถ็ๆจกๅ็ๆฌ๏ผๅฐฑๅฏ่ฝ่ขซไธขๅผๆ้ๆฐ็ๆใ่ฟไธช้ฎ้ขๅซไฝ ้ๆงๆง๏ผๆๆๆฏๆฐๆฎๅคชๆง๏ผๅทฒ็ปไธ่ฝๅพๅฅฝๅฐไปฃ่กจๅฝๅๆจกๅใ
ๆๆ ทไฟ่ฏ่ฎญ็ปๆฐๆฎๅ็กฎ๏ผ
ไฟ็ๅฎๅ จ็ธๅ็ token
่ฏญ่จๆจกๅๅฎ้ ไธไธๆฏ็ดๆฅๅค็ๅ่ฏ๏ผ่ๆฏๅค็ๆดๅฐ็ๆๅญ็ๆฎต๏ผๅซไฝ tokenใไธไธชๅ่ฏๅฏ่ฝๅฏนๅบไธไธชๆๅคไธช tokenใ
ๅค่ฝฎไปปๅกไธญ๏ผๆจกๅ็ๆถๆฏๅฏ่ฝ็ป่ฟ่ฎธๅคๅค็๏ผ
- ่ฝฌๆขๆ่ๅคฉๆ ผๅผ๏ผ
- ๅ ๅ ฅๅทฅๅ ท่ฐ็จไฟกๆฏ๏ผ
- ๅ ้คๆไบๅ ๅฎน๏ผ
- ้ๆฐๆๅๅๅฒๅฏน่ฏใ
ๅฆๆ่ฎญ็ปๅจๆๅ็ๅฐ็ token ๅๆจกๅๅฝๆถ็ๆญฃ็ๆ็ token ไธไธๆ ท๏ผ่ฎญ็ปๅฐฑๅฏ่ฝๆฏๅจๅญฆไน ไธๆฎตโๆจกๅไปๆช็ๆญฃ่ฏด่ฟ็่ฏโใ
Miles ไฝฟ็จ Token-In-Token-Out๏ผTITO๏ผ ๆบๅถๆฅ่งฃๅณ่ฟไธช้ฎ้ขใๅฎไผไฟๅญๆจกๅๅฎ้ ็ๆ็ tokenใ็ๆ่ฟไบ token ๆถ็ๆฆ็๏ผไปฅๅๅทฅๅ ท่ฐ็จไฟกๆฏใ
่ฟๆ ท๏ผ่ฎญ็ปๅจๅฏไปฅๅ็กฎๅฐ้็ฐๆจกๅๅฝๆถ็่กไธบใ
่ฟๅฐฑๅๅฝๅๆฏ่ต๏ผๅฆๆๆ็ป่ฆๅๆ่ฟๅจๅ็ๅจไฝ๏ผๅฐฑๅฟ ้กปไฝฟ็จๆฏ่ตไธญ็ๅฎๅ็็ๅฝๅ๏ผ่ไธๆฏๆ นๆฎ่ฎฐๅฟ้ๆฐๆผ็คบไธ้ใ
ๆททๅไธๅฎถๆจกๅไธญ็่ทฏ็ฑ้ฎ้ข
ๆไบๅคงๅๆจกๅๅซไฝ ๆททๅไธๅฎถๆจกๅ๏ผMoE๏ผใ่ฟ็ฑปๆจกๅๅ ้จๆ่ฎธๅคโไธๅฎถๆจกๅโ๏ผๆฏไธช token ๅชไผไบค็ปๅ ถไธญๅ ไธชไธๅฎถๅค็ใ
ๆจกๅไธญ็่ทฏ็ฑๅจไผๅณๅฎ๏ผ
่ฟไธช token ๅบ่ฏฅไบค็ปๅชๅ ไธชไธๅฎถ๏ผ
็ฑไบ็ๆ็ณป็ปๅ่ฎญ็ป็ณป็ปไฝฟ็จ็่ฎก็ฎๆนๅผๅฏ่ฝ็ฅๆไธๅ๏ผๅฎไปฌๆๆถไผๆๅไธไธช token ไบค็ปไธๅ็ไธๅฎถใ่ฟๆ ท๏ผ่ฎญ็ป่ฟ็จๅฐฑๅฏ่ฝๆดๆฐ้่ฏฏ็้จๅใ
Miles ๆไพไบ R3๏ผไนๅฐฑๆฏ rollout routing replayใๅฎไผ่ฎฐๅฝ็ๆๆถ้ๆฉไบๅชไบไธๅฎถ๏ผ่ฎญ็ปๆถๅไฝฟ็จๅฎๅ จ็ธๅ็้ๆฉใ
่ฟ็ฑปไผผไบ่ฎฉๅญฆ็ๅคไน ๆถไฝฟ็จๅ่่ฏๆถๅฎๅ จ็ธๅ็่งฃ้ขๆญฅ้ชค๏ผ่ไธๆฏๆขไธ็งๅฏ่ฝๅฏผ่ดไธๅ็ปๆ็ๆนๆณใ
็ ็ฉถไธญไฝฟ็จ็ๆๆฏๆนๆณ
่ฟ็ฏ่ฎบๆไธป่ฆๆฏไธ็ฏ็ณป็ป่ฎพ่ฎกๅๆง่ฝๆฅๅ๏ผ่ไธๆฏไผ ็ป็ๅฎ้ชๅฎคๅฎ้ชใไฝ่ ๆๅปบไบไธไธชๅฎๆด็่ฎญ็ปๅนณๅฐ๏ผๅนถๆต่ฏไบๅฎ็ๅคไธช้จๅใ
ไธป่ฆๆนๆณๅ ๆฌ๏ผ
- ไฝฟ็จ SGLang ็ๆๆจกๅๅ็ญๅๅค่ฝฎ่กๅจ๏ผ
- ไฝฟ็จ Megatron-LM ๆ PyTorch FSDP ่ฎญ็ปๆจกๅ๏ผ
- ไฝฟ็จไธๅๆนๅผๅจ็ๆ็ณป็ปๅ่ฎญ็ป็ณป็ปไน้ดๅๆญฅๆจกๅๅๆฐ๏ผ
- ไฝฟ็จ็ผๅฒๅบ็ฎก็ๅทฒ็ปๅฎๆ็่ฎญ็ปๆฐๆฎ๏ผ
- ่ฎฐๅฝ tokenใๆฆ็ๅไธๅฎถ่ทฏ็ฑ๏ผไฟๆ็ๆไธ่ฎญ็ป็ไธ่ด๏ผ
- ไฝฟ็จไฝ็ฒพๅบฆ่ฎก็ฎๅๅฐ GPU ๅ ็จๅนถๆ้ซ้ๅบฆ๏ผ
- ๆฏๆ LoRAใ็็ฃๅพฎ่ฐๅ็ฅ่ฏ่ธ้ฆ็ญ่ฎญ็ปๆนๅผ๏ผ
- ๅจไธๅ็ฑปๅ็ GPU ๅไธๅๆจกๅไธๆต่ฏ็ณป็ปใ
ๅ ถไธญ๏ผไฝ็ฒพๅบฆ่ฎก็ฎๆฏๆ็จๆดๅฐ็ๆฐๅญไฝๆฐ่กจ็คบๆฐๅญใไพๅฆ๏ผๆฎ้่ฎก็ฎๅไฝฟ็จๅพ็ฒพ็ป็ๅฐบๅญ๏ผ่ไฝ็ฒพๅบฆ่ฎก็ฎๅไฝฟ็จๅปๅบฆ่พ็ฒ็ๅฐบๅญใ็ฒๅฐบๅญ่ฎก็ฎๆดๅฟซใๅ ็จ็ฉบ้ดๆดๅฐ๏ผไฝๅฏ่ฝไธๅคๅ็กฎใ
Miles ่ฎฉ็ๆๅ่ฎญ็ป้ถๆฎตๅฐฝๅฏ่ฝไฝฟ็จ็ธๅ็ไฝ็ฒพๅบฆๆนๆณ๏ผไปฅ้ฟๅ ไธค่พน็ฎๅบไธๅ็็ปๆใ
ไธป่ฆๅ็ฐๅ็ปๆ
่ฎบๆๆฅๅไบๅ ไธช้่ฆ็ปๆใ
่ฎญ็ปๅ็ๆๅฏไปฅๅๆถ่ฟ่ก
ๅฎๅ จๅผๆญฅ่ฎญ็ป่ฝๅค่ฎฉ็ๆๅ่ฎญ็ปๅๆถ่ฟ่ก๏ผๅๅฐ GPU ็ฉบ้ฒๆถ้ดใๅฏนไบ้่ฆ้ฟๆถ้ดไฝฟ็จๅทฅๅ ท็ไปปๅก๏ผ่ฟไธ็นๅฐคๅ ถ้่ฆ๏ผๅ ไธบๆไบไปปๅกๅฏ่ฝ้่ฆๅพไน ๆ่ฝๅฎๆใ
ๅฏไปฅๆ้ซๅค่ฝฎไปปๅก็็ๆๆ็
้่ฟ่ฎฉๅไธไธชไปปๅกไธ็ด่ฟๆฅๅฐไฟๅญๅ ถๅๅฒไฟกๆฏ็ GPU๏ผMiles ๅฏไปฅ้ๅคๅฉ็จๅทฒๆ็็ผๅญใ
ๅจ่ฎบๆไธญ็ๅ่ๆต่ฏไธญ๏ผๅ็ผ็ผๅญ็ๅฝไธญ็่พพๅฐ 96%ใ่ฟๆๅณ็ๅคงๅคๆฐๆถๅ๏ผ็ณป็ปไธ้่ฆ้ๆฐๅค็ๅทฒ็ป็่ฟ็ๅฏน่ฏๅ ๅฎนใ
่ฝๅคๆฏๆๅคๆ็ๆบ่ฝไฝไปปๅก
Miles ๆฏๆๆจกๅไธ็ฏๅขไบๅจ๏ผไพๅฆ๏ผ
- ไฝฟ็จ็ป็ซฏ๏ผ
- ็ผ่พๆไปถ๏ผ
- ่ฟ่กไปฃ็ ๏ผ
- ๆฅๆถๆต่ฏ็ปๆ๏ผ
- ๆ นๆฎ็ปๆ็ปง็ปญ่กๅจใ
่ฟ่ฏดๆๅฎไธไป ้ๅ็ฎๅ็้ฎ็ญ่ฎญ็ป๏ผไน้ๅ่ฎญ็ป่ฝๅคๆง่กๅคๆญฅไปปๅก็ AI ๆบ่ฝไฝใ
ไฝ็ฒพๅบฆ่ฎก็ฎๅฏไปฅๅๅฐๆถ้ดๅ่ตๆบ
่ฎบๆๆต่ฏไบ BF16ใFP8ใMXFP8 ๅ NVFP4 ็ญๆฐๅญๆ ผๅผใไฝ่ ๆฅๅ่ฏด๏ผๅจๅทฒ็ปๆต่ฏ็้ ็ฝฎไธญ๏ผไฝ็ฒพๅบฆๆนๆณๅฏไปฅๆๆพๅๅฐ็ๆๆถ้ด๏ผๅๆถๅพๅฐไธ BF16 ๅบๅๆนๆณ็ธ่ฟ็ๅฅๅฑๆฒ็บฟใ
ไธ่ฟ๏ผMXFP8 ๅ NVFP4 ไปๅคไบๆต่ฏ้ถๆฎต๏ผๅชๅจ้จๅๆจกๅไธ้ช่ฏ่ฟ๏ผไธ่ฝไฟ่ฏๅฏนๆๆๆจกๅ้ฝๅๆ ทๆๆใ
ๅฎๆไบๅคงๅๆจกๅ็็ซฏๅฐ็ซฏๆกไพ
่ฎบๆๆๅๅฑ็คบไบไธไธชๅฎๆดๆกไพ๏ผ
- ๆจกๅ๏ผGLM-5.2 744B-A40B๏ผ
- ไปปๅก๏ผไฝฟ็จ็ป็ซฏๅฎๆ็ผ็จไปปๅก๏ผ
- ็กฌไปถ๏ผ64 ไธช NVIDIA GB300 GPU๏ผ
- ่ฎญ็ปๆนๅผ๏ผๅฎๅ จๅผๆญฅ็ๆบ่ฝไฝๅผบๅๅญฆไน ๏ผ
- ๅ 30 ไธชๆต้ๆญฅ้ชค็ไธญไฝๆถ้ด๏ผ263 ็งใ
่ฟ่ฏดๆ Miles ไธๅชๆฏไธไธช็่ฎบ่ฎพ่ฎก๏ผ่ๆฏๅฏไปฅ็จไบ้ๅธธๅคงๅ็ๅฎ้ ่ฎญ็ปไปปๅกใ
็ ็ฉถ็ปๆไธบไปไน้่ฆ๏ผ
ๅคงๅๆจกๅ่ฎญ็ป้ๅธธ้่ฆๅคง้ GPUใๆถ้ดๅ็ตๅใๅฆๆ็ณป็ป็ปๅธธ็ญๅพ ใ้ๅค่ฎก็ฎ๏ผๆ่ ๅ ไธบ็ๆๅ่ฎญ็ปไธไธ่ด่ๅบ้๏ผๆๆฌไผ้ๅธธ้ซใ
Miles ็้่ฆๆงๅจไบ๏ผๅฎ่ฏๅพๅๆถ่งฃๅณๅ ไธชๅฎ้ ้ฎ้ข๏ผ
- ่ฎฉ GPU ๆดๅฐ็ญๅพ ๏ผ
- ่ฎฉ่ฎญ็ปๆฐๆฎๆดๅ ๅ็กฎ๏ผ
- ่ฎฉๅค่ฝฎๅทฅๅ ทไฝฟ็จไปปๅกๆดๅฎนๆ่ฎญ็ป๏ผ
- ่ฎฉ้ๅธธๅคง็ๆจกๅ่ฝๅคๅๅธๅจ่ฎธๅค GPU ไธ๏ผ
- ่ฎฉ็ ็ฉถไบบๅๅฏไปฅๆดๆขๆจกๅใ็ฏๅขๅ่ฎญ็ปๆนๆณ๏ผ
- ่ฎฉ่ฎญ็ป่ฟ็จๅบ็ฐ้ฎ้ขๆถๆดๅฎนๆๆพๅฐๅๅ ใ
็ ็ฉถ็ๅฑ้
ไฝ่ ไน่ฏดๆไบ Miles ็ฎๅๅนถไธๅฎ็พ๏ผ
- ไธไบไฝ็ฒพๅบฆๆ ผๅผ่ฟๅชๅจๅฐๆฐๆจกๅไธๆต่ฏ่ฟ๏ผ
- ๆไบๆจกๅๅฎถๆ่ฟไธๆฏๆๅฎๆด็ๆ้ๅๆญฅ๏ผ
- ๅพๅๅ่ง้ข่พๅ ฅ็ฎๅไธ่ฝ้่ฟ TITO session server ๅค็๏ผ
- ไธไบๆบ่ฝไฝ็ฏๅข่ฟๆฅๅจไปๅจๅฎ้ช้ถๆฎต๏ผ
- ้จๅๆง่ฝ็ปๆๅชๆฅ่ชไธไธช็นๅฎ้ ็ฝฎ๏ผไธ่ฝไปฃ่กจๆๆ็กฌไปถๅๆจกๅ๏ผ
- ๅผๆญฅ่ฎญ็ป้่ฆไธบ็ๆๅ่ฎญ็ปๅๅคไธๅ็ GPU ่ตๆบ๏ผๅ ๆญคๅฏ่ฝ้่ฆๆดๅค็กฌไปถ๏ผ
- ๅฆๆๆฐๆฎๅจ็ผๅฒๅบไธญ็ญๅพ ๅคชไน ๏ผๅฐฑๅฏ่ฝๅๅพ่ฟๆถๅนถ่ขซไธขๅผใ
ๅฏ่ฝ็ๅฝฑๅ
ๅฆๆ Miles ็ปง็ปญๅๅฑ๏ผๅฎๅฏ่ฝ่ฎฉๆดๅค็ ็ฉถไบบๅๅไผไธๆดๅฎนๆ่ฎญ็ปๅคงๅ AI ๆบ่ฝไฝใ
ๆชๆฅ๏ผ่ฟ็ฑป็ณป็ปๅฏ่ฝๅธฎๅฉ AI ๆดๅฅฝๅฐๅฎๆ๏ผ
- ็ผ็จๅ่ฝฏไปถๆต่ฏ๏ผ
- ไฝฟ็จ็ต่ๅ็ป็ซฏ๏ผ
- ๅค็ๅคๆ็ๅทฅไฝๆต็จ๏ผ
- ไฝฟ็จๅค้จๅทฅๅ ท๏ผ
- ๅจๆจกๆ็ฏๅขไธญ่งฃๅณ้ฎ้ข๏ผ
- ่ฟ่กๅคๆญฅ้ชคๅณ็ญใ
ๆปไฝๆฅ่ฏด๏ผ่ฟ็ฏ่ฎบๆไป็ป็ไธๆฏไธ็งๆฐ็ AI ๆจกๅ๏ผ่ๆฏไธๅฅ่ฎฉๅคงๅๆจกๅ่ฎญ็ปๆดๅ ๅฟซ้ใๅ็กฎใๅฏ้ ๅๅฏๆฉๅฑ็ๅบ็ก่ฎพๆฝใๅฎๅๆฏไธๆกๆดๅฅฝ็โ่ฎญ็ป็ไบง็บฟโ๏ผๆจกๅ่ด่ดฃๅฐ่ฏ๏ผ็ฏๅข่ด่ดฃๅ้ฆ๏ผ่ฎญ็ปๅจ่ด่ดฃๅญฆไน ๏ผ่ Miles ่ด่ดฃ่ฎฉๆๆ้จๅๅ่ฐๅทฅไฝใ
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Incomplete empirical validation at frontier scale: The report provides one principal end-to-end case studyโGLM-5.2 744B-A40B on 64 GB300 GPUsโwithout establishing whether the reported performance generalizes to other model sizes, MoE configurations, GPU counts, cluster topologies, or workload types.
- Limited baseline comparisons: The paper does not provide systematic comparisons against synchronous RL, alternative asynchronous schedulers, other post-training systems, or unmodified SGLang/slime pipelines using matched hardware and workloads.
- Unclear qualityโthroughput trade-offs: The effect of asynchronous execution, trajectory staleness, group dropping, retries, and buffer size on reward quality, policy divergence, sample efficiency, and final benchmark performance is not quantified.
- No convergence analysis for stale data: The paper defines staleness operationally but does not establish theoretical or empirical bounds on how stale trajectories affect the RL objective, importance ratios, gradient bias, or training stability.
- Unresolved optimal staleness policies: It remains unclear how users should select staleness limits, buffer capacity, retry behavior, and replacement granularity for different environment latency distributions and policy-update rates.
- Insufficient analysis of trajectory-selection bias: Groups rejected because of uniform rewards, timeouts, or excessive staleness may produce a nonrepresentative training distribution, but the resulting bias in the learned policy is not measured.
- Straggler mitigation is not evaluated across workloads: The claimed benefits of sample-granularity replacement and least-loaded affinity routing are demonstrated primarily through a reference run; their effectiveness under highly heterogeneous task lengths, failures, or bursty environments remains unknown.
- Evaluation validity under asynchronous training is underexplored: The paper does not determine how delayed evaluation scores, skipped evaluations, checkpoint reuse, or mixed evaluation timing affect model-selection decisions and reported learning curves.
- No systematic study of routing affinity trade-offs: Session affinity improves KV-cache reuse but can create load imbalance; the paper does not quantify the trade-off across varying session lengths, fleet sizes, routing policies, or failure-recovery scenarios.
- Fault tolerance is insufficiently characterized: The behavior of in-flight sessions, cached prefixes, environments, buffers, weight versions, and optimizer state after worker, GPU, network, or sandbox failures is not described or experimentally evaluated.
- Reproducibility under nondeterminism is unresolved: The report does not establish whether runs can be reproduced across seeds, hardware types, kernel implementations, asynchronous schedules, or different rollout/training interleavings.
- Token-fidelity guarantees are conditional: TITO depends on registered model families, chat templates, reasoning parsers, and tool-call parsers; the paper does not quantify residual mismatch rates for supported models or explain how correctness is maintained when templates evolve.
- Unsupported multimodal workflows remain unresolved: The session server does not support image or video inputs, leaving token-exact multi-turn RL for vision-language and multimodal agent models unexplored.
- Harness-compatibility limitations are not systematically mapped: The consequences of branching histories, context compaction, retries, message rewriting, and non-verbatim replay are described qualitatively, but no benchmark measures failure rates or training-quality degradation across common agent frameworks.
- Unsafe loose matching lacks safeguards: Looser replay-comparison policies can merge histories with different tool calls or arguments, yet the paper does not provide automatic detection, provenance checks, or recovery mechanisms for such silent mismatches.
- R3 overhead and benefit are not comprehensively quantified: Although the paper estimates routing-tensor memory costs, it does not measure R3โs communication, storage, latency, and throughput overheads across sequence lengths, expert counts, batch sizes, and cluster scales.
- The necessity of R3 in asynchronous training remains unclear: The report notes that R3 may have limited effects in asynchronous settings but does not isolate routing mismatch from other sources of trainโrollout divergence or identify when R3 materially improves stability and final performance.
- No comparison of routing-replay alternatives: The paper does not compare exact routing replay with higher-precision routing, deterministic kernels, router-logit caching, selective replay, or tolerance-based approaches that might reduce memory and communication costs.
- Low-precision evidence is narrow: MXFP8 and NVFP4 are tested on only a small set of model families, and the report does not establish their robustness across longer contexts, larger MoE models, different objectives, or more diverse agentic tasks.
- Quantization effects beyond reward curves are unmeasured: The paper does not report detailed effects on calibration, log-probability error, gradient statistics, policy entropy, expert utilization, downstream task accuracy, or long-horizon training stability.
- Precision-contract coverage is incomplete: The behavior of unsupported formats, new architectures, alternative kernels, and model components held in BF16 is not systematically characterized, leaving open whether partial quantization introduces hidden trainโrollout discrepancies.
- Hardware portability is not demonstrated in the provided evidence: Although support is claimed across NVIDIA and AMD hardware, the report does not present matched performance, numerical-consistency, or stability results across vendors and GPU generations.
- Memory and offloading trade-offs are underreported: The excerpt introduces actor eviction and optimizer streaming but does not quantify their impact on step time, communication volume, disk or host-memory bandwidth, failure risk, or scalability.
- Scalability limits are unspecified: The paper does not identify bottlenecks or performance ceilings as the number of rollout engines, training ranks, environments, trajectories, or model parameters increases.
- Agent-environment cost is not separated from system cost: The case study does not decompose wall-clock time, energy use, and resource consumption among model inference, training, sandbox creation, tool execution, data movement, synchronization, and evaluation.
- Reward and environment reliability are not evaluated: The report assumes externally supplied rewards and sandbox behavior but does not study reward noise, flaky tests, environment nondeterminism, adversarial tool outputs, or reward hacking.
- Security and isolation risks are left open: Running model-generated commands in external sandboxes and allowing tool interactions introduces risks involving data exfiltration, privilege escalation, network access, and cross-episode contamination, none of which are analyzed.
- The generality of the plug-in architecture is unverified: The three connector layers are conceptually flexible, but the paper does not provide systematic integration studies showing that custom environments can preserve token fidelity, reward correctness, batching semantics, and failure handling simultaneously.
- Support for non-language diffusion post-training is unexplored: The report states that the architecture extends to diffusion models, but the provided material does not explain the adapted objectives, rollout semantics, synchronization requirements, or empirical validation.
- Production readiness is not independently established: The โproduction-readyโ characterization is not supported by long-duration tests, upgrade and rollback procedures, operational availability metrics, observability audits, or deployments beyond the reported configuration.
- The scope of the evidence is unclear because the paper text is incomplete: The supplied manuscript ends during the memory/offloading section, so conclusions about weight synchronization, additional training recipes, hardware coverage, code quality, and the full case-study methodology cannot be fully assessed.
Practical Applications
Immediate Applications
The paper presents Miles v0.1 as a production-oriented infrastructure system rather than an end-user model. Its most immediate applications are therefore in model development, evaluation, and deployment workflows where organizations already possess GPU clusters, training data, and executable environments.
- Frontier-scale reinforcement learning for LLMs โ Industry and academia; software/AI
- Organizations can use Miles to post-train LLMs with RL objectives such as GRPO, using multi-turn trajectories rather than only single completions.
- A practical workflow is: define prompts and reward functions, connect an agent environment or sandbox, generate trajectories with SGLang, train with Megatron-LM or FSDP, and synchronize updated weights back to rollout engines.
- This supports models optimized for coding, tool use, planning, customer-service interaction, and other tasks where success is measurable through an external reward.
- Dependencies: substantial GPU capacity, a reliable reward signal, compatible model architecture, and engineering expertise for distributed training.
- Training coding agents in isolated software environments โ Software engineering and developer tools
- Miles can train agents that edit files, execute shell commands, run tests, and receive a score from a repository-level or terminal-based task.
- Potential products include coding assistants that improve through test-based rewards, automated debugging systems, repository maintenance agents, and internal software-engineering copilots.
- The paperโs use of per-episode sandboxes through AgentENV, Daytona, E2B, Modal, or similar providers enables reproducible task execution and prevents one trajectory from contaminating another.
- Dependencies: secure sandboxing, deterministic or sufficiently stable test suites, well-designed task distributions, and protection against agents performing unsafe or costly operations.
- Agentic evaluation and benchmarking โ Industry, academia, and model governance
- The system can run multi-turn evaluations in which models use tools and interact with environments, rather than being judged only on static question-answering benchmarks.
- Evaluation fleets or external checkpoint backends can measure a specific model version asynchronously while training continues.
- This enables versioned comparisons of coding success, tool-use reliability, task completion, reward distributions, and failure rates.
- Dependencies: evaluation tasks must be sufficiently representative; asynchronous scores must be associated with the correct checkpoint; benchmark leakage and reward hacking must be controlled.
- Asynchronous RL infrastructure for better GPU utilization โ AI infrastructure and cloud computing
- Fully asynchronous scheduling allows rollout generation and training to run concurrently on separate GPU pools.
- This is especially useful for long-context or tool-using workloads where trajectory lengths vary substantially and synchronous training would wait for stragglers.
- The buffer metricsโqueue size, average and maximum staleness, and discarded groupsโcan be integrated into dashboards or autoscaling controllers to determine whether additional rollout or training capacity is needed.
- Dependencies: disaggregated GPU placement is required for the documented fully asynchronous mode; sufficient memory and interconnect bandwidth are also necessary.
- Cache-aware serving for multi-turn inference โ Model serving and inference systems
- Milesโs session-aware and data-parallel-rank-aware routing can be applied to agent-serving systems in which successive turns reuse a long context.
- A production tool could bind a conversation to the engine holding its KV cache, reducing repeated prompt prefilling and lowering latency and inference cost.
- The reported 96% prefix-cache hit rate in the reference configuration suggests a practical optimization for coding agents, customer-service agents, and workflow automation systems.
- Dependencies: sessions must carry stable routing keys; affinity can create load imbalance unless combined with least-loaded initial placement and monitoring.
- Reliable token accounting for multi-turn agents โ Training and observability tools
- The token-in-token-out session server can be used to ensure that the tokens sampled during rollout are exactly the tokens consumed during training.
- This is actionable for debugging training instability, auditing tool calls, reproducing model behavior, and validating log-probability calculations.
- A reusable workflow could store token IDs, log-probabilities, tool-call outputs, weight versions, and environment rewards as a complete trajectory record.
- Dependencies: the modelโs chat template and tool-call parser must be registered and verified; unsupported or vision-based inputs may require lower-level integration.
- Mixture-of-experts training consistency through routing replay โ Large-model training
- For MoE models, R3 can record the experts selected during rollout and replay those assignments during training.
- This can reduce discrepancies caused by different kernels, numerical precision, or routing decisions between serving and training.
- The technique is particularly relevant to organizations training large MoE models where small routing differences can accumulate into substantial policy drift.
- Dependencies: routing tensors increase memory and communication costs; the approach is relevant to MoE models, not dense models, and may have limited benefit when other asynchronous sources of mismatch dominate.
- Lower-cost RL with shared low-precision contracts โ AI infrastructure and cloud cost reduction
- The FP8, MXFP8, and NVFP4 workflows can reduce computation and memory requirements while maintaining a common quantization procedure across rollout, training, checkpoint conversion, and weight synchronization.
- Potential tools include low-precision RL recipes for model fine-tuning services and cluster schedulers that select precision based on supported GPUs.
- This can make post-training of large models more accessible on Hopper, Blackwell, or supported AMD hardware.
- Dependencies: support is model- and hardware-specific. MXFP8 and NVFP4 are described as beta-level recipes, and numerical behavior must be validated against BF16 baselines for each new model.
- Parameter-efficient RL with LoRA โ Industry and academic experimentation
- The paper states that Miles supports LoRA RL, allowing teams to adapt a base model using smaller trainable parameter sets.
- This can support domain-specific agents, customer-specific assistants, or rapid experiments without updating all model parameters.
- Potential products include a hosted service that trains and swaps task-specific adapters while preserving a shared base model.
- Dependencies: the paper excerpt provides limited empirical detail on LoRA performance; adapter quality, serving compatibility, and the effect of asynchronous data on adapter training require validation.
- On-policy distillation and supervised fine-tuning pipelines โ Model development
- Milesโs token-faithful trajectory handling can support supervised fine-tuning, on-policy distillation, and related post-training recipes in addition to conventional RL.
- An organization could generate responses from a teacher or current policy, preserve exact token-level information, and train a smaller or specialized student model.
- This is applicable to model compression, domain adaptation, and transferring tool-use behavior to smaller models.
- Dependencies: teacher quality, data filtering, distillation objectives, and alignment between teacher-generated tokens and the studentโs tokenizer or chat format.
- Pluggable research environments and reproducible RL experiments โ Academia
- The three nested connector layers allow researchers to replace the agent function, token-recording layer, or complete rollout orchestration without rewriting the trainer and weight-update system.
- This can shorten the time required to compare environments, reward schemes, trajectory filters, and scheduling policies.
- A laboratory could use the same training stack across coding, browser-use, simulation, and tool-calling experiments.
- Dependencies: connectors are reported as experimental and evolving; researchers must verify reward correctness, episode isolation, and compatibility with the selected tokenization policy.
- Operational monitoring for distributed post-training โ MLOps and platform engineering
- Milesโs queue and staleness metrics can form the basis of alerts and control policies.
- Examples include alerting when the rollout queue is empty, increasing rollout capacity when the trainer stalls, reducing rollout concurrency when the buffer is full, or retrying groups that become too stale.
- Weight-version verification can also be incorporated into deployment gates for evaluation and checkpoint promotion.
- Dependencies: metrics must be connected to reliable telemetry and interpreted in the context of workload variability; increasing concurrency without enough memory or sandbox capacity may worsen failures.
Long-Term Applications
The following applications are plausible extensions of the system, but they require broader validation, additional engineering, or stronger operational safeguards than are demonstrated in the paper.
- Autonomous software-engineering agents deployed in production โ Software and enterprise automation
- Miles could eventually support agents that autonomously implement features, fix defects, migrate code, update dependencies, and validate changes in CI environments.
- A mature workflow would combine task selection, isolated repositories, tool permissions, test-based rewards, human approval thresholds, and rollback mechanisms.
- Dependencies: training rewards must correlate with maintainable code rather than merely passing tests; environments need security controls, secret isolation, cost limits, and defenses against data exfiltration.
- General-purpose tool-using assistants โ Healthcare, finance, customer operations, and public services
- The combination of multi-turn token fidelity, external environments, asynchronous RL, and evaluation could train assistants that operate enterprise software, query databases, schedule actions, or execute business workflows.
- In healthcare, for example, an agent might be trained in a simulated clinical workflow; in finance, it might practice document review or compliance procedures.
- Dependencies: high-stakes deployment requires domain-specific validation, privacy protection, auditability, human oversight, and safeguards against irreversible actions. Reward functions alone are not sufficient evidence of safety.
- Robotic and embodied-agent post-training โ Robotics and autonomous systems
- The environment plug-in architecture could be extended to simulators or physical robots, with rewards based on task completion, safety, energy use, or motion quality.
- The systemโs multi-turn interaction model is relevant to robots that repeatedly observe, plan, call tools, and act in an environment.
- Dependencies: the paper does not demonstrate robotics or multimodal observation support. Real-world deployment would require image/video handling, low-latency control, sim-to-real transfer, safety certification, and robust handling of delayed or noisy rewards.
- Vision-language and computer-use RL โ Multimodal AI
- Once the session server supports image and video inputs, the same token-exact training principles could be applied to browser agents, desktop automation, visual inspection, and multimodal robotics.
- Potential products include agents that navigate graphical interfaces, inspect engineering diagrams, or operate visual workflow tools.
- Dependencies: the current session server does not support image or video inputs. Multimodal tokenization, visual state storage, screenshot replay, privacy controls, and temporal alignment would need to be developed.
- Adaptive cloud scheduling for post-training clusters โ Cloud infrastructure
- Buffer state, trajectory lengths, weight staleness, and evaluation backlog could feed an automated controller that dynamically reallocates GPUs between rollout, training, and evaluation.
- A scheduler might increase rollout capacity when the buffer is empty, increase training capacity when the buffer saturates, and reserve isolated evaluation resources when production measurements are required.
- Dependencies: reliable cost models, rapid workload migration, predictable interconnect performance, and policies preventing unstable oscillation between resource allocations.
- Large-scale distributed post-training across heterogeneous hardware โ AI platform engineering
- The support for NVIDIA and AMD hardware, together with multiple training backends and weight-transfer mechanisms, could lead to portable post-training platforms spanning different accelerator fleets.
- This may reduce dependence on a single vendor and permit organizations to use geographically distributed or opportunistically available hardware.
- Dependencies: the paper documents only selected hardware and model combinations. Kernel parity, quantization consistency, networking, fault tolerance, and cross-vendor performance require further benchmarking.
- Automated reward and trajectory-quality management โ Research tooling and AI safety
- The bounded buffer and user-replaceable selector could evolve into a quality-control layer that detects low-information groups, reward hacking, anomalous tool use, unsafe actions, or distribution shifts.
- Future systems could combine reward variance, trajectory diversity, environment outcomes, and staleness to prioritize which experiences should be trained on.
- Dependencies: filtering can inadvertently remove difficult but valuable examples; selection policies require careful statistical evaluation to avoid biasing the learned behavior.
- Near-zero-mismatch on-policy training for frontier MoE models โ Advanced model research
- Combining TITO, R3, shared quantization contracts, true-on-policy alignment, and carefully controlled weight synchronization could enable more faithful on-policy optimization at very large scale.
- This could improve the stability of RL for long-horizon agents and reduce failures caused by discrepancies between generation and training.
- Dependencies: exact replay has substantial memory and communication costs, especially for long trajectories and large routing tensors. The practical benefit must be established across diverse model families and asynchronous schedules.
- Unified post-training for language and diffusion models โ Generative media
- Because the paper states that the architecture extends to diffusion models, a longer-term application is RL or preference-based optimization of image, video, or other diffusion outputs.
- Potential objectives include adherence to prompts, visual quality, controllability, safety, latency, and domain-specific production requirements.
- Dependencies: diffusion trajectories, reward definitions, sampling steps, and weight synchronization differ from autoregressive language modeling. The paper does not provide enough evidence to establish production readiness for these workloads.
- Reproducible policy evaluation and model governance โ Public policy and regulated sectors
- Version-tagged asynchronous evaluation could support audit trails showing which policy checkpoint produced a given result and whether all evaluation requests used the intended weights.
- This may be useful for regulated deployments requiring reproducible testing, model-change documentation, and evidence that safety evaluations were not inadvertently run on mixed or outdated weights.
- Dependencies: technical weight verification does not by itself establish legal or ethical compliance. Governance frameworks would also need data lineage, access controls, human review, incident reporting, and independent audits.
- Personalized or organization-specific model adapters โ Education, enterprise, and daily productivity
- LoRA RL and supervised fine-tuning could eventually produce lightweight adapters for individual users, classrooms, departments, or companies.
- Examples include tutoring behaviors adapted to a curriculum, writing assistants aligned with organizational style, or workflow agents specialized for a teamโs software tools.
- Dependencies: user data must be collected lawfully and securely; personalization can amplify incorrect preferences or sensitive biases; adapter isolation and evaluation are necessary before deployment.
- Consumer-facing adaptive assistants โ Daily life
- In the long term, the methods could support assistants that learn from task outcomes such as successful calendar edits, completed household workflows, or user-approved tool actions.
- Multi-turn state tracking and cache-preserving routing could reduce latency for persistent personal sessions.
- Dependencies: continuous online RL is risky without strict consent, reversible actions, privacy protection, bounded exploration, and human confirmation. The paperโs infrastructure is primarily a training system and does not itself provide these consumer safety mechanisms.
Glossary
- Affinity: A routing policy that keeps related requests on the same serving engine to preserve cached context. โWe refer to this cache-preserving binding as affinity.โ
- Agentic RL: Reinforcement learning in which a model performs actions across multiple turns while interacting with an external environment. โIn agentic RL, where the model acts across multiple turns, each rollout session interacts with its own isolated environment, which executes actions and produces the reward.โ
- Attention: A neural-network mechanism that determines how strongly tokens should influence one another when producing representations. โDP-rank-aware routing further narrows that binding to an individual data-parallel rank when DP attention is enabled.โ
- Backward GEMM: A generalized matrix-multiplication operation used to compute gradients during backpropagation. โThe first, dequantized backward, touches only the training side: it runs the backward GEMMs in BF16 on operands dequantized from the NVFP4 values the forward pass used.โ
- BF16: The 16-bit brain floating-point format commonly used for efficient deep-learning computation. โBF16 train, FP8 serveโ
- Bit-exact: Producing identical numerical values at the bit representation level rather than merely approximately equal values. โBoth recipes share a bit-exact quantizer, so the training and rollout kernels see identical quantized valuesโ
- Cache locality: The tendency for computation to reuse data already stored close to the processor, thereby reducing memory-access cost. โSections~\ref{sec:rollout} covers the rollout stage, including fully asynchronous scheduling, agentic environments, token-in-token-out (TITO) sessions, and rollout routing replay.โ
- Checkpoint: A saved model state or an internally stored token-history state that can be reused later. โAfter each successful completion, the server checkpoints those prompt IDs together with the output token IDs, log-probabilities, and routed experts returned by SGLang.โ
- Colocated placement: A deployment arrangement in which training and rollout processes share the same GPUs. โMiles refuses to start a fully asynchronous run when the trainer and rollout engines share GPUs (a colocated placement).โ
- Collective operation: A synchronized operation involving multiple distributed training processes or devices. โExporting a fresh snapshot imposes a pause because it is a collective operation across the training actorsโ
- Contraction axis: The dimension over which multiplication and accumulation are performed in a tensor contraction. โThe contraction axis of some tensors, such as the final transformer layers, the shared experts, and the projections in multi-head latent attention, does not line up with a one-dimensional scaling block.โ
- Data-parallel rank: One participating replica in a distributed computation that processes a portion of the data. โDP-rank-aware routing further narrows that binding to an individual data-parallel rank when DP attention is enabled.โ
- Dequantization: The process of converting quantized numerical values back into a higher-precision representation. โIt runs the backward GEMMs in BF16 on operands dequantized from the NVFP4 values the forward pass usedโ
- Disaggregated placement: A deployment arrangement in which different pipeline stages use separate pools of hardware. โConcurrent execution requires GPU capacity for both stages at once, so the stages use separate GPU pools, a disaggregated placement.โ
- Distillation: Training a model to reproduce the behavior or probability distribution of another model. โBeyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignmentโ
- E4M3: An 8-bit floating-point representation with four exponent bits and three mantissa bits. โNVFP4 nests an E4M3 scale per block inside one FP32 scale per tensor.โ
- Expert routing: The process by which a mixture-of-experts model selects a subset of specialized subnetworks for each token. โIn a mixture-of-experts (MoE) model, rollout and training can send the same token to different expertsโ
- Fidelity: The degree to which generated training data accurately represents the model behavior and tokens that produced it. โRollout generation poses two distinct problems in agentic RL: throughput and fidelity.โ
- FlashInfer: A software library providing optimized GPU kernels for inference operations involving transformer models. โThe recipe enables it in the Transformer Engine kernels the trainer uses and in the FlashInfer kernels SGLang uses alikeโ
- Forward pass: The computation that transforms model inputs into outputs before gradients are calculated. โRollout and training run the same quantization logic on the forward pass.โ
- Fully asynchronous RL: Reinforcement learning in which rollout generation and model training proceed concurrently rather than in alternating phases. โFully asynchronous RL allows rollout generation and training to progress concurrently in Milesโ
- Generalized matrix multiplication (GEMM): A highly optimized operation that multiplies matrices or matrix-like tensors, central to neural-network computation. โBoth recipes share a bit-exact quantizer, so the training and rollout kernels see identical quantized values, apart from the per-tensor exceptions that stay in BF16.โ
- Gradient stability: The property that computed gradients remain numerically well behaved and useful for optimization. โTrading throughput for gradient stability while leaving the quantized values themselves unchanged.โ
- Group-relative centering: An operation that centers a trajectoryโs score relative to scores from other trajectories generated for the same prompt. โGroup rewards are scores that need the whole group at once, such as ranking the trajectories against one another, as distinct from GRPO's group-relative centeringโ
- HBM (high-bandwidth memory): Fast memory attached to a GPU and used to store model and training state. โMemory capacity limits a training run when a model's weights, gradients, and optimizer state exceed the GPU's available high-bandwidth memory (HBM).โ
- Importance ratio: The ratio between probabilities assigned by a current policy and the policy that generated the training data. โConsequently, the importance ratio drifts away from oneโ
- Incremental tokenization: Tokenizing only newly added text while reusing tokenized prefixes from earlier context. โIts incremental tokenization may diverge from the template's canonical output.โ
- Inference: The process of using a trained model to generate predictions or tokens. โLower-precision number formats make matrix multiplication faster.โ
- KV cache: Memory storing previously computed key and value representations so that autoregressive generation can reuse earlier context efficiently. โThe engine that served the previous turn already holds that prefix in its KV cache.โ
- LoRA (Low-Rank Adaptation): A parameter-efficient fine-tuning method that trains low-rank update matrices instead of all model parameters. โBeyond full-parameter RL, Miles also supports LoRA RLโ
- Log-probability: The logarithm of a modelโs probability assigned to a token or sequence. โThe server checkpoints those prompt IDs together with the output token IDs, log-probabilities, and routed experts returned by SGLang.โ
- Loss masking: Excluding selected tokens from contributing to the training loss. โThe sequence preserves the rollout log-probabilities while loss-masking the tokens the model did not generate.โ
- Mixture of experts (MoE): A model architecture that routes each input token through only a selected subset of specialized expert networks. โIn a mixture-of-experts (MoE) model, rollout and training can send the same token to different expertsโ
- Minibatch: A subset of training data processed in one optimization step. โThe trainer then forms each minibatch by pulling whichever trajectory groups have already finished from the buffer.โ
- NVFP4: A 4-bit floating-point format used for low-precision neural-network computation. โNVFP4 scales activations per token, which keeps quantization artifacts from depending on how a batch is composed.โ
- On-policy distillation: Distillation in which training data is generated by the current policy being optimized. โSection~\ref{sec:recipes} covers post-training paradigms beyond core RL, namely LoRA RL, on-policy distillation, and true-on-policy alignmentโ
- Optimizer state: The auxiliary parameters maintained by an optimization algorithm, such as momentum and adaptive-scaling statistics. โThe second mechanism, streaming the optimizer state, keeps optimizer state off the GPU during the training stepโ
- Prefix cache: A cache containing the previously processed beginning portion of a sequence. โAffinity and least-loaded placement together hold the prefix-cache hit rate at 96\%โ
- Quantization: The conversion of numerical values to a lower-precision representation to reduce computation or memory use. โProper quantization in RL is, however, not straightforward.โ
- Quantization-aware training: Training that accounts for the effects of quantization while learning model parameters. โTwo further options sit alongside them: quantization-aware training in INT4โ
- Rollout: The generation of a model trajectory, often through interaction with an environment, for use in training or evaluation. โSGLang engines generate trajectories.โ
- Rollout routing replay (R3): A technique that records and reuses the expert selections made during rollout instead of recomputing them during training. โR3 is a technique that mitigates this issue by treating each token's expert assignments as part of the rollout dataโ
- Sandbox: An isolated execution environment in which an agent can safely perform actions such as running commands or editing files. โA coding-agent environment, for example, provides a sandbox per task where the model runs commands, edits files, and receives a final grade from a test suite.โ
- Session server: A serving component that manages multi-turn history, tokenization, routing identity, and exact token recording. โSession server is the Miles component between the agent and the engines that takes ownership of a multi-turn trajectory.โ
- Staleness: The age of training data measured by how many weight updates separate its generating policy from the current policy. โMiles defines a group's staleness as the current trainer weight version minus the oldest weight version appearing anywhere in the group.โ
- Tensor core: Specialized GPU hardware designed to accelerate matrix operations used in deep learning. โA GPU's tensor cores roughly double their peak rate each time the precision halvesโ
- Token-in-token-out (TITO): An interface in which exact token IDs generated by the serving system are preserved and passed directly into training. โThe TITO session server closes that gap by letting the server, rather than the harness, control tokenization.โ
- Tokenization: The conversion of text or structured messages into the discrete token IDs consumed by a LLM. โOn the first turn, the server renders the selected template into token IDs.โ
- Trajectory: A complete sequence of model actions, observations, and responses produced while attempting a task. โA trajectory is one attempt at that task by the current policy.โ
- Train-rollout mismatch: A discrepancy between the probabilities or routing decisions used during rollout and those reproduced during training. โThis discrepancy appears as train-rollout mismatchโ
- Transformer Engine: A software and kernel stack optimized for transformer computation, including low-precision arithmetic. โThe recipe enables it in the Transformer Engine kernels the trainer usesโ
- Throughput: The amount of work completed per unit of time. โThroughput matters because rollout generation dominates wall-clock timeโ
- Weight synchronization: The transfer of updated model parameters from the trainer to rollout engines. โAfter each training step, Miles updates the RL policy by synchronizing the new weights with the rollout enginesโ
- Zero-KL alignment: An alignment approach designed to maintain exact or near-zero KullbackโLeibler divergence between specified policy distributions. โToken fidelity is therefore a precondition for three mechanisms described later: rollout routing replay (Section~\ref{sec:r3}), on-policy distillation (Section~\ref{sec:opd}), and true-on-policy alignment (Section~\ref{sec:zero-kl}).โ




