Scalability and generality beyond the evaluated model and task regime
Test the scalability and generality of Deep Noir beyond the evaluated model sizes and task settings, including 70B-plus models and complex reasoning, mathematics, and instruction-following tasks.
References
Testing scalability and generality on 70B+ models across complex reasoning, math, and instruction following remains future work.
— Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
(2609.20722 - Bobe et al., 17 Sep 2026) in Section 5, Limitations