Scalability and generality beyond the evaluated model and task regime

Test the scalability and generality of Deep Noir beyond the evaluated model sizes and task settings, including 70B-plus models and complex reasoning, mathematics, and instruction-following tasks.

Background

Deep Noir is evaluated primarily on binary classification tasks using models ranging from approximately 1B to 9B parameters, with limited extensions to reasoning and generation. The reasoning experiments produce limited gains because the tested models perform near randomly, leaving insufficient contrastive signal for steering.

The paper therefore leaves unresolved whether the steering-discovery framework remains effective at substantially larger scales and on more demanding capabilities such as complex reasoning, mathematical problem solving, and instruction following. Establishing this would determine whether the observed architectural and task-dependent steering behavior generalizes beyond the experimental regime.

References

Testing scalability and generality on 70B+ models across complex reasoning, math, and instruction following remains future work.

— Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models  (2609.20722 - Bobe et al., 17 Sep 2026) in Section 5, Limitations