Papers
Topics
Authors
Recent
Search
2000 character limit reached

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

Published 17 Sep 2026 in cs.AI | (2609.21139v1)

Abstract: Replacing attention in a pretrained LLM is a compatibility problem: a plausible substitute may alter representations expected by later layers. TinyCeNN-LM introduces a \emph{quality-gated post-training conversion} framework using CeNN-inspired cellular-recurrent layers with bounded local processing, compact recurrent memory, routing, fusion, and accept-or-rollback validation. Three implementations are studied: Integrated Memory, MemoryFusion, and PDelta3-GDN2-CLVR+Local32. Strict PDelta3 conversion accepts a layer only when representation and NLL criteria pass fixed thresholds. On SmolLM2-135M, layers 0-2 are accepted with cumulative ΔNLL=+0.01209Δ\mathrm{NLL}=+0.01209, while layer 3 is rejected despite acceptable NLL because representation fidelity fails. On Qwen3.5-0.8B, full-attention layers 3, 7, and 11 are accepted with final ΔNLL=+0.02073Δ\mathrm{NLL}=+0.02073. Integrated Memory keeps perplexity within −0.07%-0.07\% to +0.93%+0.93\% while reducing total cache by up to 6.01%6.01\%. A sampled 200-item downstream sanity check gives 28.5%28.5\%--32.0%32.0\% overall accuracy for converted Qwen releases. The results support conservative, quality-gated structural conversion rather than universal attention replacement or speedup.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.