Close the language-modeling performance gap in TTFS generative models

Close the substantial language-modeling performance gap between time-to-first-spike (TTFS) generative models and their artificial neural network counterparts, particularly on perplexity-sensitive tasks such as WikiText and LAMBADA.

Background

The paper reports that TTFS-GPT-2 models perform competitively with or better than artificial neural network baselines on several commonsense reasoning benchmarks, but exhibit substantially worse language-modeling results. For TTFS-GPT-2 XL, WikiText perplexity increases from 20.4 to 26.6, while LAMBADA perplexity increases from 10.6 to 23.8 and accuracy decreases from 51.2 to 39.4.

The authors attribute this gap to discretized spike timings altering logit rankings and accumulating quantization error across long contexts. They identify closing this gap as the principal unresolved challenge for TTFS-based generative models.

References

Closing this language-modeling gap is, in our view, the main open problem for TTFS-based generative models.

Large Language Models with At Most One Spike per Neuron  (2609.05151 - Zhao et al., 4 Sep 2026) in Section 4, subsection “TTFS-GPT2”