Papers
Topics
Authors
Recent
Search
2000 character limit reached

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

Published 20 Jul 2026 in cs.LG | (2607.17624v1)

Abstract: Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. Method. We present a method to optimize a transformer architecture for a given dataset, which we use as a tool to study optimal task-specific inductive biases. This method replaces the most important non-linearities (GeLUs,;softmax) with functions learned on held-out data. We then train the resulting architectures on other datasets, as a way to evaluate the compatibility between pairs of tasks. Findings. On algorithmic toy tasks, we identify new architectures with dramatic improvements in learning speed, in- and out-of-distribution generalization, and stability across seeds. The new designs prove very task-specific however, and indicate that these tasks require inductive biases very different from those of standard transformers. On code and language modeling datasets, we also find architectures with consistent, yet smaller improvements. These designs transfer much better across datasets and domains (English & computer code). Implications. Our results show that standard transformers are rarely a local optimum in the space of architectures. Simple alternatives can perform much better but sacrifice universality. This suggests that there may be room for improved architectures that better support multiple capabilities simultaneously, such as fluency and robust reasoning.

Summary

  • The paper introduces a two-stage method that learns transformer MLP activations and attention kernels as spline-based architectural biases, then freezes them for evaluation across datasets.
  • The paper finds optimized architectures converge 2–3× faster, reduce training variance, improve length and capacity generalization, and deliver the largest gains on algorithmic reasoning tasks.
  • The paper shows task-specific optimization often causes negative transfer on algorithmic tasks, while modest improvements in language and code modeling transfer broadly across datasets.

Overview

This paper, "Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks" (2607.17624), by Teney, Jiang, Saratchandran, and Lucey (Idiap Research Institute, EPFL, Adelaide University), investigates whether the standard transformer architecture encodes inductive biases that are optimal for any given task, and whether the biases suited to one task transfer to others. The authors' central methodological contribution is a procedure that replaces the two principal non-linearities of a GPT-2-style decoder — the GeLU activations in MLP layers and the softmax kernel in attention — with learnable linear splines optimized on held-out data. The resulting frozen non-linearities are then treated as fixed architectural hyperparameters and used to retrain models from scratch, both on the original dataset and on other datasets, providing a direct probe of task-specific optimality and cross-task compatibility.

The paper's headline claims are deliberately provocative relative to prevailing assumptions: standard transformers are rarely a local optimum in architecture space for any given dataset, yet the architectures that improve upon them tend to be highly task-specific, particularly for algorithmic reasoning tasks. This tension between per-task optimality and universality is the paper's core empirical finding.

Method

The method operates in two stages. In stage I, the architecture is optimized jointly with model weights on a chosen dataset D\mathbb{D}. The MLP activation ϕMLP\phi_{\mathrm{MLP}} becomes a 1D linear spline parametrized by learnable keypoints (typically 122 points over [−20,+20][-20, +20]), and attention is generalized from the exponential dot-product kernel K(x,y)=exp⁡(x⊤y/d)K(x,y) = \exp(x^\top y / \sqrt{d}) to a learned factorization K(x,y)=ϕ′(x)⊤ϕ′(y)K(x,y) = \phi'(x)^\top \phi'(y) where ϕ′\phi' is itself a spline. Linear splines are chosen because they impose minimal priors — unlike trainable-activation approaches that enforce smoothness or monotonicity — allowing representation of sharp transitions and periodic structure that hand-designed activations cannot express.

Two mechanisms prevent co-adaptation between weights and non-linearities during stage I:

  • Two-loss training: a fraction of training data (e.g., 20%) is held out exclusively for optimizing the splines, while weights train on the remainder. For length generalization experiments, the held-out split is an OOD range of sequence lengths, forcing the architecture to encode a length-generalizing bias.
  • Multi-model training: MM models (e.g., M=8M=8) with different seeds share the same splines being optimized, encouraging solutions that generalize across weight configurations.

In stage II, the splines are frozen and models are trained conventionally from scratch on any target dataset D′\mathbb{D}', making results directly comparable to baseline transformers. The authors distinguish this from prior work on adaptive activation functions [alexandridis2025adaptive], which continuously updates activations during training rather than producing reusable, hard-encoded biases.

Results on algorithmic tasks

Across eight algorithmic tasks (memorize, Dyck-language parentheses recognition, AddMod, Haystack recall, decimal addition, reversed addition, Copy, and a scaled-down Mano task), optimized architectures yield what the authors characterize as dramatic improvements:

  • Convergence speed: 2–3× faster convergence on Add and Mano, with baselines tuned to their maximum stable learning rate.
  • Variance reduction: baseline transformers exhibit large seed-to-seed variance on several tasks; optimized architectures largely eliminate this instability.
  • Generalization: on tasks where baselines fit training data but fail to reach perfect test accuracy (e.g., Mano), optimized architectures close the gap.
  • Length generalization: on Copy — described as elementary but unsolved — the baseline fails completely on unseen lengths; Alibi positional encodings provide partial relief; optimizing the Alibi architecture with the two-loss mechanism further extends accuracy to longer sequences. This is not presented as a complete solution, but it demonstrates that inappropriate base-architecture biases are one obstacle to length generalization.
  • Capacity efficiency: on some tasks (parentheses, memorize, AddReversed), optimized architectures maintain higher accuracy as width shrinks, implying that aligned architectures extract more capability per parameter.

Ablations indicate most benefits come from optimizing MLP non-linearities rather than attention; learned attention kernels were difficult to optimize and rarely beat softmax. Notably, the cross-task evaluation shows these gains are highly task-specific: few improvements transfer beyond closely related pairs (Add ↔ AddReversed), and many optimized architectures underperform the standard transformer on foreign tasks. Specialization thus comes at the cost of universality, which the authors use to explain why proposed components (alternative attention, positional encodings) rarely see adoption beyond toy settings. Whether this negative transfer is inevitable remains open; multi-task optimization is suggested as a next step.

Results on language modeling

On seven datasets — TinyStories, Shakespeare and enwik8 at character and subword levels, and CodeSearchNet Java/Python — optimized architectures produce consistent but smaller improvements than on algorithmic tasks. Key observations:

  • Optimized MLP non-linearities alone account for most gains; replacing softmax barely matches or underperforms the baseline, and alternative parametrizations initialized to mimic softmax barely move away from initialization, suggesting softmax is close to a local optimum for attention.
  • On TinyStories, learned non-linearities resemble sine wavelets, and none of the surveyed alternatives (GLU variants, ReLU², TanH, Sinc, Gaussian, polynomial attention, NormSoftmax, adaptive softmax) outperforms them. Fine details matter: symmetrized or regularized versions systematically perform worse.
  • Initializing optimization from a GeLU ("GeLU + Ours") yields intermediate performance, supporting the claim that GeLUs are neither globally nor locally optimal even for standard language modeling.
  • Improvements are larger for code than natural language, relative to the linear-vs-GeLU baseline gap, which the authors attribute to the greater systematic structure and compositionality of code — connecting code modeling to the algorithmic regime.
  • Cross-dataset compatibility is high: architectures optimized for one language or code dataset transfer well across all seven plus Mano, indicating far more uniform skill requirements than in the algorithmic setting.

Supplementary experiments scale the approach onto the NanoGPT Speedrun codebase with FineWeb data. There, the optimized spline surpasses ReLU and GeLU baselines in validation loss at depths of 4–12 layers (e.g., 3.68 vs. 3.72 at 12 layers), and a degree-18 polynomial approximation recovers the spline's performance at near-ReLU computational cost. This addresses the practical concern that exact spline evaluation becomes bandwidth-constrained in larger models.

Limitations and open questions

The paper concedes three principal limitations. First, the search space excludes complex attention forms (e.g., tropical attention) and multiplicative interactions such as GLUs in their full generality; larger spaces may yield further gains but harder optimization. Second, experimental scale is tiny relative to state-of-the-art LLMs; while the authors argue small-scale effects remain meaningful because data efficiency is the objective, they acknowledge effects may diminish with more data. Third, although efficient implementations exist, the optimized designs are not directly deployable replacements — their value is diagnostic rather than practical. Open questions include whether multi-task architecture optimization can recover universality without sacrificing per-task gains, whether improved architectures transfer to vision and speech, and whether the negative cross-task transfer observed on algorithmic tasks is fundamental or an artifact of narrow single-task optimization.

Conclusion

By gradient-optimizing transformer non-linearities against specific datasets and evaluating the resulting frozen architectures across tasks, this paper provides evidence that standard transformers sit far from local optima for algorithmic skills — where simple spline-based modifications yield large gains in speed, stability, generalization, and capacity efficiency — but lie much closer to optima for language and code modeling, where gains are modest yet consistent and transfer broadly across domains. The trade-off uncovered is explicit: better task-specific inductive biases sacrifice universality. The work reframes architecture improvement as a problem of simultaneously supporting heterogeneous capabilities, and leaves open whether joint optimization across tasks can resolve it.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.