Bayesian Distillation Framework with Tuning
- Bayesian Distillation with Behavioral Tuning (BBT) is a model-fitting procedure that combines Bayesian models' structured inductive biases with the flexibility of neural networks to predict human behavior in concept learning.
- BBT creates a robust initial hypothesis space via Bayesian distillation and then fine-tunes the network on human data to incorporate heuristics, biases, and noise, enhancing prediction accuracy in number concept learning, logical concept learning, Shepard–Hovland–Jenkins concepts, and compositional instruction learning.
- Case studies demonstrate that BBT significantly improves held-out log-likelihoods and prediction accuracy on task-specific metrics, for example BBT achieved an approximate Pearson correlation of r = 0.997 in number concept learning.
Bayesian Distillation with Behavioral Tuning (BBT) is a two-stage framework for constructing neural-network models that combine the structured inductive biases of Bayesian models with the flexibility required to predict human behavior. First, a neural network is trained on synthetic learning episodes generated by a Bayesian model, thereby approximating its posterior-predictive behavior. Second, the network is fine-tuned on human behavioral data, allowing it to acquire heuristics, biases, noise, exceptions, and representational structure absent from the original Bayesian hypothesis space. BBT was introduced for human concept-learning models and evaluated across number concepts, logical concepts, Shepard–Hovland–Jenkins concepts, and compositional instruction learning (Lake et al., 23 Aug 2026).
1. Conceptual foundations and scope
BBT addresses a tension between Bayesian cognitive models and connectionist models. Bayesian models provide explicit hypothesis spaces, interpretable priors, strong inductive biases, few-shot generalization, and inspectable posterior mass. Their limitations include hand-designed or overly narrow hypothesis spaces, strong parametric and independence assumptions, idealized accounts of human behavior, and potentially intractable posterior-predictive inference.
Neural networks provide flexible function approximation, distributed representations, efficient forward-pass prediction, and the capacity to learn heuristics, exceptions, and interactions directly from behavior. Their conventional limitations include data hunger, difficulty specifying qualitatively distinct inductive biases, challenges with systematic compositionality and symbolic generalization, and interpretability limitations.
BBT occupies an intermediate position. Bayesian modeling supplies the initial hypothesis space and inductive bias, while neural-network fine-tuning supplies representational and functional flexibility beyond the original Bayesian assumptions. The framework is a model-fitting procedure rather than a developmental theory: it does not claim that human cognition literally develops by first performing Bayesian inference and then learning through neural-network-like mechanisms.
The Bayesian starting point consists of a hypothesis space , a prior , and a likelihood . For study examples and , the posterior is
For a query , the Bayesian posterior predictive is
BBT trains a neural network to approximate this mapping:
The network therefore receives a study episode and a query and predicts a distribution over possible query outputs.
2. Bayesian distillation through synthetic episodes
Synthetic episode generation
The first stage uses forward sampling from the Bayesian generative process rather than repeated posterior inference. A synthetic episode is generated by:
- Sampling a hypothesis 0.
- Sampling study inputs and outputs under the likelihood.
- Sampling one or more query inputs and their outputs under the same hypothesis.
- Training the network to predict the sampled query outputs from the study examples.
An episode has the form
1
Forward sampling requires the prior and likelihood but not posterior inference or posterior sampling. Across many episodes, the neural network learns an amortized approximation to Bayesian inference. At test time, adaptation to a new task occurs through the network’s activations and input context rather than through weight updates.
Distillation objective
The network is trained with token-level cross-entropy:
2
For single-token binary tasks, this reduces to binary cross-entropy,
3
The objective is applied to outputs sampled from the Bayesian model. It does not directly force the network to store an explicit posterior over hypotheses; rather, the posterior-predictive mapping is encoded in the network parameters and activations.
Architecture and optimization
All four case studies used approximately the same 30-million-parameter encoder–decoder Transformer, with four encoder layers, four decoder layers, eight attention heads, 512-dimensional embeddings, 2048-dimensional feed-forward layers, GELU activations, dropout probability 4, and absolute sinusoidal positional encoding. The encoder receives a concatenated representation of 5, and the decoder produces the query output token by token.
Distillation used AdamW with learning rate 6, weight decay 7, 8, and 9. Training used 50 epochs, linear warm-up during the first epoch, a reduce-on-plateau scheduler, tenfold learning-rate reductions, gradient-norm clipping at 0, and termination after the fourth attempted learning-rate reduction.
The architecture is not conceptually mandatory. The authors identify the encoder–decoder Transformer as an implementation choice; a decoder-only architecture could also be used.
3. Behavioral tuning and the departure from Bayesian idealization
After Bayesian distillation, the same network is fine-tuned on human responses. A behavioral episode contains study information, a query, and the answer produced by a participant. The behavioral objective is again token-level cross-entropy:
1
Unlike approaches that combine Bayesian and behavioral losses in a single fixed weighted objective, BBT applies the objectives sequentially. Synthetic-data cross-entropy is used first, followed by behavioral-data cross-entropy.
Behavioral tuning is not constrained to preserve the original Bayesian prior, likelihood, hypothesis space, or inference rule. Fine-tuning may alter the relative influence of hypotheses, change the effective prior, change how data are related to hypotheses, introduce representations absent from the original hypothesis space, encode response heuristics, and learn systematic deviations, noise, exceptions, and biases.
Behavioral fine-tuning used the same optimizer and basic settings as distillation, reinitialized at the beginning of the second stage, with validation-based model selection. Networks were trained for 20 epochs, or 50 epochs for Shepard learning. Validation loss was tracked every 100 steps, and the parameters with the best validation loss were retained. For Shepard learning, block and age embeddings were additionally learned from behavioral data.
A response lapse model can be applied after prediction. For a fixed-length response with possible symbols 2,
3
where 4 is the model prediction before lapsing and 5 is the lapse probability. For variable-length sequences, the end-of-sequence token is included in 6.
The distinction between the two stages is central. Distillation transfers the Bayesian model’s structured inductive bias, while behavioral tuning permits systematic departure from Bayesian predictions when those departures improve prediction of human responses.
4. Interpretability, posterior structure, and representation change
BBT supports interpretability through comparisons between the Bayesian model, the distillation-only network, and the behaviorally tuned network. The authors use sparse combinations of human-readable hypotheses from the original Bayesian model to approximate network predictions. For deterministic binary hypotheses,
7
The network is approximated by
8
where 9, 0, and 1 is a lapse or noise parameter. The weights are encouraged to be sparse.
This analysis is diagnostic rather than constitutive. It does not prove that the neural network literally represents a posterior over the original hypotheses. It is restricted to the Bayesian candidate set and can fail to reveal genuinely new representations learned during behavioral tuning.
The authors also analyze internal embeddings using principal-component methods and hypothesis-class organization. Distillation tends to organize representations according to Bayesian hypothesis classes. Behavioral tuning can preserve this organization while changing the relative influence of hypothesis classes, expanding simpler hypotheses, blurring distinctions between previously separate classes, or incorporating hypotheses absent from the original model.
This distinction is important for interpretation. BBT preserves a Bayesian representational substrate as an initialization and learned organization, not as a hard constraint. Consequently, the final network can remain Bayesian-like while implementing a different effective prior, likelihood, or inference strategy.
5. Empirical case studies
Number concept learning
The number-game task required participants to judge whether a query number was likely to have been generated by the same program as a small study set. The Bayesian teacher was Tenenbaum’s number-game model with 5,084 hypotheses, including evens, odds, multiples, powers, mathematical hypotheses, and interval hypotheses. Distillation used 100,000 synthetic episodes.
On 51 held-out sets and 5,100 query/support combinations, the distillation-only network matched the Bayesian model with Pearson correlation
2
Held-out log-likelihoods were:
| Model | Log-likelihood |
|---|---|
| Bayesian | 3 |
| Distillation only | 4 |
| Fine-tuning only | 5 |
| BBT | 6 |
For aggregated yes probabilities, BBT achieved Pearson 7 and RMSE 8, compared with 9 and RMSE 0 for the Bayesian model.
Behavioral tuning shifted predictions toward human-endorsed regularities. For the support set 1, the Bayesian model favored multiples of 6 with posterior probability 2, whereas BBT produced predictions consistent with evens at approximately 3 and reflected “ending in 6,” a hypothesis absent from the original Bayesian model. For 4, BBT reflected “ending in 8” and evens in addition to multiples of 8.
Logical concept learning
Participants classified objects according to concepts involving color, shape, size, Boolean operators, relations, and quantifiers. The Bayesian teacher used a probabilistic context-free grammar favoring shorter expressions and included object features, Boolean operators, relations, and quantifiers.
Held-out log-likelihoods were:
| Model | Log-likelihood |
|---|---|
| Bayesian | 5 |
| Distillation only | 6 |
| Fine-tuning only | 7 |
| BBT | 8 |
BBT captured partial endorsement of simple rules such as “blue” when the formal rule was “largest blue,” uncertainty about applying “largest” in small sets, and attraction to blue or green objects under a rule involving the only blue or green object. Behavioral tuning shifted the model toward simpler hypotheses, including true, false, individual features, and Boolean rules without quantifiers.
The aggregate accuracy of BBT and the distillation-only network was both 9, indicating that the behavioral improvement was not simply better identification of the intended rule. On concepts shared with validation but absent from fine-tuning, BBT improved held-out-concept performance by at least 0 natural-log points relative to the pretrained network.
Shepard–Hovland–Jenkins concepts
The task involved six balanced category structures over three binary features. The Bayesian teacher was the Rational Rules model, using Boolean rules and a simplicity-biased probabilistic context-free grammar. Distillation used 1 synthetic episodes. Behavioral tuning incorporated block and age-group embeddings.
Held-out log-likelihoods were:
| Model | Log-likelihood |
|---|---|
| Rational Rules/Bayesian | 2 |
| Distillation only | 3 |
| Fine-tuning only | 4 |
| BBT | 5 |
BBT reproduced the human learning trajectory with 6, while the fine-tuning-only network achieved 7. Simulated accuracies were 8 for young adults and 9 for older adults, compared with human accuracies of 0 and 1, respectively. Thus, BBT captured the direction and interaction of age effects while remaining globally more accurate than participants.
After the first block, BBT increased the influence of simple Type-I hypotheses from 2 to 3 for young adults and from 4 to 5 for older adults. The authors interpret this as evidence that participants, particularly younger participants, favor simple one-feature rules more strongly than specified by the Rational Rules prior.
Compositional instruction learning
Participants learned a pseudolanguage containing primitive words and function words that transformed one or two arguments. The task required generalization to novel queries involving compositions of up to three functions.
The Bayesian teacher used a prior over interpretation grammars, including primitive rules, function rules, and a concatenation rule. Distillation used 100,000 synthetic episodes. Behavioral data were collected from 226 participants, with 166 retained for fine-tuning.
Held-out log-likelihoods were:
| Model | Log-likelihood |
|---|---|
| Distillation only | 6 |
| Fine-tuning only | 7 |
| BBT | 8 |
BBT improved over distillation alone by 9 natural-log points and slightly outperformed the earlier hand-engineered model from Lake and Baroni, with log-likelihoods of 0 and 1, respectively.
Behavioral tuning produced human-like errors. One-to-one errors accounted for 2 of BBT errors and 3 of human errors, compared with 4 for the distillation-only model. Iconic-concatenation errors involving function 3 accounted for 5 of BBT errors, 6 of human errors, and 7 of distillation-only errors. These biases were not manually encoded; they emerged through behavioral fine-tuning.
6. Limitations, extensions, and relation to broader Bayesian distillation
BBT depends strongly on the selected Bayesian model. Its behavior is shaped by the hypothesis space, prior, likelihood, and synthetic episode distribution. An overly narrow or poorly specified Bayesian teacher can provide an unhelpful initialization. Broader priors and more deterministic likelihoods often produced better BBT models according to the authors’ informal observations, but this interpretation was not systematically established.
Synthetic-data design is therefore a substantive modeling choice. Forward sampling teaches the network the structural logic embodied in the Bayesian model, while behavioral fine-tuning supplies noise, exceptions, and systematic deviations. With limited human data, validation-based model selection is necessary to reduce overfitting to participants or task-specific response sequences.
Interpretability remains restricted. Sparse decomposition can identify only hypotheses among the original Bayesian candidates and may miss new structures created by behavioral tuning. Results can also depend on the selected episodes, layers, variables, query sets, and candidate hypotheses. Individual differences are only partially modeled: the Shepard experiment included age embeddings, while participant-specific embeddings are identified as a future extension.
The framework does not provide a formal guarantee that behavioral tuning preserves Bayesian structure, nor does it specify a universal architecture or optimization procedure. It is intended to produce predictive models, not to establish a developmental account of human cognition.
BBT is related to, but distinct from, broader forms of Bayesian predictive distillation. “Bayesian Dark Knowledge” compresses a Monte Carlo posterior predictive distribution into one neural network, typically by minimizing KL divergence or cross-entropy between teacher and student predictive distributions (Korattikara et al., 2015). Generalized Bayesian Posterior Expectation Distillation extends this idea to arbitrary posterior expectations, including predictive probability, expected entropy, and posterior marginal variance (Vadera et al., 2020). These frameworks primarily preserve Bayesian predictive behavior, whereas BBT deliberately permits behavioral fine-tuning to modify that behavior.
Other distillation approaches provide complementary mechanisms. Efficient uncertainty estimation for Bayesian LLMs transfers Bayesian model-averaged predictive distributions into a deterministic student to remove test-time sampling (Vejendla et al., 16 May 2025). Bayesian teachers can also reduce optimization noise in student training when their outputs approximate Bayes class probabilities (Morad et al., 4 Jan 2026). Representation-structure distillation transfers behavior through internal relational geometry rather than only outputs, allowing transfer across tasks and architectures (Pogoncheff et al., 29 May 2025). These approaches are conceptually adjacent but do not constitute BBT unless they include the two-stage transition from Bayesian distillation to behavioral tuning.
The principal contribution of BBT is therefore the explicit separation between structured Bayesian initialization and empirically grounded behavioral adaptation. Its final models retain evidence of the Bayesian substrate while learning systematic departures that improve human-behavior prediction, including altered hypothesis preferences, age and learning effects, human-like errors, and representational structures absent from the original Bayesian formulation.