Context-Aware Incremental MLP
- Context-Aware Incremental MLP is a continual learning model that sequentially updates on tabular data streams by retrieving latent context via scaled dot-product attention.
- It integrates a fixed-size FIFO memory to store recent latent features, avoiding raw-data replay and mitigating catastrophic forgetting.
- The design emphasizes low energy, bounded memory, and modest compute, making it ideal for edge, IoT, and mobile deployments.
Context-Aware Incremental Multi-Layer Perceptron (IMLP) denotes, in the explicit usage of the 2025 tabular-stream continual-learning literature, a compact continual learner for tabular data streams that is trained sequentially on incoming segments while preserving prior knowledge without revisiting raw data from earlier segments. In this formulation, “incremental” refers to sequential no-replay training over a stream with fixed label space, and “context-aware” refers to retrieval of a learned context vector from a fixed-size FIFO latent memory using scaled dot-product attention. The resulting model remains structurally close to a small MLP, but augments the current feature vector with attended latent history in order to reduce catastrophic forgetting under low energy, bounded memory, modest compute, and simple deployment constraints, particularly on edge, IoT, and mobile systems (Wang et al., 6 Oct 2025). In adjacent literature, closely related ideas appear under different names, especially the gated context-aware neural layer and multi-layer extension CA-NN and CA-RES, which provide a broader probabilistic interpretation of context-aware feed-forward networks without using the IMLP name explicitly (Zeng, 2019).
1. Problem regime and design objective
IMLP is defined for a specific continual-learning regime: data arrive sequentially as a stream of segments or tasks; the label space stays fixed; and when segment arrives, the learner may train on that segment, possibly for multiple epochs, but cannot revisit raw data from earlier segments. The objective is to keep learning from the current segment while preserving useful knowledge from previous ones, but to do so under deployment constraints that are central in healthcare, finance, and IoT deployments: low energy, bounded memory, modest compute, and simple implementation and deployment (Wang et al., 6 Oct 2025).
This regime is narrower than generic lifelong learning and stricter than cumulative retraining. Standard replay-based continual-learning methods are treated as problematic in this setting because they preserve old knowledge by storing raw examples or synthetic surrogates and replaying them during training. In practical terms, this makes the memory footprint grow with the number of tasks or segments, increases computation as the replay set grows, and raises energy use because each update involves more data movement and more gradient steps over historical examples. IMLP therefore targets online adaptation without full retraining, without raw-sample storage, and with memory use that stays fixed over time (Wang et al., 6 Oct 2025).
Within this framing, IMLP differs from a plain segment-wise MLP in two specific ways. First, its parameters are carried forward across segments rather than reinitialized. Second, each current input is augmented by attended historical context from latent memory rather than being processed only through the weights. It also differs from replay-based methods because it does not store raw historical examples and does not perform explicit rehearsal on old data. Finally, it differs from transformer-style tabular models because it is not a full tokenized transformer over features or rows; the architecture is a small MLP plus a lightweight retrieval block over recent latent vectors (Wang et al., 6 Oct 2025).
2. Architectural specification
The input is a standard tabular feature vector . Numerical features are standardized, categorical variables are one-hot encoded, and the resulting tabular vector is used directly. There is no tokenization of columns into embeddings in the transformer sense. The latent dimension is (Wang et al., 6 Oct 2025).
For a minibatch , the model projects the current input to a query,
and attends over a sliding memory of recent latent features. Keys are
and the paper explicitly ties values to keys,
Scores are computed as
followed by
and the attended context vector
0
After squeezing the singleton dimension, 1 is concatenated with the original input,
2
This fused representation is processed by a shared two-layer MLP feature extractor and a linear classifier,
3
4
The complexity discussion indicates hidden dimensionality 5 in the feed-forward block, so the intended structure is concatenated input 6 FC1/activation 7 FC2/activation 8 256-dimensional penultimate feature 9 linear classifier (Wang et al., 6 Oct 2025).
A notable architectural subtlety is the memory definition. The text is slightly inconsistent between buffering per-sample penultimate features and storing per-segment prototypes derived from them. The clearest formal description is a FIFO buffer 0 of the last 1 segment prototypes, each a detached 256-dimensional vector. This suggests that “windowed” refers to a fixed-length FIFO temporal memory rather than local attention over positions within a single example (Wang et al., 6 Oct 2025).
3. Incremental update rule, latent memory, and complexity
The memory stores detached latent features rather than raw data. After segment 2 is processed, the model computes the mean penultimate representation for that segment,
3
and updates the FIFO buffer as
4
Optional 5 normalization may be applied before enqueueing,
6
with 7, to stabilize attention by emphasizing direction rather than magnitude (Wang et al., 6 Oct 2025).
Learning itself uses standard supervised optimization on the current segment only:
8
where 9 is the retrieved context and 0 is cross-entropy. The model is fully fine-tuned incrementally as new segments arrive. The paper does not introduce EWC, SI, distillation loss, parameter isolation, or rehearsal loss. Old knowledge is therefore preserved by weight continuity across segments and by the latent context buffer, which acts as an implicit memory rather than an explicit replay store (Wang et al., 6 Oct 2025).
The memory overhead is
1
because only 2 latent vectors of dimension 3 are stored. Since 4 is fixed, memory does not grow with the number of segments 5. For a minibatch of size 6, the paper gives the computational cost as
7
which simplifies, with fixed widths, to
8
The dominant added term from context retrieval is the key projection over the 9-entry memory, 0, plus the MLP feed-forward block. Because 1 and 2 are fixed constants, cumulative energy usage grows empirically approximately linearly with the number of segments, unlike cumulative-retraining baselines whose cost increases as historical data accumulate (Wang et al., 6 Oct 2025).
These design choices underwrite the deployment argument. The model is a small MLP plus a single attention-style retrieval block; the memory is fixed-size; no raw-data retention is required; and the method is positioned as an easy-to-deploy alternative to full retraining for tabular data streams (Wang et al., 6 Oct 2025).
4. Evaluation protocol and empirical profile
The empirical study uses 36 classification datasets from the TabZilla benchmark, selected from OpenML. Static tabular datasets are converted into stream-learning problems by partitioning each dataset chronologically into contiguous segments of size 500–1000 instances while preserving order and keeping segment sizes nearly uniform. Preprocessing is uniform across datasets: median imputation and standardization for numerical features, constant imputation and one-hot encoding for categorical features, and stratified 85%–15% train/validation and/or test splits (Wang et al., 6 Oct 2025).
The baseline set includes neural tabular models such as TabPFN v2, TabM, Real-MLP, TabR, ModernNCA, plain MLP, TabNet, DANet, ResNet, STG, and VIME, as well as tree and classical models including XGBoost, LightGBM, CatBoost, SVM, k-NN, Random Forest, Decision Tree, and linear models. These baselines are not evaluated in true streaming continual-learning mode in this study; instead, for predictive-performance fairness, they are retrained from scratch on all cumulative data seen so far at each segment. IMLP, by contrast, operates in true no-replay incremental mode (Wang et al., 6 Oct 2025).
Energy is measured with hardware instrumentation rather than software estimators, specifically an ElmorLabs PMD-USB power meter plus PCI-E slot adapter, sampling CPU and GPU power at millisecond resolution. For segment 3, total energy is
4
including both training and inference power. To summarize the accuracy–energy tradeoff, the paper introduces NetScore-T,
5
Here 6 is segment-level performance, typically balanced accuracy, and 7 is segment-level total energy. Unlike some earlier NetScore-style formulations, this version has no tunable parameters (Wang et al., 6 Oct 2025).
| Model | Balanced accuracy | Energy / NetScore-T |
|---|---|---|
| IMLP | 8 | 9 J / 0 |
| TabNet | 1 | 2 J / 3 |
| TabPFN v2 | 4 | 5 J / 6 |
| MLP | 7 | 8 J / 9 |
The central empirical claim is that IMLP is far more efficient than heavy neural tabular baselines while maintaining competitive, though not state-of-the-art, accuracy. Relative to TabNet, the average energy ratio 0 implies about 1 less energy at essentially the same average balanced accuracy, 2 versus 3. Relative to TabPFN v2, 4, so IMLP is about 5 more energy efficient in the sense of requiring that much less energy on average, with only a 6 drop in balanced accuracy. Relative to a plain MLP, IMLP yields a 7 speedup on average and a 8 reduction in energy usage, while the plain MLP remains slightly more accurate on average, 9 versus 0 (Wang et al., 6 Oct 2025).
NetScore-T favors this tradeoff. IMLP achieves the best average NetScore-T among the neural models at 1, above MLP (2), ModernNCA (3), TabR (4), TabNet (5), and TabPFN v2 (6). In the paper’s Pareto-front analysis, four out of eight Pareto-optimal tradeoffs among neural tabular models come from IMLP. Per-segment plots show balanced accuracy improving across segments like other models, but with cumulative energy that remains the lowest and most stable (Wang et al., 6 Oct 2025).
5. Functional interpretation, operating conditions, and limitations
The paper’s own interpretation is that recent latent prototypes contain information about the evolving data distribution, especially under gradual domain drift or recurring local structure. By retrieving a weighted summary of recent latent states, the model can condition current processing on stream context rather than treating each segment in isolation. Because the memory stores latent summaries rather than raw rows, it may provide some rehearsal-like benefit at lower memory and privacy cost. This suggests that IMLP is particularly well matched to tabular streams in which feature semantics remain fixed while feature distributions drift over time (Wang et al., 6 Oct 2025).
The method is therefore likely to work well when the stream exhibits temporal continuity or moderate domain drift, when recent segments remain informative for current predictions, when edge constraints make replay or full retraining impractical, and when the tabular representation is stable enough that latent summaries remain meaningful. The same discussion implies weaker suitability when dependencies are very long-range and exceed the fixed window 7, when drift is abrupt and recent prototypes become misleading, when feature spaces or schemas change substantially, when richer continual-learning mechanisms are needed to preserve much older knowledge, or when top-end accuracy matters more than energy efficiency (Wang et al., 6 Oct 2025).
Several misconceptions are explicitly precluded by the architectural description. IMLP is not a replay-based continual learner, because it does not store raw historical examples and does not perform explicit rehearsal. It is not a full transformer-style tabular model, because attention is confined to retrieval over a short latent memory rather than applied over tokenized columns or rows. Nor is it a cumulative-retraining system, because it operates in true no-replay incremental mode rather than retraining from scratch on all data seen so far (Wang et al., 6 Oct 2025).
The limitations are equally specific. The paper does not provide a dedicated forgetting metric table or a formal backward-transfer analysis, so evidence for reduced forgetting is indirect, based mainly on maintaining competitive segment-wise performance without replay. It also does not report a comprehensive ablation study in the excerpted material; the authors explicitly identify ablations on window size, feature dimension, scaling, and alternative continual-learning strategies as future work. The benchmark construction relies on segmented TabZilla datasets rather than richer real lifelong settings, and future comparison to real-life benchmarks such as TabRed is proposed rather than reported (Wang et al., 6 Oct 2025).
6. Relation to earlier context-aware and incremental neural formulations
The explicit term IMLP is attached to the tabular-stream continual learner described above, but several earlier lines of work illuminate different meanings of “context-aware” and “incremental” in MLP-adjacent architectures. The most direct conceptual precursor is the probabilistic context-aware formulation in which an output is decomposed into a context-driven transform and a context-free default state. In the CA-NN layer,
8
with gate
9
and the multi-layer CA-RES extension generalizes this gated residual mechanism across depth. This framework provides a strong context-aware, gated, residual-style MLP interpretation, but it does not define an architecture called IMLP and does not address tabular continual learning under energy constraints (Zeng, 2019).
Other neighboring strands are incremental but solve different problems. CINet frames “when to increment the number of contexts” in robot scene modeling as a supervised sequence-to-label problem over LDA-derived context statistics, with a 3-layer LSTM as the best configuration rather than an MLP (Doğan et al., 2017). NADINE is a self-constructing online MLP that adapts width and depth from streaming data via bias–variance criteria, drift detection, adaptive memory, and soft forgetting, making it an implicit context-aware incremental MLP in the sense of drift-sensitive structural adaptation rather than explicit latent context retrieval (Pratama et al., 2019). ICAL and ICAL-Mem use an autoencoder bank to discover recurring hidden contexts in streaming classification and append a context flag to a downstream classifier, but the reported predictor is tree-based CART, not an MLP, and the context detector depends on 0 rather than on features alone (Lore et al., 2019).
Taken together, these works delimit the meaning of the term. In the narrow, paper-specific sense, IMLP refers to a small continual-learning MLP for tabular streams that becomes context-aware by retrieving a learned context vector from a fixed-size FIFO memory of past latent features using scaled dot-product attention, and incremental because it updates sequentially on new segments without replaying old raw data. In the broader architectural sense, the label also evokes a family of context-conditioned feed-forward models in which hidden representations are progressively refined by gating, memory, structural growth, or discovered latent regimes. The 2025 IMLP contribution is distinctive because it combines that context-aware intuition with constant-size latent memory, no raw replay, and explicit energy-aware evaluation for tabular data streams (Wang et al., 6 Oct 2025).