RepMatch: Dataset Analysis Method
- RepMatch is a dataset-analysis method that quantifies similarity between training subsets by comparing task-specific adaptations in a shared pretrained model.
- It employs low-rank LoRA updates and geometric Grassmann similarity to enable robust dataset-level and instance-level comparisons.
- The method supports representative subset selection and dataset bias diagnosis by ranking examples based on induced adaptation similarity.
Searching arXiv for the specified paper and a small set of closely related methods mentioned in the provided data. arxiv_search(query="(Modarres et al., 2024)", max_results=5, sort_by="submittedDate") arxiv_search(query="Dataset Cartography arXiv", max_results=5, sort_by="relevance") RepMatch is a dataset-analysis method that quantifies the similarity between subsets of training instances by comparing the representation spaces learned by models trained on those subsets. It is formulated for settings in which a subset of data—ranging from a single instance to an entire dataset—is represented by the task-specific adaptation it induces in a shared pretrained model. By comparing those induced adaptations rather than individual-example statistics or fixed encoder features, RepMatch supports both dataset-to-dataset and instance-to-dataset analysis, including cross-dataset comparisons, representative-subset selection, and diagnosis of challenge-set heuristics (Modarres et al., 2024).
1. Conceptual basis and scope
RepMatch is motivated by a limitation in prevailing NLP data-analysis methods: many characterize individual examples as easy, hard, noisy, ambiguous, or atypical, and usually operate only within a single dataset. RepMatch is designed to move beyond that regime. Its central premise is that two subsets of training data are similar if fine-tuning on them induces similar task-specific adaptations in the same pretrained model (Modarres et al., 2024).
This yields a model-centered notion of similarity. A subset is not compared through its raw tokens, labels, or metadata alone, but through the knowledge encoded in the model after training on it. Because the pretrained weights are shared, identical, and frozen, the comparison is localized to the learned adaptation matrices. In the paper’s formulation, this makes it possible to compare arbitrary subsets: full datasets, selected subsets, or single instances treated as subsets of size one (Modarres et al., 2024).
The framework is also designed to avoid additional auxiliary infrastructure beyond standard fine-tuning with LoRA. This is significant for its intended use as an analysis method rather than as a separate modeling pipeline. RepMatch therefore occupies a distinct position relative to methods that rely on fixed pretrained embeddings, extra feature extractors, or separate auxiliary models.
2. Mathematical formulation and similarity score
Let and denote two subsets of data, and let a pretrained model be fine-tuned separately on each subset to obtain and . For a pretrained weight matrix in layer , fine-tuning produces an updated matrix
RepMatch uses LoRA to constrain the update to low rank:
where and 0, so the update has rank at most 1. This reduces parameter count from 2 to 3 (Modarres et al., 2024).
The comparison is geometric rather than elementwise. RepMatch compares the subspaces induced by two low-rank updates using Grassmann similarity. If the two matrices are 4 and 5, and 6 and 7 are the singular-vector matrices obtained from SVD, then the similarity between their 8-dimensional and 9-dimensional subspaces is
0
For a given layer, this produces an 1 matrix of pairwise subspace similarities. The layer-level RepMatch score is defined as the maximum entry of that matrix, and the model-level score is the average across layers:
2
where 3 is the number of layers compared (Modarres et al., 2024).
In the implementation reported in the paper, the value matrices in the attention block were found to be sufficient, and many experiments used rank 4 because the strongest similarity signal often appeared in a single dominant direction. Accordingly, RepMatch is not a direct comparison of raw parameter values; it is a comparison of the geometry of learned low-rank update subspaces.
3. Analytical modes: dataset-level and instance-level
RepMatch supports two principal modes of analysis. In dataset-level analysis, the same pretrained model is fine-tuned on two datasets and the resulting LoRA updates are compared. High RepMatch indicates that the datasets induce similar learned representations and likely share task structure. The paper illustrates this with comparisons among SST-2, SST-5, IMDB, MNLI, and SNLI, where SST-2 and SST-5 exhibit higher similarity than SST-2 and entailment datasets (Modarres et al., 2024).
In instance-level analysis, an individual example 5 is treated as a subset of size one. The model is fine-tuned on that single instance, and the resulting update is compared with the update learned from the full dataset. Instances with higher RepMatch to the full dataset are interpreted as more representative. This turns representativeness into a ranking problem defined by induced model adaptation rather than by annotation heuristics or static embedding proximity (Modarres et al., 2024).
This formulation also permits identification of outliers or potentially out-of-distribution examples. If a single instance induces an update that is weakly aligned with the update induced by the dataset as a whole, the instance can be interpreted as less representative of the dataset’s dominant task signal. This suggests a use for RepMatch in auditing dataset composition, especially where sample-level statistics do not capture subset-level task structure.
4. Experimental configuration and robustness properties
The reported evaluation spans three NLP task families: sentiment analysis on SST-2, SST-5, and IMDB; natural language inference on MNLI and SNLI; and question answering on SQuAD v1. The primary model is BERT6, with additional experiments on ELECTRA7 and LLaMA2-7B to assess architectural generality. Fine-tuning generally applies LoRA only to the value matrices while freezing all other weights (Modarres et al., 2024).
Default settings include batch size 40, 10 epochs for sentiment tasks and 5 epochs for the others, rank 8, and learning rate 9 for dataset-level analysis and 0 for instance-level experiments. Some experiments are repeated at rank 1 to examine the effect of increased rank, which the paper notes generally improves instance-selection results at higher computational cost (Modarres et al., 2024).
A central empirical result is robustness to random seeds and separation from random baselines. When two models are fine-tuned on the same dataset with different seeds, their LoRA updates remain highly similar, with average RepMatch scores above 0.7 in the SST-2 dataset-level experiment, whereas a shuffled-matrix baseline scores below 0.02. When the same instance is trained with different seeds, RepMatch remains above 0.6, while different instances from the same dataset produce much lower similarity, around 0.11 (Modarres et al., 2024).
The paper also compares RepMatch with a representation-based baseline using cosine similarity of [CLS] embeddings. RepMatch is reported as more discriminative and less susceptible to superficial structural similarity. The example given is that SNLI and STS-B may appear similar in [CLS] space because both are sentence-pair tasks, whereas RepMatch assigns them low similarity because the learned task knowledge differs (Modarres et al., 2024).
5. Representative subset selection and dataset diagnosis
A major application developed in the paper is representative subset selection. The procedure ranks instances by their similarity to the full dataset and selects the top 100 most representative examples. Training on these subsets consistently yields better accuracy or F1 than training on a random 100-example subset across SST-2, SST-5, IMDB, MNLI, SNLI, and SQuAD (Modarres et al., 2024).
The SST-2 learning-curve experiment refines this result. For subset sizes below about 400, RepMatch-selected subsets consistently outperform random subsets of equal size, although the performance gap narrows as the subset grows. The paper also reports that changing the LoRA rank from 1 to 4 generally improves performance, though not uniformly. In imbalanced tasks such as SST-5, the highest-ranked examples may skew toward a dominant label class (Modarres et al., 2024).
These results position RepMatch as an analysis tool with operational consequences: it does not merely assign a similarity score, but can induce a ranking over instances that correlates with downstream training utility. A plausible implication is that representativeness, as defined by induced low-rank update similarity, can function as a principled criterion for subset construction in low-budget training regimes.
6. Challenge datasets, related methods, and limitations
RepMatch is also used to uncover heuristics in challenge datasets. The paper studies HANS, an NLI challenge set intended to expose shallow overlap heuristics in MNLI and SNLI. It extracts subsets of MNLI and SNLI with different premise-hypothesis overlap levels but all labeled non-entailment, and compares them to HANS. The subset with full overlap and non-entailment labels has the highest RepMatch to HANS, matching the intuition that HANS was built to counter overlap bias in standard NLI datasets (Modarres et al., 2024).
This use case is broader than ordinary similarity scoring. RepMatch can identify not only whether two datasets are similar, but which region of a source dataset is most aligned with the artifacts or heuristics targeted by a challenge set. In that sense, it can help diagnose dataset construction biases and out-of-distribution vulnerabilities.
Relative to prior dataset-analysis techniques, the paper emphasizes several contrasts. Unlike training-dynamics methods such as Dataset Cartography or metadata archaeology, RepMatch is not limited to individual-example statistics such as confidence or margin trajectories. Unlike representation-only baselines, it measures what the model learned from the data rather than how examples appear under a fixed pretrained encoder. Unlike SimEx, it does not require separate autoencoder training or additional auxiliary models. Unlike coreset selection or Grad-Match, it is not primarily aimed at training efficiency; its purpose is comparative analysis and interpretation (Modarres et al., 2024).
The principal tradeoff is computational cost. RepMatch requires fine-tuning separate models on each subset, making it more expensive than instance-level statistics, especially when the number of subsets is large or the LoRA rank is increased. The paper also notes that the method does not model interactions among instances within a batch during instance-level selection, and that training randomness and batch composition can still affect results. Most experiments are on Transformer text models, so broader validation beyond that setting remains future work (Modarres et al., 2024).