---
title: 'modelDNA: Provenance Analysis for Language Models'
url: https://www.emergentmind.com/topics/modeldna
type: topic
---

# modelDNA: Provenance Analysis for Language Models

modelDNA is a provenance-analysis system for open-weight language models that performs calibrated lineage verification and merge decomposition from sampled weight fingerprints. It is designed for an ecosystem in which lineage is largely self-reported: the `base_model` metadata field on the Hugging Face Hub is optional and unverified, and more than 60% of Hub models reportedly document no parentage at all. Rather than proposing a new fingerprinting theory, modelDNA packages several existing weight-space signals into a practical workflow that fingerprints a suspect model, compares it against a reference database of foundation models, and returns a conservative verdict with a calibrated probability, preferring abstention to confident error [2607.10617].

## 1. Provenance problem and system objective

The system addresses a specific failure mode in open-weight model governance: lineage claims are often incomplete, unverifiable, or disputed, and existing weight-based lineage methods are typically distributed as paper-specific code tied to a single signal and a single experiment. In that setting, provenance disputes about reuse, fine-tuning, merging, or copying tend to be reanalyzed manually. modelDNA is presented as a reproducible alternative that standardizes the workflow into “scan, compare, calibrate, and explain” [2607.10617].

Its target question is narrow but operationally important: whether a repository that claims descent from some base model can be checked independently from its weights, even when metadata is missing, incomplete, or misleading. The design criterion is explicitly conservative. The paper identifies false accusation as the worst-case failure mode in provenance analysis and therefore prioritizes precision, background-calibrated reporting, and abstention over aggressive attribution.

This framing makes modelDNA a forensic tool rather than a training or deployment method. It is passive, static, and CPU-only; it operates on stored checkpoints and associated file formats rather than on model behavior or outputs. A plausible implication is that the tool is intended for public auditability and dispute resolution rather than for online monitoring.

## 2. Fingerprinting from partial checkpoint reads

A central engineering contribution is that modelDNA fingerprints a model without downloading the full checkpoint. For `safetensors`, the file begins with a JSON header containing tensor names, dtypes, shapes, and byte offsets. modelDNA first reads that header, computes the exact byte ranges needed for the fingerprint, and then fetches only those ranges over HTTP, coalesced per shard and read concurrently. For GGUF or llama.cpp-style quantized files, it applies the same principle through a dequantize-and-sample pipeline: sampled element positions are mapped to block-aligned byte ranges, the corresponding blocks are fetched, reference kernels dequantize them, and the desired samples are retained [2607.10617].

For a 7B model, this requires roughly 100–300 MB of ranged reads instead of about 15 GB for a full checkpoint, and the scan takes about two minutes on an ordinary connection. Because the sample positions are determined from tensor identity rather than tensor contents, the system can plan the fetches before any substantial data moves. This separation between structural planning and content retrieval is load-bearing: it is what allows the scan to remain lightweight and reproducible.

The reference side of the comparison is also compact. Candidate foundation models are fingerprinted once and cached in a local database of about 37 indexed bases spanning families such as Llama, Qwen, Mistral, Gemma, Phi, DeepSeek, Yi, Falcon, and OLMo, along with deliberate hard negatives such as OpenLLaMA, which shares Llama-like shapes but was trained independently. Each fingerprint serializes to about 1.5 MB gzipped, so the database can be distributed and used offline [2607.10617].

## 3. Signal families, sampling design, and verdict classes

modelDNA does not rely on a single similarity signal. It ensembles four published signal families, all computed over deterministically chosen sample positions:

| Signal family | Brief description |
|---|---|
| Attention-$\sigma$ curves | Per-layer standard deviations of attention and MLP projection matrices, normalized across depth |
| 1-D vectors | Norms and biases, read in full because they are small |
| PCS samples | Cosine comparisons over seeded element subsamples from 2-D projection tensors |
| Spectra | Top singular values of sampled layers |

The sampling scheme is critical. For each 2-D tensor, a seed is derived from `(global seed, canonical role, layer index)` rather than from the raw tensor name, and that seed drives 32 pseudo-random contiguous blocks of 16,384 elements. Since the positions depend only on `(seed, role, layer, tensor size)`, any two models with matching shapes are sampled at identical element positions. This makes fingerprints directly comparable across models and also underwrites the later merge-decomposition result [2607.10617].

The comparison stage is implemented as a calibrated verdict engine rather than as an unqualified similarity ranking. The classifier is a simple logistic model over four aggregate features: mean $\sigma$-curve correlation, mean vector cosine, mean PCS cosine, and mean spectral correlation. Two calibration choices are emphasized. First, missing features are imputed with unrelated-pair background means, so failure to compute a signal is not treated as positive or negative evidence. Second, coefficients are constrained to be non-negative, because all four features are oriented so that higher values indicate greater similarity [2607.10617].

The system returns one of eight verdict classes:

| Verdict class | Meaning in the system |
|---|---|
| `EXACT_COPY` | Confident copy-level match |
| `QUANTIZED_COPY` | Confident quantized-copy match |
| `FINE_TUNE` | Confident fine-tune attribution |
| `SAME_LINEAGE` | Confident same-lineage attribution |
| `LIKELY_MERGE` | Evidence favors merged ancestry |
| `SAME_FAMILY_UNRESOLVED` | Suggestive same-family evidence, but inconclusive |
| `NO_MATCH` | No candidate clears threshold |
| `INSUFFICIENT` | Comparison is structurally impossible or too incomplete |

The thresholding policy is abstention-first. A positive threshold of `p ≥ 0.90` yields a confident positive class. Scores in the `0.50–0.90` range produce `SAME_FAMILY_UNRESOLVED`. If two positive candidates in different families both look good, the system can emit `LIKELY_MERGE`. If two same-family candidates are within `Δp ≤ 0.03`, the result is treated as a tie rather than forced into a single-parent decision. This is not merely a user-interface choice; it is the mechanism by which the tool distinguishes evidence from uncertainty.

## 4. Benchmark design and lineage-verification results

The main benchmark, LineageBench, is constructed from real Hub models rather than only synthetic derivations. It contains 15 suspect models whose parentage is corroborated by the publishing organization’s own documentation rather than by the metadata field that the tool is meant to audit. The suspects span multiple derivation regimes, including community fine-tunes, continued pretrains, official instruct releases, a merge chain, a depth-upscale, and a GPTQ quantized copy. They are judged against 8 candidate bases, yielding 13 positives and 107 hard negatives [2607.10617].

The paper reports AUROC 1.0, zero false positives at the reporting threshold, and 13/13 correct top-1 parent attributions. It also reports TPR 1.0 at 1% FPR. The weakest positive scored 0.9426, while the strongest negative scored 0.6596, leaving the entire `0.50–0.90` abstention band empty on this dataset. The authors explicitly characterize this as evidence of no errors at this scale, not as a proof that errors are impossible.

Two cases are especially illustrative. SOLAR-10.7B correctly returns `NO_MATCH` because the layer count is incompatible with the candidate bases. The GPTQ copy abstains on packed int4 tensors while still ranking the true parent first on the signals it can compute. These examples clarify that modelDNA’s objective is not universal forced attribution but calibrated judgment under structural constraints.

The benchmark construction also encodes a substantive methodological choice: same-family wrong-parent pairs are intentionally excluded from the hard-negative set because they often share lineage and are therefore not clean negatives. This makes the evaluation more aligned with the actual semantics of lineage verification, where proximity within a family may be real evidence rather than label noise.

## 5. Merge decomposition from fingerprints

The second major contribution is merge decomposition. The paper argues that mainstream weight-merging methods are linear or near-linear per tensor and that, because fingerprint sample positions are deterministic functions of tensor identity, the fingerprint of a merged model is approximately the same linear combination of its parents’ fingerprints. If $y$ denotes the target fingerprint and $x_i$ the candidate-parent fingerprints, the method assumes

$$
y \approx \sum_i \alpha_i x_i .
$$

The mixture weights can then be recovered from fingerprints alone by sum-to-one constrained least squares, without downloading full checkpoints [2607.10617].

The sum-to-one constraint is treated as the key geometric device. Naive regression is ill-conditioned because merge parents typically share a base and are nearly collinear. By choosing a pivot parent $x_0$ and eliminating its coefficient, the problem becomes

$$
y - x_0 \approx \sum_i \beta_i (x_i - x_0),
$$

so the regressors become task-vector differences that are much closer to orthogonal. The paper emphasizes that this is not a cosmetic reformulation: it is what makes the regression numerically stable.

The theoretical justification is matched to contemporary merge practice. Task arithmetic and linear interpolation are linear by definition. Slerp is not exactly linear in general, but for fine-tunes of a shared base with pairwise cosines around 0.99 or above, the small-angle approximation makes it effectively coincide with linear interpolation up to a small quadratic correction. TIES and DARE break exact per-element linearity through sign selection and sparsification, but are described as approximately linear in expectation, with residuals absorbing the deviations.

Empirically, the results are unusually concrete. For the published slerp merge NeuralPipe-7B-slerp, modelDNA recovers pooled weights of 0.644 / 0.356 with reconstruction cosine 0.99996, removing 78% of the residual left by the best single parent. At the per-layer level, the fit reconstructs the published interpolation curves with `r = 0.999` for attention and `r = 0.970` for MLP, and the largest error is 0.048 in attention. For the DARE-style merge Monarch-7B, the recovered weights are 0.371 / 0.347 / 0.291 against published weights 0.36 / 0.34 / 0.30, with maximum error 0.011 and cross-role spread below 0.01; the base column is near zero, around −0.01 [2607.10617].

The method is careful not to overstate these reconstructions. Monarch does not always trigger an explicit merge verdict: despite recovering the published weights closely, it receives `AMBIGUOUS` because residual shrink falls below the merge threshold, which is intentionally conservative to avoid false merge calls on ordinary sibling fine-tunes. AlphaMonarch-7B supplies the complementary negative case. It is a chain-closure case rather than a merge, and when decomposed against its immediate ancestor plus decoys, the fit places essentially all mass on the true parent and returns `SINGLE_PARENT`. This establishes that decomposition answers a conditional question—how the model decomposes over the supplied candidates—rather than reconstructing the full historical tree automatically.

## 6. Public atlas, significance, and limitations

modelDNA is intended to make provenance analysis reproducible and public. The paper states that all fingerprints, benchmark caches, and the inferred lineage graph of 55 models are public and reproducible offline. It also reports a public atlas of 55 models, 475 depth-compatible pairwise comparisons, and 116 edges, each carrying its evidence [2607.10617].

This reproducibility claim is part of the system’s significance. Provenance disputes in the open-weight ecosystem are often less about whether a signal exists than about whether the analysis can be rerun independently. modelDNA is designed so that a third party can repeat the scan from a laptop, without privileged access and without downloading every candidate checkpoint in full. This suggests a shift from ad hoc authorship disputes toward auditable, cached, and shareable lineage evidence.

The paper is equally explicit about scope boundaries. modelDNA does not solve distillation, which is invisible to weight-space methods by construction. It does not claim adversarial completeness, and deliberate reparameterizations can defeat some direct similarity signals. Probability magnitudes are less validated than rankings because the calibrator was fit on synthetic data and the real benchmark is too small to fully validate calibration curves. These caveats are not peripheral; they define the methodological stance of the system.

In that sense, modelDNA’s distinctive feature is not maximal attribution aggressiveness but operational conservatism. It combines partial-checkpoint fingerprinting, multi-signal comparison, calibrated abstention, and constrained merge decomposition into a single provenance workflow. Within its stated limits, it provides a practical mechanism for verifying lineage claims and, in some cases, reconstructing merge recipes from fingerprints alone [2607.10617].

Source: https://www.emergentmind.com/topics/modeldna