---
title: 'traceSDD: Spec-Driven LLM Code Generation'
url: https://www.emergentmind.com/topics/tracesdd
type: topic
---

# traceSDD: Spec-Driven LLM Code Generation

traceSDD is a Spec-Driven Development (SDD) framework for LLM-powered code generation that enforces mandatory per-line requirement citations using hierarchical `REQ-XXX.Y.Z` identifiers. In the cited condition, every nontrivial source-code line must carry an inline comment indicating exactly which atomic requirement it implements; if a cited identifier does not exist in the specification, the resulting “orphan” requisition is automatically flagged as a hallucination [2606.30689]. The framework was evaluated in a controlled empirical study against `traceSDD_uncited`, `Spec Kit`, and `OpenSpec`, with the central finding that citation annotations trade determinism for verifiability: mandatory citations reduce output determinism but uniquely enable automated hallucination detection with nonzero detection rates and zero false positive rate across the reported studies [2606.30689].

## 1. Formal specification and citation discipline

traceSDD adopts a three-tier hierarchical requirement numbering scheme:

- Tier 1: high-level feature groups (`REQ-XXX`)
- Tier 2: subfeatures (`REQ-XXX.Y`)
- Tier 3: atomic requirement points (`REQ-XXX.Y.Z`)

A traceSDD specification consists of a structured tree of `REQ-XXX.Y.Z` entries. Under the framework’s cited condition, an LLM must attach an inline comment to every nontrivial source-code line of the form `# [REQ-XXX.Y.Z]`, thereby establishing line-level traceability between the specification and the generated implementation [2606.30689].

The paper gives a representative example in which a function signature is annotated as follows:

```python
def validate_email(email: str) -> bool:  # [REQ-012.3.1]
```

This citation discipline is stricter than artifact-level traceability. `Spec Kit` uses user stories and acceptance criteria, while `OpenSpec` relies on post-hoc external trace maps in YAML. By contrast, traceSDD localizes traceability directly in the generated code, and `traceSDD_uncited` isolates the effect of the citation mechanism by supplying the same `REQ-XXX.Y.Z` specification without enforcing inline comments.

A key operational property is the treatment of nonexistent identifiers. If the model cites an ID not present in the specification, that orphan citation is automatically flagged as a hallucination. This makes the identifier space itself part of the verification surface rather than merely a documentation aid.

## 2. Metrics and operational definitions

The empirical study evaluates traceSDD using two primary outcomes: output determinism and automated hallucination detection rate [2606.30689].

Output determinism is measured by the Lexical Similarity Score (LSS). For a given task and condition, three independent cold-start outputs, `Run₁`, `Run₂`, and `Run₃`, are generated. Comments are stripped and whitespace is normalized, after which pairwise Levenshtein Set Similarity is computed via Python’s `difflib.SequenceMatcher.ratio()`. The task-level score is defined as

```latex
\mathrm{LSS}_i = \mathrm{mean}\{ \mathrm{LSM}(\mathrm{Run}_1,\mathrm{Run}_2), \mathrm{LSM}(\mathrm{Run}_1,\mathrm{Run}_3) \}
```

with

```latex
\mathrm{LSM}(A,B) = \frac{\text{length of matching blocks between A and B}}{\max(|A|,|B|)}.
```

The grand-mean score for a condition is

```latex
\overline{\mathrm{LSS}} = \frac{1}{N}\sum_{i=1}^{N}\mathrm{LSS}_i.
```

Automated hallucination detection is measured by the detection rate `TDR`. The study injects exactly three types of hallucinations per task—`H-SCOPE`, `H-PRIOR`, and `H-OVER`—and in `traceSDD_cited` runs these are cited with fake IDs such as `REQ-099.1.1`. The rate is defined as

```latex
\mathrm{TDR} =
\frac{|\{\text{injected lines detected via orphan-REQ}\}|}
     {|\{\text{injected hallucinated lines}\}|}
\times 100\%.
```

The study also records `FPR`, the fraction of non-hallucinated code lines in correct implementations that nevertheless cite an ID absent from the specification. Pairwise LSS comparisons use paired two-tailed $t$-tests with $\alpha=0.05$ and Bonferroni–Holm correction, backed up by Wilcoxon signed-rank tests; effect sizes are reported as Cohen’s $d$.

These definitions are notable because they separate two distinct properties of generated code. LSS quantifies reproducibility across independent sessions after comment removal, whereas TDR quantifies whether hallucinated requirements can be automatically detected from the citation structure itself.

## 3. Comparative framework design

The study compares four conditions across two LLMs. The differences concern how traceability is imposed and where citations appear.

| Condition | Specification basis | Traceability mechanism |
|---|---|---|
| traceSDD (cited) | `REQ-XXX.Y.Z` tree | Inline per-line citations |
| traceSDD_uncited | Same `REQ-XXX.Y.Z` tree | No inline comments |
| Spec Kit | User stories `FR-XXX` | Artifact-level traceability |
| OpenSpec | External trace map | YAML trace map |

The distinction between `traceSDD` and `traceSDD_uncited` is analytically central. Both conditions use the same structured requirement format, but only the cited condition mandates inline comments. The paper therefore treats their comparison as a citation-isolation test: it measures the effect of citation discipline itself rather than the effect of structured requirements in general [2606.30689].

This comparison also clarifies the scope of traceability. In `Spec Kit`, traceability resides at the artifact level through user stories and acceptance criteria. In `OpenSpec`, it is externalized into YAML trace maps. In traceSDD, traceability is embedded into the implementation line by line. A plausible implication is that the framework shifts verification from external reconciliation to local syntactic inspection of source-code annotations.

## 4. Experimental configuration

The evaluation comprises two controlled studies on frontier LLMs [2606.30689].

In Study 1, Claude Sonnet 4.6 is tested on 20 benchmark Python tasks: 15 single-file tasks of approximately `50–100 LOC` and 5 multi-file tasks of approximately `600–800 LOC`. In Study 2, GLM-5-turbo is tested on 50 novel tasks spanning 8 domains, 3 difficulty levels, and 2 size classes, specifically 37 small and 13 large tasks.

Each model is evaluated under four conditions:

- `A.` traceSDD (cited)
- `B.` traceSDD_uncited
- `C.` Spec Kit
- `D.` OpenSpec

For every model–task–condition combination, the study uses 3 independent cold-start sessions. This yields:

- Claude: `20 tasks × 4 conds × 3 runs = 240 implementations`
- GLM: `50 tasks × 4 conds × 3 runs = 600 implementations`

Metrics collected per implementation are `LSS`, `TDR`, `True Positive Rate (TPR on correctness tests)`, and `FPR`.

The design is explicitly cross-model and replicated. The paper characterizes Claude Sonnet 4.6 and GLM-5-turbo as architecturally distinct LLMs and uses this replication to test whether the observed trade-off is specific to a particular model family or attributable to the citation discipline itself.

## 5. Determinism results

The mean LSS values reported for the four conditions establish a consistent ordering across both studies [2606.30689].

| Model | Condition | Mean LSS |
|---|---|---|
| Claude Sonnet 4.6 | traceSDD (cited) | `0.535 (sd=0.167)` |
| Claude Sonnet 4.6 | traceSDD (uncited) | `0.745 (sd=0.194)` |
| Claude Sonnet 4.6 | Spec Kit | `0.460 (sd=0.221)` |
| Claude Sonnet 4.6 | OpenSpec | `0.487 (sd=0.248)` |
| GLM-5-turbo | traceSDD (cited) | `0.510 (sd=0.147)` |
| GLM-5-turbo | traceSDD (uncited) | `0.644 (sd=0.160)` |
| GLM-5-turbo | Spec Kit | `0.434 (sd=0.174)` |
| GLM-5-turbo | OpenSpec | `0.480 (sd=0.160)` |

The citation-isolation comparison is statistically significant in both studies. For Claude, traceSDD cited versus uncited yields `t=-3.26`, `p=0.004`, and `Cohen’s d=-0.73`, with Wilcoxon `W=56.0`, `p=0.001`, `r=0.50`. For GLM, the same comparison yields `t=-5.09`, `p<0.001`, `d=-0.72`, with Wilcoxon `W=211.0`, `p<0.001`, `r=0.58`. The paper summarizes these nearly identical effect sizes as establishing that mandatory citations consistently reduce determinism.

The framework also outperforms `Spec Kit` on determinism in the cited condition: for Claude, `d=0.47`, `p=0.049`; for GLM, `d=0.42`, `p=0.003`. By contrast, the difference between traceSDD cited and `OpenSpec` is non-significant in both studies, with `Claude p=0.44`, `GLM p=0.32`, and effect sizes reported as approximately `0.16–0.18`.

The reported mechanism for the determinism penalty is “citation-placement variability.” At each session, the model must decide which lines to annotate and how many identifiers to place on a line. The paper states that this injects lexical noise, lowering LSS by approximately `0.21` for Claude and `0.13` for GLM relative to the uncited condition. Because LSS is computed after comment stripping and whitespace normalization, this observation suggests that citation discipline affects not only superficial formatting but also the structure of the generated implementation.

## 6. Verifiability, hallucination detection, and interpretation

The central empirical result is that only traceSDD in the cited condition achieves nonzero automated hallucination detection [2606.30689]. For Claude, `TDR=86.4%` with `FPR=0%`; for GLM, `TDR=88.0%` with `FPR=0%`. All other conditions—`traceSDD_uncited`, `Spec Kit`, and `OpenSpec`—score `TDR=0%`.

The Claude study additionally reports `100% for each H-Scope/Prior/Over` within the cited condition. The paper also states that `True Positive Rate (TPR) for functionality was 100 % in every condition (H3)`. Accordingly, the verifiability advantage of traceSDD does not coincide with a reported loss in correctness-test pass rate within the experimental setup.

These results motivate the paper’s formulation of a determinism–verifiability trade-off. Mandatory inline citations impose a reproducible determinism penalty, but they uniquely enable “an automated, language-agnostic set-difference check that flags orphan REQ IDs.” The study summarizes this capability as yielding `TDR≈87 %` with zero false alarms.

A common misconception would be to treat traceability annotations as merely documentary. The reported results contradict that interpretation in this setting: the annotations are operationally coupled to an automated detection mechanism. Another possible misconception would be that any structured specification should suffice for automated hallucination detection. The comparison shows otherwise, since `traceSDD_uncited` uses the same `REQ-XXX.Y.Z` specification yet still yields `TDR=0%`. This indicates that the decisive factor is not only structured requirements, but the enforced embedding of identifiers into the code.

The paper further states that, in regulated or safety-critical domains, the moderate determinism cost is outweighed by near-perfect hallucination detection, whereas for rapid prototyping one may disable citations to maximize consistency while retaining a structured REQ-format spec anchor. This suggests a deployment choice rather than a universal optimum: traceSDD can be used either as a verifiability-first workflow in its cited form or as a consistency-oriented workflow in its uncited variant, depending on whether line-level audibility or higher LSS is the governing constraint.

Source: https://www.emergentmind.com/topics/tracesdd