Papers
Topics
Authors
Recent
Search
2000 character limit reached

traceSDD: Spec-Driven LLM Code Generation

Updated 5 July 2026
  • The paper demonstrates that mandatory inline citations in traceSDD reduce output determinism while achieving nearly 87% automated hallucination detection with 0% false positives.
  • It employs rigorous empirical studies comparing traceSDD with uncited and alternative frameworks using metrics like Lexical Similarity Score (LSS) and True Detection Rate (TDR).
  • The framework embeds hierarchical REQ-XXX.Y.Z identifiers directly into code, ensuring precise traceability and operationally coupling verification to the implementation.

traceSDD is a Spec-Driven Development (SDD) framework for LLM-powered code generation that enforces mandatory per-line requirement citations using hierarchical REQ-XXX.Y.Z identifiers. In the cited condition, every nontrivial source-code line must carry an inline comment indicating exactly which atomic requirement it implements; if a cited identifier does not exist in the specification, the resulting “orphan” requisition is automatically flagged as a hallucination (Panda, 28 Jun 2026). The framework was evaluated in a controlled empirical study against traceSDD_uncited, [Spec Kit](https://www.emergentmind.com/topics/spec-kit), and [OpenSpec](https://www.emergentmind.com/topics/openspec), with the central finding that citation annotations trade determinism for verifiability: mandatory citations reduce output determinism but uniquely enable automated hallucination detection with nonzero detection rates and zero false positive rate across the reported studies (Panda, 28 Jun 2026).

1. Formal specification and citation discipline

traceSDD adopts a three-tier hierarchical requirement numbering scheme:

  • Tier 1: high-level feature groups (REQ-XXX)
  • Tier 2: subfeatures (REQ-XXX.Y)
  • Tier 3: atomic requirement points (REQ-XXX.Y.Z)

A traceSDD specification consists of a structured tree of REQ-XXX.Y.Z entries. Under the framework’s cited condition, an LLM must attach an inline comment to every nontrivial source-code line of the form # [REQ-XXX.Y.Z], thereby establishing line-level traceability between the specification and the generated implementation (Panda, 28 Jun 2026).

The paper gives a representative example in which a function signature is annotated as follows:

1
def validate_email(email: str) -> bool:  # [REQ-012.3.1]

This citation discipline is stricter than artifact-level traceability. Spec Kit uses user stories and acceptance criteria, while OpenSpec relies on post-hoc external trace maps in YAML. By contrast, traceSDD localizes traceability directly in the generated code, and traceSDD_uncited isolates the effect of the citation mechanism by supplying the same REQ-XXX.Y.Z specification without enforcing inline comments.

A key operational property is the treatment of nonexistent identifiers. If the model cites an ID not present in the specification, that orphan citation is automatically flagged as a hallucination. This makes the identifier space itself part of the verification surface rather than merely a documentation aid.

2. Metrics and operational definitions

The empirical study evaluates traceSDD using two primary outcomes: output determinism and automated hallucination detection rate (Panda, 28 Jun 2026).

Output determinism is measured by the Lexical Similarity Score (LSS). For a given task and condition, three independent cold-start outputs, Run₁, Run₂, and Run₃, are generated. Comments are stripped and whitespace is normalized, after which pairwise Levenshtein Set Similarity is computed via Python’s difflib.SequenceMatcher.ratio(). The task-level score is defined as

1
\mathrm{LSS}_i = \mathrm{mean}\{ \mathrm{LSM}(\mathrm{Run}_1,\mathrm{Run}_2), \mathrm{LSM}(\mathrm{Run}_1,\mathrm{Run}_3) \}

with

1
\mathrm{LSM}(A,B) = \frac{\text{length of matching blocks between A and B}}{\max(|A|,|B|)}.

The grand-mean score for a condition is

1
\overline{\mathrm{LSS}} = \frac{1}{N}\sum_{i=1}^{N}\mathrm{LSS}_i.

Automated hallucination detection is measured by the detection rate [TDR](https://www.emergentmind.com/topics/tdr). The study injects exactly three types of hallucinations per task—H-SCOPE, H-PRIOR, and H-OVER—and in traceSDD_cited runs these are cited with fake IDs such as REQ-099.1.1. The rate is defined as

1
2
3
4
\mathrm{TDR} =
\frac{|\{\text{injected lines detected via orphan-REQ}\}|}
     {|\{\text{injected hallucinated lines}\}|}
\times 100\%.

The study also records FPR, the fraction of non-hallucinated code lines in correct implementations that nevertheless cite an ID absent from the specification. Pairwise LSS comparisons use paired two-tailed tt-tests with α=0.05\alpha=0.05 and Bonferroni–Holm correction, backed up by Wilcoxon signed-rank tests; effect sizes are reported as Cohen’s dd.

These definitions are notable because they separate two distinct properties of generated code. LSS quantifies reproducibility across independent sessions after comment removal, whereas TDR quantifies whether hallucinated requirements can be automatically detected from the citation structure itself.

3. Comparative framework design

The study compares four conditions across two LLMs. The differences concern how traceability is imposed and where citations appear.

Condition Specification basis Traceability mechanism
traceSDD (cited) REQ-XXX.Y.Z tree Inline per-line citations
traceSDD_uncited Same REQ-XXX.Y.Z tree No inline comments
Spec Kit User stories FR-XXX Artifact-level traceability
OpenSpec External trace map YAML trace map

The distinction between traceSDD and traceSDD_uncited is analytically central. Both conditions use the same structured requirement format, but only the cited condition mandates inline comments. The paper therefore treats their comparison as a citation-isolation test: it measures the effect of citation discipline itself rather than the effect of structured requirements in general (Panda, 28 Jun 2026).

This comparison also clarifies the scope of traceability. In Spec Kit, traceability resides at the artifact level through user stories and acceptance criteria. In OpenSpec, it is externalized into YAML trace maps. In traceSDD, traceability is embedded into the implementation line by line. A plausible implication is that the framework shifts verification from external reconciliation to local syntactic inspection of source-code annotations.

4. Experimental configuration

The evaluation comprises two controlled studies on frontier LLMs (Panda, 28 Jun 2026).

In Study 1, Claude Sonnet 4.6 is tested on 20 benchmark Python tasks: 15 single-file tasks of approximately 50–100 LOC and 5 multi-file tasks of approximately 600–800 LOC. In Study 2, GLM-5-turbo is tested on 50 novel tasks spanning 8 domains, 3 difficulty levels, and 2 size classes, specifically 37 small and 13 large tasks.

Each model is evaluated under four conditions:

  • A. traceSDD (cited)
  • B. traceSDD_uncited
  • C. Spec Kit
  • D. OpenSpec

For every model–task–condition combination, the study uses 3 independent cold-start sessions. This yields:

  • Claude: 20 tasks × 4 conds × 3 runs = 240 implementations
  • GLM: 50 tasks × 4 conds × 3 runs = 600 implementations

Metrics collected per implementation are LSS, TDR, True Positive Rate (TPR on correctness tests), and FPR.

The design is explicitly cross-model and replicated. The paper characterizes Claude Sonnet 4.6 and GLM-5-turbo as architecturally distinct LLMs and uses this replication to test whether the observed trade-off is specific to a particular model family or attributable to the citation discipline itself.

5. Determinism results

The mean LSS values reported for the four conditions establish a consistent ordering across both studies (Panda, 28 Jun 2026).

Model Condition Mean LSS
Claude Sonnet 4.6 traceSDD (cited) 0.535 (sd=0.167)
Claude Sonnet 4.6 traceSDD (uncited) 0.745 (sd=0.194)
Claude Sonnet 4.6 Spec Kit 0.460 (sd=0.221)
Claude Sonnet 4.6 OpenSpec 0.487 (sd=0.248)
GLM-5-turbo traceSDD (cited) 0.510 (sd=0.147)
GLM-5-turbo traceSDD (uncited) 0.644 (sd=0.160)
GLM-5-turbo Spec Kit 0.434 (sd=0.174)
GLM-5-turbo OpenSpec 0.480 (sd=0.160)

The citation-isolation comparison is statistically significant in both studies. For Claude, traceSDD cited versus uncited yields t=-3.26, p=0.004, and Cohen’s d=-0.73, with Wilcoxon W=56.0, p=0.001, r=0.50. For GLM, the same comparison yields t=-5.09, p\<0.001, d=-0.72, with Wilcoxon W=211.0, p\<0.001, r=0.58. The paper summarizes these nearly identical effect sizes as establishing that mandatory citations consistently reduce determinism.

The framework also outperforms Spec Kit on determinism in the cited condition: for Claude, d=0.47, p=0.049; for GLM, d=0.42, p=0.003. By contrast, the difference between traceSDD cited and OpenSpec is non-significant in both studies, with Claude p=0.44, GLM p=0.32, and effect sizes reported as approximately 0.16–0.18.

The reported mechanism for the determinism penalty is “citation-placement variability.” At each session, the model must decide which lines to annotate and how many identifiers to place on a line. The paper states that this injects lexical noise, lowering LSS by approximately 0.21 for Claude and 0.13 for GLM relative to the uncited condition. Because LSS is computed after comment stripping and whitespace normalization, this observation suggests that citation discipline affects not only superficial formatting but also the structure of the generated implementation.

6. Verifiability, hallucination detection, and interpretation

The central empirical result is that only traceSDD in the cited condition achieves nonzero automated hallucination detection (Panda, 28 Jun 2026). For Claude, TDR=86.4% with FPR=0%; for GLM, TDR=88.0% with FPR=0%. All other conditions—traceSDD_uncited, Spec Kit, and OpenSpec—score TDR=0%.

The Claude study additionally reports 100% for each H-Scope/Prior/Over within the cited condition. The paper also states that True Positive Rate (TPR) for functionality was 100 % in every condition (H3). Accordingly, the verifiability advantage of traceSDD does not coincide with a reported loss in correctness-test pass rate within the experimental setup.

These results motivate the paper’s formulation of a determinism–verifiability trade-off. Mandatory inline citations impose a reproducible determinism penalty, but they uniquely enable “an automated, language-agnostic set-difference check that flags orphan REQ IDs.” The study summarizes this capability as yielding TDR≈87 % with zero false alarms.

A common misconception would be to treat traceability annotations as merely documentary. The reported results contradict that interpretation in this setting: the annotations are operationally coupled to an automated detection mechanism. Another possible misconception would be that any structured specification should suffice for automated hallucination detection. The comparison shows otherwise, since traceSDD_uncited uses the same REQ-XXX.Y.Z specification yet still yields TDR=0%. This indicates that the decisive factor is not only structured requirements, but the enforced embedding of identifiers into the code.

The paper further states that, in regulated or safety-critical domains, the moderate determinism cost is outweighed by near-perfect hallucination detection, whereas for rapid prototyping one may disable citations to maximize consistency while retaining a structured REQ-format spec anchor. This suggests a deployment choice rather than a universal optimum: traceSDD can be used either as a verifiability-first workflow in its cited form or as a consistency-oriented workflow in its uncited variant, depending on whether line-level audibility or higher LSS is the governing constraint.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to traceSDD.