UserTrace: Automated Requirements Traceability
- UserTrace is a multi-agent system that bridges user goals with code elements by generating live trace links from repository artifacts.
- It constructs both component and file dependency graphs through static AST analysis to synthesize user-level and implementation-level requirements.
- Its workflow leverages specialized agents for code review, domain search, writing, and verification to improve trace-link recovery and validate AI-generated software.
UserTrace is a multi-agent system for automatically generating user-level requirements and recovering live trace links from software repositories. In the system’s terminology, the recovered chain is from user-level requirements to implementation-level requirements to code, with the stated aim of supporting user understanding, repository maintainability, and validation of whether AI-generated software aligns with user intent. It addresses two gaps identified in prior work: automated code summarization has mainly produced implementation-level, developer-oriented descriptions for fine-grained code units, and requirements traceability methods have often neglected project evolution (Jin et al., 14 Sep 2025).
1. Definition and scope
UserTrace distinguishes three artifacts. User-level requirements are high-level descriptions of what a system must accomplish from the end-user perspective; implementation-level requirements are developer-oriented descriptions derived from program behavior and logic; and live trace links are explicit connections of the form URs → IRs → code that can be regenerated or updated when the repository changes. The paper formalizes these notions as
for user-level requirements, where denotes user goals, contextual constraints, and domain-specific business rules, and
for implementation-level requirements, where denotes program logic and contextual constraints (Jin et al., 14 Sep 2025).
The motivating contrast is between implementation-facing and user-facing abstraction. An implementation-level requirement such as “checkPassword() compares user input against the stored hash in the DB” remains close to program logic, whereas a user-level requirement such as “The system shall allow users to securely log into their accounts by verifying their credentials to access personal services” abstracts multiple code behaviors into a higher-level goal (Jin et al., 14 Sep 2025).
Despite its name, UserTrace is not a runtime tracing framework. It operates over repository structure and repository artifacts rather than execution logs. This is materially different from systems such as PreciseTracer, which reconstructs per-request causal paths in multi-tier black-box services from SEND and RECEIVE activities (Sang et al., 2010), and from model-based trace-checking, which replays execution traces against formal models using tools such as Spin and Pro-B (Howard et al., 2011). In UserTrace, “traceability” refers to trace links across requirements abstractions and code, not event streams emitted by a running program.
2. Repository representation and trace-link model
UserTrace begins with static analysis of the repository and constructs a dual-level dependency representation. At the component level, it builds a Component Dependency Graph
where is the set of components with attributes including type, file path, depends_on, and code, and contains directed dependencies. At the file level, it builds a File Dependency Graph
0
where 1 is the set of files and 2 contains inter-file dependencies induced by component relations. A mapping
3
assigns each component to its file (Jin et al., 14 Sep 2025).
The system extracts these graphs from AST parsing. For Java, the parsed units are class, interface, enum, and method; for Python, they are class, function, and method. Because repository dependency graphs can contain cycles, UserTrace computes strongly connected components and breaks cycles by removing minimal edges to obtain DAGs suitable for topological processing (Jin et al., 14 Sep 2025).
The trace-link model is constructive rather than retrieval-only. Component-level IRs are first generated from code and dependency context, then aggregated to file-level IRs, and then used to synthesize URs over graph-derived file communities. As a result, UR → IR links are attached to the file-level IRs used during UR synthesis, and IR → code links are inherited from the aggregation path from components to files. This construction is the basis for the paper’s notion of live trace links (Jin et al., 14 Sep 2025).
A common misconception is that trace recovery here means ranking arbitrary textual pairs. UserTrace does compare itself against ranking-based traceability baselines, but its own primary mechanism is graph-structured generation and aggregation rather than pure similarity search.
3. Multi-agent architecture and workflow
UserTrace coordinates four specialized agents through a three-phase pipeline: repository structuring, IR derivation, and UR synthesis with verification (Jin et al., 14 Sep 2025).
| Agent | Role | Output |
|---|---|---|
| Code Reviewer | Derives dependency-aware IRs | Component-level and file-level IRs |
| Searcher | Retrieves domain/business knowledge | Domain concepts and terminology |
| Writer | Abstracts IRs into URs | Structured use cases |
| Verifier | Evaluates and refines URs | Acceptance or revision directives |
The first phase structures repository dependencies. Static analysis builds the CDG and FDG, breaks cycles, and topologically orders both DAGs. The second phase derives implementation-level requirements. For each component in topological order, the Code Reviewer receives the component code and the IRs of its dependencies and produces a component-level IR. For each file in topological order, the system aggregates component-level IRs into a file-level IR. Dependency-aware prompting is deliberately restricted to one-hop IRs to reduce context length (Jin et al., 14 Sep 2025).
The third phase synthesizes user-level requirements. The FDG is partitioned with Leiden community detection so that UR synthesis operates on semantically cohesive file groups rather than on the repository as a whole. For each community, the Searcher retrieves domain context from project artifacts and external sources, including readme files, issues, wiki pages, online documentation, and general web search. The Writer then abstracts the community’s file-level IRs into URs in use-case form with name, actors, description, preconditions, flow of events, and exit conditions. The Verifier evaluates these URs against three criteria: business context value, completeness, and detail level, and triggers iterative revision when needed (Jin et al., 14 Sep 2025).
This division of labor is central to the system’s design. The Code Reviewer is intended to remain close to program semantics, the Searcher injects domain knowledge not explicit in code, the Writer performs abstraction, and the Verifier constrains hallucination and under-specification. The paper’s qualitative analysis states that Code Reviewer alone yields good IRs but insufficient URs, that Searcher significantly improves helpfulness and correctness, and that Verifier is key to reducing hallucinations and improving completeness (Jin et al., 14 Sep 2025).
4. Trace-link recovery, formal criteria, and evaluation measures
UserTrace evaluates both requirement quality and trace-link recovery. For UR generation, it uses Precision, Recall, and F1 over generated and ground-truth UR sets:
4
5
6
For trace links, it uses group-based precision, recall, and F1 over generated links judged correct (Jin et al., 14 Sep 2025).
Correctness of a generated requirement-to-code link is defined groupwise. If a generated requirement links to code set 7 and the ground-truth requirement links to code set 8, then the generated link is correct if
9
with 0 in the experiments (Jin et al., 14 Sep 2025). This criterion tolerates partial overlap but requires majority agreement at the linked-code-set level.
The baseline retrieval methods used for comparison include VSM, LSI, COMET, FTLR, and LiSSA. For textual baselines such as VSM and LSI, cosine similarity is written as
1
UR quality is also assessed by an LLM-as-a-judge framework using completeness, correctness, and helpfulness. Completeness means that the UR document covers all or almost all ground-truth requirements; correctness means that it avoids hallucinations and that each UR is supported by repository content; helpfulness means that it goes beyond restating code elements and clarifies purpose, usage context, actors, and expected behavior (Jin et al., 14 Sep 2025).
The paper’s “live” terminology is partly architectural and partly aspirational. It states that links are amenable to incremental updates when the repository changes and sketches an incremental update concept: recompute diffs at file and component level, update CDG and FDG locally, rerun Code Reviewer for affected nodes, propagate updates to file-level IRs and impacted UR communities, and re-verify the revised URs. At the same time, it explicitly notes that the evaluation uses static snapshots and that commit history analysis and dynamic runtime traces are not part of the implemented pipeline (Jin et al., 14 Sep 2025).
5. Experimental results and reported outcomes
The experiments use three primary systems: eTour, eAnci, and SMOS. These are Java repositories in tourism, governance, and education, respectively. The reported dataset statistics are 58 use cases, 116 code units, and 308 links for eTour; 139 use cases, 55 code units, and 567 links for eAnci; and 67 use cases, 100 code units, and 1044 links for SMOS. AST preprocessing includes identifier translation to the predominant project language to reduce cross-language vocabulary mismatch (Jin et al., 14 Sep 2025).
| Dataset | UR F1, HieSum/Claude 3 → UserTrace/Claude 3 | Best baseline trace F1 → UserTrace trace F1 |
|---|---|---|
| eTour | 0.81 → 0.89 | LiSSA ≈ 0.278 → GPT-4 0.357 |
| eAnci | 0.31 → 0.41 | FTLR ≈ 0.157 → GPT-4 0.179 |
| SMOS | 0.87 → 0.91 | ≤0.193 → GPT-4 0.146 |
For UR quality, Claude 3 is reported as the strongest model among the tested LLMs. On eTour, HieSum with Claude 3 achieves F1 = 0.81, while UserTrace with Claude 3 reaches F1 = 0.89; on eAnci the corresponding scores are 0.31 and 0.41; on SMOS they are 0.87 and 0.91. The document-level LLM-judge scores also improve: for eTour, completeness/correctness/helpfulness are reported as 5, 4, 4 for UserTrace versus 4, 4, 4 for HieSum (Jin et al., 14 Sep 2025).
For trace-link recovery, the results are mixed but generally precision-favoring. On eTour, the best baseline is LiSSA at F1 ≈ 0.278, while UserTrace reaches F1 = 0.357 with GPT-4 and 0.336 with Claude 3. On eAnci, the best baseline is FTLR at F1 ≈ 0.157, while UserTrace with GPT-4 reaches 0.179. On SMOS, baselines are low at or below 0.193, and UserTrace with GPT-4 reports F1 = 0.146. The paper characterizes UserTrace as consistently increasing precision and overall F1 compared to VSM and LSI and as often surpassing COMET, FTLR, and LiSSA, while also emphasizing that its constructive linking is “by design” rather than solely similarity-ranked (Jin et al., 14 Sep 2025).
The user study is small but concrete. It involves 3 practitioners with at least 4 years of experience, who wrote intents and then validated repositories generated via Kiro with and without UserTrace. The reported outcomes are a reduction in time cost from 292 s to 188 s, an increase in validation accuracy from 0.63 to 0.85, satisfaction from 4.0 to 4.7, and confidence from 3.3 to 4.7. The paper notes that no statistical tests are reported (Jin et al., 14 Sep 2025).
The qualitative case study on SMOS describes a generated UR, “Manage Cultural Heritage Asset,” that abstracts multiple IRs such as getCulturalHeritage and addTagCulturalHeritage into a structured use case with actors, preconditions, flows, and exit conditions. This is representative of the system’s intended abstraction step from repository-level implementation fragments to user-facing functionality (Jin et al., 14 Sep 2025).
6. Limitations, interpretation, and relation to broader tracing research
The paper identifies several limitations. Generalizability is constrained by the evaluation set: three systems, Java code, and use-case style URs. It explicitly notes that use-case representation may not fit all projects, especially libraries and data pipelines. UserTrace also depends on repository structure and supporting artifacts; sparse comments or missing documentation can hinder IR derivation. LLM hallucination, abstraction errors, vocabulary mismatch across languages, and sensitivity to community partitioning are all listed as error sources. Most importantly, the “live” traceability claim is motivated by the graph construction, but empirical evidence for evolution handling is limited because the evaluation is conducted on static snapshots rather than commit histories or dynamically evolving repositories (Jin et al., 14 Sep 2025).
A further interpretive issue is terminological. In adjacent research areas, “trace” generally denotes observed execution behavior. Model-based trace-checking instruments programs, records rich execution traces, and checks them against formal models using Spin and Pro-B, allowing archived traces to be re-checked for new properties (Howard et al., 2011). PreciseTracer reconstructs exact per-request causal paths for multi-tier black-box services using OS-level send and receive events together with component activity graphs and dominated causal path patterns (Sang et al., 2010). Aggregate distributed-trace visualization groups traces by service overlap, structural similarity, graph depth, and latency in order to reason about whole trace datasets rather than individual traces (Samanta et al., 2024). UserTrace belongs to a different lineage: its traces are traceability links across repository abstractions.
This distinction matters because it clarifies what UserTrace does not do. It does not instrument GUI events as JETracer does for AWT, Swing, and SWT (Molnar, 2017); it does not reconstruct volatile database activity from memory snapshots as MemTraceDB does for MySQL (Nissan, 7 Sep 2025); and it does not watermark tool-using agent trajectories as TRACE does for reseller-controlled logs (Gao et al., 9 Jul 2026). UserTrace instead treats repository dependencies, generated IRs, synthesized URs, and their explicit links as the relevant evidence structure.
This contrast suggests a broader research agenda. Repository-derived traceability and execution-derived traces solve different observability problems: one connects user intent to code structure, the other connects behavior to runtime phenomena. A plausible implication is that future systems could combine UserTrace’s UR → IR → code links with execution-trace systems that recover causal paths or formally check behavioral traces, thereby relating user-facing requirements to actual executions. Such an integration is not part of the implemented UserTrace pipeline, but the surrounding trace literature shows that the underlying artifacts and analysis methods already exist in adjacent domains (Jin et al., 14 Sep 2025).