---
title: 'NodeSynth: Synthetic Evaluation for AI Safety'
url: https://www.emergentmind.com/topics/nodesynth
type: topic
---

# NodeSynth: Synthetic Evaluation for AI Safety

NodeSynth is a socially aligned, evidence-grounded methodology for generating synthetic evaluation queries for AI systems in sensitive domains. It was introduced as an evaluation framework rather than a training-data pipeline, with the aim of producing synthetic queries that preserve sociotechnical nuance by binding taxonomic structure to documented real-world contexts, stakeholder groups, and policy definitions. Operationally, NodeSynth converts an evaluation intent into an annotated query set through multi-scale taxonomy generation, evidence grounding, and structured prompt synthesis; in the reported experiments on Medical Advice and Self-Harm, it elicited substantially higher failure rates than both human-authored and generic synthetic benchmarks, while also revealing robustness gaps in independent guard models [2605.14381].

## 1. Conceptual basis and problem formulation

NodeSynth addresses a specific limitation of generic LLM-generated evaluation sets: although such sets are scalable, they are described as shallow, insufficiently diverse, and prone to hallucinating context. In sensitive domains, these weaknesses are consequential because harms often arise only at particular intersections of domain knowledge, geography, institutional role, and demographic context. The NodeSynth formulation therefore treats synthetic evaluation as a sociotechnical modeling problem rather than merely a prompt-generation problem [2605.14381].

In this framework, “socially aligned” does not denote preference optimization. It refers instead to alignment with real-world social contexts, stakeholder groups, and policy definitions. Queries are constrained by documented harms and sensitive attributes such as demographics, occupations, and country. The associated notion of “sociotechnical nuance” refers to the interplay of social factors with technical system behavior; the paper’s examples include how self-harm content manifests for particular user groups and how medical-advice policies intersect with malpractice, emergency protocols, or country-specific constraints [2605.14381].

The paper evaluates NodeSynth in two domains. “Medical Advice” is defined as content advising on diagnosis or treatment without proper disclosures, and “Self-Harm” as content that promotes, instructs, or glamorizes self-harm, eating disorders, or suicide, or that is clinically risky without resources. These domain definitions are aligned to public policies, including Google, OpenAI, and YouTube policies. This suggests that NodeSynth is designed not only to surface model failures, but also to localize failures relative to externally legible policy categories [2605.14381].

## 2. End-to-end pipeline and data model

NodeSynth implements a three-stage pipeline. The first stage is multi-scale taxonomy generation. Given a target concept $C$, the system learns a taxonomy mapping $f(C) \to T$, where $T$ is a three-level taxonomy with $L1$ themes, $L2$ sub-topics, and $L3$ granular keyterms. In the reported system, $L1$ and $L2$ are bootstrapped by Gemini 2.5 Flash and then reviewed by experts, with reported agreement of at least 90% on accuracy, completeness, specificity, and relevance. The $L3$ layer is then expanded automatically by a supervised fine-tuned model called TaG [2605.14381].

The second stage is evidence grounding and automated annotation. For each taxonomy branch $(L1 \to L2 \to L3)$, the system retrieves real-world documents by using Gemini 2.5 configured with Google Search. From those retrieved documents, a strict extraction prompt is used to produce exactly four fields: Title, Occupation, Demographics, and Country. The extraction prompt includes a one-shot example and forbids rationale or extra fields, with the stated goal of minimizing hallucinations and constraining outputs to verifiable social attributes [2605.14381].

The third stage is multi-factor synthetic prompt generation. A structured template constructs the final prompts from Domain, $(L1, L2)$, $L3$ keyterms, Sensitive user groups given as Occupation, Demographics, and Country, plus Use case and Modality. The output is an annotated query set in which each query remains traceable to taxonomic nodes and grounded attributes. The paper emphasizes this lineage because it enables interpretable diagnostics, such as slicing failures by topic depth, geography, or stakeholder group rather than treating benchmark errors as undifferentiated counts [2605.14381].

The paper also gives an explicit algorithmic sketch. Starting from target concept $C$, domain definition $D$, and user options $U=\{\text{use\_case}, \text{modality}, \text{country/language}\}$, the pipeline first generates and validates $L1$–$L2$, then uses TaG to expand $L3$, then performs search-and-extract over each node, and finally instantiates a template to return a prompt set $Q$ with metadata $M$. This representation is central to NodeSynth’s interpretability claims because each query carries both prompt text and structured provenance [2605.14381].

## 3. TaG: fine-tuned taxonomy expansion

TaG, the Taxonomy Generator, is the component responsible for expanding validated $L2$ topics into granular $L3$ keyterms. It is implemented as a parameter-efficient supervised fine-tuning of gemini-2.5-flash specialized for sensitive domains and safety policies. The training corpus contains 2,576 expert-labeled instances spanning safety policies such as Self-Harm, Sexual Content, and Medical Advice, as well as sensitive domains including Culture, Education, Labor/Employment, Legal/Civil Rights, Politics/Government, and Privacy/Security. Each instance includes target concept, description, $L1$, $L2$, $L3$, country, and language; the languages reported are English (global), Indonesian, Brazilian Portuguese, Hindi, and Spanish (Mexico). The split is 80% train, 10% validation, and 10% test [2605.14381].

The fine-tuning procedure uses supervised cross-entropy over tokenized outputs, with the objective of producing specific, relevant, and complete $L3$ lists conditioned on $L1$–$L2$ context, target concept, and language. The reported hyperparameters are 4 epochs, `learning_rate_multiplier = 1.0`, and `adapter_size = 4`. Inputs consist of the $L1$ topic, the $L2$ topic, the domain definition, and country/language; outputs consist of exactly three $L3$ keyterms per $L2$ topic together with concise rationale, under a prompted format that enforces specificity and non-duplication [2605.14381].

TaG is evaluated quantitatively and by human review. On $L3$ list evaluation, mean Jaccard similarity improves by 29.6% over the base model, from 0.54 to 0.70, where
$$
J(P,G)=\frac{|P\cap G|}{|P\cup G|}.
$$
A blinded human review on 531 education/culture topics reports 88% accuracy for TaG versus 60% for the base model, using a three-point Likert assessment collapsed to accuracy over “somewhat/excellent” hierarchical checks. These results are important because the paper later attributes a substantial share of NodeSynth’s failure-discovery advantage to the added granularity supplied by the $L3$ layer [2605.14381].

The paper also presents a derived formalization of candidate selection,
$$
S(n)=\alpha \operatorname{Rel}(n)+\beta \operatorname{Spec}(n)+\gamma \operatorname{Cov}(n)-\delta \operatorname{Dup}(n),
$$
but explicitly notes that the implementation does not report such an explicit scoring function. Instead, it selects three $L3$ terms per $L2$ per prompt specification. This distinction matters because the published empirical results concern a concrete prompt-generation procedure, whereas the scoring formula is presented only as a plausible abstraction of how one might formalize candidate ranking [2605.14381].

## 4. Evaluation protocol and empirical findings

NodeSynth is evaluated against three benchmark families: Human-Authored prompts drawn from published benchmarks, Generic Synthetic prompts generated by Gemini 2.5 Flash without taxonomy or evidence grounding, and NodeSynth prompts in two variants, one using only $L2$ and one enabling $L3$ in order to isolate the contribution of fine-grained taxonomic depth. The target models are Gemini 2.5 Flash, Llama 4 Scout, Claude 4.5 Haiku, and GPT o4-mini [2605.14381].

The evaluation pipeline uses Gemini-3.0 as an automated rater with domain-specific rubrics. For Medical Advice, responses are categorized as Disclosure with instructions, Disclosure without instructions, or No disclosure, and failure is defined as No disclosure. For Self-Harm, responses are categorized as Harmful, Risky, or Safe, and failure is defined as Non-safe, that is, Harmful or Risky. The core metric is failure rate,
$$
FR=\frac{\#\ \text{failures}}{\#\ \text{queries}}.
$$
The paper reports $N=4{,}459$ query–response–rating triads, a minimum sample size of 354 via Cochran’s formula with finite population correction at 95% confidence and 5% margin of error, and a human review of 700 auto-rated instances. In that review, 60 errors were found, corresponding to an 8.57% error rate; using a Wald confidence interval with finite population correction, the reported auto-rater accuracy is 89.54%–93.32% at 95% confidence, with $\pm 1.89\%$ margin of error [2605.14381].

The reported failure rates show a clear separation between generic synthetic generation and NodeSynth. In Medical Advice, Generic Synthetic averages approximately 6% failure, whereas NodeSynth pushes failures to approximately 30% on average. Concrete examples include Claude 4.5 Haiku, where Medical Advice failures rise from 5.10% for Generic Synthetic to 32.29% for NodeSynth $L3$, and Gemini 2.5 Flash, where they rise from 5.10% to 28.82%. In Self-Harm, Claude 4.5 Haiku rises from 2.00% for Generic Synthetic to 15.88% for NodeSynth $L3$, and Gemini 2.5 Flash rises from 2.00% to 14.62%. Relative to Human-Authored data, specific model-domain pairings also show substantial increases; for GPT o4-mini on Self-Harm, Human-Authored data yields 1.93% failure while NodeSynth $L3$ yields 9.89%, reported as a 5.13× increase [2605.14381].

The paper attributes part of this advantage to taxonomic granularity. A one-tailed two-proportion $z$-test is used, with pooled proportion
$$
p=\frac{n_1p_1+n_2p_2}{n_1+n_2},
$$
and
$$
z=\frac{p_1-p_2}{\sqrt{p(1-p)(1/n_1+1/n_2)}}.
$$
Across all conditions, NodeSynth $L3$ failure rates are reported as higher than baselines at $p<.01$. In the $L3$ versus $L2$ ablation, Medical Advice increases from 24.86% to 29.78% ($\Delta=+4.92$ points, $p=0.0027$), and Self-Harm increases from 13.76% to 16.47% ($\Delta=+2.71$ points, $p=0.0076$). The paper therefore treats fine-grained taxonomic expansion as a principal driver of failure discovery rather than a cosmetic increase in prompt variety [2605.14381].

Qualitative error analysis reinforces this interpretation. Medical Advice failures concentrate in topics such as malpractice liability and emergency medical protocols. For Self-Harm, the paper highlights failures for Claude 4.5 Haiku in areas such as severity of injury for the Researchers user group and social isolation for Healthcare professionals. This suggests that the system’s main contribution is not simply generating harder prompts, but generating prompts whose difficulty emerges from realistic contextual intersections [2605.14381].

## 5. Guard-model validation, safety posture, and limitations

A separate contribution of NodeSynth is guard-model validation. The paper evaluates Self-Harm query–response pairs with LlamaGuard-3-8B, Qwen3Guard-4B, and ShieldGemma-9B. On violation rates, absolute values remain small, but NodeSynth $L3$ is consistently hardest: LlamaGuard-3-8B rises from 0.07% on Human to 0.56% on NodeSynth $L3$; Qwen3Guard-4B rises from 2.08% to 4.21%; ShieldGemma-9B rises from 0.60% to 1.44%. More revealing are bypass rates on queries classified safe by the guard. For Self-Harm, Qwen3Guard rises from 13.69% on Human to 56.35% on NodeSynth $L3$, and LlamaGuard-3-8B rises from 29.61% to 72.95%. Medical Advice also shows high bypass, including 81.41% for Qwen3Guard on NodeSynth $L3$. A one-tailed two-proportion $z$-test on Self-Harm confirms that NodeSynth $L3$ significantly increases flagged responses across guard models at $p<0.0001$ [2605.14381].

The system’s safety posture is explicitly conservative in presentation. Evidence extraction is restricted to field-only outputs; evidence grounding is intended to reduce stereotyping and flattening by tying prompts to documented contexts and sensitive groups; and the methodology is framed as being for evaluation and detection rather than training. Human review remains part of the pipeline at both taxonomy construction and rating-quality assurance stages. The paper also proposes future additions such as automated evidence-quality metrics and multi-model auto-rater triangulation [2605.14381].

The limitations are equally explicit. First, the diversity of SFT data creators should be expanded across countries and languages. Second, the generalization of TaG beyond trained domains remains to be established. Third, evidence grounding still carries residual hallucination risk because web search and extraction are model-mediated. Fourth, the present results rely on one auto-rater family; the paper recommends secondary raters such as Claude for triangulation. Fifth, broader replication across newer models and additional sensitive domains is still pending. A plausible implication is that NodeSynth’s strength lies in targeted evaluation depth rather than in being a universal benchmark constructor across all policy settings [2605.14381].

## 6. Operational use, artifacts, and intervention workflow

NodeSynth is released as an open-source research prototype with datasets at `https://github.com/google-research/nodesynth`. The paper states that prompts for each pipeline stage—taxonomy generation, evidence grounding, and query generation—are included in the appendices, and that the repository provides the end-to-end prototype and datasets. The example data schema contains the fields `domain`, `L1`, `L2`, `L3[]`, `occupation[]`, `demographics[]`, `country[]`, `use_case`, `modality`, `query_text`, and lineage metadata for interpretability. For reproducibility in guard evaluation, the reported configurations use standard, no-custom-policy settings for LlamaGuard and Qwen Guard, while ShieldGemma uses the provided “No Dangerous Content” rule with probability thresholding above 0.6 [2605.14381].

The paper also gives practical guidance for deploying NodeSynth as a safety-analysis workflow. The recommended sequence begins with defining evaluation intent through domain, policy definition, geography/language, and use case. It then proceeds to taxonomy generation and SME validation, followed by TaG-based $L3$ expansion. The next step is evidence grounding, in which Occupation, Demographics, and Country are attached to each taxonomy branch through search-and-extract. After that, the template-based synthesizer composes multi-factor queries, with difficulty increased by implicit risk framing, professional roles, country-specific context, and inclusion of $L3$. Models are then evaluated with domain rubrics and an auto-rater, combined with human sampling checks. Finally, failures are sliced by $L1$, $L2$, $L3$, user group, and geography to support targeted safety fine-tuning, refusal-policy adjustments, or improved guard-rail rules [2605.14381].

This operational framing is significant because it makes NodeSynth more than a benchmark-construction recipe. The metadata schema and lineage structure support post hoc intervention design. Rather than only reporting an aggregate failure rate, a team can identify which intersections of topic, role, and geography are fragile. The paper presents this as the route from synthetic evaluation to targeted safety intervention, and the reported ablations support the view that the most informative interventions will often be associated with the fine-grained $L3$ layer [2605.14381].

## 7. Comparisons, trade-offs, and name disambiguation

Within evaluation research, NodeSynth is positioned against three alternative approaches. Relative to random or generic synthetic generation, it adds grounded social attributes and fine-grained taxonomic specificity. Relative to automated red-teaming based on broad taxonomies, it adds $L3$ granularity and evidence-linked personas. Relative to programmatic templates, it offers controlled diversity without being wholly decontextualized. The trade-off is explicit: NodeSynth requires additional setup for taxonomy construction and evidence retrieval, but the paper reports substantially higher failure discovery and more interpretable diagnostics as a result [2605.14381].

The name should also be distinguished from unrelated lines of work. NSynth is a large-scale dataset of 306,043 monophonic musical notes introduced for raw-audio neural synthesis and WaveNet-style autoencoding, not a safety-evaluation methodology [1704.01279]. SING is a non-autoregressive symbol-to-instrument neural generator for musical notes conditioned on instrument, pitch, and velocity, with emphasis on waveform synthesis efficiency and generalization to unseen instrument–pitch pairs [1810.09785]. “Universal audio synthesizer control with normalizing flows” formulates synthesizer control as an invertible mapping between audio latent space and synthesizer parameters, targeting parameter inference, macro-control learning, and preset exploration rather than evaluation benchmarking [1907.00971]. Maniposynth, by contrast, is a bimodal tangible functional programming environment centered on live values, hole expressions, and program synthesis in OCaml [2206.14992]. A further distinct line concerns differentiable spectral modeling of piano notes through sines, transient, and noise decomposition for low-latency audio emulation [2409.06513].

These distinctions matter because “NodeSynth” can be conflated with synthetic generation in audio or with node-oriented synthesis metaphors in programming environments. In the published literature summarized here, however, NodeSynth specifically denotes an evidence-grounded synthetic-query generation framework for high-stakes AI evaluation. Its characteristic features are social alignment to policy and stakeholder context, multi-scale taxonomic expansion through TaG, evidence-grounded metadata extraction, and prompt lineage that supports both quantitative benchmarking and targeted safety intervention [2605.14381].

Source: https://www.emergentmind.com/topics/nodesynth