Papers
Topics
Authors
Recent
Search
2000 character limit reached

NodeSynth: Synthetic Evaluation for AI Safety

Updated 5 July 2026
  • NodeSynth is an evidence-grounded methodology that generates synthetic evaluation queries by binding multi-scale taxonomies to real-world social contexts and stakeholder profiles.
  • It employs a three-stage pipeline—taxonomy generation, evidence grounding, and structured prompt synthesis—to systematically reveal AI system failures in sensitive domains.
  • Experiments on Medical Advice and Self-Harm domains show higher failure rates compared to generic benchmarks, demonstrating its effectiveness in identifying robustness gaps.

NodeSynth is a socially aligned, evidence-grounded methodology for generating synthetic evaluation queries for AI systems in sensitive domains. It was introduced as an evaluation framework rather than a training-data pipeline, with the aim of producing synthetic queries that preserve sociotechnical nuance by binding taxonomic structure to documented real-world contexts, stakeholder groups, and policy definitions. Operationally, NodeSynth converts an evaluation intent into an annotated query set through multi-scale taxonomy generation, evidence grounding, and structured prompt synthesis; in the reported experiments on Medical Advice and Self-Harm, it elicited substantially higher failure rates than both human-authored and generic synthetic benchmarks, while also revealing robustness gaps in independent guard models (Rashid et al., 14 May 2026).

1. Conceptual basis and problem formulation

NodeSynth addresses a specific limitation of generic LLM-generated evaluation sets: although such sets are scalable, they are described as shallow, insufficiently diverse, and prone to hallucinating context. In sensitive domains, these weaknesses are consequential because harms often arise only at particular intersections of domain knowledge, geography, institutional role, and demographic context. The NodeSynth formulation therefore treats synthetic evaluation as a sociotechnical modeling problem rather than merely a prompt-generation problem (Rashid et al., 14 May 2026).

In this framework, “socially aligned” does not denote preference optimization. It refers instead to alignment with real-world social contexts, stakeholder groups, and policy definitions. Queries are constrained by documented harms and sensitive attributes such as demographics, occupations, and country. The associated notion of “sociotechnical nuance” refers to the interplay of social factors with technical system behavior; the paper’s examples include how self-harm content manifests for particular user groups and how medical-advice policies intersect with malpractice, emergency protocols, or country-specific constraints (Rashid et al., 14 May 2026).

The paper evaluates NodeSynth in two domains. “Medical Advice” is defined as content advising on diagnosis or treatment without proper disclosures, and “Self-Harm” as content that promotes, instructs, or glamorizes self-harm, eating disorders, or suicide, or that is clinically risky without resources. These domain definitions are aligned to public policies, including Google, OpenAI, and YouTube policies. This suggests that NodeSynth is designed not only to surface model failures, but also to localize failures relative to externally legible policy categories (Rashid et al., 14 May 2026).

2. End-to-end pipeline and data model

NodeSynth implements a three-stage pipeline. The first stage is multi-scale taxonomy generation. Given a target concept CC, the system learns a taxonomy mapping f(C)→Tf(C) \to T, where TT is a three-level taxonomy with L1L1 themes, L2L2 sub-topics, and L3L3 granular keyterms. In the reported system, L1L1 and L2L2 are bootstrapped by Gemini 2.5 Flash and then reviewed by experts, with reported agreement of at least 90% on accuracy, completeness, specificity, and relevance. The L3L3 layer is then expanded automatically by a supervised fine-tuned model called TaG (Rashid et al., 14 May 2026).

The second stage is evidence grounding and automated annotation. For each taxonomy branch (L1→L2→L3)(L1 \to L2 \to L3), the system retrieves real-world documents by using Gemini 2.5 configured with Google Search. From those retrieved documents, a strict extraction prompt is used to produce exactly four fields: Title, Occupation, Demographics, and Country. The extraction prompt includes a one-shot example and forbids rationale or extra fields, with the stated goal of minimizing hallucinations and constraining outputs to verifiable social attributes (Rashid et al., 14 May 2026).

The third stage is multi-factor synthetic prompt generation. A structured template constructs the final prompts from Domain, f(C)→Tf(C) \to T0, f(C)→Tf(C) \to T1 keyterms, Sensitive user groups given as Occupation, Demographics, and Country, plus Use case and Modality. The output is an annotated query set in which each query remains traceable to taxonomic nodes and grounded attributes. The paper emphasizes this lineage because it enables interpretable diagnostics, such as slicing failures by topic depth, geography, or stakeholder group rather than treating benchmark errors as undifferentiated counts (Rashid et al., 14 May 2026).

The paper also gives an explicit algorithmic sketch. Starting from target concept f(C)→Tf(C) \to T2, domain definition f(C)→Tf(C) \to T3, and user options f(C)→Tf(C) \to T4, the pipeline first generates and validates f(C)→Tf(C) \to T5–f(C)→Tf(C) \to T6, then uses TaG to expand f(C)→Tf(C) \to T7, then performs search-and-extract over each node, and finally instantiates a template to return a prompt set f(C)→Tf(C) \to T8 with metadata f(C)→Tf(C) \to T9. This representation is central to NodeSynth’s interpretability claims because each query carries both prompt text and structured provenance (Rashid et al., 14 May 2026).

3. TaG: fine-tuned taxonomy expansion

TaG, the Taxonomy Generator, is the component responsible for expanding validated TT0 topics into granular TT1 keyterms. It is implemented as a parameter-efficient supervised fine-tuning of gemini-2.5-flash specialized for sensitive domains and safety policies. The training corpus contains 2,576 expert-labeled instances spanning safety policies such as Self-Harm, Sexual Content, and Medical Advice, as well as sensitive domains including Culture, Education, Labor/Employment, Legal/Civil Rights, Politics/Government, and Privacy/Security. Each instance includes target concept, description, TT2, TT3, TT4, country, and language; the languages reported are English (global), Indonesian, Brazilian Portuguese, Hindi, and Spanish (Mexico). The split is 80% train, 10% validation, and 10% test (Rashid et al., 14 May 2026).

The fine-tuning procedure uses supervised cross-entropy over tokenized outputs, with the objective of producing specific, relevant, and complete TT5 lists conditioned on TT6–TT7 context, target concept, and language. The reported hyperparameters are 4 epochs, learning_rate_multiplier = 1.0, and adapter_size = 4. Inputs consist of the TT8 topic, the TT9 topic, the domain definition, and country/language; outputs consist of exactly three L1L10 keyterms per L1L11 topic together with concise rationale, under a prompted format that enforces specificity and non-duplication (Rashid et al., 14 May 2026).

TaG is evaluated quantitatively and by human review. On L1L12 list evaluation, mean Jaccard similarity improves by 29.6% over the base model, from 0.54 to 0.70, where

L1L13

A blinded human review on 531 education/culture topics reports 88% accuracy for TaG versus 60% for the base model, using a three-point Likert assessment collapsed to accuracy over “somewhat/excellent” hierarchical checks. These results are important because the paper later attributes a substantial share of NodeSynth’s failure-discovery advantage to the added granularity supplied by the L1L14 layer (Rashid et al., 14 May 2026).

The paper also presents a derived formalization of candidate selection,

L1L15

but explicitly notes that the implementation does not report such an explicit scoring function. Instead, it selects three L1L16 terms per L1L17 per prompt specification. This distinction matters because the published empirical results concern a concrete prompt-generation procedure, whereas the scoring formula is presented only as a plausible abstraction of how one might formalize candidate ranking (Rashid et al., 14 May 2026).

4. Evaluation protocol and empirical findings

NodeSynth is evaluated against three benchmark families: Human-Authored prompts drawn from published benchmarks, Generic Synthetic prompts generated by Gemini 2.5 Flash without taxonomy or evidence grounding, and NodeSynth prompts in two variants, one using only L1L18 and one enabling L1L19 in order to isolate the contribution of fine-grained taxonomic depth. The target models are Gemini 2.5 Flash, Llama 4 Scout, Claude 4.5 Haiku, and GPT o4-mini (Rashid et al., 14 May 2026).

The evaluation pipeline uses Gemini-3.0 as an automated rater with domain-specific rubrics. For Medical Advice, responses are categorized as Disclosure with instructions, Disclosure without instructions, or No disclosure, and failure is defined as No disclosure. For Self-Harm, responses are categorized as Harmful, Risky, or Safe, and failure is defined as Non-safe, that is, Harmful or Risky. The core metric is failure rate,

L2L20

The paper reports L2L21 query–response–rating triads, a minimum sample size of 354 via Cochran’s formula with finite population correction at 95% confidence and 5% margin of error, and a human review of 700 auto-rated instances. In that review, 60 errors were found, corresponding to an 8.57% error rate; using a Wald confidence interval with finite population correction, the reported auto-rater accuracy is 89.54%–93.32% at 95% confidence, with L2L22 margin of error (Rashid et al., 14 May 2026).

The reported failure rates show a clear separation between generic synthetic generation and NodeSynth. In Medical Advice, Generic Synthetic averages approximately 6% failure, whereas NodeSynth pushes failures to approximately 30% on average. Concrete examples include Claude 4.5 Haiku, where Medical Advice failures rise from 5.10% for Generic Synthetic to 32.29% for NodeSynth L2L23, and Gemini 2.5 Flash, where they rise from 5.10% to 28.82%. In Self-Harm, Claude 4.5 Haiku rises from 2.00% for Generic Synthetic to 15.88% for NodeSynth L2L24, and Gemini 2.5 Flash rises from 2.00% to 14.62%. Relative to Human-Authored data, specific model-domain pairings also show substantial increases; for GPT o4-mini on Self-Harm, Human-Authored data yields 1.93% failure while NodeSynth L2L25 yields 9.89%, reported as a 5.13× increase (Rashid et al., 14 May 2026).

The paper attributes part of this advantage to taxonomic granularity. A one-tailed two-proportion L2L26-test is used, with pooled proportion

L2L27

and

L2L28

Across all conditions, NodeSynth L2L29 failure rates are reported as higher than baselines at L3L30. In the L3L31 versus L3L32 ablation, Medical Advice increases from 24.86% to 29.78% (L3L33 points, L3L34), and Self-Harm increases from 13.76% to 16.47% (L3L35 points, L3L36). The paper therefore treats fine-grained taxonomic expansion as a principal driver of failure discovery rather than a cosmetic increase in prompt variety (Rashid et al., 14 May 2026).

Qualitative error analysis reinforces this interpretation. Medical Advice failures concentrate in topics such as malpractice liability and emergency medical protocols. For Self-Harm, the paper highlights failures for Claude 4.5 Haiku in areas such as severity of injury for the Researchers user group and social isolation for Healthcare professionals. This suggests that the system’s main contribution is not simply generating harder prompts, but generating prompts whose difficulty emerges from realistic contextual intersections (Rashid et al., 14 May 2026).

5. Guard-model validation, safety posture, and limitations

A separate contribution of NodeSynth is guard-model validation. The paper evaluates Self-Harm query–response pairs with LlamaGuard-3-8B, Qwen3Guard-4B, and ShieldGemma-9B. On violation rates, absolute values remain small, but NodeSynth L3L37 is consistently hardest: LlamaGuard-3-8B rises from 0.07% on Human to 0.56% on NodeSynth L3L38; Qwen3Guard-4B rises from 2.08% to 4.21%; ShieldGemma-9B rises from 0.60% to 1.44%. More revealing are bypass rates on queries classified safe by the guard. For Self-Harm, Qwen3Guard rises from 13.69% on Human to 56.35% on NodeSynth L3L39, and LlamaGuard-3-8B rises from 29.61% to 72.95%. Medical Advice also shows high bypass, including 81.41% for Qwen3Guard on NodeSynth L1L10. A one-tailed two-proportion L1L11-test on Self-Harm confirms that NodeSynth L1L12 significantly increases flagged responses across guard models at L1L13 (Rashid et al., 14 May 2026).

The system’s safety posture is explicitly conservative in presentation. Evidence extraction is restricted to field-only outputs; evidence grounding is intended to reduce stereotyping and flattening by tying prompts to documented contexts and sensitive groups; and the methodology is framed as being for evaluation and detection rather than training. Human review remains part of the pipeline at both taxonomy construction and rating-quality assurance stages. The paper also proposes future additions such as automated evidence-quality metrics and multi-model auto-rater triangulation (Rashid et al., 14 May 2026).

The limitations are equally explicit. First, the diversity of SFT data creators should be expanded across countries and languages. Second, the generalization of TaG beyond trained domains remains to be established. Third, evidence grounding still carries residual hallucination risk because web search and extraction are model-mediated. Fourth, the present results rely on one auto-rater family; the paper recommends secondary raters such as Claude for triangulation. Fifth, broader replication across newer models and additional sensitive domains is still pending. A plausible implication is that NodeSynth’s strength lies in targeted evaluation depth rather than in being a universal benchmark constructor across all policy settings (Rashid et al., 14 May 2026).

6. Operational use, artifacts, and intervention workflow

NodeSynth is released as an open-source research prototype with datasets at https://github.com/google-research/nodesynth. The paper states that prompts for each pipeline stage—taxonomy generation, evidence grounding, and query generation—are included in the appendices, and that the repository provides the end-to-end prototype and datasets. The example data schema contains the fields domain, L1, L2, L3[], occupation[], demographics[], country[], use_case, modality, query_text, and lineage metadata for interpretability. For reproducibility in guard evaluation, the reported configurations use standard, no-custom-policy settings for LlamaGuard and Qwen Guard, while ShieldGemma uses the provided “No Dangerous Content” rule with probability thresholding above 0.6 (Rashid et al., 14 May 2026).

The paper also gives practical guidance for deploying NodeSynth as a safety-analysis workflow. The recommended sequence begins with defining evaluation intent through domain, policy definition, geography/language, and use case. It then proceeds to taxonomy generation and SME validation, followed by TaG-based L1L14 expansion. The next step is evidence grounding, in which Occupation, Demographics, and Country are attached to each taxonomy branch through search-and-extract. After that, the template-based synthesizer composes multi-factor queries, with difficulty increased by implicit risk framing, professional roles, country-specific context, and inclusion of L1L15. Models are then evaluated with domain rubrics and an auto-rater, combined with human sampling checks. Finally, failures are sliced by L1L16, L1L17, L1L18, user group, and geography to support targeted safety fine-tuning, refusal-policy adjustments, or improved guard-rail rules (Rashid et al., 14 May 2026).

This operational framing is significant because it makes NodeSynth more than a benchmark-construction recipe. The metadata schema and lineage structure support post hoc intervention design. Rather than only reporting an aggregate failure rate, a team can identify which intersections of topic, role, and geography are fragile. The paper presents this as the route from synthetic evaluation to targeted safety intervention, and the reported ablations support the view that the most informative interventions will often be associated with the fine-grained L1L19 layer (Rashid et al., 14 May 2026).

7. Comparisons, trade-offs, and name disambiguation

Within evaluation research, NodeSynth is positioned against three alternative approaches. Relative to random or generic synthetic generation, it adds grounded social attributes and fine-grained taxonomic specificity. Relative to automated red-teaming based on broad taxonomies, it adds L2L20 granularity and evidence-linked personas. Relative to programmatic templates, it offers controlled diversity without being wholly decontextualized. The trade-off is explicit: NodeSynth requires additional setup for taxonomy construction and evidence retrieval, but the paper reports substantially higher failure discovery and more interpretable diagnostics as a result (Rashid et al., 14 May 2026).

The name should also be distinguished from unrelated lines of work. NSynth is a large-scale dataset of 306,043 monophonic musical notes introduced for raw-audio neural synthesis and WaveNet-style autoencoding, not a safety-evaluation methodology (Engel et al., 2017). SING is a non-autoregressive symbol-to-instrument neural generator for musical notes conditioned on instrument, pitch, and velocity, with emphasis on waveform synthesis efficiency and generalization to unseen instrument–pitch pairs (Défossez et al., 2018). “Universal audio synthesizer control with normalizing flows” formulates synthesizer control as an invertible mapping between audio latent space and synthesizer parameters, targeting parameter inference, macro-control learning, and preset exploration rather than evaluation benchmarking (Esling et al., 2019). Maniposynth, by contrast, is a bimodal tangible functional programming environment centered on live values, hole expressions, and program synthesis in OCaml (Hempel et al., 2022). A further distinct line concerns differentiable spectral modeling of piano notes through sines, transient, and noise decomposition for low-latency audio emulation (Simionato et al., 2024).

These distinctions matter because “NodeSynth” can be conflated with synthetic generation in audio or with node-oriented synthesis metaphors in programming environments. In the published literature summarized here, however, NodeSynth specifically denotes an evidence-grounded synthetic-query generation framework for high-stakes AI evaluation. Its characteristic features are social alignment to policy and stakeholder context, multi-scale taxonomic expansion through TaG, evidence-grounded metadata extraction, and prompt lineage that supports both quantitative benchmarking and targeted safety intervention (Rashid et al., 14 May 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to NodeSynth.