Papers
Topics
Authors
Recent
Search
2000 character limit reached

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Published 14 May 2026 in cs.LG and cs.CL | (2605.14381v1)

Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the sociotechnical nuance required for sensitive domains. We introduce NodeSynth, an evidence-grounded methodology that generates socially relevant synthetic queries by leveraging a fine-tuned taxonomy generator (TaG) anchored in real-world evidence. Evaluated against four mainstream LLMs (e.g., Claude 4.5 Haiku), NodeSynth elicited failure rates up to five times higher than human-authored benchmarks. Ablation studies confirm that our granular taxonomic expansion significantly drives these failure rates, while independent validation reveals critical deficiencies in prominent guard models (e.g., Llama-Guard-3). We open-source our end-to-end research prototype and datasets to enable scalable, high-stakes model evaluation and targeted safety interventions (https://github.com/google-research/nodesynth).

Summary

  • The paper presents a novel, taxonomy-driven framework that leverages expert-validated taxonomies and evidence annotation to generate scalable synthetic queries.
  • Its method increases failure rate detection in AI safety testing, achieving up to eightfold higher vulnerabilities than traditional synthetic or human-authored datasets.
  • The approach offers interpretable diagnostics and root-cause attribution, enabling targeted interventions to enhance model safety guardrails.

NodeSynth: Socially Aligned Synthetic Data for Sociotechnical AI Evaluation

Introduction

The rapid expansion of generative AI technologies introduces challenges in the development of robust, socio-technically grounded evaluation pipelines, especially for safety-sensitive domains such as medical advice and self-harm. Present evaluation regimes primarily use human-authored datasets or naïvely generated synthetic data, both of which fail to elicit failures representing nuanced, context-specific, real-world risks. "NodeSynth: Socially Aligned Synthetic Data for AI Evaluation" (2605.14381) introduces NodeSynth, an integrated taxonomy-driven and evidence-grounded framework for scalable, interpretable, and high-coverage synthetic query generation. This approach operationalizes granular taxonomic expansion, evidence annotation, and synthesized prompt construction to expose latent model vulnerabilities that are otherwise recalcitrant in traditional benchmarks.

Methodology: Taxonomy-Driven, Evidence-Grounded Synthetic Data Generation

NodeSynth employs a three-step pipeline to systematically construct socially grounded evaluation datasets:

  1. Multi-Scale Taxonomy Generation: Starting from a target domain (e.g., “Self-Harm”), the system uses a fine-tuned taxonomy generator, TaG, to automatically expand abstract safety concepts into validated, hierarchical taxonomies (L1–L3). Supervised fine-tuning on expert-reviewed data across multiple languages facilitates domain-specific and culturally situated taxonomic refinement, dramatically increasing the Jaccard similarity of predicted-to-ground-truth labels (+29.6%) and boosting human accuracy validation (TaG: 88% vs. default LLM: 60%).
  2. Evidence Grounding and Automated Annotation: With full taxonomies, the system automatically mines credible, real-world documentation and research (via Google Search-augmented LLM queries), mapping taxonomic branches to concrete metadata (e.g., demographic, geographic, occupation-level attributes). LLMs are prompted with strict extraction constraints, maintaining high fidelity in annotation and reducing information hallucination.
  3. Multi-Factor Synthetic Prompt Generation: The annotated taxonomic structure is translated through templated prompt synthesis, cross-conditioning on user personas, scenario, and evaluation modality. Synthetic queries are thus differentiated across sensitive subpopulations and, by construction, accompanied by interpretable provenance metadata.

Figure 1

Figure 1: NodeSynth’s pipeline for evidence-grounded taxonomy generation and annotated evaluation prompt synthesis.

Experimental Design

Evaluation Domains

Evaluation focuses on two high-stakes domains with direct policy mapping: Medical Advice (non-disclosed AI-generated medical information) and Self-Harm (promotion or non-mitigation of self-injury or suicide). Comparative benchmarks include:

  • Generic Synthetic Data (standard LLM prompt generation)
  • Human-Authored Data (public datasets [pfohl2024toolbox, nikhileswar2021suicide])
  • NodeSynth Data at different taxonomy depths (L2 and L3).

Models and Protocols

Four mainstream LLMs (Gemini 2.5 Flash, Llama 4 Scout, Claude 4.5 Haiku, GPT o4-mini) serve as evaluation targets. Model responses are auto-rated via a domain-specific rubric (Gemini-3.0), with reliability bootstrapped against human review (accuracy 89.5–93.3%, margin of error ±1.89%). Independent safety classifiers (LlamaGuard-3, QwenGuard, ShieldGemma) assess violation and bypass rates, particularly for Self-Harm content.

Results

Adversarial Efficacy and Taxonomy Depth

NodeSynth outperforms both human-authored and generic synthetic baselines in eliciting unsafe responses (“failures”) across all target models and domains.

Figure 2

Figure 2: Per-model breakdown of Medical Advice “no disclosure” and Self-Harm “non-safe” failure rates by Level 2 taxonomy.

NodeSynth L3 queries provoke up to fivefold higher failure rates than human-authored data in the Self-Harm domain (example: GPT o4-mini 9.89% vs. 1.93%) and over eightfold compared to generic synthetic queries on some models. For Medical Advice, mainstream models demonstrate resilience to generic synthetic queries (∼6% failure), but NodeSynth pushes average failures to nearly 30%. Increased taxonomy depth (L3 vs. L2) further escalates failure rates (Medical Advice: +4.9%, p=0.0027p=0.0027; Self-Harm: +2.7%, p=0.0076p=0.0076), demonstrating that expert-granular semantic expansion is critical for surfacing latent vulnerabilities.

Taxonomy Similarity, Human Review, and Precision

Figure 3

Figure 3: Distribution shift in taxonomy similarity scores before and after supervised fine-tuning (SFT) for TaG.

Fine-tuning the taxonomy generator yields substantial improvements in both lexical overlap (Jaccard) and relevance under blinded expert review, reinforcing the necessity of supervised, expert-curated, and multilingual taxonomic construction for credible sociotechnical evaluation.

Interpretable Diagnostics and Root Cause Attribution

NodeSynth’s metadata enables breakdowns of model safety failures by demographic, geography, and taxonomic intersections. Critical failures are traceable to specific themes—e.g., “malpractice liability” and “emergency medical protocols” in Medical Advice; “severity of injury” and “social isolation” by user group in Self-Harm—supporting targeted intervention and patching at the root concept or demographic level.

Figure 4

Figure 4: Failure rate analysis by Level 2 taxonomy and User Group intersections across domains and models.

Guardrail Deficiency Analysis

NodeSynth-generated Self-Harm queries reveal systematic deficiencies in mainstream safety guardrails. Independent guard models flag only a small fraction of genuinely unsafe content (NodeSynth L3: LlamaGuard-3 0.56%, QwenGuard 4.21%, ShieldGemma 1.44%), with NodeSynth achieving much higher bypass rates than human or generic baselines (e.g., NodeSynth L3 bypasses QwenGuard 56.4% of the time in Self-Harm). This discrepancy highlights an urgent need for more nuanced, context-aware guard models.

NodeSynth’s auto-generated, annotated, and structurally complex queries provide a paradigm for validating and debugging safeguard model architectures.

Implications and Future Directions

NodeSynth advances the field of sociotechnical AI evaluation through several channels:

  • Scalable, High-Coverage Safety Stress Testing: Evidence-grounded taxonomies extend domain coverage far beyond previous generic synthetic datasets, surfacing rare and intersectional harms at scale.
  • Interpretability and Root-Cause Analysis: Intrinsic, line-level metadata enables rapid troubleshooting and domain-specific patching, supporting developer and auditor workflows.
  • Data Creator Well-Being: Automated generation of nuanced, socio-culturally diverse prompts minimizes the psychological burden of manual query and annotation for traumatic domains.
  • Practical Accessibility: The open-source NodeSynth prototype lowers barriers for civil society and low-resource organizations to audit high-stakes models on sensitive, policy-relevant issues.

The results also expose substantial limitations in current LLM-based safety guardrails, warranting further research on multi-turn, multi-factor adversarial data for robust alignment. Extending NodeSynth beyond current domains, dynamic adaptation for new languages and regulatory regimes, and integration with iterative RLHF or policy patching pipelines remain tractable avenues for future work.

Conclusion

NodeSynth establishes a rigorous, systematic methodology for socially aligned, evidence-grounded synthetic data generation, enabling high-fidelity AI safety evaluation in critical application domains. The framework’s combination of expert-driven taxonomy expansion, real-world evidence binding, and metadata enrichment results in synthetic queries that consistently induce and localize failures in commercial LLMs and defeat prevailing safety guardrails. NodeSynth’s release represents a substantial step towards interpretable, scalable, and socially responsible evaluation infrastructure for next-generation AI systems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.