- The paper presents a unified HI rotation curve dataset aggregating 8,963 measurements from 438 galaxies to address the mass discrepancy in disk galaxies.
- It harmonizes data from SPARC, THINGS, LITTLE THINGS, and WALLABY DR2 by standardizing units, annotations, and metadata, facilitating robust computational analysis.
- The corpus supports multi-component mass modeling and automated quality assessments, paving the way for advanced dark matter and astrophysical studies.
A Unified HI Rotation Curve Corpus for Computational Astrophysics
Context and Motivation
The mass discrepancy problem in disk galaxies, highlighted by the divergence between observed rotation velocities and those predicted from baryonic matter, remains central to contemporary astrophysics. Robust analysis depends critically on high-fidelity, spatially resolved HI rotation curve data. Historically, the four dominant HI surveys—SPARC, THINGS, LITTLE THINGS, and WALLABY DR2—have distributed their data across disparate infrastructures, with heterogeneous formatting and inconsistent conventions. This fragmentation inhibits reproducible computational analysis, systematic meta-studies, and the deployment of pipeline or LLM-driven workflows.
The presented work introduces the Unified HI Rotation Curve Corpus (v7.0), aggregating 8,963 rotation curve measurements from 423 galaxies (438 total catalog entries including 15 metadata-only THINGS galaxies). The corpus harmonizes unit conventions, schema, quality annotations, and metadata, providing a machine-readable dataset optimized for both traditional analysis and LLM/RAG pipelines. This enables sample selection, cross-survey comparison, and automated inference with minimal preprocessing.
Corpus Construction and Survey Integration
The Unified Corpus features rigorous data ingestion and cross-verification against primary tables. SPARC (175 galaxies) delivers HI/Hα rotation curves and full baryonic decomposition at 3.6 μm, with per-point uncertainties and independently verified distance and inclination parameters. THINGS (34 galaxies; 19 with per-point data) provides high-resolution VLA-tilted ring fits, with remaining entries offering validated metadata. LITTLE THINGS (26 galaxies) contributes rotation velocities with uncertainties for dwarf irregular systems. WALLABY DR2 (203 galaxies) integrates ASKAP pipeline outputs (3DBarolo + FAT) with a cautionary tier-2 designation due to the absence of per-ring uncertainties and limitation from beam-smearing effects at Vrot​<50 km/s.
Two-tier quality annotations distinguish between hand-curated (tier 1) and automated (tier 2) products. All radii are standardized to kiloparsecs and velocities to km/s. Baryonic decomposition adopts Υ=1, with sign-preserving quadrature supporting negative gas velocities at inner radii. The corpus schema accommodates survey-specific observables, balancing structural unification with retention of original quantities.
Figure 1: DDO~161 (SPARC Tier~1) rotation curve and baryonic decomposition, showing mass discrepancy and omega correction application.
Data Accessibility and Computational Interface
Three complementary formats facilitate broad accessibility:
- Master JSON: Encapsulates the entire corpus structure, metadata, and per-galaxy entries.
- Flat CSV: Provides 438 rows with 29 columns summarizing essential catalog-level parameters for rapid selection and filtering.
- Per-Galaxy ZIP Archive: Contains 438 self-contained JSON files, optimized for retrieval-augmented generation and LLM ingestion.
All formats are versioned, self-describing, and maintain full floating-point precision for round-trip programmatic fidelity.
Corpus-Level Statistical Properties
The corpus offers unmatched coverage—spanning low-mass dwarf irregulars to massive spirals with outer radii up to ∼100 kpc. Distributions of peak rotation velocities, coverage depth, and morphology types demonstrate broad parameter space representation. Crossmatched galaxies enable complementary analysis across pipeline products and baryonic decomposition.
Figure 2: Peak rotation velocity distribution and Rmax​ vs.\ Vrot,max​ parameter space for all surveys, colored by quality tier.
Exemplary Analyses
The paper details canonical analyses achievable within minimal Python using the corpus:
LLM-Driven Inference and RAG Integration
A distinguishing feature is the corpus's optimization for LLM-based inference within RAG architectures. Each per-galaxy JSON is sufficiently self-describing for automated consumption. Usability tests across Gemini Pro, Claude, AstroSage-Llama, and Copilot Pro confirm robust generation of syntactically correct Python workflows for rotation curve plotting, baryonic quadrature, and omega correction without reference to external documentation.
The corpus's explicit nomenclature, unit annotation, and quality flags support reliable code generation, facilitating migration of astrophysical workflows toward retrieval-augmented, context-driven analysis environments.
Limitations and Caveats
Key limitations include:
- Metadata-only entries for THINGS galaxies lacking usable rotation curve data.
- Absence of coordinates and systemic velocities for SPARC entries due to original publication constraints.
- Lack of per-ring uncertainties and baryonic decomposition in WALLABY DR2.
- Partial harmonization of JSON schema across surveys to respect original data structures.
- Maintenance of floating-point precision potentially exceeding instrumental uncertainties.
These caveats are meticulously documented, ensuring transparent provenance and reproducibility for downstream analytical efforts.
Theoretical and Practical Implications
By standardizing HI rotation curve data, the corpus enables theoretical advances in dark matter halo modeling, modified gravity frameworks, and empirical correction schemes—removing preprocessing barriers and fostering reproducible cross-survey analysis. The explicit support for LLM-based workflows anticipates accelerating adoption of context-driven computational astrophysics, bridging archival data with inference pipelines. The corpus's extensible schema positions it as a template for future standardization efforts across legacy HI surveys and broader multi-wavelength studies.
Conclusion
The Unified HI Rotation Curve Corpus delivers an authoritative, meticulously curated dataset integrating four major HI surveys with rigorous provenance and quality annotation. Its harmonized schema, optimized data formats, and explicit support for LLM/RAG workflows lower the technical threshold for high-quality computational analysis and automated inference. With broad survey coverage and documented limitations, it provides a foundational tool for both traditional astrophysical modeling and emerging AI-driven research. The corpus's public availability and reproducibility standards position it as a central resource for future studies of galaxy kinematics and mass modeling.