DINGO: A Multidomain Research Overview
- DINGO is a cross-domain research label that represents distinct frameworks in semantic funding, radio astronomy, gravitational-wave inference, and more.
- It integrates varied methodologies such as linked-data ontologies, deep Bayesian estimation, and distributed optimization to address domain-specific challenges.
- Implementations include the Data INtegration for Grants Ontology, ASKAP HI surveys, innovative neutron imaging setups, and dynamic decoding in diffusion LLMs.
DINGO is a reused research label rather than a single scientific object. In the arXiv literature it denotes, among other things, the Data INtegration for Grants Ontology in semantic-web research, the Deep Investigation of Neutral Gas Origins H I survey in radio astronomy, the Deep INference for Gravitational-wave Observations family of simulation-based inference systems, the DIstribution Network GeneratOr power-grid dataset, a distributed Newton-type method for gradient-norm optimization, a fine-grained instruction-following benchmark, a constrained inference algorithm for diffusion LLMs, and the Dingo neutron imaging beamline at ANSTO’s OPAL reactor (Chialva et al., 2020, Duffy et al., 2012, Dax et al., 2021, Abbas et al., 2 Sep 2025, Crane et al., 2019, Gu et al., 2024, Suresh et al., 29 May 2025, Kingston et al., 26 Feb 2025).
1. Principal meanings in the literature
| Domain | Meaning of DINGO | Representative source |
|---|---|---|
| Semantic web | Data INtegration for Grants Ontology | (Chialva et al., 2020) |
| Radio astronomy | Deep Investigation of Neutral Gas Origins | (Duffy et al., 2012) |
| Gravitational-wave inference | Deep INference for Gravitational-wave Observations | (Dax et al., 2021) |
| Power systems | DIstribution Network GeneratOr | (Abbas et al., 2 Sep 2025) |
| Optimization | Distributed Newton-Type Method for Gradient-Norm Optimization | (Crane et al., 2019) |
| LLM evaluation | Diverse and Fine-grained Instruction-Following evaluation dataset | (Gu et al., 2024) |
| Diffusion LLM inference | Constrained inference algorithm named DINGO | (Suresh et al., 29 May 2025) |
| Neutron imaging | Dingo beamline / instrument at OPAL | (Kingston et al., 26 Feb 2025) |
The term therefore functions as a cross-domain homonym. Some uses define a sustained lineage—most notably the gravitational-wave DINGO family, which later includes Dingo-T1, Dingo-Pop, DINGO-lensing, and related lensing and LISA extensions—whereas the ontology, radio-survey, optimization, dataset, and neutron-imaging uses are separate developments (Dax et al., 2021, Kofler et al., 2 Dec 2025, Leyde et al., 11 May 2026, Chan et al., 18 Dec 2025).
2. DINGO as a research-funding ontology
In ontology engineering, DINGO denotes the Data INtegration for Grants Ontology, an OWL-DL ontology comprising 40 classes and 68 properties for representing projects, grants, funding, actors, and especially funding policies as linked data (Chialva et al., 2020). It was introduced to address fragmentation, heterogeneity, and poor interoperability in funding-related information, with the stated goal of providing a machine-readable, extensible framework for semantically enabled applications in the research landscape and beyond, including domains such as arts or cultural conservation where funding is central (Chialva et al., 2020).
Its core modeling choices are explicitly domain-driven. A Project is defined as “an organised endeavour (collective or individual) planned to reach a particular aim or achieve a result,” whereas a Grant is “a disbursed fund paid to a recipient or beneficiary and the process for it” (Chialva et al., 2020). DINGO separates these classes because the relation is not one-to-one: a project may be funded by one or more grants, and a grant may fund one or several projects (Chialva et al., 2020). It likewise separates project participation from grant beneficiary status, allowing grants to be awarded to persons and/or organisations, and projects to be participated in by persons and/or organisations, without forcing those sets to coincide (Chialva et al., 2020).
The ontology’s distinctive feature is its explicit treatment of funding policies and instruments. Policy is modeled primarily through FundingScheme and Criterion. A FundingScheme is described as a funding instrument accompanied by specifications such as grant coverage, eligibility, reimbursement rates, specific criteria for funding, and target populations; these specifications are represented as Criterion instances and subclasses (Chialva et al., 2020). The paper characterizes this as the source of DINGO’s “high modeling power and elasticity,” meaning the ontology can represent diverse funding practices without collapsing agency-specific distinctions or requiring redesign from scratch (Chialva et al., 2020).
Development followed a mixed middle-out and bottom-up approach grounded in datasets from the EU Framework Programmes, the Australian Research Council, the Swiss National Science Foundation, the Croatian Science Foundation, NIH, NSF, and UK research councils (Chialva et al., 2020). The design guidelines included practical usability, interoperability “from the inception” with graphs such as Wikidata and Schema.org, sufficient granularity for monitoring and evaluation, sufficient generality to accommodate potentially all funding data, and straightforward extensibility (Chialva et al., 2020). DINGO is distributed in RDF-Turtle, documented on the web, accompanied by a Shape Expressions (ShEx) validation model, and mapped to Wikidata, schema.org, and FRAPO, with conservative use of owl:equivalentClass and owl:equivalentProperty (Chialva et al., 2020).
The paper also reports substantial uptake. DINGO was first presented publicly in 2018 at the “Wikidata for research” workshop in Berlin; it inspired the grants/funding part of schema.org, was adopted for the European Commission knowledge base through the OpenAire LOD service, and served as one of the bases for the Crossref GRANTID initiative schema (Chialva et al., 2020). Maintenance is described as continuous and evolutive, with many extensions requiring only subclassing for new concepts (Chialva et al., 2020).
3. DINGO as an H I survey and its observational program
In radio astronomy, DINGO denotes the Deep Investigation of Neutral Gas Origins, an ASKAP H I survey designed to trace the evolution of neutral atomic hydrogen beyond the local Universe. Forecasts for the proposed survey described a two-tier design: DINGO DEEP, with five non-contiguous fields covering over to $0.26$ and $500$ hr per field, and DINGO UDEEP, with two ultra-deep fields covering over to $0.43$ and $2500$ hr per field (Duffy et al., 2012). The survey was framed as part of a “wedding-cake” strategy with WALLABY, with DINGO providing the depth required to follow H I over the last $4$–$5$ billion years of cosmic evolution and to measure the evolution of the H I mass function and cosmic H I density (Duffy et al., 2012).
The 2012 forecasting study estimated that DINGO would detect roughly 0 galaxies in H I. At 1, the predicted counts were 2 galaxies for DEEP and 3 for UDEEP in the conservative fixed-4 model, with UDEEP rising to 5 in an alternative fixed-6 model (Duffy et al., 2012). The same study argued that the higher-resolution 36-antenna ASKAP configuration was particularly important for DINGO because it reduced maximum source confusion at the survey edge from about 7 to 8 (Duffy et al., 2012).
Early-science ASKAP commissioning data established the practical stacking program. Using 35.5 hr of ASKAP-12 observations over about 9 in the GAMA 23h field, the early-science study reported seven direct H I detections at $0.26$0—six previously known sources and one new source—and used H I spectral stacking of 3799 galaxies over $0.26$1 to measure colour-, environment-, and scaling-relation trends (Rhee et al., 2022). The reported cosmic H I densities were $0.26$2 at $0.26$3 and $0.26$4 at $0.26$5, consistent with other low-redshift measurements (Rhee et al., 2022). The same paper found that group central galaxies have larger average H I masses than satellite and isolated galaxies but lower H I gas fractions, and that ASKAP stacking reproduces known H I scaling relations while extending them to lower stellar masses and stellar surface densities (Rhee et al., 2022).
A VLA pathfinder, DINGO-VLA, was used to establish low-redshift methodological baselines for interferometric H I stacking. That study stacked 3622 galaxies extracted from 267 VLA pointings in the GAMA G09 field, obtained a $0.26$6 H I mass measurement, inferred an average H I mass of $0.26$7, and reported $0.26$8 at $0.26$9 (Chen et al., 2021). It explicitly framed this as a methodological pathfinder for the broader DINGO program and concluded that low-redshift measurements are consistent with little or no significant evolution in $500$0 over the past $500$1 Gyr (Chen et al., 2021).
A later ASKAP pilot analysis connected DINGO to halo-scale gas statistics. Using DINGO pilot 100h data with GAMA and WAVES, the 2026 study measured an H I–halo mass relation over $500$2, reported a double power-law form with turnover near $500$3, found that central galaxies dominate the halo H I budget below $500$4, and that satellites dominate above that scale (Dev et al., 29 Apr 2026). Including WAVES photometric members increased the measured H I content in halos above $500$5 by a factor of $500$6–$500$7, which the paper interprets as evidence that gas-rich faint satellites are important in the group and cluster regime (Dev et al., 29 Apr 2026).
4. DINGO as a gravitational-wave inference family
In gravitational-wave data analysis, DINGO denotes Deep INference for Gravitational-wave Observations, a deep-learning framework for rapid Bayesian parameter estimation (Dax et al., 2021). The original system uses neural posterior estimation with a conditional normalizing flow to learn a neural approximation $500$8 to the posterior, amortizing the cost of simulation and waveform generation across future events (Dax et al., 2021). A major contribution of the 2021 paper is conditioning not only on the strain data but also on a representation of the detector-noise PSD, allowing inference to adapt across events with different detector-noise conditions (Dax et al., 2021).
The baseline 2021 implementation targeted the full 15-dimensional binary black hole parameter space for precessing quasicircular BBHs under an IMRPhenomPv2 waveform model (Dax et al., 2021). It introduced GNPE: group equivariant neural posterior estimation to handle detector coalescence times and used a conditional flow with 30 coupling transforms and 5 residual blocks per transform (Dax et al., 2021). On eight GWTC-1 BBH events, DINGO produced 50,000 posterior samples in about 20 seconds with 30 GNPE iterations, compared with $500$9 for standard analyses, and achieved a mean Jensen–Shannon divergence of 0.0009 nat relative to LALInference MCMC, only slightly above the 0.0007 nat variation between repeated LALInference runs (Dax et al., 2021).
Subsequent work expanded deployability. A 2022 paper modeled future PSD distributions with a latent probabilistic model so that DINGO could be trained for a new observing run using O2 data plus a single O3 PSD, rather than waiting for the entire run to complete (Wildberger et al., 2022). On 37 real BBH events from O3, that synthetic-PSD training regime achieved average JSD 0 nat, close to an oracle model trained on real O3 PSDs and much better than a naive early-run-only baseline (Wildberger et al., 2022).
The framework later acquired a flexible transformer-based encoder in Dingo-T1, which replaced the fixed-dimensional residual encoder with a tokenized transformer over variable-length detector-frequency inputs (Kofler et al., 2 Dec 2025). That model used multibanded frequency-domain detector data and PSDs, detector identity and frequency-bound metadata, and a masking-based training objective to amortize over missing detectors, altered frequency ranges, and localized notches (Kofler et al., 2 Dec 2025). On 48 O3 events spanning 17 different detector/frequency configurations, Dingo-T1 improved median sample efficiency from 1.4% for the baseline fixed-setting Dingo NPE model to 4.2%, while enabling detector-subset studies and inspiral-merger-ringdown consistency tests with a single trained model (Kofler et al., 2 Dec 2025).
The DINGO family also expanded into specialized regimes. For LISA, a 2026 paper adapted DINGO to massive black-hole binaries in the high-mass, short-duration regime, using a conditional normalizing flow trained on IMRPhenomXHM waveforms and a low-frequency approximation to the detector response (Spadaro et al., 20 Mar 2026). That implementation produced 1 posterior samples in under a minute, remained accurate up to roughly SNR 2, and still yielded useful unbiased proposals at SNR 3 despite much lower importance-sampling efficiency (Spadaro et al., 20 Mar 2026).
Lensing became another major branch. A 2025 proof-of-principle combined DINGO with wave-optics lensing calculations from GLoW for microlensed GW signals, using a 17-parameter lensed network that extended the standard 15 source parameters by the impact parameter 4 and 5 for an isolated point-mass lens (Caldarola et al., 11 Nov 2025). The study concluded that DINGO plus importance sampling could provide efficient estimation of the background Bayes-factor distribution required for significance assessment, but warned that foreground lensed events can cause sampling efficiency to collapse when analyzed by an unlensed network (Caldarola et al., 11 Nov 2025). A later paper, DINGO-lensing, built on this line to reanalyze GW231123 and argued that its statistical significance cannot exceed 6, while reporting that 8% of GW231123-like nonlensed simulations and 58% of GW231123-like lensed simulations yield larger support for lensing than the real event (Chan et al., 18 Dec 2025).
At the catalog level, Dingo-Pop extended the DINGO philosophy from single-event inference to end-to-end population inference from gravitational-wave strain using transformers (Leyde et al., 11 May 2026). Each event is first embedded by a pretrained Dingo encoder into a low-dimensional token, then a transformer aggregates a variable-size catalog and conditions a normalizing flow over the population hyperparameters (Leyde et al., 11 May 2026). The model was trained for catalog sizes from 25 to 1000 events, passed calibration tests with a combined KS 7-value of 0.20, and produced population posteriors in about one second without per-event Monte Carlo sampling noise (Leyde et al., 11 May 2026).
5. DINGO in power systems and neutron imaging
In power-systems machine learning, DINGO denotes the DIstribution Network GeneratOr, a large open collection of synthetic medium-voltage grids used as a realism benchmark for graph generation (Abbas et al., 2 Sep 2025). The 2025 VGAE study describes it as containing 2,722 medium-voltage grid districts (MVGDs), typically corresponding to districts supplied by a single HV-MV substation, with networks of up to 40,000 nodes per grid (Abbas et al., 2 Sep 2025). The dataset is graph-structured and is used primarily for topology generation, not rich node-feature prediction (Abbas et al., 2 Sep 2025). In that paper DINGO functions as the “hard” dataset: even the best tested model, an Iterative-GCN VGAE, generated synthetic DINGO graphs with average degree 2.5300 versus the real mean 1.9986, synthetic standard deviation 1.4651 versus real 0.0115, and normalized-Laplacian Wasserstein distance 0.5072, revealing disconnected components and repeated motifs (Abbas et al., 2 Sep 2025).
In neutron imaging, Dingo is the neutron radiography and tomography beamline at ANSTO’s OPAL research reactor. A 2018 tomography paper used DINGO data to test a convex-algorithm statistical image reconstruction framework and reported that it achieved image quality similar to ramp-filtered back-projection using only 12.5% of the projections, implying a potential eight-fold increase in throughput for facilities such as DINGO (Brown et al., 2018). The beamline configuration in that study used a thermal neutron spectrum with maximum intensity at 8, beam divergence on the order of 1 mrad, and flux 9 n/(cm0s) in the high-resolution setup (Brown et al., 2018).
The beamline also served as the platform for a 2025 proof-of-principle on ghost projection for neutron beam shaping (Kingston et al., 26 Feb 2025). In that experiment, the Dingo beamline operated in a high-intensity configuration with a thermal spectrum peaking at 1, divergence 2, and average neutron flux 3 (Kingston et al., 26 Feb 2025). A 10 mm × 10 mm fractal gadolinium mask was translated over a 4 grid of positions across a 2 mm × 2 mm field of view; 2307 measured basis patterns were retained after data loss and used in a nonnegative least-squares synthesis procedure 5, with integerized weights mapped to repeated 15 s exposures (Kingston et al., 26 Feb 2025). Six target beam shapes, including a dingo paw print and dingo silhouette, were successfully projected, demonstrating programmable neutron beam shaping with a universal translatable mask (Kingston et al., 26 Feb 2025).
6. DINGO in optimization, language technology, and bioinformatics
In optimization, DINGO denotes the DIstributed Newton-type method for Gradient-norm Optimization, a communication-efficient distributed second-order algorithm for minimizing a finite-sum objective 6 (Crane et al., 2019). Its defining idea is to optimize the surrogate 7 rather than 8 directly, following the Newton-MR perspective (Crane et al., 2019). The algorithm uses three cases based on local pseudo-inverse or regularized least-squares directions, together with an Armijo-type line search on the gradient norm, and the paper proves strict reduction in 9 at every iteration regardless of the selected hyperparameters (Crane et al., 2019). The worker-side subproblems are linear least-squares or SPD solves, and the method is explicitly presented as being applicable beyond convexity and under arbitrary data partitioning (Crane et al., 2019).
In instruction-following evaluation for LLMs, DINGO is the Diverse and Fine-grained Instruction-Following evaluation dataset (Gu et al., 2024). It contains 5,026 samples organized under a 4-level, 130-node manually annotated category tree derived from 7,265 ShareGPT seed samples (Gu et al., 2024). The top level contains six categories—Language Understanding, Code, Knowledge Utilization, Creation, Language Generation, and Mathematics and Reasoning—and the dataset introduces diversity in style, attitude, and language, with a generation-and-filtering pipeline using GPT-4 and a ROUGE-L diversity threshold $0.43$0 (Gu et al., 2024). The paper argues that DINGO exposes task-level weaknesses hidden by coarse instruction-following benchmarks and shows that DINGO-style reformulations are harder than their underlying “basic questions” (Gu et al., 2024).
In diffusion language modeling, DINGO is a dynamic-programming-based constrained decoding strategy for diffusion LLMs (Suresh et al., 29 May 2025). The paper formulates constrained block decoding under regular-expression constraints as an exact dynamic program over token positions and automaton states, yielding the highest-probability valid output block under the model’s factorized blockwise distribution (Suresh et al., 29 May 2025). On GSM-Symbolic and JSON-Mode-Eval, DINGO achieved 100% parse and accuracy rates in several JSON settings and up to a 68 percentage point improvement over unconstrained inference on standard symbolic math and JSON generation benchmarks (Suresh et al., 29 May 2025). The method is explicitly presented as both efficient and provably correct for regular-language constraints in diffusion-style parallel generation (Suresh et al., 29 May 2025).
In bioinformatics, DINGO appears as an existing RNA-seq differential network method used as a comparator rather than as the subject of development (Ahn et al., 2022). A paper introducing PRANA describes DINGO as a method for finding differentially connected genes in subnetworks corresponding to different pathways between two patient groups, and benchmarks it as a univariable method that is not equipped to adjust for available covariates such as patient-age (Ahn et al., 2022). In that comparison DINGO was sometimes competitive in unconfounded simulations, but the paper emphasizes its computational burden and its low precision in an age-confounded scenario (Ahn et al., 2022).
7. Nomenclature, lineages, and recurrent themes
A common misconception is to read “DINGO” as a single framework with many applications. The literature does not support that interpretation. Only the gravitational-wave sequence forms a direct methodological lineage, beginning with DINGO for BBH parameter estimation and extending through PSD-shift adaptation, flexible transformer inference, lensing, LISA, and catalog-level population inference (Dax et al., 2021, Wildberger et al., 2022, Kofler et al., 2 Dec 2025, Spadaro et al., 20 Mar 2026, Caldarola et al., 11 Nov 2025, Chan et al., 18 Dec 2025, Leyde et al., 11 May 2026). The ontology, ASKAP survey, optimizer, power-grid dataset, instruction-following benchmark, diffusion-LLM decoder, and neutron beamline are independent uses of the same name (Chialva et al., 2020, Duffy et al., 2012, Crane et al., 2019, Abbas et al., 2 Sep 2025, Gu et al., 2024, Suresh et al., 29 May 2025, Kingston et al., 26 Feb 2025).
Despite that disunity, several recurrent design themes are visible. Many DINGO systems emphasize machine-readable structure or amortization: the ontology encodes grants and policy as linked data; the gravitational-wave DINGO family amortizes posterior inference; Dingo-Pop amortizes over catalog size; the diffusion-LLM DINGO replaces heuristic constrained decoding with exact dynamic programming; and the instruction-following benchmark DINGO converts real-world task diversity into a structured taxonomy (Chialva et al., 2020, Dax et al., 2021, Leyde et al., 11 May 2026, Suresh et al., 29 May 2025, Gu et al., 2024). Other instances serve primarily as benchmarking or infrastructure names, such as the DINGO MV-grid dataset and the Dingo neutron beamline (Abbas et al., 2 Sep 2025, Kingston et al., 26 Feb 2025).
This suggests that domain qualifiers are essential in scholarly use. “DINGO” in semantic-web funding-data integration, ASKAP neutral-hydrogen astronomy, GW simulation-based inference, diffusion decoding, distributed optimization, and neutron instrumentation refers to distinct research objects, and accurate interpretation depends on the surrounding field-specific context (Chialva et al., 2020, Duffy et al., 2012, Dax et al., 2021, Suresh et al., 29 May 2025, Crane et al., 2019, Kingston et al., 26 Feb 2025).