RGZ EMU: Hybrid AI and Citizen Science for EMU
- Radio Galaxy Zoo EMU is a hybrid workflow that combines volunteer classifications with deep learning to accurately identify extended and complex radio sources.
- It leverages crowdsourced labels and expert golden samples to train active-learning models, thereby improving catalog reliability and host associations.
- The system integrates multi-wavelength data and a semantic tagging taxonomy to overcome challenges in automated cataloging, enhancing discovery of rare radio morphologies.
Radio Galaxy Zoo EMU (RGZ EMU) is a hybrid citizen-science and artificial-intelligence workflow developed in the ASKAP Evolutionary Map of the Universe (EMU) program to identify, assemble, and classify extended and morphologically complex radio sources that automated cataloging pipelines handle incompletely. Its stated role is to generate high-quality classifications from EMU data, use those classifications to train deep-learning models, and feed the resulting products into the open-science EMU catalogue EMUCAT, with particular emphasis on radio structures and host-galaxy associations that remain difficult for fully automated methods (Vardoulaki et al., 24 Sep 2025).
1. Survey context and scientific rationale
RGZ EMU is embedded in the observational regime defined by ASKAP and EMU. ASKAP consists of 36 × 12 m antennas fitted with phased-array feeds, yielding a ∼30 deg² instantaneous field of view at GHz frequencies. EMU is a wide-field continuum survey covering the entire southern sky up to +30° declination and is intended to detect tens of millions of radio sources, including AGN, star-forming galaxies, relics, cluster emission, and Odd Radio Circles. For the EMU Pilot Survey data used in Phase I, the reported parameters are rms 25–30 μJy beam⁻¹, resolution 11–18 arcsec, and 287 555 radio components found by Selavy (Vardoulaki et al., 24 Sep 2025).
Within that setting, RGZ EMU is defined as a complement to EMUCAT rather than a replacement for it. A launch-stage project description states that EMUCAT performs very well on compact radio sources, yet “struggles with extended objects,” whereas the Phase I paper states more generally that automatic pipelines face challenges in recognising complex radio structures and accurately associating radio emission with their host galaxies. The mission of RGZ EMU is therefore to produce high-quality classifications of extended and morphologically complex radio sources and feed them into EMUCAT (Tang et al., 19 Jun 2025).
The project also inherits a methodological problem identified in earlier Radio Galaxy Zoo machine-learning work. Alger et al. recast radio/host cross-identification as a binary classification and weighted-ranking task, showed that a nearest-galaxy heuristic can be competitive on many sources, and concluded that there were not enough complex examples in the ATLAS data to learn accurate cross-identification for extended morphologies. They further showed that training on Radio Galaxy Zoo crowd labels gave comparable results to training on expert labels. This established two premises that RGZ EMU adopts directly: crowdsourced labels can be scientifically useful, and substantially larger labeled sets are required for extended-source cataloging at EMU scale (Alger et al., 2018).
2. Phase I sample construction and subject presentation
Phase I of RGZ EMU selected 6 230 extended sources, provided as 6′×6′ cutouts spanning simple to highly irregular morphologies. A project description of the sample-selection pipeline states that the process began with ∼2 × 10⁵ entries from the EMU PS 1 Selavy component catalog, generated 6′ × 6′ radio-continuum cutouts centred on each Selavy component, computed a “complexity” metric , retained images with to obtain 37 578 cutouts, and then required that the principal Selavy component have major axis ≥ 20″, reducing the Phase I set to 6 230 subjects. The Phase I paper summarizes the same filtering more concisely, stating that subjects are pre-filtered by “image complexity” and Selavy angular size so that volunteers focus on the most challenging cases (Tang et al., 19 Jun 2025).
The complexity measure is described in two related but differently granular ways across the project literature. One account calls it a “compression-based proxy” for complexity and says that sources are scored for “complexity” to prioritize human inspection; another describes it as a coarse-grained Kolmogorov complexity proxy. Neither description supplies an explicit anomaly-scoring formula in the Phase I paper, which instead refers readers to Segal et al. (2023) for methodological details. This suggests that RGZ EMU uses anomaly-style triage primarily as a work-allocation mechanism rather than as a published end-product metric in Phase I (Vardoulaki et al., 24 Sep 2025).
Published descriptions of the interface reflect different project stages. The Phase I paper states that volunteers are shown 6′×6′ JPEG cutouts combining ASKAP radio, DES optical, and WISE 3.4 μm infrared on the Zooniverse platform. A launch-stage framework paper describes a three-scale interface in which each volunteer sees 3′, 6′, and 12′ overlaid views containing ASKAP EMU radio contours, WISE 3.4 μm infrared, and DES optical; it also notes SKAO and NADC mirrors. The coexistence of these descriptions indicates an evolving presentation layer while preserving the same core multi-wavelength visual context (Vardoulaki et al., 24 Sep 2025).
3. Citizen-science workflow, training, and consensus formation
The volunteer workflow centers on three tasks. For each subject, participants group related radio components into single sources, identify the host galaxy at optical or infrared wavelengths, and assign one or more simple morphological tags drawn from a 10-term subset of a 22-term semantic taxonomy. A parallel project description formulates the same workflow as radio-source assembly, morphology tagging, and optional free-text discussion on a Talkboard. In that formulation, volunteers draw regions associating lobes and jets and connect them to an optical or infrared host, then choose from 10 high-level semantic tags, with examples including “FR I,” “bent-double,” “diffuse halo,” and “relic” (Vardoulaki et al., 24 Sep 2025).
The training burden imposed on volunteers is intentionally modest. The Phase I paper states that no formal expert-level training is required beyond an illustrated tutorial and examples on the platform. At the same time, a launch-stage description says that a small set of “golden samples” classified by experts is interleaved to measure volunteer accuracy and calibrate consensus thresholds. These two statements are compatible: the platform does not require expert training of participants, but it does embed expert-labeled control material in the workflow (Tang et al., 19 Jun 2025).
Participation statistics reported in project documents are stage-dependent. The Phase I paper reports that at least 2 500 volunteers have contributed ∼97 000 independent classifications, equivalent to ≈39 per source if evenly distributed, although the paper notes that actual per-subject counts vary. A launch-stage report, explicitly described as ≈1 month post-launch, reports 1 435 volunteers and more than 53 000 classifications. The published record therefore captures the same project at different maturity levels rather than a disagreement about scope (Vardoulaki et al., 24 Sep 2025).
Consensus formation is described at a functional rather than fully algorithmic level. The Phase I paper states that once “a sufficient number of independent classifications is obtained, the collective accuracy of volunteers approaches expert consensus,” in a manner consistent with Willett et al. (2016), but it does not specify the exact consensus algorithm. A companion framework description states that individual annotations are aggregated by majority vote or weighted consensus, that only subjects whose consensus surpasses a predefined level are retired, and that their aggregated labels enter the EMUCAT supplementary catalog. The absence of a published threshold or finalized aggregation rule in the Phase I paper is therefore an explicit limitation rather than an omission to be inferred away (Tang et al., 19 Jun 2025).
4. Semantic morphology taxonomy and the role of NLP
A distinctive feature of RGZ EMU is its semantic tagging system. The underlying taxonomy was derived through an NLP-driven process based on 8,486 plain English annotations collected from 299 EMU Pilot Survey cutouts, supplemented by 1,257 expert classifications across a 22-class scheme. The methodology embedded each annotation with SpaCy’s large English model, averaged token embeddings into sentence vectors, clustered semantically similar annotations using cosine similarity with threshold , mapped aggregated vectors back to nearest tokens, represented each source by a binary tag vector, and used one-vs-rest random forests plus SHAP values to weight tag importance across expert-defined science classes (Bowles et al., 2023).
The resulting taxonomy comprises 22 semantic tags. The 10 tags proposed for citizen-science use are amorphous, bent, bridge, core, hourglass, jet, lobe, merger, plume, and tail. The 12 tags intended for algorithmic assignment are asymmetric brightness, asymmetric structure, compact, diffuse, double, edge brightened, extended, faint, host, peak, small, and traces host galaxy. Because each source may carry any subset of these 22 tags, the taxonomy supports possible tag combinations, allowing rare feature conjunctions to be represented without forcing a source into a single exclusive morphology class (Bowles et al., 2023).
This compositional view of morphology is central to RGZ EMU’s design. The taxonomy paper argues that plain-English tags are more flexible and more sensitive to rare feature combinations than traditional rigid class labels. A plausible implication is that RGZ EMU treats morphology as a multi-label semantic space that can serve both human annotation and downstream machine learning. The published descriptions are not completely identical, however: the taxonomy paper places “traces host galaxy” among the algorithmically assigned tags, whereas the Phase I paper lists it as an example within the volunteer tagging workflow. That discrepancy is best understood as a difference between the original taxonomy design and later interface practice, not as a settled redefinition (Vardoulaki et al., 24 Sep 2025).
5. Machine-learning architecture, anomaly triage, and validation status
RGZ EMU is designed as a human-in-the-loop training pipeline. The Phase I paper states that the ∼97 000 volunteer classifications on 6 230 sources form the initial labeled training set for active-learning loops, and that as machine-learning models improve they will select new sources for human labeling in a quasi-automated framework. A companion framework paper expresses the supervised objective as
where are the network parameters, is the pixel-region plus multi-wavelength input, and is the consensus label, such as a morphology tag or host coordinate. The same paper advocates an active-learning “foundational + downstream” paradigm in which foundational models learn general source segmentation and host association, while downstream models specialize to narrower tasks (Tang et al., 19 Jun 2025).
The Phase I paper is explicit about what has not yet been disclosed. It states that RGZ EMU will use deep-learning architectures trained on citizen-science “golden samples,” but it does not provide specific model architectures, such as CNN layer layouts or loss functions beyond the generic supervised formulation. For anomaly detection, it states that sources are scored for “complexity” via a compression-based proxy, again without explicit autoencoder or mixture-model formulas in the paper itself. For NLP, it notes that the semantic taxonomy was derived through an NLP-driven process, but it does not specify embedding models or clustering choices there; those details appear in the taxonomy work. Planned enhancements include more sophisticated deep networks, explicitly including multi-scale CNNs with attention, improved anomaly detection with explicit autoencoders or Gaussian-mixture latent-space models, and iterative active learning in which model uncertainties seed new volunteer tasks (Vardoulaki et al., 24 Sep 2025).
Validation claims are similarly stratified by publication stage. The Phase I paper does not report precision, recall, , or confusion-matrix results, although it gives the standard definitions
0
and states that a forthcoming paper by Tang & Vardoulaki will present these metrics. It adds that early-stage comparisons between 2 500 volunteers and automated pipelines demonstrate improved recovery of complex and rare morphologies, but provides no numerical values or uncertainty estimates. By contrast, the launch-stage framework paper reports volunteer accuracies in excess of 90 % on simple compact sources and ∼80 % on moderately extended double sources, while noting that exact confusion matrices are still being compiled; it also cites early RG-CAT tests in pilot fields suggesting completeness of ∼95 % and purity of ∼92 % for FR-II lobed sources down to a 10σ surface-brightness threshold. These figures should therefore be read as preliminary project-stage results rather than finalized Phase I catalog metrics (Tang et al., 19 Jun 2025).
6. EMUCAT integration, open-science products, and projected scale
RGZ EMU’s catalog outputs are intended for direct incorporation into EMUCAT. The Phase I paper states that the final products will be integrated into EMUCAT and that, although exact schemas are still under development, the intention is to deliver science-ready tables in VO Table and FITS binary table formats containing radio–optical associations, morphology tags, and quality flags. It also states that data-release protocols, API endpoints, and formats will follow EMU open-science policies. A complementary framework description says that all RGZ EMU retired labels and regions will be ingested into EMUCAT as a dedicated “Extended Sources Supplement” alongside the standard Selavy-derived tables (Vardoulaki et al., 24 Sep 2025).
The project’s scientific contribution is framed in terms of catalog completeness, host-association reliability, and discovery space. By leveraging human pattern recognition on ∼6 000 challenging sources, RGZ EMU is intended to produce greater completeness for extended and complex sources, higher reliability in host association, and enhanced discovery potential for unusual objects such as ORCs. The same logic extends to external survey products: RGZ EMU is designed to ingest and cross-match value-added data from POSSUM, PEGASUS, WALLABY, and Euclid. In that broader scheme, cross-identification is formulated probabilistically as
1
with 2 denoting radio regions and 3 an external catalog. This indicates that RGZ EMU is not limited to morphology tagging but is also conceived as a multi-wavelength association framework (Tang et al., 19 Jun 2025).
Scaling arguments are central to the project’s rationale. Over five years, RGZ EMU aims to classify ∼4 million extended sources in the full EMU survey footprint. A launch-stage paper states that pure citizen classification of ∼4 × 10⁶ subjects would take ≃157 years at the volunteer rate of the original Radio Galaxy Zoo effort, hence the requirement for AI assistance. The same paper argues that retiring high-confidence subjects early and sending only the most difficult ∼10–20 % through active-learning cycles reduces volunteer workload by an order of magnitude. Infrastructure scaling in compute, storage, and server load is already being tested, and planned compute resources include a hybrid cloud/HPC deployment at Jodrell Bank and NADC with on the order of 100 GPUs for periodic retraining every few weeks (Tang et al., 19 Jun 2025).
Broader scientific applications are also explicitly identified. Project descriptions mention integration with POSSUM polarisation, WALLABY H I, PEGASUS, and Euclid for multi-wavelength studies of AGN feedback, galaxy evolution, and large-scale structure, as well as education and outreach via the RADIIO programme. This suggests that RGZ EMU functions simultaneously as a catalog-production system, a training-data generator for radio-astronomical machine learning, and an open-science interface between large survey infrastructures and human visual expertise (Vardoulaki et al., 24 Sep 2025).