FrankenBlast: Fast Host-Galaxy Analysis
- The paper introduces FrankenBlast, a high-throughput pipeline that integrates probabilistic host association, multi-survey aperture photometry, and simulation-based SED inference.
- FrankenBlast employs the novel Pröst module, which improves host matching by incorporating brightness, redshift, and angular offset to derive robust galaxy associations.
- The pipeline achieves a ~500× speedup over traditional Bayesian methods, making it ideal for Rubin-era surveys and large-scale transient host characterization.
Searching arXiv for FrankenBlast and related tools to anchor citations. FrankenBlast is a fast, end-to-end host-galaxy analysis pipeline for optical transients, presented as a customized and improved successor to the Blast web application. Its operational scope is to take a newly discovered transient, identify the most likely host galaxy, measure that host’s broadband photometry from archival survey imaging, and infer key stellar population properties—especially redshift, stellar mass, star-formation rate, metallicity, age, and dust extinction—on a timescale of minutes per object. In the reported implementation, FrankenBlast was tested on 14,432 supernovae and constrained host properties for 9,262 events, with the study framed explicitly around the Rubin Observatory and Roman regimes, in which most transients will be photometrically classified and full spectroscopic follow-up will be infeasible (Nugent et al., 10 Sep 2025).
1. Origins, scope, and relation to Blast
FrankenBlast is described as a “customized and improved version” of Blast, retaining the latter’s role as a live web application for transient host characterization while updating several core components. Blast originally used GHOST-style host association and the SBI++ stellar population inference framework. FrankenBlast modifies the host-association machinery, extends and refines the SED-fitting workflow, and adds a photometric-redshift-capable model. The authors state that these changes will be folded back into future Blast versions, and they present FrankenBlast itself as publicly available together with all data products and a developing Rubin-focused version (Nugent et al., 10 Sep 2025).
The motivation is scalability rather than methodological novelty in isolation. FrankenBlast integrates probabilistic host association, multi-survey aperture photometry, and ultra-fast SED-based inference into one automated workflow. A plausible implication is that the pipeline is intended not merely as an analysis environment for retrospective samples, but as an operational system for high-throughput transient streams.
The paper also places FrankenBlast in a practical lineage. Relative to earlier Blast behavior, the main changes are not cosmetic: host selection becomes probabilistic, redshift can be incorporated directly in the association step, and stellar population inference is extended to a redshift range intended to better match both the current transient-host sample and Rubin-era requirements.
2. Host-galaxy association and photometric extraction
FrankenBlast begins with host-galaxy association using Pröst, a Python host-association code that improves on older methods by incorporating galaxy brightness and redshift probabilistically in addition to angular or offset information. The earlier Blast association approach was based on GHOST and used a modified directional light radius (DLR) method, but did not treat host selection probabilistically or include redshift as a prior. In FrankenBlast, the DLR-based fractional offset is the transient’s angular offset divided by the galaxy’s directional light radius, defined as the half-light radius in the direction of the transient. Pröst samples candidate hosts in a search cone and constructs a Monte Carlo posterior from priors and likelihoods over redshift, fractional offset, and host absolute magnitude; it also accounts explicitly for the possibility that the true host is unobserved by integrating over a simulated population of missing hosts and renormalizing the total posterior (Nugent et al., 10 Sep 2025).
For spectroscopic SN hosts, the adopted setup uses a uniform redshift prior over , a 5% nominal SN redshift error, a uniform fractional-offset prior between 0 and 10 DLR, a Gamma likelihood in offset with , and a broad magnitude prior or likelihood from mag. For photometric SN hosts, redshift dominates so strongly that the analysis effectively relies only on fractional-offset and magnitude terms. Associations are attempted against GLADE, Pan-STARRS, and DECaLS DR10, and the highest-probability host across catalogs is retained.
After association, FrankenBlast performs global aperture photometry from archival imaging in GALEX, Pan-STARRS, DECaLS DR9, 2MASS, and WISE, for up to 17 bands total: GALEX FUV/NUV, PS1 , DECaLS , 2MASS , and WISE 1–4. Apertures are constructed with Astropy Photutils, using Kron apertures from segmentation maps for sources detected at with at least 10 connected pixels. If a band is not detected, the code estimates an upper limit by rescaling the aperture from a neighboring filter by that filter’s FWHM. The authors contrast this with Blast, which measured the aperture in a single optical band and rescaled all others from there, and they argue that FrankenBlast’s strategy should reduce background noise and improve , while Blast’s approach may recover more total flux in non-optical bands.
Photometric uncertainties are treated as Poisson-dominated and calibrated from filter zeropoints, with an additional correlated-noise correction for WISE apertures. The pipeline does not explicitly deblend or correct for contamination; the paper treats this as acceptable for the relatively sparse surveys used in the study, while noting that deblending will become critical for Rubin and Roman. This suggests that source confusion, rather than host-association latency, may become the dominant technical bottleneck in denser future imaging.
3. Stellar population inference with SBI++
The stellar population modeling stage is built around SBI++, a simulation-based inference framework using neural posterior estimation. FrankenBlast adopts SBI++ because conventional Bayesian SED-fitting tools such as Prospector, BEAGLE, or BAGPIPES can require hours to days per object, which is incompatible with large transient-host samples. The paper highlights two specific advantages: handling of missing and noisy photometry, and a speedup of roughly relative to traditional Bayesian SED fitting, with runtime less than minutes on a single CPU (Nugent et al., 10 Sep 2025).
The training set consists of 2 million mock galaxies generated from Prospector models with MIST stellar evolution tracks, MILES stellar libraries, FSPS/python-FSPS, a Chabrier IMF, the Kriek & Conroy dust law, nebular emission, AGN torus emission, Draine & Li dust IR emission, and a non-parametric continuity SFH with seven age bins. The model samples redshift 0, total mass formed 1, stellar metallicity 2, dust optical depths 3 and 4, the dust-law offset dust_index, SFH bin-to-bin SFR ratios, AGN parameters 5 and 6, and dust emission parameters 7, 8, and 9. A notable update relative to Blast is extension of the redshift range to 0 instead of 1.
Training photometry is “noised up” to match observed-host 2, with a 1% error floor. The model is trained with the sbi package, yielding two inference networks: a spec-3 model for fixed or known redshift and a photo-4 model for simultaneous redshift inference. Training requires about 2 days for the spec-5 model and 5 days for the photo-6 model on a single CPU.
FrankenBlast treats observational incompleteness explicitly. For noisy bands, it samples simulated noisy realizations, truncates them to the neighborhood of similar training SEDs, fits each realization, and averages the posteriors. For missing bands, it uses nearest-neighbor SEDs to build a KDE prior over plausible missing photometry, samples from that prior, fits each sample, and again averages the posteriors. If both missing and noisy data are present, the procedures are combined. The implementation generates 50 realizations per noisy or missing band and 50 posterior draws per realization, yielding 2500 posterior draws per object; the same 2500-draw convention is used even for clean fits.
Derived host properties include an SFH, a present-day SFR integrated over the last 100 Myr, a mass-weighted age 7, stellar mass 8, and 9-band extinction 0. The paper defines sSFR as the ratio of the present-day SFR to 1, and gives
2
The stellar mass is described as being converted from 3 using the inferred SFH, IMF, and metallicity 4. The training set also incorporates the Gallazzi et al. mass-metallicity prior for 5 as a function of 6.
4. Validation, throughput, and sample scale
The paper validates FrankenBlast at several stages. Pröst typically returns a host in a few seconds, with latency dominated by remote catalog queries. Across the full sample of 14,432 supernovae, FrankenBlast finds confident host associations for 13,111 objects. On the YSE DR1 validation subset, it successfully associates 1,897 of 1,975 events, and comparison with earlier work suggests only 7 disagreement where both methods found hosts. For stellar population inference, the authors compare 100 randomly chosen hosts from the spectroscopic sample against full Prospector nested-sampling fits using dynesty; they conclude that SBI++ and Prospector broadly agree, but that Prospector can lock onto unrealistic metallicities or underestimate UV flux, producing biased 8, 9, and redshift-dependent quantities, whereas SBI++ gives broader but better-calibrated redshift posteriors in the photo-0 comparison (Nugent et al., 10 Sep 2025).
The study scale is summarized by the following reported counts:
| Stage | Reported quantity | Count |
|---|---|---|
| Input sample | Total transients | 14,432 |
| Input sample | Spectroscopically classified SNe | 6,676 |
| Input sample | Photometrically classified SNe | 7,756 |
| Host association | Confident host associations | 13,111 |
| Photometry | Usable photometry for stellar population modeling | 11,153 |
| Inference | Host properties successfully constrained | 9,262 |
The underlying transient set comes mainly from ZTF and YSE DR1. The ZTF spectroscopic sample comprises 6,199 SNe after removing non-SNe and hydrogen-rich SLSNe; the ZTF photometric sample contains 6,283 SNe-like transients; and YSE DR1 contributes 1,975 transients total, with 477 spectroscopic and 1,485 photometric objects.
The remaining inference failures are attributed mainly to hosts lacking usable optical photometry or being too sparse, noisy, or faint for reliable SED fitting. A plausible implication is that incompleteness in derived host-property catalogs is driven primarily by data quality constraints rather than by failure of the probabilistic association formalism.
5. Supernova host demographics and class dependence
A principal scientific result is that photometrically classified supernovae are systematically at higher redshift than spectroscopically classified supernovae, which by itself shifts redshift-dependent host-property distributions such as stellar mass and sSFR. The paper further reports that photometric samples have systematically higher host stellar masses than spectroscopic samples. Using Anderson-Darling tests on 2500 posterior realizations, the authors reject the null hypothesis that the distributions are the same for essentially all classes. When the photometric sample is refined by requiring classification pseudo-probability 1, high host-association confidence, and a new host-choice probability cut, the mass distributions move closer to those of the spectroscopic sample, but residual differences remain. The paper interprets these discrepancies as arising from a combination of misclassified transients contaminating the photometric sample, mis-associated or undetected faint hosts, and redshift-driven selection effects (Nugent et al., 10 Sep 2025).
This point directly addresses a common interpretive error: differences between spectroscopic and photometric host distributions are not treated as straightforward evidence for intrinsic environmental differences. The paper argues instead that contamination and selection are primary drivers. A particularly explicit case concerns photometric SNe Ib/c, for which some high-mass, high-sSFR, dusty hosts resemble SN Ia environments more than Ib/c environments; the authors argue that these events are likely misclassified SNe Ia, with dust-reddened light curves making them appear Ib/c-like.
Using the spectroscopic sample as the cleanest basis for class-by-class comparison, the paper concludes that all SN populations seem to depend on both 2 and SFR, but with different relative weights. SNe II and IIn are described as somewhat more SFR-dependent than SNe Ia and Ib/c, while SNe Ia are more 3-dependent than all other classes.
For SNe Ia, hosts are reported to be more massive, older, more metal-rich, and less star-forming than CCSN hosts, and to occupy more mass-weighted SFR-4 distributions. The paper relates this to delayed white-dwarf progenitor scenarios and long delay times.
For SNe Ib/c versus SNe II, the distinction is emphasized as especially intriguing. SNe Ib/c hosts are more massive and less star-forming than SNe II hosts, and they also prefer more mass-weighted SFR-5 distributions. The authors speculate that SNe Ib/c must be more dependent on higher 6 and more evolved environments for the right conditions for progenitor formation. They note that a single-star Wolf-Rayet channel could naturally produce such an environment dependence via metallicity-driven winds, whereas a binary channel is harder to reconcile with a strong mass or metallicity dependence, though not impossible.
For SNe II versus SNe IIn, the hosts are broadly similar in mass, sSFR, age, and metallicity, but SNe IIn hosts are dustier. The paper finds little evidence that their mass or sSFR distributions differ, but reports a statistically significant difference in 7, with IIn hosts being dustier. The authors note that this may reflect either a real metallicity or dust preference or an observational bias, since SNe IIn and Ia are generally brighter and can be detected in dustier environments more readily than SNe II and Ib/c.
The study also reports a small but nonzero fraction of CCSNe in quiescent hosts: 9% of SNe Ib/c hosts, 3% of SNe II hosts, and only about one SNe IIn host. The interpretation offered is residual star formation in apparently quiescent galaxies rather than the necessity of exotic progenitor channels.
6. Public release and Rubin-era significance
FrankenBlast is presented as a practical high-throughput system: host association occurs in seconds, full host-property inference in minutes, and the SED-fitting stage is roughly 8 faster than conventional Bayesian methods. The paper states that FrankenBlast and all data products are public, with a GitHub repository and a Zenodo archive, and it also mentions a developing Rubin Observatory version built with Rubin 9 filters and cosmoDC2-based SNR training (Nugent et al., 10 Sep 2025).
The stated scientific use cases are photometric transient classification, anomaly discovery, host-based demographic studies, and progenitor inference in Rubin- and Roman-era surveys. In that sense, FrankenBlast is best understood as a production-grade probabilistic host-characterization pipeline: it merges host association, broadband host photometry, and fast galaxy SED inference into a single workflow that scales to survey-sized samples.
Its broader significance lies in the combination of methodological and astrophysical results. Methodologically, the pipeline demonstrates that probabilistic association and simulation-based stellar population inference can be executed at throughput compatible with large transient surveys. Astrophysically, its main conclusion is that differences between spectroscopic and photometric supernova samples are driven by both contamination and redshift selection, while intrinsic class dependence of host environments follows a structured hierarchy: SNe Ia occupy more massive, older, less star-forming hosts than CCSNe; SNe Ib/c are associated with more massive and more evolved environments than SNe II; and SNe IIn differ most clearly through dustier hosts.