Papers
Topics
Authors
Recent
Search
2000 character limit reached

Inferring the Dark from the Observable: Estimating Halo Masses Using Galaxy Properties

Published 19 Aug 2026 in astro-ph.GA and astro-ph.CO | (2608.19154v1)

Abstract: The formation and evolution of galaxies are interconnected with that of their host halos. Given this closely related evolutionary history, it should follow that galaxy properties are correlated with their host halo mass. Previous works have shown that star formation rates of galaxies in cluster halos differ systematically from those in field halos, suggesting that host halo mass plays a significant role in the history of galaxy evolution. Here, we examine the correlations between observable galaxy properties - such as stellar mass, star formation rate, half-stellar radius, r−r-band magnitude, and color - and the underlying host halo mass of galaxies. Central galaxy observables are used to probe host halo mass, while satellite galaxy properties are used to estimate subhalo mass. We perform this analysis using random forest, ordinary least squares (OLS), and symbolic regression. Random forest regression can assess the relative significance of different observable galaxy properties, while OLS and symbolic regression can quantify their relationship with the host halo mass. Our results show that gas mass is generally the most significant property in central galaxies of all halo mass ranges, followed by stellar mass and observed r−r-band magnitude. Adding parameters such as the colour and magnitude of galaxies improves host halo mass estimations compared to standard stellar-to-halo mass relations. These results show that halo mass likely has a multi-dimensional dependence on galaxy properties and open the avenue of finding a simple way to constrain halo mass based on what is directly observable in galaxy surveys. Finding the observables that have the strongest correlation with halo mass can also motivate future surveys on which areas to concentrate in detail.

Authors (2)

Summary

  • The paper trains three estimators (random forest, OLS, and symbolic regression) on observable galaxy properties to predict dark matter halo masses, achieving $R^2$ values up to 0.97 for group/cluster halos
  • Gas mass is consistently the most important predictor for halo masses of the trained set. Formation redshift had a negligible impact on host-halo connections, except for group–cluster predictions.
  • When assess against observational data, measurement reveals inconsistencies which suggests a larger sample or improvements of simulation techniques would add to this analysis.

Overview

This paper by Chen and Ahad investigates the galaxy–halo connection using the TNG100 magnetohydrodynamical simulation, asking how well observable galaxy properties can predict the dark matter mass of their host halos. Rather than relying on the classical stellar-to-halo mass relation (SHMR), the authors train three estimators — random forest regression, ordinary least squares (OLS), and symbolic regression via PySR — on a suite of galaxy properties (stellar mass, gas mass, half-stellar radius, specific star formation rate, formation redshift, rr-band magnitude, and g−rg-r color) for central and satellite galaxies. Central galaxy properties are used to infer host halo mass; satellite properties are used to infer subhalo mass. The sample is split into field halos (1012−1013 M⊙10^{12} - 10^{13}\ M_\odot; 1988 halos) and group/cluster halos (≥1013 M⊙\geq 10^{13}\ M_\odot; 215 halos), with galaxies required to have M∗≥109 M⊙M_\ast \geq 10^9\ M_\odot and stellar mass within 30 kpc apertures at most 20% of total subhalo mass.

Feature importance: gas mass dominates

The random forest analysis consistently identifies gas mass as the most important predictor of host halo mass across both environments and all training sets that include it, followed by stellar mass, g−rg-r color, and rr-band magnitude. This is physically consistent: baryons constitute a large fraction of halo mass, and only roughly ≤20%\leq 20\% of accreted gas cools into stars. PySR symbolic regression independently corroborates this ranking — when two methods agree, the authors treat it as evidence of genuine correlation rather than an artifact of one algorithm.

Two secondary findings deserve attention. First, removing gas mass from the feature set substantially elevates the importance of stellar mass, particularly for field centrals, consistent with the known peak of the stellar mass fraction at ∼1012 M⊙\sim 10^{12}\ M_\odot. Second, formation redshift contributes little predictive power for halos below 1014 M⊙10^{14}\ M_\odot, implying that present-day halo mass is weakly tied to assembly history in this mass regime. For satellites, however, the ordering differs: in field subhalos, stellar mass outranks gas mass, suggesting a genuinely different galaxy–subhalo relation than the central–host case.

Predictive performance

The headline quantitative result is that multi-dimensional relations outperform the standard SHMR everywhere:

Sample Features Method g−rg-r0
Field Stellar mass only OLS 0.64
Field Observables only Random forest 0.68
Field Observables + gas Random forest 0.88
Group+cluster Stellar mass only OLS 0.81
Group+cluster All parameters OLS 0.96
Group+cluster All parameters PySR 0.97

For group/cluster halos, adding observables raises g−rg-r1 from 0.81 to 0.85–0.97 depending on method; PySR consistently edges out OLS by 2–5 points, indicating the underlying scaling relations are nonlinear. Bootstrap error analysis shows predicted-to-true mass ratios centered near 1.0 for field halos but 1.3 for group/cluster predictions with observables-only training — a difference exceeding g−rg-r2. The authors attribute the higher field scatter to greater diversity in galaxy properties, while the group/cluster bias likely reflects small-sample effects.

A notable exploratory result concerns cluster-scale halos specifically: within the 17 clusters available, gas mass importance drops and formation redshift importance rises sharply relative to groups. The authors explicitly caution that this is not statistically significant, though they flag it as behavior worth testing in larger-volume simulations.

The richness test adds an important caveat to any purely photometric approach: when the number of member galaxies is included as a feature, it becomes by far the most important predictor of host halo mass in groups and clusters, consistent with established mass–richness relations.

Comparison against observational data

When applied to group/cluster halo masses from Yang et al. catalog measurements, the TNG-trained model matches observed masses only for part of the sample and systematically underpredicts the rest. Cross-comparison of property distributions between TNG100, EAGLE, and observations reveals systematic offsets of order g−rg-r3 dex in color-based quantities, plausibly attributable to over-cooling in simulations inflating sSFRs and stellar masses. The paper concedes plainly that it cannot determine whether the discrepancy stems from the trained model or from observational mass uncertainties (including possible misclassification of centrals). Only a fraction of overlapping predictions supports feasibility, not validated accuracy, of the simulation-to-survey transfer.

Limitations and open questions

Several limitations bound the results. The intrinsic scatter in predicted versus true masses is substantial (g−rg-r4 across bootstrap samples), which the authors acknowledge as "not ideal" and attribute to environmental/assembly-history variance and degeneracies among correlated features such as sSFR–color and magnitude–color; disentangling these degeneracies remains unresolved. Cluster-scale conclusions rest on only 17 objects and may reflect simulation bias. Subhalo mass estimates show a radial trend — overprediction near host centers, underprediction farther out — whose cause is unidentified. Halo boundary definition ambiguity and observational incompleteness in richness estimation are flagged as unsolved complications. Finally, no field-halo validation against observations was performed because weak-lensing signals are too weak and no large rotation-curve mass samples exist publicly.

Conclusion

This work demonstrates empirically that host halo mass depends multidimensionally on galaxy properties: gas mass is the dominant tracer, with stellar mass, color, and magnitude providing complementary information, and multi-parameter models uniformly outperform SHMR-based inference (e.g., g−rg-r5 vs. 0.81 for groups/clusters). The main open questions left by the paper are whether the anomalous cluster-scale feature importances hold with larger cluster samples, how to correct simulation–observation offsets in color and stellar mass distributions before deploying these relations on survey data, and whether richness can be incorporated into a unified multi-dimensional mass estimator given observational completeness limits.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.