- The paper trains three estimators (random forest, OLS, and symbolic regression) on observable galaxy properties to predict dark matter halo masses, achieving $R^2$ values up to 0.97 for group/cluster halos
- Gas mass is consistently the most important predictor for halo masses of the trained set. Formation redshift had a negligible impact on host-halo connections, except for group–cluster predictions.
- When assess against observational data, measurement reveals inconsistencies which suggests a larger sample or improvements of simulation techniques would add to this analysis.
Overview
This paper by Chen and Ahad investigates the galaxy–halo connection using the TNG100 magnetohydrodynamical simulation, asking how well observable galaxy properties can predict the dark matter mass of their host halos. Rather than relying on the classical stellar-to-halo mass relation (SHMR), the authors train three estimators — random forest regression, ordinary least squares (OLS), and symbolic regression via PySR — on a suite of galaxy properties (stellar mass, gas mass, half-stellar radius, specific star formation rate, formation redshift, r-band magnitude, and g−r color) for central and satellite galaxies. Central galaxy properties are used to infer host halo mass; satellite properties are used to infer subhalo mass. The sample is split into field halos (1012−1013 M⊙​; 1988 halos) and group/cluster halos (≥1013 M⊙​; 215 halos), with galaxies required to have M∗​≥109 M⊙​ and stellar mass within 30 kpc apertures at most 20% of total subhalo mass.
Feature importance: gas mass dominates
The random forest analysis consistently identifies gas mass as the most important predictor of host halo mass across both environments and all training sets that include it, followed by stellar mass, g−r color, and r-band magnitude. This is physically consistent: baryons constitute a large fraction of halo mass, and only roughly ≤20% of accreted gas cools into stars. PySR symbolic regression independently corroborates this ranking — when two methods agree, the authors treat it as evidence of genuine correlation rather than an artifact of one algorithm.
Two secondary findings deserve attention. First, removing gas mass from the feature set substantially elevates the importance of stellar mass, particularly for field centrals, consistent with the known peak of the stellar mass fraction at ∼1012 M⊙​. Second, formation redshift contributes little predictive power for halos below 1014 M⊙​, implying that present-day halo mass is weakly tied to assembly history in this mass regime. For satellites, however, the ordering differs: in field subhalos, stellar mass outranks gas mass, suggesting a genuinely different galaxy–subhalo relation than the central–host case.
The headline quantitative result is that multi-dimensional relations outperform the standard SHMR everywhere:
| Sample |
Features |
Method |
g−r0 |
| Field |
Stellar mass only |
OLS |
0.64 |
| Field |
Observables only |
Random forest |
0.68 |
| Field |
Observables + gas |
Random forest |
0.88 |
| Group+cluster |
Stellar mass only |
OLS |
0.81 |
| Group+cluster |
All parameters |
OLS |
0.96 |
| Group+cluster |
All parameters |
PySR |
0.97 |
For group/cluster halos, adding observables raises g−r1 from 0.81 to 0.85–0.97 depending on method; PySR consistently edges out OLS by 2–5 points, indicating the underlying scaling relations are nonlinear. Bootstrap error analysis shows predicted-to-true mass ratios centered near 1.0 for field halos but 1.3 for group/cluster predictions with observables-only training — a difference exceeding g−r2. The authors attribute the higher field scatter to greater diversity in galaxy properties, while the group/cluster bias likely reflects small-sample effects.
A notable exploratory result concerns cluster-scale halos specifically: within the 17 clusters available, gas mass importance drops and formation redshift importance rises sharply relative to groups. The authors explicitly caution that this is not statistically significant, though they flag it as behavior worth testing in larger-volume simulations.
The richness test adds an important caveat to any purely photometric approach: when the number of member galaxies is included as a feature, it becomes by far the most important predictor of host halo mass in groups and clusters, consistent with established mass–richness relations.
Comparison against observational data
When applied to group/cluster halo masses from Yang et al. catalog measurements, the TNG-trained model matches observed masses only for part of the sample and systematically underpredicts the rest. Cross-comparison of property distributions between TNG100, EAGLE, and observations reveals systematic offsets of order g−r3 dex in color-based quantities, plausibly attributable to over-cooling in simulations inflating sSFRs and stellar masses. The paper concedes plainly that it cannot determine whether the discrepancy stems from the trained model or from observational mass uncertainties (including possible misclassification of centrals). Only a fraction of overlapping predictions supports feasibility, not validated accuracy, of the simulation-to-survey transfer.
Limitations and open questions
Several limitations bound the results. The intrinsic scatter in predicted versus true masses is substantial (g−r4 across bootstrap samples), which the authors acknowledge as "not ideal" and attribute to environmental/assembly-history variance and degeneracies among correlated features such as sSFR–color and magnitude–color; disentangling these degeneracies remains unresolved. Cluster-scale conclusions rest on only 17 objects and may reflect simulation bias. Subhalo mass estimates show a radial trend — overprediction near host centers, underprediction farther out — whose cause is unidentified. Halo boundary definition ambiguity and observational incompleteness in richness estimation are flagged as unsolved complications. Finally, no field-halo validation against observations was performed because weak-lensing signals are too weak and no large rotation-curve mass samples exist publicly.
Conclusion
This work demonstrates empirically that host halo mass depends multidimensionally on galaxy properties: gas mass is the dominant tracer, with stellar mass, color, and magnitude providing complementary information, and multi-parameter models uniformly outperform SHMR-based inference (e.g., g−r5 vs. 0.81 for groups/clusters). The main open questions left by the paper are whether the anomalous cluster-scale feature importances hold with larger cluster samples, how to correct simulation–observation offsets in color and stellar mass distributions before deploying these relations on survey data, and whether richness can be incorporated into a unified multi-dimensional mass estimator given observational completeness limits.