---
title: Sample Abundance in Diverse Scientific Domains
url: https://www.emergentmind.com/topics/sample-abundance
type: topic
---

# Sample Abundance in Diverse Scientific Domains

Across the cited literature, the expression **sample abundance** appears in several technically distinct senses. In one-bit and few-bit signal processing, it denotes a regime in which very large numbers of low-precision measurements compensate for coarse quantization and convert recovery into an overdetermined linear-feasibility problem. In astronomy, it denotes chemical abundances derived for a defined sample of stars, H II regions, AGN narrow-line regions, or open clusters. In ecology, paleobiology, and microbiomics, it denotes either the statistical recovery of abundance from finite counts or the explicit use of relative abundance as part of the sample representation itself. The common structure is not a single formalism but a shared concern with how abundance information survives sampling, quantization, and aggregation [2308.00695] [1701.07850] [2406.10921] [2508.11075].

## 1. One-bit sensing and the sample-abundance regime

In one-bit sensing, sample abundance is the regime in which one-bit ADCs collect **many more measurements than the ambient dimension or the number of unknown degrees of freedom**. The one-bit measurement model is
\[
r_k=\operatorname{sgn}(y_k-\tau_k),
\]
with time-varying thresholds \(\tau_k\). Each sign observation yields the inequality
\[
r_k(y_k-\tau_k)\ge 0,
\]
and, for linear observations \(\mathbf y=A\mathbf x\), the stacked constraints become a large linear system such as
\[
P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).
\]
The feasible set is a **one-bit polyhedron**: the intersection of many half-spaces induced by thresholded sign measurements. In the sample-abundant regime, the polyhedron shrinks to a small cell around the target, so that expensive constraints such as PSD, rank, or sparsity can become effectively redundant in practice [2308.00695] [2507.19415].

This reduction is especially explicit in one-bit quadratic compressed sensing. With quadratic measurements
\[
y_j=x^{\mathrm H}A_jx,
\]
lifting gives \(X=xx^{\mathrm H}\) and
\[
y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).
\]
After one-bit thresholding,
\[
r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},
\]
so recovery becomes the search for a point in a large polyhedron rather than the solution of a semidefinite or other nonconvex program. The associated algorithmic emphasis shifts to RKA, SKM, PrSKM, and Block SKM. The theoretical backbone is the **Finite Volume Property**: for isotropic sensing, forcing the feasible polyhedron into an \(\epsilon\)-ball requires \(m=O(\epsilon^{-3})\), while structured sparse or low-rank classes admit \(m=O(\epsilon^{-2})\). The later overview literature terms the abrupt computational transition above this threshold the **sample abundance singularity**. Numerically, the quadratic-compressed-sensing study reports that Block SKM reached NMSE \(3.1072\times 10^{-7}\) in \(0.0026\) s, versus GESPAR’s \(2.4382\times 10^{-5}\) in \(0.0041\) s in the reported setting [2303.09594] [2507.19415].

## 2. Stellar abundance determinations in astronomical samples

In stellar spectroscopy, sample abundance usually means element-by-element abundances derived for a defined stellar sample under a homogeneous analysis. A representative example is the high-resolution study of **38 solar analogues**, including **11** previously flagged SMR candidates. It measured equivalent widths for **34 lines** of **Mg, Al, Si, Ca, Ti, Fe, and Ni** at \(\mathcal R\sim 80{,}000\), modeled each star with an individual ATLAS12 atmosphere, and computed abundances with MOOG. The paper uses
\[
[{\rm X/H}] = A({\rm X})_{\rm star} - A({\rm X})_{\odot},
\qquad
[{\rm X/Fe}] = [{\rm X/H}] - [{\rm Fe/H}],
\]
and also evaluates the refractory index [Ref] from Mg, Si, and Fe. It confirms the super-metal-rich status of **6** stars; **BD+60 600** is the most iron-rich object with \([{\rm Fe/H}] = +0.35\) dex and \([{\rm Ref}] = +0.42\), while **HD 166991** has \([{\rm Fe/H}] = -0.53\) dex. For the 25 stars with [Ref], **BD+60 600** and **BD+28 3198** are explicitly highlighted as especially promising targets for giant-planet searches [1701.07850].

Larger Galactic-abundance samples generalize this usage from detailed case studies to population stratification. The HARPS-GTO analysis derives abundances for the heavy elements **Cu, Zn, Sr, Y, Zr, Ba, Ce, Nd, and Eu** for a **homogeneous sample of 1059 stars**, and restricts age analyses to **377** stars whose Hipparcos- and Gaia-DR1-based ages differ by **less than 1 Gyr** and have uncertainty **< 2 Gyr**. It finds that **thick disk stars are chemically distinct for Zn and Eu**, that the **high-\(\alpha\) metal-rich** population is overabundant in **Cu and Zn** and underabundant in **Y and Ba** relative to thin-disk stars, and that several abundance ratios correlate significantly with age only for chemically separated thin-disk stars. At supersolar metallicity, the abundance–age trends weaken for several elements [1707.05156].

Very metal-poor samples show a different aspect of sample abundance: the distribution of abundance ratios within chemically primitive populations. The TOPoS study analyzes **65 metal-poor turn-off stars** spanning approximately \([{\rm Fe/H}] \approx -2.38\) to \(-4.11\). It confirms super-solar \([{\rm Mg/Fe}]\), \([{\rm Si/Fe}]\), and \([{\rm Ca/Fe}]\) on average, but also a significant spread, including several stars with sub-solar \([{\rm Ca/Fe}]\). Strontium was measured in **12** stars and is typically around the solar value; barium was detected in only **2** stars, with **SDSS J114424−004658** showing \([{\rm Sr/Fe}] = 1.28\) and \([{\rm Ba/Fe}] = 1.03\) [1811.00035].

Other stellar studies extend sample-abundance methodology to abundances that are inaccessible to standard optical spectroscopy. Asteroseismic helium measurements for **38** stars in the Kepler LEGACY sample infer envelope helium from the helium-ionization acoustic glitch and, after accounting for settling, derive preliminary estimates \(Y_p = 0.244\pm0.019\) and \(\Delta Y/\Delta Z = 1.226\pm0.849\). At the opposite end of the periodic table, HST/STIS far-UV spectroscopy of **six metal-poor warm stars** uses the practically unblended **B I 2089.6 Å** line to obtain boron abundances with unprecedented precision, finding
\[
{\rm A(B)} = 1.644\,[{\rm Fe/H}] + 3.781,
\]
a slope significantly larger than 1 over \(-2.6 < [{\rm Fe/H}] < -1.0\), and an inferred break in boron enrichment near \([{\rm Fe/H}] \approx -1\) [1812.02751] [2510.11594].

A survey-oriented counterpart is the APOKASC re-analysis, which uses fixed \(T_{\rm eff}\), asteroseismic \(\log g\), BACCHUS, Turbospectrum, and MARCS atmospheres to derive metallicity, microturbulence, broadening, and abundances for up to **21 elements** in the APOGEE/Kepler red-giant sample. The catalogue is explicitly line-by-line and differential with respect to Arcturus, and the paper attributes its improved metal-poor performance to line selection, differential analysis, and externally constrained stellar parameters [1604.08800].

## 3. Nebular, galactic, and AGN abundance patterns

For emission-line systems, sample abundance refers to nebular abundance ratios inferred from direct \(T_e\)-based methods, strong-line calibrations, or photoionization modeling. In Seyfert 2 nuclei, one study uses Cloudy models for a literature-compiled sample and obtains successful solutions for **44** objects. It reports nitrogen abundances from about **0.3** to **7.5** times solar and derives
\[
\log(N/H) = (1.05 \pm0.09)\,[\log(O/H)] -(0.35 \pm 0.33),
\]
interpreting the trend as secondary-nitrogen behavior for \(12+\log(O/H) > 8.0\). A separate Seyfert 2 study derives **Ar/H** for **64** local nuclei with an AGN-specific \(T_e\)-based method and Cloudy-calibrated ICFs, finding argon abundances from about **0.1** to **3** times solar, a mean \(12+\log(Ar/H)=6.14\pm0.32\), and about **75%** of the sample below solar Ar abundance. When combined with H II galaxies, the fitted trend is
\[
\log(Ar/O) = (-0.11\pm0.02)\,x - (1.51\pm0.15),
\qquad x=\log(O/H),
\]
so Ar/O decreases slightly with increasing metallicity [1703.03250] [2109.10590].

Low-metallicity star-forming galaxies provide a broader baseline for abundance-pattern work. A VLT compilation analyzes **121 spectra** of H II regions in **46** galaxies over \(12+\log(O/H)\simeq 7.2\)–8.4 using the classical direct method. It presents the first empirical low-metallicity relations linking \(t_e({\rm SIII})\) and \(t_e({\rm NII})\) to \(t_e({\rm OIII})\), finds that **Ne/O increases** mildly with oxygen abundance, and that **Fe/O** drops from roughly solar at the lowest metallicities to about one-tenth solar by \(12+\log({\rm O/H})\sim 8.5\), which is interpreted as increasing depletion of iron into dust. It also reports a possible N/O upturn below \(12+\log(O/H)<7.5\), confirms the RL–CEL abundance discrepancy through the ADF, and finds that **C/O increases with O/H** [1111.1392].

At the scale of spiral disks, the same abundance problem becomes radial rather than purely demographic. The large H II-region compilation of **2831** measurements in **51** nearby spirals derives direct abundances for **610** regions and compares them with multiple strong-line methods. A central methodological result is that **O3N2**, especially the **Marino et al. (2013)** calibration, compresses direct-method oxygen abundances into a narrower interval, roughly **8.2–8.7 dex** rather than **7.7–8.9 dex**, and thereby tends to flatten the steepest radial metallicity profiles. This is a substantive caution against reading apparent gradient universality as purely astrophysical rather than partly calibration-driven [2007.12289].

## 4. Counting design and absolute abundance in paleobiological samples

In species-rich assemblages, sample abundance is a statistical design problem before it is an interpretive one. For microfossil counts, the minimum sample size \(n\) required to detect taxa at a specified confidence must be chosen **a priori** from the desired confidence level, the assumed population proportion \(p\), and the number of taxa \(K\) to be detected simultaneously. For a single taxon,
\[
P(Y_1=0)=(1-p)^n<\alpha,
\qquad
n=\left\lfloor \frac{\ln(\alpha)}{\ln(1-p)} \right\rfloor+1.
\]
For concurrent detection of several taxa, the correct framework is multinomial rather than binomial. For equal \(p\), the conservative approximation
\[
\left(1-e^{-np}\right)^K \ge 1-\alpha
\]
gives
\[
n=\left\lceil \frac{-\ln\!\left(1-(1-\alpha)^{1/K}\right)}{p} \right\rceil.
\]
The resulting sample-size inflation is large: **300 specimens is not enough to detect more than one taxon at 95% confidence**, and **500** is not enough at **99%**. The paper gives, among other examples, \(n=366\) for \(K=2\), \(p=1\%\), 95% confidence; \(n=527\) for the same \(K\) and \(p\) at 99%; \(n=599\) for \(K=4\), \(p=1\%\), 99%; and \(n=477\) for \(K=6\), \(p=1\%\), 95% [1804.11226].

Absolute abundance estimation from spatial count data uses a different but related logic: infer concentration from a count subsample plus a known reference. The refined exotic-marker framework writes concentration as
\[
c = \frac{x\, N_1\, Y_1}{n\, V},
\]
where \(x\) is the counted target abundance, \(n\) is the counted marker abundance, \(N_1\) is the number of added marker doses, \(Y_1\) is the mean number of markers per dose, and \(V\) is sample mass, area, or volume. The paper’s main methodological contribution is **field-of-view subsampling (FOVS)**. Calibration counts estimate the mean targets per field of view, \(Y_{3x}\), after which the target total in the full-count region is extrapolated as
\[
x = Y_{3x} \times N_{3F},
\]
yielding
\[
C_{Fx} = \frac{x\, N_1\, Y_1}{n\, V}.
\]
Simulations and Permian–Triassic terrestrial organic microfossil case studies show that, in almost all cases, FOVS delivers higher precision than the linear method at equivalent effort, and the authors estimate that achieving **10%** error with the linear method would require about **seven times** the effort needed by FOVS in the empirical setting [2406.10921].

## 5. Species richness and abundance inference from sparse ecological data

Ecological abundance estimation often operates under data scarcity rather than sample abundance in the signal-processing sense. One Bayesian response is the **triple Poisson** hierarchy for trace counts:
\[
Y|T,G,\nu \sim \mbox{Poisson}(\alpha T \nu), \qquad
T|G \sim \mbox{Poisson}(G\lambda_N), \qquad
G \sim \mbox{Poisson}(\lambda_G).
\]
Here \(T\) is total abundance, \(G\) is the number of groups, \(\lambda_N\) is mean group size, \(\nu\) is coverage, and \(\alpha\) absorbs vestige production and decay. The framework is implemented in JAGS and can be extended with a Negative Binomial observation model. Simulation results show accurate performance even when data are very scarce, with relative mean bias values around \(0.1\)–\(0.2\) for the triple-Poisson variants in the main scarce-data scenario. In case studies, the model estimated **44 collared peccaries** with 95% credible interval **(16, 87)** from only two transects with counts \((7,1)\), and **3304 red foxes** with 95% credible interval **(2603, 4442)** from the Italian trace-count dataset [2206.05944].

A distinct statistical tradition approaches abundance through occupancy laws rather than latent group processes. In Gibbs–Poisson abundance models, iid species abundances \(X_1,\dots,X_n\) are conditioned on the total sample size \(k\), producing the exchangeable occupancy vector \(\mathbf K_{n,k}\). The number of observed species,
\[
P_{n,k}:=\sum_{m=1}^n \mathbf 1_{\{K_{n,k}(m)>0\}},
\]
is the canonical summary for unseen-species inference. The framework derives explicit occupancy and frequency-of-frequencies laws through Bell polynomials \(B_{k,p}(\theta_\bullet)\), and it extends naturally to the infinite-species limit
\[
n\to\infty,\qquad x\to 0,\qquad nx\to\gamma>0,
\]
where \(\gamma\) becomes a richness or diversity parameter. In that limit,
\[
\mathbb P^*(P_k=p)=\frac{\gamma^p}{\theta_k(\gamma)}\,B_{k,p}(\theta_\bullet),
\]
and the MLE \(\widehat\gamma\) solves
\[
P=\widehat\gamma\,\frac{\theta_k'(\widehat\gamma)}{\theta_k(\widehat\gamma)}.
\]
The paper further shows that a large class of these models admits a continuum interpretation as sampling from a random partition of unity, typically with bias by the total length in the normalization [1307.3000].

## 6. Abundance-aware sample representations in microbiomics

In microbiome representation learning, sample abundance is neither a chemical abundance ratio nor a latent ecological parameter; it is the explicit use of taxa relative abundance in the construction of a sample embedding. The abundance-aware Set Transformer framework starts from DNABERT-2 sequence embeddings and argues that simple averaging or majority-vote aggregation obscures taxa abundance. For a sample \(S\) with sequence embeddings \(\mathbf e_i\) and abundances \(a_i\), the weighted average embedding is
\[
\mathbf{z}_S = \sum_{i=1}^N \alpha_i \mathbf{e}_i,
\qquad
\alpha_i = \frac{a_i}{\sum_{j=1}^N a_j}.
\]
The abundance-aware Set Transformer applies either weighted output pooling,
\[
\mathbf{z}_S = \sum_{i=1}^{N} \alpha_i \mathbf{o}_i,
\]
or a repetition-based strategy in which sequence embeddings are replicated in proportion to abundance before attention. The architecture remains permutation-invariant through ISAB, PMA, and SAB, and the final sample representation is a fixed **768-dimensional** embedding [2508.11075].

The empirical setting comprises three Qiita-based classification tasks. On the bladder-microbiota task, **Weighted Set Transformer + RF** achieved **0.5833** accuracy and **0.5804** macro F1, exceeding average pooling and the unweighted Set Transformer. On the Acanthamoeba–Leptospira co-occurrence task, the abundance-aware model reached **1.0000** accuracy and **1.0000** macro F1 with both FCNN and RF, whereas non-abundance-aware baselines were lower. Under cross-study soil versus non-soil prediction, **Weighted Set Transformer + FCNN** achieved **0.5882** accuracy and **0.3704** macro F1, versus **0.4118** and **0.2917** for the baseline methods. The stated interpretation is that abundance-aware aggregation preserves quantitative sample structure and can recover signals carried by rare or dominant taxa that uniform pooling suppresses [2508.11075].

In aggregate, the literature does not support a single universal definition of sample abundance. It instead supports a family of domain-specific meanings organized around the same technical tension: abundance must be inferred, preserved, or exploited under finite observation. In one-bit sensing, abundant samples reduce optimization complexity; in astronomy, homogeneous samples stabilize abundance scales and reveal population structure; in paleobiology and ecology, count design and hierarchical modeling determine whether abundance is identifiable at all; and in microbiomics, abundance-aware aggregation changes the geometry of the sample representation itself.

Source: https://www.emergentmind.com/topics/sample-abundance