Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sample Abundance in Diverse Scientific Domains

Updated 7 July 2026
  • Sample Abundance is a multifaceted concept defining how quantitative information is preserved and inferred from coarse or limited measurements across various scientific domains.
  • It enables the transformation of numerous low-precision or sparse counts into overdetermined problems in one-bit sensing, detailed chemical analyses in astronomy, and statistical inferences in ecology and paleobiology.
  • Innovative methods including polyhedral constraint reduction, Bayesian hierarchical modeling, and abundance-aware aggregation significantly improve accuracy and efficiency in abundance estimation.

Across the cited literature, the expression sample abundance appears in several technically distinct senses. In one-bit and few-bit signal processing, it denotes a regime in which very large numbers of low-precision measurements compensate for coarse quantization and convert recovery into an overdetermined linear-feasibility problem. In astronomy, it denotes chemical abundances derived for a defined sample of stars, H II regions, AGN narrow-line regions, or open clusters. In ecology, paleobiology, and microbiomics, it denotes either the statistical recovery of abundance from finite counts or the explicit use of relative abundance as part of the sample representation itself. The common structure is not a single formalism but a shared concern with how abundance information survives sampling, quantization, and aggregation (Eamaz et al., 2023, López-Valdivia et al., 2017, Mays et al., 2024, Yoo et al., 14 Aug 2025).

1. One-bit sensing and the sample-abundance regime

In one-bit sensing, sample abundance is the regime in which one-bit ADCs collect many more measurements than the ambient dimension or the number of unknown degrees of freedom. The one-bit measurement model is

rk=sgn(ykτk),r_k=\operatorname{sgn}(y_k-\tau_k),

with time-varying thresholds τk\tau_k. Each sign observation yields the inequality

rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,

and, for linear observations y=Ax\mathbf y=A\mathbf x, the stacked constraints become a large linear system such as

Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).

The feasible set is a one-bit polyhedron: the intersection of many half-spaces induced by thresholded sign measurements. In the sample-abundant regime, the polyhedron shrinks to a small cell around the target, so that expensive constraints such as PSD, rank, or sparsity can become effectively redundant in practice (Eamaz et al., 2023, Eamaz et al., 25 Jul 2025).

This reduction is especially explicit in one-bit quadratic compressed sensing. With quadratic measurements

yj=xHAjx,y_j=x^{\mathrm H}A_jx,

lifting gives X=xxHX=xx^{\mathrm H} and

yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).

After one-bit thresholding,

rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},

so recovery becomes the search for a point in a large polyhedron rather than the solution of a semidefinite or other nonconvex program. The associated algorithmic emphasis shifts to RKA, SKM, PrSKM, and Block SKM. The theoretical backbone is the Finite Volume Property: for isotropic sensing, forcing the feasible polyhedron into an ϵ\epsilon-ball requires τk\tau_k0, while structured sparse or low-rank classes admit τk\tau_k1. The later overview literature terms the abrupt computational transition above this threshold the sample abundance singularity. Numerically, the quadratic-compressed-sensing study reports that Block SKM reached NMSE τk\tau_k2 in τk\tau_k3 s, versus GESPAR’s τk\tau_k4 in τk\tau_k5 s in the reported setting (Eamaz et al., 2023, Eamaz et al., 25 Jul 2025).

2. Stellar abundance determinations in astronomical samples

In stellar spectroscopy, sample abundance usually means element-by-element abundances derived for a defined stellar sample under a homogeneous analysis. A representative example is the high-resolution study of 38 solar analogues, including 11 previously flagged SMR candidates. It measured equivalent widths for 34 lines of Mg, Al, Si, Ca, Ti, Fe, and Ni at τk\tau_k6, modeled each star with an individual ATLAS12 atmosphere, and computed abundances with MOOG. The paper uses

τk\tau_k7

and also evaluates the refractory index [Ref] from Mg, Si, and Fe. It confirms the super-metal-rich status of 6 stars; BD+60 600 is the most iron-rich object with τk\tau_k8 dex and τk\tau_k9, while HD 166991 has rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,0 dex. For the 25 stars with [Ref], BD+60 600 and BD+28 3198 are explicitly highlighted as especially promising targets for giant-planet searches (López-Valdivia et al., 2017).

Larger Galactic-abundance samples generalize this usage from detailed case studies to population stratification. The HARPS-GTO analysis derives abundances for the heavy elements Cu, Zn, Sr, Y, Zr, Ba, Ce, Nd, and Eu for a homogeneous sample of 1059 stars, and restricts age analyses to 377 stars whose Hipparcos- and Gaia-DR1-based ages differ by less than 1 Gyr and have uncertainty < 2 Gyr. It finds that thick disk stars are chemically distinct for Zn and Eu, that the high-rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,1 metal-rich population is overabundant in Cu and Zn and underabundant in Y and Ba relative to thin-disk stars, and that several abundance ratios correlate significantly with age only for chemically separated thin-disk stars. At supersolar metallicity, the abundance–age trends weaken for several elements (Mena et al., 2017).

Very metal-poor samples show a different aspect of sample abundance: the distribution of abundance ratios within chemically primitive populations. The TOPoS study analyzes 65 metal-poor turn-off stars spanning approximately rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,2 to rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,3. It confirms super-solar rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,4, rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,5, and rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,6 on average, but also a significant spread, including several stars with sub-solar rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,7. Strontium was measured in 12 stars and is typically around the solar value; barium was detected in only 2 stars, with SDSS J114424−004658 showing rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,8 and rk(ykτk)0,r_k(y_k-\tau_k)\ge 0,9 (François et al., 2018).

Other stellar studies extend sample-abundance methodology to abundances that are inaccessible to standard optical spectroscopy. Asteroseismic helium measurements for 38 stars in the Kepler LEGACY sample infer envelope helium from the helium-ionization acoustic glitch and, after accounting for settling, derive preliminary estimates y=Ax\mathbf y=A\mathbf x0 and y=Ax\mathbf y=A\mathbf x1. At the opposite end of the periodic table, HST/STIS far-UV spectroscopy of six metal-poor warm stars uses the practically unblended B I 2089.6 Å line to obtain boron abundances with unprecedented precision, finding

y=Ax\mathbf y=A\mathbf x2

a slope significantly larger than 1 over y=Ax\mathbf y=A\mathbf x3, and an inferred break in boron enrichment near y=Ax\mathbf y=A\mathbf x4 (Verma et al., 2018, Spite et al., 13 Oct 2025).

A survey-oriented counterpart is the APOKASC re-analysis, which uses fixed y=Ax\mathbf y=A\mathbf x5, asteroseismic y=Ax\mathbf y=A\mathbf x6, BACCHUS, Turbospectrum, and MARCS atmospheres to derive metallicity, microturbulence, broadening, and abundances for up to 21 elements in the APOGEE/Kepler red-giant sample. The catalogue is explicitly line-by-line and differential with respect to Arcturus, and the paper attributes its improved metal-poor performance to line selection, differential analysis, and externally constrained stellar parameters (Hawkins et al., 2016).

3. Nebular, galactic, and AGN abundance patterns

For emission-line systems, sample abundance refers to nebular abundance ratios inferred from direct y=Ax\mathbf y=A\mathbf x7-based methods, strong-line calibrations, or photoionization modeling. In Seyfert 2 nuclei, one study uses Cloudy models for a literature-compiled sample and obtains successful solutions for 44 objects. It reports nitrogen abundances from about 0.3 to 7.5 times solar and derives

y=Ax\mathbf y=A\mathbf x8

interpreting the trend as secondary-nitrogen behavior for y=Ax\mathbf y=A\mathbf x9. A separate Seyfert 2 study derives Ar/H for 64 local nuclei with an AGN-specific Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).0-based method and Cloudy-calibrated ICFs, finding argon abundances from about 0.1 to 3 times solar, a mean Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).1, and about 75% of the sample below solar Ar abundance. When combined with H II galaxies, the fitted trend is

Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).2

so Ar/O decreases slightly with increasing metallicity (Dors et al., 2017, Monteiro et al., 2021).

Low-metallicity star-forming galaxies provide a broader baseline for abundance-pattern work. A VLT compilation analyzes 121 spectra of H II regions in 46 galaxies over Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).3–8.4 using the classical direct method. It presents the first empirical low-metallicity relations linking Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).4 and Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).5 to Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).6, finds that Ne/O increases mildly with oxygen abundance, and that Fe/O drops from roughly solar at the lowest metallicities to about one-tenth solar by Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).7, which is interpreted as increasing depletion of iron into dust. It also reports a possible N/O upturn below Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).8, confirms the RL–CEL abundance discrepancy through the ADF, and finds that C/O increases with O/H (Guseva et al., 2011).

At the scale of spiral disks, the same abundance problem becomes radial rather than purely demographic. The large H II-region compilation of 2831 measurements in 51 nearby spirals derives direct abundances for 610 regions and compares them with multiple strong-line methods. A central methodological result is that O3N2, especially the Marino et al. (2013) calibration, compresses direct-method oxygen abundances into a narrower interval, roughly 8.2–8.7 dex rather than 7.7–8.9 dex, and thereby tends to flatten the steepest radial metallicity profiles. This is a substantive caution against reading apparent gradient universality as purely astrophysical rather than partly calibration-driven (Zurita et al., 2020).

4. Counting design and absolute abundance in paleobiological samples

In species-rich assemblages, sample abundance is a statistical design problem before it is an interpretive one. For microfossil counts, the minimum sample size Pyxvec(Ry)vec(Γ).P_y\mathbf x \succeq \operatorname{vec}(R_y)\odot \operatorname{vec}(\Gamma).9 required to detect taxa at a specified confidence must be chosen a priori from the desired confidence level, the assumed population proportion yj=xHAjx,y_j=x^{\mathrm H}A_jx,0, and the number of taxa yj=xHAjx,y_j=x^{\mathrm H}A_jx,1 to be detected simultaneously. For a single taxon,

yj=xHAjx,y_j=x^{\mathrm H}A_jx,2

For concurrent detection of several taxa, the correct framework is multinomial rather than binomial. For equal yj=xHAjx,y_j=x^{\mathrm H}A_jx,3, the conservative approximation

yj=xHAjx,y_j=x^{\mathrm H}A_jx,4

gives

yj=xHAjx,y_j=x^{\mathrm H}A_jx,5

The resulting sample-size inflation is large: 300 specimens is not enough to detect more than one taxon at 95% confidence, and 500 is not enough at 99%. The paper gives, among other examples, yj=xHAjx,y_j=x^{\mathrm H}A_jx,6 for yj=xHAjx,y_j=x^{\mathrm H}A_jx,7, yj=xHAjx,y_j=x^{\mathrm H}A_jx,8, 95% confidence; yj=xHAjx,y_j=x^{\mathrm H}A_jx,9 for the same X=xxHX=xx^{\mathrm H}0 and X=xxHX=xx^{\mathrm H}1 at 99%; X=xxHX=xx^{\mathrm H}2 for X=xxHX=xx^{\mathrm H}3, X=xxHX=xx^{\mathrm H}4, 99%; and X=xxHX=xx^{\mathrm H}5 for X=xxHX=xx^{\mathrm H}6, X=xxHX=xx^{\mathrm H}7, 95% (Haidar et al., 2018).

Absolute abundance estimation from spatial count data uses a different but related logic: infer concentration from a count subsample plus a known reference. The refined exotic-marker framework writes concentration as

X=xxHX=xx^{\mathrm H}8

where X=xxHX=xx^{\mathrm H}9 is the counted target abundance, yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).0 is the counted marker abundance, yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).1 is the number of added marker doses, yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).2 is the mean number of markers per dose, and yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).3 is sample mass, area, or volume. The paper’s main methodological contribution is field-of-view subsampling (FOVS). Calibration counts estimate the mean targets per field of view, yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).4, after which the target total in the full-count region is extrapolated as

yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).5

yielding

yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).6

Simulations and Permian–Triassic terrestrial organic microfossil case studies show that, in almost all cases, FOVS delivers higher precision than the linear method at equivalent effort, and the authors estimate that achieving 10% error with the linear method would require about seven times the effort needed by FOVS in the empirical setting (Mays et al., 2024).

5. Species richness and abundance inference from sparse ecological data

Ecological abundance estimation often operates under data scarcity rather than sample abundance in the signal-processing sense. One Bayesian response is the triple Poisson hierarchy for trace counts: yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).7 Here yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).8 is total abundance, yj=Tr(AjX)=vec(Aj)vec(X).y_j=\operatorname{Tr}(A_jX)=\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X).9 is the number of groups, rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},0 is mean group size, rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},1 is coverage, and rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},2 absorbs vestige production and decay. The framework is implemented in JAGS and can be extended with a Negative Binomial observation model. Simulation results show accurate performance even when data are very scarce, with relative mean bias values around rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},3–rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},4 for the triple-Poisson variants in the main scarce-data scenario. In case studies, the model estimated 44 collared peccaries with 95% credible interval (16, 87) from only two transects with counts rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},5, and 3304 red foxes with 95% credible interval (2603, 4442) from the Italian trace-count dataset (Mimnagh et al., 2022).

A distinct statistical tradition approaches abundance through occupancy laws rather than latent group processes. In Gibbs–Poisson abundance models, iid species abundances rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},6 are conditioned on the total sample size rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},7, producing the exchangeable occupancy vector rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},8. The number of observed species,

rj()vec(Aj)vec(X)rj()τj(),r_j^{(\ell)}\,\operatorname{vec}(A_j^\top)^\top \operatorname{vec}(X)\ge r_j^{(\ell)}\tau_j^{(\ell)},9

is the canonical summary for unseen-species inference. The framework derives explicit occupancy and frequency-of-frequencies laws through Bell polynomials ϵ\epsilon0, and it extends naturally to the infinite-species limit

ϵ\epsilon1

where ϵ\epsilon2 becomes a richness or diversity parameter. In that limit,

ϵ\epsilon3

and the MLE ϵ\epsilon4 solves

ϵ\epsilon5

The paper further shows that a large class of these models admits a continuum interpretation as sampling from a random partition of unity, typically with bias by the total length in the normalization (Huillet et al., 2013).

6. Abundance-aware sample representations in microbiomics

In microbiome representation learning, sample abundance is neither a chemical abundance ratio nor a latent ecological parameter; it is the explicit use of taxa relative abundance in the construction of a sample embedding. The abundance-aware Set Transformer framework starts from DNABERT-2 sequence embeddings and argues that simple averaging or majority-vote aggregation obscures taxa abundance. For a sample ϵ\epsilon6 with sequence embeddings ϵ\epsilon7 and abundances ϵ\epsilon8, the weighted average embedding is

ϵ\epsilon9

The abundance-aware Set Transformer applies either weighted output pooling,

τk\tau_k00

or a repetition-based strategy in which sequence embeddings are replicated in proportion to abundance before attention. The architecture remains permutation-invariant through ISAB, PMA, and SAB, and the final sample representation is a fixed 768-dimensional embedding (Yoo et al., 14 Aug 2025).

The empirical setting comprises three Qiita-based classification tasks. On the bladder-microbiota task, Weighted Set Transformer + RF achieved 0.5833 accuracy and 0.5804 macro F1, exceeding average pooling and the unweighted Set Transformer. On the Acanthamoeba–Leptospira co-occurrence task, the abundance-aware model reached 1.0000 accuracy and 1.0000 macro F1 with both FCNN and RF, whereas non-abundance-aware baselines were lower. Under cross-study soil versus non-soil prediction, Weighted Set Transformer + FCNN achieved 0.5882 accuracy and 0.3704 macro F1, versus 0.4118 and 0.2917 for the baseline methods. The stated interpretation is that abundance-aware aggregation preserves quantitative sample structure and can recover signals carried by rare or dominant taxa that uniform pooling suppresses (Yoo et al., 14 Aug 2025).

In aggregate, the literature does not support a single universal definition of sample abundance. It instead supports a family of domain-specific meanings organized around the same technical tension: abundance must be inferred, preserved, or exploited under finite observation. In one-bit sensing, abundant samples reduce optimization complexity; in astronomy, homogeneous samples stabilize abundance scales and reveal population structure; in paleobiology and ecology, count design and hierarchical modeling determine whether abundance is identifiable at all; and in microbiomics, abundance-aware aggregation changes the geometry of the sample representation itself.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sample Abundance.