Papers
Topics
Authors
Recent
Search
2000 character limit reached

Selection Functions: Concepts and Applications

Updated 12 July 2026
  • Selection Functions are formal devices that specify how elements from a possibility space are selected, bridging latent models and observed outcomes.
  • In astronomy and cluster cosmology, selection functions quantify inclusion probabilities and threshold criteria to model survey completeness and calibration accurately.
  • Beyond observational sciences, selection functions guide decision-making in axiomatic choice theory, automated reasoning, and machine learning by structuring choices and optimizing outcomes.

Searching arXiv for relevant papers on “selection functions” across major usages (astronomy, logic/game theory, optimization/ML). Searching arXiv for “selection functions astronomy survey completeness selection functions”. Selection functions are formal devices that specify how elements of a space become selected, retained, or acted upon. Across disciplines, the term denotes different mathematical objects: a survey-inclusion probability S(q)S(\vec q), a cluster-detection threshold S(M,z)S(M,z), a choice map C:QQC:Q\to Q_{\emptyset}, a higher-order functional JRX=(XR)XJ_R X=(X\to R)\to X, a nondeterministic selector JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X), or a learned internal score sθ(x)s_\theta(x) used for replacement and survival. What unifies these constructions is not a single formalism but a common role: each mediates between an underlying space of possibilities and an observed subset, selected action, or continuation (Rix et al., 2021, Ascaso et al., 2016, Bock, 2020, Escardo et al., 2014, Bolt et al., 2018, Frans et al., 2021).

1. Conceptual scope and formal heterogeneity

The term has no field-independent canonical definition. In astronomical catalog modeling, a selection function is the probability that an object with observables xx or q\vec q is included in a catalog or sample, typically written S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x) or S(q)S(\vec q) (Rix et al., 2021, Everall et al., 2021). In cluster cosmology, the operational definition becomes a threshold object: the minimum halo mass as a function of redshift for which completeness and purity exceed prescribed levels, formalized as

S(M,z)S(M,z)0

with S(M,z)S(M,z)1 in the baseline construction (Ascaso et al., 2016). In axiomatic choice theory, by contrast, a selection or choice function is a map S(M,z)S(M,z)2 sending each non-empty feasible set S(M,z)S(M,z)3 to a chosen subset S(M,z)S(M,z)4, possibly empty (Bock, 2020). In higher-order semantics, a selection function is a functional of type S(M,z)S(M,z)5, and the corresponding selection monad supports binary products and iterated products with direct proof-theoretic and game-theoretic interpretations (Escardo et al., 2014). In evolutionary search, a selection function maps individual-level features such as fitness, novelty, age, and rank to deterministic tournament outcomes or selection probabilities, thereby controlling exploration and exploitation (Frans et al., 2021). In DatalogS(M,z)S(M,z)6, a selection function maps a program S(M,z)S(M,z)7 to a subset of chase positions S(M,z)S(M,z)8, marking where finitely many values appear (Bertossi et al., 2021).

Domain Formal object Primary role
Astronomical catalogs S(M,z)S(M,z)9, C:QQC:Q\to Q_{\emptyset}0 Inclusion probability
Cluster surveys C:QQC:Q\to Q_{\emptyset}1 Completeness–purity threshold
Choice theory C:QQC:Q\to Q_{\emptyset}2 Chosen subset from feasible set
Higher-order computation C:QQC:Q\to Q_{\emptyset}3 Optimization-aware choice
Nondeterministic games C:QQC:Q\to Q_{\emptyset}4 Set-valued admissible moves
Evolutionary search C:QQC:Q\to Q_{\emptyset}5 or tournament rule Survival and replacement
DatalogC:QQC:Q\to Q_{\emptyset}6 C:QQC:Q\to Q_{\emptyset}7 Finite-position marking

A recurrent misconception is that a selection function must always be probabilistic. The literature shows otherwise: some are probabilities, some are threshold indicators, some are set-valued maps, and some are higher-type continuations.

2. Inclusion probabilities, completeness, and calibration in astronomy

In astronomical survey modeling, selection functions are the link between a population model and catalog incidence. For a catalog or curated sample, C:QQC:Q\to Q_{\emptyset}8 is the probability that an object with attributes C:QQC:Q\to Q_{\emptyset}9 appears in the sample, and it enters forward modeling multiplicatively, with nuisance-variable dependence removed by integration. The standard effective-volume construction is

JRX=(XR)XJ_R X=(X\to R)\to X0

which yields JRX=(XR)XJ_R X=(X\to R)\to X1. In the Gaia white dwarf example, the paper argues that strict volume-limited sampling is often sub-optimal: for JRX=(XR)XJ_R X=(X\to R)\to X2 mag, the accessible volume ratio in a flux-limited survey is JRX=(XR)XJ_R X=(X\to R)\to X3, and after explicit selection modeling the inferred white dwarf luminosity–color function peak shifts by JRX=(XR)XJ_R X=(X\to R)\to X4 magnitudes relative to raw counts (Rix et al., 2021).

Cluster cosmology uses a more operational notion. In stage-IV optical and infrared surveys, completeness is defined as JRX=(XR)XJ_R X=(X\to R)\to X5, purity as JRX=(XR)XJ_R X=(X\to R)\to X6, and the selection function is the minimum halo mass satisfying threshold cuts in both quantities. Using the Bayesian Cluster Finder on realistic mock catalogs, Euclid-Optimistic reaches JRX=(XR)XJ_R X=(X\to R)\to X7 completeness and purity down to JRX=(XR)XJ_R X=(X\to R)\to X8 for JRX=(XR)XJ_R X=(X\to R)\to X9, and JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)0 up to JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)1. Euclid-Pessimistic has a similar shape with JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)2 higher mass limits, while LSST has JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)3 higher mass limits than Euclid-Optimistic up to JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)4 and reaches JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)5 by JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)6. The same work fits

JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)7

with JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)8, finding intrinsic scatters JRPX=(XR)P(X)J_R^{\mathcal P}X=(X\to R)\to \mathcal P(X)9, sθ(x)s_\theta(x)0, and sθ(x)s_\theta(x)1 dex for Euclid-Optimistic, Euclid-Pessimistic, and LSST, respectively, and negligible redshift evolution within uncertainties (Ascaso et al., 2016).

For Milky Way spectroscopic surveys, empirical selection functions are commonly estimated relative to a photometric parent catalog. In observed space, sθ(x)s_\theta(x)2, with sθ(x)s_\theta(x)3 typically containing sky position, magnitude, and color. The seestar framework models both the parent photometric density and the field-level spectroscopic selection function with Gaussian mixture models, uses a Poisson point-process likelihood, and combines overlapping fields via

sθ(x)s_\theta(x)4

It also defines an intrinsic selection function sθ(x)s_\theta(x)5, obtained by composing the observable-space selector with geometry, stellar models, and extinction maps (Everall et al., 2019).

Gaia EDR3 subset selection functions sharpen the same principle. The paper on Gaia EDR3 models the probability of having astrometry, sθ(x)s_\theta(x)6, or a reported radial velocity as conditional probabilities relative to inclusion in the EDR3 source catalog, with dependencies on sθ(x)s_\theta(x)7, sky position, and, for RUWE and RVS, color sθ(x)s_\theta(x)8. It uses HEALPix sθ(x)s_\theta(x)9, spherical needlets up to xx0, and binomial likelihoods per HEALPix–magnitude–color bin. The inferred functions show strong scanning-law and crowding dependence, and the paper emphasizes that ignoring this structure risks biased density and kinematic inference (Everall et al., 2021).

Two further lines of work address calibration and precision. A median-division binning algorithm for spectroscopic surveys recursively partitions the xx1 versus xx2 plane until each cell contains between xx3 and xx4 spectroscopic stars, then estimates

xx5

with the empirical calibration xx6 (Mints et al., 2018). Separately, injection-based calibration of empirically measured selection functions shows that the effective number of found injections must satisfy xx7 for a well-behaved posterior, and that the number of found injections scales linearly with the number of detected objects in the population (Farr, 2019).

Selection functions also enter weak-lensing systematics. For galaxy samples with complex, implicit targeting, the effective slope controlling magnification bias is calibrated from realistic mocks rather than from a nominal flux limit. For BOSS-like samples, the calibrated values are xx8 for xx9 and q\vec q0 for q\vec q1, values then inserted into the standard magnification term q\vec q2 (Wietersheim-Kramsta et al., 2021).

3. Selection effects in dynamical and spatio-temporal inference

In stellar dynamics, the crucial distinction is between purely spatial selection and phase-space selection. If q\vec q3 is the true tracer DF and the survey selection is purely spatial, then

q\vec q4

so at fixed q\vec q5,

q\vec q6

This invariance fails when selection depends on velocity or on phase-space quantities. The conditional Deep Potential method exploits exactly this cancellation: it learns q\vec q7, the gravitational potential q\vec q8, and the full tracer density q\vec q9 under the stationary collisionless Boltzmann equation without explicitly modeling the spatial selection function. In mock tests with complex dust-induced angular incompleteness, it recovers S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)0 to S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)1 accuracy across all S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)2, recovers S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)3 to S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)4 down to S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)5, and uses S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)6 thousand parameters for S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)7, compared with S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)8 million parameters for an explicit selection-function network. The same paper stresses that the approach fails without modification under kinematic or magnitude-dependent kinematic selection (Kalda et al., 1 Dec 2025).

In animal movement, selection functions appear as statistical weights and as mechanistic objects. Resource selection functions model locations through a weighted distribution over space; step selection functions condition availability on the previous location. In the mechanistic construction derived from ecological diffusion, the selection term is

S(x)=P(included in samplex)S(x)=P(\text{included in sample}\mid x)9

and the availability kernel is Gaussian, so the stepwise point-process model coincides with the PDE fundamental solution. This yields a direct link between covariates, movement probability S(q)S(\vec q)0, and residence time

S(q)S(\vec q)1

The paper’s mountain lion application interprets the fitted model in terms of movement speed through ridges, slopes, and lower-elevation habitats rather than merely relative habitat preference (Hooten et al., 2019).

4. Choice, games, and higher-order semantics

In axiomatic decision theory, selection functions appear as choice functions over feasible sets. Let S(q)S(\vec q)2 be a real vector space and S(q)S(\vec q)3 the set of non-empty subsets of S(q)S(\vec q)4. A choice function is S(q)S(\vec q)5, where S(q)S(\vec q)6. The paper on strict partial orders studies “proper orders” S(q)S(\vec q)7 satisfying irreflexivity, transitivity, positive homogeneity, and translation invariance, and defines the maximality-based selector

S(q)S(\vec q)8

For a family S(q)S(\vec q)9 of proper orders,

S(M,z)S(M,z)00

which coincides with E-admissibility in imprecise probability. The representation theory proceeds through desirability cones S(M,z)S(M,z)01, rejection kernels S(M,z)S(M,z)02, and axioms PCS(M,z)S(M,z)03–PCS(M,z)S(M,z)04, yielding a general characterization of proper choice functions generated by sets of strict partial orders (Bock, 2020).

In higher-order computation, a selection function has type

S(M,z)S(M,z)05

Escardó and Oliva show that S(M,z)S(M,z)06 forms a monad with unit

S(M,z)S(M,z)07

and a bind operation that threads the continuation S(M,z)S(M,z)08. More importantly, the binary product of selection functions encodes backward induction, and its iterated forms yield the explicitly controlled product and implicitly controlled product. Over System S(M,z)S(M,z)09, the explicitly controlled product is S(M,z)S(M,z)10-equivalent to Spector’s bar recursion, while the implicitly controlled product is S(M,z)S(M,z)11-equivalent to modified bar recursion; both suffice, with System S(M,z)S(M,z)12, to interpret full classical analysis via the dialectica interpretation and modified realizability (Escardo et al., 2014).

The 2025 work on handling the selection monad reintroduces this higher-order object into programming-language semantics. It takes the selection monad

S(M,z)S(M,z)13

as the semantic basis for effect handlers that, in addition to delimited continuations, receive choice continuations exposing possible future losses. In the operational semantics of S(M,z)S(M,z)14, a handler for an operation receives both S(M,z)S(M,z)15 and S(M,z)S(M,z)16, enabling programmer-defined optimization heuristics that inspect future loss rather than relying on an externally fixed argmin or argmax. The paper establishes progress, type soundness, and, under a mild hierarchical constraint on effect signatures, termination, then gives a denotational semantics in an augmented selection monad and proves soundness and adequacy (Plotkin et al., 4 Apr 2025).

Game theory supplies a further interpretation. For the finite nonempty powerset monad S(M,z)S(M,z)17, a nondeterministic selection function has type

S(M,z)S(M,z)18

Its product computes sets of admissible plays in sequential games of perfect information. The paper identifies two key structural properties—witnessing and upwards closure—and proves that they characterize when the product coincides with the set of rational plays. It also proves a negative result: no nondeterministic selection function computes the set of all subgame perfect Nash equilibrium plays unless it is constant. A positive substitute is obtained via strict-dominance selection functions S(M,z)S(M,z)19, whose products compute sequential versions of iterated removal of strictly dominated strategies (Bolt et al., 2018).

5. Adaptive and learned selection in optimization and machine learning

In evolutionary algorithms, selection functions determine which individuals survive and reproduce. Sel4Sel replaces hand-crafted selectors with a learned internal fitness S(M,z)S(M,z)20 computed by a three-layer feedforward network over the features Underlying Fitness, Rank, Age, Novelty, Noise, and Generation. Inner-loop replacement is deterministic tournament selection: S(M,z)S(M,z)21 The outer loop uses evolution strategies to optimize S(M,z)S(M,z)22 for evolvability, measured as final population fitness after S(M,z)S(M,z)23 generations. On Convex Bits, Hashed Bits, and Deceptive Bits, Sel4Sel achieves S(M,z)S(M,z)24, S(M,z)S(M,z)25, and S(M,z)S(M,z)26, respectively, matching greedy fitness on smooth landscapes and outperforming novelty-only, minimal-criterion, and random-drift baselines on deceptive or uncorrelated ones. Feature correlations reveal an emergent “explore first, exploit later” schedule: high positive correlation with novelty early, then near-perfect correlation with fitness and rank later (Frans et al., 2021).

In functional data analysis, adaptive basis selection is cast as a spike-and-slab problem. The representation

S(M,z)S(M,z)27

uses Bernoulli indicators S(M,z)S(M,z)28 to zero coefficients with positive probability, thereby selecting both the number of bases and which bases are active. The Gibbs sampler updates S(M,z)S(M,z)29, S(M,z)S(M,z)30, S(M,z)S(M,z)31, S(M,z)S(M,z)32, and S(M,z)S(M,z)33 from closed-form conditionals. In simulations with B-spline and Fourier bases, the method recovers the true zero and nonzero coefficients closely, and in the trigonometric experiment it selects only two bases out of S(M,z)S(M,z)34 in the Fourier family, consistent with the generating curve S(M,z)S(M,z)35. The paper reports that the proposed model can outperform Bayesian LASSO in this functional, multi-curve setting while also quantifying uncertainty in the selection process (Sousa et al., 2022).

Selection functions also appear in differentiable subset selection through submodular modeling. FLEXSUBNET learns monotone, non-monotone, and monotone S(M,z)S(M,z)36-submodular set functions by recursively applying learned concave functions to modular functions. For the monotone model,

S(M,z)S(M,z)37

For S(M,z)S(M,z)38-submodularity, the paper proves that if S(M,z)S(M,z)39 is increasing and satisfies

S(M,z)S(M,z)40

then S(M,z)S(M,z)41 is monotone S(M,z)S(M,z)42-submodular for S(M,z)S(M,z)43. The learned selector is then paired with an order-sensitive probabilistic greedy rule and an adversarial Gumbel–Sinkhorn permutation relaxation to learn from S(M,z)S(M,z)44 supervision. On synthetic submodular targets and on Amazon baby-registry subset selection, FLEXSUBNET outperforms Set Transformer, Deep Sets, DSF, SubMix, and non-trainable Facility Location, DPP, and Disparity-Min baselines (De et al., 2022).

6. Selection as search control in logic and automated reasoning

In DatalogS(M,z)S(M,z)45, selection functions mark chase positions where only finitely many values arise. A selection function is any map S(M,z)S(M,z)46 with S(M,z)S(M,z)47. The extreme cases are S(M,z)S(M,z)48 and S(M,z)S(M,z)49; the weakly-sticky choice is S(M,z)S(M,z)50, and the paper introduces S(M,z)S(M,z)51 from existential dependency ranks. Stickiness modulo S(M,z)S(M,z)52 defines the semantic class S(M,z)S(M,z)53, and the query-answering procedure S(M,z)S(M,z)54 is terminating, sound, and complete for S(M,z)S(M,z)55. In particular, conjunctive query answering is in PTIME for Sticky, WS, and the new JWS class, and JWS—unlike WS—is closed under magic-sets rewriting (Bertossi et al., 2021).

Automated theorem proving uses a different notion: literal selection in superposition calculi. The paper “Selecting the Selection” distinguishes classical completeness-preserving selection from intentionally incomplete strategies that reduce search-space growth. It formulates the completeness condition as “select either a negative literal or all maximal literals with respect to S(M,z)S(M,z)56,” then proposes two refinements. Quality-based selection composes literal preorders such as heavier weight, fewer variables, fewer top-level variables, avoidance of positive equality, and preference for negative polarity. Lookahead selection estimates the number of immediate children a clause would produce if a given literal were selected, using

S(M,z)S(M,z)57

over the current term indexes S(M,z)S(M,z)58. In experiments with Vampire on S(M,z)S(M,z)59 TPTP FOF and CNF problems, S(M,z)S(M,z)60 were solved by at least one strategy with AVATAR enabled, and the incomplete lookahead strategy 1011 was the strongest overall. The paper reports selection overheads of about S(M,z)S(M,z)61 of proof time for quality selection, S(M,z)S(M,z)62 for complete lookahead, and S(M,z)S(M,z)63 for incomplete lookahead, arguing that restricting inference proliferation is often more useful in practice than preserving completeness at every step (Reger et al., 2016).

Taken together, these literatures show that “selection function” is best understood as a family resemblance term. In some fields it denotes inclusion probabilities needed for unbiased inference; in others, admissible-choice operators, monadic continuations, or search-control policies. The underlying mathematical forms differ sharply, but the recurring problem is the same: how a system maps latent possibilities into realized observations, decisions, or derivations under explicit structural constraints.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Selection Functions.