Selection Functions: Concepts and Applications
- Selection Functions are formal devices that specify how elements from a possibility space are selected, bridging latent models and observed outcomes.
- In astronomy and cluster cosmology, selection functions quantify inclusion probabilities and threshold criteria to model survey completeness and calibration accurately.
- Beyond observational sciences, selection functions guide decision-making in axiomatic choice theory, automated reasoning, and machine learning by structuring choices and optimizing outcomes.
Searching arXiv for relevant papers on “selection functions” across major usages (astronomy, logic/game theory, optimization/ML). Searching arXiv for “selection functions astronomy survey completeness selection functions”. Selection functions are formal devices that specify how elements of a space become selected, retained, or acted upon. Across disciplines, the term denotes different mathematical objects: a survey-inclusion probability , a cluster-detection threshold , a choice map , a higher-order functional , a nondeterministic selector , or a learned internal score used for replacement and survival. What unifies these constructions is not a single formalism but a common role: each mediates between an underlying space of possibilities and an observed subset, selected action, or continuation (Rix et al., 2021, Ascaso et al., 2016, Bock, 2020, Escardo et al., 2014, Bolt et al., 2018, Frans et al., 2021).
1. Conceptual scope and formal heterogeneity
The term has no field-independent canonical definition. In astronomical catalog modeling, a selection function is the probability that an object with observables or is included in a catalog or sample, typically written or (Rix et al., 2021, Everall et al., 2021). In cluster cosmology, the operational definition becomes a threshold object: the minimum halo mass as a function of redshift for which completeness and purity exceed prescribed levels, formalized as
0
with 1 in the baseline construction (Ascaso et al., 2016). In axiomatic choice theory, by contrast, a selection or choice function is a map 2 sending each non-empty feasible set 3 to a chosen subset 4, possibly empty (Bock, 2020). In higher-order semantics, a selection function is a functional of type 5, and the corresponding selection monad supports binary products and iterated products with direct proof-theoretic and game-theoretic interpretations (Escardo et al., 2014). In evolutionary search, a selection function maps individual-level features such as fitness, novelty, age, and rank to deterministic tournament outcomes or selection probabilities, thereby controlling exploration and exploitation (Frans et al., 2021). In Datalog6, a selection function maps a program 7 to a subset of chase positions 8, marking where finitely many values appear (Bertossi et al., 2021).
| Domain | Formal object | Primary role |
|---|---|---|
| Astronomical catalogs | 9, 0 | Inclusion probability |
| Cluster surveys | 1 | Completeness–purity threshold |
| Choice theory | 2 | Chosen subset from feasible set |
| Higher-order computation | 3 | Optimization-aware choice |
| Nondeterministic games | 4 | Set-valued admissible moves |
| Evolutionary search | 5 or tournament rule | Survival and replacement |
| Datalog6 | 7 | Finite-position marking |
A recurrent misconception is that a selection function must always be probabilistic. The literature shows otherwise: some are probabilities, some are threshold indicators, some are set-valued maps, and some are higher-type continuations.
2. Inclusion probabilities, completeness, and calibration in astronomy
In astronomical survey modeling, selection functions are the link between a population model and catalog incidence. For a catalog or curated sample, 8 is the probability that an object with attributes 9 appears in the sample, and it enters forward modeling multiplicatively, with nuisance-variable dependence removed by integration. The standard effective-volume construction is
0
which yields 1. In the Gaia white dwarf example, the paper argues that strict volume-limited sampling is often sub-optimal: for 2 mag, the accessible volume ratio in a flux-limited survey is 3, and after explicit selection modeling the inferred white dwarf luminosity–color function peak shifts by 4 magnitudes relative to raw counts (Rix et al., 2021).
Cluster cosmology uses a more operational notion. In stage-IV optical and infrared surveys, completeness is defined as 5, purity as 6, and the selection function is the minimum halo mass satisfying threshold cuts in both quantities. Using the Bayesian Cluster Finder on realistic mock catalogs, Euclid-Optimistic reaches 7 completeness and purity down to 8 for 9, and 0 up to 1. Euclid-Pessimistic has a similar shape with 2 higher mass limits, while LSST has 3 higher mass limits than Euclid-Optimistic up to 4 and reaches 5 by 6. The same work fits
7
with 8, finding intrinsic scatters 9, 0, and 1 dex for Euclid-Optimistic, Euclid-Pessimistic, and LSST, respectively, and negligible redshift evolution within uncertainties (Ascaso et al., 2016).
For Milky Way spectroscopic surveys, empirical selection functions are commonly estimated relative to a photometric parent catalog. In observed space, 2, with 3 typically containing sky position, magnitude, and color. The seestar framework models both the parent photometric density and the field-level spectroscopic selection function with Gaussian mixture models, uses a Poisson point-process likelihood, and combines overlapping fields via
4
It also defines an intrinsic selection function 5, obtained by composing the observable-space selector with geometry, stellar models, and extinction maps (Everall et al., 2019).
Gaia EDR3 subset selection functions sharpen the same principle. The paper on Gaia EDR3 models the probability of having astrometry, 6, or a reported radial velocity as conditional probabilities relative to inclusion in the EDR3 source catalog, with dependencies on 7, sky position, and, for RUWE and RVS, color 8. It uses HEALPix 9, spherical needlets up to 0, and binomial likelihoods per HEALPix–magnitude–color bin. The inferred functions show strong scanning-law and crowding dependence, and the paper emphasizes that ignoring this structure risks biased density and kinematic inference (Everall et al., 2021).
Two further lines of work address calibration and precision. A median-division binning algorithm for spectroscopic surveys recursively partitions the 1 versus 2 plane until each cell contains between 3 and 4 spectroscopic stars, then estimates
5
with the empirical calibration 6 (Mints et al., 2018). Separately, injection-based calibration of empirically measured selection functions shows that the effective number of found injections must satisfy 7 for a well-behaved posterior, and that the number of found injections scales linearly with the number of detected objects in the population (Farr, 2019).
Selection functions also enter weak-lensing systematics. For galaxy samples with complex, implicit targeting, the effective slope controlling magnification bias is calibrated from realistic mocks rather than from a nominal flux limit. For BOSS-like samples, the calibrated values are 8 for 9 and 0 for 1, values then inserted into the standard magnification term 2 (Wietersheim-Kramsta et al., 2021).
3. Selection effects in dynamical and spatio-temporal inference
In stellar dynamics, the crucial distinction is between purely spatial selection and phase-space selection. If 3 is the true tracer DF and the survey selection is purely spatial, then
4
so at fixed 5,
6
This invariance fails when selection depends on velocity or on phase-space quantities. The conditional Deep Potential method exploits exactly this cancellation: it learns 7, the gravitational potential 8, and the full tracer density 9 under the stationary collisionless Boltzmann equation without explicitly modeling the spatial selection function. In mock tests with complex dust-induced angular incompleteness, it recovers 0 to 1 accuracy across all 2, recovers 3 to 4 down to 5, and uses 6 thousand parameters for 7, compared with 8 million parameters for an explicit selection-function network. The same paper stresses that the approach fails without modification under kinematic or magnitude-dependent kinematic selection (Kalda et al., 1 Dec 2025).
In animal movement, selection functions appear as statistical weights and as mechanistic objects. Resource selection functions model locations through a weighted distribution over space; step selection functions condition availability on the previous location. In the mechanistic construction derived from ecological diffusion, the selection term is
9
and the availability kernel is Gaussian, so the stepwise point-process model coincides with the PDE fundamental solution. This yields a direct link between covariates, movement probability 0, and residence time
1
The paper’s mountain lion application interprets the fitted model in terms of movement speed through ridges, slopes, and lower-elevation habitats rather than merely relative habitat preference (Hooten et al., 2019).
4. Choice, games, and higher-order semantics
In axiomatic decision theory, selection functions appear as choice functions over feasible sets. Let 2 be a real vector space and 3 the set of non-empty subsets of 4. A choice function is 5, where 6. The paper on strict partial orders studies “proper orders” 7 satisfying irreflexivity, transitivity, positive homogeneity, and translation invariance, and defines the maximality-based selector
8
For a family 9 of proper orders,
00
which coincides with E-admissibility in imprecise probability. The representation theory proceeds through desirability cones 01, rejection kernels 02, and axioms PC03–PC04, yielding a general characterization of proper choice functions generated by sets of strict partial orders (Bock, 2020).
In higher-order computation, a selection function has type
05
Escardó and Oliva show that 06 forms a monad with unit
07
and a bind operation that threads the continuation 08. More importantly, the binary product of selection functions encodes backward induction, and its iterated forms yield the explicitly controlled product and implicitly controlled product. Over System 09, the explicitly controlled product is 10-equivalent to Spector’s bar recursion, while the implicitly controlled product is 11-equivalent to modified bar recursion; both suffice, with System 12, to interpret full classical analysis via the dialectica interpretation and modified realizability (Escardo et al., 2014).
The 2025 work on handling the selection monad reintroduces this higher-order object into programming-language semantics. It takes the selection monad
13
as the semantic basis for effect handlers that, in addition to delimited continuations, receive choice continuations exposing possible future losses. In the operational semantics of 14, a handler for an operation receives both 15 and 16, enabling programmer-defined optimization heuristics that inspect future loss rather than relying on an externally fixed argmin or argmax. The paper establishes progress, type soundness, and, under a mild hierarchical constraint on effect signatures, termination, then gives a denotational semantics in an augmented selection monad and proves soundness and adequacy (Plotkin et al., 4 Apr 2025).
Game theory supplies a further interpretation. For the finite nonempty powerset monad 17, a nondeterministic selection function has type
18
Its product computes sets of admissible plays in sequential games of perfect information. The paper identifies two key structural properties—witnessing and upwards closure—and proves that they characterize when the product coincides with the set of rational plays. It also proves a negative result: no nondeterministic selection function computes the set of all subgame perfect Nash equilibrium plays unless it is constant. A positive substitute is obtained via strict-dominance selection functions 19, whose products compute sequential versions of iterated removal of strictly dominated strategies (Bolt et al., 2018).
5. Adaptive and learned selection in optimization and machine learning
In evolutionary algorithms, selection functions determine which individuals survive and reproduce. Sel4Sel replaces hand-crafted selectors with a learned internal fitness 20 computed by a three-layer feedforward network over the features Underlying Fitness, Rank, Age, Novelty, Noise, and Generation. Inner-loop replacement is deterministic tournament selection: 21 The outer loop uses evolution strategies to optimize 22 for evolvability, measured as final population fitness after 23 generations. On Convex Bits, Hashed Bits, and Deceptive Bits, Sel4Sel achieves 24, 25, and 26, respectively, matching greedy fitness on smooth landscapes and outperforming novelty-only, minimal-criterion, and random-drift baselines on deceptive or uncorrelated ones. Feature correlations reveal an emergent “explore first, exploit later” schedule: high positive correlation with novelty early, then near-perfect correlation with fitness and rank later (Frans et al., 2021).
In functional data analysis, adaptive basis selection is cast as a spike-and-slab problem. The representation
27
uses Bernoulli indicators 28 to zero coefficients with positive probability, thereby selecting both the number of bases and which bases are active. The Gibbs sampler updates 29, 30, 31, 32, and 33 from closed-form conditionals. In simulations with B-spline and Fourier bases, the method recovers the true zero and nonzero coefficients closely, and in the trigonometric experiment it selects only two bases out of 34 in the Fourier family, consistent with the generating curve 35. The paper reports that the proposed model can outperform Bayesian LASSO in this functional, multi-curve setting while also quantifying uncertainty in the selection process (Sousa et al., 2022).
Selection functions also appear in differentiable subset selection through submodular modeling. FLEXSUBNET learns monotone, non-monotone, and monotone 36-submodular set functions by recursively applying learned concave functions to modular functions. For the monotone model,
37
For 38-submodularity, the paper proves that if 39 is increasing and satisfies
40
then 41 is monotone 42-submodular for 43. The learned selector is then paired with an order-sensitive probabilistic greedy rule and an adversarial Gumbel–Sinkhorn permutation relaxation to learn from 44 supervision. On synthetic submodular targets and on Amazon baby-registry subset selection, FLEXSUBNET outperforms Set Transformer, Deep Sets, DSF, SubMix, and non-trainable Facility Location, DPP, and Disparity-Min baselines (De et al., 2022).
6. Selection as search control in logic and automated reasoning
In Datalog45, selection functions mark chase positions where only finitely many values arise. A selection function is any map 46 with 47. The extreme cases are 48 and 49; the weakly-sticky choice is 50, and the paper introduces 51 from existential dependency ranks. Stickiness modulo 52 defines the semantic class 53, and the query-answering procedure 54 is terminating, sound, and complete for 55. In particular, conjunctive query answering is in PTIME for Sticky, WS, and the new JWS class, and JWS—unlike WS—is closed under magic-sets rewriting (Bertossi et al., 2021).
Automated theorem proving uses a different notion: literal selection in superposition calculi. The paper “Selecting the Selection” distinguishes classical completeness-preserving selection from intentionally incomplete strategies that reduce search-space growth. It formulates the completeness condition as “select either a negative literal or all maximal literals with respect to 56,” then proposes two refinements. Quality-based selection composes literal preorders such as heavier weight, fewer variables, fewer top-level variables, avoidance of positive equality, and preference for negative polarity. Lookahead selection estimates the number of immediate children a clause would produce if a given literal were selected, using
57
over the current term indexes 58. In experiments with Vampire on 59 TPTP FOF and CNF problems, 60 were solved by at least one strategy with AVATAR enabled, and the incomplete lookahead strategy 1011 was the strongest overall. The paper reports selection overheads of about 61 of proof time for quality selection, 62 for complete lookahead, and 63 for incomplete lookahead, arguing that restricting inference proliferation is often more useful in practice than preserving completeness at every step (Reger et al., 2016).
Taken together, these literatures show that “selection function” is best understood as a family resemblance term. In some fields it denotes inclusion probabilities needed for unbiased inference; in others, admissible-choice operators, monadic continuations, or search-control policies. The underlying mathematical forms differ sharply, but the recurring problem is the same: how a system maps latent possibilities into realized observations, decisions, or derivations under explicit structural constraints.