Relative Size Framework: Cross-disciplinary Overview
- Relative Size Framework is a cross-disciplinary concept that defines 'size' relative to a reference structure such as filters, ranks, or distributions.
- It unifies diverse methodologies from semigroup theory to language models, balancing algebraic characterizations, statistical scaling, and operational calibrations.
- Applications span semantic segmentation, speaker diarization, and replication analysis, enabling robust comparisons and improved model performance.
Searching arXiv for recent and foundational uses of “relative size framework” and closely related formulations. The expression “Relative Size Framework” appears in several non-equivalent research literatures. In semigroup theory, it denotes filter-parametrized notions of largeness, thickness, and prethickness; in language modeling, a rank-based scaling framework centered on the probability that the correct token lies in the top-; in semantic segmentation, supervision by approximate relative object-size distributions; and in on-device speaker diarization, a clustering rule in which minimum cluster size is scaled by the number of embeddings in a recording. Other uses concern retail shelf reasoning, commonsense object-size inference, replication analysis through the relative effect size , controlled-ratio sampling for relative risk and odds ratio, and small-size relative -approximations in range spaces (Protasov et al., 2015, Yue et al., 23 Oct 2025, Fan et al., 10 Mar 2025, Yamaguchi, 7 Jun 2026).
1. Scope and recurring meaning
Across these literatures, the term does not name a single standardized theory. It names a family of constructions in which “size” is not treated absolutely but is conditioned on a reference object such as a filter, a rank threshold, an image-level size prior, an embedding budget, or a prescribed ratio of sampling effort. This suggests a common design principle: a quantity is declared large, small, or adequate only relative to an ambient structure.
| Domain | Relative quantity | Reference structure |
|---|---|---|
| Semigroups | -large, -thick, -prethick | filter |
| LLMs | and its scaling with | top- rank threshold |
| Semantic segmentation | class-size distribution 0 | image-level relative area |
| Speaker diarization | 1 | embedding count 2 |
| Replication analysis | 3 | original study effect |
| Binary-response estimation | average sample-size ratio | prescribed population-allocation ratio |
The same phrase is therefore best understood as a cross-disciplinary label for reference-conditioned size rather than as a single doctrine. In some fields the reference is algebraic, in others statistical, geometric, or algorithmic.
2. Filter-relative largeness in semigroups
In the semigroup literature, the framework is explicit and foundational. Given a semigroup 4 and a filter 5 on 6, the paper "Relative size of subsets of a semigroup" (Protasov et al., 2015) defines 7-large, 8-thick, 9-prethick, and 0-small subsets by inserting 1 into the classical notions of syndetic, thick, piecewise syndetic, and small. The basic definitions are: 2
3
4
When 5, these collapse to the classical absolute notions. The same paper introduces 6-extrathick sets, namely sets that belong to every ultrafilter extending 7, and develops ultrafilter characterizations in 8, including: 9 is 0-large iff 1, and 2 is 3-thick iff 4 for some 5 (Protasov et al., 2015).
The algebraic structure becomes sharper when 6 is a semigroup filter. Then 7 is a compact subsemigroup, minimal left ideals of 8 become available, and 9-prethick sets are characterized by their interaction with the union 0 of minimal left ideals. The paper proves that, for a left inverse invariant filter, 1 is 2-prethick iff 3, and that the family of all 4-prethick subsets is partition regular. In groups, 5-prethick sets are exactly the sets that are not 6-small (Protasov et al., 2015).
A more abstract reformulation is given in "Algebraic characterizations of some relative notions of size" (Christopherson et al., 2021). There, largeness is modeled by stacks, filters, grills, and ultrafilters, and the mesh operator
7
is used to formalize duality. For stacks 8, the paper defines 9-syndetic and 0-thick sets and proves the exact duality
1
It also introduces relative piecewise syndeticity, proves a relative Brown-type partition regularity theorem, and shows that classical piecewise syndeticity can be recovered as a composition of relative syndetic and relative thick notions. In this setting, the framework is not merely terminological: it is a calculus of size notions indexed by filters, duals, and semigroup products (Christopherson et al., 2021).
3. Mathematical and biological size-structured formulations
A distinct mathematical use appears in computational geometry. "Small-Size Relative 2-Approximations for Well-Behaved Range Spaces" (Ezra, 2012) studies finite range spaces 3 and asks for a sample 4 such that for every range 5, the empirical measure 6 approximates 7 relatively when 8 and absolutely when 9. The relative 0-approximation condition is: 1 with additive error at most 2 otherwise. For general VC-dimension, the known bound is 3, but for well-behaved range spaces—those in which the number of ranges of size at most 4 is 5—the paper improves this to
6
and shows that such approximations can be constructed in expected polynomial time. For constant 7, the result yields 8-nets that are also relative approximations, and for points with axis-parallel boxes in two and three dimensions, and for points with fat triangles in the plane, the resulting bound matches the optimal bound for 9-nets (Ezra, 2012).
A different discrete-algebraic meaning appears in "On the relative size of toric bases" (Tatakis et al., 2019). There the objects compared are the Graver basis, universal Gröbner basis, a Markov basis, and the set of circuits of a toric ideal. The main theorem states that if 0 and 1 are any two of these bases with 2, then there is no polynomial on the size or on the maximal degree of the elements of 3 which bounds the size or the maximal degree of the elements of 4. In this literature, “relative size” refers to asymptotic incomparability between canonical bases rather than to an ambient filter or a sampling rule (Tatakis et al., 2019).
In ecology, the phrase names a size-structured dynamical framework. "Food web framework for size-structured populations" (Hartvig et al., 2010) makes body size and size at maturation 5 the organizing coordinates of a food web model. Each species is represented as a size spectrum 6, species identity enters only through the trait 7, and parameters are made species independent through scaling with individual body size and size at maturation. Predation is driven by predator–prey mass ratios, allocation to reproduction depends on 8, and the analytical approximation assumes a power-law community spectrum. Here relative size is neither purely geometric nor purely combinatorial; it is the state variable of the biological system itself (Hartvig et al., 2010).
4. Relative ordering and model scaling in LLMs
In neural language modeling, the framework is explicitly rank-based. "Relative-Based Scaling Law for Neural LLMs" (Yue et al., 23 Oct 2025) argues that cross-entropy is an absolute-based metric: it measures the probability mass assigned to the correct token but ignores its ranking among alternatives. The paper therefore defines Relative-Based Probability
9
where 0 is the rank of the ground-truth token. 1 is the fraction of positions where greedy decoding outputs the correct token, and 2 is the fraction where the correct token lies in the top-3 predictions.
The proposed scaling law is
4
or equivalently
5
with 6 the number of non-embedding parameters. For 7, the exponent is reported as 8–9 depending on dataset and model family; for moderate 0, the same power law holds and 1 increases with 2. Empirically, the law is tested on Pythia, GPT-2, OPT, and Qwen2.5, across datasets including Wikipedia, C4, Github/HumanEval, HotpotQA, Open Australian Legal Corpus, allenai/C4, and pile-uncopyrighted. For 3, 4 versus 5 exhibits straight lines in log–log plots, typically with 6–0.99; for 7, 8 usually stays above 0.9; and the behavior breaks down when 9 approaches the vocabulary size (Yue et al., 23 Oct 2025).
The paper emphasizes that cross-entropy and 00 show numerically very close scaling, with slope differences often below 01 and 02, yet the interpretation differs. Cross-entropy tracks absolute mass on the correct token, whereas RBP tracks rank-based accessibility under greedy or top-03 decoding. This distinction is used to model emergence: under independence and stationarity assumptions, sequence-level success over 04 tokens is
05
so smooth token-level power-law scaling yields sigmoid-like sequence-level success curves. The paper also proposes a lognormal rank distribution hypothesis to explain why cross-entropy scaling and RBP scaling have nearly identical slopes (Yue et al., 23 Oct 2025).
5. Visual learning, scene understanding, and physical reasoning
In semantic segmentation, the framework is a supervision scheme based on approximate relative object-size distributions. "Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation" (Fan et al., 10 Mar 2025) replaces pixel masks with an image-level categorical distribution
06
where 07 is the approximate fraction of image pixels belonging to class 08. A standard segmentation network outputs per-pixel softmax scores 09, and the average prediction
10
is interpreted as the predicted relative size of class 11. Training uses the forward KL-divergence
12
whose zero-avoiding property prevents tagged classes from vanishing. The simplest objective is 13, optionally augmented with partial cross-entropy for scribbles or seeds. On PASCAL VOC with DeepLabv3+, exact size targets give 14 validation mIoU and 15 test mIoU with ResNet101, while synthetically corrupted targets with 16 and WR38 give 17 validation and 18 test mIoU. Human size annotation for cat, dog, and bird yields mean relative errors of 19, 20, and 21, with annotation times of about 22, 23, and 24 seconds per image; the average human 25 is about 26. The method is reported to remain accurate up to about 27 target noise and, for some classes, to perform slightly better than full pixel-level supervision (Fan et al., 10 Mar 2025).
A retail-shelf version appears in "Machine Learning approaches to do size based reasoning on Retail Shelf objects to classify product variants" (Srivastava et al., 2021). The pipeline is modular: object detection on shelf images, brand or product-group classification on crops, and then a size-reasoning stage that uses bounding-box geometry and the context of other products in the same image. Absolute area is treated as unreliable because of viewpoint and scale variation, so the key features are relative area ratios 28, aspect ratios, and co-occurring group labels. The paper proposes per-group XGBoost classifiers for cleaner facings and a GMM-plus-neural-network model for noisy or irregular stacks. In this usage, a relative size framework is a scene-level inference layer added downstream of ordinary vision models (Srivastava et al., 2021).
A broader visual commonsense variant is given in "Are Elephants Bigger than Butterflies? Reasoning about Sizes of Objects" (Bagherinezhad et al., 2016). There object categories are nodes of a size graph, each object size is modeled as log-normal, textual observations provide noisy absolute sizes, and images provide noisy relative size ratios via depth-adjusted bounding boxes. The joint model is trained by maximizing a combined likelihood over textual and visual observations. On a relative size dataset of 41 physical objects and 486 ordered pairwise comparisons, the full model reaches 29 accuracy, versus 30 for a language-only baseline, 31 for a vision-only baseline, 32 for the model using only textual observations, and 33 for the model using only visual observations. Here “relative size” refers to probabilistic comparison between categories rather than to an architectural hyperparameter (Bagherinezhad et al., 2016).
6. Adaptive thresholds, inferential ratios, and measurement design
In speaker diarization, the phrase is used in an explicitly operational sense. "Fast and Robust On-Device Speaker Diarization: Relative Minimum Cluster Size for Stride-Accelerated Pipelines" (Yamaguchi, 7 Jun 2026) studies a Pyannote 3.1-based pipeline accelerated by coarser segmentation stride and per-chunk embedding. On AMI, this recipe is largely DER-neutral and reaches up to 34 speedup on MPS over the CAM++ baseline, but on VoxConverse it causes a DER increase from 35 to 36. The degradation is traced to speaker under-counting in agglomerative clustering, caused by a fixed minimum cluster size interacting with the reduced number of embeddings per speaker. The proposed correction is a relative minimum cluster size
37
with 38 the number of embeddings in the recording. A single value 39 recovers VoxConverse DER to 40, about 41 of the lost accuracy, while keeping AMI essentially flat (Yamaguchi, 7 Jun 2026).
A statistical use of relative size appears in replication methodology. "The assessment of replication success based on relative effect size" (Held et al., 2020) centers analysis on
42
the ratio of replication to original effect estimate. The reverse-Bayes criterion implies a minimum admissible relative effect size 43, and at the proposed golden level a borderline significant original study can achieve replication success only if the replication effect estimate is larger than the original one. The paper argues that this recalibration penalizes shrinkage more appropriately than the two-trials rule requiring significance in both studies, while still allowing conditional power for replication success to take any desired value when the original study is significant and the replication sample size is large enough (Held et al., 2020).
A related design-based notion is developed in "Estimation of relative risk, odds ratio and their logarithms with guaranteed accuracy and controlled sample size ratio" (Mendo, 6 Mar 2025). For two Bernoulli populations with parameters 44 and 45, the paper constructs estimators of 46, 47, and their logarithms, such that the relative mean-square error for RR and OR, or the mean-square error for the logarithms, is below a target value for every 48. Simultaneously, the ratio of average sample sizes from the two populations is kept close to a prescribed value, and the same framework extends to group sampling. Efficiency with respect to the Cramér–Rao bound is reported to be good, and close to 49 for small target error (Mendo, 6 Mar 2025).
Astrophysical usage is again distinct. "Measuring The Soft Excess Region Size Relative to the Corona in AGN With NICER" (Zoghbi et al., 2023) uses variability time scales to compare the size of the soft excess region to the hard X-ray corona in active galactic nuclei. The reported result is source-dependent: for TON S180 the soft excess region is comparable in size to the corona, whereas for MRK 335 and 1H0707-495 the soft excess region is larger than the corona by a factor of 50–51. The paper emphasizes that this is the first time these relative sizes are quantified independently of assumptions of the spectral models (Zoghbi et al., 2023).
7. Common architecture, distinctions, and limitations
These literatures are not terminologically unified, but they exhibit a recurring structural move. An absolute threshold is replaced by a reference-conditioned quantity: a filter 52, a pair of stacks 53, a top-54 rank threshold, an image-level size distribution 55, an embedding count 56, a relative effect size 57, or a prescribed sample-size ratio. This suggests that the phrase “Relative Size Framework” is best interpreted as a methodological pattern rather than a discipline-specific term (Protasov et al., 2015, Yue et al., 23 Oct 2025, Fan et al., 10 Mar 2025, Yamaguchi, 7 Jun 2026).
The same pattern also clarifies why the frameworks are domain-specific. In semigroup theory, the gain is algebraic: ultrafilters, minimal left ideals, and partition regularity become available. In LLMs, the gain is operational: model size is tied to decoding success through token ranks rather than through cross-entropy alone. In semantic segmentation, the gain is supervisory: global class proportions replace masks while retaining strong mIoU. In diarization, the gain is robustness under stride acceleration: a fixed cluster-size threshold is replaced by one scaled to the embedding budget. In replication analysis and controlled-ratio Bernoulli estimation, the gain is inferential calibration: effect sizes and sample allocations are constrained relatively rather than absolutely (Christopherson et al., 2021, Yue et al., 23 Oct 2025, Held et al., 2020, Mendo, 6 Mar 2025).
The limitations are equally heterogeneous. In the language-model setting, the relative-based scaling law is robust for 58 but breaks down when 59 approaches vocabulary size, and the emergence analysis assumes independence and stationarity across positions (Yue et al., 23 Oct 2025). In segmentation, the method addresses semantic rather than instance segmentation, and size estimation becomes harder for very small or highly variable objects (Fan et al., 10 Mar 2025). In diarization, relative minimum cluster size corrects the VoxConverse failure mode but has only marginal effect on MSDWild (Yamaguchi, 7 Jun 2026). In the semigroup literature, the strongest structural results require semigroup filters, left inverse invariance, or extrathickness assumptions (Protasov et al., 2015, Christopherson et al., 2021). In computational geometry, the improved bounds require well-behaved range spaces rather than arbitrary VC classes (Ezra, 2012).
Taken together, these works show that “relative size” can mean relative largeness in an algebraic compactification, relative rank in a token distribution, relative area in an image, relative cluster cardinality in a recording, relative effect magnitude across studies, or relative allocation of sampling effort across populations. The phrase therefore denotes a class of frameworks in which scale is anchored to context, and in which the reference object is mathematically part of the definition rather than an after-the-fact normalization.