Luce-Distributed Permutations
- Luce-distributed permutations are probability models that generate rankings by sequentially sampling items with probabilities proportional to positive weights.
- These models encompass standard, extended, and inverse variants, enabling flexible formulations in ranking, Bayesian inference, and pairwise comparison frameworks.
- Recent research focuses on efficient computational methods, asymptotic analysis, and applications in rank aggregation, causal learning, and secretary-type online decision problems.
Luce-distributed permutations are probability distributions on permutations generated by sequential Luce choices: at each step, one samples without replacement from the remaining items with probabilities proportional to prescribed positive weights. In the statistical literature, the standard Plackett–Luce model, the Extended Plackett–Luce model, and inverse-Luce variants are all instances of this general idea, while broader Gibbs-posterior formulations recover Luce-family noisy-comparison models such as Bradley–Terry–Luce. Recent work studies these distributions from several directions: parametric ranking models, Bayesian inference, message-passing approximations, asymptotic limit theory, stochastic-gradient methods for latent permutations, and secretary-type online decision problems (Johnson et al., 2020, Cantwell et al., 2021, Borga et al., 9 Sep 2025, Pinsky, 11 May 2026, Gadetsky et al., 2019).
1. Sequential-choice definition and equivalent continuous representations
In the standard Plackett–Luce formulation, a complete ranking of entities is a permutation , where is the entity assigned to rank . Each entity has a positive worth parameter , and the model assigns probability
This is a discrete distribution on all permutations. In the logit parameterization used for black-box optimization, for a permutation and score vector ,
0
Both expressions encode the same stagewise mechanism: the first item is chosen from all items, the second from the remaining items, and each selection obeys Luce’s choice axiom (Johnson et al., 2020, Gadetsky et al., 2019).
A weight-based formulation emphasizes sampling without replacement. For positive weights 1 with total mass
2
define
3
The law on 4 is then
5
Equivalently, if 6 is the draw order, then 7, so 8 is the draw time of label 9 (Borga et al., 9 Sep 2025).
Two continuous representations are central. In one, if 0 are independent exponentials with mean 1, then comparisons among the 2’s encode the Luce order. In the other, if
3
then the event 4 has probability exactly equal to the Plackett–Luce probability. These representations are the basis of both asymptotic analysis and low-variance gradient estimators (Borga et al., 9 Sep 2025, Gadetsky et al., 2019).
The mode has a simple form in the standard model. Under standard Plackett–Luce, the modal ordering is
5
and, in the score parameterization, the mode is the permutation that sorts scores in descending order (Johnson et al., 2020, Gadetsky et al., 2019).
2. Extended, reverse, and inverse Luce models
The Extended Plackett–Luce (EPL) model generalizes standard Plackett–Luce by introducing a choice order
6
which specifies the order in which ranks are filled during the ranking process. Its probability mass function is
7
Equivalently, it is the standard Plackett–Luce probability applied to permuted data
8
The model therefore induces a flexible discrete distribution on 9 in which the sequence of Luce choices is itself permutation-valued (Johnson et al., 2020).
Two special cases are explicit. If 0, EPL reduces to the standard forward Plackett–Luce model. If 1, it becomes the reverse Plackett–Luce model. The worth parameters 2 still govern selection probabilities, but for general 3 their direct interpretation as “preference for rank 4, rank 5, \dots” is no longer immediate because ranking stages are traversed in the order 6, not necessarily from best to worst. The modal ranking under EPL is
7
so the choice order re-labels the ranking stages (Johnson et al., 2020).
The parameterization is mixed discrete/continuous:
- discrete component: 8,
- continuous component: 9.
The likelihood is invariant under scalar multiplication,
0
so the model is identifiable only up to scale. The cited work notes that this is not a major substantive issue, but it can affect MCMC mixing (Johnson et al., 2020).
A distinct but closely related variant is the inverse-Luce model. For positive weights 1, the direct Luce distribution on 2 is
3
Its inverse-Luce version is the pushforward under inversion 4: 5 The inverse-Luce model is the one that becomes tractable in the secretary problem when the smallest number is the top rank (Pinsky, 11 May 2026).
3. Statistical inference and computational methods
Bayesian inference for EPL starts from data 6 and likelihood
7
The prior is factored as
8
For the choice order, the cited work uses a Plackett–Luce prior
9
which is uniform over 0 if 1 for all 2. Conditional on 3, the worth parameters have independent Gamma priors,
4
A central feature is a mode-preserving prior predictive framework: the hyperparameters 5 are chosen so that the modal ordering of the prior predictive distribution is preserved across different choice orders (Johnson et al., 2020).
Posterior sampling is carried out with Metropolis-coupled MCMC (MC6) / parallel tempering because the posterior over 7 can have local modes separated by large distances in permutation space. The algorithm alternates between updating 8, updating 9, rescaling 0, and swapping states between tempered chains. The paper recommends geometric temperature spacing and swapping adjacent chains only. Predictive inference is explicit: 1 The posterior predictive mode is interpreted as the aggregate ranking, while discrepancies
2
are used to detect lack of fit (Johnson et al., 2020).
A different inference perspective is given by Gibbs posteriors on permutations from pairwise comparison data. With a directed comparison graph 3, a uniform prior over permutations, and edge factor 4, the posterior is
5
This is interpreted as
6
The framework includes a hard step-function Hamiltonian for partial orders and a Bradley–Terry–Luce noisy-comparison model with
7
Belief propagation computes approximate marginal position distributions, while the Bethe free energy is used both to approximate the number of linear extensions and to perform model selection between competing ranking models such as the step-function model, Bradley–Terry–Luce, and related alternatives (Cantwell et al., 2021).
For optimization over latent permutations, low-variance score-function estimators have been developed specifically for the Plackett–Luce distribution. The task is
8
where 9 is a permutation and 0 may be non-differentiable or only available as a black box. The cited work extends REBAR and RELAX to permutations by using the Gumbel representation 1, deriving a conditional distribution 2 as a sequence of truncated Gumbels, and constructing PL-REBAR and PL-RELAX control variates. The support size is 3, but Plackett–Luce uses 4 parameters and can be sampled in 5, making it usable in causal structure learning and other black-box permutation problems (Gadetsky et al., 2019).
4. Global and local asymptotic theory
A recent asymptotic theory studies permutations 6 through their permuton and local limits. Set
7
and assume
8
for some positive, finite, measurable function 9. Define
0
The limiting permuton 1 is the law of
2
where 3, 4, independent. Under these assumptions, 5 converges in probability in the permuton sense to the deterministic permuton 6 (Borga et al., 9 Sep 2025).
The limit is explicit. The permuton is absolutely continuous with density
7
For Sukhatme weights 8, one has 9, and the density is asymmetric and singular near 0. Pattern densities are also identified: for 1,
2
where 3 are the order statistics of 4 i.i.d. uniform variables. Thus the limiting pattern law is an average of exact Luce laws with random weights. For Sukhatme weights, 5 and 6, showing a marked global bias (Borga et al., 9 Sep 2025).
The same work emphasizes that exact Luce permutations and permutations sampled from the limiting permuton can differ significantly on fine-scale statistics, even though they share the same global permuton limit. A key example is the distribution of the first or top values: a naive approximation replacing random weights 7 by deterministic weights is false in general. This suggests a separation between global convergence in the permuton sense and finer local or top-rank behavior (Borga et al., 9 Sep 2025).
Local convergence is formulated through consecutive patterns. For fixed 8 and 9, if
00
exists, then
01
in probability. For Sukhatme weights 02, the limiting local frequencies are uniform,
03
for every fixed 04. Hence the model is locally uniform even though it is globally non-uniform. The paper further proves a Berry–Esseen-type CLT for consecutive pattern occurrences and a CLT for inversions. For any positive weights,
05
with an explicit variance formula, and for Sukhatme weights
06
where 07 (Borga et al., 9 Sep 2025).
5. Applications: rank aggregation, causal structure learning, and secretary problems
Empirical studies of EPL emphasize cases where the ranking process is not naturally forward-only. In the Song data example, the posterior puts almost all mass on a nonstandard choice order,
08
with only tiny mass on the standard or reverse Plackett–Luce choice orders. The posterior predictive distribution under EPL assigns observed rankings much more plausible probability than standard PL, and the modal predictive ranking differs between EPL and SPL. In Formula 1 data, with 09 and 10, the posterior suggests a choice order closer to reverse PL in the later stages of the ranking process; EPL gives much more reasonable predictive counts for wins and podiums than standard PL, and its posterior predictive aggregate ranking is closer to the final championship standings in meaningful ways (Johnson et al., 2020).
For causal structure learning, Luce-distributed permutations are used as latent topological orders. A toy objective minimizes
11
and the cited experiments report that REINFORCE has too much variance to be useful, PL-REBAR improves substantially, and PL-RELAX performs best due to the learned control variate. In DAG learning from continuous data, optimization is over permutation matrices 12 through a black-box score
13
while for discrete Bayesian networks a non-differentiable quotient NML score is used. The reported implication is that the Plackett–Luce family is practically useful for black-box permutation optimization, including non-differentiable objectives (Gadetsky et al., 2019).
The secretary problem provides a different application. Here the arrival order of the 14 ranked items is not uniform, but follows either a Luce distribution or a Mallows distribution on 15. For inverse-Luce permutations, when the smallest label is best, the success probability of the threshold rule 16 is
17
For exponential weights 18, the cited work proves the exact identity
19
and further states that, for every 20 and every strategy, the secretary-problem probabilities when the smallest number is treated as highest rank under inverse-Luce weights 21 coincide with those for the corresponding Mallows distribution (Pinsky, 11 May 2026).
The same paper derives asymptotically optimal thresholds in several regimes. For 22 with 23,
24
For 25, 26,
27
again with limiting success probability 28. For fixed 29,
30
and the limiting success probability is
31
For Sukhatme weights 32,
33
while for reverse Sukhatme weights 34,
35
and in both cases the limiting success probability is 36 (Pinsky, 11 May 2026).
6. Scope of the term and distinction from unrelated permutation coordinatizations
The phrase “Luce-distributed permutations” refers in these sources to probabilistic models on 37: sequential weighted sampling without replacement, parametric ranking distributions such as Plackett–Luce and EPL, inverse-Luce laws, and Gibbs posteriors that include Bradley–Terry–Luce as a special case (Johnson et al., 2020, Cantwell et al., 2021, Borga et al., 9 Sep 2025, Pinsky, 11 May 2026, Gadetsky et al., 2019).
It should be distinguished from deterministic combinatorial uses of permutations that do not define probability distributions. A clear example is the theory of join-distributive lattices of length 38 and join-width at most 39. In that setting, there are two descriptions by 40 permutations acting on an 41-element set: the Edelman–Jamison construction inside a powerset lattice and a lattice-theoretic coordinatization by eligible tuples. The key result is
42
and the paper also characterizes join-distributive lattices by trajectories. This is a theory of finite join-distributive lattices, semimodularity, meet-semidistributivity, and trajectory classes of prime intervals, rather than a Luce-type stochastic law on permutations (Czédli et al., 2012).
The distinction matters because both areas use permutation data, but for different purposes. In Luce-type models, permutations are random outcomes whose probabilities are determined by worth parameters, choice orders, or pairwise interaction factors. In the join-distributive lattice setting, permutations serve as coordinates for a deterministic representation theorem. The two subjects therefore share notation but not mathematical content (Czédli et al., 2012).