---
title: Individual Tendency Learning (ITL)
url: https://www.emergentmind.com/topics/individual-tendency-learning-itl
type: topic
---

# Individual Tendency Learning (ITL)

Individual Tendency Learning (ITL) is a research term that has acquired at least two distinct technical meanings in recent arXiv literature. In one line of work, ITL denotes a framework for modeling human exploration in social dilemmas as a two-stage process of trial and reflection that coevolves with social imitation; this formulation is used to explain asymmetric exploration and its consequences for cooperation on networks [2505.07022]. In another line of work, ITL denotes a paradigm in multi-annotator learning that models annotator-specific labeling behavior rather than collapsing disagreement into a single consensus label, and is accompanied by evaluation criteria designed to test whether a model truly captures individual tendencies and their explanations [2508.10393]. A related but terminologically distinct usage appears in cooperative multi-agent reinforcement learning, where “action tendency” refers to agent-specific policy-value structure and is regularized toward consistency through intrinsic rewards [2406.18152]. Across these settings, the common theme is that individual-level behavioral regularities are treated as structured signals rather than as noise.

## 1. Terminological scope and conceptual core

Recent work uses “Individual Tendency Learning” to designate methods that explicitly preserve and model heterogeneity at the level of the individual. In the multi-annotator setting, the contrast is drawn directly against Consensus-oriented Learning (CoL), whose objective is to aggregate multiple annotators’ labels into a single “ground-truth” label per instance, treating all disagreements as noise; typical CoL methods listed in this context include majority voting, probabilistic EM models (Dawid–Skene), confusion-matrix models (MaDL), and Gaussian fitting (PADL) [2508.10393]. ITL, by contrast, explicitly models each annotator’s unique labeling “tendency”—the systematic pattern of how and when they agree or disagree—and produces annotator-conditional predictions and often annotator-specific explanations [2508.10393].

In the social-dilemma literature, ITL is defined differently but retains the same emphasis on individual-level structure. There, the framework developed in Hou et al. models human “exploration” as neither pure random noise nor a fixed mutation rate, but as a multi-step cognitive process of trial and reflection, embedded alongside social imitation [2505.07022]. The central claim is that exploration intensity can coevolve with social learning, and that this interdependence materially affects the emergence of cooperation.

A useful synthesis is that both usages reject the treatment of deviations from aggregate behavior as unstructured perturbations. In one case, disagreement among annotators is informative about background, bias, expertise, or subjective interpretation; in the other, exploratory deviations in strategic behavior are informative about payoff-sensitive trial-and-error and satisficing [2508.10393] [2505.07022]. This suggests that ITL is best understood not as a single method, but as a family of modeling commitments centered on individual-level regularity.

## 2. ITL in social dilemmas: trial, reflection, and imitation

In the ITL model for cooperation on networks, each update event selects a focal player who, with probability $p_{IL}$, performs individual learning and, with complementary probability $1-p_{IL}$, updates by social imitation [2505.07022]. Individual learning is formalized as a two-stage process.

Suppose that at time $t$ the focal player has current strategy $s(t)\in\{C,D\}$ with immediate payoff $\pi(t)$. Under individual learning, the player first reverses the current strategy to $1-s(t)$, then engages in $u$ further rounds of play, receiving $\pi(t+1),\pi(t+2),\dots,\pi(t+u)$. The player computes an “experiential cognition”
$$
E(t,\mu)=\sum_{\ell=1}^{\mu}\lambda_\ell \pi(t+\ell),
$$
where $(\lambda_1,\lambda_2,\dots,\lambda_u)$ is a weighting kernel over future payoffs [2505.07022]. If $E(t,\mu)>\pi(t)$, the player retains the reversed strategy; otherwise the player reverts to the original one. Two special cases are identified: Terminal-payoff (TP) focus, where $(\lambda_1,\dots,\lambda_u)=(0,\dots,0,1)$ so that $E=\pi(t+u)$, and Equal-payoff (EP) focus, where $\lambda_\ell=1/u$ for all $\ell$ [2505.07022].

If individual learning is not selected, the player updates by the standard pairwise-comparison rule. The realized payoff $\pi_i(t)$ is first rescaled into a social value
$$
f_i(t)=1-w+w\pi_i(t),\qquad w\ll 1,
$$
and a random neighbor $j$ is sampled; the focal player adopts $j$’s strategy with probability
$$
P_{i\to j}=\frac{f_j(t)}{f_i(t)+f_j(t)}.
$$
This embeds ITL within weak-selection imitation dynamics while inserting, with probability $p_{IL}$, a multi-step reversal-and-reflection procedure [2505.07022].

The resulting framework is explicitly coupled. Individual learning alters the local composition of cooperators around the focal player, which modifies imitative weight under pairwise comparison; social imitation, conversely, shapes the payoffs that enter future trials. Under weak selection, the average change in cooperator density $x$ is written as
$$
\Delta x \propto (1-\bar p_{IL})\cdot \sum_i \beta_i(b_i-d_i)
\;+\;
\bar p_{IL}\cdot [\Psi_{D\to C}-\Psi_{C\to D}],
$$
where $\beta_i$ is the probability that $i$ is imitated, $b_i-d_i$ is the net bias in the pairwise-comparison rule, and $\Psi_{s\to s'}$ is the probability that a focal player with strategy $s$ ends reflection in $s'$ [2505.07022]. Under TP focus, the transition probabilities are further expressed in matrix form:
$$
\Psi_{D\to C}(u)=P(t)\cdot (G\odot P^u)\cdot 1,\qquad
\Psi_{C\to D}(u)=P(t)\cdot (H\odot P^u)\cdot 1.
$$
This formalism makes the feedback between individual and social learning explicit [2505.07022].

## 3. Asymmetric exploration and the promotion of cooperation

A central contribution of the social-dilemma ITL framework is its treatment of exploration-probability as payoff-dependent. To capture the empirical observation that “satisfied” players explore less, the individual-learning probability is written as
$$
p_{IL}^i=P_0\cdot g(\pi_i),\qquad g'(\cdot)<0.
$$
In the simplest simulation implementation,
$$
p_{IL}^i=P_0\cdot (n_c/k),
$$
where $n_c$ is the number of cooperator-neighbors [2505.07022]. Because payoff grows with $n_c$, this implements $p_{IL}^i\propto 1/\pi_i$, so highly paid individuals are less likely to engage in trial-and-error. This negative payoff–exploration correlation produces asymmetric exploration rates $f_{C\to D}\neq f_{D\to C}$ and is linked in the paper to observations reported by Traulsen et al. and Rand et al. [2505.07022].

The analytical and simulation findings distinguish between constant and adaptive exploration. When $p_{IL}$ is constant, increasing reflection depth $u$ causes the critical benefit-to-cost ratio for cooperation, $(b/c)^*$, to decrease monotonically and approach the ideal network-reciprocity threshold $b/c>k$ [2505.07022]. Under TP focus, the paper reports an approximate power-law decay
$$
(b/c)^*-k\sim u^{-\gamma},\qquad \gamma\approx 0.5.
$$
Deeper reflection therefore makes cooperation easier, but any nonzero constant exploration still raises $(b/c)^*$ above $k$ [2505.07022].

The adaptive case changes the qualitative conclusion. When $p_{IL}^i$ is negatively correlated with payoff, the individual-learning term in the weak-selection decomposition no longer uniformly penalizes cooperation. For $u\approx 500$ and $k=4$, the paper reports $(b/c)^*$ dropping even below $k/2$, which is interpreted as stable cooperation emerging in very harsh environments [2505.07022]. The proposed intuition is asymmetric: defectors in low-cooperator clusters, having higher payoffs, explore less, whereas cooperators in low-payoff contexts explore more vigorously, which breaks up defector clusters and seeds new cooperator domains.

The broader significance is that exploration is modeled as an active, payoff-sensitive process rather than blind mutation. The framework thereby addresses mixed evidence for network reciprocity, moody conditional cooperation, and payoff-dependent mutation rates through a unified mechanism in which trial outcomes alter imitation bias and imitation, in turn, shapes future trial payoffs [2505.07022].

## 4. ITL in multi-annotator learning: modeling annotator-specific behavior

In multi-annotator learning, ITL is defined as an alternative to consensus-oriented aggregation. Rather than producing one unified prediction, ITL methods model each annotator’s unique labeling tendency and output annotator-conditional models $f_\theta(x,k)$, one per annotator $k$, often together with annotator-specific explanations such as attention over frames or image regions [2508.10393]. The motivating claim is that disagreements encode valuable information about personal background, bias, expertise, or subjective interpretation, and that collapsing them into consensus can obscure behavioral structure [2508.10393].

The paper "A Unified Evaluation Framework for Multi-Annotator Tendency Learning" introduces evaluation metrics intended to test whether ITL methods actually recover such structure [2508.10393]. For annotators indexed by $k=1,\dots,M$, with true labels $Y_k=\{y_i^{(k)}:i\in S_k\}$ and overlap sets $S_{kl}=S_k\cap S_l$ satisfying $|S_{kl}|\ge \tau$, the ground-truth inter-annotator consistency matrix $M\in\mathbb{R}^{M\times M}$ is defined by
$$
m_{kl}=\kappa\bigl(Y_k|_{S_{kl}},Y_l|_{S_{kl}}\bigr),
$$
where $\kappa(\cdot,\cdot)$ is Cohen’s kappa [2508.10393]. Given model predictions $\hat Y_k=\{f_\theta(x_i,k):i\in S_k\}$, the predicted consistency matrix $M'$ is
$$
m'_{kl}=\kappa\bigl(\hat Y_k|_{S_{kl}},\hat Y_l|_{S_{kl}}\bigr).
$$
The Difference of Inter-annotator Consistency (DIC) metric is then
$$
\mathrm{DIC}=\frac{\|M-M'\|_F}{\|M\|_F},
$$
with lower values indicating better preservation of the true consistency structure [2508.10393].

The second metric, Behavior Alignment Explainability (BAE), evaluates whether explanations reflect true behavioral similarity. Ground-truth behavioral similarity is defined by kappa:
$$
S^{\text{true}}_{ij}=\kappa\bigl(Y_i|_{S_{ij}},Y_j|_{S_{ij}}\bigr).
$$
Model-derived similarity can be computed at the feature level using average feature embeddings
$$
F_i^{avg}=(1/|S_i|)\sum_{x\in S_i} f_{\text{feature}}(x,i),
$$
with
$$
S^{\text{feature}}_{ij}=\cos(F_i^{avg},F_j^{avg}),
$$
or at the region level for attention-based models using average attention maps
$$
A_i^{avg}=(1/|S_i|)\sum_{x\in S_i}\mathrm{Attention}(x,i),
$$
with
$$
S^{\text{region}}_{ij}=\cos(A_i^{avg},A_j^{avg}).
$$
For either model-derived similarity,
$$
\mathrm{BAE}=1-\frac{\|S^{\text{model}}-S^{\text{true}}\|_F}{\|S^{\text{true}}\|_F},
$$
so higher BAE indicates better alignment [2508.10393]. Multidimensional Scaling (MDS) is then used to project distances $(1-\text{similarity})$ into 2D for visual inspection of annotator clustering [2508.10393].

This formulation makes ITL in annotation settings a two-part problem: preserving the relational structure of annotator behaviors in predictions, and preserving that same structure in explanations. The explicit separation between tendency capture and explanation faithfulness is one of the paper’s main conceptual contributions [2508.10393].

## 5. Benchmarks, baselines, and empirical findings in multi-annotator ITL

The unified evaluation framework is tested on two datasets: STREET, consisting of urban scene images labeled by 10 annotators across five “impression” dimensions—Happiness, Healthiness, Safety, Liveliness, and Orderliness—and AMER, a video-based emotion recognition dataset with 13 annotators and temporal emotion labels [2508.10393]. Representative ITL baselines are D-LEMA, PADL, MaDL, and QuMAB. Input preprocessing uses resize to $224\times 224$ and normalization; training follows original protocols per method with the same hyperparameters where applicable; hardware is listed as $4\times$ NVIDIA V100 GPUs [2508.10393].

The main reported quantitative results show that DIC and BAE differentiate methods more sharply than conventional metrics. On DIC, QuMAB obtains the lowest values across all listed settings, including STREET-Safety at $0.24\pm 0.03$ and AMER at $0.23\pm 0.02$, compared with D-LEMA at $0.36\pm 0.02$ and $0.42\pm 0.03$, PADL at $0.32\pm 0.02$ and $0.36\pm 0.01$, and MaDL at $0.29\pm 0.01$ and $0.31\pm 0.01$ [2508.10393]. The paper further states that, on Street (Safety) and AMER, DIC exhibits a larger discriminative range than Accuracy, Fleiss’ $\kappa$, and Pearson Corr [2508.10393].

On BAE, QuMAB again scores highest in the tabulated results, with feature-level BAE of $0.54\pm 0.02$ on STREET-Safety and $0.52\pm 0.02$ on AMER; its region-level values are $0.57\pm 0.02$ and $0.55\pm 0.02$, respectively [2508.10393]. The paper also compares BAE with alternative explainability metrics on Safety and AMER, stating that BAE shows greater variance, with standard deviation approximately $0.05$–$0.06$, than Cosine or Gradient [2508.10393].

| Component | Definition or result | Source |
|---|---|---|
| DIC | $\|M-M'\|_F / \|M\|_F$; lower better | [2508.10393] |
| BAE | $1-\|S^{\text{model}}-S^{\text{true}}\|_F / \|S^{\text{true}}\|_F$; higher better | [2508.10393] |
| Qualitative finding | QuMAB heatmaps and MDS clusters more closely match ground truth | [2508.10393] |

Qualitatively, the paper reports that Figure 3 heatmaps for STREET Safety show QuMAB’s predicted $\kappa$-matrix closely resembling ground truth, while D-LEMA exhibits large structural distortions; Figure 4 MDS projections show QuMAB’s feature-level and region-level embeddings forming clusters that match true high-agreement groupings more closely than competing methods [2508.10393]. The paper presents these visual analyses as confirmation that DIC and BAE are measuring meaningful structural properties rather than only pointwise predictive accuracy.

The stated limitations are also important for delimiting the applicability of the framework. DIC requires sufficient overlap $\tau$ between annotators; very sparse labels may weaken reliability. Region-level BAE is only applicable to attention-based models and yields modest gains. MDS-based interpretability depends on similarity-to-distance conversion choices [2508.10393]. Proposed improvements include incorporating human-derived signals such as eye-tracking or cursor paths, exploring alignment methods beyond MDS such as Procrustes analysis, and extending the framework to streaming or evolving annotation scenarios with dynamic tendencies [2508.10393].

## 6. Relation to “action tendency” in cooperative multi-agent reinforcement learning

A related but separate research thread studies “action tendency” in cooperative multi-agent reinforcement learning. In "Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement Learning," the tendency of agent $i$ at time $t$ is its vector of Q-values,
$$
Q_i(o_i^t,\cdot)\in\mathbb{R}^{|\mathcal A|},
$$
and divergent action tendencies among agents are identified as an obstacle to the training efficiency of CTDE-based value-decomposition methods such as VDN and QMIX [2406.18152].

The proposed mechanism is not labeled ITL, but it shares the emphasis on modeling individual-level tendency explicitly. Each agent $i$ maintains an action model
$$
\mathcal F_i^{AM}(o_{i_j}^t,\cdot;\omega_i)\approx Q_i(o_i^t,\cdot;\theta_i),
$$
where $o_{i_j}^t$ is the imagined observation of $i$ from neighbor $j$’s perspective [2406.18152]. The action model is trained by supervised regression to the agent’s own Q-values, and its prediction error is converted into a cooperative intrinsic reward
$$
r_i^I
=
-\frac{1}{|S(i)|}\sum_{j\in S(i)}
\left\|
\mathcal F_i^{AM}(o_{i_j},\cdot;\omega_i)-Q_i(o_i,\cdot;\theta_i)
\right\|_2.
$$
This intrinsic reward penalizes disagreement between an agent’s actual action tendency and what its neighbors predict it will do [2406.18152].

To integrate these rewards, the paper introduces Reward-Additive CTDE (RA-CTDE), a factorization of the global TD-loss into per-agent losses and proves gradient-level equivalence to standard CTDE. The augmented reward for agent $i$ becomes
$$
r_i^{tot}=r^{ext}+\beta r_i^I,
$$
and the resulting IAM loss is
$$
\mathcal L_i^{IAM}(\theta_i,\phi)
=
\mathbb E_{\mathcal D}
\bigl[
r^{ext}+\beta r_i^I+\gamma \max_{a'}Q_T(\tau',a')-Q_{tot}(\tau,a;\theta,\phi)
\bigr]^2
$$
[2406.18152].

Empirically, the paper reports that embedding IAM into QMIX and VDN yields large gains on SMAC, Google Research Football Academy, and Multi-Agent Particle Environment. The examples given include median win-rate increases on $3s5z\_vs\_3s6z$ from approximately $60\%$ to above $95\%$, on $8m\_vs\_9m$ from approximately $40\%$ to approximately $85\%$, average return on GRF CPT from approximately $0.3$ to $0.8$, and occupancy rate in MPE from approximately $0.6$ to $0.9$ [2406.18152]. Although this work does not define ITL as such, it shows that the broader research program of explicit tendency modeling extends beyond annotation and social-dilemma settings.

A plausible implication is that the term “tendency” is serving a unifying methodological role across different areas: it names latent, individual-specific behavioral regularities that can be modeled directly, aligned across agents or annotators, and evaluated structurally rather than only through aggregate task reward or consensus accuracy.

## 7. Interpretive synthesis, misconceptions, and open directions

The most immediate misconception is that ITL denotes a single established framework. Recent arXiv usage does not support that reading. Instead, the term appears in at least two different technical senses: a social-dilemma framework for trial-and-reflection learning coupled to imitation [2505.07022], and a multi-annotator learning paradigm centered on annotator-specific prediction and explanation [2508.10393]. Related work on action tendency consistency in MARL strengthens the case that the underlying idea is broader than any one formalization, but it does not by itself make the terminology uniform [2406.18152].

A second misconception is that individual tendencies are merely nuisance variation. All three cited lines of work reject that premise. In the annotation setting, disagreement is explicitly treated as informative about bias, expertise, and interpretation rather than as noise to be removed [2508.10393]. In social dilemmas, exploratory deviations are modeled as payoff-sensitive, cognitively grounded behavior rather than random mutation [2505.07022]. In MARL, divergence in action tendencies is treated as a coordination bottleneck that can be regularized through intrinsic rewards [2406.18152].

The main methodological commonality is structural evaluation. In the cooperation setting, the relevant structure is the coupled dynamic between imitation bias and trial outcomes. In multi-annotator ITL, it is the inter-annotator consistency matrix and the alignment between behavioral similarity and explanations. In MARL, it is the relation between an agent’s actual Q-vector and neighbor-conditioned predictions of that vector [2505.07022] [2508.10393] [2406.18152]. This suggests that “learning tendencies” is less about fitting isolated labels or actions than about preserving relational organization.

Several open directions are already stated in the underlying works. In multi-annotator ITL, these include richer behavioral ground truth such as eye-tracking and cursor paths, alternative manifold-alignment methods beyond MDS, and dynamic extensions for streaming or evolving tendencies [2508.10393]. In the social-dilemma setting, the reported dependence on reflection depth $u$, payoff-adaptive exploration, and kernel choice indicates that further study of attention span, satisficing, and environment-dependent exploration remains central [2505.07022]. Taken together, the literature portrays ITL as an emerging family of approaches in which individual-specific regularities are formal model objects, explanatory targets, and, increasingly, primary evaluation units.

Source: https://www.emergentmind.com/topics/individual-tendency-learning-itl