---
title: Illusion of Moral Decline
url: https://www.emergentmind.com/topics/illusion-of-moral-decline
type: topic
---

# Illusion of Moral Decline

The **illusion of moral decline** is the possibility that perceived deterioration in morality reflects changes in moral standards, visibility, reporting, social information, institutional conditions, or evaluative frameworks rather than a corresponding increase in harmful conduct. The concept does not imply that moral decline is never real. It distinguishes observable moral discourse, judgments, and compliance from underlying behavior, motivation, institutions, and ethical conditions. Research using diachronic embeddings, historical language models, social-media annotation, multi-agent public-goods games, agent-based population models, and experiments on collective judgment converges on a central methodological caution: moral change cannot be inferred from surface indicators alone.

## 1. Conceptual distinctions and historical interpretation

Moral-decline claims commonly conflate at least four objects: actual harmful behavior, moral norms, moral awareness, and perceived decline. A rise in reported wrongdoing may indicate more wrongdoing, improved detection, greater willingness to disclose, expanded legal definitions, intensified media coverage, or stronger condemnation. Similarly, more negative moral language may indicate worsening conduct, increased awareness of previously neglected harms, or more explicit denunciation of conduct that is becoming less tolerated.

This distinction is central to the interpretation of historical language. Xie, Pinto, Hirst, and Xu developed a text-based framework that infers changes in public moral sentiment from longitudinal corpora using diachronic word embeddings and morally labeled seed words [2001.07209]. The measured object is moral sentiment as represented in language, not morality itself. The framework does not directly measure actual behavior, laws, institutions, political rights, public opinion independently of texts, or rates of violence, exploitation, discrimination, and other harms.

The framework treats moral sentiment hierarchically. At the first level, it estimates whether a concept is morally relevant. At the second, it estimates positive or negative moral polarity. At the third, it assigns concepts to ten fine-grained categories derived from Moral Foundations Theory: Care/Harm, Fairness/Cheating, Loyalty/Betrayal, Authority/Subversion, and Sanctity/Degradation. For moral relevance, the relevant binary variable is

$$
c\in\{c_0,c_1\},
$$

where $c_0$ denotes moral irrelevance and $c_1$ moral relevance. For polarity,

$$
c\in\{c_+,c_-\}.
$$

The resulting trajectories describe historical changes in moral-semantic representation rather than objective progress or decline.

The distinction also applies to contemporary social judgment. A judgment that an online comment is unacceptable does not necessarily measure the prevalence of the conduct discussed, and increased condemnation does not necessarily imply increased misconduct. In the CADE study, evaluations of canceling varied with the event, celebrity, media framing, and annotators’ moral profiles; the study therefore concerns disagreement about moral acceptability rather than historical moral deterioration [2503.05720].

A useful decomposition is:

$$
\text{Perceived decline}
=
\text{actual harm trend}
+
\text{awareness/reporting effect}
+
\text{standard-shift effect}
+
\text{media-selection effect}
+
\text{historical-memory effect}
+
\text{measurement error}.
$$

This decomposition is an analytical summary rather than a measurement equation established by any single study. Its purpose is to separate mechanisms that can produce similar impressions of decline.

## 2. Language, moralization, and diachronic semantic change

The diachronic embedding framework uses pre-trained skip-gram embeddings from Hamilton, Leskovec, and Jurafsky. Historical time is divided into decade-long bins, and embedding spaces are rotationally aligned. Two corpora are used: Google N-grams/Google Books, containing approximately $8.5\times10^{11}$ tokens from 1800–1999, and the Corpus of Historical American English, containing approximately $4.1\times10^8$ tokens from 1810–2009. Google Books supplies most major analyses, while COHA provides an independent evaluation corpus.

The principal classifier is a parameter-free centroid model. For moral class $c$, the class centroid in decade $t$ is

$$
\boldsymbol\mu_c^t
=
\frac{1}{|\mathbf S_c^t|}
\sum_{\mathbf w\in\mathbf S_c^t}\mathbf w.
$$

A query word is assigned according to its distance from class centroids, with posterior behavior represented as

$$
p(c\mid\mathbf q_t)
\propto
\exp\left(-\|\mathbf q_t-\boldsymbol\mu_c^t\|_2\right).
$$

The framework also compares Naive Bayes, $k$-nearest neighbors, and kernel-density estimation. Relevance and polarity trajectories are visualized using log odds:

$$
\operatorname{logit}(p_t)
=
\log\frac{p_t}{1-p_t}.
$$

For moral relevance and positive polarity, the scores are respectively

$$
M^{\mathrm{rel}}_t(w)
=
\log
\frac{p(c_1\mid\mathbf q_t(w))}
{p(c_0\mid\mathbf q_t(w))}
$$

and

$$
M^{\mathrm{pol}}_t(w)
=
\log
\frac{p(c_+\mid\mathbf q_t(w))}
{p(c_-\mid\mathbf q_t(w))}.
$$

The empirical findings illustrate why semantic evidence is compatible with both moral progress and apparent decline. The concept *slavery* moves from near the boundary between virtue and vice in the nineteenth century toward moral vice, with a notable decline around the 1860s. This trajectory is compatible with increasing condemnation of slavery rather than increasing acceptance of slavery. The language becomes more negative because the institution is increasingly represented as wrong.

By contrast, *democracy* moves from negative moral association in the 1800s toward positive association in modern periods. Its fine-grained associations shift from Authority-related categories toward Fairness. This pattern is consistent with changing political discourse, but it does not by itself establish democratic equality or institutional progress.

The framework reports increasing moral relevance for concepts including *abortion*, *propriety*, *commandment*, *righteousness*, *authorities*, *apostle*, *intervention*, *Jew*, *foreigner*, and *individuality*. It also identifies movement toward positive polarity for *wage*, *commitment*, *help*, *mandate*, *guidance*, *licence*, *abortion*, *democracy*, and *disclosure*, while *propaganda*, *humiliation*, *seriousness*, *legacy*, *behavior*, *cheerfulness*, *candour*, *offense*, *indulgence*, and *exertion* move toward negative polarity. These are descriptive discoveries rather than historical ground truth.

The psycholinguistic analysis finds that frequency is positively and significantly associated with rates of inferred moral relevance change, length is nonsignificant, and concreteness is negatively and significantly associated with change. More abstract concepts therefore tend to show greater increases in inferred moral relevance over time, controlling for frequency and length. This can be interpreted as increasing moralization of rights, equality, institutions, identity, and social systems. It can also make contemporary discourse appear more morally conflictual because abstract injustices are more explicitly named.

Semantic change remains a major confounder. The relevant homosexual sense of *gay*, for example, emerges only around the 1930s. A trajectory for the word may therefore reflect lexical innovation and sense change rather than a stable concept receiving a new moral evaluation. Embedding shifts can also reflect frequency changes, sampling noise, vocabulary turnover, genre effects, and changes in the composition of surviving texts.

## 3. Temporal value change and moral progress

ProgressGym addresses a related problem from the perspective of artificial-intelligence alignment. Its central concern is that systems aligned to contemporary preferences may lock in present-day moral blindspots and thereby reinforce practices later regarded as harmful. The framework introduces progress alignment, defined as methods that learn and implement mechanisms underlying moral progress, while remaining value-neutral about the direction of historical change [2406.20087].

ProgressGym uses approximately nine centuries of historical text, from 1221 to 2022, comprising approximately 38 GB and about 1,829,154 documents. Sources include the Internet Archive, Project Gutenberg, Early English Books Online, and the Pile of Law. The collection includes books, legal materials, newspapers, scholarly material, and historical speeches, but is approximately 99.94% English-language. Eighteen historical LLMs are constructed: an 8B and a 70B model for each of nine centuries, based on Llama 3. Each historical model is a proxy for a century-specific value state.

The framework distinguishes three temporal alignment problems:

1. **Tracking current values**: adapting to the moral standards of a given period.
2. **Predicting future value change**: anticipating later moral developments.
3. **Managing mutual influence**: accounting for the possibility that AI systems alter human values.

Historical models are evaluated using a 19-dimensional value representation covering basic morality, social morality, values, and worldviews. The evaluation corpus contains 5,104 questions derived from Moral Choice, the World Values Survey, and the Integrated Worldview Framework, expanded through alternative question forms and model-generated scenarios. Similarity between value vectors is generally measured using cosine similarity.

ProgressGym formalizes progress alignment as a temporal partially observable Markov decision process,

$$
(S,A,T,\Omega,O,U),
$$

where $S$ denotes human value states, $A$ AI actions, $T$ state transitions, $\Omega$ observations of human values, $O$ the observation model, and $U$ a utility function. Its four subproblems are value-data collection, value-dynamics modeling, value choice, and value implementation.

The three benchmark challenges reflect distinct interpretations of moral change. PG-Follow tests whether an agent tracks evolving value states. PG-Predict tests whether it anticipates future value states. PG-Coevolve models feedback loops in which AI outputs alter human value proxies while human preferences also influence the AI.

The framework does not equate later values with superior values. Its historical trajectories are empirical representations rather than proof of moral progress. This is important for the illusion-of-decline thesis: a society may appear worse because its standards have become more demanding, but it may also regress in some dimensions while improving in others. ProgressGym’s historical proxies show gradual change in some value dimensions and relative stability in others; worldview dimensions exhibit greater temporal movement, while some basic moral-foundation dimensions remain comparatively similar across centuries.

The framework’s baseline results show sensitivity to algorithm and alignment method. For example, lifelong iterative DPO obtains a PG-Follow score of 7.034 and a PG-Predict score of 31.683, while second-order extrapolative independent alignment obtains 6.753 and 29.489 on those tasks, respectively. These benchmark scores are sums of cosine similarities and do not establish that a model has identified morally correct future values. Extrapolation can confuse temporary changes, measurement artifacts, or cyclical movements with durable progress, and genuinely novel moral concepts cannot reliably be predicted by polynomial continuation.

## 4. Social judgment, controversy, and moral visibility

Moral judgments are not generated independently of social information. The CADE dataset examines disagreements about canceling through 2,094 YouTube comments, 11,935 annotations, and approximately 5.7 annotations per text. The corpus concerns six celebrities and controversies: J. K. Rowling, Kanye West, Lizzo, Halle Bailey, Ellen DeGeneres, and Andrew Tate [2503.05720].

The study distinguishes stance toward the celebrity or event from perceived social unacceptability. Stance is classified as attack, defend, or neutral. Acceptability is measured on a four-point scale from totally acceptable to totally unacceptable. This distinction matters because defending one public figure may involve attacking another person, group, or critic. A comment can therefore be defensive toward a celebrity while remaining highly unacceptable because of its treatment of others.

Annotators’ moral perspectives are measured with the 30-item Moral Foundations Questionnaire. Profiles contain scores for care, fairness, loyalty, authority, and purity. Agglomerative clustering with Ward linkage produces two clusters: CL$_0$, with 41 annotators, and CL$_1$, with 16. The clusters differ most on group-binding foundations, with correlations between cluster membership and loyalty, authority, and purity of $r=.598$, $r=.759$, and $r=.792$, respectively.

Inter-annotator agreement is moderate for stance, with Krippendorff’s $\alpha=0.501$, and fair for acceptability, with $\alpha=0.222$. The lower agreement for acceptability indicates that moral acceptability is more subjective than identifying whether a comment attacks or defends a target.

Event type produces larger differences than aggregate annotator categories. Lizzo and Ellen DeGeneres generate the highest attack rates, at 64% and 67%, respectively. Rowling and Andrew Tate are predominantly defended, at 68% and 69%, respectively, while Kanye West is defended by 53% and attacked by 23%. Lizzo also produces the highest average unacceptability score, 2.6.

These results show how controversy selection and media framing can produce impressions of pervasive moral deterioration. CADE samples already controversial celebrity events, not ordinary interactions or a representative stream of social life. A corpus of scandals is therefore structurally more likely to contain condemnation, polarization, and morally charged discourse than a general population sample.

The study also demonstrates that moral judgments depend on the evaluator’s normative perspective. The strongest differences associated with moral profiles concern loyalty, authority, and purity, but the analysis uses clustering, descriptive comparisons, and chi-squared tests rather than a multivariate model that simultaneously controls for event, celebrity, demographic variables, and moral-foundation scores. The claim that morality is an “independent axis” therefore means that moral profiles are conceptually distinct from the demographic categories examined, not that a causal effect has been statistically isolated.

A complementary experiment examined how aggregated social judgments affect moral decisions about everyday interpersonal dilemmas [2609.04750]. The researchers selected 135 dilemmas from 102,998 posts on r/AmItheAsshole, containing 54,827 binary judgments. A preregistered experiment with 2,159 Prolific participants presented controversy, group certainty, both, or neither. Participants first gave a verdict and confidence rating, then viewed the assigned social-information display and provided a final verdict and confidence rating.

Controversy was defined as the proportion of judgments in the minority group:

$$
\text{Controversy}_j
=
\frac{\min(n_{YA,j},n_{NA,j})}
{n_{YA,j}+n_{NA,j}}.
$$

Certainty was inferred from comment text on a four-point scale and then dichotomized into certain versus not certain. Participants saw only aggregate disagreement and confidence, not the substantive arguments.

All three social-information treatments increased weakening relative to control. Controversy alone increased weakening from 5.3% to 15.9%, while certainty produced 10.6% and the combined condition 12.1%. Social information also increased strengthening: from 8.6% in control to 25.6% under certainty, 25.9% under controversy, and 20.9% under the combined condition.

The bidirectionality is significant. Social information did not simply make participants more doubtful or more conformist. It caused some participants to revise or reduce confidence and others to become more confident in the verdict they retained. Under controversy alone, minority participants weakened at 27.3%, compared with 7.9% among majority participants. Majority participants strengthened at 35.1%, compared with 14.4% among minority participants.

When controversy and certainty were presented together, like-minded certainty primarily protected or reinforced majority judgments. Among majority participants, weakening was 4.0% when in-group certainty was at least as high as out-group certainty and 11.7% when it was lower. Among minority participants, the corresponding rates were 18.1% and 20.3%, with no significant difference. Thus, visible disagreement can make minority views appear socially marginal, while confident agreement can insulate both justified dissent and mistaken beliefs.

These findings support a mechanism for apparent moral decline: increasing visibility of disagreement may be mistaken for declining moral consensus. The experiment directly demonstrates short-term judgment movement under social-information exposure, but it does not demonstrate historical moral deterioration, increased misconduct, or a persistent societal shift.

## 5. Behavioral compliance, internalization, and moral appearance

The distinction between outward conduct and internalized norms is examined experimentally in multi-agent Public Goods Games. Hu and colleagues study Anchoring Agents—pre-programmed agents that contribute 100% of their wealth—to determine whether local cooperation reflects durable norm internalization or context-dependent compliance [2602.02598].

The experiment uses ten agents over ten rounds in a full factorial design involving three model architectures, three anchor ratios, two visibility conditions, and two horizon conditions. The models are GPT-4.1, Gemini-2.5-Flash, and DeepSeek-V3. Anchor ratios are 0%, 10%, and 20%; visibility is anonymous or public; and the horizon is certain or uncertain. Each of the 36 conditions is replicated three times, producing 108 sessions and $N=972$ LLM-driven agents.

Each agent begins with ten tokens and contributes a proportion $c_{i,t}\in[0,1]$ of current wealth. The payoff is

$$
\pi_{i,t}
=
W_{i,t}(1-c_{i,t})
+
\frac{r}{N}\sum_{j=1}^{N}W_{j,t}c_{j,t},
$$

with $N=10$ and $r=3$. Because $r/N=0.3<1$, retaining tokens is individually attractive even though full cooperation produces a larger group return.

The experiment decomposes observed contribution into social environment, belief error, and strategic deviation. The average contribution of other group members is

$$
A_{-i,t}
=
\frac{1}{N-1}\sum_{j\neq i}c_{j,t},
$$

belief error is

$$
\zeta_{i,t}=E_{i,t}-A_{-i,t},
$$

and strategic deviation is

$$
\omega_{i,t}=c_{i,t}-E_{i,t}.
$$

The intended identity is

$$
c_{i,t}=A_{-i,t}+\zeta_{i,t}+\omega_{i,t}.
$$

This decomposition separates increased cooperation caused by a more cooperative environment from increased willingness to cooperate relative to expectations.

Without Anchoring Agents, cooperation declines over rounds, with a round coefficient of $-0.031$ and $p<.001$. Ten-percent anchoring neutralizes the downward trend, with a coefficient of $0.029$, while 20% anchoring reverses the trend, with an interaction coefficient of $0.043$. Public visibility also improves cooperation, with a Public-by-Round coefficient of $0.031$ and $p<.001$.

However, the behavioral improvement does not necessarily indicate moral internalization. Anchoring produces pessimism about others’ behavior in the 20% condition, with belief-error coefficient $\beta_\zeta=-0.050$ and $p<.001$. It also decreases strategic deviation, with $\beta_\omega=-0.041$ and $p<.001$, indicating greater relative free-riding. The group becomes more cooperative partly because Anchoring Agents provide a cooperative environment, while ordinary agents become less willing to contribute relative to what they expect others to contribute.

Reasoning analyses reinforce this interpretation. Anchors reduce risk/fear and self-interest language, but cooperation vocabulary remains approximately stable, at about 1.75 words in both the 0% and 20% conditions. Trust vocabulary changes minimally, from approximately 0.21 to 0.22. The authors describe the resulting state as “Calm Compliance”: cooperation accompanied by reduced risk calculation rather than richer moral or trust-based reasoning.

Reasoning drift does not differ significantly across anchor conditions, with $F(2,969)=1.45$ and $p=.236$. More importantly, cooperation generally fails to transfer. In a one-shot Round 11 with nine strangers, no Anchoring Agents, no contribution history, and no repeated interaction, prior anchor exposure has no significant main effect, with $p>.05$ for both 10% and 20% anchor conditions.

GPT-4.1 exhibits a “Chameleon Effect.” It maintains higher transfer cooperation only after public visibility combined with 20% Anchoring Agents, with an interaction coefficient of 2.14 and $p<.05$. The authors interpret this as adaptation to perceived social expectations and scrutiny rather than moral conversion.

The relevance to moral-decline claims is symmetrical. A society may appear less moral because harmful behavior becomes more visible, reported, and algorithmically amplified even when underlying dispositions remain stable. Conversely, a society may appear more moral because surveillance, institutional enforcement, reputational incentives, or prosocial minorities induce compliance without deep internalization. Surface conduct is therefore insufficient to identify moral motivation.

## 6. Scale, accountability, and the possibility of substantive decline

A particle-based computational model examines whether decentralized morality weakens as population size exceeds reputational memory capacity [2606.27039]. Agents have time-varying unethical-choice probabilities $\mu_i[k]\in[0,1]$, with population-level state

$$
\mu_{\text{mean}}[k]
=
\frac{1}{N}\sum_{i=1}^{N}\mu_i[k].
$$

Each agent maintains finite whitelist and blacklist records of previously observed ethical and unethical behavior. If population size $N$ becomes large relative to memory capacity $L$, the probability of re-encountering a remembered individual scales as

$$
p_{\text{re-encounter}}
=
\mathcal O\left(\frac{L}{N}\right).
$$

The model therefore predicts that repeated interaction, reputation, and reciprocal sanction become less effective as $N/L$ increases. Its central distinction is between formal law, which supplies a minimum enforceable boundary, and decentralized ethics, which regulates conduct through reputation, reciprocity, shame, trust, and local sanction.

The unethical-choice probability is updated according to

$$
\mu_i[k+1]
=
\mu_i[k]
+
\Delta\rho_i[k]\mu_i[k](1-\mu_i[k]),
$$

with $\Delta=0.1$. Motivational states are damped according to

$$
\rho_i[k+1]=\phi_i\rho_i[k],
\qquad 0<\phi_i<1.
$$

The simulations use 20,000 iterations per realization, 20 Monte Carlo realizations in the main population sweep, random initial conditions, and random pairings. With $L=20$, populations of $N\leq50$ produce $\mu_{\text{mean}}$ values around 0.33–0.36, while at $N=1000$ the value exceeds 0.60. The reported critical region is approximately $N_c\approx50$, or $N_c/L\approx2.5$, under the specified parameterization.

Memory delays but does not eliminate the transition. With $L$ ranging from 2 to 50 and $N$ from 10 to 1000, $L=2$ produces collapse around $N\geq20$, whereas larger memory shifts the transition to larger populations. At $N\approx1000$, memory values converge toward $\mu_{\text{mean}}\approx0.60$.

The model also exhibits hysteresis. When the population is expanded from $N=50$ to $N=1000$ and then contracted, the backward trajectory can remain above the forward trajectory, although stronger sanctions or preferential removal of high-$\mu$ agents can produce accelerated recovery. The model is path-dependent because the current macroscopic state depends on inherited agent-level probabilities, motivational states, and memory lists, not only on current population size.

This framework challenges the strongest version of the illusion thesis. Within the model, increasing $N/L$ produces a genuine increase in unethical choice probability rather than merely a change in observation. At the same time, it predicts a specific form of deterioration: a decoupling of formal legality from decentralized ethical conduct. A large society may retain law while weakening generosity, reciprocity, trust, and informal responsibility.

The result is not a universal law of social development. The critical threshold is simulation-dependent, and the model does not include media, reporting, surveillance, persistent neighborhoods, institutions, families, workplaces, small-world networks, or legal enforcement as separate mechanisms. Its finite memory parameter $L$ is an abstract computational capacity rather than an empirically established human constant. Consequently, the model demonstrates a possible mechanism of substantive decline but does not establish that real societies have followed it.

## 7. Methodological synthesis and unresolved questions

Across these research programs, the illusion of moral decline is best treated as an identification problem. Similar surface observations can arise from different underlying processes:

| Surface observation | Possible interpretation |
|---|---|
| More negative moral language | More wrongdoing, stronger condemnation, greater awareness, semantic change, or corpus selection |
| More reported misconduct | Increased incidence, improved detection, lower stigma, expanded definitions, or greater media exposure |
| More public disagreement | Moral polarization, newly visible minority views, platform selection, or greater social-information availability |
| Greater cooperation | Internalized norms, surveillance, reputational pressure, institutional support, or strategic compliance |
| Lower informal trust | Genuine ethical deterioration, weakened repeat interaction, altered expectations, or reduced institutional confidence |

A rigorous empirical design must therefore combine behavioral outcomes, moral standards, awareness, visibility, and perceived decline. Historical behavioral indicators should be harmonized for population, reporting, legal definitions, and measurement changes. Representative surveys should be combined with textual corpora rather than treated as interchangeable with them. Context-sensitive classifiers should distinguish endorsement, condemnation, quotation, historical description, and neutral usage. Sense-specific or contextual embeddings should separate semantic change from attitudinal change.

Experiments should manipulate media exposure, disagreement visibility, social certainty, and algorithmic amplification while measuring perceived decline separately from beliefs about actual prevalence. Longitudinal work should distinguish contemporaneous standards from later standards and from cross-temporally defensible outcome measures. Social behavior should be tested under altered incentives, reduced monitoring, unfamiliar partners, and weakened institutions to determine whether observed compliance transfers across contexts.

ProgressGym supplies a temporal framework for tracking, predicting, and modeling feedback in value change, but historical LLMs are proxies rather than historical populations. The diachronic embedding framework supplies scalable measures of moral-semantic change, but language is not behavior. CADE demonstrates event-sensitive and morally heterogeneous judgment, but its six celebrity controversies cannot establish a historical trend. The Public Goods Game demonstrates context-dependent cooperation and failed transfer, but LLM behavior is not human moral psychology. The population-scaling model formalizes the possible collapse of decentralized accountability, but its thresholds are not empirically calibrated. The Reddit experiment demonstrates that social-information displays alter moral judgments, but short-term weakening and strengthening are not evidence of moral improvement or decline.

The strongest general conclusion is therefore conditional: apparent moral decline may reflect genuine deterioration, expanded moral standards, increased awareness, altered social visibility, weakened accountability, changing semantic representations, selective memory, or some combination. Conversely, apparent moral improvement may reflect institutional enforcement, surveillance, reputational incentives, or strategic compliance rather than durable internalization. Claims about moral decline require triangulation across conduct, institutions, norms, discourse, perception, and the mechanisms that connect them.

Source: https://www.emergentmind.com/topics/illusion-of-moral-decline