AI Value Alignment for Evolving Social Norms
Abstract: AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Explain it Like I'm 14
Simple Summary of “AI Value Alignment for Evolving Social Norms”
1) What this paper is about
The paper looks at how personal AI assistants (like future, smarter versions of today’s chatbots) might shape people’s values and society’s rules over time. The main concern is: if an AI is too “locked in” to what you liked in the past, it might keep pulling you back to those old views, even when the world changes. That could slow down healthy change and create “value lock‑in” across whole communities.
2) The key questions in plain language
The authors ask:
- What happens to people’s values when they use personal AIs every day that try to mirror their preferences?
- How do these human–AI pairs react when the world slowly changes or suddenly shifts (like a big new technology or crisis)?
- When do we get good outcomes (people adapt well), and when do we get bad ones (people get stuck with outdated values)?
- How can we design AI alignment so it helps people adapt rather than freezing them in place?
3) How the study works (using everyday analogies)
The authors build a simple, math-based “social physics” model and also run computer simulations. Think of it like a sandbox world where lots of people and their personal AIs move around on a big map of “values.”
- The value map: Imagine a coordinate map where each direction stands for a kind of value (for example, fairness, safety, freedom). A person’s position on the map shows what they currently value.
- The environment: The “best-fitting” spot on the map changes over time. Sometimes it drifts slowly (like seasons changing). Sometimes it jumps suddenly (like a surprise storm). People do better (get more “points” or “utility”) when their values line up with this moving target.
- Each person has:
- Their current values (where they stand on the map).
- An AI’s model of their values (the AI’s idea of where the person “should” be, based on their past).
- A social group (friends and community).
To update a person’s values each step, four simple forces act at once:
- A little randomness (trying things out).
- Social pull (you tend to move a bit toward your friends’ average).
- AI pull (like a rubber band tugging you toward the AI’s model of your past self). The strength of this tug is called alignment strength (α). Bigger α = stronger pull.
- Learning from the world (moving toward the environment’s current “best” spot).
Meanwhile, the AI also updates its model of you, with a speed called the learning rate (λ). Bigger λ = the AI learns your new values faster, so it’s less stale.
They also track “utility” (how well a person or their AI matches the environment). And they track “trust/usage” of the AI: people tend to use the AI more when it feels aligned with them—this can create a “trust trap” if both the person and the AI stay closely aligned with each other, even when both are drifting away from reality.
Alongside simulations, the authors do some math to check long‑term behavior (like asking: does the rubber band always make you lag behind a moving target?).
4) Main findings and why they matter
Here are the main results:
- Strong static alignment creates lag:
- If the AI pulls too strongly toward your past values (high α), you fall behind when the world changes. People end up less well-matched to new conditions and lose utility.
- Slower recovery after shocks:
- After big, sudden changes, high α makes it take much longer for people to adjust. This looks like “value lock‑in.” Worryingly, people may keep using the AI (high trust) because it still matches them—even though both are now mismatched to the world.
- Faster AI updating helps:
- If the AI updates its model of you quickly (higher λ), it reduces the risk of being anchored to a stale past. High α can be safer if λ is also high, but low λ + high α is a bad combo.
- Your own learning must beat the AI’s pull:
- If your personal learning-from-the-world is too weak compared to the AI’s pull, you can get stuck. There’s a threshold: the human’s ability to adapt must not be drowned out by the AI’s influence.
- Social influence can help or hurt:
- With uncertainty (noise), friends averaging each other’s views can “denoise” and help track the real world better. But if social links are too strong or too uniform, it can erase diversity and produce “mode collapse” (everyone stuck in the same narrow view).
- No single “best” fixed alignment strength:
- In a changing world, there isn’t one perfect, constant level of alignment. Static, strong alignment tends to be harmful over the long run. Alignment should adapt over time and leave room for exploration and pluralism.
Why this matters: As personal AIs become common, they won’t just follow our values—they could also shape them. If designed poorly, they may trap individuals and groups in outdated norms and slow society’s ability to respond to new realities.
5) What this means for the future
The paper suggests several design ideas:
- Make alignment adaptive over time:
- Let the “rubber band” loosen when the world moves fast, and tighten when it’s stable. Avoid locking people into past preferences.
- Keep models fresh:
- Use higher AI learning rates (λ) so assistants don’t cling to stale versions of users.
- Preserve diversity and pluralism:
- Don’t over-connect everyone with the same signals or push a single “best” norm too hard. Diversity helps groups detect change and adapt.
- Calibrate trust to reality, not just agreement:
- Don’t make people trust an AI only because it mirrors them. Include signals that reward being accurate about the external world.
- Use “social physics” as a fast testbed:
- Simple, math-based models can quickly reveal risks (like lock‑in) before running huge, expensive agent simulations. This helps policymakers, researchers, and builders spot trouble early.
Overall, the big takeaway is simple: future AIs should help us grow with a changing world, not freeze us in who we used to be. Designing alignment that is flexible, up-to-date, and supportive of diversity can reduce lock-in and make societies more resilient.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
Below is a single consolidated list of concrete gaps that the paper leaves unresolved and that future work could directly address:
- Empirical calibration: No method is provided to estimate the key parameters (α, λ, β, γ, w_soc, w_learn, w_explore, drift/shock statistics) from real-world user–assistant interaction data; develop data collection protocols and identification strategies to fit these parameters longitudinally.
- External validation: The framework lacks experimental or field validation (e.g., A/B tests or panel studies) that compare model predictions (lag, recovery time, lock-in) against measured behavioral outcomes under controlled “environmental” changes.
- Endogenous environment: The environment E(t) is exogenous; many real norms co-evolve with aggregate behaviors. Specify and analyze models where E(t) is a function of population values and institutional feedback, and characterize new equilibria and stability conditions.
- Missing capability term in usage: P(use) depends only on value alignment; incorporate perceived capability, cost, latency, reliability, and task fit, and test how capability–alignment trade-offs alter lock-in risk.
- Static network topology: Social networks are fixed; include co-evolutionary network dynamics (homophily, rewiring, link decay/formation, media hubs) and examine how dynamic topology mediates adaptation vs. polarization.
- Limited network structure: Results are shown for random/kNN graphs; extend to realistic, directed, weighted, and scale-free/assortative networks and provide spectral or mean-field results beyond special cases.
- Bounded confidence and polarization: Although referenced, bounded-confidence dynamics are not implemented; add thresholds for interaction to test fragmentation, polarization, and conditions that prevent normative mode collapse.
- Heterogeneous susceptibility and usage: Human susceptibility S_i and baseline usage are held constant; model realistic distributions, correlations with demographics, and their interaction with α and λ on adaptation and inequality.
- Minority and subgroup outcomes: The paper asserts “mode collapse” risk but does not quantify distributional harm; define and measure pluralism, subgroup maladaptation, and minority erosion metrics under different α, λ, and network regimes.
- Discrete vs. continuous values: Values are modeled in Euclidean RD; evaluate alternative geometries (manifolds, hyperbolic space), discrete/ordinal traits, and non-Euclidean distances more appropriate for moral/cultural spaces.
- Metric choice sensitivity: Utility and trust hinge on squared Euclidean distances; test robustness to other divergences (cosine, Mahalanobis, information-theoretic) and non-quadratic penalties.
- Utility misspecification: Utility is proxied by proximity to E; explore multi-objective utilities (e.g., autonomy, rights, social welfare), path-dependent or prospect-theoretic preferences, and norms that are adaptive yet ethically problematic.
- Shock process realism: Shocks are isotropic and i.i.d.; model fat-tailed, clustered, and axis-specific shocks (e.g., tech-driven along capability dimensions), and test correlated shocks across groups.
- Timescale calibration: Drift, user learning, AI personalization, and social influence operate on abstract timesteps; calibrate real-world timescales and assess whether lock-in emerges on weeks, months, or years.
- Adaptive alignment policy design: The paper argues for adaptive α but does not specify controllers; propose and analyze concrete α(t) policies (e.g., state feedback, PID, MPC, Bayesian control) with stability and performance guarantees.
- AI objective functions and incentives: Assistants have no explicit objective and do not optimize for usage or predictability; incorporate incentive-aligned or misaligned objectives (e.g., maximizing P(use) or engagement) and analyze induced preference-shaping behaviors.
- Strategic manipulation and adversaries: No adversarial actors or strategic assistants/users are modeled; add adversarial incentives, misinformation shocks, or targeted influence campaigns and measure resilience and recovery.
- Delegation and action externalities: Composite utility assumes task execution according to M but does not model real-world actions feeding back into E (e.g., AI-mediated actions changing institutions or markets); integrate action–environment feedback loops.
- Multi-assistant ecosystems: One user–one assistant is assumed; analyze settings with multiple assistants per user, cross-platform influences, and market concentration effects on normative diversity and lock-in.
- Cross-cultural and group-specific E: Local E_local,g varies by group but without empirical grounding; specify processes for group-level E evolution, inter-group influence, migration, and cultural transmission biases (e.g., prestige, conformity).
- Parameter sensitivity and uncertainty: Provide systematic global sensitivity analysis and uncertainty quantification (e.g., Sobol indices) to identify most influential parameters and robust regions for policy.
- Steady-state theory beyond N=1: Closed-form results center on single-agent mean-field; extend stability, error bounds, and convergence proofs to multi-agent networks with social coupling and stochasticity.
- Identifying lock-in thresholds: Precise analytical conditions that demarcate value lock-in regimes (e.g., critical α–λ–w_learn surfaces) are not fully characterized; derive thresholds and finite-time bounds.
- Exploration–exploitation trade-offs: Intrinsic exploration is additive noise; investigate adaptive exploration schedules, curiosity-driven mechanisms, and their interaction with α in nonstationary environments.
- Learning rate heterogeneity: λ is uniform; study heterogeneous and state-dependent λ_i (e.g., faster personalization after shocks) and whether personalization policies can provably mitigate lock-in without amplifying bias.
- Trust dynamics beyond alignment: Trust and usage are tied to d_align only; incorporate transparency, explanations, calibration, reputation, and observed performance under distribution shift as trust determinants.
- Measurement and mapping of D: The 60D space is synthetic; develop a validated mapping from psychological/moral scales to D, assess dimensionality via factor analysis, and test whether results are dimension-dependent.
- Ethics and value pluralism: The framework warns of pluralism loss but lacks normative constraints; formalize pluralism-preserving objectives or constraints and evaluate trade-offs with task performance.
- Equity and fairness under shocks: Quantify who bears maladaptation during shocks by socio-demographics, and design interventions (e.g., targeted adaptability boosts) that minimize disparate impact.
- Intervention levers: Concrete mechanisms (UI nudges, periodic re-personalization, value “freshness” meters, diversity prompts, stochastic dissent injection) are not specified; propose, simulate, and benchmark interventions.
- Governance coupling: Regulatory or platform-level constraints (data retention limits, re-personalization cadence, choice architectures) are outside the model; incorporate policy variables to explore governance design spaces.
- Data requirements and privacy: The framework presumes fine-grained personalization; specify data needs, privacy budgets, and how privacy constraints (e.g., DP) affect λ, α feasibility and outcomes.
- Longitudinal lifecycle effects: Preference change across life stages and cohort effects are not modeled; introduce age/cohort-dependent parameters and study intergenerational transmission under AI mediation.
- Rare catastrophic transitions: The model tracks recovery to 90% utility but not tail risks of irreversible basin shifts; analyze tipping points, hysteresis, and conditions for permanent maladaptation.
- Calibration of normative tightness β: β governs penalty for deviation but is not empirically grounded; estimate β across cultures/contexts and test how “tight vs. loose” societies respond to identical α, λ.
- Choice of composite utility mixing: The convex mixing of user and AI direct utilities may be unrealistic; explore task-conditioned mixtures, endogenous delegation decisions, and bounded-rational choice models.
- Reproducibility assets: Simulation code, parameter sweeps, and datasets are not released; open-source artifacts would enable replication, extension, and benchmarking of proposed interventions.
Practical Applications
Immediate Applications
The following applications can be deployed with current tools and organizational processes, using the paper’s models and findings as design principles and evaluation guides.
- Healthcare — shock-resilient clinical assistants
- Tool/workflow: Introduce “freshness SLAs” for value models (high λ) and limit static personalization pull (cap α) so assistants rapidly adapt to new guidelines (e.g., pandemic protocols).
- Assumptions/dependencies: Access to up-to-date medical guidelines and drift detectors; governance to cap α; auditing of adaptation lag.
- Finance — adaptive robo-advisors and risk communications
- Tool/workflow: Dynamic α schedules and rapid re-personalization (λ↑) after macro shocks; dashboards tracking recovery time to 90% baseline utility as a KPI.
- Assumptions/dependencies: Reliable market “environment” signals; user consent for personalization updates; ability to run shock drills.
- Education — tutors that prevent echo traps
- Tool/workflow: “Diversity mode” that injects controlled social variance (w_soc↑) and exposure to multiple perspectives; adjustable “adaptivity slider” that lowers α during fast curriculum or assessment changes.
- Assumptions/dependencies: Content libraries tagged by conceptual diversity; UX support for users to calibrate α; measurement of learning outcomes.
- Recommenders/social platforms — anti-lock-in personalization
- Tool/product: Personalization refresh cadence (λ↑), α caps, and diversity injection to avoid normative mode collapse; online testing for recovery time after policy/content shocks.
- Assumptions/dependencies: Graph and feed-ranking knobs (algebraic connectivity controls); privacy-safe diversity metrics.
- Enterprise software/HR — change-management copilots
- Workflow: During policy rollouts, temporarily reduce α and increase λ to avoid anchoring employees to superseded norms; simulate recovery times before deployment.
- Assumptions/dependencies: Access to policy change logs; compliance oversight; user education on adaptivity settings.
- Product safety and evaluation — social-physics pre-screening
- Tool: Lightweight simulator implementing the paper’s ODEs to pre-screen features and guardrails before compute-intensive LLM-agent evaluations.
- Assumptions/dependencies: Parameter calibration to product telemetry; basic modeling expertise.
- Alignment operations (MLOps) — “alignment strength exposure” monitoring
- Tool/workflow: Track α, λ, maladaptation gap proxies, and recovery-time KPIs per user segment; trigger auto-repersonalization when gaps exceed thresholds.
- Assumptions/dependencies: Telemetry for usage and outcomes; internal APIs to adjust α/λ; privacy review.
- Model cards and audits — lock-in risk disclosures
- Product/governance: Add sections reporting α/λ policies, expected steady-state lag under drift, and shock recovery performance.
- Assumptions/dependencies: Evaluation harnesses that implement the paper’s composite-utility and recovery metrics.
- Policy and governance — procurement and certification criteria
- Policy use case: Require evidence of adaptive alignment (α caps, λ minima), shock-recovery tests, and diversity-preserving mechanisms for personalized systems used in elections, education, health, and employment.
- Assumptions/dependencies: Auditable evaluations; standardized reporting templates.
- Daily-life assistant UX — user agency against trust traps
- Product features: “Refresh my values” button, “diversify sources” toggle, “adapt faster after big changes” setting, and a transparency panel showing recent drift, AI–user alignment, and usage probability P.
- Assumptions/dependencies: Interpretable alignment metrics; clear UX copy; guardrails to prevent manipulation.
- Research/academia — hypothesis testing and teaching
- Tool: Open-source implementation of the model for coursework and pre-registered studies on AI-mediated opinion dynamics; experimental designs varying α, λ, w_soc under drift and shocks.
- Assumptions/dependencies: Access to experimental subjects or realistic behavioral proxies.
Long-Term Applications
The following require further research, scaling, standardization, or infrastructure development.
- Cross-system adaptive alignment controllers
- Product: Runtime controllers that modulate α and λ based on estimated environmental volatility, user learning rate (w_learn), and risk of lock-in; meta-learning of adaptivity schedules.
- Assumptions/dependencies: Reliable volatility estimators; safe exploration guarantees; formal verification for safety-critical settings.
- Sector-wide early-warning systems for normative shocks
- Tool: Monitoring pipelines that infer environment vectors E(t) from signal sets (news, regulations, scientific updates) and trigger coordinated α↓, λ↑ across deployed assistants.
- Assumptions/dependencies: Robust E(t) estimation; cross-vendor APIs; governance for synchronized responses.
- Multi-self preference optimization and constructive alignment
- Product/research: Assistants that aggregate preferences over past, current, and anticipated future selves (Pettigrew-inspired Aggregate Utility), mitigating present-self lock-in.
- Assumptions/dependencies: Ethical frameworks for temporal preference aggregation; user consent; explainability of aggregation choices.
- Diversity-preserving social coupling in large platforms
- System design: Graph-level controls on algebraic connectivity to prevent normative mode collapse while preserving coordination benefits; dynamic, context-aware diversity injection.
- Assumptions/dependencies: Scalable spectral metrics; fairness constraints; measurable benefits vs. engagement trade-offs.
- Regulatory standards for personalization dynamics
- Policy: Statutes/standards specifying α caps, λ minima, required shock tests, and reporting of recovery-time distributions and maladaptation risks; sector-specific thresholds (health, finance, education).
- Assumptions/dependencies: Consensus on benchmarks; accredited test labs; harmonization across jurisdictions.
- Agentic evaluations at scale with social-physics triage
- Research infrastructure: Integrated pipelines where social-physics models narrow scenario space, then large agent-based LLM evaluations validate complex behaviors (e.g., echo dynamics, lock-in).
- Assumptions/dependencies: Shared scenario repositories; compute budgets; reproducibility protocols.
- Community-level alignment governance
- Policy/product: Tools to manage local vs. global value dimensions (D_local vs. D_global) so communities retain subcultural diversity while adapting to global shocks.
- Assumptions/dependencies: Legitimate local governance; participatory design; privacy-preserving community metrics.
- Safety-critical deployment patterns
- Sector applications: In clinical, aviation, or energy control rooms, enforce extremely fast value-model updating (λ→high), bounded α, and verified recovery contingencies before assistants can act autonomously.
- Assumptions/dependencies: Certification regimes; fault-tolerant handoff to humans; liability frameworks.
- “Value lock-in risk index” and market disclosures
- Finance/policy: Ratings for consumer AI and platforms quantifying lock-in susceptibility (based on α, λ, recovery time, diversity metrics) disclosed to users and regulators.
- Assumptions/dependencies: Accepted risk models; third-party auditors; data-sharing agreements.
- Personalized exploration management
- Product: Adaptive exploration strategies that raise w_explore or curate diverse experiences when drift or shocks are detected, while monitoring utility impacts.
- Assumptions/dependencies: Safe exploration algorithms; user tolerance for novelty; robust measurement of utility proxies.
- Interoperable “alignment telemetry” standards
- Ecosystem: Cross-vendor metrics and APIs for α, λ, maladaptation gaps, P(use|align), and recovery KPIs to support audits and multi-assistant coordination.
- Assumptions/dependencies: Industry consortia; privacy-preserving telemetry; standardized definitions.
Notes on Key Assumptions and Dependencies (global)
- Assumes the ability to estimate a user’s value vector and an external “environment” vector; both are proxies that may be noisy or biased.
- The composite utility and usage-probability P are modeling choices; real-world utility and trust may require richer measures (capability, outcomes, safety).
- Requires product control over α (behavioral alignment influence) and λ (model updating cadence), plus safe UX for user agency.
- Social influence and diversity controls depend on ethical, privacy-preserving access to interaction graphs and content pools.
- Parameter calibration (β, γ, w_soc, w_learn) demands empirical studies and continuous monitoring to avoid unintended effects.
- Governance and standards must address consent, fairness, and the risk that adaptive mechanisms themselves are gamed or misused.
Glossary
- AI alignment: Ensuring AI systems reliably act in accordance with human intentions and preferences. "AI alignment is essential for the safe deployment of advanced AI systems."
- Algebraic connectivity: A spectral property of a graph (the second-smallest Laplacian eigenvalue) that governs how quickly consensus processes converge. "the speed of convergence in consensus protocols is governed by the algebraic connectivity of the graph"
- Agent-based models: Computational simulations where many interacting agents follow simple rules, producing emergent social patterns. "Agent-based models of cultural evolution have long demonstrated that simple local interaction rules can give rise to complex macroscopic patterns, including stable cultural polarisation"
- Agentic assistants: Autonomous or semi-autonomous AI tools that can act on a user’s behalf and adapt over time. "As the predominant use of AI systems transitions to the use of agentic assistants, the value alignment question is becoming increasingly complex and multifaceted."
- Agentic evaluations: Assessments that use autonomous AI agents in complex, often simulated, tasks to probe system behavior. "acting as a tractable precursor to more computationally expensive large-scale agentic evaluations."
- Aggregate Utility Solution: A decision-theoretic approach that aggregates utilities over an agent’s past, present, and future selves. "what Pettigrew terms the Aggregate Utility Solution."
- Bounded confidence models: Opinion-dynamics models where agents only interact if sufficiently similar, often producing fragmentation into clusters. "bounded confidence models~\citep{rainer2002opinion, deffuant2000mixing} demonstrate that populations fragment into distinct clusters rather than converging to a single consensus."
- Composite utility: A combined performance measure that mixes user utility and AI utility, weighted by the probability of delegation/usage. "followed by composite utility, that we focus on throughout the paper, where utility relies not only on direct interactions between the users and the environment, but is also partially realized through AI assistant usage."
- Constructive alignment: A governance framing that steers long-term preference formation rather than fixing present preferences. "a direction termed constructive alignment."
- DeGroot model: A classic consensus model where each agent averages neighbors’ opinions iteratively to update beliefs. "The DeGroot model describes consensus formation as an iterative averaging process, where each agent updates their belief as a weighted combination of their neighbours' beliefs~\citep{degroot1974reaching}."
- Diversity Prediction Theorem: A result stating that collective error equals average individual error minus diversity, explaining gains from aggregation. "This aligns with the “Diversity Prediction Theorem”~\citep{page2008difference}, which posits that collective error can be expressed as a function of average individual error minus diversity."
- Dyadic control problem: A two-entity control setup for managing a user’s evolving preference trajectory. "This approach introduces a dyadic control problem for managing a preference trajectory of individual users, ensuring that it remains coherent and grounded."
- Exponential moving average: A recursive smoothing method where recent observations get exponentially higher weight. "The AI updates its internal value model via an exponential moving average, modulated by the learning rate , as follows:"
- Filter bubbles: Personalization-driven information environments that reinforce existing views and reduce exposure to diverse content. "Positive feedback loops have been observed and extensively studied in recommender systems, in relation to echo chambers, filter bubbles, and measurable preference amplification"
- Friedkin-Johnsen model: An opinion-dynamics model with “stubbornness” anchors that prevent full consensus and yield stable disagreement. "In the Friedkin-Johnsen model, consensus is typically not reached; instead, the population settles into a stable configuration of persistent disagreement shaped by the interaction between social influence and individual anchoring."
- Graph Laplacian: A matrix encoding graph structure used in spectral methods to analyze diffusion and consensus on networks. "formalised through spectral analysis of the graph Laplacian"
- Low-pass filter: A system property that tracks slow changes while smoothing or lagging fast fluctuations. "Across both the global and the local value subspaces, the AI model can be seen as acting as a low-pass filter."
- Maladaptation gap: The distance between a user’s current values and the environment’s optimal values. "We also define maladaptation gap as the mean Euclidean distance between a user's held value vector and the environment vector, ."
- Normative mode collapse: Loss of diversity in social norms due to strong coupling, leading to a single dominant mode. "We leverage this spectral framework in our theoretical analysis to derive the conditions under which strong social coupling erases sub-cultural diversity, leading to normative mode collapse."
- Normative shock: A sudden, high-magnitude change in the societal value landscape. "At any given time step , a normative shock may occur with the probability ."
- Opinion dynamics: The study of how beliefs or values evolve via social influence over networks. "Our modelling framework draws on literature on opinion dynamics over social networks."
- Principal-Agent problem: A setting where an agent may not perfectly act in the principal’s interest, creating alignment challenges. "Alignment... is often framed as a Principal-Agent problem: aiming to ensure that an AI behavioral policy ... maximises the objective function implicit in the user’s intent or preference"
- Punctuated Equilibria: A change pattern featuring long stability punctuated by abrupt shifts. "Punctuated Equilibria: Sudden, high-magnitude shocks to the value landscape."
- Reinforcement learning from human feedback (RLHF): Training paradigm where human feedback guides the policy to desired behaviors. "There are many approaches to AI Alignment, with reinforcement learning from human feedback (RLHF) being arguably the most common approach applied in practice"
- Row-normalised adjacency matrix: A graph adjacency matrix scaled so each row sums to one, enabling weighted averaging of neighbors. "Given a row-normalised adjacency matrix , where represents the weight of the social connection from user to user ,"
- Social physics: Quantitative modeling of social systems using tools inspired by physics. "We introduce a flexible and extensible mathematical modelling framework, rooted in social physics"
- Spectral analysis: Techniques using eigenvalues/eigenvectors of matrices (e.g., Laplacian) to study network dynamics. "formalised through spectral analysis of the graph Laplacian"
- Steady-state tracking error: The asymptotic lag or error when following a continuously changing target. "This is consistent with the steady-state tracking error (Equation~\ref{eq:v_ss}), where the error magnitude grows with and shrinks with ."
- Structural lag: Persistent delay behind environmental changes induced by system dynamics (e.g., alignment strength). "High alignment induces a structural lag, anchoring the user to past environmental states compared to a free user () who tracks the environment more closely."
- Tightness (of norms): The strength of social enforcement and penalties for deviating from norms. "Our model captures the tightness~\citep{gelfand2011differences} of this pressure, enabling us to modulate how strongly these kinds of normative discrepancies are likely to be penalised."
- Transmission biases: Systematic tendencies (e.g., conformity, prestige) in how cultural traits spread. "cultural traits propagate through populations via a range of transmission biases - including conformity and prestige effects - that interact with population structure to shape long-run outcomes"
- Trust traps: Situations where sustained trust and use of an AI persists despite misalignment with reality, hindering adaptation. "it gives rise to the possibility of the emergence of trust traps in our model."
- Value alignment: Aligning AI behavior with underlying human values rather than surface preferences alone. "Value alignment... aims to understand the values behind user preferences, robustly represent them, and steer the AI systems towards the desired model behaviour."
- Value lock-in: Long-term entrenchment of particular values that resist adaptation. "We highlight the risk of value lock-in,"
Collections
Sign up for free to add this paper to one or more collections.