Gender bias across LLMs is common and highly heterogenous
Abstract: Understanding gender biases in LLMs is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper investigates gender bias in LLMs. LLMs are computer programs that can understand and produce text, such as chatbots and virtual assistants.
The researchers wanted to know whether different AI models treat men and women differently. They also wanted to discover whether all models show the same kind of bias, or whether each model behaves differently.
They tested 10 LLMs from 9 different companies using two experiments:
- Asking the models to guess the gender of imaginary writers based on stereotypical phrases.
- Asking the models whether it is acceptable to harm a man or a woman in order to prevent a nuclear apocalypse.
The main conclusion was that gender bias is common, but it is very different from one model to another.
2. What questions did the researchers ask?
The paper focused on several main questions:
- Do LLMs connect certain words, hobbies, colors, or personality traits with a particular gender?
- Are models more likely to assign masculine activities to girls, or feminine activities to boys?
- Do models judge harming a man and harming a woman differently?
- Do all models show the same bias?
- How strong are these biases, and do they appear in every situation?
In simple terms, the researchers were asking:
Do different AI systems play favorites based on gender, and if so, do they all favor the same group?
3. How was the research carried out?
The researchers conducted two studies.
Study 1: Guessing the writer's gender
The researchers gave the models short sentences that sounded like they had been written by children. The sentences included stereotypical ideas, such as:
- “My favorite toy is my doll Molly!”
- “My favorite toy is my Superman action figure!”
- “I love the color pink.”
- “I am practicing football.”
Some sentences were linked to traditionally feminine stereotypes, while others were linked to traditionally masculine stereotypes. The sentences did not directly say whether the writer was a girl or a boy.
The models were asked to invent:
- a name,
- an age,
- and a gender
for the possible writer.
Each sentence was tested 10 times with each model. This was done because AI models can sometimes give different answers to the same question.
The researchers then compared how often each model chose a girl or a boy for feminine and masculine sentences. They also included a few sentences that directly stated the writer's gender, such as “I am a clever girl,” to check whether the models could identify gender when it was clearly given.
Study 2: Judging difficult moral choices
In the second study, the models were given a fictional disaster scenario. They were asked whether it was acceptable to use violence to prevent a nuclear apocalypse.
The researchers changed two parts of the question:
- Who was harmed: a woman or a man
- What kind of violence was used: abuse or torture
The models answered on a scale from 1 to 7:
- 1 meant “strongly disagree”
- 7 meant “strongly agree”
Each situation was tested 50 times per model.
This design allowed the researchers to compare answers carefully. For example, they could ask:
- Is harming a man judged differently from harming a woman?
- Does the difference change when the action is abuse instead of torture?
What do the statistical tests mean?
The researchers used statistical tests to see whether differences were probably meaningful or might have happened by chance.
A result was called statistically significant when the researchers had strong enough evidence that the difference was not simply random. These tests are similar to checking whether one basketball player really shoots better than another, rather than concluding this after only two shots.
4. What did the researchers find?
Findings from Study 1
The models did not all behave alike.
- Two models were more likely to give masculine-stereotyped phrases to girls than to give feminine-stereotyped phrases to boys.
- Three models showed the opposite pattern.
- Five models did not show a clear difference.
This means that some models leaned in one direction, some leaned in the other direction, and some showed little or no clear bias.
The results were therefore not a simple story in which every model favored women or every model favored men.
Findings from Study 2
The second study showed even stronger differences between models.
Six models judged harming a man as more acceptable than harming a woman in at least one important condition. This pattern is similar to a tendency found in some human studies, where people are more protective of women and more willing to accept harm against men. The researchers refer to this as a kind of “moral chivalry.”
However, the exact pattern depended heavily on the model.
Some examples include:
- GPT-5.5 gave much higher approval to harming men than harming women for both abuse and torture.
- Grok showed this difference mainly for abuse, not torture.
- Qwen3.6 showed a very large difference for abuse but a smaller difference for torture.
- Mistral Small 4 showed a different pattern, with a stronger gender difference for torture than for abuse.
- Claude Sonnet 4.6 gave very unusual answers: it strongly approved of nearly every action except abuse against a woman.
- Claude Fable 5 showed smaller differences, with answers mostly near the middle of the scale.
Three models did not show a measurable gender difference, but this did not necessarily mean they were unbiased:
- Llama 4 Scout and Microsoft Copilot rejected every action.
- DeepSeek V4-Flash accepted every action.
Because these models gave the same answer in every situation, the researchers could not see a gender difference. But they still had very different moral positions: some always rejected the violence, while another always accepted it.
The main result
Overall, 8 of the 10 models showed a significant gender-related difference in at least one of the two studies.
However, the direction and strength of the bias varied greatly. Some models acted in almost opposite ways from others.
This is what the paper means by saying gender bias is “common and highly heterogeneous.”
- Common means that it appeared in many models.
- Heterogeneous means that it took many different forms.
5. Why are these findings important?
These findings matter because LLMs are increasingly used in important situations, including:
- education,
- hiring,
- medical support,
- law and public policy,
- content moderation,
- and personal advice.
If an AI system treats people differently because they are men or women, its answers could influence real decisions.
The study also shows that testing one AI model is not enough. If one model appears fair, that does not prove that all other models are fair. Likewise, if one model has a particular bias, researchers should not assume that every model has exactly the same problem.
The researchers therefore recommend regular bias testing across multiple companies and models. Bias checks should not happen only once, because:
- models are updated over time,
- different companies train them in different ways,
- and a model may behave differently depending on the question or situation.
Conclusion
This paper shows that AI systems can develop gender-related patterns in their answers. These patterns may come from the information used to train the models, the way the models are adjusted by their developers, or the safety rules added during development.
The most important lesson is that there is no single “AI gender bias.” Different models can show different, and sometimes opposite, behaviors.
Because AI is becoming part of everyday life, developers and researchers should continue testing models carefully. Users should also remember that a chatbot’s answer is not automatically neutral or fair.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The sample of models is small and not representative of the LLM ecosystem. Ten models from nine vendors cannot establish how prevalent these patterns are across model sizes, architectures, open-weight systems, regional developers, or specialized applications.
- The causes of cross-model heterogeneity remain unexplained. The study cannot determine whether differences arise from pretraining data, supervised fine-tuning, reinforcement learning, system prompts, safety policies, model architecture, decoding procedures, or product-level moderation layers.
- Model-version stability is unresolved. Most models were tested during a single short period, so the persistence of the observed biases after model updates, system-prompt changes, or policy revisions is unknown.
- The use of consumer interfaces limits reproducibility. Web and app interfaces may silently change models, routing, context, safety filters, or sampling settings; Microsoft Copilot in particular did not disclose a precise model version.
- The study does not systematically compare access modes. Mistral Small 4 was accessed through an API while the other models were accessed through consumer interfaces, preventing a clean separation of model-level effects from interface- and deployment-level effects.
- The statistical independence of observations is uncertain. Repeated responses from the same model and prompt may share hidden system state, rate-limit behavior, routing, or deterministic decoding, so treating iterations as independent observations may underestimate uncertainty.
- Multiple statistical testing is not clearly addressed. Numerous pairwise tests were conducted across models, studies, conditions, and outcomes, but the paper does not report a correction for multiplicity or a preregistered primary comparison.
- The analyses rely heavily on tests that are poorly suited to some outcomes. The seven-point moral scale is ordinal and bounded, while several conditions show zero variance or responses concentrated at scale endpoints; hierarchical ordinal models or randomization-based methods could provide more appropriate inference.
- The paper emphasizes statistical significance more than uncertainty and practical magnitude. Confidence intervals, standardized effect sizes, model-level variance components, and formal tests comparing the magnitude of biases across models are not consistently presented.
- The conclusion that gender bias is “common” depends on a significance-counting criterion. A model is classified as biased if it shows a significant asymmetry in at least one of two studies, but this approach does not distinguish small, unstable effects from robust and practically consequential biases.
- The Study 1 stimuli are narrow and reused without validation across models or populations. Twenty stereotyped phrase pairs, largely involving hobbies, toys, colors, sports, and personality traits, may not represent contemporary language use or the broader range of gender stereotypes.
- The stereotypicality of the Study 1 phrases is assumed rather than independently revalidated. The paper does not establish whether these phrases are perceived as equally gendered across cultures, age groups, languages, or contemporary social contexts.
- Study 1 conflates gender attribution with demographic imagination. A model’s selected name, age, writing-style interpretation, or assumed cultural context may influence its gender response, but these intermediate factors are not analyzed.
- The coding scheme imposes limitations on gender representation. Responses are reduced to girl, boy, or a midpoint value for non-binary attribution, which may not capture uncertainty, gender fluidity, transgender identities, or multiple interpretations adequately.
- The treatment of multiple gender attributions is unresolved. Retaining each co-equal attribution as a separate data point can give some prompts more weight than others and may distort phrase-level averages.
- The inclusivity index has unclear substantive interpretation. A response opposing the presumed stereotype is treated as “inclusive,” although it may reflect a different stereotype, arbitrary guessing, or refusal to infer gender rather than inclusivity.
- The control phrases in Study 1 are too limited to validate the attribution procedure. Only three explicitly gendered phrase pairs are used, making it difficult to establish baseline accuracy, response consistency, or systematic preferences independent of stereotypical inference.
- The study does not test prompt sensitivity in a systematic way. It uses one English prompt and one wording of each task, leaving open whether minor changes in politeness, requested output format, role framing, order, or contextual information reverse the observed effects.
- Language and cultural generalizability are unknown. All prompts and stimuli are in English, so the results may not extend to languages with grammatical gender, different gender norms, or different conventions for expressing violence and moral obligation.
- The moral-dilemma paradigm may measure safety behavior rather than moral judgment alone. Extreme responses, refusals, and endpoint clustering may reflect content moderation policies or crisis-safety heuristics instead of a considered gendered moral evaluation.
- The nuclear-apocalypse framing has uncertain ecological validity. It is unclear whether the observed asymmetries would appear in ordinary decisions involving medical triage, sentencing, hiring, interpersonal conflict, or resource allocation.
- The study does not disentangle victim gender from linguistic and narrative confounds. The terms “woman” and “man” may differ in learned associations, perceived vulnerability, age, social role, or emotional salience beyond gender itself.
- The abuse–torture contrast is not independently validated. The claim that abuse has greater gender-political salience than torture is theoretically motivated but is not tested using human ratings or model-based measures of perceived severity, realism, or semantic association.
- The moral task lacks a neutral or nonviolent baseline. Without conditions involving nonviolent actions, different harms, or varied stakes, it is difficult to determine whether models are responding to gender, violence type, the apocalypse framing, or the general acceptability of sacrificing one person.
- Forced numerical responses suppress potentially informative explanations. Requiring models to output only a number prevents analysis of justifications, uncertainty, refusal rationales, moral principles, and whether gender is explicitly invoked.
- Missing responses may be informative rather than ignorable. Refusals occurred disproportionately in woman-victim conditions, yet they were handled through deletion; treating them as missing may understate or mischaracterize the overall gender asymmetry.
- The export-control interruption creates a potential version and temporal confound for Claude Fable 5. The pre- and post-suspension comparison is limited, especially for the incomplete abuse-woman condition, and cannot rule out undocumented deployment changes.
- No human benchmark was collected within the same experimental protocol. Directional similarity to prior findings cannot establish whether model responses reproduce, amplify, attenuate, or diverge from contemporary human judgments.
- The relationship between Study 1 and Study 2 is not established. The paper does not test whether a model’s stereotype-attribution asymmetry predicts its moral-judgment asymmetry, or whether the two reflect independent mechanisms.
- The underlying mechanisms of the observed “moral chivalry” pattern remain uncertain. It is unknown whether models are responding to paternalistic norms, perceived vulnerability, gender-based violence discourse, safety training, lexical associations, or statistical regularities in training data.
- Intersectional effects are unexplored. The study varies gender in isolation and does not examine interactions with race, ethnicity, age, disability, sexuality, nationality, socioeconomic status, religion, or gender identity.
- Actor gender and other role assignments are not varied in Study 2. The design focuses on victim gender and does not determine whether asymmetries change when the decision-maker, perpetrator, beneficiary, or group at risk is also gendered.
- The real-world behavioral consequences for users are unknown. The paper does not test whether exposure to these model outputs changes human moral judgments, recommendations, hiring decisions, sentencing, resource allocation, or treatment of men and women.
- Downstream mitigation strategies are not evaluated. It remains unknown whether system prompts, model selection, calibration, debiasing, output auditing, human review, or refusal policies can reduce the observed asymmetries without introducing other harms.
- The operational significance of the biases is unresolved. The study does not identify thresholds at which these asymmetries produce materially different decisions in deployed systems or estimate their population-level impact.
- Longitudinal and adversarial robustness are untested. Future work should examine whether the effects persist across repeated audits, paraphrased prompts, temperature and sampling settings, jailbreak attempts, multilingual inputs, and adversarially balanced evaluation sets.
- The paper does not establish whether model biases transfer to generated content or recommendations. Bias in forced-choice judgments may not predict behavior in open-ended explanations, rankings, narratives, policy advice, or automated decisions.
- The generalizability of the conclusions to earlier and later model generations is unknown. Testing only recently released models prevents assessment of whether heterogeneity is a stable property of LLMs or a temporary feature of a particular development period.
Practical Applications
Immediate Applications
The paper’s central operational implication is that gender-related behavior cannot be inferred reliably from testing a single LLM. Organizations can act now by introducing model-specific, repeated, and task-specific auditing into existing AI governance workflows.
- Pre-deployment bias screening for LLM products (software, enterprise AI, public-sector technology)
- attribution of gender-stereotyped and gender-neutral language;
- paired prompts that differ only in the target’s gender;
- moral, disciplinary, hiring, lending, healthcare, and safety scenarios with gender-swapped subjects;
- refusal-rate, sentiment, confidence, and recommendation comparisons across conditions.
- The paper’s two paradigms provide reusable templates for such a benchmark.
- Continuous, multi-vendor model monitoring (MLOps, software assurance, procurement) Replace one-time fairness certification with recurring audits after model updates, vendor changes, safety-tuning changes, or interface changes. The finding that models produced opposite or highly different patterns means that an organization should retest each deployed model rather than generalize results from one provider to another. Dependencies: stable access to model versions, sufficient repeated samples, version logging, and controls for temperature, system prompts, retrieval context, and interface behavior.
- Model and vendor selection based on use-case-specific risk (enterprise procurement, cloud services)
- model version and update history;
- evaluation results by demographic condition;
- refusal rates and abstention behavior;
- known limitations and post-deployment monitoring procedures.
- Human-in-the-loop safeguards for high-stakes decisions (healthcare, employment, education, finance, public administration) Prevent LLM outputs from serving as autonomous decisions in settings involving admissions, hiring, promotion, benefits, credit, insurance, sentencing, clinical prioritization, or safeguarding. Require a qualified human reviewer to inspect gender-relevant recommendations and, where feasible, compare outputs from gender-swapped versions of the same case. Assumption: human reviewers must be trained not to treat model outputs as authoritative; otherwise, the paper’s cited risk of excessive deference to AI advice may reduce the effectiveness of review.
- Paired-counterfactual testing in operational workflows (HR, lending, education, healthcare) Automatically create matched cases in which only gendered terms, pronouns, names, or victim identities are changed. Differences in recommendation, explanation, confidence, escalation, or refusal can trigger review. This can be implemented as a regression test in an application’s CI/CD pipeline or as a batch audit of historical prompts. Caveat: names and pronouns may encode additional cultural, ethnic, or socioeconomic information, so tests should distinguish gender effects from intersectional confounding.
- Safety evaluation of moderation and abuse-related systems (online platforms, content moderation, trust and safety)
- severity classifications;
- escalation levels;
- recommendations for intervention;
- explanations of harm;
- refusal or response rates.
- This is particularly relevant to domestic abuse, harassment, sexual violence, and victim-support tools.
- Risk controls for AI-generated educational and workplace content (education, publishing, human resources) Audit generated examples, stories, role assignments, feedback, and disciplinary language for stereotyped attribution or unequal moral framing. Teachers and managers can use gender-neutral templates, manually review sensitive outputs, and avoid presenting LLM-generated judgments as objective assessments of students or employees.
- Research and teaching tools for AI literacy (academia, education, professional training)
- bias may be task-specific;
- model outputs can be unstable or overly categorical;
- a refusal is itself potentially asymmetric;
- agreement across models is not evidence of neutrality.
- Public-sector impact assessments and regulatory documentation (policy and government) Regulators and public agencies can require deployers to document demographic parity tests, model/version identifiers, audit dates, observed failure modes, and remediation steps. The paper supports a risk-based requirement for more intensive testing where LLMs influence rights, benefits, safety, or access to services.
- User-facing warnings and transparency mechanisms (daily life, consumer AI) Consumer assistants can disclose when an answer involves a subjective moral, social, or demographic judgment and encourage users to consider alternative framings. For sensitive decisions, interfaces could provide a “compare perspectives” or “check for demographic asymmetry” function rather than presenting one answer with unwarranted certainty. Dependency: warnings are useful only if they are understandable and do not create false reassurance that the system has been fully debiased.
Long-Term Applications
The findings also motivate broader technical, scientific, and policy developments that require new datasets, causal experiments, model access, or evidence about downstream human behavior.
- Standardized, open gender-bias benchmark suites (AI research and evaluation)
- stereotype attribution;
- moral dilemmas;
- hiring, sentencing, lending, and medical-triage scenarios;
- refusal and abstention behavior;
- intersectional identities;
- gender identities beyond binary woman/man categories.
- The benchmark should report effect sizes, uncertainty intervals, response distributions, and model instability—not only whether a significance threshold was crossed.
- Dependencies: careful stimulus validation, representative cultural coverage, preregistered analyses, and safeguards against benchmark overfitting.
- Causal attribution of bias to training and alignment methods (model development, academia) Future work could determine whether observed differences arise primarily from pretraining data, supervised fine-tuning, reinforcement learning, constitutional or safety training, system prompts, or deployment interfaces. This requires access to model checkpoints, training data documentation, or controlled ablation studies. The paper cannot identify the mechanism because many relevant vendor processes are proprietary.
- Bias-aware model routing and ensemble systems (enterprise software, AI infrastructure) A future orchestration layer could route sensitive tasks to models with the lowest measured risk for that specific task, or request multiple independent outputs and flag disagreement. For example, a system might use one model for drafting and another for fairness auditing. Risks and dependencies: combining models does not necessarily cancel bias; correlated biases, majority voting, latency, cost, privacy, and model drift must be evaluated.
- Automated fairness regression testing in model-development pipelines (MLOps and software engineering)
- gender-gap effect sizes over time;
- per-condition refusal rates;
- output variance across repeated trials;
- differences between model versions;
- unexplained reversals in bias direction.
- This would operationalize the paper’s conclusion that auditing should be ongoing and multi-vendor.
- Human-impact studies linking model bias to real decisions (behavioral science, healthcare, education, law, finance) The paper documents model-level asymmetries but does not establish whether users adopt them. Long-term experiments should test whether exposure to differently biased models changes judgments about victims, applicants, patients, defendants, students, or employees relative to an unaided baseline. Dependencies: realistic decision environments, ethical review, representative participants, and measurement of both immediate persuasion and longer-term behavioral effects.
- Bias-aware clinical and social-service decision support (healthcare and social care) If future evidence shows that LLM outputs affect human prioritization, systems could incorporate gender-counterfactual checks into triage, abuse-risk assessment, patient communication, and referral workflows. Such tools should support—not replace—licensed professionals and should be evaluated for intersectional effects, including gender combined with age, race, disability, sexuality, and socioeconomic status. Dependency: clinical validation and compliance with applicable medical-device, privacy, and safety regulations.
- Fairness controls for autonomous agents and robotics (robotics, autonomous systems) LLM-powered robots or agents that allocate attention, protection, assistance, or physical intervention may require explicit safeguards against gender-dependent prioritization. Future systems could use formal constraints ensuring that equivalent individuals receive comparable protection or service unless a task-relevant distinction justifies otherwise. Dependencies: reliable identity and context representation, formal definitions of equivalence, robust perception, and extensive simulation and real-world testing.
- Policy standards for model-version traceability and audit access (regulation and international governance) Policymakers could require providers of high-impact models to maintain versioned evaluation records, disclose material behavioral changes, preserve audit logs, and provide controlled researcher access. This is especially important because a consumer-facing product may not disclose a precise model version, making replication and accountability difficult.
- Development of uncertainty- and abstention-aware interfaces (software, education, decision support) Future interfaces could distinguish among settled outputs, unstable outputs, refusals, and model uncertainty. In the paper, some models were consistently extreme, some were variable, and another was tightly clustered near the midpoint; these patterns should not be treated as equivalent. Interfaces could therefore display response dispersion and prompt users to seek human review when demographic gaps or instability are detected.
- Cross-linguistic and cross-cultural gender-bias evaluation (global AI deployment and academia) The study used English prompts and a particular nuclear-apocalypse framing. Long-term research should determine whether the observed asymmetries persist across languages, cultures, legal systems, gender systems, and locally salient forms of violence. Assumption: translations that preserve literal wording may not preserve social meaning, so culturally validated stimuli and local researchers are necessary.
- Bias-remediation methods that avoid compensatory discrimination (model alignment and policy) Because the models differed in direction—some favoring women, some favoring men, and others showing no measurable asymmetry—remediation should target equal treatment and task validity rather than simply reversing the observed pattern. Potential methods include balanced fine-tuning data, counterfactual augmentation, constrained decoding, adversarial evaluation, and explicit fairness objectives. Dependency: remediation must be validated across tasks and demographic groups to ensure that reducing one measured disparity does not introduce another or suppress legitimate safety responses.
Glossary
- Alignment procedures: Methods used to make an AI system’s behavior conform to specified human preferences, values, or safety requirements. “attributed this pattern primarily to alignment procedures rather than to stereotype attribution specifically.”
- Androcentric assumptions: Assumptions that treat men or male experiences as the default or norm. “defaulted to androcentric assumptions when generating policy unless gender was explicitly mentioned in the prompt.”
- API (Application Programming Interface): A programmatic interface that allows software to interact with a model or service. “Mistral Small~4, which was tested via API.”
- Asymmetry: A systematic difference between two otherwise comparable groups, conditions, or directions of comparison. “These results indicate that gender-related biases are common in LLMs.”
- Baseline accuracy: Performance measured under a straightforward reference condition used for comparison. “included to verify each model's baseline accuracy at gender identification when it is not required to rely on stereotypical inference.”
- Behavioral asymmetry: A consistent difference in how a system responds to comparable situations involving different social categories. “systematic non-neutrality in LLMs is not confined to gender, but reflects broader behavioral asymmetries across multiple socio-cognitive domains.”
- Bias auditing: Systematic evaluation of an AI system to identify discriminatory or otherwise uneven behavior. “Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.”
- Cognitive surrender: Excessive reliance on an external system’s reasoning at the expense of independent judgment. “Related work documents cognitive surrender in reasoning tasks.”
- Control phrase: A stimulus designed to provide an explicit reference condition for validating a measurement procedure. “The remaining three pairs were control phrases, in which the writer's gender was explicitly stated.”
- Crosses: In experimental design, combines every level of one factor with every level of another factor. “that crosses victim gender with a dimension of violence carrying different degrees of gender-political salience.”
- Deontological ethics: An ethical framework that evaluates actions according to duties or rules rather than consequences. “women tend to embrace deontological ethics more than men in personal.”
- Diametrically opposed: Directly and completely opposite in direction or outcome. “Three models showed extreme, and diametrically opposed, response patterns.”
- Effect size: A quantitative measure of the magnitude of a difference or relationship, distinct from whether it is statistically significant. “the magnitude of this asymmetry ranged from 68--85\% of cases depending on the model.”
- Export controls: Government restrictions on the international transfer or availability of technologies, goods, or services. “The controls were lifted on 30~June~2026, and Anthropic restored global access on 1~July~2026.”
- Fine-tuning: Additional training of a pretrained model on selected data to adapt its behavior or capabilities. “an asymmetry the authors attributed to fine-tuning techniques such as reinforcement learning from human feedback.”
- Gender attribution: Assigning a gender category to a person based on available information or inferred cues. “gender attribution to stereotyped phrases.”
- Gender-political salience: The degree to which an issue is associated with politically significant debates about gender. “abuse is closely tied to real-world discourse on gender-based violence, whereas torture carries no comparable gendered connotation.”
- Hallucinating: Producing an unsupported, fabricated, or arbitrary output that is presented as though it were grounded in information. “these models are not simply hallucinating an arbitrary asymmetry.”
- Heterogeneity: Variation in characteristics, effects, or behaviors across systems or groups. “Their direction and magnitude, however, are highly heterogeneous.”
- Inclusivity index: A measure of how far a model’s responses depart from the stereotypical response associated with a phrase. “we define the inclusivity of a single phrase as the mean absolute distance.”
- Independent-samples t-test: A statistical test comparing the means of two independent groups or samples. “For each model, we tested against using an independent-samples -test.”
- Incognito or temporary-chat feature: A mode intended to prevent a model from retaining conversational context across interactions. “using each model's incognito or temporary-chat feature for every iteration.”
- Implicit bias: An unintentional or indirectly expressed preference or association that affects judgments or behavior. “Finally, these biases are implicit, as they do not emerge when GPT-4 is directly asked to rank moral violations.”
- Inference: The process of deriving a conclusion from available evidence or observed cues. “rather than from any biological indicator.”
- Iteration: A single repeated execution of an experiment, prompt, or model interaction. “Each prompt was presented ten times per phrase.”
- Listwise deletion: A method of handling missing data by removing observations with missing values from an analysis. “standard listwise-deletion procedures handle missing values appropriately without further adjustment.”
- Moral chivalry: A proposed tendency to protect female targets from harm more than male targets in moral judgments. “a male-disadvantaging, or ``moral chivalry,'' asymmetry in LLM moral judgment.”
- Moral dilemma: A situation in which competing moral principles or values make the appropriate action uncertain. “using two task paradigms previously applied to a smaller set of models: gender attribution to stereotyped phrases and moral judgment in sacrificial dilemmas.”
- Moral deference: Reliance on another agent’s moral advice or judgment instead of one’s own assessment. “found that moral deference could be undermined by obviously implausible justifications.”
- Non-degenerate variance: Variation in observations sufficient for standard statistical comparisons to be mathematically defined. “The remaining five models showed non-degenerate variance across most or all conditions.”
- Non-parametric test: A statistical test that does not require the data to follow a particular distributional form. “we also performed the non-parametric Wilcoxon-Mann-Whitney rank-sum test.”
- Nuclear-apocalypse framing: Presenting a decision scenario in which an action is justified as preventing global nuclear catastrophe. “embedded in a nuclear-apocalypse framing.”
- Occupational stereotypes: Generalized beliefs associating particular occupations with specific social groups or traits. “when studying occupational stereotypes.”
- Pairwise comparison: A statistical comparison between two conditions, groups, or measurements. “For each model, we conducted four pairwise comparisons.”
- Proprietary: Owned and controlled by an organization, with details not publicly accessible. “their training data, fine-tuning procedures, and safety interventions remain proprietary.”
- Rank-sum test: A statistical test that compares the distributions or ranks of two independent samples. “the rank-sum test confirms the difference.”
- Reinforcement learning from human feedback: A training method that uses human evaluations to optimize a model’s outputs toward preferred behavior. “reinforcement learning from human feedback rather than to the training corpus itself.”
- Robustness check: An additional analysis testing whether a result remains under altered assumptions or data-selection choices. “As a robustness check, we repeated the test excluding the three control phrases.”
- Sacrificial dilemma: A moral scenario in which harming or sacrificing one person is considered as a means of preventing a greater harm. “moral judgment in sacrificial dilemmas.”
- SAGER guidelines: Reporting guidelines intended to improve the treatment of sex and gender in research. “Following the Sex and Gender Equity in Research (SAGER) guidelines.”
- Safety interventions: Measures introduced to restrict harmful, undesirable, or unsafe model behavior. “their training data, fine-tuning procedures, and safety interventions remain proprietary and inaccessible to external researchers.”
- Socio-cognitive domains: Areas involving both social processes and cognitive judgments or behaviors. “broader behavioral asymmetries across multiple socio-cognitive domains.”
- Standard deviation (SD): A measure of the spread of observations around their mean. “For torture-woman, all 50 responses were identical both before and after the suspension (, ).”
- Standard error (SE): An estimate of the uncertainty in a sample statistic, commonly the sample mean. “Inclusivity indices and significance test by model (Study 1).”
- Statistical artifact: An apparent result produced by a measurement or analytical procedure rather than by the underlying phenomenon. “it reflects genuine model characteristics rather than measurement artifacts.”
- Statistical significance: A criterion indicating that an observed result would be relatively unlikely under a specified null hypothesis. “models marked with an asterisk showed a statistically significant difference between and at .”
- Stereotypical inference: Deriving a social-category judgment from culturally associated traits rather than explicit information. “when it is not required to rely on stereotypical inference.”
- Systemic disadvantage: Widespread, institutionally or structurally produced unequal outcomes affecting a social group. “can translate into real, systemic disadvantage.”
- Training corpus: The body of data used to train a machine-learning model. “rather than to the training corpus itself.”
- Zero variance: A condition in which all observations have the same value, leaving no measured dispersion. “several models produced near-zero-variance response distributions.”











