Papers
Topics
Authors
Recent
Search
2000 character limit reached

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

Published 26 May 2026 in cs.CL and cs.AI | (2605.26397v1)

Abstract: LLMs are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-stakes settings affecting autistic communities. While previous research has identified disability-related biases in LLMs, it remains unclear how they conceptualize ableism or detect it in text. We introduce a bias-aware evaluation framework targeting anti-autistic ableist language with a psychometrically-weighted, community-proximate ground truth anchored in annotator positionality. This framework constitutes a stricter standard than conventional majority-vote aggregation which significantly and consistently underweights autistic and autism-accepting perspectives. We find that LLMs frequently produce harmful outputs, mislabel community-reclaimed language as ableist, and express more negative attitudes toward autistic people when assessment instruments are masked. Our error analysis reveals that models rely on surface-level keyword matching rather than contextual factors such as speaker identity, and whether the language fosters in-group solidarity or inflicts out-group harm.

Summary

  • The paper introduces a composite psychometric weighting method using implicit bias, explicit attitudes, and autistic traits to build stricter, community-aware labels, producing significantly lower agreement with LLMs than majority voting (p < .001, d = 0.869).
  • LLM performance remains poor against the weighted ground truth, with mean weighted κ = 0.118 and F1 = 0.237; safety-focused models perform best, while most systems over-classify content as ableist because recall exceeds precision.
  • The paper finds that models rarely reason about insider–outsider context or reclaimed language, and masked psychometric testing suggests recognizable instruments can conceal anti-autistic attitudes through socially desirable responses.

This paper introduces a bias-aware evaluation framework for detecting anti-autistic ableist language, built on a psychometrically weighted ground truth that treats annotator positionality as signal rather than noise (2605.26397). The work addresses a methodological gap in harmful-language detection: conventional majority-vote aggregation of annotator labels systematically underweights the perspectives of autistic and autism-accepting annotators, and LLMs evaluated against such aggregated labels may appear better aligned than they actually are. The authors combine instruments measuring implicit bias (IAT), explicit attitudes (SATA), and autistic traits (AQ) into a composite annotator reliability score, evaluate twelve LLMs under seven prompting strategies, and administer recognition-controlled psychometric instruments to the models themselves.

Motivation and problem framing

The paper is grounded in two observations about anti-autistic ableism that distinguish it from other forms of toxic speech. First, the harmfulness of autism-related language is context-dependent in an insider/outsider sense: terms such as "autie" or "aspie" are self-affirming when used within autistic communities but ableist when used by outsiders, creating systematic misclassification risk for both human raters and automated systems. Second, prior NLP research has framed autism primarily through diagnosis, cure, or analogy for model behavior, while the ability of LLMs to recognize and reason about ableism remains largely unexamined. The authors build on the AUTALIC dataset of autism-related social media posts annotated for ableism, but depart from its majority-vote ground truth by weighting annotators according to measured bias and community proximity.

The study poses three research questions: how psychometric weighting affects estimates of human–LLM agreement; how LLM reasoning differs from human reasoning on this task and whether prompting closes the gap; and whether masking instrument-recognition cues changes LLM psychometric self-assessment.

Psychometrically weighted ground truth

Annotators completed the SATA scale, the AQ questionnaire, and an autism-adapted IAT. The authors compute a composite trust score RiR_i as the mean of min–max normalized scores, with IAT inverted so higher values indicate lower implicit bias, then normalize RiR_i relative to each annotation team's mean to produce weights WiW_i in [0.881,1.168][0.881, 1.168]. Ground-truth labels are weighted means of annotator labels binarized at 0.5.

A key empirical finding motivates the composite design: autistic annotators (n=3n=3) show substantially greater dispersion in implicit bias (IAT SD = 0.344) than non-autistic annotators (n=6n=6, SD = 0.120), despite comparable explicit-attitude consistency. Two annotators can share high AQ scores yet diverge sharply in implicit bias, so single-instrument selection is insufficient. This is confirmed by robustness analysis: AQ-only weighting fails to produce a ground truth significantly stricter than majority vote (p=.651p = .651, d=0.096d = 0.096), whereas the full composite does so consistently (p<.001p < .001, d=0.869d = 0.869). Across all 56 model–condition combinations, psychometrically weighted Cohen's RiR_i0 was lower than unweighted RiR_i1 (mean RiR_i2, negative in 84% of combinations), establishing that majority vote overestimates human–LLM agreement. Model rankings remain stable across five weighting schemes (RiR_i3), so the weighting affects ground-truth stringency without changing which models perform best.

The task itself is genuinely difficult: intra-group human–human RiR_i4 reaches only 0.208 for the high-AQ group and 0.016 for the high-bias group. One high-bias annotator assigned non-neutral labels to 99.3% of items, justifying down-weighting biased raters.

Classification performance

Performance across eight models and seven prompting conditions is low: mean weighted RiR_i5 and mean F1 = 0.237 against the weighted ground truth. Three structural patterns emerge:

  • Precision–recall imbalance: recall substantially exceeds precision in 51 of 56 combinations (mean recall 0.531 vs. precision 0.173), reflecting systematic over-classification of content as ableist. Given AUTALIC's class imbalance (~13% ableist), accuracy (mean 0.661) is misleading.
  • Two performance tiers: safety-oriented classifiers outperform general-purpose models regardless of parameter count. LLaMA Guard 3 (mean RiR_i6) and GPT-OSS 20B (0.175) exceed all instruction-tuned models (range 0.013–0.145). Task-specific safety training appears to encode sensitivity to disability-related harm signals that prompting alone cannot recover.
  • Condition effects: zero-shot and chain-of-thought (CoT) prompting yield the highest performance (mean F1 = 0.273 each); CoT benefits are model-specific (largest for GPT-OSS at RiR_i7). In-context learning performs below zero-shot for most models, with exemplars destabilizing strong models more than helping weak ones (StarlingLM RiR_i8). Persona prompts fall below zero-shot and CoT across all three identity-language framings, with no consistent identity-first-over-person-first advantage.

On reference-group alignment, models agree significantly more with reliable annotators than with autistic-proximate annotators (RiR_i9, WiW_i0, WiW_i1; the sole comparison surviving Bonferroni correction), validating the composite weighting scheme. Autistic-proximate and non-autistic reference groups are statistically indistinguishable, suggesting either that LLMs fail to align with any community perspective or that the small autistic-proximate group limits discriminative power. LLaMA Guard 3 is an outlier: it achieves its highest WiW_i2 against high-bias annotators under persona conditions, likely because systematic over-labeling coincidentally matches the high-bias rater who labeled 99.3% of items non-neutral — implying its positive WiW_i3 reflects over-labeling rather than genuine alignment.

Reasoning failures

A triangulated error analysis of 92 sampled model outputs identifies surface keyword matching as the dominant failure mode, persisting even under explicit CoT scaffolding instructing models to consider speaker identity and reclaimed terms. Four findings stand out:

  • Insider/outsider blind spot: wording-driven reasoning appeared in 41% of excerpts; stereotype invocation reached 34% (thematic lens) and 35% (linguistic lens), the only pattern achieving cross-lens convergence. Models flagged community-reclaimed terms ("aspie," "autie") as ableist in 68% of instances where autistic-proximate annotators did not, while deficit-based framings lacking negatively valenced keywords were frequently labeled non-ableist.
  • Positionality reasoning is rare and condition-dependent: speaker-identity reasoning appeared in 10% of excerpts overall and 13% under CoT, with zero instances in zero-shot or persona conditions — a default behavior gap rather than a hard capability limit.
  • CoT redistributes rather than repairs errors: topic deflection dropped from 25% (zero-shot) to 0% (CoT), but positionality reasoning rose only from 0% to 13%. Structured scaffolding eliminates question-evasion without producing contextual judgment.
  • Hallucination and deflection: representative errors include a model asserting a sentence contains "no reference to autism" when it explicitly mentions neurodivergence, another flagging hyperbolic in-group frustration as a literal violent threat, and one hallucinating an entirely absent topic (anti-Semitism).

Masked psychometric evaluation of LLMs

To counter recognition-based responding — since published instruments likely appear in training data — the authors administered a purpose-built rewrite combining SATA and AQ items with randomized ordering, varied phrasing, filler items from the International Personality Item Pool, and 52 validity dummies. Nine scoreable models showed a consistent directional divergence: six of nine declined on the rewritten SATA (significant, paired WiW_i4, WiW_i5, WiW_i6), with GPT-OSS 20B and DeepSeek-R1 1.5B each dropping a full tier, while eight of nine increased on the AQ (directional trend only, WiW_i7), raising clinical-range scorers from four to six. The interpretation is that recognizable instruments elicit socially desirable responding concealing latent anti-autistic attitudes. Because the masking manipulation simultaneously alters item wording, structure, and filler composition, the shift can only be attributed to instrument-level recognition broadly, not any single design change. Several models were excluded or flagged for refusing items or endorsing dummies indiscriminately.

Limitations

The paper is candid about several constraints. Each participant completed each psychometric instrument only once, though these tools are more reliable with repeated administration. The error analysis is designed for pattern identification, not frequency estimation: roughly 1–2 excerpts per model–condition combination means qualitative findings illustrate reproducible patterns rather than corpus-level rates, and one analyst used open-ended characterization rather than a predefined coding scheme. The segmentation analysis rests on a small annotator pool — the autistic-proximate group has WiW_i8, and high-bias and unreliable groups are identical by construction — so key comparisons, though statistically significant, should be read in light of sample size. Participants were predominantly college students in computing programs in a Western context, and the authors do not claim generalizability to multilingual or cross-cultural settings. They also acknowledge that the standardized instruments themselves are rooted in the medical model of disability, terminology that many autistic people reject, and that developing community-informed alternatives lies outside the paper's scope. Finally, whether LLaMA Guard 3's anomalous alignment pattern reflects over-labeling remains an unverified hypothesis.

Conclusion

This paper demonstrates that ground-truth construction is not neutral: majority-vote aggregation measurably overstates human–LLM agreement on anti-autistic ableism detection, and a composite psychometric weighting anchored in implicit bias, explicit attitudes, and autistic traits yields a stricter, more community-proximate standard without altering model rankings. Mean F1 of 0.237 indicates near-chance performance on a task deployed in consequential moderation decisions affecting autistic communities, leading the authors to argue that task-specific fine-tuning on community-labeled data is a precondition for deployment-grade performance rather than an optional enhancement. The framework — annotator screening, reliability weighting, weighted ground truth, and recognition-controlled psychometrics — offers a replicable template for community-grounded evaluation in harmful-language detection, while leaving open questions about scaling the approach beyond small annotator pools and verifying the mechanisms behind safety-classifier behavior.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.