Papers
Topics
Authors
Recent
Search
2000 character limit reached

Identifying, Explaining, and Correcting Ableist Language with AI

Published 23 Feb 2026 in cs.HC | (2602.19560v1)

Abstract: Ableist language perpetuates harmful stereotypes and exclusion, yet its nuanced nature makes it difficult to recognize and address. Artificial intelligence could serve as a powerful ally in the fight against ableist language, offering tools that detect and suggest alternatives to biased terms. This two-part study investigates the potential of LLMs, specifically ChatGPT, to rectify ableist language and educate users about inclusive communication. We compared GPT-4o generations with crowdsourced annotations from trained disability community members, then invited disabled participants to evaluate both. Participants reported equal agreement with human and AI annotations but significantly preferred the AI, citing its narrative consistency and accessible style. At the same time, they valued the emotional depth and cultural grounding of human annotations. These findings highlight the promise and limits of LLMs in handling culturally sensitive content. Our contributions include a dataset of nuanced ableism annotations and design considerations for inclusive writing tools.

Summary

  • The paper compares GPT-4o with trained disability-community annotators across identification, explanation, and correction of nuanced ableist language, finding equivalent sentence-level agreement averaging 72.3%.
  • GPT-4o was preferred overall by 43.9% of participants versus 23.4% for the human baseline, largely because of clearer explanations, consistent formatting, plain language, and non-judgmental corrections.
  • The findings support AI as a scalable complement to disabled expertise, while emphasizing community-centered design, contextual customization, preservation of identity-relevant language, and avoidance of prescriptive or culturally insensitive edits.

Overview and motivation

This paper examines whether LLMs can serve as effective annotators of nuanced ableism—subtle, often normalized language that perpetuates stereotypes about disability without explicit slurs. The authors, spanning Carnegie Mellon University, Columbia University, and Microsoft Research, conduct a two-part study comparing GPT-4o annotations against crowdsourced annotations from trained members of the disability community, then having disabled participants evaluate both side by side. The work is motivated by the observation that ableist language is less publicly recognized than racism or sexism, more socially normalized, and therefore less likely to be flagged by general hate-speech tools (2602.19560). The authors position AI assistance not as a replacement for disabled expertise but as a scalable complement that reduces the emotional labor placed on marginalized communities to continuously correct others' language.

The study addresses three research questions: (RQ1) whether humans prefer AI or human annotators for identifying, explaining, and correcting ableist language; (RQ2) which qualities of each annotation style make them agreeable or disagreeable; and (RQ3) how AI annotators can improve.

Study 1: Collecting community annotations

In Study 1, the authors used GPT-4o to generate one-paragraph fictional stories about individuals across seven disability categories (vision impairment, hearing impairment, mental health conditions, intellectual/learning disabilities, neurological disabilities, autism, and reduced mobility), following a prompting strategy from prior work showing that LLMs produce ableist content when asked to write about disabled people completing tasks. Stories were piloted with in-group participants to verify that nuanced ableism was present according to criteria synthesized from UN, APA, and National Center on Disability and Journalism guidelines. A total of 110 participants—each holding at least a bachelor's degree with prior DEI or disability-related training—annotated their assigned story at both sentence level and passage level across three tasks: identification, explanation, and correction. This yielded 276 annotations (210 sentence-level, 66 story-level).

Two findings from this phase deserve emphasis. First, inter-annotator agreement was low: Fleiss' Kappa ranged from 0.175 (reduced mobility) to 0.396 (intellectual/learning disabilities). The authors argue this is not a methodological failure but itself a finding, reflecting genuine subjectivity in interpreting subtle ableism and diversity of perspective within disability communities. Second, despite sentence-level variability, a majority of participants agreed that the overall stories contained ableism in five of seven surveys, validating the stimulus design. The choice of AI-generated narratives is a deliberate trade-off: it standardizes complexity and length across seven disability groups, but limits generalizability to human-authored texts with broader stylistic and cultural variation—a limitation the authors state plainly.

Constructing the human baseline and AI annotations

Rather than treating any single annotation as ground truth, the authors constructed composite human baseline annotations by including phrases that at least two participants independently identified as ableist for similar reasons. Explanations and corrections were synthesized through an AI-assisted, human-led process: GPT-4o produced first-pass summaries of small annotation sets (n = 4–10), which researchers then manually edited under four documented rules (retain participant wording; add only source-grounded phrases; remove unsupported claims; resolve disagreements by inclusion rather than omission). The authors acknowledge that using generative AI to build a "human" baseline introduces risk of overgeneralization, mitigated by these rules and by tagging all annotations omitted from the final baseline.

AI annotations were produced via chained GPT-4o prompts after earlier attempts using curated anti-ableism guidelines yielded inconsistent, overly critical output. Notably, the human and AI annotators received different instructions—the humans were constrained toward spell-checker-like objectivity, while the AI prompts were tuned until outputs naturally followed similar conventions. The authors justify this asymmetry as investigating how AI "naturally" annotates rather than simulating a disabled annotator, though it means the comparison is not fully controlled on instruction format.

Study 2 results: equivalent agreement, significant preference for AI

Study 2 surveyed 106–108 disability community members who evaluated both annotations under the impression that both were AI-generated, with annotator order randomized. The central quantitative result is a divergence between per-item agreement and overall preference:

  • Sentence-level agreement was statistically equivalent across identification (Z = 1.47, p = 0.14), explanation (Z = −0.05, p = 0.96), and correction (Z = −0.75, p = 0.45), averaging 72.3% for both annotators.
  • Overall preference significantly favored the AI: 43.9% preferred the AI versus 23.4% for the human baseline, with 32.7% expressing no preference (χ² = 6.80, p = 0.0333; two-sample Z-test p = 0.0011).

Task-level patterns showed the human annotator rated slightly higher on identification in 6 of 7 surveys while the AI was favored for corrections in 6 of 7. Preliminary subgroup analysis found lower agreement for autism, blindness, and Deafness stories, particularly on identification tasks, suggesting contested representations in those communities warrant culturally grounded handling.

The implication is that perceived annotation quality depends on more than correctness: clarity, tone, formatting consistency, and cognitive load shaped preferences even when accuracy was comparable. Thematic analysis attributed the AI's advantage to narrative consistency, plain language, well-aligned explanations and corrections, and non-judgmental framing of obvious harms. Participants valued the human annotations' emotional depth, advocacy vocabulary ("inspiration porn," "infantilization"), cultural grounding, and justice-oriented reframing—but criticized them for wordiness, inconsistent logic between explanations and corrections, and grammatical awkwardness.

Risks identified by participants

Participant feedback surfaced substantive concerns beyond tool quality. Several warned against over-sanitization and erasure: removing mentions of Deafness diminished a story's portrayal of lived experience, and some argued that diluting depictions of struggle is itself ableist ("Diluting it is ableist because it reduces the struggle of the people going through it"). Others rejected the premise of prescriptive language policing ("Stop trying to force people to use one set of 'acceptable' language") and questioned outsider authority over community-specific norms ("Unless you've lived in an autistic mind, you really shouldn't have much of an opinion"). These responses indicate that annotation tools carry epistemic risks when they frame bias judgments as objective fact or override identity-affirming language, reinforcing the need for genre-aware, context-sensitive deployment.

Design guidelines

From the qualitative analysis, the authors derive guidelines organized into four pillars: education and empathy building (grounding feedback in historical context, framing corrections as suggestions rather than mandates); community-centered design (centering marginalized perspectives in training data, minimizing epistemic harm, positioning tools as co-authors rather than enforcers); practical revision support (balancing correction with preservation of voice, ensuring explanations align with corrections, making thematic feedback modular); and context and customization (adapting annotation thresholds to genre and intent, avoiding static word lists that ignore reclamation and intra-group variation). These guidelines respond directly to observed failure modes: the AI's tendency to strip identity-relevant detail, and mismatches between what an explanation identifies and what a correction resolves.

Limitations and open questions

The authors are explicit about several constraints. All stimuli were short AI-generated fictional stories, so findings may not generalize to nonfiction, journalism, medical writing, or autobiographical text—and participants noted that ableist tropes in fiction may be intentional, making blanket correction inappropriate. The participant pool consisted of college-educated, DEI-trained disabled people, which is not representative of the broader disability community; the paper leaves open who within marginalized communities should count as an "expert" contributor. The human baseline involved editorial synthesis with AI assistance, raising reproducibility concerns that are mitigated but not eliminated by the documented editing rules. The dataset, while described as the first crowdsourced collection of nuanced ableism annotations, remains too small to train generalizable bias-aware models. Finally, the study measures immediate reactions only; retention, behavioral change, and long-term effects on writing practice remain untested, as do effects on non-disabled users who would be primary consumers of such tools.

Conclusion

This paper provides empirical evidence that GPT-4o can identify, explain, and correct nuanced ableist language at parity with trained disability community annotators on per-item agreement, while being significantly preferred overall for its clarity, consistency, and accessible formatting. At the same time, it documents where LLM-based annotation falls short: preserving narrative integrity, respecting identity-affirming language, and delivering the cultural depth and affective resonance that human annotators provide. The contributions—a community-sourced ableism annotation dataset, a controlled comparison of AI and human annotation styles, and design guidelines for culturally sensitive writing tools—position AI annotation as a viable educational complement to, rather than substitute for, disabled expertise, contingent on participatory design and careful attention to context.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.