- The paper demonstrates that LLMs systematically underrepresent religious perspectives in ethical queries despite high human expectations.
- The authors introduce the AllFaith Religious Representation Benchmark with 150 curated questions to assess model performance.
- Findings reveal a significant mismatch between human religious expectations and LLM outputs, stressing the need for more inclusive alignment protocols.
Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making
Introduction and Motivation
This work addresses a fundamental gap in the evaluation of LLMs: the systematic omission of religious perspectives when responding to everyday ethical queries. Previous benchmarks have focused extensively on the presence of bias, stereotypes, or partisan attitudes. In contrast, this study introduces and benchmarks omissive bias—the absence of relevant perspectives, specifically religious, in LLM outputs when such frameworks are contextually pertinent.
As LLMs increasingly become the default sources for personal, moral, and existential advice, their handling of religion becomes consequential. With a majority of the global population identifying with some religion, religions serve not merely as philosophical systems but as embedded sources of guidance on grief, forgiveness, family, and purpose. The hypothesis is that current LLMs, largely aligned via Western, secular-rationalist paradigms, systematically undervalue or omit religious perspectives, particularly in practical life domains.
Benchmark Construction and Methodology
The authors present the AllFaith Religious Representation Benchmark, consisting of 150 open-ended, ethically salient questions. The selection process began with the WildChat-1M corpus—containing real-world user-LLM dialogues—which was filtered through automated and manual processes to focus on questions where religion could plausibly be an integral part of a good answer. Supplementary questions, curated by diverse faith communities, further ensured coverage of topics frequently addressed by religious traditions.
Figure 1: The pipeline for sourcing and curating questions for the benchmark from WildChat-1M and supplementary faith community contributions.
Rather than assess the correctness or quality of religious content, the rubric set an intentionally low bar: any mention of religion, religious practice, or religious leader sufficed for full credit. Benchmark evaluation followed the LLM-as-judge paradigm, employing another LLM (Gemini 3.1 Pro Preview) to robustly and systematically annotate religious content.
To establish a ground truth for expectation, the study conducted a large-scale human survey (n=1,125, 11,250 ratings), wherein participants scored the likelihood that an answer to a given question would reference religious ideas, practices, or leaders.
Experimental Results: Systematic Omission
A broad array of 27 leading LLMs, spanning both commercial and open-source ecosystems, were evaluated on the benchmark. Across all major question categories, LLMs displayed a markedly low propensity for religious representation, with an average score of 0.084 (binarized; max 1, min 0). This result persists despite significant human expectation for religious input, which remains stable across categories (Figure 2).
Figure 2: Mean human religious expectation scores by question category versus LLM performance, indicating a persistent and substantial omission gap.
Drilling into subcategories and models revealed some variation, but the general trend of neglecting religious perspectives was robust (Figure 3). Models tended to invoke religion more readily for abstract existential questions—meaning, death, and truth—rather than practical matters such as grief, marriage, addiction, or family conflict. This asymmetry suggests that LLMs treat religion as a philosophical relic instead of a daily resource.
Figure 3: Mean LLM religious relevance scores across subcategories and AI models, demonstrating uniformly low representation except in a narrow band of existential themes.
A granular comparison between human and model scores (Figure 4) confirms the weak alignment: models routinely underperform relative to human expectation on religious mention, with little calibration to the situations where religious input is contextually salient.
Figure 4: Categorical comparison of human and model ratings on question religious relevance, with the retained benchmark focusing on items with the greatest misalignment.
Concrete model outputs further illustrate the nature of the omission: LLM responses to questions with high human religious expectations were often entirely secular, even when secular approaches alone may be insufficient for many users seeking moral or existential counsel.
Human Expectation versus LLM Behavior
The expectation-behavior comparison produces two key insights. First, humans—including respondents with varying degrees of religiosity—generally expect some religious dimension in answers to a wide range of ethical questions (Figure 5). Second, LLMs' behavior is uncorrelated with these expectations (r = 0.257), failing both to introduce religion when widely expected and to calibrate to contexts of heightened religious salience.
Figure 5: Human religious relevance scores across question categories, stratified by participant religiosity.
Survey demography (Figure 6) reveals a balanced and pluralistic distribution, mitigating concerns that results are idiosyncratic to particular religious or ideological subgroups.
Figure 6: Demographic distributions of survey respondents, including religious affiliation, age, political ideology, and religiosity.
Practical and Theoretical Implications
These findings have immediate implications for the design and alignment of LLMs. First, the systematic omission of religion constitutes representational harm, a form under-theorized and under-benchmarked in current AI evaluation practices. Second, the over-emphasis on anti-stereotypic and neutral alignment protocols results in a secular baseline that may itself be non-neutral for a significant portion of global users.
This omissive bias is not limited to technical oversight. Analysis of alignment protocols (e.g., OpenAI Model Spec, Claude Constitution) reveals a near-total absence of explicit religious policies. The behavior thus appears to be an emergent property of atheoretical, safety-oriented alignment processes rather than conscious exclusion, though design choices favoring generic, secular advice have predictable impact.
Pragmatically, these results suggest that current LLMs misrepresent the pluralism of user worldviews, diminish religion’s practical relevance, and elevate philosophical abstraction over actionable, tradition-rooted frameworks. The risk is the quiet marginalization of the religious perspectives that structure many users’ ethical reasoning and life choices. Theoretical work on representational bias, symbolic annihilation, and erasure in media and algorithmic systems underlines the long-term cultural and societal stakes.
Limitations and Prospective Directions
The benchmark, though diverse and grounded in real-world interactions, is US- and Western-skewed; global religious variation remains underrepresented. Question selection, despite robust pipelines, may inadvertently inherit classifier and manual curation biases. LLM evaluation via automated judging introduces possible secondary model bias. These constraints provide a foundation but not a ceiling for future work.
Future research could address dynamic adaptation to user religiosity and tradition, assess the quality and appropriateness of religious content rather than mere presence, and extend the analysis to multi-turn and contextualized dialogues. Rigorous pluralistic alignment protocols, explicit policy development on religion, and faith-community involvement in LLM design emerge as necessary next steps.
Conclusion
The AllFaith Religious Representation Benchmark exposes an empirically robust omissive bias in current LLMs: models rarely invoke religion as a practical or ethical resource in everyday decision-making, even when human users expect such input. This omission is most pronounced in practical domains—grief, guilt, forgiveness, family, and purpose—where religious traditions historically play central roles. Addressing this gap is essential for value-plural, genuinely representative AI systems that faithfully reflect the lived frameworks guiding human action. Future LLM development must incorporate explicit strategies for egalitarian religious representation, not merely avoid negative stereotyping, if alignment is to achieve genuine inclusivity and distributive fairness.