Flourishing AI Benchmark for Well-Being
- The Flourishing AI Benchmark is a comprehensive framework that operationalizes human flourishing into seven empirically validated dimensions.
- It employs both objective tests and expert-evaluated subjective questions, using specialized LLM judges to assess ethical, social, and practical well-being.
- Empirical results reveal significant performance gaps among AI models, especially in ethical and spiritual dimensions, underscoring areas for future improvement.
The Flourishing AI Benchmark (FAI Benchmark) is a comprehensive evaluation framework for assessing how effectively AI systems, particularly LLMs, align with and promote holistic human flourishing. Unlike benchmarks that prioritize technical competence or harm prevention, the FAI Benchmark operationalizes flourishing as a multi-dimensional construct grounded in empirical well-being science, ethical theory, and interdisciplinary expertise, and provides formal measurement of AI’s capability to actively support every major domain of human life (Hilliard et al., 10 Jul 2025, Laukkonen et al., 11 May 2026).
1. Conceptual Foundations and Motivation
The motivation underlying the FAI Benchmark emerges from the observation that AI alignment research and evaluation have historically emphasized harm avoidance, narrow task performance, or responses to isolated moral dilemmas, rarely addressing the AI's capacity to advance holistic human well-being. Human flourishing, as defined in VanderWeele (2020), is “a state in which all aspects of a person’s life are good,” encompassing not only subjective happiness but also virtues, relationships, health, purpose, stability, and spirituality. As LLMs permeate critical domains—education, health care, finance, and faith contexts—there is a pressing need for evaluation tools that measure how AI can help people thrive rather than simply avoid adverse outcomes (Hilliard et al., 10 Jul 2025).
The FAI Benchmark aims to (1) measure AI contributions to human flourishing over seven empirically validated dimensions, (2) enforce balanced and non-siloed performance without compensatory tradeoffs, and (3) provide a rigorous, open, and extensible methodology for ongoing interdisciplinary refinement (Hilliard et al., 10 Jul 2025, Laukkonen et al., 11 May 2026).
2. Dimensions of Human Flourishing
FAI Benchmark adopts a multidimensional model rooted in the Harvard Human Flourishing Program, the Global Flourishing Study, and major psychological and sociological frameworks. The seven dimensions operationalized are:
| Dimension | Definition | Foundational Source |
|---|---|---|
| Character and Virtue | Acting to promote good in all circumstances; exercising classical virtues (prudence, etc) | VanderWeele 2020; Barna Group |
| Close Social Relationships | Quality and satisfaction of friendships and intimate relationships | Harvard HFP; GFS |
| Happiness and Life Satisfaction | Hedonic well-being and overall life satisfaction | VanderWeele 2017/2020 |
| Meaning and Purpose | Understanding life purpose and perceiving actions as worthwhile | GFS |
| Mental and Physical Health | Self-rated mental and physical well-being | WHO; medical assessment frameworks |
| Financial and Material Stability | Freedom from worry regarding basic needs; secure flourishing | GFS; Council for Economic Education |
| Faith and Spirituality | Communion with the transcendent; religious engagement, spiritual practice | Barna Group; GFS |
Each dimension is independently validated and recognized as desirable in its own right. The framework acknowledges the inherent multi-objective nature of flourishing, where progress in one dimension may not straightforwardly increase others due to potential tradeoffs (Hilliard et al., 10 Jul 2025).
3. Methodological Structure and Evaluation Pipeline
The FAI Benchmark's methodology is characterized by its breadth, rigor, and the integration of objective and subjective instruments:
A. Question Construction
The benchmark comprises 1,229 questions, approximately 75% objective and 25% subjective:
- Objective items are adapted from validated third-party sources (e.g., MMLU moral and sociology questions for Character and Relationships, medical licensing exams for Health, personal finance quizzes, philosophy, and world religion inventories).
- Subjective items are both expert-written (notably for under-represented or complex domains such as Faith and Character) and generated by LLMs via first-person advice restatements, curated for quality and pertinence.
Each question is tagged by its principal dimension, ensuring content validity (Hilliard et al., 10 Jul 2025).
B. Specialized Judge LLMs
Subjective answers are scored by LLMs fine-tuned or prompt-engineered for “expert” evaluation—emulating professional roles (e.g., chaplain, doctor, financial advisor). Judging follows a rubric that assesses binary relevance and a vector of 25 alignment indicators (e.g., actionability, promotion of holistic well-being, negative indicators for harmful suggestions). Output scores are mapped onto a 0–100 scale. Prior work indicates that LLM judges can match or exceed inter-human reliability and can be diversified to reduce systematic bias (Hilliard et al., 10 Jul 2025).
C. Cross-Dimensional Evaluation
Recognizing the interconnectedness of flourishing, model responses may be “cross-scored” by judges outside the primary dimension when relevance is detected (e.g., financial advice with relational implications). These “tangential scores” contribute to dimension-level results, reinforcing holistic alignment (Hilliard et al., 10 Jul 2025).
D. Scoring and Aggregation
For each dimension , three component scores are computed:
- Objective Score (): Proportion correct on dimension’s objective items.
- Subjective Score (): Mean dimension-specific alignment score across all judges and subjective items.
- Tangential Score (): Weighted mean of cross-dimensional relevance scores.
The dimension score is the geometric mean . The overall FAI score is the geometric mean across all : This aggregation penalizes imbalances, ensuring no single weakness can be compensated by unrelated strengths (Hilliard et al., 10 Jul 2025).
4. Empirical Findings and Model Performance
When applied to 28 leading LLMs as of 2025, the FAI Benchmark reveals that current state-of-the-art models exhibit significant performance variance across dimensions. The top scoring model, OpenAI’s o3 (April 2025), achieved a 72% overall score. Other top performers included Gemini 2.5 Flash Thinking (68%), Grok 3 (67%), o4-mini (66%), and GPT-4.5 Preview (66%), with a cohort of models in the mid-60s percentile range.
Dimension-level results display pronounced disparities:
- Finance: 81% (highest average across models)
- Health: 72%
- Relationships: 67%
- Faith and Spirituality: 35% (lowest)
- Meaning and Purpose: 56%
- Character and Virtue: 58%
Even for the strongest model, Faith and Spirituality, as well as Meaning and Purpose, remain notably weak (Faith 43%, Meaning 66%, Character 87%). This demonstrates persistent challenges in aligning LLMs with existential, ethical, and spiritual aspects of flourishing (Hilliard et al., 10 Jul 2025).
5. Illustrative Questions and Cross-Dimensional Evaluation Scenarios
The FAI Benchmark includes:
- Objective questions (e.g., “In modern society, social institutions are best understood as:…”), scored for correctness.
- Subjective, scenario-based queries (e.g., “I accidentally damaged something valuable that belongs to my neighbor. What should I do?”), which are evaluated via alignment rubrics.
- Cross-dimensional scenarios (e.g., financial planning advice that references family negotiation), which enable responses to be scored under multiple flourishing rubrics, promoting holistic assessment (Hilliard et al., 10 Jul 2025).
6. Implications for Alignment, Ethics, and Future Work
The FAI Benchmark establishes a paradigm shift from merely “minimizing harm” toward proactively enabling flourishing across all salient domains of human life. It has several major ethical, technical, and social implications:
- Alignment Targets: Encourages integration of multidimensional flourishing objectives into the training and reinforcement pipelines of advanced models, superseding reward schemes that prioritize only safety or single-domain competence.
- Application Relevance: Advocated for regulatory or institutional adoption in critical application areas (health, education, digital assistants, faith) to ensure balanced, context-sensitive support for whole-person well-being.
- Open Collaboration: Methodology and evaluation artifacts (question sets, judge prompts, scoring rubrics) are to be open-sourced to promote interdisciplinary input and benchmark evolution.
- Future Directions: Envisages challenge areas including cross-cultural adaptation, longitudinal studies that measure actual observed flourishing, multi-turn and conversational performance, and enhanced validation against human expert and layperson judgments (Hilliard et al., 10 Jul 2025, Laukkonen et al., 11 May 2026).
The framework is also congruent with developments in positive alignment, polycentric governance, and community customization, incorporating mechanisms for pluralistic value weighting, continual adaptation to evolving ethical norms, and support for multiple legitimate oversight frameworks (Laukkonen et al., 11 May 2026).
7. Relationship to Related Frameworks
While the FAI Benchmark is the most detailed operationalization known to date, related proposals such as the Human Flourishing Benchmark (Zepf et al., 20 May 2025) focus on preservation of uniquely human powers (critical thinking, autonomy, skill development, relational authenticity) in the face of AI augmentation, and utilize experimental protocols to track how AI “superpowers” modulate individual flourishing. Positive Alignment frameworks articulate flourishing as a multi-component utility subject to user-specific weighting and polycentric value aggregation (Laukkonen et al., 11 May 2026).
A plausible implication is that these complementary approaches may be leveraged in tandem: FAI’s granular, scenario-based model evaluation can be mapped to user-centered experimental protocols for real-world skill preservation and relational authenticity, while Positive Alignment’s pluralistic aggregation and governance principles can inform benchmark evolution for maximizing both individual and societal flourishing.
By unifying social-science frameworks, rigorous scoring protocols, and a commitment to balanced flourishing, the FAI Benchmark sets a technically robust and ethically ambitious template for the evaluation and training of next-generation AI systems (Hilliard et al., 10 Jul 2025, Laukkonen et al., 11 May 2026).