Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

Published 17 Jul 2025 in cs.AI and cs.CL | (2507.13302v1)

Abstract: The evaluation of LLMs is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of different topics. However, this method has certain limitations, being the most concerning, the poor correlation with the humans. An alternative approach, is to have humans evaluate the LLMs. This poses scalability issues as there is a large and growing number of models to evaluate making it impractical (and costly) to run traditional studies based on recruiting a number of evaluators and having them rank the responses of the models. An alternative approach is the use of public arenas, such as the popular LM arena, on which any user can freely evaluate models on any question and rank the responses of two models. The results are then elaborated into a model ranking. An increasingly important aspect of LLMs is their energy consumption and, therefore, evaluating how energy awareness influences the decisions of humans in selecting a model is of interest. In this paper, we present GEA, the Generative Energy Arena, an arena that incorporates information on the energy consumption of the model in the evaluation process. Preliminary results obtained with GEA are also presented, showing that for most questions, when users are aware of the energy consumption, they favor smaller and more energy efficient models. This suggests that for most user interactions, the extra cost and energy incurred by the more complex and top-performing models do not provide an increase in the perceived quality of the responses that justifies their use.

Summary

  • The paper presents a two-step evaluation protocol that first assesses blind quality and then incorporates energy awareness to shift user preferences.
  • It shows that disclosing energy consumption leads to a 46% change in choices and a 75% preference for smaller, energy-efficient models.
  • The study underscores the need for sustainable AI practices by integrating energy metrics into evaluation protocols for real-world LLM deployment.

The Generative Energy Arena (GEA): Integrating Energy Awareness into LLM Human Evaluation

The paper introduces the Generative Energy Arena (GEA), a novel framework for evaluating LLMs that explicitly incorporates energy consumption information into the human evaluation process. The motivation stems from the increasing computational and environmental costs associated with LLM deployment, and the observation that traditional evaluation methodologies—whether automated benchmarks, LLM-as-a-judge, or human preference arenas—do not account for energy efficiency as a factor in model selection.

Methodological Contributions

GEA is designed as an open, web-based arena (implemented on Hugging Face) where users compare responses from pairs of LLMs within the same model family but of different sizes. The key innovation is a two-step evaluation protocol:

  1. Blind Quality Assessment: Users are first presented with two anonymized model responses and asked to select the preferred answer based solely on quality.
  2. Energy-Aware Reconsideration: If the user initially selects the higher-energy model, they are then informed that the alternative response was generated with lower energy consumption and asked if they would reconsider their choice, assuming a potential loss in quality.

This protocol is intended to decouple initial quality judgments from energy considerations, thereby enabling a direct measurement of the impact of energy awareness on user preferences.

Energy consumption data is presented in relative terms (e.g., "Model A uses more energy than Model B"), avoiding the need for absolute measurements, which are often unavailable or hardware-dependent. Only models from the same family are compared to control for confounding factors such as training data and architecture.

Experimental Setup and Results

The GEA was deployed in the context of a MOOC, with the majority of participants being students with some familiarity with LLMs. The evaluation involved 694 questions, with a significant portion designed to probe both general and technical LLM capabilities.

Key findings include:

  • Substantial Impact of Energy Awareness: On average, 46% of users changed their initial preference after being informed of the energy differential, with the range across model families spanning 41% to 52%.
  • Preference Shift Toward Smaller Models: After energy information was disclosed, smaller, more energy-efficient models were selected over their larger counterparts in more than 75% of cases, despite initial win rates being nearly balanced.
  • Model Family Variability: While some families (e.g., Llama3) showed an initial preference for larger models, the energy-aware step consistently shifted user choices toward smaller models across all families.

These results suggest that, for the majority of user interactions, the marginal quality improvements offered by larger, more energy-intensive models do not justify their increased resource consumption from the perspective of informed users.

Implications

The findings have several practical and theoretical implications:

  • Evaluation Methodology: The study demonstrates that energy consumption is a salient factor in human model selection and should be integrated into future LLM evaluation protocols, especially as sustainability becomes a central concern in AI deployment.
  • Model Development and Deployment: The results indicate that smaller models may suffice for a wide range of user queries, with larger models reserved for specialized or high-stakes tasks. This has direct implications for resource allocation, cost optimization, and environmental impact in real-world LLM serving.
  • User-Centric AI: The two-step protocol provides a template for incorporating additional contextual factors (e.g., latency, privacy) into human-in-the-loop evaluation, enabling more nuanced trade-off analyses.

Limitations and Future Directions

The study acknowledges several limitations:

  • Sample Size and Diversity: The user base was limited, predominantly composed of Spanish-speaking MOOC students, and the number of evaluated questions and models was modest.
  • Language and Task Coverage: Only Spanish-language queries were analyzed, and no breakdown by question type was performed.
  • Model Scope: The arena was restricted to a handful of model families from major providers.

Future work should address these limitations by scaling up the number and diversity of users, expanding to additional languages and model families, and conducting fine-grained analyses by task type. There is also scope for integrating more granular energy metrics and exploring the interplay between energy, latency, and other operational constraints.

Conclusion

The Generative Energy Arena provides empirical evidence that energy awareness significantly influences human preferences in LLM evaluation. The results challenge the prevailing assumption that larger models are always preferable and highlight the importance of sustainability considerations in both model development and evaluation. The GEA framework offers a practical, extensible approach for integrating energy efficiency into the broader discourse on responsible AI deployment, with potential to inform both research and industry best practices.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 2 tweets with 7 likes about this paper.