Papers
Topics
Authors
Recent
Search
2000 character limit reached

Human-MME: Human-Centric MLLM Benchmark

Updated 14 July 2026
  • Human-MME is a human-centric image benchmark designed to evaluate multimodal large language models across eight dimensions from fine-grained perception to higher-level reasoning.
  • It encompasses detailed annotations using 16,765 high-quality images across diverse visual domains, focusing on facial, body, and human-object interactions.
  • The benchmark employs a hybrid automatic-plus-manual annotation pipeline, facilitating rigorous evaluation through metrics like accuracy, IoU, and F1 scores.

Human-MME is a human-centric image benchmark for multimodal LLMs (MLLMs), introduced as “a holistic evaluation benchmark for human-centric multimodal LLMs.” It is designed to evaluate models “from granular human-oriented perception to the higher-dimensional multi-target and causal reasoning” across eight dimensions, using real-world images, diverse question formats, and a hybrid automatic-plus-manual annotation pipeline. The benchmark contains 16,765 high-quality images and 19,945 real-world image-question pairs, spans 4 primary visual domains, 15 secondary domains, and 43 sub-fields, and extends single-target understanding to multi-person and multi-image mutual understanding through choice, short-answer, grounding, ranking, and judgment components (Liu et al., 30 Sep 2025).

1. Conceptual scope and motivation

Human-MME is motivated by a specific gap in multimodal evaluation: existing MLLM benchmarks are described as insufficient for human-centric scene understanding because they do not adequately cover both granular perception of people and higher-level reasoning about people. The benchmark therefore targets scenes in which the central objects of interpretation are human bodies, faces, clothing, human-object interactions, and inter-person relations, together with abstract inferences about intention, cause, and emotion (Liu et al., 30 Sep 2025).

The paper identifies three reasons for this gap. First, existing evaluation settings are described as overly simplistic for the complexity of human scenes. Second, existing benchmarks rarely combine fine-grained human detail recognition—such as eyebrows, accessories, body parts, pose, and left/right distinctions—with higher-level spatial, social, and causal reasoning. Third, annotation quality is difficult to maintain because human-centric evaluation requires precise labels for face, body, hands, feet, facial parts, clothing items, interaction objects, and person-specific attributes (Liu et al., 30 Sep 2025).

This design places Human-MME in a broader benchmark landscape but with a specialized target. The original MME benchmark introduced a comprehensive perception-and-cognition evaluation for MLLMs across 14 subtasks using paired yes/no questions, but it did not introduce or define Human-MME (Fu et al., 2023). Human-MME instead focuses specifically on human-centered scenes, with richer output types and a progression from low-level human perception to higher-dimensional reasoning (Liu et al., 30 Sep 2025).

2. Dataset composition and scene coverage

Human-MME is built from two sources. The raw collection contains 54,122 images from Pexels and Pixabay, and 47,776 images from HICO-DET. After filtering, annotation, and de-duplication, the final benchmark image pool contains 8,010 images from free media sites and 8,755 images from the open-source dataset, for a total of 16,765 high-quality images (Liu et al., 30 Sep 2025).

The benchmark is organized into 4 main domains:

  1. Daily Life
  2. Work
  3. Study
  4. Entertainment

It also contains 15 secondary domains and 43 sub-fields. An explicitly named example is given under Work, where Office includes Meetings, Desk Work, and Reports (Liu et al., 30 Sep 2025).

Scene diversity is treated as a primary design goal. The benchmark includes images ranging from below 480p to over 4K, with over 30% of images at 4K (3840×2160) or higher. It includes single-person and multi-person scenes, and its HOI coverage involves hundreds of distinct objects, with examples including skateboards, books, laptops, knives. The automated annotation pipeline produces 13 different types of bounding boxes, 42 binary facial features, 4 person-level attributes, 8 categories of clothing, human-object interaction annotations, and 3 types of higher-dimensional features (Liu et al., 30 Sep 2025).

The benchmark’s hierarchy suggests a deliberate attempt to cover both ordinary and structured human activity. This suggests that Human-MME is not intended as a narrow face or pose benchmark, but as a scene-level benchmark in which the human figure remains the principal anchor of multimodal interpretation.

3. Progressive evaluation framework

Human-MME is organized around 8 evaluation dimensions, ordered from fine-grained perception to higher-dimensional reasoning (Liu et al., 30 Sep 2025).

Dimension Abbrev. Main focus
Face Understanding FU Facial parts and facial attributes
Body Understanding BU Body parts and clothing details
HOI Understanding HU Human-object interaction
Multi-Image Understanding MIU Cross-image comparison across four images
Multi-Person Reasoning MPR Identification and reasoning in multi-person scenes
Intention Discrimination ID Intention inference
Causal Discrimination CD Past and future scene causality
Emotion Discrimination ED Identity-neutral emotional analysis

Face Understanding (FU) evaluates single-person images through Face Grounding and Face Choice. It targets the face box FiF_i, facial-part boxes FPinameFP_i^{\text{name}} such as mouth, nose, left eye, right eye, left eyebrow, and right eyebrow, and facial attributes FAinameFA_i^{\text{name}}, with examples such as wavy hair, rosy cheeks, has goatee, and mouth slightly open (Liu et al., 30 Sep 2025).

Body Understanding (BU) focuses on the full body, body parts, and clothing. Its question forms are Body Grounding, Wearing Choice, and Wearing Short-Answer. Targets include the full body box BiB_i, body-part boxes PinameP_i^{\text{name}} such as left hand, right hand, left foot, and right foot, and clothing items Wi(j)W_i^{(j)} with type, color, and name (Liu et al., 30 Sep 2025).

HOI Understanding (HU) tests the semantics of Human-Object Interaction, including the interacted object, the action, the body part involved, and the object location. Its forms are HOI Choice, HOI Short-Answer, and HOI Grounding. The benchmark explicitly stresses distractors involving left hand / right hand, object substitutions, and action confusion (Liu et al., 30 Sep 2025).

Multi-Image Understanding (MIU) uses four images and requires ranking or selection across them. Multi-Face ranks images by the number of target facial attributes present in a single person; Multi-Wearing does the same for clothing; Multi-HOI asks which of four images best matches a given interaction description (Liu et al., 30 Sep 2025).

Multi-Person Reasoning (MPR) is the richest dimension. It includes Identify Short-Answer, Identify Bounding Box, Identify Open HOI, Identify Choice, Judgment Short-Answer, Judgment Bounding Box, and Common Choice. Its logic is based on selecting a target person via a feature set Ti\mathcal{T}_i, then either answering about that person or abstaining when no matching person exists. For Judgment tasks, the required abstention outputs are explicit: “unknown” for J+SA and [-1,-1,-1,-1] for J+BB (Liu et al., 30 Sep 2025).

Intention Discrimination (ID) evaluates identity-neutral intention through Intention Choice. Distractors are drawn from visually similar images retrieved with CLIP, making them intentionally hard (Liu et al., 30 Sep 2025).

Causal Discrimination (CD) evaluates scene-level causality through Causal Choice, which requires two selections: one for past and one for future. Each image is annotated with CpastC^{\text{past}} and CfutureC^{\text{future}}, and distractors are taken from visually similar images found by CLIP (Liu et al., 30 Sep 2025).

Emotion Discrimination (ED) evaluates identity-neutral emotional analysis. Distractors come from visually similar images where at least one person shares the same raw emotion label AiemotionA_i^{\text{emotion}}, which makes the choices deliberately subtle (Liu et al., 30 Sep 2025).

The paper summarizes the difficulty ordering of the three abstract reasoning dimensions as

FPinameFP_i^{\text{name}}0

where “>” denotes easier or higher accuracy (Liu et al., 30 Sep 2025).

4. Question paradigms and annotation pipeline

Human-MME supports 21 question types and a set of answer paradigms broader than standard multiple choice. The five basic components are Choice, Short-Answer, Bounding Box / Grounding, Ranking, and Judgment. Composite forms include Judgment + Short-Answer, Judgment + Bounding Box, and Short-Answer + Bounding Box. The appendix summary lists 7 answer forms: single choice, ranking, short-answer, bounding box, judgment + short-answer, judgment + bounding box, and short-answer + bounding box (Liu et al., 30 Sep 2025).

The benchmark uses structured output templates. Examples include:

  • SC: Analyze: <your analysis> followed by Answer: A/B/C/D
  • DC: Analyze: <your analysis> followed by Past: A/B/C/D and Future: A/B/C/D
  • R: ordered outputs from First to Fourth
  • BB: Answer: [x1,y1,x2,y2]
  • J+SA: answer normally, but if no match, answer unknown
  • J+BB: answer normally, but if no match, answer [-1,-1,-1,-1]
  • SA+BB: separate Name and Box fields (Liu et al., 30 Sep 2025)

The automated annotation pipeline has five steps (Liu et al., 30 Sep 2025).

Step 1: Face/body/body-part detection and alignment. The pipeline uses YOLOv11 for candidate face/body boxes and DWPose for whole-body pose estimation with 134 keypoints, including 18 body keypoints, 21 for each hand, 68 facial keypoints, and 3 for each foot. The matched body box is defined as

FPinameFP_i^{\text{name}}1

where FPinameFP_i^{\text{name}}2 is the set of body boxes from YOLOv11 and FPinameFP_i^{\text{name}}3 is the keypoint set for person FPinameFP_i^{\text{name}}4. Face matching is analogous. Part-level boxes are created for left hand, right hand, left foot, and right foot (Liu et al., 30 Sep 2025).

Step 2: Person-level attributes, clothing, and HOI. Each isolated person instance is sent to Qwen2.5-VL-72B using a JSON template to extract general attributes (age group, gender, race, emotion), wearing attributes (type, color, name), and object interactions (object name, action, interacting body part) (Liu et al., 30 Sep 2025).

Step 3: HOI object grounding. The original image and object names are sent to Grounding DINO to predict HOI object bounding boxes. If Step 2 returns “single hand,” the system refines this by comparing object proximity to left-hand versus right-hand keypoints and replacing it with the closer one (Liu et al., 30 Sep 2025).

Step 4: Facial attributes and facial-part grounding. Enlarged face crops are processed with FaceXFormer to obtain 40 binary facial attributes, head pose, and 68 facial landmarks. The pipeline yields 42 binary facial features, comprising the 40 FaceXFormer attributes plus 2 head-pose-derived binary values. Pitch and yaw are discretized with ±15° thresholds into up/down and left/right image-side orientation. Facial-part boxes are derived for nose, mouth, left eye, right eye, left eyebrow, and right eyebrow, and are cross-checked against DWPose geometry (Liu et al., 30 Sep 2025).

Step 5: Higher-dimensional semantics. Each person is highlighted and queried by Qwen2.5-VL-72B to produce intermediate emotional analysis FPinameFP_i^{\text{name}}5 and behavior/intention description FPinameFP_i^{\text{name}}6. Qwen3 is then used to remove identity-revealing details and generate identity-neutral emotion FPinameFP_i^{\text{name}}7, identity-neutral intention FPinameFP_i^{\text{name}}8, and scene-level FPinameFP_i^{\text{name}}9 and FAinameFA_i^{\text{name}}0. The paper states that this cross-model setup is intended to reduce same-model bias (Liu et al., 30 Sep 2025).

A manual Gradio-based annotation platform is then used for cluster-level de-duplication and instance-level correction. Experts select images to maximize diversity, annotation quality, complexity, successful HOI annotation, and accurate boxes, and can remove or adjust boxes, edit wearing JSON, correct HOI annotations, and accept, discard, or reselect images (Liu et al., 30 Sep 2025).

5. Evaluation protocol and metrics

Human-MME evaluates 17 state-of-the-art MLLMs, including 15 open-source and 2 closed-source systems. The open-source models are deployed with vLLM; the closed-source models are accessed via official APIs. Format-enforcing prompts are used to ensure structured outputs, and final answers are parsed with regular expressions (Liu et al., 30 Sep 2025).

The evaluated models are:

  • Open-source: GLM-4.5V, GLM-4.1V-9B, Qwen2.5-VL-72B, Qwen2.5-VL-32B, Qwen2.5-VL-7B, InternVL3-78B, InternVL3.5-38B, Intern-S1, MiniCPM-V-4.5, Gemma3-27B, LLaVA-NeXT-72B, Aya-vision-32B, Kimi-VL-A3B, Llama-4-Scout, Phi-4
  • Closed-source: GPT-4o, Gemini-2.5-Pro (Liu et al., 30 Sep 2025)

Metrics are defined by task type.

For Choice questions, the metric is Accuracy: FAinameFA_i^{\text{name}}1

For Short-Answer questions, the benchmark uses BERT F1, Cosine Similarity, and Keyword Coverage: FAinameFA_i^{\text{name}}2

FAinameFA_i^{\text{name}}3

FAinameFA_i^{\text{name}}4

with the composite score

FAinameFA_i^{\text{name}}5

For Ranking questions, the metric is Kendall’s Tau: FAinameFA_i^{\text{name}}6 where FAinameFA_i^{\text{name}}7 and FAinameFA_i^{\text{name}}8 are concordant and discordant pairs.

For Bounding Box questions, the metric is IoU: FAinameFA_i^{\text{name}}9

For Judgment questions, the metric is F1: BiB_i0

The paper reports results by dimension, by question component, and as an average across dimensions. It does not give an explicit formal equation for the final aggregate average in the excerpted formulation, although the tables report “Avg.” over the eight dimensions (Liu et al., 30 Sep 2025).

6. Empirical findings, failure modes, and benchmark position

Human-MME reports a clear separation among model families. The best open-source overall model is GLM-4.5V with an overall average of 76.0, followed by Qwen2.5-VL-72B at 72.8. Among closed-source systems, Gemini-2.5-Pro scores 68.9, outperforming GPT-4o at 59.0. The weakest overall model is Phi-4 at 42.9 (Liu et al., 30 Sep 2025).

At the dimension level, GLM-4.5V leads most perception-heavy dimensions, scoring 61.6 on FU, 77.4 on BU, 82.5 on HU, and 71.5 on MPR. Qwen2.5-VL-72B is strongest on high-level reasoning dimensions, with 88.1 on Intention Discrimination and 86.3 on Causal Discrimination. Intern-S1 is notable for strong performance on MIU at 79.8 and ED at 68.3 among open-source systems, despite weaker grounding. Gemini-2.5-Pro performs strongly on MIU (83.6), CD (86.1), short-answer (83.9), ranking (90.9), and judgment (72.0), but is weaker on the bounding-box component at 23.5 (Liu et al., 30 Sep 2025).

By task component, GLM-4.5V leads Bounding Box with 66.3 and Short-Answer with 83.5. Qwen2.5-VL-72B has the best Judgment score at 71.3. Intern-S1 and InternVL3-78B perform well on choice, short-answer, and ranking, but poorly on bounding box tasks. This suggests that explicit grounding ability depends more on grounding-oriented training and output alignment than on general model size alone (Liu et al., 30 Sep 2025).

The paper summarizes six benchmark-level findings (Liu et al., 30 Sep 2025):

  1. Stronger scaling effects in Choice and Ranking tasks: these correlate more strongly with model size than other metrics.
  2. Grounding depends more on training data than scale: bounding-box performance depends heavily on grounding-specific training data and output alignment format.
  3. Left-right body-part discrimination is hard: all evaluated models struggle with left vs right hands and left vs right feet, while performing much better on left-right facial parts.
  4. Judgment tasks expose hallucination: models typically show high recall but lower precision, often answering when they should abstain.
  5. Adding Judgment makes tasks harder: Judgment versions reduce performance relative to corresponding Identify tasks.
  6. Difficulty hierarchy in abstract reasoning: Intention is easiest, Emotion hardest.

Several qualitative failure modes recur. The paper highlights left-right confusion for hands and feet, difficulty matching multiple person-specific conditions in MPR, hallucination in Judgment tasks, and grounding errors that remain severe for models without explicit grounding training. It also notes a specific failure mode in which Qwen2.5-VL-72B sometimes refuses to output a full body box when the body is partially occluded, even though the instruction is to annotate “all visible body” (Liu et al., 30 Sep 2025).

Within the benchmark comparison presented by Human-MME, the benchmark is described as the only one among MME, Seed-Bench, HV-MMBench, HumanVBench, Face-Human-Bench, HumaniBench, and Human-MME that covers fine-grained grounding, face features, body features, human-object interaction, multi-image understanding, multi-person understanding, and high-level abstract features simultaneously (Liu et al., 30 Sep 2025). In that comparison, the original MME is characterized as an image benchmark with 2.8K QA and true/false format, covering body features and multi-person settings but lacking the broader set of human-centric capabilities emphasized by Human-MME (Liu et al., 30 Sep 2025). The original MME paper itself defines only MME and does not define Human-MME (Fu et al., 2023).

Human-MME does not present a separate limitations section in the summarized material, but the paper implicitly acknowledges several constraints: fine-grained human annotation is difficult, automatic annotation can be noisy and requires expert correction, same-model bias is a concern in higher-level text generation, and high-level tasks such as emotion remain difficult and potentially subjective (Liu et al., 30 Sep 2025). A plausible implication is that Human-MME functions less as a saturated leaderboard benchmark than as a diagnostic instrument for exposing where current MLLMs remain weak in human grounding, person-specific disambiguation, abstention behavior, and socially grounded reasoning.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Human-MME.