---
title: 'Human-MME: Human-Centric MLLM Benchmark'
url: https://www.emergentmind.com/topics/human-mme
type: topic
---

# Human-MME: Human-Centric MLLM Benchmark

Human-MME is a human-centric image benchmark for multimodal large language models (MLLMs), introduced as “a holistic evaluation benchmark for human-centric multimodal large language models.” It is designed to evaluate models “from granular human-oriented perception to the higher-dimensional multi-target and causal reasoning” across eight dimensions, using real-world images, diverse question formats, and a hybrid automatic-plus-manual annotation pipeline. The benchmark contains **16,765 high-quality images** and **19,945 real-world image-question pairs**, spans **4 primary visual domains**, **15 secondary domains**, and **43 sub-fields**, and extends single-target understanding to **multi-person** and **multi-image mutual understanding** through **choice, short-answer, grounding, ranking, and judgment** components [2509.26165].

## 1. Conceptual scope and motivation

Human-MME is motivated by a specific gap in multimodal evaluation: existing MLLM benchmarks are described as insufficient for **human-centric scene understanding** because they do not adequately cover both **granular perception** of people and **higher-level reasoning** about people. The benchmark therefore targets scenes in which the central objects of interpretation are human bodies, faces, clothing, human-object interactions, and inter-person relations, together with abstract inferences about **intention**, **cause**, and **emotion** [2509.26165].

The paper identifies three reasons for this gap. First, existing evaluation settings are described as overly simplistic for the complexity of human scenes. Second, existing benchmarks rarely combine **fine-grained human detail recognition**—such as **eyebrows, accessories, body parts, pose, and left/right distinctions**—with **higher-level spatial, social, and causal reasoning**. Third, annotation quality is difficult to maintain because human-centric evaluation requires precise labels for **face**, **body**, **hands**, **feet**, **facial parts**, **clothing items**, **interaction objects**, and **person-specific attributes** [2509.26165].

This design places Human-MME in a broader benchmark landscape but with a specialized target. The original MME benchmark introduced a comprehensive perception-and-cognition evaluation for MLLMs across **14 subtasks** using paired yes/no questions, but it did **not** introduce or define Human-MME [2306.13394]. Human-MME instead focuses specifically on **human-centered scenes**, with richer output types and a progression from low-level human perception to higher-dimensional reasoning [2509.26165].

## 2. Dataset composition and scene coverage

Human-MME is built from two sources. The raw collection contains **54,122 images** from **Pexels** and **Pixabay**, and **47,776 images** from **HICO-DET**. After filtering, annotation, and de-duplication, the final benchmark image pool contains **8,010 images** from free media sites and **8,755 images** from the open-source dataset, for a total of **16,765 high-quality images** [2509.26165].

The benchmark is organized into **4 main domains**:

1. **Daily Life**
2. **Work**
3. **Study**
4. **Entertainment**

It also contains **15 secondary domains** and **43 sub-fields**. An explicitly named example is given under **Work**, where **Office** includes **Meetings**, **Desk Work**, and **Reports** [2509.26165].

Scene diversity is treated as a primary design goal. The benchmark includes images ranging from **below 480p to over 4K**, with **over 30% of images at 4K (3840×2160) or higher**. It includes **single-person** and **multi-person** scenes, and its HOI coverage involves **hundreds of distinct objects**, with examples including **skateboards, books, laptops, knives**. The automated annotation pipeline produces **13 different types of bounding boxes**, **42 binary facial features**, **4 person-level attributes**, **8 categories of clothing**, human-object interaction annotations, and **3 types of higher-dimensional features** [2509.26165].

The benchmark’s hierarchy suggests a deliberate attempt to cover both ordinary and structured human activity. This suggests that Human-MME is not intended as a narrow face or pose benchmark, but as a scene-level benchmark in which the human figure remains the principal anchor of multimodal interpretation.

## 3. Progressive evaluation framework

Human-MME is organized around **8 evaluation dimensions**, ordered from fine-grained perception to higher-dimensional reasoning [2509.26165].

| Dimension | Abbrev. | Main focus |
|---|---|---|
| Face Understanding | FU | Facial parts and facial attributes |
| Body Understanding | BU | Body parts and clothing details |
| HOI Understanding | HU | Human-object interaction |
| Multi-Image Understanding | MIU | Cross-image comparison across four images |
| Multi-Person Reasoning | MPR | Identification and reasoning in multi-person scenes |
| Intention Discrimination | ID | Intention inference |
| Causal Discrimination | CD | Past and future scene causality |
| Emotion Discrimination | ED | Identity-neutral emotional analysis |

**Face Understanding (FU)** evaluates **single-person images** through **Face Grounding** and **Face Choice**. It targets the face box \(F_i\), facial-part boxes \(FP_i^{\text{name}}\) such as **mouth**, **nose**, **left eye**, **right eye**, **left eyebrow**, and **right eyebrow**, and facial attributes \(FA_i^{\text{name}}\), with examples such as **wavy hair**, **rosy cheeks**, **has goatee**, and **mouth slightly open** [2509.26165].

**Body Understanding (BU)** focuses on the full body, body parts, and clothing. Its question forms are **Body Grounding**, **Wearing Choice**, and **Wearing Short-Answer**. Targets include the full body box \(B_i\), body-part boxes \(P_i^{\text{name}}\) such as **left hand**, **right hand**, **left foot**, and **right foot**, and clothing items \(W_i^{(j)}\) with **type**, **color**, and **name** [2509.26165].

**HOI Understanding (HU)** tests the semantics of **Human-Object Interaction**, including the interacted object, the action, the body part involved, and the object location. Its forms are **HOI Choice**, **HOI Short-Answer**, and **HOI Grounding**. The benchmark explicitly stresses distractors involving **left hand / right hand**, object substitutions, and action confusion [2509.26165].

**Multi-Image Understanding (MIU)** uses **four images** and requires ranking or selection across them. **Multi-Face** ranks images by the number of target facial attributes present in a single person; **Multi-Wearing** does the same for clothing; **Multi-HOI** asks which of four images best matches a given interaction description [2509.26165].

**Multi-Person Reasoning (MPR)** is the richest dimension. It includes **Identify Short-Answer**, **Identify Bounding Box**, **Identify Open HOI**, **Identify Choice**, **Judgment Short-Answer**, **Judgment Bounding Box**, and **Common Choice**. Its logic is based on selecting a target person via a feature set \(\mathcal{T}_i\), then either answering about that person or abstaining when no matching person exists. For **Judgment** tasks, the required abstention outputs are explicit: **“unknown”** for J+SA and **[-1,-1,-1,-1]** for J+BB [2509.26165].

**Intention Discrimination (ID)** evaluates **identity-neutral intention** through **Intention Choice**. Distractors are drawn from **visually similar images retrieved with CLIP**, making them intentionally hard [2509.26165].

**Causal Discrimination (CD)** evaluates scene-level causality through **Causal Choice**, which requires two selections: one for **past** and one for **future**. Each image is annotated with \(C^{\text{past}}\) and \(C^{\text{future}}\), and distractors are taken from visually similar images found by CLIP [2509.26165].

**Emotion Discrimination (ED)** evaluates **identity-neutral emotional analysis**. Distractors come from visually similar images where at least one person shares the same raw emotion label \(A_i^{\text{emotion}}\), which makes the choices deliberately subtle [2509.26165].

The paper summarizes the difficulty ordering of the three abstract reasoning dimensions as
\[
\text{Intention} > \text{Cause} > \text{Emotion}
\]
where “>” denotes easier or higher accuracy [2509.26165].

## 4. Question paradigms and annotation pipeline

Human-MME supports **21 question types** and a set of answer paradigms broader than standard multiple choice. The five basic components are **Choice**, **Short-Answer**, **Bounding Box / Grounding**, **Ranking**, and **Judgment**. Composite forms include **Judgment + Short-Answer**, **Judgment + Bounding Box**, and **Short-Answer + Bounding Box**. The appendix summary lists **7 answer forms**: single choice, ranking, short-answer, bounding box, judgment + short-answer, judgment + bounding box, and short-answer + bounding box [2509.26165].

The benchmark uses structured output templates. Examples include:

- **SC**: `Analyze: <your analysis>` followed by `Answer: A/B/C/D`
- **DC**: `Analyze: <your analysis>` followed by `Past: A/B/C/D` and `Future: A/B/C/D`
- **R**: ordered outputs from `First` to `Fourth`
- **BB**: `Answer: [x1,y1,x2,y2]`
- **J+SA**: answer normally, but `if no match, answer unknown`
- **J+BB**: answer normally, but `if no match, answer [-1,-1,-1,-1]`
- **SA+BB**: separate `Name` and `Box` fields [2509.26165]

The automated annotation pipeline has **five steps** [2509.26165].

**Step 1: Face/body/body-part detection and alignment.** The pipeline uses **YOLOv11** for candidate face/body boxes and **DWPose** for whole-body pose estimation with **134 keypoints**, including **18 body keypoints**, **21 for each hand**, **68 facial keypoints**, and **3 for each foot**. The matched body box is defined as
\[
B_i = \arg\max_{B \in \mathcal{B}} \mathrm{IoU}\big(B,\ \mathrm{MBR}(K_i)\big)
\]
where \(\mathcal{B}\) is the set of body boxes from YOLOv11 and \(K_i\) is the keypoint set for person \(i\). Face matching is analogous. Part-level boxes are created for **left hand**, **right hand**, **left foot**, and **right foot** [2509.26165].

**Step 2: Person-level attributes, clothing, and HOI.** Each isolated person instance is sent to **Qwen2.5-VL-72B** using a JSON template to extract general attributes (**age group**, **gender**, **race**, **emotion**), wearing attributes (**type**, **color**, **name**), and object interactions (**object name**, **action**, **interacting body part**) [2509.26165].

**Step 3: HOI object grounding.** The original image and object names are sent to **Grounding DINO** to predict HOI object bounding boxes. If Step 2 returns “single hand,” the system refines this by comparing object proximity to left-hand versus right-hand keypoints and replacing it with the closer one [2509.26165].

**Step 4: Facial attributes and facial-part grounding.** Enlarged face crops are processed with **FaceXFormer** to obtain **40 binary facial attributes**, head pose, and **68 facial landmarks**. The pipeline yields **42 binary facial features**, comprising the 40 FaceXFormer attributes plus **2 head-pose-derived binary values**. Pitch and yaw are discretized with **±15° thresholds** into up/down and left/right image-side orientation. Facial-part boxes are derived for **nose**, **mouth**, **left eye**, **right eye**, **left eyebrow**, and **right eyebrow**, and are cross-checked against DWPose geometry [2509.26165].

**Step 5: Higher-dimensional semantics.** Each person is highlighted and queried by **Qwen2.5-VL-72B** to produce intermediate emotional analysis \(E_i^+\) and behavior/intention description \(I_i^+\). **Qwen3** is then used to remove identity-revealing details and generate **identity-neutral emotion** \(E_i\), **identity-neutral intention** \(I_i\), and scene-level \(C^{\text{past}}\) and \(C^{\text{future}}\). The paper states that this cross-model setup is intended to reduce **same-model bias** [2509.26165].

A manual **Gradio-based** annotation platform is then used for **cluster-level de-duplication** and **instance-level correction**. Experts select images to maximize **diversity**, **annotation quality**, **complexity**, successful HOI annotation, and accurate boxes, and can remove or adjust boxes, edit wearing JSON, correct HOI annotations, and accept, discard, or reselect images [2509.26165].

## 5. Evaluation protocol and metrics

Human-MME evaluates **17 state-of-the-art MLLMs**, including **15 open-source** and **2 closed-source** systems. The open-source models are deployed with **vLLM**; the closed-source models are accessed via official APIs. Format-enforcing prompts are used to ensure structured outputs, and final answers are parsed with regular expressions [2509.26165].

The evaluated models are:

- **Open-source**: GLM-4.5V, GLM-4.1V-9B, Qwen2.5-VL-72B, Qwen2.5-VL-32B, Qwen2.5-VL-7B, InternVL3-78B, InternVL3.5-38B, Intern-S1, MiniCPM-V-4.5, Gemma3-27B, LLaVA-NeXT-72B, Aya-vision-32B, Kimi-VL-A3B, Llama-4-Scout, Phi-4
- **Closed-source**: GPT-4o, Gemini-2.5-Pro [2509.26165]

Metrics are defined by task type.

For **Choice** questions, the metric is **Accuracy**:
\[
\text{Accuracy} = \frac{\text{Number of correct selections}}{\text{Total number of questions}}
\]

For **Short-Answer** questions, the benchmark uses **BERT F1**, **Cosine Similarity**, and **Keyword Coverage**:
\[
\text{BERT F1} = \frac{2 \cdot P_\text{BERT} \cdot R_\text{BERT}}{P_\text{BERT} + R_\text{BERT}}
\]
\[
\text{CosineSim} = \frac{\mathbf{v}_\text{pred} \cdot \mathbf{v}_\text{gt}}{\|\mathbf{v}_\text{pred}\| \, \|\mathbf{v}_\text{gt}\|}
\]
\[
\text{KeywordCoverage} = \frac{|\text{Keywords}_{\text{pred}} \cap \text{Keywords}_{\text{gt}}|}{|\text{Keywords}_{\text{gt}}|}
\]
with the composite score
\[
\text{Composite Score} = 0.5 \cdot \text{BERT F1} + 0.3 \cdot \text{CosineSim} + 0.2 \cdot \text{KeywordCoverage}
\]

For **Ranking** questions, the metric is **Kendall’s Tau**:
\[
\tau = \frac{C - D}{\frac{1}{2} n(n-1)}
\]
where \(C\) and \(D\) are concordant and discordant pairs.

For **Bounding Box** questions, the metric is **IoU**:
\[
\text{IoU} = \frac{\text{Area}(B_p \cap B_{gt})}{\text{Area}(B_p \cup B_{gt})}
\]

For **Judgment** questions, the metric is **F1**:
\[
\text{F1} = \frac{2 \cdot \text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}
\]

The paper reports results by **dimension**, by **question component**, and as an **average** across dimensions. It does not give an explicit formal equation for the final aggregate average in the excerpted formulation, although the tables report “Avg.” over the eight dimensions [2509.26165].

## 6. Empirical findings, failure modes, and benchmark position

Human-MME reports a clear separation among model families. The best open-source overall model is **GLM-4.5V** with an overall average of **76.0**, followed by **Qwen2.5-VL-72B** at **72.8**. Among closed-source systems, **Gemini-2.5-Pro** scores **68.9**, outperforming **GPT-4o** at **59.0**. The weakest overall model is **Phi-4** at **42.9** [2509.26165].

At the dimension level, **GLM-4.5V** leads most perception-heavy dimensions, scoring **61.6** on **FU**, **77.4** on **BU**, **82.5** on **HU**, and **71.5** on **MPR**. **Qwen2.5-VL-72B** is strongest on high-level reasoning dimensions, with **88.1** on **Intention Discrimination** and **86.3** on **Causal Discrimination**. **Intern-S1** is notable for strong performance on **MIU** at **79.8** and **ED** at **68.3** among open-source systems, despite weaker grounding. **Gemini-2.5-Pro** performs strongly on **MIU** (**83.6**), **CD** (**86.1**), **short-answer** (**83.9**), **ranking** (**90.9**), and **judgment** (**72.0**), but is weaker on the **bounding-box** component at **23.5** [2509.26165].

By task component, **GLM-4.5V** leads **Bounding Box** with **66.3** and **Short-Answer** with **83.5**. **Qwen2.5-VL-72B** has the best **Judgment** score at **71.3**. **Intern-S1** and **InternVL3-78B** perform well on **choice**, **short-answer**, and **ranking**, but poorly on bounding box tasks. This suggests that explicit grounding ability depends more on grounding-oriented training and output alignment than on general model size alone [2509.26165].

The paper summarizes six benchmark-level findings [2509.26165]:

1. **Stronger scaling effects in Choice and Ranking tasks**: these correlate more strongly with model size than other metrics.
2. **Grounding depends more on training data than scale**: bounding-box performance depends heavily on grounding-specific training data and output alignment format.
3. **Left-right body-part discrimination is hard**: all evaluated models struggle with **left vs right hands** and **left vs right feet**, while performing much better on left-right facial parts.
4. **Judgment tasks expose hallucination**: models typically show **high recall** but **lower precision**, often answering when they should abstain.
5. **Adding Judgment makes tasks harder**: Judgment versions reduce performance relative to corresponding Identify tasks.
6. **Difficulty hierarchy in abstract reasoning**: **Intention** is easiest, **Emotion** hardest.

Several qualitative failure modes recur. The paper highlights **left-right confusion for hands and feet**, difficulty matching multiple person-specific conditions in **MPR**, hallucination in **Judgment** tasks, and grounding errors that remain severe for models without explicit grounding training. It also notes a specific failure mode in which **Qwen2.5-VL-72B** sometimes refuses to output a full body box when the body is partially occluded, even though the instruction is to annotate **“all visible body”** [2509.26165].

Within the benchmark comparison presented by Human-MME, the benchmark is described as the only one among **MME**, **Seed-Bench**, **HV-MMBench**, **HumanVBench**, **Face-Human-Bench**, **HumaniBench**, and **Human-MME** that covers **fine-grained grounding**, **face features**, **body features**, **human-object interaction**, **multi-image understanding**, **multi-person understanding**, and **high-level abstract features** simultaneously [2509.26165]. In that comparison, the original **MME** is characterized as an image benchmark with **2.8K QA** and **true/false** format, covering body features and multi-person settings but lacking the broader set of human-centric capabilities emphasized by Human-MME [2509.26165]. The original MME paper itself defines only MME and does not define Human-MME [2306.13394].

Human-MME does not present a separate limitations section in the summarized material, but the paper implicitly acknowledges several constraints: fine-grained human annotation is difficult, automatic annotation can be noisy and requires expert correction, **same-model bias** is a concern in higher-level text generation, and high-level tasks such as **emotion** remain difficult and potentially subjective [2509.26165]. A plausible implication is that Human-MME functions less as a saturated leaderboard benchmark than as a diagnostic instrument for exposing where current MLLMs remain weak in **human grounding**, **person-specific disambiguation**, **abstention behavior**, and **socially grounded reasoning**.

Source: https://www.emergentmind.com/topics/human-mme