Papers
Topics
Authors
Recent
Search
2000 character limit reached

MMTB: Multimedia Terminal Agent Benchmark

Updated 5 July 2026
  • MMTB is a benchmark that evaluates terminal agents on multimedia workflows by integrating audio and video perception with terminal operations.
  • It employs a controlled environment with fixed tools and interfaces, simulating real-world tasks from media production to compliance analysis.
  • Empirical findings indicate that native multimodal access improves performance, while proxy workflows lead to higher execution costs and lower success rates.

MultiMedia-TerminalBench (MMTB) is a benchmark for evaluating terminal agents on multimedia-file workflows: tasks in which an agent must inspect, process, and act on audio and video files stored in a terminal workspace, then produce a verifiable artifact (Heo et al., 8 May 2026). It was introduced to target a gap between audio-visual benchmarks, which mainly test perception and reasoning, and computer-use or terminal benchmarks, which primarily evaluate action and artifact creation over text, code, or structured files. In MMTB, multimedia is not merely contextual input for question answering; it is the object of work. The benchmark is paired with Terminus-MM, a multimedia harness that extends Terminus-KIRA with audio and video perception, enabling controlled study of how modality access affects media-grounded terminal performance (Heo et al., 8 May 2026).

1. Definition, motivation, and problem setting

MMTB is defined around media-grounded terminal work. The benchmark tests whether an agent can perceive relevant evidence in audio and video files, translate that evidence into terminal actions, and produce a checkable output artifact (Heo et al., 8 May 2026). This differs from settings in which a model is asked only to describe or answer questions about media. The terminal agent must instead complete a filesystem workflow under a specified interface and satisfy an artifact evaluator.

The motivation is explicitly practical. The benchmark is grounded in workflows such as subtitle and localization work, broadcast or social clip production, music or acting feedback, meeting or compliance analysis, and dataset annotation (Heo et al., 8 May 2026). These are presented as workflows that require decisions to be grounded in auditory and visual evidence across files and then executed through terminal operations. A plausible implication is that MMTB evaluates the coupling of perception and action more directly than benchmarks whose outputs are only textual.

The paper situates MMTB at an intersection that earlier benchmark families do not fully cover. Audio-visual benchmarks are described as testing perception and reasoning, while terminal benchmarks test action and artifact creation; MMTB fills the missing intersection by requiring content-aware work over multimedia files in a terminal environment (Heo et al., 8 May 2026). The paper emphasizes three distinguishing properties: audio and video files are central rather than incidental, the agent must inspect media content rather than rely on metadata alone, and success is defined by completion of a persistent file workflow rather than generation of a text answer (Heo et al., 8 May 2026).

2. Benchmark scope, taxonomy, and task construction

MMTB contains 105 tasks across 5 meta-categories, 16 fine-grained categories, 536 media files, and 6 h 54 min of timed audio and video content; the median per-task duration is 1 m 20 s (Heo et al., 8 May 2026). The five meta-categories are Media Production, Performance / Coaching, Enterprise / Compliance, Personal / Education, and Operations / Research (Heo et al., 8 May 2026). Fine-grained categories include Broadcast/Film Production, Subtitling/Localization, Game QA/Esports, Audio Engineering/Podcast Production, Music Coaching, Acting/Casting, Corporate Workflows/Meetings, Lecture Content, and Dataset/ML Annotation (Heo et al., 8 May 2026).

The benchmark media include audio files, video files, image files, and some supporting static files such as PDFs or text or metadata files (Heo et al., 8 May 2026). Across all tasks, the reported average media volume is about 5.10 files per task, comprising 0.75 image files, 1.94 video files, and 1.89 audio files, with average timed duration 3 m 57 s per task (Heo et al., 8 May 2026). These statistics indicate that many tasks involve mixed-media workspaces rather than isolated single-file inputs.

Task construction is based on public sources reflecting paid practitioner workflows, described as predominantly Upwork and Fiverr gig listings, with some casting calls, practitioner forums, and industry standards documents (Heo et al., 8 May 2026). The construction pipeline began with 163 candidate scenarios, scaffolded them into Harbor tasks, populated them with license-compatible external media or controlled synthetic or derivative assets, and filtered them through automated checks, baseline checks, and manual review before releasing 105 tasks (Heo et al., 8 May 2026). Each task records provenance in media.toml, including source descriptions, licensing, and hashes (Heo et al., 8 May 2026). This suggests an emphasis on traceability and reproducibility at the level of task assets as well as evaluation.

The benchmark is annotated with multi-label capability tags rather than a partition. These tags include Reference Resolution, Spatial Reasoning, Cross-File Comparison, Temporal Localization, Audio-Visual Alignment, Music Understanding, Speech Prosody, Speaker/Voice Identity, Non-Speech Audio, On-Screen Text, Speech Understanding, and Visual Perception (Heo et al., 8 May 2026). In aggregate analyses, the paper also distinguishes perception-heavy tasks, which require identifying cues in media, from reasoning-heavy tasks, which require inferring the correct edit, annotation, or workflow from those cues (Heo et al., 8 May 2026).

Aspect Reported value Notes
Tasks 105 Released benchmark tasks
Meta-categories 5 With 16 fine-grained categories
Media files 536 Audio, video, image, and supporting files
Timed audio/video content 6 h 54 min Corpus-level total
Median per-task duration 1 m 20 s Task-level statistic

3. Task format, execution environment, and evaluation protocol

MMTB uses the Harbor format from Terminal-Bench. Each task has five components: instruction, workspace, allowed terminal or tool interface, output schema, and artifact evaluator (Heo et al., 8 May 2026). The instruction specifies the user goal without giving away the answer; the workspace is a containerized filesystem containing media and optional supporting files; the interface constrains available terminal tools; the output schema specifies the required artifact path and format; and the artifact evaluator scores the produced artifact (Heo et al., 8 May 2026).

The benchmark evaluates only the final state. A task is successful only if the final artifact at the required path satisfies the evaluator, and evaluation is based on the final workspace state rather than the model’s reasoning or command trace (Heo et al., 8 May 2026). Expected outputs vary by task and include selected file names, timestamps or intervals, structured JSON or CSV, edit lists, and processed media artifacts such as edited clips (Heo et al., 8 May 2026). This final-state criterion is central to the benchmark’s design: it treats terminal agency as successful artifact construction, not merely plausible intermediate reasoning.

All systems are evaluated in the same Harbor-style environment with fixed workspace, fixed tools, fixed evaluator, fixed logging, and a 10-minute interaction budget (Heo et al., 8 May 2026). Agents may use ordinary terminal tools such as ffmpeg, ffprobe, ASR/transcription, OCR, silence detection, and signal-processing scripts (Heo et al., 8 May 2026). The benchmark does not forbid proxy workflows; instead, it measures whether the agent can succeed either through native perception or by reconstructing evidence through terminal tools (Heo et al., 8 May 2026). This design makes it possible to compare direct multimodal access against terminal-only reduction pipelines under a common evaluation scaffold.

The benchmark defines a task-level verifier score in which task TiT_i receives a partial score si=Vi(yi;A,Ti)s_i = V_i(y_i; A, T_i), where si[0,1]s_i \in [0,1], and evaluation depends only on the final workspace state yiy_i, not the trajectory (Heo et al., 8 May 2026). It then reports binary and partial aggregate metrics:

Binary(A;T)=1Ni=1NI[siτi],Partial(A;T)=1Ni=1Nsi,Binary(A;\mathcal{T}) = \frac{1}{N}\sum_{i=1}^{N} \mathbb{I}\left[s_i \geq \tau_i\right], \qquad Partial(A;\mathcal{T}) = \frac{1}{N}\sum_{i=1}^{N} s_i,

where τi\tau_i is the task-specific acceptance threshold (Heo et al., 8 May 2026). The paper also reports mean API cost per task and mean agent execution time per task (Heo et al., 8 May 2026).

4. Terminus-MM and workspace-aware multimedia access

MMTB is paired with Terminus-MM, a multimedia terminal-agent harness that extends Terminus-KIRA, which itself extends the minimal terminal harness Terminus-2 by adding native image access (Heo et al., 8 May 2026). Terminus-MM adds native audio and video perception on top of that. The harness family is reported as Terminus-2 (text-only shell), Terminus-KIRA (text + image), Terminus-A (text + audio), Terminus-IV (text + image + video), Terminus-IA (text + image + audio), and Terminus-MM (text + image + audio + video) (Heo et al., 8 May 2026).

A distinctive design choice is workspace-aware tool routing. Terminus-MM scans the workspace, checks file extensions, and exposes only the perception tools relevant to the media actually present (Heo et al., 8 May 2026). The routing logic maps .wav, .mp3, .ogg, .flac, .aac, .m4a to audio; .mp4, .webm, .avi, .mov, .mkv to video; and .png, .jpg, .jpeg, .gif, .webp to image (Heo et al., 8 May 2026). It then retains execute_commands, task_complete, and the relevant modality tools (Heo et al., 8 May 2026). The paper also notes a specific rule: view_image is included whenever any media file is present, so frames or spectrograms produced by shell tools still have a visual inspection path (Heo et al., 8 May 2026).

The benchmark compares Terminus-MM with a version without modality masking in which all perception tools are always exposed (Heo et al., 8 May 2026). The routed version performs better, because irrelevant tools can induce redundant verification over absent modalities (Heo et al., 8 May 2026). The paper attributes the improvement to reduction of wasted inspection, repeated loops, budget exhaustion, and imprecise outputs. A plausible implication is that the benchmark is not only evaluating model capability but also the quality of tool-interface design for terminal agents operating over heterogeneous media.

Harness Modalities
Terminus-2 text-only shell
Terminus-KIRA text + image
Terminus-A text + audio
Terminus-IV text + image + video
Terminus-IA text + image + audio
Terminus-MM text + image + audio + video

5. Empirical findings and failure modes

The central empirical result is that full multimodal access substantially improves terminal-agent performance on MMTB (Heo et al., 8 May 2026). For Gemini-3.1-Pro, the reported scores are 0.124 binary and 0.162 partial for Terminus-2; 0.105 binary and 0.159 partial for Terminus-KIRA; 0.333 binary and 0.406 partial for Terminus-IA; 0.333 binary and 0.432 partial for Terminus-IV; and 0.371 binary and 0.469 partial for Terminus-MM (Heo et al., 8 May 2026). The paper interprets these results as indicating that the main bottlenecks are often speech, sound events, motion, timing boundaries, and audio-visual alignment rather than static visual cues alone.

The modality-ladder ablation shows that audio-only and video-only access both help, adding image access helps further, and the full text-plus-image-plus-audio-plus-video setting is best overall (Heo et al., 8 May 2026). The paper therefore argues that multimedia terminal tasks benefit from integrated media access rather than from a single additional modality. This suggests that the relevant competence is not simply “multimodality” in the abstract, but coordinated evidence acquisition across modalities within a filesystem workflow.

The paper also compares native access with proxy reconstruction. When a required modality is missing, agents often fall back to extracted frames, transcripts, OCR dumps, or signal features (Heo et al., 8 May 2026). On matched successful tasks, the measured overhead relative to Terminus-MM includes average API-cost ratios from 1.63× to 7.72×, with worst cases up to 30.11× or 42.49× depending on the missing modality (Heo et al., 8 May 2026). The stated conclusion is that native access shortens the evidence-acquisition path and avoids lossy intermediate representations.

Different systems solve different subsets of tasks. The strongest full-modality controlled harness and Codex CLI solve overlapping but non-nested sets: 28 tasks are solved only by Terminus-MM, 6 only by Codex CLI, 11 by both, and 60 by neither (Heo et al., 8 May 2026). The paper uses this to argue that native multimodal grounding helps on many tasks, while strong terminal scripting and pipeline construction can still outperform native perception on some tasks where media can be reduced to stable intermediate signals (Heo et al., 8 May 2026).

Failure analysis is differentiated by access type. Terminus-MM failures are reported as more often due to model reasoning, over-checking, or choosing the wrong media evidence, whereas Codex CLI failures are more often due to tool or setup loops and lossy proxy reconstruction (Heo et al., 8 May 2026). The qualitative evidence analysis further suggests that native multimodal agents rely on raw audio, raw video, joint audio-visual alignment, visual text on screen, temporal cues, and speaker identity and prosody, while terminal-only or proxy-based agents rely on transcripts, OCR, extracted frames, signal energy or waveform features, and deterministic edit pipelines (Heo et al., 8 May 2026). This distinction indicates that MMTB evaluates not only whether a task is solved, but also which forms of evidence are operationally available to the agent.

6. Position within the benchmark landscape and naming distinctions

MMTB belongs to a benchmark lineage concerned with multimodal evaluation, but its task design is materially different from benchmarks that measure perception or reasoning through question answering. For example, MT-Video-Bench is described as a holistic benchmark for evaluating multimodal LLMs in multi-turn video dialogues, with 987 curated dialogues and 5,805 QA pairs across six competencies such as Object Reference, Memory Recall, Content Summary, Answer Refusal, Topic Shifting, and Proactive Interaction (Pan et al., 20 Oct 2025). Its emphasis is multi-turn video-grounded dialogue using curated golden context, whereas MMTB requires agents to operate in a terminal workspace and produce filesystem artifacts (Pan et al., 20 Oct 2025, Heo et al., 8 May 2026). The distinction is between dialogue evaluation and executable workflow evaluation.

MMBench is likewise a distinct benchmark family. It is a bilingual, objective benchmark for large vision-LLMs built around image-based multiple-choice questions across 20 fine-grained abilities and evaluated with CircularEval and LLM-based answer extraction (Liu et al., 2023). MMTB does not use multiple-choice image QA, bilingual evaluation, or CircularEval. Its primary unit is the Harbor task with an artifact evaluator rather than the image-question quadruple used in MMBench (Liu et al., 2023, Heo et al., 8 May 2026).

A further source of confusion is the similarly named MMTBENCH for multimodal table reasoning. That benchmark studies 500 real-world multimodal tables and 4,021 human-written question-answer pairs, with evaluation based on Exact Match, Substring Match, and F1 over interleaved text-image tables (Titiya et al., 27 May 2025). Despite partial naming overlap, MMTBENCH concerns table reasoning rather than terminal agents operating on multimedia files (Titiya et al., 27 May 2025). In the terminology of the cited papers, MMTB is specifically the benchmark introduced as “Evaluating Terminal Agents on Multimedia-File Tasks” (Heo et al., 8 May 2026).

Within this landscape, MMTB’s distinctive contribution is the evaluation of content-aware audio and video work under persistent file workflows, with success determined by verifiable artifacts rather than human-judged responses or benchmark-internal answer strings (Heo et al., 8 May 2026). This suggests that MMTB is best understood as a terminal-agent benchmark with multimedia grounding, rather than as a conventional multimodal perception benchmark.

7. Limitations, implications, and research direction

The reported results show that even the strongest configurations leave substantial headroom. In the overlap analysis, 60 tasks are solved by neither Terminus-MM nor Codex CLI (Heo et al., 8 May 2026). The paper’s failure analyses attribute remaining errors to model reasoning failures, over-checking, wrong evidence selection, tool or setup loops, and lossy proxy reconstruction (Heo et al., 8 May 2026). These findings indicate that multimedia terminal performance depends on both perception quality and reliable terminal workflow construction.

The benchmark also exposes a methodological point about access pathways. Because MMTB permits ordinary terminal tools such as ffmpeg, ffprobe, ASR/transcription, OCR, silence detection, and signal-processing scripts, it does not measure only native end-to-end multimodal inference; it also measures whether an agent can reconstruct sufficient evidence through terminal pipelines (Heo et al., 8 May 2026). The paper’s comparison between native and proxy routes therefore frames multimedia terminal agency as a question of evidence acquisition strategy as much as raw model competence.

The broader significance of MMTB lies in its operational definition of multimedia competence. Rather than asking whether a system can describe audio or video, it asks whether the system can use audio-video evidence to construct executable workflows that satisfy task-specific evaluators (Heo et al., 8 May 2026). The benchmark is therefore aligned with settings in which practitioners need selected file names, timestamps, JSON or CSV annotations, edit lists, or processed clips rather than prose answers. A plausible implication is that MMTB can serve as a bridge between multimodal reasoning research and applied agentic systems research, particularly for workflows in media production, compliance, education, and annotation.

In summary, MMTB formalizes a distinct evaluation regime: media-centered terminal tasks with artifact-based scoring, multimodal access ablations, and workspace-aware routing through Terminus-MM (Heo et al., 8 May 2026). Its results indicate that native multimedia perception substantially improves both effectiveness and efficiency, but they also show that robust media-grounded terminal agency remains unsolved.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MultiMedia-TerminalBench (MMTB).