- AgenticVBench showcases that frontier vision-language models (VLMs) complete only 31% of real-world video post-production tasks, compared to expert human performance of 81–95%.
- The design of agent harnesses, especially their influence on task outcomes, is identified as being equally important to model choice, shifting performance by up to 20% for some task families.
- Failure modes of agents are task-specific, including problems like long-context information loss, temporal reasoning, and modality misalignment, which holds implications for future improvements in agentic post-production tools.
AgenticVBench is a benchmark of 100 agentic tasks spanning real-world video post-production, authored by 20 industry experts and evaluated across frontier vision-LLMs (VLMs) and multiple agent harnesses. Its central findings are stark: the best model-harness combination reaches only about 31% mean score against expert human performance of 81–95% per family, and the choice of harness alone can shift a fixed model's score by as much as 20 percentage points. The paper argues that harness design is a first-order variable in agentic video production, not merely scaffolding around the model.
Motivation and positioning
Existing multimodal benchmarks evaluate perception and reasoning through single-shot QA over images or videos (Liu et al., 26 Jun 2025, Sakshi et al., 2024, Fu et al., 6 Apr 2026), while agentic benchmarks concentrate on software engineering, browser control, and desktop use (Zhou et al., 2023), [(Zheng et al., 2024)-adjacent works cited: koh2024visualwebarena, xie2024osworld, jimenez2024swebench, merrill2026terminalbench]. Neither measures whether an agent can plan over many steps, manipulate media files, preserve narrative continuity, and produce a finished audiovisual artifact. AgenticVBench fills this gap by evaluating end-to-end production capability rather than understanding in isolation.
Benchmark construction
The four task families were derived bottom-up from expert workflows. Twenty experts—drawn from traditional film studios (avg. 8 years experience), AI film studios (6), independent creators (10), and video AI companies (4)—drafted end-to-end task briefs from their daily work; families were retained only if they corresponded to long-horizon production tasks and admitted verifiable evaluation via programmatic tests or atomic binary rubrics. The families map onto established stages of post-production:
| Family |
Post-production stage |
Tasks |
Evaluation |
| Assembly |
Rough-cut construction |
18 |
Programmatic, chance-adjusted per-slot accuracy |
| Repair |
Review and finishing |
18 |
Programmatic reward vs. golden reference |
| Sequencing |
Narrative reorganization |
28 |
Multiplicative ND * LIS * ADJ |
| Repurpose |
End-to-end compression |
36 |
~30 binary rubrics (~36 pts) + format verifier |
Assembly provides a storyboard of 3–6 slots defined by shot size, camera angle, lens size, and camera movement, plus candidate clips containing one golden clip and AI-generated distractors that differ along exactly one cinematic dimension (generated with Nano Banana Pro frame regeneration followed by Seedance 2.0 image-to-video). Audio is stripped to prevent shortcut exploitation. Scores are chance-adjusted so random selection yields zero. Source material comes from the 2025 Runway AI Film Festival, whose AI-generated films make distractors naturally similar to goldens.
Repair injects defects into one time window of professional source video across three groups: audio defects (noise, echo, distortion), visual defects (color shifts, blur), and timeline defects (swapped shots, filler words, repeated frames, factually wrong sentences). The verifier interpolates the agent's output between the broken input (score 0) and a golden reference (score 1); timeline tasks additionally require an honesty check confirming the rendered output matches the reported edit ranges, gated to zero on mismatch.
Sequencing requires restoring the order of 7–20 shuffled clips given a brief synopsis. The multiplicative score (1−ND)⋅LIS⋅ADJ sharply penalizes partial or random orderings, since a weak dimension cannot be hidden by strong performance elsewhere.
Repurpose is fully expert-authored because deliverables have no single correct answer. Agents receive source videos from 4 minutes to 3 hours—a talk, narrative short, sports broadcast, or music performance—plus a creative brief specifying runtime, aspect ratio, container, tone, and pacing. Rubrics span five pillars: Format (programmatic), Visual, Narrative, Sound (expert binary rubrics), and Penalty items for critical violations.
Quality control runs every candidate through four gates (authoring, asset, expert, verifier QC), including adversarial submissions to detect reward leakage and parser failures. Inter-grader agreement on Repurpose subjective pillars after calibration is high: 96.9% Visual, 96.4% Narrative, 98.2% Sound. The paper concedes that rubric sensitivity and robustness to shallow, format-only solutions remain open concerns.
Main results
Seven frontier VLMs (Claude Opus 4.7, Claude Sonnet 4.6, GPT-5.5, GPT-5.4-mini, Gemini 3.1 Pro, Gemini 3 Flash, Qwen3-VL-235B-A22B-Instruct) were run in 20 model-harness combinations across five harnesses (Claude Code CLI, Codex CLI, Gemini CLI, OpenCode, OpenClaw), each cell with K=3 rollouts under held-constant settings. GPT-5.5 wins every task family; Codex is the winning harness on three of four, with OpenClaw winning Sequencing.
The headline gap is large. On Repurpose—the largest gap at 65 percentage points—the best stack scores 0.30 versus an expert human baseline of 0.95. Assembly shows the smallest gap at 43 pp (0.38 vs. 0.81). Repair localization tops out at 0.30. Agents do produce watchable deliverables, but far below trained editors under identical scoring rules; the human baseline was collected from three film-program editors per task on a stratified subset, graded with the same verifiers and rubrics used for agents.
Failure modes are task-dependent
Trajectory analysis surfaces three behavioral archetypes: smart parallelizers (GPT-5.5) build composite frame strips inspected in single calls; extreme detailists (Claude Opus/Sonnet) over-read frames then self-correct before submission; direct executors (Gemini 3.1 Pro, Qwen3-VL) act before reading context and never recheck. Failure reasons differ sharply by family:
| Failure reason |
Repurpose (%) |
Repair (%) |
| Long-context information loss |
83 |
0 |
| Temporal reasoning |
1 |
65 |
| Modality misalignment |
10 |
24 |
| Hallucinated grounding |
6 |
11 |
In Repurpose, agents exhaust rollout budgets on full-source transcription or repeated thumbnailing without ever reaching assembly. In Repair, agents render valid .mp4 files but cut the wrong window, with median start-time errors of 15–100+ seconds. The authors emphasize that reporting failures as a single aggregate rate would misroute engineering effort.
Harness as a first-order variable
Two results establish harness effects as comparable in magnitude to model differences. First, holding the model fixed, varying the harness moves GPT-5.5's Assembly score by 20 pp (38/37/18 on Codex/OpenCode/OpenClaw)—comparable to the gap between adjacent models on the leaderboard. Second, Qwen3-VL-235B scores 0.009 on Assembly with OpenCode versus 0.073 with OpenClaw, an eight-fold gap attributable to the harness on a fixed model, though the paper notes it lacks Assembly traces to pinpoint the routing difference.
Harness architecture explains these patterns. Codex interleaves reasoning blocks with tool calls, favoring deliberative loops. OpenClaw routes modality work to sub-models (vision sub-VLM, TTS), which wins Sequencing through cleaner narrative-beat reasoning but halves Assembly performance because the planner never sees pixels directly. Claude Code's Plan/TodoWrite tools enforce state-tracking that keeps Opus competitive. OpenCode's cache-first design keeps long ffmpeg pipelines stable. Only OpenClaw ships typed multimodal primitives (image, tts, music_generate, video_generate); the other four expose generic shell/file tools, forcing agents to compose multimodal pipelines from scratch. A striking case study shows the highest-scoring Repurpose rollout producing a structurally complete 60-second narrated recap without ever ingesting a single source frame—narration synthesized via gTTS, music via sine-wave synthesis, clips cut at fixed offsets—illustrating how agents satisfy structural requirements while bypassing visual grounding entirely.
Ablations identify per-family bottlenecks
Oracle interventions isolate what limits each family. In Repair, providing oracle defect locations lifts mean score by +13 pp, showing localization and domain-specific tool use (color grading, super-resolution, audio restoration) are joint bottlenecks. In Sequencing, prepending a full narrative description lifts score by only +22 pp, indicating visual-temporal reasoning remains a genuine constraint even when the task reduces to caption matching. In Assembly, stripping the prose description drops the mean by 27 pp while stripping camera_movement moves it by 1 pp: agents rely on description rather than cinematic shorthand, whereas human editors treat cinematic fields as primary variables—a direct divergence from professional practice. In Repurpose, prepending an editor's reference document yields +23 pp, most pronounced for films with complex logical structure such as non-linear narratives.
Limitations
The evaluation covers a focused set of models and harnesses constrained by compute, cost, and access; coverage of languages, regions, genres, and production conventions is limited by the expert pool and source collection. Reproducibility depends on public media assets and evolving provider APIs, requiring versioned releases and stable verifier code. Subjective editorial dimensions rest on expert rubrics whose sensitivity, inter-rater robustness, and vulnerability to format-only gaming warrant further study. The human baseline is a matched-task reference from film-school editors, not an upper bound achievable by professional studio teams. Finally, the authors note labor-market implications for the production roles represented in their expert pool, recommending human-supervised deployment given the current gap.
Conclusion
AgenticVBench contributes a rigorously quality-controlled, expert-authored evaluation suite for agentic video post-production, with programmatic verifiers, calibrated rubrics, and a matched human baseline. Its empirical conclusions are twofold: current frontier agents complete these tasks at well under half of expert performance, with failure modes that vary sharply by task family; and harness design—including typed multimodal primitives, planning artifacts, and modality routing—moves scores by margins comparable to model choice. Progress on this benchmark therefore depends jointly on stronger multimodal long-context reasoning and on workflow-aware harness engineering, and the released rubrics and verifiers provide the diagnostic substrate for both.