Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

This presentation examines Timeline-Bench, a benchmark that evaluates whether AI agents can complete professional video-editing assignments from start to finish. Rather than testing isolated operations like shot selection or GUI manipulation, it assesses end-to-end editorial capability across 56 real-world tasks spanning narrative scenes, documentaries, advertisements, and more. The benchmark reveals a striking gap: agents can usually satisfy explicit requirements and produce technically valid videos, but they consistently fail to match the editorial quality that professional editors achieve, particularly in pacing, shot selection, and overall polish.
Script
Can an AI agent turn 12 minutes of raw footage into a professional 2-minute edit that an experienced editor would call finished? Timeline-Bench puts 16 frontier models to that test across 56 real production assignments, and the results expose a fundamental gap between technical correctness and editorial craft.
The benchmark contains 56 licensed assignments drawn from EditStock packages, Cinestudy projects, and commissioned commercial work. Collectively, these tasks span 33 hours of source footage and require scripting, take selection, music editing, captions, graphics, color matching, and dual-system synchronization.
Every agent output must pass delivery tests, content integrity checks, and brief compliance before reaching the quality panel. Three multimodal judges score each edit on story, pacing, picture, sound, and graphics, but the judges are calibrated against 2,582 assessments from 43 professional editors to ensure the quality threshold reflects human editorial standards.
The strongest agent, GPT-6 Astra with curated guidance, resolves just 15 of 56 tasks. Claude Opus 5 follows at 23.2 percent, and the aggregate resolution rate across all 896 runs is only 14 percent. Human editors prefer the professional reference edit in 83.5 percent of pairwise comparisons, and even the best agent loses more often than it wins or ties on 54 of 56 tasks.
Here is where the benchmark exposes the core problem: of 771 unresolved runs, 562 pass every delivery, content, and brief test but fail only on quality. Agents produce technically valid videos that meet explicit requirements, yet human editors consistently identify deficiencies in pacing, shot selection, sound design, and overall polish. The agents are clean assemblers, not finished editors.
Trajectory analysis reveals why agents fall short: they devote 83 percent of their perceptual effort to inspecting frames and transcripts before ever rendering the edit, and 98 percent of source perception happens before the first render. Agents verify form, checking for silence or frozen frames, but they rarely evaluate the emergent narrative and temporal qualities that determine whether an edit communicates effectively. If you want to explore this research further or create your own video summaries, visit EmergentMind.com.