Papers
Topics
Authors
Recent
Search
2000 character limit reached

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

Published 28 Sep 2026 in cs.CV, cs.AI, and cs.MM | (2609.35143v1)

Abstract: AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.

Summary

  • The paper evaluates a benchmark dataset of 56 professionally varied, complex video editing tasks using 16 agents, with a primary focus on gaps between completed, correctly formatted tasks and professional- quality projects, such as assessments on story, pacing, photo, sound, and graphics, where professional accuracy significantly outweighs agent accuracy; agent accuracy leads to consistent gaps in AI comprehension of tasks.
  • The study analyzed the failure models of 16 agents, revealing that while 562 out of 771 uncompleted tasks pass technical requirements, they fall short in quality standards, highlighting AI's inability to produce professional-level editorial decisions in video editing.
  • The work details that the Keypoint 3 highlights how each of the 16 agents struggled at aligning with major quality components such as story assembly, pacing, sound and more where human decision-making wins 83.5% of the time.
  • follow-up_questions

Benchmark scope and motivation

“Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut” (2609.35143) evaluates whether AI agents can complete end-to-end professional video-editing assignments rather than isolated operations such as shot selection, temporal localization, or GUI manipulation. The benchmark is motivated by a substantive distinction between satisfying explicit production requirements and producing an edit that experienced editors regard as professionally successful. Video editing is particularly suitable for exposing this distinction because editorial decisions are temporally interdependent: shot selection affects narrative interpretation, duration affects pacing, and changes to picture can require corresponding revisions to dialogue, music, ambience, graphics, and sound design.

The benchmark contains 56 assignments drawn from 11 licensed EditStock packages, 15 Cinestudy projects, and 30 commissioned projects divided between user-generated-content and commercial collections. Each task supplies a brief, raw audiovisual assets, supporting paperwork, a reproducible execution environment, and delivery tests. The reference edit is held out from the agent. Tasks include narrative scenes, interview-based stories, documentaries, trailers, advertisements, personal-branding videos, and travel films. Collectively, they contain approximately 33 hours of primary source footage, with a median of 12.4 minutes per task and delivery durations ranging from 25 to 315 seconds. The dataset includes both landscape and portrait outputs and requires a broad combination of editorial skills, including scripting, take selection, music, dialogue editing, captions, graphics, color matching, reframing, and, in selected cases, dual-system synchronization and compositing.

This design places the benchmark closer to outcome-graded professional-work evaluations than to conventional video-understanding benchmarks. It also extends related agent benchmarks that evaluate software artifacts or knowledge-work deliverables by requiring a finished audiovisual artifact whose quality cannot be fully specified as a binary checklist. The benchmark’s central methodological contribution is therefore not only the task collection but the integration of mechanical verification with human-calibrated quality assessment.

Task construction and evaluation protocol

A task is resolved only if the rendered video passes every required test. The evaluation pipeline separates technical delivery, explicit brief compliance, generic content integrity, and editorial quality. Four delivery tests verify the existence of the output, codec and frame rate, duration, and audio properties. Six shared content tests reject degenerate outputs, including excessive silence, frozen imagery, repeated footage, dead air, channel imbalance, and excessive face cropping. Brief tests check requirements such as scene ordering, use of supplied voice-over, inclusion of specified content, and required on-screen text.

Most brief tests are programmatic. The verifier uses frame and soundtrack matching, OCR, and speech transcription to establish whether required source material, text, and spoken content appear in the appropriate form. The remaining visible-content tests are decided by a two-of-three majority among multimodal model judges. The authors validate the tests against reference edits and deliberately defective controls, such as muted audio, reordered sections, black openings, and frozen frames. This is important because a benchmark can otherwise produce misleading failure rates through under-specified or over-permissive verifiers.

Quality is evaluated separately from compliance. Three multimodal model judges score each edit on story and assembly, pacing, picture, sound, and graphics, together with an overall rating. One judge watches the video with sound; two judges receive contact sheets and file-level measurements rather than continuous audiovisual playback. Scores are standardized within judge and aggregated into a panel score. An agent output passes the quality test only when its panel score exceeds that of the held-out reference edit by a collection-specific tie margin. The margins are calibrated against 2,582 assessable judgments from 43 professional video editors, who performed blind pairwise comparisons between agent outputs and reference edits.

This procedure has a deliberate asymmetry. The reference edit is not treated as unique ground truth, since multiple valid editorial solutions may satisfy a brief. Instead, human preference establishes an empirical tolerance for the model-based quality test. The calibration is effective at the aggregate agent level: quality-test resolution rates correlate strongly with human win-or-tie rates across agents, with Spearman correlation ρ=0.93\rho = 0.93 and Pearson correlation r=0.93r = 0.93. However, the paper reports substantially weaker validity at the individual-edit level: agreement with the human majority is only κ=0.15\kappa = 0.15, and per-task correlations reach only ρ=0.32\rho = 0.32. Thus, the quality test is suitable for comparing agents over many tasks, not for treating every individual pass or failure as a reliable human-equivalent judgment.

Agents and experimental conditions

The study evaluates 16 agents formed from frontier models and three execution settings. Fifteen are coding or terminal agents operating in Linux containers with FFmpeg, Python, Node.js, Remotion, HyperFrames, OCR, transcription, and other media-processing utilities. The models are run through OpenCode, Codex CLI, or Claude Code. A sixteenth condition uses GPT-6. Astra in Codex computer-use mode with DaVinci Resolve on macOS, without shell access.

Each agent receives one attempt on each of the 56 tasks, producing 896 runs in total. Coding runs have a 300-minute limit, 32 virtual CPUs, and 256 GB of memory. The benchmark therefore evaluates sustained planning, asset inspection, construction, rendering, verification, and repair rather than short interactive responses. One additional condition supplies GPT-6. Astra with 1,909 words of general editorial guidance covering brief interpretation, footage inspection, planning, construction, review, and delivery. The guidance is task-independent and does not disclose solutions.

The matched comparisons are informative about the relative contribution of harness and interaction mode. Switching between OpenCode and developer-specific harnesses does not produce statistically reliable differences in either human preference or task resolution. Curated guidance improves the point estimate for GPT-6. Astra from 21.4% to 26.8%, but the matched human win-or-tie difference is only +0.6+0.6 percentage points and is not significant. By contrast, computer use substantially degrades performance: GPT-6. Astra resolves 2 of 53 jointly evaluated tasks versus 12 for the coding condition, with a human win-or-tie difference of −16.4-16.4 percentage points and adjusted p<0.001p < 0.001. This result does not establish that GUI interaction is intrinsically inferior; it establishes that this particular computer-use configuration, which lacked shell and code tools and could not directly access audio through the same workflow, was markedly less effective on these assignments.

Main performance results

The benchmark produces a low absolute resolution rate across all evaluated agents. The strongest condition, GPT-6. Astra in Codex CLI with curated guidance, resolves 15 of 56 tasks, or 26.8%. Claude Opus 5 in Claude Code follows at 23.2%, while GPT-6. Astra in either OpenCode or unguided Codex CLI resolves 21.4%. Across all 896 runs, only 125 tasks are resolved, for an aggregate resolution rate of 14.0% with a 95% confidence interval of 11.7–16.4%.

Agent condition Tasks resolved Resolution rate Human win-or-tie rate
GPT-6. Astra, Codex CLI + guidance 15/56 26.8% 24.4%
Claude Opus 5, Claude Code 13/56 23.2% 23.2%
GPT-6. Astra, Codex CLI 12/56 21.4% 23.8%
GPT-6. Astra, OpenCode 12/56 21.4% 22.3%
Claude Fable 5.1, Claude Code 10/56 17.9% 23.5%
Claude Opus 5, OpenCode 10/56 17.9% 19.0%
GPT-5.6. Sol, OpenCode 10/56 17.9% 17.6%
All agents 125/896 14.0% 16.5%

The confidence intervals are wide because each agent receives only one run per task; for individual agents they are approximately ±11\pm 11 percentage points. Consequently, the paper appropriately cautions that several top-ranked agents cannot be reliably distinguished. Runtime also does not explain success: median runtime ranges from 16 to 61 minutes, with Spearman correlation ρ=0.15\rho = 0.15 between runtime and resolution. Estimated cost is more positively associated with performance, with ρ=0.79\rho = 0.79 among 12 agents with cost records, although the best-performing guided condition costs approximately $9.76 per run rather than being the most expensive condition.

The human study establishes a stricter baseline for editorial quality. Across 2,582 judgments, editors prefer the professional reference edit in 83.5% of comparisons, prefer the agent edit in 11.1%, and report no meaningful preference in 5.5%. The aggregate human win-or-tie rate across agents is 16.5%. The low single-judgment reliability, $r = 0.93$0, reflects the subjectivity of pairwise editorial evaluation, but averaging three judgments per pair yields substantially more stable agent-level estimates, with estimated reliability of 0.84. The ranking remains highly stable when individual editors are removed.

Compliance is substantially stronger than editorial craft

The most consequential result concerns where agents fail. Of 771 unresolved runs, 562, or 73%, pass delivery, content, and brief tests but fail only the quality test. For 15 of the 16 agents, quality-test failure accounts for 61–90% of unresolved runs. Only six of the 863 delivered videos fail a delivery test, and 762 pass every brief test. This produces a strong and potentially contradictory claim: the evaluated agents are generally capable of producing technically valid videos that satisfy explicit requirements, but they are not reliably capable of producing edits that professional editors consider finished.

The distinction has direct implications for benchmark interpretation. A checklist-only evaluation would substantially overstate competence. The agents can render valid files, preserve required scene order, include specified dialogue or graphics, and avoid many forms of timeline corruption. Yet these capabilities do not imply adequate pacing, shot selection, sound design, visual finishing, or overall polish. The quality test is therefore not an optional aesthetic supplement; it determines whether the benchmark measures the intended professional deliverable rather than merely a syntactically valid video.

The quantitative editing differences support the human judgments. Relative to the reference edits, agents cut at 0.72 times the reference cut rate, use 0.76 times as many shots per edit, produce median shots 1.59 times longer, and hold their longest static shots 1.50 times longer. Human editors penalize slower pacing but do not show an analogous penalty for cutting faster than the reference. The paper estimates that an agent loses 6.6 percentage points of human win-or-tie rate for each halving of its cut rate below the reference rate, with a 95% confidence interval of 3.2–10.0 points.

Human comments identify overall polish, shot selection, graphics, transitions, story structure, sound design, and color as the principal strengths of the reference edits. In contrast, explicit technical defects are relatively infrequent among the strongest agents. For the six highest-rated agents, only 10% of rejected outputs contain a human-noted defect that a content test could plausibly detect, compared with 19% for the three weakest coding agents. The computer-use condition is a clear outlier: 28% of rejected edits contain such defects, including slates, crew dialogue, or retakes in 18.5% of losses.

The pattern of preferences varies by collection. Narrative quality drives 84–86% of tagged judgments in narrative collections, whereas graphics determine 91% of tagged UGC judgments. This result indicates that benchmark difficulty is partly collection-dependent and that a single aggregate score conceals different failure modes across production genres.

Process analysis and the missing audiovisual feedback loop

Trajectory analysis provides a mechanistic account of the quality gap. Coding agents devote 54% of their 111,636 native actions to source perception, 17% to verification, and only 9% to building. Eighty-three percent of source-perception reads involve frames or contact sheets. The first render occurs late, between 72% and 90% of the way through a run, and 98% of source-perception activity occurs before that first render. Agents therefore form plans from still images, transcripts, and measurements before observing the assembled audiovisual result.

This workflow is not merely inefficient; it limits the type of errors that agents can detect. Agents inspect their renders in 86–100% of runs for most coding conditions, and verification correlates with resolution across agents at r=0.93r = 0.931. However, only 5.5% of the problems they report after checking concern pacing, story, or shot choice, whereas 66% of human editors’ notes concern those dimensions. Agents predominantly verify form: file properties, repeated footage, synchronization, silence, and other mechanically observable defects. They rarely evaluate the emergent temporal and narrative properties that determine whether an edit communicates effectively.

The mismatch between self-assessment and actual performance is particularly severe. Agents claim full success in 93% of runs, including 95.5% of outputs that fail at least one benchmark test. This indicates that their verification routines do not provide calibrated uncertainty about either compliance or quality. The problem is not simply that agents lack an editing heuristic; they also lack a reliable evaluator capable of identifying when the constructed edit requires substantive revision.

The paper’s characterization is consequently precise: current agents are “clean assemblers” rather than finished editors. They are relatively reliable at delivery and explicit requirements, increasingly capable at basic narrative assembly, and weak at rhythm, shot selection, finishing, and self-judgment. The evidence supports a need for native audiovisual perception and iterative review of rendered sequences, but it does not establish which combination of continuous video encoders, audio-conditioned reasoning, editorial planning representations, or learned quality critics would solve the problem.

Task difficulty, model ability, and benchmark structure

A Rasch analysis shows that task difficulty varies more than agent ability. The estimated task standard deviation is 0.83 on the logit scale, compared with 0.51 for agents, yielding a variance ratio of 2.6. Thirty-five tasks are resolved by no agent, whereas only approximately five would be expected under random allocation of successes. Even the strongest agent is more likely to lose than to win or tie against the reference edit on 54 of the 56 tasks.

Collection membership explains approximately half of the variation in task difficulty. Human win-or-tie rates range from 28.7% for Cinestudy to 6.8% for UGC. Within collections, however, no individual task descriptor or required editing skill predicts difficulty after correction for multiple comparisons. Likewise, after task difficulty is removed, no agent demonstrates a reliable specialization for particular required skills; the joint permutation test yields r=0.93r = 0.932.

These findings constrain claims about model capability. The benchmark does not support a simple interpretation in which one model is consistently superior for, for example, interviews, graphics, or sound effects. Performance is strongly conditioned by the assignment itself, and the current sample does not identify stable agent-by-skill interactions. The benchmark is therefore better suited to measuring robust end-to-end reliability over heterogeneous tasks than to constructing a fine-grained skill profile for each model.

Limitations and open questions

The principal statistical limitation is that every agent runs only once per task. The reported confidence intervals quantify uncertainty across tasks but omit run-to-run variance caused by stochastic model behavior, tool-use trajectories, and rendering choices. Consequently, small differences in resolution rates should not be interpreted as stable model rankings.

The quality test has a further dependence on its calibration procedure. It is calibrated using the same pool of human judgments against which aggregate behavior is assessed, although the reported results use leave-one-agent-out margins to reduce circularity. The paper explicitly states that the test is validated per agent rather than per edit. In addition, professional reference edits are exemplars rather than certified ground truth; a reference may be less suitable than an alternative edit for a particular audience even when it serves as the benchmark comparison.

The human sample is experienced but heterogeneous: 43 freelance editors report an average of at least two years of experience, with most falling in the one-to-three-year band. Single judgments are noisy, and only aggregate rates are released to protect participants. The source media are also access-controlled, which supports licensing and privacy requirements but limits fully independent reproduction outside approved research teams. Finally, the benchmark samples 56 tasks from four collections, and collection effects are large. The absence of within-collection predictors should therefore be treated as a result about this task set, not as evidence that editorial skill requirements are generally unrelated to difficulty.

Several specific questions remain open. It is unclear whether agents would improve primarily through continuous audiovisual observation, earlier rendering and iterative revision, stronger editorial planning, specialized video-editing tools, or a calibrated critic trained on professional preferences. It is also unresolved whether repeated stochastic runs would materially increase resolution, whether agent ensembles could improve shot selection and pacing, and whether the current quality-test calibration would remain valid as agents produce edits substantially unlike the reference distribution.

Conclusion

TIMELINE-BENCH evaluates video editing as a complete professional workflow rather than as a collection of isolated media operations. Across 16 agents, the best condition resolves only 15 of 56 tasks, while the aggregate resolution rate is 14.0%. The dominant failure is not delivery or explicit brief compliance: 562 of 771 unresolved runs fail only the quality test. Human editors prefer the reference edits in 83.5% of judgments and identify deficiencies in polish, shot selection, pacing, graphics, sound, transitions, and finishing.

The benchmark consequently demonstrates a persistent separation between executable correctness and editorial quality. The evaluated agents can usually produce valid, requirement-compliant videos, but their perceptual and evaluative loops remain focused on detectable defects rather than narrative rhythm and viewing experience. By releasing the tasks, verifier, frozen evaluation constants, and per-run results, the paper provides an outcome-oriented test of whether future editing agents can close that specific gap.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces TIMELINE-BENCH, a test designed to measure how well AI agents can perform real video-editing jobs.

Instead of asking an AI a question or giving it a small editing task, the researchers give it:

  • Raw video and audio files
  • A written project brief
  • Editing tools
  • Rules about the final video

The AI must then create a complete video, such as a short film scene, interview, advertisement, or travel video.

The main question is: Can AI make a finished video that is not only technically correct, but also interesting and professional to watch?

2. What did the researchers want to find out?

The researchers focused on several questions:

  1. Can AI agents complete realistic video-editing assignments from beginning to end?
  2. Can they understand a project brief and choose the right video clips, sounds, music, and graphics?
  3. Can they create a video with a good story, rhythm, and professional style?
  4. How do different AI models and computer tools compare?
  5. When AI fails, does it fail because of technical mistakes or because the video simply does not feel well edited?

This last question is especially important. A video may have the correct length and file type but still be boring, confusing, or poorly paced.

3. How was the research carried out?

The video-editing tasks

The researchers created 56 different editing tasks. These tasks came from several sources, including professional editing projects and publicly available practice projects.

The tasks covered many kinds of videos:

  • Short narrative films
  • Interviews and documentaries
  • Advertisements
  • Trailers
  • Travel videos
  • Social-media videos
  • Product and brand videos

Each task included raw materials and instructions. For example, an AI might be told to create a 60-second advertisement using certain product shots, a voice-over, music, and text on screen.

The AI had to:

  1. Look through the available footage and audio
  2. Decide which parts to use
  3. Put the clips in the right order
  4. Add music, sound effects, captions, or graphics
  5. Export the final video in the required format

The AI systems

The study tested 16 AI agents. An agent is more than just an AI model: it is an AI model connected to tools that let it read files, write code, run programs, and create videos.

Some agents worked through coding tools such as:

  • Codex
  • Claude Code
  • OpenCode

One version also received extra advice about professional video editing. Another version used a computer mouse and keyboard in DaVinci Resolve, a popular video-editing program.

The tests

The researchers used several kinds of tests.

Technical tests checked things such as:

  • Does the video file exist?
  • Is it in the correct format?
  • Is it the right length?
  • Does it contain working sound?
  • Is the video free of frozen images, long silent sections, or repeated footage?

Brief tests checked whether the AI followed the instructions. For example:

  • Did it use the required voice-over?
  • Did it keep scenes in the correct order?
  • Did it include the requested logo or text?

Some of these checks were done by computer programs. For example, speech-recognition software checked what was said, and OCR software read words appearing on screen. OCR is a tool that turns writing in an image into computer-readable text.

Human and AI judges

Technical tests cannot fully decide whether a video is good. There can be many correct ways to edit the same project.

Therefore, 43 professional video editors compared the AI videos with reference videos made by humans. They watched both versions without knowing which one was made by AI and chose the version they preferred.

The researchers also used three AI judges to score video quality. These AI judges looked at things such as:

  • Story and structure
  • Pacing, meaning how quickly the video moves
  • Picture quality
  • Sound
  • Graphics and titles

The researchers adjusted the AI judging system so that its results matched human opinions as closely as possible.

4. What were the main findings?

AI agents completed relatively few tasks

The strongest system completed only 15 of the 56 tasks, or 26.8%.

Across all the systems, the average success rate was about 14%. This means that, on average, an AI successfully completed only about one task out of seven.

The best results came from GPT-6 Astra using Codex with extra editing advice. However, the researchers warn that the top systems were close enough that it is difficult to say that one was definitely better than the others.

Most failures were about quality, not basic correctness

The most important result was that 562 of 771 unsuccessful attempts passed the technical, content, and instruction tests but failed the quality test.

In simple terms, many AI systems created videos that:

  • Had the right file format
  • Used the required material
  • Followed the basic instructions
  • Did not contain obvious technical problems

But the videos still did not feel as polished or professional as the human reference edits.

This shows that making a video is not only about following rules. It also requires creative judgment.

Human editors strongly preferred the human reference videos

Professional editors preferred the human-made reference video in 83.5% of their judgments. They preferred the AI version in only 11.1% of cases, while the remaining judgments were ties.

Human editors especially noticed differences in:

  • Overall polish
  • Choosing the best shots
  • Transitions
  • Graphics and text
  • Sound design
  • Color
  • Pacing

The AI often produced a clean basic assembly, but it did not make the many small creative decisions that make a video feel finished.

AI videos were often too slow

Compared with the reference edits, AI videos changed shots only about 72% as often. Their longest still images were about 1.5 times longer.

This suggests that AI often left shots on screen too long. The result could feel slow or less energetic, even when all the required scenes were included.

AI had difficulty judging its own work

The agents usually checked their own videos, but they mostly looked for technical problems:

  • Missing files
  • Broken audio
  • Incorrect video length
  • Repeated footage
  • Rendering errors

Only about 5.5% of the problems they reported involved deeper creative issues such as story, pacing, or shot choice. By contrast, human editors mentioned these creative issues in about 66% of their comments.

Even more concerning, the agents claimed they had succeeded in 93% of their runs, including many videos that actually failed the tests. This means the agents were often unable to recognize that their own editing was weak.

The way AI viewed video was a major limitation

Most coding agents did not watch the footage in the same way a human editor does. They often examined:

  • Still images taken from the video
  • Contact sheets, which are collections of small preview images
  • Written transcripts
  • Audio volume measurements

This is like trying to understand a movie by looking at one photograph per second and reading a script, without properly watching and listening to the whole thing.

Because of this, the agents had trouble understanding:

  • Natural timing
  • Facial expressions
  • Emotional moments
  • The rhythm of a conversation
  • Whether a cut felt smooth
  • How music and images worked together

5. Why are these findings important?

The paper shows that current AI agents are fairly good at following explicit instructions and producing technically valid files. They can perform many basic editing operations.

However, they are still weak at the part of editing that human professionals contribute most: creative judgment.

A human editor does not simply place clips in an order. They decide:

  • Which take has the best performance
  • When to cut to another shot
  • How long to hold a moment
  • How to build emotion and suspense
  • How music should support the story
  • Whether the finished video feels natural and engaging

The research suggests that future video-editing AI will need several improvements:

  • Better ability to watch and listen to complete videos
  • Stronger understanding of emotion and storytelling
  • Better tools for professional editing, sound, color, and graphics
  • More useful ways to review its own work
  • The ability to recognize creative problems, not only technical errors

Conclusion

TIMELINE-BENCH provides a realistic way to test AI video editors. It shows that current AI agents can often create videos that meet basic instructions, but they rarely match the quality of professional human editors.

The main lesson is that correctness is not the same as quality. An AI may deliver a video with the right length, clips, and file format while still making poor choices about pacing, storytelling, and style.

The benchmark could help researchers measure whether future AI systems genuinely improve. For now, the results suggest that AI can assist with parts of video editing, but human editors are still needed to make the final video feel polished, clear, and emotionally effective.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Limited task sample size: The benchmark contains only 56 tasks, so resolution-rate differences of a few tasks are not statistically distinguishable; larger task collections are needed for reliable agent ranking.
  • No estimate of run-to-run variability: Each agent performs each task only once, leaving unresolved how much results depend on stochastic model behavior, sampling, transient tool failures, or different planning trajectories.
  • Unclear generalizability beyond the sampled collections: Tasks come from EditStock, Cinestudy, commissioned projects, and UGC-style productions, but it is unknown whether the findings transfer to feature films, television, news, educational media, social-media content, live events, or other professional workflows.
  • Potential source and selection bias: The composition of the benchmark, including which publicly available projects and reference edits were selected, may overrepresent particular genres, production standards, visual styles, or editorial conventions.
  • Insufficient coverage of editing skills: Some capabilities, such as dual-system synchronization and VFX compositing, appear in only seven tasks, making it difficult to assess agent competence in these areas separately.
  • No systematic analysis of task attributes: Within collections, the paper does not identify which measurable properties—footage volume, number of takes, dialogue density, ambiguity of the brief, audiovisual complexity, or required effects—cause task difficulty.
  • Unresolved distinction between model and harness effects: Comparisons confound model capability with harness design, tool availability, prompting, context management, and provider-specific optimizations; controlled factorial experiments are needed to isolate these factors.
  • Limited evaluation of computer-use agents: Only one computer-use configuration is tested, and it differs from coding agents in operating system, application, tool access, and interaction modality, preventing a clean comparison between GUI and programmatic editing.
  • No systematic ablation of editorial guidance: The curated guidance condition changes only one agent and one harness, so it remains unclear which specific instructions, workflow recommendations, or perception strategies produce the observed improvement.
  • Unclear causal role of audiovisual perception: The analysis associates reliance on still frames and transcripts with weak editorial craft, but does not experimentally test whether native video/audio playback, shot-level browsing, or richer temporal representations improve editing quality.
  • No intervention study on iterative reviewing: Agents render late and check mainly for technical defects, but the benchmark does not test whether mandatory early previews, structured editorial review cycles, or targeted pacing/story checklists improve outcomes.
  • Quality evaluation depends on reference edits: Reference edits are treated as comparison targets even though they are not certified ground truth and may represent only one valid creative interpretation of a brief.
  • Possible mismatch between reference quality and task acceptability: The quality test requires an agent edit to exceed a collection-specific margin relative to the reference, which may penalize coherent alternatives that differ substantially in style, pacing, or structure.
  • Calibration and validation are not fully independent: The model-based quality test is calibrated using the same human judgments used to validate it, creating a risk of overestimating agreement despite the leave-one-agent-out procedure.
  • Quality-test validity is established per agent rather than per edit: The paper acknowledges that the evaluator is validated at the agent level, leaving uncertainty about whether it reliably classifies individual outputs, especially unusual or highly creative edits.
  • Heavy reliance on multimodal LLM judges: Three model judges determine quality and some brief requirements, but their susceptibility to visual-style bias, familiar editing conventions, prompt sensitivity, and systematic disagreement with expert editors remains insufficiently characterized.
  • Asymmetric judge inputs: One quality judge watches the videos with sound, whereas two judges receive contact sheets and file-level measurements; this heterogeneous observation protocol may distort the panel score and obscure which sensory information is necessary for evaluation.
  • Human preference is measured against a one-line project description rather than the full brief: Editors may judge overall appeal without knowing explicit narrative, technical, or client requirements, so human win-or-tie rates may not fully represent task success.
  • Low individual inter-rater agreement: Krippendorff’s α=0.10\alpha=0.10 indicates substantial judgment ambiguity, but the study does not investigate which task types or editorial dimensions produce disagreement.
  • Limited expert-panel diversity: The study uses 43 editors with an average of at least two years of experience, but does not report detailed variation in seniority, specialization, geography, genre expertise, or professional standards.
  • No comparison with novice, client, or non-editor audiences: It remains unknown whether the benchmark’s expert preferences align with client satisfaction, audience engagement, or judgments from less specialized viewers.
  • Potential reference-order and presentation effects remain incompletely explored: Although order effects were reported as negligible, the hidden seek bar, forced full viewing, one-line description, and pairwise reference comparison may not reflect real commissioning or review workflows.
  • No evaluation of edit usefulness in downstream production: The benchmark measures final-video tests and preference, but does not assess whether outputs are practically usable by clients, require substantial human revision, or reduce total production time and cost.
  • No human–agent collaboration condition: The study evaluates autonomous agents, leaving open whether agents can meaningfully assist professional editors through candidate selection, rough cuts, revision suggestions, or targeted automation.
  • Unclear relationship between technical cleanliness and creative quality: The paper reports that agents pass many delivery and content tests while failing quality, but does not quantify how much each technical defect contributes to human rejection or whether technical reliability enables later creative improvement.
  • No evaluation of revision capability: Agents are not tested on responding to editorial feedback, client changes, deadline pressure, or successive rounds of revision—central properties of real editing work.
  • No robustness testing under altered briefs or assets: It remains unknown whether agents can handle ambiguous, incomplete, contradictory, noisy, or deliberately revised briefs, nor how they respond to missing media, mislabeled files, poor transcripts, or corrupted assets.
  • No adversarial or contamination analysis beyond reference isolation: The paper checks that agents did not retrieve reference edits, but does not examine training-data contamination, familiarity with benchmark source projects, or memorization of publicly available footage.
  • Temporal reproducibility is uncertain: The evaluated models, harnesses, transcription systems, and model judges are rapidly changing, and provider outputs may not be stable, complicating longitudinal comparisons and replication.
  • Cost–quality trade-offs are underexplored: The reported positive correlation between cost and resolution is based on few agents and observational comparisons; controlled studies are needed to determine whether additional inference time, tokens, or compute causally improve editorial performance.
  • The benchmark does not establish a ceiling for human performance: Human editors compare outputs with references but are not evaluated on independently completing the same tasks, so the gap between agents, professional editors, and the best possible solution remains unknown.
  • No assessment of creative diversity: Since the evaluation emphasizes similarity or superiority relative to reference edits, it does not measure whether agents can produce multiple valid stylistic solutions or adapt edits to different audiences and brand identities.
  • Limited analysis of error propagation: The paper does not trace how early decisions—such as asset selection, transcript interpretation, or story outline—cause downstream problems in pacing, sound, graphics, and finishing.
  • Open question about scalable quality verification: The benchmark’s human-calibrated quality procedure is expensive and tied to fixed collections; it remains unresolved whether reliable, collection-independent evaluators can be built for new tasks without fresh human judgments.

Practical Applications

Immediate Applications

The paper’s results support applications that use video agents as assistive, compliance-oriented, and quality-controlled tools, rather than as fully autonomous replacements for professional editors. The strongest current capabilities are delivery-format compliance, explicit brief requirements, transcription, basic assembly, and technical validation.

  • Automated first cuts for media-production teams (film, documentary, advertising, and digital media)
    • Workflow: ingest dailies → transcribe and index footage → generate a script-compliant assembly → run delivery and content tests → hand the timeline to a human editor.
    • Assumptions/dependencies: footage must be well organized; the brief must state explicit requirements; human editors remain responsible for pacing, shot selection, graphics, sound design, and final approval.
  • Automated technical quality assurance for video deliverables (broadcasting, streaming, advertising, education, and corporate communications)
    • Potential product: a CI-style “video build verifier” integrated with editing systems or media asset-management platforms.
    • Assumptions/dependencies: automated tests must be customized for each organization’s delivery specifications and validated against known-good and defective examples.
  • Brief-compliance checking for post-production workflows (advertising, branded content, social media, and e-learning)
    • Example: before publishing a commercial, the system confirms that the supplied voice-over is audible, the legal disclaimer appears on screen, the product is shown, and the final cut meets the duration window.
    • Assumptions/dependencies: transcription and OCR errors must be monitored, especially for accents, noisy recordings, stylized text, multilingual content, and overlapping dialogue.
  • Automated subtitle, caption, and social-format preparation (education, accessibility, marketing, and creator platforms)
    • Potential tools: batch generation of captioned vertical clips, platform-specific aspect-ratio exports, and compliance checks for caption presence and timing.
    • Assumptions/dependencies: captions require human or automated review for accuracy, speaker attribution, timing, accessibility standards, and sensitive content.
  • Media-asset logging and editorial search (production companies, newsrooms, universities, and archives)
    • Workflow: automatically create transcripts, thumbnails, shot boundaries, speaker labels, and metadata before a human begins editing.
    • Assumptions/dependencies: this improves retrieval more reliably than creative judgment; visual understanding remains incomplete when agents inspect footage primarily through still images.
  • Human-in-the-loop rough-cut assistance (small studios, agencies, independent creators, and corporate communications)
    • Practical division of labor: the agent handles synchronization, organization, format conversion, basic scene assembly, and defect detection; the editor handles story, pace, performance, shot quality, sound design, and visual polish.
    • Assumptions/dependencies: the interface must preserve editable timelines and source links rather than returning only a flattened video.
  • Benchmark-based procurement and regression testing for media agents (software vendors, production houses, and research organizations)
    • Potential workflow: evaluate candidate systems on representative tasks, track delivery-test and quality-test performance separately, and rerun the suite after model or tool updates.
    • Assumptions/dependencies: the benchmark is a limited sample of 56 tasks; one run per task omits run-to-run variability, and the reference edits are not uniquely correct ground truth.
  • Research training and evaluation for multimedia-agent developers (academia and industrial AI research)
    • Useful metrics: resolution rate, human win-or-tie rate, brief-test failures, quality-test margin, cut frequency, static-shot duration, and the proportion of fixes addressing creative rather than technical defects.
    • Assumptions/dependencies: model-judge scores should remain calibrated against independent human judgments and should not be treated as a perfect substitute for expert review.
  • Personal and small-business video production (daily life, influencers, educators, and local organizations)
    • Best current use: produce a publishable draft for low-risk content or a starting point for a human review.
    • Assumptions/dependencies: users must check privacy, copyright, factual accuracy, unwanted speech, music licensing, and accidental inclusion of slates, retakes, or private material.
  • Public-sector and institutional media compliance (government, healthcare communications, universities, and nonprofits)
    • Potential product: an institutional media-governance pipeline that blocks noncompliant exports before publication.
    • Assumptions/dependencies: legal and accessibility requirements vary by jurisdiction and must be encoded and periodically reviewed by domain specialists.

Long-Term Applications

The paper indicates that broader applications require advances in native audiovisual perception, creative self-evaluation, iterative editing, and human-calibrated quality assessment. The current resolution rate—26.8% for the strongest reported agent, with humans preferring reference edits in 83.5% of judgments—does not support unsupervised professional editing.

  • Autonomous end-to-end video editors (film, advertising, television, and digital media)
    • Required developments: continuous audio-visual viewing rather than primarily still frames and transcripts; multimodal timeline representations; stronger temporal reasoning; and the ability to generate and compare multiple editorial alternatives.
    • Dependencies: reliable creative-quality evaluation, rights-aware asset handling, robust performance across genres, and acceptance by professional editors and clients.
  • Interactive co-editors with creative critique and revision loops (professional editing software and collaborative production)
    • Potential workflow: render early draft → watch the full result with synchronized sound → critique pacing, story, graphics, and audio → generate two or more alternatives → request targeted human approval.
    • Dependencies: evaluations must measure whether proposed changes improve human preference, not merely whether they remove technical defects.
  • Native audio-visual agents for long-form footage (documentary, news, sports, and reality production)
    • Potential tools: synchronized audio-video embeddings, searchable performance-level indexes, speaker and emotion continuity tracking, and timeline-aware scene graphs.
    • Assumptions/dependencies: models must avoid overinterpreting emotion or identity and must support privacy protections for people appearing in footage.
  • Automated sound-design and finishing systems (cinema, advertising, games, podcasts with video, and immersive media)
    • Expected benefit: reduce the finishing gap identified in the study, where human reviewers frequently cited polish, sound design, color, graphics, and transitions as advantages of reference edits.
    • Dependencies: access to licensed music and effects libraries, accurate loudness and mixing standards, and safeguards against stylistic homogenization.
  • Personalized multi-version content generation (marketing, education, healthcare communication, and accessibility)
    • Examples: a long documentary condensed into classroom modules; a product video adapted for portrait social media; or a health-information video accompanied by captioned, translated, and audio-described variants.
    • Dependencies: human review for factual claims, translation quality, cultural appropriateness, accessibility compliance, and preservation of required warnings or disclaimers.
  • Policy and procurement standards for creative AI systems (government, media regulators, broadcasters, and enterprise buyers)
    • Possible policy requirements: disclose benchmark coverage, report failure modes, preserve audit logs, test for unauthorized use of reference material, and require human approval for high-impact or public-facing content.
    • Dependencies: benchmark representativeness, independent replication, protection of licensed media, and avoidance of treating one professional editing style as universal quality.
  • Education and training simulators for editors (film schools, vocational programs, and corporate training)
    • Potential product: a simulated post-production studio with briefs, raw footage, timeline tools, verifier tests, and structured critiques of story, pacing, sound, graphics, and color.
    • Dependencies: expert-authored rubrics, diverse genres and cultures, careful separation between objective errors and legitimate stylistic choices, and protection against training students to optimize only for benchmark scores.
  • Multi-agent production pipelines (large studios, advertising networks, and distributed creative teams)
    • Potential workflow: asset analysis → narrative planning → rough assembly → sound and graphics passes → independent quality critique → human sign-off.
    • Dependencies: shared timeline representations, conflict resolution, provenance tracking, consistent stylistic direction, and safeguards against compounding errors across agents.
  • Creative-quality evaluation infrastructure for AI-generated media (research labs, platforms, and model developers)
    • Potential capability: continuously updated quality panels that measure pacing, narrative coherence, visual continuity, audio quality, graphics, and audience suitability across task collections.
    • Dependencies: independent human validation, protection against judge drift, representative reviewers, transparent uncertainty estimates, and recognition that quality is audience- and purpose-dependent.
  • High-stakes automated communication production (healthcare, emergency response, finance, and public information)
    • Required safeguards: approved-source retrieval, factual and legal verification, version control, clear provenance, accessibility review, and mandatory expert approval before release.
    • Feasibility constraint: the paper demonstrates capability on editorial tasks, not factual safety or high-stakes communication; deployment in these sectors therefore requires separate domain-specific validation.

Glossary

  • Adaptive detector: An algorithm that adjusts its detection criteria based on the characteristics of the input, such as changes in video scenes. “our re-implementation of PySceneDetect’s adaptive detector”
  • Ambience: Background environmental sound that establishes the acoustic setting of a scene. “Ambience / room tone”
  • Audio-visual perception: The ability to interpret visual and auditory information jointly. “The results suggest that creative benchmarks need human-calibrated quality evaluation, while capable agents need native audio-visual perception”
  • Blind pairwise study: An evaluation in which participants compare two alternatives without knowing their identities or experimental conditions. “In a blind pairwise study, 43 professional video editors compared each delivered output with its task’s reference edit.”
  • Color matching: Adjusting the colors of different shots so that they appear visually consistent. “every UGC task requires captions, color matching and reframing for portrait delivery”
  • Contact sheet: A collection of representative images, usually arranged in a grid, used to inspect visual content efficiently. “they read contact sheets of one frame per second instead”
  • Container: An isolated software environment that packages an application, its dependencies, and execution settings. “A TIMELINE-BENCH task consists of an edit brief, source assets, a Docker image, a set of tests, and a time limit”
  • Content test: An automated or model-based check that verifies whether a video contains required material and avoids defective output. “Six content tests, shared by all tasks, reject degenerate edits such as mostly silent, frozen or looped ones”
  • Cut rate: The frequency at which a video changes shots, commonly measured in cuts per minute. “agents change shots only 0.72 times as often as the reference edit”
  • Dailies: Unedited footage recorded during a production, including multiple takes and production material. “Klug Brand Story asks for a sixty-second, interview-driven brand-story commercial from about four hours of dailies.”
  • Degenerate edit: A technically produced but unusable video, such as one that is mostly silent, frozen, or repetitive. “Six content tests, shared by all tasks, reject degenerate edits”
  • Delivery specification: The required technical properties of a media file, including codec, frame rate, duration, and audio format. “The tests assess whether the rendered video meets the technical specifications”
  • Dual-system sync: Synchronizing separately recorded production audio with corresponding camera footage. “Dual-system sync”
  • FFmpeg: A software suite and command-line framework for processing, converting, and analyzing audio and video. “The coding agents run in Linux containers with the tools an editor working in code needs: FFmpeg”
  • Frozen input: An evaluation input that remains identical across runs to ensure comparability and reproducibility. “checklists for agent benchmarks call for frozen inputs and validated evaluators”
  • Harbor format: A standardized packaging format for agent-evaluation tasks, including inputs, instructions, and tests. “Tasks are packaged in the Harbor format”
  • Harness: A software system that supplies an AI model with tools and executes its commands. “An agent is a model running in a harness, the program that gives the model its tools and executes its commands.”
  • Headless browser: A web browser controlled programmatically without a visible graphical interface. “Python and Node, the programmatic video frameworks Remotion and HyperFrames with a headless browser”
  • Integrated loudness: A measurement of the average perceived loudness of an entire audio program. “integrated loudness (EBU R128, measured with FFmpeg)”
  • Inter-rater reliability: The degree to which different evaluators produce consistent judgments when assessing the same material. “per-agent rates are reliable (Spearman–Brown reliability 0.84 with three human editors per pair)”
  • Leave-one-agent-out: An evaluation procedure that excludes an agent’s own data when fitting parameters used to score that agent. “each agent is scored with margins refitted per collection without its own judgments (leave-one-agent-out).”
  • Long-horizon task: A task requiring sustained planning and execution across many sequential operations. “AI agents increasingly carry out long-horizon professional work”
  • Mechanical oracle: A deliberately constructed reference implementation or output used to validate whether automated tests behave correctly. “Each task therefore ships a mechanical oracle”
  • Multimodal LLM: A LLM capable of processing multiple information modalities, such as text, images, audio, or video. “Three multimodal LLM judges, Gemini 3.8. Flash, GPT-6. Astra and Claude Opus 5.5”
  • Non-linear editing: Video editing in which clips can be accessed and rearranged independently rather than processed sequentially from start to finish. “DaVinci Resolve 21, 2026. URL … Non-linear editing, color, effects and audio post-production software.”
  • OCR: Optical character recognition, the automated conversion of text appearing in images or video frames into machine-readable text. “reads its on-screen text with OCR (PP-OCR via RapidOCR”
  • Panel score: A combined evaluation score calculated by aggregating the normalized scores of multiple judges. “The panel score of an edit is the average of its three z-scores.”
  • Pacing: The temporal rhythm of an edited video, including the speed and duration of shots and narrative progression. “only 5.5% of the problems agents report after a check concern pacing, story or shot choice”
  • Picture lock: The stage at which the visual edit is considered final and no further picture changes are expected. “coordinating picture and sound over time”
  • Portrait delivery: A video output designed for a vertically oriented display format. “the 41 landscape and 15 portrait deliverables”
  • Post-production: The stage of media creation in which recorded material is edited, enhanced, mixed, and prepared for delivery. “post-production operations and GUI trajectories in media software”
  • Programmatic test: A test executed by software rather than manually by a human evaluator. “Of these, 153 are programmatic”
  • Quality margin: The minimum score difference required for an output to be considered at least as good as a reference. “The test passes when the agent edit’s panel score exceeds the reference edit’s by at least the tie margin”
  • Rasch model: A statistical model that estimates the difficulty of tasks and the ability of evaluators or agents from response data. “If we model each human vote as depending on the agent’s skill and the task’s difficulty (a Rasch model)”
  • Reference edit: A completed video used as a comparison point for evaluating an agent’s output. “Human editors prefer the reference edit in 83.5% of judgments.”
  • Reframing: Altering the composition or crop of a shot to fit a different aspect ratio or emphasize a subject. “every UGC task requires captions, color matching and reframing for portrait delivery”
  • Room tone: The characteristic background sound recorded in a location when no dialogue or intentional action is occurring. “Ambience / room tone”
  • Shot-boundary detector: An algorithm that identifies transitions between separate shots in video. “cuts per minute from a shot-boundary detector”
  • Still-image sequence: An edited passage composed primarily of static photographs rather than moving footage. “built largely from still-image sequences, photo holds and controlled crop movement”
  • Take selection: The process of choosing the preferred recorded version of a scene or line from multiple takes. “Take selection”
  • Temporal reasoning: The ability to understand and coordinate events, durations, ordering, and relationships across time. “Together, these assignments test audiovisual understanding, temporal reasoning, sustained tool use”
  • Tie margin: The score difference used to determine whether an output is sufficiently better than a reference to pass a quality test. “a margin is fitted on a collection’s study pairs”
  • Timeline hygiene: The avoidance of technical or organizational problems in an edit, such as repeated footage, outtakes, or unwanted gaps. “Timeline hygiene”
  • Transcription: The conversion of spoken audio into written text. “transcribes its speech (Whisper small via faster-whisper”
  • VFX compositing: The combination of visual-effects elements with filmed footage to create a unified image. “VFX / compositing”
  • Win-or-tie rate: The proportion of evaluations in which an agent’s output is preferred or judged equivalent to the comparison output. “An agent’s win-or-tie rate, as in GDPval”
  • Z-score: A standardized score indicating how many standard deviations a value lies above or below a reference mean. “Because judges use the 1-to-10 scale differently, we convert each judge’s scores to z-scores”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 895 likes about this paper.