Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
Abstract: AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces TIMELINE-BENCH, a test designed to measure how well AI agents can perform real video-editing jobs.
Instead of asking an AI a question or giving it a small editing task, the researchers give it:
- Raw video and audio files
- A written project brief
- Editing tools
- Rules about the final video
The AI must then create a complete video, such as a short film scene, interview, advertisement, or travel video.
The main question is: Can AI make a finished video that is not only technically correct, but also interesting and professional to watch?
2. What did the researchers want to find out?
The researchers focused on several questions:
- Can AI agents complete realistic video-editing assignments from beginning to end?
- Can they understand a project brief and choose the right video clips, sounds, music, and graphics?
- Can they create a video with a good story, rhythm, and professional style?
- How do different AI models and computer tools compare?
- When AI fails, does it fail because of technical mistakes or because the video simply does not feel well edited?
This last question is especially important. A video may have the correct length and file type but still be boring, confusing, or poorly paced.
3. How was the research carried out?
The video-editing tasks
The researchers created 56 different editing tasks. These tasks came from several sources, including professional editing projects and publicly available practice projects.
The tasks covered many kinds of videos:
- Short narrative films
- Interviews and documentaries
- Advertisements
- Trailers
- Travel videos
- Social-media videos
- Product and brand videos
Each task included raw materials and instructions. For example, an AI might be told to create a 60-second advertisement using certain product shots, a voice-over, music, and text on screen.
The AI had to:
- Look through the available footage and audio
- Decide which parts to use
- Put the clips in the right order
- Add music, sound effects, captions, or graphics
- Export the final video in the required format
The AI systems
The study tested 16 AI agents. An agent is more than just an AI model: it is an AI model connected to tools that let it read files, write code, run programs, and create videos.
Some agents worked through coding tools such as:
- Codex
- Claude Code
- OpenCode
One version also received extra advice about professional video editing. Another version used a computer mouse and keyboard in DaVinci Resolve, a popular video-editing program.
The tests
The researchers used several kinds of tests.
Technical tests checked things such as:
- Does the video file exist?
- Is it in the correct format?
- Is it the right length?
- Does it contain working sound?
- Is the video free of frozen images, long silent sections, or repeated footage?
Brief tests checked whether the AI followed the instructions. For example:
- Did it use the required voice-over?
- Did it keep scenes in the correct order?
- Did it include the requested logo or text?
Some of these checks were done by computer programs. For example, speech-recognition software checked what was said, and OCR software read words appearing on screen. OCR is a tool that turns writing in an image into computer-readable text.
Human and AI judges
Technical tests cannot fully decide whether a video is good. There can be many correct ways to edit the same project.
Therefore, 43 professional video editors compared the AI videos with reference videos made by humans. They watched both versions without knowing which one was made by AI and chose the version they preferred.
The researchers also used three AI judges to score video quality. These AI judges looked at things such as:
- Story and structure
- Pacing, meaning how quickly the video moves
- Picture quality
- Sound
- Graphics and titles
The researchers adjusted the AI judging system so that its results matched human opinions as closely as possible.
4. What were the main findings?
AI agents completed relatively few tasks
The strongest system completed only 15 of the 56 tasks, or 26.8%.
Across all the systems, the average success rate was about 14%. This means that, on average, an AI successfully completed only about one task out of seven.
The best results came from GPT-6 Astra using Codex with extra editing advice. However, the researchers warn that the top systems were close enough that it is difficult to say that one was definitely better than the others.
Most failures were about quality, not basic correctness
The most important result was that 562 of 771 unsuccessful attempts passed the technical, content, and instruction tests but failed the quality test.
In simple terms, many AI systems created videos that:
- Had the right file format
- Used the required material
- Followed the basic instructions
- Did not contain obvious technical problems
But the videos still did not feel as polished or professional as the human reference edits.
This shows that making a video is not only about following rules. It also requires creative judgment.
Human editors strongly preferred the human reference videos
Professional editors preferred the human-made reference video in 83.5% of their judgments. They preferred the AI version in only 11.1% of cases, while the remaining judgments were ties.
Human editors especially noticed differences in:
- Overall polish
- Choosing the best shots
- Transitions
- Graphics and text
- Sound design
- Color
- Pacing
The AI often produced a clean basic assembly, but it did not make the many small creative decisions that make a video feel finished.
AI videos were often too slow
Compared with the reference edits, AI videos changed shots only about 72% as often. Their longest still images were about 1.5 times longer.
This suggests that AI often left shots on screen too long. The result could feel slow or less energetic, even when all the required scenes were included.
AI had difficulty judging its own work
The agents usually checked their own videos, but they mostly looked for technical problems:
- Missing files
- Broken audio
- Incorrect video length
- Repeated footage
- Rendering errors
Only about 5.5% of the problems they reported involved deeper creative issues such as story, pacing, or shot choice. By contrast, human editors mentioned these creative issues in about 66% of their comments.
Even more concerning, the agents claimed they had succeeded in 93% of their runs, including many videos that actually failed the tests. This means the agents were often unable to recognize that their own editing was weak.
The way AI viewed video was a major limitation
Most coding agents did not watch the footage in the same way a human editor does. They often examined:
- Still images taken from the video
- Contact sheets, which are collections of small preview images
- Written transcripts
- Audio volume measurements
This is like trying to understand a movie by looking at one photograph per second and reading a script, without properly watching and listening to the whole thing.
Because of this, the agents had trouble understanding:
- Natural timing
- Facial expressions
- Emotional moments
- The rhythm of a conversation
- Whether a cut felt smooth
- How music and images worked together
5. Why are these findings important?
The paper shows that current AI agents are fairly good at following explicit instructions and producing technically valid files. They can perform many basic editing operations.
However, they are still weak at the part of editing that human professionals contribute most: creative judgment.
A human editor does not simply place clips in an order. They decide:
- Which take has the best performance
- When to cut to another shot
- How long to hold a moment
- How to build emotion and suspense
- How music should support the story
- Whether the finished video feels natural and engaging
The research suggests that future video-editing AI will need several improvements:
- Better ability to watch and listen to complete videos
- Stronger understanding of emotion and storytelling
- Better tools for professional editing, sound, color, and graphics
- More useful ways to review its own work
- The ability to recognize creative problems, not only technical errors
Conclusion
TIMELINE-BENCH provides a realistic way to test AI video editors. It shows that current AI agents can often create videos that meet basic instructions, but they rarely match the quality of professional human editors.
The main lesson is that correctness is not the same as quality. An AI may deliver a video with the right length, clips, and file format while still making poor choices about pacing, storytelling, and style.
The benchmark could help researchers measure whether future AI systems genuinely improve. For now, the results suggest that AI can assist with parts of video editing, but human editors are still needed to make the final video feel polished, clear, and emotionally effective.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited task sample size: The benchmark contains only 56 tasks, so resolution-rate differences of a few tasks are not statistically distinguishable; larger task collections are needed for reliable agent ranking.
- No estimate of run-to-run variability: Each agent performs each task only once, leaving unresolved how much results depend on stochastic model behavior, sampling, transient tool failures, or different planning trajectories.
- Unclear generalizability beyond the sampled collections: Tasks come from EditStock, Cinestudy, commissioned projects, and UGC-style productions, but it is unknown whether the findings transfer to feature films, television, news, educational media, social-media content, live events, or other professional workflows.
- Potential source and selection bias: The composition of the benchmark, including which publicly available projects and reference edits were selected, may overrepresent particular genres, production standards, visual styles, or editorial conventions.
- Insufficient coverage of editing skills: Some capabilities, such as dual-system synchronization and VFX compositing, appear in only seven tasks, making it difficult to assess agent competence in these areas separately.
- No systematic analysis of task attributes: Within collections, the paper does not identify which measurable properties—footage volume, number of takes, dialogue density, ambiguity of the brief, audiovisual complexity, or required effects—cause task difficulty.
- Unresolved distinction between model and harness effects: Comparisons confound model capability with harness design, tool availability, prompting, context management, and provider-specific optimizations; controlled factorial experiments are needed to isolate these factors.
- Limited evaluation of computer-use agents: Only one computer-use configuration is tested, and it differs from coding agents in operating system, application, tool access, and interaction modality, preventing a clean comparison between GUI and programmatic editing.
- No systematic ablation of editorial guidance: The curated guidance condition changes only one agent and one harness, so it remains unclear which specific instructions, workflow recommendations, or perception strategies produce the observed improvement.
- Unclear causal role of audiovisual perception: The analysis associates reliance on still frames and transcripts with weak editorial craft, but does not experimentally test whether native video/audio playback, shot-level browsing, or richer temporal representations improve editing quality.
- No intervention study on iterative reviewing: Agents render late and check mainly for technical defects, but the benchmark does not test whether mandatory early previews, structured editorial review cycles, or targeted pacing/story checklists improve outcomes.
- Quality evaluation depends on reference edits: Reference edits are treated as comparison targets even though they are not certified ground truth and may represent only one valid creative interpretation of a brief.
- Possible mismatch between reference quality and task acceptability: The quality test requires an agent edit to exceed a collection-specific margin relative to the reference, which may penalize coherent alternatives that differ substantially in style, pacing, or structure.
- Calibration and validation are not fully independent: The model-based quality test is calibrated using the same human judgments used to validate it, creating a risk of overestimating agreement despite the leave-one-agent-out procedure.
- Quality-test validity is established per agent rather than per edit: The paper acknowledges that the evaluator is validated at the agent level, leaving uncertainty about whether it reliably classifies individual outputs, especially unusual or highly creative edits.
- Heavy reliance on multimodal LLM judges: Three model judges determine quality and some brief requirements, but their susceptibility to visual-style bias, familiar editing conventions, prompt sensitivity, and systematic disagreement with expert editors remains insufficiently characterized.
- Asymmetric judge inputs: One quality judge watches the videos with sound, whereas two judges receive contact sheets and file-level measurements; this heterogeneous observation protocol may distort the panel score and obscure which sensory information is necessary for evaluation.
- Human preference is measured against a one-line project description rather than the full brief: Editors may judge overall appeal without knowing explicit narrative, technical, or client requirements, so human win-or-tie rates may not fully represent task success.
- Low individual inter-rater agreement: Krippendorff’s indicates substantial judgment ambiguity, but the study does not investigate which task types or editorial dimensions produce disagreement.
- Limited expert-panel diversity: The study uses 43 editors with an average of at least two years of experience, but does not report detailed variation in seniority, specialization, geography, genre expertise, or professional standards.
- No comparison with novice, client, or non-editor audiences: It remains unknown whether the benchmark’s expert preferences align with client satisfaction, audience engagement, or judgments from less specialized viewers.
- Potential reference-order and presentation effects remain incompletely explored: Although order effects were reported as negligible, the hidden seek bar, forced full viewing, one-line description, and pairwise reference comparison may not reflect real commissioning or review workflows.
- No evaluation of edit usefulness in downstream production: The benchmark measures final-video tests and preference, but does not assess whether outputs are practically usable by clients, require substantial human revision, or reduce total production time and cost.
- No human–agent collaboration condition: The study evaluates autonomous agents, leaving open whether agents can meaningfully assist professional editors through candidate selection, rough cuts, revision suggestions, or targeted automation.
- Unclear relationship between technical cleanliness and creative quality: The paper reports that agents pass many delivery and content tests while failing quality, but does not quantify how much each technical defect contributes to human rejection or whether technical reliability enables later creative improvement.
- No evaluation of revision capability: Agents are not tested on responding to editorial feedback, client changes, deadline pressure, or successive rounds of revision—central properties of real editing work.
- No robustness testing under altered briefs or assets: It remains unknown whether agents can handle ambiguous, incomplete, contradictory, noisy, or deliberately revised briefs, nor how they respond to missing media, mislabeled files, poor transcripts, or corrupted assets.
- No adversarial or contamination analysis beyond reference isolation: The paper checks that agents did not retrieve reference edits, but does not examine training-data contamination, familiarity with benchmark source projects, or memorization of publicly available footage.
- Temporal reproducibility is uncertain: The evaluated models, harnesses, transcription systems, and model judges are rapidly changing, and provider outputs may not be stable, complicating longitudinal comparisons and replication.
- Cost–quality trade-offs are underexplored: The reported positive correlation between cost and resolution is based on few agents and observational comparisons; controlled studies are needed to determine whether additional inference time, tokens, or compute causally improve editorial performance.
- The benchmark does not establish a ceiling for human performance: Human editors compare outputs with references but are not evaluated on independently completing the same tasks, so the gap between agents, professional editors, and the best possible solution remains unknown.
- No assessment of creative diversity: Since the evaluation emphasizes similarity or superiority relative to reference edits, it does not measure whether agents can produce multiple valid stylistic solutions or adapt edits to different audiences and brand identities.
- Limited analysis of error propagation: The paper does not trace how early decisions—such as asset selection, transcript interpretation, or story outline—cause downstream problems in pacing, sound, graphics, and finishing.
- Open question about scalable quality verification: The benchmark’s human-calibrated quality procedure is expensive and tied to fixed collections; it remains unresolved whether reliable, collection-independent evaluators can be built for new tasks without fresh human judgments.
Practical Applications
Immediate Applications
The paper’s results support applications that use video agents as assistive, compliance-oriented, and quality-controlled tools, rather than as fully autonomous replacements for professional editors. The strongest current capabilities are delivery-format compliance, explicit brief requirements, transcription, basic assembly, and technical validation.
- Automated first cuts for media-production teams (film, documentary, advertising, and digital media)
- Workflow: ingest dailies → transcribe and index footage → generate a script-compliant assembly → run delivery and content tests → hand the timeline to a human editor.
- Assumptions/dependencies: footage must be well organized; the brief must state explicit requirements; human editors remain responsible for pacing, shot selection, graphics, sound design, and final approval.
- Automated technical quality assurance for video deliverables (broadcasting, streaming, advertising, education, and corporate communications)
- Potential product: a CI-style “video build verifier” integrated with editing systems or media asset-management platforms.
- Assumptions/dependencies: automated tests must be customized for each organization’s delivery specifications and validated against known-good and defective examples.
- Brief-compliance checking for post-production workflows (advertising, branded content, social media, and e-learning)
- Example: before publishing a commercial, the system confirms that the supplied voice-over is audible, the legal disclaimer appears on screen, the product is shown, and the final cut meets the duration window.
- Assumptions/dependencies: transcription and OCR errors must be monitored, especially for accents, noisy recordings, stylized text, multilingual content, and overlapping dialogue.
- Automated subtitle, caption, and social-format preparation (education, accessibility, marketing, and creator platforms)
- Potential tools: batch generation of captioned vertical clips, platform-specific aspect-ratio exports, and compliance checks for caption presence and timing.
- Assumptions/dependencies: captions require human or automated review for accuracy, speaker attribution, timing, accessibility standards, and sensitive content.
- Media-asset logging and editorial search (production companies, newsrooms, universities, and archives)
- Workflow: automatically create transcripts, thumbnails, shot boundaries, speaker labels, and metadata before a human begins editing.
- Assumptions/dependencies: this improves retrieval more reliably than creative judgment; visual understanding remains incomplete when agents inspect footage primarily through still images.
- Human-in-the-loop rough-cut assistance (small studios, agencies, independent creators, and corporate communications)
- Practical division of labor: the agent handles synchronization, organization, format conversion, basic scene assembly, and defect detection; the editor handles story, pace, performance, shot quality, sound design, and visual polish.
- Assumptions/dependencies: the interface must preserve editable timelines and source links rather than returning only a flattened video.
- Benchmark-based procurement and regression testing for media agents (software vendors, production houses, and research organizations)
- Potential workflow: evaluate candidate systems on representative tasks, track delivery-test and quality-test performance separately, and rerun the suite after model or tool updates.
- Assumptions/dependencies: the benchmark is a limited sample of 56 tasks; one run per task omits run-to-run variability, and the reference edits are not uniquely correct ground truth.
- Research training and evaluation for multimedia-agent developers (academia and industrial AI research)
- Useful metrics: resolution rate, human win-or-tie rate, brief-test failures, quality-test margin, cut frequency, static-shot duration, and the proportion of fixes addressing creative rather than technical defects.
- Assumptions/dependencies: model-judge scores should remain calibrated against independent human judgments and should not be treated as a perfect substitute for expert review.
- Personal and small-business video production (daily life, influencers, educators, and local organizations)
- Best current use: produce a publishable draft for low-risk content or a starting point for a human review.
- Assumptions/dependencies: users must check privacy, copyright, factual accuracy, unwanted speech, music licensing, and accidental inclusion of slates, retakes, or private material.
- Public-sector and institutional media compliance (government, healthcare communications, universities, and nonprofits)
- Potential product: an institutional media-governance pipeline that blocks noncompliant exports before publication.
- Assumptions/dependencies: legal and accessibility requirements vary by jurisdiction and must be encoded and periodically reviewed by domain specialists.
Long-Term Applications
The paper indicates that broader applications require advances in native audiovisual perception, creative self-evaluation, iterative editing, and human-calibrated quality assessment. The current resolution rate—26.8% for the strongest reported agent, with humans preferring reference edits in 83.5% of judgments—does not support unsupervised professional editing.
- Autonomous end-to-end video editors (film, advertising, television, and digital media)
- Required developments: continuous audio-visual viewing rather than primarily still frames and transcripts; multimodal timeline representations; stronger temporal reasoning; and the ability to generate and compare multiple editorial alternatives.
- Dependencies: reliable creative-quality evaluation, rights-aware asset handling, robust performance across genres, and acceptance by professional editors and clients.
- Interactive co-editors with creative critique and revision loops (professional editing software and collaborative production)
- Potential workflow: render early draft → watch the full result with synchronized sound → critique pacing, story, graphics, and audio → generate two or more alternatives → request targeted human approval.
- Dependencies: evaluations must measure whether proposed changes improve human preference, not merely whether they remove technical defects.
- Native audio-visual agents for long-form footage (documentary, news, sports, and reality production)
- Potential tools: synchronized audio-video embeddings, searchable performance-level indexes, speaker and emotion continuity tracking, and timeline-aware scene graphs.
- Assumptions/dependencies: models must avoid overinterpreting emotion or identity and must support privacy protections for people appearing in footage.
- Automated sound-design and finishing systems (cinema, advertising, games, podcasts with video, and immersive media)
- Expected benefit: reduce the finishing gap identified in the study, where human reviewers frequently cited polish, sound design, color, graphics, and transitions as advantages of reference edits.
- Dependencies: access to licensed music and effects libraries, accurate loudness and mixing standards, and safeguards against stylistic homogenization.
- Personalized multi-version content generation (marketing, education, healthcare communication, and accessibility)
- Examples: a long documentary condensed into classroom modules; a product video adapted for portrait social media; or a health-information video accompanied by captioned, translated, and audio-described variants.
- Dependencies: human review for factual claims, translation quality, cultural appropriateness, accessibility compliance, and preservation of required warnings or disclaimers.
- Policy and procurement standards for creative AI systems (government, media regulators, broadcasters, and enterprise buyers)
- Possible policy requirements: disclose benchmark coverage, report failure modes, preserve audit logs, test for unauthorized use of reference material, and require human approval for high-impact or public-facing content.
- Dependencies: benchmark representativeness, independent replication, protection of licensed media, and avoidance of treating one professional editing style as universal quality.
- Education and training simulators for editors (film schools, vocational programs, and corporate training)
- Potential product: a simulated post-production studio with briefs, raw footage, timeline tools, verifier tests, and structured critiques of story, pacing, sound, graphics, and color.
- Dependencies: expert-authored rubrics, diverse genres and cultures, careful separation between objective errors and legitimate stylistic choices, and protection against training students to optimize only for benchmark scores.
- Multi-agent production pipelines (large studios, advertising networks, and distributed creative teams)
- Potential workflow: asset analysis → narrative planning → rough assembly → sound and graphics passes → independent quality critique → human sign-off.
- Dependencies: shared timeline representations, conflict resolution, provenance tracking, consistent stylistic direction, and safeguards against compounding errors across agents.
- Creative-quality evaluation infrastructure for AI-generated media (research labs, platforms, and model developers)
- Potential capability: continuously updated quality panels that measure pacing, narrative coherence, visual continuity, audio quality, graphics, and audience suitability across task collections.
- Dependencies: independent human validation, protection against judge drift, representative reviewers, transparent uncertainty estimates, and recognition that quality is audience- and purpose-dependent.
- High-stakes automated communication production (healthcare, emergency response, finance, and public information)
- Required safeguards: approved-source retrieval, factual and legal verification, version control, clear provenance, accessibility review, and mandatory expert approval before release.
- Feasibility constraint: the paper demonstrates capability on editorial tasks, not factual safety or high-stakes communication; deployment in these sectors therefore requires separate domain-specific validation.
Glossary
- Adaptive detector: An algorithm that adjusts its detection criteria based on the characteristics of the input, such as changes in video scenes. “our re-implementation of PySceneDetect’s adaptive detector”
- Ambience: Background environmental sound that establishes the acoustic setting of a scene. “Ambience / room tone”
- Audio-visual perception: The ability to interpret visual and auditory information jointly. “The results suggest that creative benchmarks need human-calibrated quality evaluation, while capable agents need native audio-visual perception”
- Blind pairwise study: An evaluation in which participants compare two alternatives without knowing their identities or experimental conditions. “In a blind pairwise study, 43 professional video editors compared each delivered output with its task’s reference edit.”
- Color matching: Adjusting the colors of different shots so that they appear visually consistent. “every UGC task requires captions, color matching and reframing for portrait delivery”
- Contact sheet: A collection of representative images, usually arranged in a grid, used to inspect visual content efficiently. “they read contact sheets of one frame per second instead”
- Container: An isolated software environment that packages an application, its dependencies, and execution settings. “A TIMELINE-BENCH task consists of an edit brief, source assets, a Docker image, a set of tests, and a time limit”
- Content test: An automated or model-based check that verifies whether a video contains required material and avoids defective output. “Six content tests, shared by all tasks, reject degenerate edits such as mostly silent, frozen or looped ones”
- Cut rate: The frequency at which a video changes shots, commonly measured in cuts per minute. “agents change shots only 0.72 times as often as the reference edit”
- Dailies: Unedited footage recorded during a production, including multiple takes and production material. “Klug Brand Story asks for a sixty-second, interview-driven brand-story commercial from about four hours of dailies.”
- Degenerate edit: A technically produced but unusable video, such as one that is mostly silent, frozen, or repetitive. “Six content tests, shared by all tasks, reject degenerate edits”
- Delivery specification: The required technical properties of a media file, including codec, frame rate, duration, and audio format. “The tests assess whether the rendered video meets the technical specifications”
- Dual-system sync: Synchronizing separately recorded production audio with corresponding camera footage. “Dual-system sync”
- FFmpeg: A software suite and command-line framework for processing, converting, and analyzing audio and video. “The coding agents run in Linux containers with the tools an editor working in code needs: FFmpeg”
- Frozen input: An evaluation input that remains identical across runs to ensure comparability and reproducibility. “checklists for agent benchmarks call for frozen inputs and validated evaluators”
- Harbor format: A standardized packaging format for agent-evaluation tasks, including inputs, instructions, and tests. “Tasks are packaged in the Harbor format”
- Harness: A software system that supplies an AI model with tools and executes its commands. “An agent is a model running in a harness, the program that gives the model its tools and executes its commands.”
- Headless browser: A web browser controlled programmatically without a visible graphical interface. “Python and Node, the programmatic video frameworks Remotion and HyperFrames with a headless browser”
- Integrated loudness: A measurement of the average perceived loudness of an entire audio program. “integrated loudness (EBU R128, measured with FFmpeg)”
- Inter-rater reliability: The degree to which different evaluators produce consistent judgments when assessing the same material. “per-agent rates are reliable (Spearman–Brown reliability 0.84 with three human editors per pair)”
- Leave-one-agent-out: An evaluation procedure that excludes an agent’s own data when fitting parameters used to score that agent. “each agent is scored with margins refitted per collection without its own judgments (leave-one-agent-out).”
- Long-horizon task: A task requiring sustained planning and execution across many sequential operations. “AI agents increasingly carry out long-horizon professional work”
- Mechanical oracle: A deliberately constructed reference implementation or output used to validate whether automated tests behave correctly. “Each task therefore ships a mechanical oracle”
- Multimodal LLM: A LLM capable of processing multiple information modalities, such as text, images, audio, or video. “Three multimodal LLM judges, Gemini 3.8. Flash, GPT-6. Astra and Claude Opus 5.5”
- Non-linear editing: Video editing in which clips can be accessed and rearranged independently rather than processed sequentially from start to finish. “DaVinci Resolve 21, 2026. URL … Non-linear editing, color, effects and audio post-production software.”
- OCR: Optical character recognition, the automated conversion of text appearing in images or video frames into machine-readable text. “reads its on-screen text with OCR (PP-OCR via RapidOCR”
- Panel score: A combined evaluation score calculated by aggregating the normalized scores of multiple judges. “The panel score of an edit is the average of its three z-scores.”
- Pacing: The temporal rhythm of an edited video, including the speed and duration of shots and narrative progression. “only 5.5% of the problems agents report after a check concern pacing, story or shot choice”
- Picture lock: The stage at which the visual edit is considered final and no further picture changes are expected. “coordinating picture and sound over time”
- Portrait delivery: A video output designed for a vertically oriented display format. “the 41 landscape and 15 portrait deliverables”
- Post-production: The stage of media creation in which recorded material is edited, enhanced, mixed, and prepared for delivery. “post-production operations and GUI trajectories in media software”
- Programmatic test: A test executed by software rather than manually by a human evaluator. “Of these, 153 are programmatic”
- Quality margin: The minimum score difference required for an output to be considered at least as good as a reference. “The test passes when the agent edit’s panel score exceeds the reference edit’s by at least the tie margin”
- Rasch model: A statistical model that estimates the difficulty of tasks and the ability of evaluators or agents from response data. “If we model each human vote as depending on the agent’s skill and the task’s difficulty (a Rasch model)”
- Reference edit: A completed video used as a comparison point for evaluating an agent’s output. “Human editors prefer the reference edit in 83.5% of judgments.”
- Reframing: Altering the composition or crop of a shot to fit a different aspect ratio or emphasize a subject. “every UGC task requires captions, color matching and reframing for portrait delivery”
- Room tone: The characteristic background sound recorded in a location when no dialogue or intentional action is occurring. “Ambience / room tone”
- Shot-boundary detector: An algorithm that identifies transitions between separate shots in video. “cuts per minute from a shot-boundary detector”
- Still-image sequence: An edited passage composed primarily of static photographs rather than moving footage. “built largely from still-image sequences, photo holds and controlled crop movement”
- Take selection: The process of choosing the preferred recorded version of a scene or line from multiple takes. “Take selection”
- Temporal reasoning: The ability to understand and coordinate events, durations, ordering, and relationships across time. “Together, these assignments test audiovisual understanding, temporal reasoning, sustained tool use”
- Tie margin: The score difference used to determine whether an output is sufficiently better than a reference to pass a quality test. “a margin is fitted on a collection’s study pairs”
- Timeline hygiene: The avoidance of technical or organizational problems in an edit, such as repeated footage, outtakes, or unwanted gaps. “Timeline hygiene”
- Transcription: The conversion of spoken audio into written text. “transcribes its speech (Whisper small via faster-whisper”
- VFX compositing: The combination of visual-effects elements with filmed footage to create a unified image. “VFX / compositing”
- Win-or-tie rate: The proportion of evaluations in which an agent’s output is preferred or judged equivalent to the comparison output. “An agent’s win-or-tie rate, as in GDPval”
- Z-score: A standardized score indicating how many standard deviations a value lies above or below a reference mean. “Because judges use the 1-to-10 scale differently, we convert each judge’s scores to z-scores”