Papers
Topics
Authors
Recent
Search
2000 character limit reached

Doki: Text-Native Generative Video Authoring

Updated 3 July 2026
  • Doki is a text-native interface for generative video authoring that consolidates asset definitions, scene structuring, and audio integration into one cohesive, editable document.
  • Its design leverages natural language for narrative construction, using @mentions and #hashtags to ensure parameterized, consistent, and gradual content refinement.
  • The modular architecture with LLM-driven prompt rewriting and generative pipelines accelerates ideation and streamlines multimodal video editing.

Doki is a text-native interface for generative video authoring that reconceptualizes the act of video creation as an interaction with a structured but freeform document. By centering the entire workflow in text—defining assets, structuring scenes, instantiating shots, refining edits, and embedding audio—Doki collapses the fragmented, tool-heavy landscape of traditional video editing into a contiguous, parameterized, and executable video script. This approach aims to leverage the affordances of writing as a creative medium and aligns video authoring practices with the input modalities most natural to large language and generative models (Liu et al., 10 Mar 2026).

1. Motivation and Design Principles

Conventional generative video workflows are characterized by a multiplicity of disconnected tools and formats: script editors, image reference managers, prompt engineering GUIs, and timeline-based non-linear editors (NLEs). These workflows present three core challenges: fragmentation, an overemphasis on prompt engineering at the expense of narrative, and drift in visual or stylistic consistency due to a lack of parameterized asset management.

Doki's text-first paradigm is guided by several explicit design principles:

  • Naturalness: Text-centric editing mirrors established narrative and creative practices, requiring no specialized video tooling expertise.
  • Model-Nativity: Text is already the lingua franca of generative systems; aligning authoring with this makes the system inherently compatible with current and future AI models.
  • Gradual Enrichment: Authors can begin with skeletal outlines, progressively refining content in situ.
  • Unification: Asset definitions, narrative structure, shot annotation, and audio spec are consolidated into a single document.
  • Parameterization: Reusable definitions (via @mentions and #hashtags) ensure content and style coherence.
  • Simplicity of Interface: The UI emphasizes minimalism, with text and structured blocks replacing complex panes and timelines (Liu et al., 10 Mar 2026).

2. System Architecture and Document Grammar

Doki is architected as a modular pipeline, with key components including a rich-text editor (TipTap), a parser and reference resolver, a prompt rewriter powered by an LLM, a visual reference manager, a generative image–video pipeline, state management (Zustand), agentic AI editing modules, and export/playback modules implemented in Node.js and FFmpeg.

The core data structure is a rich-text document parsed according to a grammar with constructs for headings, definitions, shot markers, paragraphs, and audio blocks. The canonical BNF-style grammar is:

1
2
3
4
5
6
7
8
9
10
11
12
13
Document       ::= Element*
Element        ::= Heading | Definition | Paragraph | AudioBlock
Heading        ::= ("#"^1-6) WS Text newline
Definition     ::= MentDef | HashDef
MentDef        ::= "@" Identifier "=" Text [InlinePreview] newline
HashDef        ::= "#" Identifier "=" Text newline
InlinePreview  ::= "[" "image://" URL_or_ID "]"
Paragraph      ::= Sentence+ newline
Sentence       ::= [ShotMarker] Text
ShotMarker     ::= "▢shot"
AudioBlock     ::= "[" AudioSpec "]"
Identifier     ::= ASCIIα (ASCIIα | ASCIIdigit | "_")*
Text           ::= any UTF-8 string not containing newline

Practical authoring occurs via slash-menu commands, which insert these constructs interactively. Definitions parameterize properties of characters, scenes, and styles with @mentions and #hashtags, propagating updates consistently across all dependent shots.

3. Generative Pipeline and Prompt Orchestration

Doki orchestrates several off-the-shelf generative models:

  • Prompt Rewriter: Processes structured prompts via Gemini 2.5 Flash (LLM).
  • Image Generation: Flux Kontext Pro (accepts text and up to four image references).
  • Video Generation: Veo 3 Fast (generating up to 8-second clips, optionally conditioned on audio).

Each shot ss undergoes a two-stage generative process:

  1. Static frame generation: I^s∼pimg(I∣prompts,{Ri})\hat I_s \sim p_{\mathrm{img}}(I \mid \mathrm{prompt}_s, \{R_i\})
  2. Video clip generation: V^s∼pvid(V∣I^s,prompts,{Ri})\hat V_s \sim p_{\mathrm{vid}}(V \mid \hat I_s, \mathrm{prompt}_s, \{R_i\})

Where prompts\mathrm{prompt}_s is the rewritten, LLM-processed shot prompt and {Ri}\{R_i\} is a set of visual references drawn from user-defined assets, previous shots, or LLM-inferred relevance. Prompt structuring follows a pipeline: initial user prompt → reference-resolved structured prompt → LLM-rewritten output. This workflow injects context, style, and definitions, preserving global project coherence.

4. Editing, State Management, and Refinement Operations

All edits are text-based or agentic: commands modify the document tree, and updates are reflected across the generative pipeline. Editing primitives include inserting shots (/shot), generating variants (/variant), embedding audio (/audio), and creating definitions (e.g., /define character). Inline refinement leverages the Inline Agent, offering sentence enhancement, creation of definitions from inline text, and executing arbitrary edit requests. A Sidebar Conversational Agent enables scoped, document-level transformations.

Document state is synchronized using explicit status flags for each shot node: image-generating, image-ready, video-generating, video-ready, or outdated. Edits to definitions propagate state invalidation, ensuring all downstream nodes are marked for regeneration. Regeneration logic is managed via an internal dependency-driven algorithm triggered by user interaction with paragraph handles.

5. Audio Annotation and Integration

Audio is specified via in-text markup blocks—e.g., [VOICE: ‘All aboard!’ in cheerful tone] [MUSIC: light ukulele strum]. When supported by the target video model (e.g., Veo 3 Fast), these are appended to prompts and passed as audio conditioning. Absence of explicit audio annotation results in default (usually ambient) audio generation. This embedded annotation approach maintains document primacy and reduces the context-switching inherent in timeline-based audio editing.

6. Empirical Study and Evaluation

A five-day diary study (N=10; UX designers, filmmakers, animators, and related backgrounds; mean age 32.4) evaluated Doki's usability and expressiveness. The procedure included an onboarding phase, daily diary logs (with satisfaction surveys, feature usage, pain points, and interaction telemetry), and an exit interview with System Usability Scale (SUS) administration. Observed metrics:

Metric Value
Avg. session length 91.7 min/session
Images per session 45.5
Videos per session 20.3
Mean SUS 81.2 (Excellent category)
Mean daily satisfaction 7.48/10

Qualitative findings:

  • Accelerated ideation: Participants required ~15 minutes to draft a 1-minute video, markedly outpacing traditional or prompt-based workflows.
  • Enhanced narrative comprehension: The structural document view provided semantic clarity over legacy timelines.
  • Consistency benefits: Reusable parameterization via @mentions and #hashtags mitigated prompt drift.
  • Accessibility: Novice users could author complex multi-shot stories with minimal learning curve.
  • Limitations: The approach revealed deficits in fine-grained camera control, temporal compositional expressivity (e.g., cross-dissolves, J/L-cuts), artifact remediation, and model output unpredictability (Liu et al., 10 Mar 2026).

7. Future Directions and Open Challenges

Doki’s authors highlight several directions and open design considerations:

  • Narrative scaffolding and diagnostics: Enhanced support for narrative structures, plot coherence, and pacing could raise the ceiling on story quality, bridging the gap between accessibility and sophisticated storytelling.
  • Temporal expressivity: The linear document model cannot natively express concurrency, cross-cuts, or non-overlapping audio events. Proposed extensions include inline timing syntax and explicit transition primitives (e.g., #CrossDissolve[12f]).
  • Unified document paradigm: The human–AI shared workspace—rendered as a dynamic, editable document—is posited as a general design pattern for future multimodal generative environments.
  • Comparative paradigms: Doki’s single-pane, document-centric model contrasts with compositional "bento-box" systems (script+storyboard+timeline), with differential advantages in immediacy, versioning, and parameterized co-authoring, but potential trade-offs in low-level control.

A plausible implication is that embracing text-native authoring as the primary interaction paradigm could fundamentally realign the accessibility, collaborative potential, and creative fluidity of generative video platforms (Liu et al., 10 Mar 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Doki.