---
title: 'Doki: A Text-Native Video Interface'
url: https://www.emergentmind.com/topics/doki-a-text-native-interface
type: topic
---

# Doki: A Text-Native Video Interface

Doki is a text-native interface for generative video authoring, positioning free-form text as the primary medium for constructing, editing, and rendering visual narratives. Operating through a single document paradigm, Doki allows users to specify assets, structure scenes, define shots, insert and synchronize audio, and direct generative AI agents entirely via expressive textual commands. This approach fundamentally contrasts with traditional video editing paradigms, which rely on subtractive, multi-pane visual interfaces and separate scripting, editing, and rendering workflows. Doki's methodology synchronizes the natural act of text-based storytelling with advanced generative models, offering a unified, parameterized, and document-centered video creation pipeline [2603.09072].

## 1. Motivation and Background

Conventional video authoring environments—ranging from professional non-linear editors (e.g., Premiere Pro, Final Cut) to “text-in-timeline” systems like Descript—partition the creative process across disparate panels, timelines, and asset libraries. This fragmentation is shown to increase user cognitive load (Chandler & Sweller 1991) and disrupt narrative flow, especially when text serves only as an auxiliary input (e.g., dialogue editing) rather than as the structural substrate.

Contemporary generative video systems (e.g., Veo 3, Runway Gen4, Pika, Kling) inherit this divide by relegating text to single-shot prompts, after which manual curation and assembly are required to achieve narrative structure. Consistency across shots—regarding characters, lighting, and staging—often degrades as each generation remains isolated. Doki builds on prior advances in “dynamic documents” (Potluck, Embark) and novel video interfaces (VideOrigami, ExpressEdit) to redefine this model: the entire creative arc, from assets to the full timeline, is expressed and manipulated natively in text [2603.09072].

## 2. Design Principles

Doki’s design is guided by four explicit principles:

1. **Text-First Interaction (D1):** All elements—definitions, shots, audio—are written inline as part of the document. Both human and AI edits are directly visible and modifiable; prompt engineering remains fully embedded within story construction.
2. **Single Canonical Representation (D2):** Scripts, prompts, previews, timeline order, and audio cues reside in one document. Rearranging the document text (e.g., paragraph order) alters the rendered video sequence without the need for a separate timeline view.
3. **Consistency through Parameterization (D3):** Key elements (characters, styles, objects) are encapsulated as named definitions (e.g., @Hero, #wideShot), referenced throughout. Changes to definitions propagate globally, maintaining coherence in identity and aesthetics.
4. **Minimalist Interaction (D4):** All authoring occurs in a single editor (TipTap), with a lightweight slash menu, AI agents (conversational sidebar and inline bubble), and inline media previews supplanting multicomponent toolchains.

Formally, Doki’s DSL is specified in BNF as
```
<Document>      ::= <Definition>* <HeadingOrPara>*
<Definition>    ::= ( "@" <ID> "=" <Text> ) | ( "#" <ID> "=" <Text> )
<HeadingOrPara> ::= <Heading> | <Paragraph>
<Paragraph>     ::= <Sentence>+
<Sentence>      ::= <ShotNode> <Text>
<ShotNode>      ::= "/shot"
```
A mapping $f$: ShotPrompt × Context $\to$ VideoClip defines generation, where Context aggregates resolved definitions and preceding shot images [2603.09072].

## 3. System Architecture and User Workflow

Doki’s architecture comprises five primary components:

| Component            | Role                                          | Technologies/Models      |
|----------------------|-----------------------------------------------|--------------------------|
| Text Editor Frontend | Rich React/TipTap editor, slash menu, agents  | TipTap, React            |
| Parsing & Reference Resolver | Extracts/prompts, resolves mentions/hashtags  | Custom parser           |
| Prompt Rewriter      | Produces model-optimized prompts              | Gemini 2.5 Flash (LLM)   |
| Asset Manager        | Manages image/shot assets, propagates context | Internal, attached to editor |
| Video & Audio Renderer| Generates images, videos, inline audio       | Flux Kontext Pro, Veo 3 Fast |

The user defines assets as in:
```
@corgi = a cute corgi with golden-brown and white fur
@airport = a small regional airport terminal
```
and constructs video by writing paragraphs with `/shot` nodes, e.g.:
```
/shot @corgi arrives at the @airport with luggage
/shot @corgi then boards the plane
```
The parser resolves references, builds structured prompts (JSON trees of definition IDs and descriptions), and invokes the LLM-based prompt rewriter. Initial frame images are generated (Flux Kontext Pro) and sequenced into videos (Veo 3 Fast), with audio specified by bracketed annotation ([“narration”], [“SFX…”]) passed into the renderer [2603.09072].

## 4. Textual DSL, Parameterization, and Commands

Doki’s textual domain-specific language leverages slash commands and clear syntactic structures:

- `@ID = ...` for character, scene, object, or reference frame definitions
- `#ID = ...` for style, framing, camera move, mood
- `/shot ...` to begin a new shot prompt
- `/audio [...]` to insert synchronized audio annotations

References may be nested or parameterized. Context menus permit rapid generation of variants (e.g., alternate shot styles), and rearrangement of text equates to reordering rendered sequences. This document-as-video logic is central to maintaining narrative and visual coherence, with prompt reparametrization ensuring consistency throughout long-form stories. Bracketed audio directly synchronizes narration, music, or SFX at shot or paragraph granularity.

## 5. Generative Video and Audio Engine

Doki employs a pipeline built on state-of-the-art generative models:

- **Text/Prompt Rewriting:** Gemini 2.5 Flash (LLM)
- **Image Generation:** Flux Kontext Pro (flow-matching diffusion)
- **Video Synthesis:** Veo 3 Fast (latent diffusion with synchronized audio)

Objective functions are sequenced:
1. Minimize $\mathbb{L}_{img} = L_{diffusion}(r(p_{struct}), I_{gen})$ for image realism and semantic fidelity.
2. Minimize $\mathbb{L}_{vid} = L_{diffusion}(I_{prev}, V_{gen})$ for video continuity, using the latest generated image as seed.

Negative prompts are used to suppress undesired generative artifacts (e.g., on-screen text, random speech). *A plausible implication is that prompt quality and strict parameterization directly impact shot-to-shot consistency and the suppression of generative noise* [2603.09072].

## 6. Empirical Evaluation and Case Studies

A five-day diary study engaged 10 participants (6 F, 4 M; ages 24–57, spanning designers, animators, filmmakers), each using Doki ca. 50 min/day to author 2–3 videos:

- 46 videos (mean length 67.1 s, $\sigma=36$ s) produced
- Mean session duration: 91.7 min ($\sigma=45.5$)
- Per session: avg. 45.5 images, 20.3 videos generated

Feature usage frequencies were high: shot gen 90%, export 96%, hashtags 82%, mentions 78%, sidebar agent 74%, inline agent 50%. The System Usability Scale score was 81.2 ($\sigma=8.9$), considered “Excellent” (Bangor et al. 2008).

Qualitative results indicated rapid idea-to-draft conversion—novices produced five 30 s videos in ca. 10 min each. The unified document view aided global narrative comprehension, and parameterized definitions ensured referential consistency. Limitations were noted in model unpredictability (e.g., generative dropout such as “dogs would vanish mid-run”), difficulties in specifying cross-shot or temporal audio, and precise compositional control [2603.09072].

## 7. Benefits, Limitations, and Future Directions

Doki delivers several benefits:

- **Low entry barrier:** Text-literate users can quickly prototype videos without specialized training.
- **Single-document workflow:** Eliminates context switching and streamlines versioning.
- **Coherent parameterization:** Enforces narrative and visual consistency spanning multiple shots and scenes.
- **Human–AI collaboration:** Users author, AI agents execute—all within a unified text medium.

Limitations include generative model noise, limited temporal expressivity (challenging cross-shot sound or transitions), and insufficient precision for professional-grade needs (e.g., color grading, exact frames). Future work aims to introduce narrative scaffolding (e.g., three-act templates, pacing tools), richer temporal DSL constructs (inline timing, shot transitions), “compositional+text” hybrid interfaces for micro-editing, and robust intermediate representations to enable richer human–AI agency.

By making text not merely a prompt but the documentary ground truth of the video, Doki enables generative video to be constructed in the same domain as traditional literary storytelling—a shift in authoring modality with implications for both productivity and quality in AI-assisted content generation [2603.09072].

Source: https://www.emergentmind.com/topics/doki-a-text-native-interface