LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
Abstract: Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: https://ziqi-ma.github.io/logo-website/
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces LoGo, a method for making AI-generated videos look more like a stable, believable 3D world.
Some AI video models can follow a moving camera, but they often make mistakes as the camera moves. For example:
- A door may change shape.
- A wall may suddenly tilt.
- An object may disappear and reappear.
- A chair may face a different direction.
- A whole room may change into another scene.
These problems are especially common in long videos, where the camera travels through a scene and later returns to places it has already seen.
The main idea of LoGo is to give the AI two kinds of feedback:
- Global feedback: Is the whole video consistent?
- Local feedback: Exactly which object or small area is inconsistent?
The name LoGo comes from combining local and global rewards.
2. What questions are the researchers asking?
The researchers want to answer questions such as:
- How can AI video models keep objects and buildings stable while the camera moves?
- Why do existing training methods miss small problems, such as a disappearing door?
- Is it better to give the model feedback about specific 3D areas instead of only one score for the entire video?
- Can a model become more consistent without making the video blurry or lowering its visual quality?
- Does this approach still work for long videos and complicated camera movements?
A major problem with older methods is that they give one score to the entire video. Imagine a student taking a test with 99 correct answers and 1 wrong answer. If the teacher only gives an overall score, the single mistake may not receive enough attention. LoGo instead tries to identify where the mistake happened.
3. How did the researchers do it?
Giving the video a 3D structure
First, the researchers used computer vision tools to estimate:
- How far away each part of the scene is, called depth.
- Where the camera was located and how it was moving.
- Which pixels in different video frames belong to the same part of the 3D scene.
They then placed the scene into many small 3D boxes, called voxels. A voxel is like a tiny cube in a 3D grid, similar to how a picture is made of small square pixels.
For example, a room might be divided into thousands of tiny cubes. Some cubes would contain parts of a wall, a table, or a plant.
Measuring mistakes in each 3D area
The researchers checked whether the same 3D area looked consistent from different camera positions. They compared:
- Color and appearance: Does the wall still have the same color and pattern?
- Depth: Is the wall still at the same distance and shape?
If a voxel had a large error, it might indicate that something had changed or gone wrong in that area.
This creates a kind of error map showing where the video has problems. For example, the map might highlight a door that changed shape while leaving the rest of the room unmarked.
Combining local and global rewards
The researchers used these error measurements as rewards during post-training.
Here, a reward is a score that tells the AI which results are better. The AI is trained to produce videos with higher rewards.
LoGo combines:
- A local reward, which focuses on individual 3D regions.
- A global reward, which checks the video as a whole.
Using only the local reward caused a problem: the model sometimes made videos smoother, blurrier, and less colorful because those changes reduced certain errors. This is called reward hacking—the model finds an easy way to improve its score without truly becoming better.
Using only the global reward also caused problems because small mistakes were hidden by the rest of the video.
LoGo solves this by blending both rewards. The researchers also took turns using rewards for:
- 3D consistency,
- Good-looking video,
- Correct camera movement.
Testing the method
The researchers tested LoGo on three different video-generation models:
- Lingbot2
- Lyra2
- UniWorld
They also compared it with earlier methods that use only a single score for the whole video.
They created a new test set called TrajectoryBench. It contains 2,000 examples with camera movements that:
- Travel far from the starting view.
- Return to previously seen areas from different angles.
- Pass through doors or other transitions into new spaces.
- Combine several camera movements into long videos.
The researchers also tested the method on the DL3DV dataset, which contains many real 3D scenes.
4. What did they find?
LoGo improved 3D consistency
Across the three video models, LoGo made objects and environments more stable.
It helped reduce problems such as:
- Objects appearing or disappearing.
- Walls and buildings changing position.
- Floating or distorted shapes.
- Objects changing their appearance.
- Entire scenes unexpectedly changing.
The improvements were especially strong on difficult, long camera paths.
For example, in one experiment, LoGo reduced a measurement called epipolar error by as much as 37% on the DL3DV dataset. Lower epipolar error means that the scene is more geometrically consistent when viewed from different camera positions.
On videos up to 400 frames long, LoGo’s advantage became even larger. Its reduction in epipolar error grew from about 24% at 80 frames to 62% at 400 frames in one analysis.
It worked better than global-only methods
The researchers compared LoGo with:
- The original video models.
- Earlier post-training methods.
- A version of their own method using only global feedback.
LoGo usually produced better results than all of these alternatives.
The results showed that local feedback was important because it allowed the model to focus on the exact location of a mistake instead of letting the mistake disappear inside an overall score.
It preserved video quality
Using only local rewards improved consistency but sometimes made videos blurry, gray, or less attractive.
LoGo avoided much of this problem by combining local and global rewards and by sometimes rewarding visual quality and camera accuracy.
In many tests, LoGo preserved or even improved:
- Video sharpness and attractiveness.
- Correct camera movement.
- Overall scene quality.
Local 3D boxes worked better than 2D patches
The researchers also tested other ways of giving local feedback, such as:
- Giving feedback to each video frame.
- Giving feedback to 2D image patches.
- Giving feedback to groups of frames.
- Giving feedback to 3D voxels.
The 3D voxel method worked best. This makes sense because the problem is really about keeping the same 3D object or place consistent, not just keeping the same rectangle in a picture consistent.
5. Why is this important?
The paper shows that AI video models need more than a single overall score. Small problems can be very important, even when most of the video looks correct.
LoGo is important because it could help create more reliable:
- Virtual worlds.
- Video games.
- Virtual reality experiences.
- 3D design tools.
- Training environments for robots.
- Interactive movie or animation scenes.
For example, if someone explores an AI-generated virtual museum, the paintings, doors, and walls should remain in the same places when the person walks around. LoGo could help make that experience feel more like a real world and less like a collection of changing pictures.
The new TrajectoryBench test set is also useful because it gives researchers a harder way to compare video models. It tests long journeys, revisiting places, and moving into new spaces instead of only testing simple camera movements.
Conclusion
This paper’s main message is that AI video models need specific feedback about where mistakes happen. A single score for the whole video can miss a problem with one object or part of a room.
LoGo combines:
- Global feedback to protect overall quality.
- Local 3D feedback to fix specific inconsistencies.
- Additional rewards to preserve camera control and attractive video.
The method produced more stable and believable scenes across several video models, especially for long and complicated camera movements. However, it still has limitations: videos longer than about 400 frames remain difficult, and the method mainly handles still scenes rather than scenes with moving people or objects.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Very long-horizon scalability remains unresolved: The evaluation reaches at most 400 frames, so it is unclear whether LoGo can maintain consistency over substantially longer rollouts or whether error accumulation eventually overwhelms the reward design.
- Memory and representation bottlenecks are not isolated from reward effects: The paper suggests that horizons beyond 400 frames may require new memory architectures, but it does not determine whether failures arise primarily from inadequate model memory, autoregressive drift, imperfect geometry estimation, or reward optimization.
- Dynamic scenes are outside the method’s scope: LoGo assumes static environments and does not address moving objects, articulated subjects, lighting changes, nonrigid deformation, or interactions between camera motion and scene motion.
- The method’s dependence on pretrained geometry estimators is insufficiently characterized: VGGT and Depth Anything 3 provide the depth, pose, and correspondence information used for rewards, but the paper does not quantify how their estimation errors affect training stability, reward reliability, or final consistency.
- Reward bias and reward hacking remain only partially addressed: Local rewards can encourage blurrier, less vibrant, or desaturated outputs, and the proposed global-local mixture does not establish whether other forms of geometric reward hacking—such as texture removal, object simplification, or scene alteration—remain possible.
- The local-global weighting is not systematically studied: The paper does not provide a broad sensitivity analysis of the blending coefficients, normalization strategy, clipping range, or reward-interleaving schedule, leaving unclear how these choices should be selected across models, datasets, and trajectory lengths.
- Voxel resolution has limited evaluation coverage: Voxel granularity is examined over a restricted range, but the effects of adaptive voxel sizes, nonuniform spatial resolution, sparse or highly distant scenes, and scenes containing objects at very different scales remain unexplored.
- Credit assignment is not evaluated under severe camera-pose estimation errors: The claim that 3D voxel localization is superior to frame-, patch-, and window-level localization may not hold when camera following is poor, poses are ambiguous, or the generated views have little overlap.
- Causal contributions of LoGo’s components are not fully disentangled: Although local rewards, global rewards, and reward interleaving are ablated, the study does not establish how much of the gain comes from 3D voxelization, reward normalization, additional training exposure, the particular post-training algorithm, or interactions among these components.
- Generalization beyond the evaluated base models is uncertain: Results are reported for three models, but all are large camera-controlled video systems with related training assumptions; transfer to smaller models, diffusion-only architectures, autoregressive systems with different memory mechanisms, or models trained on substantially different data is unknown.
- Generalization beyond the evaluated datasets and scene distributions is limited: TrajectoryBench and DL3DV emphasize selected indoor, outdoor, and transition scenarios, leaving performance on urban-scale environments, cluttered scenes, reflective or transparent materials, low-light settings, extreme depth ranges, and culturally diverse visual content unclear.
- TrajectoryBench’s coverage and validity require independent verification: The benchmark is newly introduced and partly constructed from existing trajectories and sourced images; its representativeness, difficulty calibration, annotation quality, and resistance to benchmark-specific optimization have not yet been established through external studies.
- Benchmark metrics may not fully reflect perceptual or physical consistency: Reprojection PSNR, depth metrics, Gaussian reconstruction, epipolar error, and learned video-quality scores can reward smooth or geometrically plausible outputs that are semantically wrong; the relationship between these metrics and human judgments of persistent objects and navigable worlds remains unresolved.
- Human evaluation is largely absent: The paper does not report systematic human assessments of object permanence, scene identity, visual realism, camera controllability, or preference tradeoffs, making it difficult to determine whether metric improvements correspond to meaningful user benefits.
- Semantic and relational consistency are not directly measured: The evaluation focuses primarily on geometric consistency and does not test whether object identity, attributes, affordances, spatial relationships, text, or object interactions remain stable across revisits.
- Failure modes are not comprehensively quantified: Qualitative examples show disappearing objects, hallucinations, floaters, and appearance changes, but the paper does not report per-category error rates, localization accuracy, or the frequency with which LoGo introduces new artifacts while fixing existing ones.
- Performance on unseen camera trajectories is unclear: The training and benchmark construction use structured trajectory types, but robustness to user-generated, irregular, discontinuous, highly accelerated, backward, rotational, or interactive camera commands is not established.
- Online and interactive use has not been evaluated: The method is assessed in offline post-training settings; latency, reward computation during deployment, adaptation to user corrections, and consistency under continually changing camera commands remain open questions.
- Computational and environmental costs are incompletely reported: The stated LoGo overhead excludes major costs such as rollout generation and repeated VGGT inference, so the total training energy, monetary cost, memory use, and scaling behavior for longer videos or larger datasets are not clear.
- Statistical robustness is limited in scope: Standard errors are reported, but the paper does not provide multi-seed training results, confidence intervals for all comparisons, or tests of whether improvements persist across different prompts, initial images, and random rollout samples.
- The interaction between LoGo and sampling diversity is unexplored: It is unclear whether optimizing localized consistency reduces the model’s ability to generate diverse valid scene interpretations or causes mode collapse toward conservative, low-detail reconstructions.
- The method’s behavior under ambiguous or inconsistent conditioning is unknown: The paper does not examine cases where the initial image, text prompt, and camera trajectory imply incompatible scene geometry or where the initial view provides insufficient information for later viewpoints.
- Dynamic-scene extensions lack a defined reward formulation: Although the conclusion suggests using 4D reconstruction methods, it does not specify how LoGo would distinguish legitimate motion from inconsistency, assign credit to moving regions, or handle occlusion and topology changes.
- Reproducibility of the full training pipeline remains uncertain: The paper releases code and checkpoints, but key practical details—including exact reward schedules, sampling procedures, hyperparameter selection, estimator versions, and data preprocessing—may be necessary to reproduce the reported gains across all three base models.
Practical Applications
Immediate Applications
- More reliable 3D-aware video-generation products — Software, media, and entertainment
- Integrate LoGo-style post-training into camera-controlled video generators to reduce disappearing objects, geometry shifts, floating artifacts, and inconsistent object appearance during long camera movements.
- Potential products include virtual cinematography tools, controllable storyboarding systems, game-asset previsualization, advertising-video generators, and interactive scene-generation platforms.
- The local-global reward combination is particularly useful because it improves consistency without the quality degradation observed when using only localized rewards.
- Dependencies: A compatible camera-controlled video model, estimated camera poses, depth prediction, and a 3D reconstruction component such as VGGT or an equivalent system. The reported gains are demonstrated mainly for static scenes and trajectories up to 400 frames.
- Automated quality control for generated videos — Media platforms and model providers
- Use voxel-level reprojection errors as a diagnostic system that identifies where a generated video is inconsistent, rather than assigning only one quality score to the entire clip.
- A production workflow could generate a 3D error heatmap, flag problematic objects or regions, and trigger regeneration or targeted post-processing.
- This could support human reviewers by prioritizing clips containing hallucinated objects, scene changes, geometry drift, or local artifacts.
- Dependencies: Reliable depth and camera-pose estimation; error thresholds must be calibrated for different visual styles, scene types, and generation lengths.
- Benchmarking and regression testing for world models — Academia and industrial R&D
- Adopt TrajectoryBench-style evaluations in model-development pipelines to test expansion away from the initial view, scene transitions, revisits from new angles, and complex multi-stage camera paths.
- Teams can use the benchmark as a release gate: a model must meet minimum thresholds for 3D consistency, camera following, and perceptual video quality before deployment.
- The benchmark is also suitable for tracking whether model updates improve long-horizon behavior or merely optimize short clips.
- Dependencies: Benchmark licensing and reproducibility, consistent evaluation models for depth and geometry, and safeguards against overfitting to the released trajectories.
- Post-training improvements for existing video models — AI infrastructure
- Apply LoGo as a modular reward layer rather than redesigning the entire video-generation architecture. The paper reports gains across three base models and several post-training algorithms.
- This makes the method suitable for model vendors that already use reinforcement learning, DPO-like optimization, DiffusionNFT, Flow-GRPO, or differentiable fine-tuning.
- The reported computational overhead of the localized reward calculation is small relative to rollout, geometry estimation, and backpropagation, although total training remains resource-intensive.
- Dependencies: Access to model weights and training infrastructure; the approach still requires substantial GPU resources for rollout and geometry estimation.
- Virtual production and camera-path previsualization — Film, television, and advertising
- Generate navigable rough scenes from a reference image and planned camera trajectory while preserving walls, furniture, props, and other spatial relationships.
- Directors and cinematographers could inspect alternative dolly, crane, orbit, or room-transition shots before physical production or expensive 3D asset creation.
- Localized consistency rewards are valuable when a small but important object—such as a sign, product, door, or actor prop—must remain stable throughout a shot.
- Dependencies: Static or near-static environments, sufficiently accurate camera trajectories, and human review before using generated content for final production.
- Educational and research tooling for 3D vision — Academia and education
- Use the voxelized reward pipeline as a teaching and experimentation framework for structure-from-motion, depth estimation, reprojection, reinforcement learning, and credit assignment.
- Students can compare global, framewise, patchwise, window-based, and voxel-based rewards and observe why 3D-space localization better identifies persistent object errors.
- TrajectoryBench can support coursework and reproducible research on long-horizon world models.
- Dependencies: Availability of code, checkpoints, datasets, and sufficient compute; evaluation results may depend on the selected depth and geometry estimators.
- Interactive virtual tours and visualization — Real estate, tourism, museums, and retail
- Generate navigable previews of interiors, galleries, retail spaces, or architectural concepts from limited visual input while maintaining more stable room layouts and object identities.
- Potential workflows include a user selecting a camera path, receiving a generated walkthrough, and revising the path interactively.
- The method may reduce distracting visual failures that undermine trust in virtual tours.
- Dependencies: The generated environment should not be treated as a metrically accurate survey without external validation. Privacy, consent, and copyright issues also apply to reference imagery.
- Consumer-facing image-to-video and virtual-world tools — Daily life and creative work
- Improve applications that animate a photograph into a navigable scene, such as family-memory visualizations, hobbyist filmmaking, room redesign previews, or game-world ideation.
- Users could specify commands such as “move through the doorway,” “orbit the table,” or “return to the original viewpoint” with fewer object and layout changes.
- Dependencies: User-facing systems require low latency, robust failure detection, and clear labeling that generated geometry may be visually plausible but not physically accurate.
Long-Term Applications
- Embodied AI and robotics simulation — Robotics
- Use LoGo-trained world models to create more persistent visual environments for robot navigation, manipulation planning, and policy training.
- A robot could simulate revisiting a room, moving through doors, or approaching an object from multiple angles while maintaining a more coherent spatial representation.
- Localized 3D rewards could help identify failures around task-critical regions, such as a handle, tool, obstacle, or object being grasped.
- Dependencies: The current work focuses on static visual generation rather than physical dynamics, action-conditioned prediction, collision validity, or metric accuracy. Integration with 4D dynamic-scene models and physically grounded simulators is required.
- Synthetic-data generation for autonomous vehicles and drones — Robotics and transportation
- Generate long-horizon camera trajectories through roads, buildings, or outdoor environments while preserving scene geometry and object permanence.
- Such data could supplement training and testing for visual odometry, mapping, route planning, and camera-control systems, especially for unusual viewpoints or difficult revisit paths.
- Dependencies: The method must be extended to dynamic objects, accurate lighting and weather variation, sensor realism, and safety-critical validation. Generated scenes cannot substitute for real-world testing without establishing strong sim-to-real evidence.
- Persistent interactive world models — Games, metaverse platforms, and spatial computing
- Build virtual environments in which users can move freely, revisit locations, and interact with objects without the scene changing unexpectedly.
- LoGo’s global reward can preserve overall camera behavior and visual quality, while its local reward can maintain specific objects and surfaces across extended navigation.
- Potential products include AI-generated game levels, virtual classrooms, collaborative 3D workspaces, and persistent augmented-reality environments.
- Dependencies: Horizons beyond 400 frames remain difficult according to the paper. Real-time memory, state management, controllable editing, multi-user synchronization, and dynamic-object modeling are unresolved requirements.
- Medical and scientific visualization — Healthcare and research
- In the longer term, consistency-aware generation could support exploratory navigation of reconstructed anatomy, laboratory environments, archaeological sites, or scientific simulations.
- A localized 3D reward could help preserve clinically or scientifically important structures when generating alternative viewpoints or educational walkthroughs.
- Dependencies: These applications require validated measurements, provenance, uncertainty estimates, and strict separation between visualization and diagnosis or scientific inference. Hallucinated structures would be unacceptable in clinical workflows.
- Construction, architecture, and urban planning — Architecture and engineering
- Generate and inspect camera-controlled previews of proposed buildings, interiors, infrastructure, or urban redesigns while retaining spatial relationships across complex walkthroughs.
- LoGo-like rewards could provide automated checks for disappearing rooms, shifting walls, inconsistent openings, or unstable design elements in generative design systems.
- Dependencies: Visual consistency is not equivalent to engineering correctness. Integration with CAD/BIM constraints, surveyed geometry, lighting simulation, building codes, and human architectural review would be necessary.
- Dynamic-scene consistency through 4D rewards — Future research
- Extend the voxelized local-global reward from static 3D scenes to time-varying 4D representations so that moving people, vehicles, animals, fluids, and deformable objects remain temporally coherent.
- A future system could distinguish legitimate motion from hallucinated appearance changes and assign credit to the specific moving regions that fail.
- Dependencies: The paper explicitly identifies dynamic scenes as an open direction. This requires reliable motion decomposition, dynamic-scene reconstruction, temporal identity tracking, and reward designs that do not penalize valid motion.
- Long-horizon agent planning and verification — AI agents
- Use localized visual rewards as intermediate feedback for agents that act through simulated or generated environments.
- Instead of rewarding only whether an entire episode succeeds, the system could identify which spatial region or step caused an environmental inconsistency, improving credit assignment for navigation and manipulation policies.
- Dependencies: This requires linking voxel-level visual errors to actions, task success, and causal agent behavior. Visual consistency alone does not guarantee correct planning or physical interaction.
- Policy and standards for evaluating generative world models — Public policy and industry governance
- Establish standardized reporting requirements for long-horizon generation, including separate scores for local object permanence, global layout stability, camera following, perceptual quality, and performance under revisits and scene transitions.
- TrajectoryBench’s structure could inform procurement standards, safety evaluations, and disclosure practices for systems marketed as “3D-consistent” or “world-generating.”
- Dependencies: Metrics based on estimated depth and reconstruction models can contain evaluator bias. Governance frameworks would need human studies, cross-model validation, robustness testing, and explicit distinction between visual plausibility and physical truth.
Glossary
- Autoregressive rollout: Sequential generation in which a model produces later content conditioned on previously generated content. “Recent efforts enable long-horizon generation via autoregressive rollout of flow-matching chunks”
- Backpropagation: An optimization procedure that computes gradients through a model to update its parameters. “Training time is dominated by the video rollout, VGGT invocation and backpropagation”
- Camera pose: The position and orientation of a camera in a scene. “Let P be the scene point cloud, cj the camera pose of frame j”
- Camera trajectory: A sequence of camera poses or movements used to control viewpoint changes during generation. “TrajectoryBench, a new benchmark for long-horizon video generation with complex camera trajectories”
- Camera following: The degree to which generated video conforms to a specified camera movement. “This local reward improves 3D consistency, and when combined with camera and aesthetic rewards it preserves or even improves video quality and camera following.”
- Camera-controlled video model: A video-generation model whose output is guided by camera movements or poses. “Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control.”
- Credit assignment: The process of determining which parts of a model’s output or actions contributed to a reward. “We propose LoGo, which enhances the global reward with a spatially localized reward.”
- Depth reprojection: The projection of estimated scene depth into another viewpoint for geometric comparison. “They apply DPO (Rafailov et al., 2023) or Flow-GRPO (Liu et al., 2026a) with geometry rewards such as epipolar constraints (Kupyn et al., 2025), depth reprojection (An et al., 2026; Du et al., 2026), or Gaussian reconstruction”
- Diffusion model: A generative model that produces data by progressively denoising a random representation. “We post-train three state-of-the-art camera-controlled video generation base models”
- DiffusionNFT: A post-training method that optimizes diffusion models using reward-based reinforcement learning. “Our main results use DiffusionNFT (Zheng et al., 2026)”
- DPO (Direct Preference Optimization): A method for fine-tuning models directly from preference data without explicitly training a separate reinforcement-learning policy. “They apply DPO (Rafailov et al., 2023) or Flow-GRPO (Liu et al., 2026a)”
- Epipolar constraint: A geometric relationship restricting where a corresponding point may appear between two camera views. “They apply DPO (Rafailov et al., 2023) or Flow-GRPO (Liu et al., 2026a) with geometry rewards such as epipolar constraints”
- Epipolar error: A measure of disagreement with the geometric relationship expected between corresponding points in different views. “LoGo effectively reduces local object shifts, artifacts, and global scene changes”
- Flow matching: A generative modeling framework that learns a vector field for transporting noise into data. “Recent efforts enable long-horizon generation via autoregressive rollout of flow-matching chunks”
- Flow-GRPO: A reinforcement-learning post-training algorithm designed for flow-matching models. “The FlowGRPO (Liu et al., 2026a) objective can be decomposed similarly.”
- Gaussian reconstruction: Reconstruction of a scene using 3D Gaussian primitives representing spatial appearance and geometry. “These approaches all collapse the reward to a single value for the entire generated video”
- Gaussian splatting: A rendering and scene-reconstruction technique that represents scenes with collections of three-dimensional Gaussian primitives. “They propose geometry-based reward models, such as epipolar constraint in Epipolar-DPO (Kupyn et al., 2025), depth reprojection in VG-GRPO (An et al., 2026), VideoGPA (Du et al., 2026), VIGOR (Yin et al., 2026), and GeoVideo (Bai et al., 2026), or Gaussian splatting (Kerbl et al., 2023) in World-R1”
- Global reward: A scalar evaluation computed over an entire generated video or scene. “In contrast, a global reward, defined as the average reprojection error across all pixels and frames, proves more robust to this degradation.”
- HPSv3: A learned human-preference score used to evaluate generated video quality. “The local reward (orange) sacrifices video quality (measured by HPSv3 (Ma et al., 2025))”
- Hallucinated object: An object generated without corresponding support in the intended scene or source imagery. “Our reward localization scheme can correctly capture each category of local inconsistency that we observe, including appearance inconsistency of the plant in the first example, hallucinated objects in the second example”
- Latent spatio-temporal patch: A localized region in a model’s internal representation indexed across spatial and temporal dimensions. “With LoGo, we assign a different r per latent spatio-temporal patch”
- Long-horizon generation: Generation of sequences substantially longer than the short clips typically used for evaluation or training. “To address the lack of 3D consistency evaluation for long-horizon generation, we introduce TrajectoryBench”
- MVCS: A depth-based metric for measuring consistency across multiple views. “For 3D consistency, we measure RGBD reprojection PSNR (Bai et al., 2026; Du et al., 2026), its depth-only variant MVCS”
- Optimality probability: A normalized reward value representing the estimated quality or preference of an output. “the optimality probability in DiffusionNFT (Zheng et al., 2026).”
- Pareto frontier: The boundary representing the best achievable trade-offs among competing objectives. “This blended local-global reward expands the Pareto frontier in the consistency-quality tradeoff”
- Pixel–voxel correspondence: The association between image pixels and the 3D voxels into which they are projected. “LoGo adds minimal overhead – only 2.1% of the total reward calculation time and 0.5% of the total per-step training time.”
- Point cloud: A collection of points representing the geometry of a 3D scene. “Given a video, we first construct a scene point cloud by passing keyframes to VGGT-Ω”
- Post-training: Additional optimization performed on a pretrained model, often using task-specific rewards or preferences. “We propose LoGo, a post-training approach that combines global and spatially localized rewards”
- Reprojection error: The discrepancy between an observed image measurement and the projection of a 3D estimate into the image. “each voxel’s error is defined as the average reprojection error of every pixel unprojected into the voxel.”
- Relative pose error (RPE): The difference between estimated and ground-truth relative camera motion, including translation and rotation. “For camera control, we report relative pose error (RPE): the VGGT-estimated camera translation and rotation between frames versus the ground truth.”
- Reward hacking: Optimization of a reward in a way that increases its numerical value while undermining the intended overall objective. “We attribute this to reward hacking”
- Reward interleaving: Alternating training steps that optimize different rewards for different desired properties. “We adopt a cyclic schedule: k steps with rLoGo, then m steps of aesthetic reward – HPSv3 (Ma et al., 2025) – and n”
- RGBD reprojection PSNR: A peak-signal-to-noise measure based on reprojection errors in both color and depth. “For 3D consistency, we measure RGBD reprojection PSNR (Bai et al., 2026; Du et al., 2026)”
- Scene point cloud: A 3D point representation of a scene constructed from video imagery and estimated geometry. “Given a video, we first construct a scene point cloud by passing keyframes to VGGT-Ω”
- Sparse 3D cache: A selectively stored three-dimensional representation used to support retrieval during generation. “Lyra-2 (Shen et al., 2026), a 14B model that generates 80 frames per chunk, with a sparse 3D cache that supports camera-based retrieval”
- Spatially localized reward: A reward assigned to specific spatial regions rather than to an entire generated sequence. “We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models.”
- Unprojection: The inverse projection of image pixels into 3D space using estimated depth and camera geometry. “unproject each pixel using the predicted depth into a shared coordinate system.”
- VGGT: A visual-geometry model used in the paper to estimate scene structure, depth, and camera information. “For a 240-frame video, this yields a median of roughly 3,000 voxels per scene”
- Voxel: A volumetric pixel representing a discrete region of three-dimensional space. “we voxelize this shared 3D space into a voxel grid”
- Voxel grid: A regular three-dimensional partition of space into discrete volumetric cells. “we voxelize this shared 3D space into a voxel grid, and calculate a spatially localized error per voxel”
- World model: A generative model intended to represent and produce coherent environments or scenes over time. “Very long-horizon generation (beyond 400 frames) remains challenging and might require new memory design, not just better reward design”