Bi-FlowGS: Bidirectional Flow for Sparse-View 3D Reconstruction
- Bi-FlowGS is a sparse-view 3D reconstruction framework that couples generative video completion with explicit 3D Gaussian geometry optimization to reduce geometry cheating.
- Its V2G module distills reliable optical flow from restored videos into differentiable depth and Gaussian-position updates, while G2V uses current geometry to guide future video restoration.
- The method improves both appearance and geometry across several benchmarks, reporting up to 0.43 dB higher PSNR on Mip-NeRF 360 and roughly 48–51% geometry gains in complete-model ablations.
Bi-FlowGS is a framework for sparse-view 3D scene reconstruction that couples generative novel-view completion with explicit 3D Gaussian geometry optimization. Built on 3D Gaussian Splatting (3DGS), it addresses the underconstrained nature of reconstruction from a small number of images by using optical flow as an interface between restored videos and Gaussian geometry. Its two principal components are Video-to-Geometry Flow Distillation (V2G), which transfers temporal correspondence information from restored videos into differentiable Gaussian-depth optimization, and Geometry-to-Video Flow-Guided Restoration (G2V), which uses the current Gaussian geometry to condition video restoration. Their repeated interaction forms an implicit bidirectional co-refinement process intended to reduce the discrepancy between visually plausible renderings and geometrically incorrect scenes, a failure mode termed Geometry Cheating (Wang et al., 15 Sep 2026).
1. Problem formulation and conceptual basis
Sparse-view 3DGS is inherently underconstrained. A scene represented by anisotropic, colored, semi-transparent 3D Gaussians can contain erroneous Gaussian positions or depths while producing plausible RGB renderings. Opacity, scale, anisotropy, view-dependent appearance, overlapping splats, and neighboring Gaussian arrangements can compensate for incorrect geometry. Bi-FlowGS refers to this discrepancy between rendering quality and actual scene geometry as Geometry Cheating.
The problem is particularly pronounced in wide-baseline sparse-view settings, unbounded scenes, unseen regions, thin structures, occlusion boundaries, and repetitive or textureless areas. Conventional geometric regularization based on depth, semantics, smoothness, structure, or correspondence can constrain observed views, but does not directly supervise large unseen portions of a scene. RGB pseudo-supervision from generated views increases viewpoint coverage, yet RGB agreement alone can remain compatible with incorrect geometry.
Bi-FlowGS instead exploits the temporal correspondence structure of restored videos. A video generated along a camera trajectory contains not only RGB appearance but also camera-induced motion. If a restored video predicts that an image point moves by a particular amount between two viewpoints, the corresponding 3D geometry should induce a compatible reprojection displacement. Optical flow therefore provides a mechanism for converting generated-view motion into explicit constraints on Gaussian depth and spatial arrangement.
The framework is organized around two reciprocal mappings:
and
V2G transfers motion information from video to geometry, while G2V uses geometry to improve the temporal and structural consistency of video restoration. Repeated application is summarized as
The terminology distinguishes Bi-FlowGS from a general bidirectional optical-flow system: the bidirectionality refers to co-refinement between generative view completion and Gaussian geometry, rather than merely to forward and backward flow estimation.
2. 3D Gaussian representation and reconstruction pipeline
A Bi-FlowGS scene is represented as
where is the center of Gaussian , is its covariance, is opacity, and denotes color or view-dependent appearance parameters. The Gaussian density is written as
For camera 0, differentiable rasterization produces an RGB image and a depth image:
1
The depth image is central to V2G because it determines the geometry-induced correspondence between viewpoints.
Given sparse input images and known camera poses,
2
the system first constructs and optimizes an initial Gaussian scene. It then selects a camera trajectory and renders RGB and depth sequences. The rendered video is restored by G2V, and the restored frames are used both as RGB pseudo-views and as inputs to optical-flow estimation. V2G compares flow estimated from the restored frames with flow induced by the current Gaussian depth. The Gaussian parameters are then updated using real-view photometric loss, restored-view pseudo-view loss, and V2G flow alignment.
A complete iteration consists of:
- Initializing 3DGS from sparse images and poses.
- Rendering a camera trajectory from the current Gaussian scene.
- Restoring the rendered video with G2V.
- Using restored frames for RGB pseudo-view supervision and optical-flow estimation.
- Estimating teacher flow with frozen WAFT.
- Computing differentiable student flow by projecting Gaussian geometry between the same frame pair.
- Optimizing Gaussian parameters.
- Re-rendering the updated scene.
- Using improved depth, flow, and structural cues for subsequent video restoration.
The initial Gaussian optimization uses 7,000 iterations. During subsequent feedback optimization, rendered clips, depth maps, and Flow-Geometry conditions are periodically regenerated from the latest Gaussian scene. Dataset-dependent trajectories include elliptical trajectories for Mip-NeRF 360 and Tanks and Temples, and bounded circular trajectories for CO3D and DL3DV. The trajectory phase varies between rounds, while height variation is gradually reduced for more conservative refinement.
3. Video-to-Geometry Flow Distillation
Teacher-flow construction
Let a restored video along a known camera trajectory be
3
For a selected pair of frames 4, frozen WAFT estimates teacher optical flow:
5
For a source pixel 6, forward–backward consistency is assessed using
7
The teacher-validity mask rejects invalid warps, excessively large flow, and excessive cycle error. The appendix specifies rejection of flow magnitudes above 8 pixels, cycle errors above 9, and confidence values below 0. Teacher flow, masks, and confidence are detached during Gaussian optimization so that the Gaussian scene cannot modify the supervision signal.
Geometry-induced student flow
For source camera 1 and target camera 2, a source pixel 3 is back-projected using the rendered source depth 4, transformed into the target camera, and projected into the target image:
5
The geometry-induced student flow is
6
With camera poses fixed, the main variable controlling 7 is rendered depth. Consequently, flow disagreement propagates through reprojection and differentiable depth rendering to Gaussian parameters, particularly Gaussian centers:
8
A geometric-validity mask excludes invalid source depths, target projections outside the image, occluded correspondences, and other invalid reprojection cases. Target depth is used for validity and occlusion checks, while source depth remains differentiable.
Reliability-aware loss
The reliability weight is
9
where 0 is teacher reliability, 1 is geometric validity, 2 is teacher confidence, and 3. The V2G loss is
4
Here 5 is Smooth-6, applied to horizontal and vertical flow components, and 7 prevents division by zero. The reported configuration uses pixel-unit flow, Smooth-8 parameter 9, one sampled frame pair per V2G update, and
0
The V2G contribution is warmed up over the first 100 valid updates and capped at 1 times the current base reconstruction loss. These measures address the possibility that restored-video flow is unreliable because of hallucinated content, temporal inconsistencies, occlusions, textureless regions, repetitive patterns, or diffusion artifacts.
4. Geometry-to-Video Flow-Guided Restoration
G2V modifies video diffusion so that restored videos are conditioned on the current Gaussian scene’s geometry and motion. The restoration model is initialized from CogVideoX-5B-I2V and fine-tuned in latent space. The standard configuration uses 49 frames at 2 resolution, 50 DDIM denoising steps, and classifier-free guidance scale 3.
For a rendered clip
4
the first and last frames serve as endpoint references 5 and 6. G2V receives the rendered control video, endpoint DINOv2 reference features, and a Flow-Geometry condition.
Flow-Geometry conditioning
Adjacent rendered-frame flow is estimated with frozen WAFT:
7
Depth Anything 3 DA3MONO-LARGE estimates monocular depth. The Flow-Geometry condition contains a depth-consistency mask 8, structural boundaries 9 derived from normalized monocular-depth gradients, and normalized flow components:
0
A FlowEncoder initialized from FloVD produces multiscale features
1
The depth-consistency mask identifies regions in which rendered geometry is trustworthy or unreliable, while the gradient map highlights boundaries and structural transitions.
Frozen DINOv2 ViT-L/14 extracts endpoint reference tokens:
2
These features preserve semantic identity and fine texture details from real endpoint observations.
Motion-consistent feature injection
Let 3 denote Transformer tokens at block 4. Flow features are injected before self-attention:
5
where 6 controls the strength of flow conditioning. This mechanism encourages spatiotemporal propagation to follow camera-induced motion.
After self-attention, a global branch attends to all endpoint reference tokens:
7
where
8
A local branch uses flow-derived correspondence biases:
9
The temporal interpolation weights are
0
Thus, earlier frames rely more on the first endpoint and later frames more on the last. Global attention provides broad semantic and appearance consistency, whereas local attention retrieves correspondence-aware texture.
The fused feature is modulated by resized depth consistency:
1
2
The second factor increases reference conditioning in regions where rendered geometry is unreliable.
The diffusion model is trained with standard latent-space 3-prediction:
4
No additional pixel, perceptual, or optical-flow loss is used to train the diffusion restoration model.
5. Joint optimization and bidirectional co-refinement
The photometric loss applied to both real and restored views is
5
The complete Gaussian objective is
6
Here 7 compares renders with the original sparse input images, 8 compares renders with restored pseudo-views, 9 is a time-dependent annealed pseudo-view weight, and 0 aligns geometry-induced and teacher flow.
The teacher flow is detached, whereas the geometry-induced flow remains differentiable. This separation prevents the Gaussian optimizer from manipulating the teacher while allowing flow disagreement to update geometry.
Bi-FlowGS uses V2G as a plug-and-play module. It was tested with GenFusion, ViewCrafter, and GSFixer, and requires a restored video, known camera poses, rendered Gaussian depth, an optical-flow estimator, and differentiable reprojection. The complete system also uses the 3DGS renderer and optimizer, DDIM sampling, classifier-free guidance, CogVideoX-5B-I2V, WAFT, FloVD FlowEncoder, DINOv2 ViT-L/14, Depth Anything 3 DA3MONO-LARGE, and BLIP-2 for text-prompt generation from the first reference frame. WAFT and DINOv2 are frozen in the relevant stages, while the G2V diffusion model is fine-tuned in two stages.
Training uses 800 scenes sampled from DL3DV-10K, initial reconstructions from 3, 6, or 9 views, two non-overlapping 49-frame clips per scene, and 1 resolution. G2V training runs for 10,000 steps: during the first 2,000 steps, conditioning modules are trained while the diffusion backbone remains frozen; during the remaining 8,000 steps, Transformer self-attention and cross-normalization layers are additionally unfrozen. Optimization uses AdamW, BF16 precision, effective batch size 8, learning rates of 2 for ordinary trainable parameters and 3 for new scale and gating parameters, and two NVIDIA RTX PRO 6000 Blackwell GPUs.
6. Evaluation, results, and limitations
Bi-FlowGS is evaluated on DL3DV-Benchmark, Mip-NeRF 360, Tanks and Temples, and CO3D. Baselines include 3DGS, 2DGS, FSGS, ViewCrafter, GenFusion, Difix3D+, GSFixer, ZeroNVS, and ReconFusion in selected comparisons. Rendering is assessed with PSNR, SSIM, and LPIPS. Geometry is evaluated with depth metrics including AbsRel, SqRel, RMSE, RMSE-log, MedianRel, and threshold accuracy 4. Tanks and Temples additionally uses point-cloud precision, recall, and F-score. Fixed SfM-track reprojection error is visualized to expose Geometry Cheating when RGB renderings appear similar but cross-view geometric consistency differs.
On Mip-NeRF 360, the reported results are:
| Views | PSNR | SSIM | LPIPS | |---:|---:|---:| | 3 | 16.08 | 0.388 | 0.541 | | 6 | 17.70 | 0.446 | 0.468 | | 9 | 19.06 | 0.501 | 0.413 |
The reported PSNR improvements over the strongest baseline are 5 dB with 3 views, 6 dB with 6 views, and 7 dB with 9 views. Across Tanks and Temples, DL3DV-Benchmark, and CO3D, Bi-FlowGS ranks first in PSNR over the reported settings, with an average gain of approximately 8 dB over the strongest baselines. It also obtains the best SSIM and LPIPS in most settings. Qualitative improvements are reported for thin structures, signs, railings, chair geometry, storefront details, and regions where incorrect depth causes cross-view inconsistency.
Ablations support the separation of roles between V2G and G2V. Inserting V2G into GenFusion, ViewCrafter, and GSFixer produces relative geometry gains ranging from 9 to 0, depending on baseline and view count, with an additional reported cost of approximately 1–2 ms per step. For the complete model, geometry gains are reported as 3, 4, and 5 for 3, 6, and 9 views, respectively.
On Mip-NeRF 360 with six views, removing V2G reduces PSNR from 6 to 7, SSIM from 8 to 9, and produces LPIPS degradation from 0 to 117.70217.5030.00740.0095\lambda_{\mathrm{V2G}}=0.014$; larger values may improve some geometry metrics while harming RGB quality and SSIM when teacher flow is inaccurate.
Several assumptions constrain interpretation. The reprojection formulation assumes a static scene and known camera poses; dynamic or nonrigid objects can be interpreted as geometric inconsistency. Camera-pose errors can corrupt both geometry-induced flow and trajectory generation. The method depends on large pretrained models and GPU-intensive latent video processing, including 50-step DDIM inference for 49-frame clips. Restored videos may contain hallucinated or temporally inconsistent content, and confidence filtering cannot eliminate all erroneous supervision. Because the co-refinement process is an alternating feedback system rather than a jointly convex optimization, no formal convergence guarantee is provided.
Bi-FlowGS is therefore best characterized as a correspondence-level coupling between generative view completion and differentiable Gaussian geometry. Its principal methodological distinction is that generated views are not used solely as RGB pseudo-targets: their motion structure is distilled into geometric supervision, while the evolving Gaussian scene simultaneously conditions subsequent video restoration. This design targets Geometry Cheating directly by imposing cross-view correspondence constraints on Gaussian depth and spatial arrangement, particularly in sparse-view, wide-baseline, and unbounded 360° reconstruction settings.