Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bi-FlowGS: Bidirectional Flow for Sparse-View 3D Reconstruction

Updated 18 September 2026
  • Bi-FlowGS is a sparse-view 3D reconstruction framework that couples generative video completion with explicit 3D Gaussian geometry optimization to reduce geometry cheating.
  • Its V2G module distills reliable optical flow from restored videos into differentiable depth and Gaussian-position updates, while G2V uses current geometry to guide future video restoration.
  • The method improves both appearance and geometry across several benchmarks, reporting up to 0.43 dB higher PSNR on Mip-NeRF 360 and roughly 48–51% geometry gains in complete-model ablations.

Bi-FlowGS is a framework for sparse-view 3D scene reconstruction that couples generative novel-view completion with explicit 3D Gaussian geometry optimization. Built on 3D Gaussian Splatting (3DGS), it addresses the underconstrained nature of reconstruction from a small number of images by using optical flow as an interface between restored videos and Gaussian geometry. Its two principal components are Video-to-Geometry Flow Distillation (V2G), which transfers temporal correspondence information from restored videos into differentiable Gaussian-depth optimization, and Geometry-to-Video Flow-Guided Restoration (G2V), which uses the current Gaussian geometry to condition video restoration. Their repeated interaction forms an implicit bidirectional co-refinement process intended to reduce the discrepancy between visually plausible renderings and geometrically incorrect scenes, a failure mode termed Geometry Cheating (Wang et al., 15 Sep 2026).

1. Problem formulation and conceptual basis

Sparse-view 3DGS is inherently underconstrained. A scene represented by anisotropic, colored, semi-transparent 3D Gaussians can contain erroneous Gaussian positions or depths while producing plausible RGB renderings. Opacity, scale, anisotropy, view-dependent appearance, overlapping splats, and neighboring Gaussian arrangements can compensate for incorrect geometry. Bi-FlowGS refers to this discrepancy between rendering quality and actual scene geometry as Geometry Cheating.

The problem is particularly pronounced in wide-baseline sparse-view settings, unbounded scenes, unseen regions, thin structures, occlusion boundaries, and repetitive or textureless areas. Conventional geometric regularization based on depth, semantics, smoothness, structure, or correspondence can constrain observed views, but does not directly supervise large unseen portions of a scene. RGB pseudo-supervision from generated views increases viewpoint coverage, yet RGB agreement alone can remain compatible with incorrect geometry.

Bi-FlowGS instead exploits the temporal correspondence structure of restored videos. A video generated along a camera trajectory contains not only RGB appearance but also camera-induced motion. If a restored video predicts that an image point moves by a particular amount between two viewpoints, the corresponding 3D geometry should induce a compatible reprojection displacement. Optical flow therefore provides a mechanism for converting generated-view motion into explicit constraints on Gaussian depth and spatial arrangement.

The framework is organized around two reciprocal mappings:

restored video→teacher optical flowGaussian geometry,\text{restored video} \xrightarrow{\text{teacher optical flow}} \text{Gaussian geometry},

and

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.

V2G transfers motion information from video to geometry, while G2V uses geometry to improve the temporal and structural consistency of video restoration. Repeated application is summarized as

better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.

The terminology distinguishes Bi-FlowGS from a general bidirectional optical-flow system: the bidirectionality refers to co-refinement between generative view completion and Gaussian geometry, rather than merely to forward and backward flow estimation.

2. 3D Gaussian representation and reconstruction pipeline

A Bi-FlowGS scene is represented as

G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},

where μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3} is the center of Gaussian nn, Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3} is its covariance, αn\alpha_n is opacity, and cn\mathbf{c}_n denotes color or view-dependent appearance parameters. The Gaussian density is written as

Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).

For camera Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.0, differentiable rasterization produces an RGB image and a depth image:

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.1

The depth image is central to V2G because it determines the geometry-induced correspondence between viewpoints.

Given sparse input images and known camera poses,

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.2

the system first constructs and optimizes an initial Gaussian scene. It then selects a camera trajectory and renders RGB and depth sequences. The rendered video is restored by G2V, and the restored frames are used both as RGB pseudo-views and as inputs to optical-flow estimation. V2G compares flow estimated from the restored frames with flow induced by the current Gaussian depth. The Gaussian parameters are then updated using real-view photometric loss, restored-view pseudo-view loss, and V2G flow alignment.

A complete iteration consists of:

  1. Initializing 3DGS from sparse images and poses.
  2. Rendering a camera trajectory from the current Gaussian scene.
  3. Restoring the rendered video with G2V.
  4. Using restored frames for RGB pseudo-view supervision and optical-flow estimation.
  5. Estimating teacher flow with frozen WAFT.
  6. Computing differentiable student flow by projecting Gaussian geometry between the same frame pair.
  7. Optimizing Gaussian parameters.
  8. Re-rendering the updated scene.
  9. Using improved depth, flow, and structural cues for subsequent video restoration.

The initial Gaussian optimization uses 7,000 iterations. During subsequent feedback optimization, rendered clips, depth maps, and Flow-Geometry conditions are periodically regenerated from the latest Gaussian scene. Dataset-dependent trajectories include elliptical trajectories for Mip-NeRF 360 and Tanks and Temples, and bounded circular trajectories for CO3D and DL3DV. The trajectory phase varies between rounds, while height variation is gradually reduced for more conservative refinement.

3. Video-to-Geometry Flow Distillation

Teacher-flow construction

Let a restored video along a known camera trajectory be

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.3

For a selected pair of frames Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.4, frozen WAFT estimates teacher optical flow:

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.5

For a source pixel Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.6, forward–backward consistency is assessed using

Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.7

The teacher-validity mask rejects invalid warps, excessively large flow, and excessive cycle error. The appendix specifies rejection of flow magnitudes above Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.8 pixels, cycle errors above Gaussian geometry→flow and depth conditioningrestored video.\text{Gaussian geometry} \xrightarrow{\text{flow and depth conditioning}} \text{restored video}.9, and confidence values below better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.0. Teacher flow, masks, and confidence are detached during Gaussian optimization so that the Gaussian scene cannot modify the supervision signal.

Geometry-induced student flow

For source camera better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.1 and target camera better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.2, a source pixel better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.3 is back-projected using the rendered source depth better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.4, transformed into the target camera, and projected into the target image:

better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.5

The geometry-induced student flow is

better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.6

With camera poses fixed, the main variable controlling better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.7 is rendered depth. Consequently, flow disagreement propagates through reprojection and differentiable depth rendering to Gaussian parameters, particularly Gaussian centers:

better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.8

A geometric-validity mask excludes invalid source depths, target projections outside the image, occluded correspondences, and other invalid reprojection cases. Target depth is used for validity and occlusion checks, while source depth remains differentiable.

Reliability-aware loss

The reliability weight is

better video⟹better teacher flow⟹better Gaussian geometry⟹better geometric guidance⟹better video.\text{better video} \Longrightarrow \text{better teacher flow} \Longrightarrow \text{better Gaussian geometry} \Longrightarrow \text{better geometric guidance} \Longrightarrow \text{better video}.9

where G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},0 is teacher reliability, G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},1 is geometric validity, G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},2 is teacher confidence, and G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},3. The V2G loss is

G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},4

Here G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},5 is Smooth-G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},6, applied to horizontal and vertical flow components, and G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},7 prevents division by zero. The reported configuration uses pixel-unit flow, Smooth-G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},8 parameter G={(μn,Σn,αn,cn)}n=1N,\mathcal{G} = \left\{ (\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n) \right\}_{n=1}^{N},9, one sampled frame pair per V2G update, and

μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}0

The V2G contribution is warmed up over the first 100 valid updates and capped at μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}1 times the current base reconstruction loss. These measures address the possibility that restored-video flow is unreliable because of hallucinated content, temporal inconsistencies, occlusions, textureless regions, repetitive patterns, or diffusion artifacts.

4. Geometry-to-Video Flow-Guided Restoration

G2V modifies video diffusion so that restored videos are conditioned on the current Gaussian scene’s geometry and motion. The restoration model is initialized from CogVideoX-5B-I2V and fine-tuned in latent space. The standard configuration uses 49 frames at μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}2 resolution, 50 DDIM denoising steps, and classifier-free guidance scale μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}3.

For a rendered clip

μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}4

the first and last frames serve as endpoint references μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}5 and μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}6. G2V receives the rendered control video, endpoint DINOv2 reference features, and a Flow-Geometry condition.

Flow-Geometry conditioning

Adjacent rendered-frame flow is estimated with frozen WAFT:

μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}7

Depth Anything 3 DA3MONO-LARGE estimates monocular depth. The Flow-Geometry condition contains a depth-consistency mask μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}8, structural boundaries μn∈R3\boldsymbol{\mu}_n\in\mathbb{R}^{3}9 derived from normalized monocular-depth gradients, and normalized flow components:

nn0

A FlowEncoder initialized from FloVD produces multiscale features

nn1

The depth-consistency mask identifies regions in which rendered geometry is trustworthy or unreliable, while the gradient map highlights boundaries and structural transitions.

Frozen DINOv2 ViT-L/14 extracts endpoint reference tokens:

nn2

These features preserve semantic identity and fine texture details from real endpoint observations.

Motion-consistent feature injection

Let nn3 denote Transformer tokens at block nn4. Flow features are injected before self-attention:

nn5

where nn6 controls the strength of flow conditioning. This mechanism encourages spatiotemporal propagation to follow camera-induced motion.

After self-attention, a global branch attends to all endpoint reference tokens:

nn7

where

nn8

A local branch uses flow-derived correspondence biases:

nn9

The temporal interpolation weights are

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}0

Thus, earlier frames rely more on the first endpoint and later frames more on the last. Global attention provides broad semantic and appearance consistency, whereas local attention retrieves correspondence-aware texture.

The fused feature is modulated by resized depth consistency:

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}1

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}2

The second factor increases reference conditioning in regions where rendered geometry is unreliable.

The diffusion model is trained with standard latent-space Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}3-prediction:

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}4

No additional pixel, perceptual, or optical-flow loss is used to train the diffusion restoration model.

5. Joint optimization and bidirectional co-refinement

The photometric loss applied to both real and restored views is

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}5

The complete Gaussian objective is

Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}6

Here Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}7 compares renders with the original sparse input images, Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}8 compares renders with restored pseudo-views, Σn∈R3×3\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}9 is a time-dependent annealed pseudo-view weight, and αn\alpha_n0 aligns geometry-induced and teacher flow.

The teacher flow is detached, whereas the geometry-induced flow remains differentiable. This separation prevents the Gaussian optimizer from manipulating the teacher while allowing flow disagreement to update geometry.

Bi-FlowGS uses V2G as a plug-and-play module. It was tested with GenFusion, ViewCrafter, and GSFixer, and requires a restored video, known camera poses, rendered Gaussian depth, an optical-flow estimator, and differentiable reprojection. The complete system also uses the 3DGS renderer and optimizer, DDIM sampling, classifier-free guidance, CogVideoX-5B-I2V, WAFT, FloVD FlowEncoder, DINOv2 ViT-L/14, Depth Anything 3 DA3MONO-LARGE, and BLIP-2 for text-prompt generation from the first reference frame. WAFT and DINOv2 are frozen in the relevant stages, while the G2V diffusion model is fine-tuned in two stages.

Training uses 800 scenes sampled from DL3DV-10K, initial reconstructions from 3, 6, or 9 views, two non-overlapping 49-frame clips per scene, and αn\alpha_n1 resolution. G2V training runs for 10,000 steps: during the first 2,000 steps, conditioning modules are trained while the diffusion backbone remains frozen; during the remaining 8,000 steps, Transformer self-attention and cross-normalization layers are additionally unfrozen. Optimization uses AdamW, BF16 precision, effective batch size 8, learning rates of αn\alpha_n2 for ordinary trainable parameters and αn\alpha_n3 for new scale and gating parameters, and two NVIDIA RTX PRO 6000 Blackwell GPUs.

6. Evaluation, results, and limitations

Bi-FlowGS is evaluated on DL3DV-Benchmark, Mip-NeRF 360, Tanks and Temples, and CO3D. Baselines include 3DGS, 2DGS, FSGS, ViewCrafter, GenFusion, Difix3D+, GSFixer, ZeroNVS, and ReconFusion in selected comparisons. Rendering is assessed with PSNR, SSIM, and LPIPS. Geometry is evaluated with depth metrics including AbsRel, SqRel, RMSE, RMSE-log, MedianRel, and threshold accuracy αn\alpha_n4. Tanks and Temples additionally uses point-cloud precision, recall, and F-score. Fixed SfM-track reprojection error is visualized to expose Geometry Cheating when RGB renderings appear similar but cross-view geometric consistency differs.

On Mip-NeRF 360, the reported results are:

| Views | PSNR | SSIM | LPIPS | |---:|---:|---:| | 3 | 16.08 | 0.388 | 0.541 | | 6 | 17.70 | 0.446 | 0.468 | | 9 | 19.06 | 0.501 | 0.413 |

The reported PSNR improvements over the strongest baseline are αn\alpha_n5 dB with 3 views, αn\alpha_n6 dB with 6 views, and αn\alpha_n7 dB with 9 views. Across Tanks and Temples, DL3DV-Benchmark, and CO3D, Bi-FlowGS ranks first in PSNR over the reported settings, with an average gain of approximately αn\alpha_n8 dB over the strongest baselines. It also obtains the best SSIM and LPIPS in most settings. Qualitative improvements are reported for thin structures, signs, railings, chair geometry, storefront details, and regions where incorrect depth causes cross-view inconsistency.

Ablations support the separation of roles between V2G and G2V. Inserting V2G into GenFusion, ViewCrafter, and GSFixer produces relative geometry gains ranging from αn\alpha_n9 to cn\mathbf{c}_n0, depending on baseline and view count, with an additional reported cost of approximately cn\mathbf{c}_n1–cn\mathbf{c}_n2 ms per step. For the complete model, geometry gains are reported as cn\mathbf{c}_n3, cn\mathbf{c}_n4, and cn\mathbf{c}_n5 for 3, 6, and 9 views, respectively.

On Mip-NeRF 360 with six views, removing V2G reduces PSNR from cn\mathbf{c}_n6 to cn\mathbf{c}_n7, SSIM from cn\mathbf{c}_n8 to cn\mathbf{c}_n9, and produces LPIPS degradation from Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).0 to Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).117.70Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).217.50Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).30.007Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).40.009Gn(x)=exp⁡(−12(x−μn)⊤Σn−1(x−μn)).G_n(\mathbf{x}) = \exp\left( -\frac{1}{2} (\mathbf{x}-\boldsymbol{\mu}_n)^\top \boldsymbol{\Sigma}_n^{-1} (\mathbf{x}-\boldsymbol{\mu}_n) \right).5\lambda_{\mathrm{V2G}}=0.014$; larger values may improve some geometry metrics while harming RGB quality and SSIM when teacher flow is inaccurate.

Several assumptions constrain interpretation. The reprojection formulation assumes a static scene and known camera poses; dynamic or nonrigid objects can be interpreted as geometric inconsistency. Camera-pose errors can corrupt both geometry-induced flow and trajectory generation. The method depends on large pretrained models and GPU-intensive latent video processing, including 50-step DDIM inference for 49-frame clips. Restored videos may contain hallucinated or temporally inconsistent content, and confidence filtering cannot eliminate all erroneous supervision. Because the co-refinement process is an alternating feedback system rather than a jointly convex optimization, no formal convergence guarantee is provided.

Bi-FlowGS is therefore best characterized as a correspondence-level coupling between generative view completion and differentiable Gaussian geometry. Its principal methodological distinction is that generated views are not used solely as RGB pseudo-targets: their motion structure is distilled into geometric supervision, while the evolving Gaussian scene simultaneously conditions subsequent video restoration. This design targets Geometry Cheating directly by imposing cross-view correspondence constraints on Gaussian depth and spatial arrangement, particularly in sparse-view, wide-baseline, and unbounded 360° reconstruction settings.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Bi-FlowGS.