---
title: 'Bi-FlowGS: Bidirectional Flow for Sparse-View 3D Reconstruction'
url: https://www.emergentmind.com/topics/bi-flowgs
type: topic
---

# Bi-FlowGS: Bidirectional Flow for Sparse-View 3D Reconstruction

Bi-FlowGS is a framework for sparse-view 3D scene reconstruction that couples generative novel-view completion with explicit 3D Gaussian geometry optimization. Built on 3D Gaussian Splatting (3DGS), it addresses the underconstrained nature of reconstruction from a small number of images by using optical flow as an interface between restored videos and Gaussian geometry. Its two principal components are **Video-to-Geometry Flow Distillation (V2G)**, which transfers temporal correspondence information from restored videos into differentiable Gaussian-depth optimization, and **Geometry-to-Video Flow-Guided Restoration (G2V)**, which uses the current Gaussian geometry to condition video restoration. Their repeated interaction forms an implicit bidirectional co-refinement process intended to reduce the discrepancy between visually plausible renderings and geometrically incorrect scenes, a failure mode termed **Geometry Cheating** [2609.17039].

## 1. Problem formulation and conceptual basis

Sparse-view 3DGS is inherently underconstrained. A scene represented by anisotropic, colored, semi-transparent 3D Gaussians can contain erroneous Gaussian positions or depths while producing plausible RGB renderings. Opacity, scale, anisotropy, view-dependent appearance, overlapping splats, and neighboring Gaussian arrangements can compensate for incorrect geometry. Bi-FlowGS refers to this discrepancy between rendering quality and actual scene geometry as Geometry Cheating.

The problem is particularly pronounced in wide-baseline sparse-view settings, unbounded scenes, unseen regions, thin structures, occlusion boundaries, and repetitive or textureless areas. Conventional geometric regularization based on depth, semantics, smoothness, structure, or correspondence can constrain observed views, but does not directly supervise large unseen portions of a scene. RGB pseudo-supervision from generated views increases viewpoint coverage, yet RGB agreement alone can remain compatible with incorrect geometry.

Bi-FlowGS instead exploits the temporal correspondence structure of restored videos. A video generated along a camera trajectory contains not only RGB appearance but also camera-induced motion. If a restored video predicts that an image point moves by a particular amount between two viewpoints, the corresponding 3D geometry should induce a compatible reprojection displacement. Optical flow therefore provides a mechanism for converting generated-view motion into explicit constraints on Gaussian depth and spatial arrangement.

The framework is organized around two reciprocal mappings:

\[
\text{restored video}
\xrightarrow{\text{teacher optical flow}}
\text{Gaussian geometry},
\]

and

\[
\text{Gaussian geometry}
\xrightarrow{\text{flow and depth conditioning}}
\text{restored video}.
\]

V2G transfers motion information from video to geometry, while G2V uses geometry to improve the temporal and structural consistency of video restoration. Repeated application is summarized as

\[
\text{better video}
\Longrightarrow
\text{better teacher flow}
\Longrightarrow
\text{better Gaussian geometry}
\Longrightarrow
\text{better geometric guidance}
\Longrightarrow
\text{better video}.
\]

The terminology distinguishes Bi-FlowGS from a general bidirectional optical-flow system: the bidirectionality refers to co-refinement between generative view completion and Gaussian geometry, rather than merely to forward and backward flow estimation.

## 2. 3D Gaussian representation and reconstruction pipeline

A Bi-FlowGS scene is represented as

\[
\mathcal{G}
=
\left\{
(\boldsymbol{\mu}_n,\boldsymbol{\Sigma}_n,\alpha_n,\mathbf{c}_n)
\right\}_{n=1}^{N},
\]

where $\boldsymbol{\mu}_n\in\mathbb{R}^{3}$ is the center of Gaussian $n$, $\boldsymbol{\Sigma}_n\in\mathbb{R}^{3\times3}$ is its covariance, $\alpha_n$ is opacity, and $\mathbf{c}_n$ denotes color or view-dependent appearance parameters. The Gaussian density is written as

\[
G_n(\mathbf{x})
=
\exp\left(
-\frac{1}{2}
(\mathbf{x}-\boldsymbol{\mu}_n)^\top
\boldsymbol{\Sigma}_n^{-1}
(\mathbf{x}-\boldsymbol{\mu}_n)
\right).
\]

For camera $c_i$, differentiable rasterization produces an RGB image and a depth image:

\[
I_i^G=R_{\mathrm{rgb}}(\mathcal{G},c_i),
\qquad
D_i^G=R_{\mathrm{depth}}(\mathcal{G},c_i).
\]

The depth image is central to V2G because it determines the geometry-induced correspondence between viewpoints.

Given sparse input images and known camera poses,

\[
\{I_i\}_{i=1}^{K},
\qquad
\{P_i\}_{i=1}^{K},
\]

the system first constructs and optimizes an initial Gaussian scene. It then selects a camera trajectory and renders RGB and depth sequences. The rendered video is restored by G2V, and the restored frames are used both as RGB pseudo-views and as inputs to optical-flow estimation. V2G compares flow estimated from the restored frames with flow induced by the current Gaussian depth. The Gaussian parameters are then updated using real-view photometric loss, restored-view pseudo-view loss, and V2G flow alignment.

A complete iteration consists of:

1. Initializing 3DGS from sparse images and poses.
2. Rendering a camera trajectory from the current Gaussian scene.
3. Restoring the rendered video with G2V.
4. Using restored frames for RGB pseudo-view supervision and optical-flow estimation.
5. Estimating teacher flow with frozen WAFT.
6. Computing differentiable student flow by projecting Gaussian geometry between the same frame pair.
7. Optimizing Gaussian parameters.
8. Re-rendering the updated scene.
9. Using improved depth, flow, and structural cues for subsequent video restoration.

The initial Gaussian optimization uses 7,000 iterations. During subsequent feedback optimization, rendered clips, depth maps, and Flow-Geometry conditions are periodically regenerated from the latest Gaussian scene. Dataset-dependent trajectories include elliptical trajectories for Mip-NeRF 360 and Tanks and Temples, and bounded circular trajectories for CO3D and DL3DV. The trajectory phase varies between rounds, while height variation is gradually reduced for more conservative refinement.

## 3. Video-to-Geometry Flow Distillation

### Teacher-flow construction

Let a restored video along a known camera trajectory be

\[
\widehat{\mathbf V}
=
\{\widehat I_t\}_{t=1}^{T}.
\]

For a selected pair of frames $i,j$, frozen WAFT estimates teacher optical flow:

\[
\mathbf F^T_{i\rightarrow j}
=
\mathcal{F}_{\mathrm{WAFT}}(\widehat I_i,\widehat I_j).
\]

For a source pixel $\mathbf p$, forward–backward consistency is assessed using

\[
e_{\mathrm{cycle}}(\mathbf p)
=
\left\|
\mathbf F_{i\rightarrow j}(\mathbf p)
+
\mathbf F_{j\rightarrow i}
\left(
\mathbf p+\mathbf F_{i\rightarrow j}(\mathbf p)
\right)
\right\|_2.
\]

The teacher-validity mask rejects invalid warps, excessively large flow, and excessive cycle error. The appendix specifies rejection of flow magnitudes above $96$ pixels, cycle errors above $0.08\min(H,W)$, and confidence values below $0.4$. Teacher flow, masks, and confidence are detached during Gaussian optimization so that the Gaussian scene cannot modify the supervision signal.

### Geometry-induced student flow

For source camera $c_i$ and target camera $c_j$, a source pixel $\mathbf p$ is back-projected using the rendered source depth $D_i^G(\mathbf p)$, transformed into the target camera, and projected into the target image:

\[
\mathbf p'
=
\pi_j
\left(
T_jT_i^{-1}
\pi_i^{-1}
\left(
\mathbf p,D_i^G(\mathbf p)
\right)
\right).
\]

The geometry-induced student flow is

\[
\mathbf F^S_{i\rightarrow j}(\mathbf p)
=
\mathbf p'-\mathbf p.
\]

With camera poses fixed, the main variable controlling $\mathbf F^S$ is rendered depth. Consequently, flow disagreement propagates through reprojection and differentiable depth rendering to Gaussian parameters, particularly Gaussian centers:

\[
\frac{\partial \mathcal L_{\mathrm{V2G}}}
{\partial \boldsymbol{\Theta}_{\mathcal G}}
=
\frac{\partial \mathcal L_{\mathrm{V2G}}}
{\partial \mathbf F^S_{i\rightarrow j}}
\frac{\partial \mathbf F^S_{i\rightarrow j}}
{\partial D_i^G}
\frac{\partial D_i^G}
{\partial \boldsymbol{\Theta}_{\mathcal G}}.
\]

A geometric-validity mask excludes invalid source depths, target projections outside the image, occluded correspondences, and other invalid reprojection cases. Target depth is used for validity and occlusion checks, while source depth remains differentiable.

### Reliability-aware loss

The reliability weight is

\[
\mathbf W(\mathbf p)
=
\mathbf M^T(\mathbf p)
\mathbf M^G(\mathbf p)
\left(\mathbf C^T(\mathbf p)\right)^\gamma,
\]

where $\mathbf M^T$ is teacher reliability, $\mathbf M^G$ is geometric validity, $\mathbf C^T$ is teacher confidence, and $\gamma=1.5$. The V2G loss is

\[
\mathcal L_{\mathrm{V2G}}
=
\frac{
\displaystyle
\sum_{\mathbf p}
\mathbf W(\mathbf p)
\rho
\left(
\mathbf F^S_{i\rightarrow j}(\mathbf p)
-
\mathbf F^T_{i\rightarrow j}(\mathbf p)
\right)
}{
\displaystyle
\sum_{\mathbf p}\mathbf W(\mathbf p)+\epsilon
}.
\]

Here $\rho$ is Smooth-$L_1$, applied to horizontal and vertical flow components, and $\epsilon$ prevents division by zero. The reported configuration uses pixel-unit flow, Smooth-$L_1$ parameter $\beta=1$, one sampled frame pair per V2G update, and

\[
\lambda_{\mathrm{V2G}}=0.014.
\]

The V2G contribution is warmed up over the first 100 valid updates and capped at $0.9$ times the current base reconstruction loss. These measures address the possibility that restored-video flow is unreliable because of hallucinated content, temporal inconsistencies, occlusions, textureless regions, repetitive patterns, or diffusion artifacts.

## 4. Geometry-to-Video Flow-Guided Restoration

G2V modifies video diffusion so that restored videos are conditioned on the current Gaussian scene’s geometry and motion. The restoration model is initialized from CogVideoX-5B-I2V and fine-tuned in latent space. The standard configuration uses 49 frames at $480\times720$ resolution, 50 DDIM denoising steps, and classifier-free guidance scale $6.0$.

For a rendered clip

\[
\mathbf V=\{I_t\}_{t=1}^{T},
\]

the first and last frames serve as endpoint references $I^0$ and $I^1$. G2V receives the rendered control video, endpoint DINOv2 reference features, and a Flow-Geometry condition.

### Flow-Geometry conditioning

Adjacent rendered-frame flow is estimated with frozen WAFT:

\[
\mathbf F_t
=
(\mathbf F_{x,t},\mathbf F_{y,t}),
\qquad
t=1,\ldots,T-1.
\]

Depth Anything 3 DA3MONO-LARGE estimates monocular depth. The Flow-Geometry condition contains a depth-consistency mask $\mathbf M^D$, structural boundaries $\mathbf G^D$ derived from normalized monocular-depth gradients, and normalized flow components:

\[
\mathbf Q^F
=
\operatorname{Concat}
\left(
\mathbf M^D,
\mathbf G^D,
\mathbf F_x,
\mathbf F_y
\right).
\]

A FlowEncoder initialized from FloVD produces multiscale features

\[
\{\boldsymbol{\Phi}_l^F\}_{l=1}^{S}
=
E_{\mathrm{flow}}(\mathbf Q^F).
\]

The depth-consistency mask identifies regions in which rendered geometry is trustworthy or unreliable, while the gradient map highlights boundaries and structural transitions.

Frozen DINOv2 ViT-L/14 extracts endpoint reference tokens:

\[
\mathbf S^r
=
E_{\mathrm{DINO}}(I^r),
\qquad
r\in\{0,1\}.
\]

These features preserve semantic identity and fine texture details from real endpoint observations.

### Motion-consistent feature injection

Let $\mathbf H_l$ denote Transformer tokens at block $l$. Flow features are injected before self-attention:

\[
\mathbf H_l^{\mathrm{SA}}
=
\mathbf H_l
+
\operatorname{Attention}^{\mathrm{SA}_l}
\left(
\operatorname{Norm}_1(\mathbf H_l)
+
\gamma_l\boldsymbol{\Phi}_l^F
\right),
\]

where $\gamma_l$ controls the strength of flow conditioning. This mechanism encourages spatiotemporal propagation to follow camera-induced motion.

After self-attention, a global branch attends to all endpoint reference tokens:

\[
\mathbf A_l^{\mathrm{global}}
=
\operatorname{CrossAttention}_l
\left(
\mathbf X_l,[\mathbf S^0;\mathbf S^1]
\right),
\]

where

\[
\mathbf X_l
=
\operatorname{Norm}_{\mathrm{cross}}
\left(
\mathbf H_l^{\mathrm{SA}}
\right).
\]

A local branch uses flow-derived correspondence biases:

\[
\mathbf A_{l,t}^{\mathrm{local}}
=
\sum_{r=0}^{1}
w_t^r
\operatorname{CrossAttention}_l
\left(
\mathbf X_{l,t},
\mathbf S^r;
\mathbf B_{l,t}^{r,F}
\right).
\]

The temporal interpolation weights are

\[
\tau_t=\frac{t-1}{T-1},
\qquad
w_t^0=1-\tau_t,
\qquad
w_t^1=\tau_t.
\]

Thus, earlier frames rely more on the first endpoint and later frames more on the last. Global attention provides broad semantic and appearance consistency, whereas local attention retrieves correspondence-aware texture.

The fused feature is modulated by resized depth consistency:

\[
\widehat{\mathbf M}_l^D
=
\operatorname{Resize}(\mathbf M^D),
\]

\[
\mathbf A_l
=
\left[
\mathbf A_l^{\mathrm{global}}
+
\alpha_l^{\mathrm{local}}
\left(
\mathbf A_l^{\mathrm{local}}
-
\mathbf A_l^{\mathrm{global}}
\right)
\right]
\odot
\left[
1+
\alpha_l^{\mathrm{conf}}
\left(
1-\widehat{\mathbf M}_l^D
\right)
\right].
\]

The second factor increases reference conditioning in regions where rendered geometry is unreliable.

The diffusion model is trained with standard latent-space $v$-prediction:

\[
\mathcal L_{\mathrm{diff}}
=
\mathbb E_{\mathbf z_0,\boldsymbol{\epsilon},t}
\left[
\left\|
v_\theta(\mathbf z_t,t,\mathcal C)-v_t
\right\|_2^2
\right].
\]

No additional pixel, perceptual, or optical-flow loss is used to train the diffusion restoration model.

## 5. Joint optimization and bidirectional co-refinement

The photometric loss applied to both real and restored views is

\[
\mathcal L_{\mathrm{photo}}
=
(1-\lambda_{\mathrm{SSIM}})\mathcal L_1
+
\lambda_{\mathrm{SSIM}}
(1-\operatorname{SSIM}).
\]

The complete Gaussian objective is

\[
\mathcal L
=
\mathcal L_{\mathrm{real}}
+
\lambda_{\mathrm{pseudo}}(t)\mathcal L_{\mathrm{pseudo}}
+
\lambda_{\mathrm{V2G}}\mathcal L_{\mathrm{V2G}}.
\]

Here $\mathcal L_{\mathrm{real}}$ compares renders with the original sparse input images, $\mathcal L_{\mathrm{pseudo}}$ compares renders with restored pseudo-views, $\lambda_{\mathrm{pseudo}}(t)$ is a time-dependent annealed pseudo-view weight, and $\mathcal L_{\mathrm{V2G}}$ aligns geometry-induced and teacher flow.

The teacher flow is detached, whereas the geometry-induced flow remains differentiable. This separation prevents the Gaussian optimizer from manipulating the teacher while allowing flow disagreement to update geometry.

Bi-FlowGS uses V2G as a plug-and-play module. It was tested with GenFusion, ViewCrafter, and GSFixer, and requires a restored video, known camera poses, rendered Gaussian depth, an optical-flow estimator, and differentiable reprojection. The complete system also uses the 3DGS renderer and optimizer, DDIM sampling, classifier-free guidance, CogVideoX-5B-I2V, WAFT, FloVD FlowEncoder, DINOv2 ViT-L/14, Depth Anything 3 DA3MONO-LARGE, and BLIP-2 for text-prompt generation from the first reference frame. WAFT and DINOv2 are frozen in the relevant stages, while the G2V diffusion model is fine-tuned in two stages.

Training uses 800 scenes sampled from DL3DV-10K, initial reconstructions from 3, 6, or 9 views, two non-overlapping 49-frame clips per scene, and $480\times720$ resolution. G2V training runs for 10,000 steps: during the first 2,000 steps, conditioning modules are trained while the diffusion backbone remains frozen; during the remaining 8,000 steps, Transformer self-attention and cross-normalization layers are additionally unfrozen. Optimization uses AdamW, BF16 precision, effective batch size 8, learning rates of $2\times10^{-5}$ for ordinary trainable parameters and $1\times10^{-4}$ for new scale and gating parameters, and two NVIDIA RTX PRO 6000 Blackwell GPUs.

## 6. Evaluation, results, and limitations

Bi-FlowGS is evaluated on DL3DV-Benchmark, Mip-NeRF 360, Tanks and Temples, and CO3D. Baselines include 3DGS, 2DGS, FSGS, ViewCrafter, GenFusion, Difix3D+, GSFixer, ZeroNVS, and ReconFusion in selected comparisons. Rendering is assessed with PSNR, SSIM, and LPIPS. Geometry is evaluated with depth metrics including AbsRel, SqRel, RMSE, RMSE-log, MedianRel, and threshold accuracy $\delta_k$. Tanks and Temples additionally uses point-cloud precision, recall, and F-score. Fixed SfM-track reprojection error is visualized to expose Geometry Cheating when RGB renderings appear similar but cross-view geometric consistency differs.

On Mip-NeRF 360, the reported results are:

| Views | PSNR | SSIM | LPIPS |
|---:|---:|---:|
| 3 | 16.08 | 0.388 | 0.541 |
| 6 | 17.70 | 0.446 | 0.468 |
| 9 | 19.06 | 0.501 | 0.413 |

The reported PSNR improvements over the strongest baseline are $+0.43$ dB with 3 views, $+0.34$ dB with 6 views, and $+0.38$ dB with 9 views. Across Tanks and Temples, DL3DV-Benchmark, and CO3D, Bi-FlowGS ranks first in PSNR over the reported settings, with an average gain of approximately $0.71$ dB over the strongest baselines. It also obtains the best SSIM and LPIPS in most settings. Qualitative improvements are reported for thin structures, signs, railings, chair geometry, storefront details, and regions where incorrect depth causes cross-view inconsistency.

Ablations support the separation of roles between V2G and G2V. Inserting V2G into GenFusion, ViewCrafter, and GSFixer produces relative geometry gains ranging from $10.70\%$ to $47.77\%$, depending on baseline and view count, with an additional reported cost of approximately $21.71$–$26.51$ ms per step. For the complete model, geometry gains are reported as $48.01\%$, $51.49\%$, and $50.78\%$ for 3, 6, and 9 views, respectively.

On Mip-NeRF 360 with six views, removing V2G reduces PSNR from $17.70$ to $17.35$, SSIM from $0.446$ to $0.428$, and produces LPIPS degradation from $0.468$ to $0.473`; geometry metrics also degrade substantially. Removing FlowCond reduces PSNR from $17.70$ to approximately $17.50$ dB, indicating the contribution of geometry-derived motion guidance to video restoration. Removing DINOv2 reduces texture and perceptual quality, with SSIM decreasing by roughly $0.007$ and LPIPS increasing by approximately $0.009$ relative to the corresponding ablation. The selected V2G weight is $\lambda_{\mathrm{V2G}}=0.014$; larger values may improve some geometry metrics while harming RGB quality and SSIM when teacher flow is inaccurate.

Several assumptions constrain interpretation. The reprojection formulation assumes a static scene and known camera poses; dynamic or nonrigid objects can be interpreted as geometric inconsistency. Camera-pose errors can corrupt both geometry-induced flow and trajectory generation. The method depends on large pretrained models and GPU-intensive latent video processing, including 50-step DDIM inference for 49-frame clips. Restored videos may contain hallucinated or temporally inconsistent content, and confidence filtering cannot eliminate all erroneous supervision. Because the co-refinement process is an alternating feedback system rather than a jointly convex optimization, no formal convergence guarantee is provided.

Bi-FlowGS is therefore best characterized as a correspondence-level coupling between generative view completion and differentiable Gaussian geometry. Its principal methodological distinction is that generated views are not used solely as RGB pseudo-targets: their motion structure is distilled into geometric supervision, while the evolving Gaussian scene simultaneously conditions subsequent video restoration. This design targets Geometry Cheating directly by imposing cross-view correspondence constraints on Gaussian depth and spatial arrangement, particularly in sparse-view, wide-baseline, and unbounded 360° reconstruction settings.

Source: https://www.emergentmind.com/topics/bi-flowgs