Papers
Topics
Authors
Recent
Search
2000 character limit reached

RemEdit Diffusion Framework

Updated 1 February 2026
  • The paper introduces a novel diffusion-based framework that leverages Riemannian manifold navigation and dual-SLERP blending to achieve high-fidelity and controllable image edits.
  • It employs goal-aware prompt enrichment via a vision-language model and task-specific attention pruning to enhance semantic consistency and accelerate inference.
  • Empirical benchmarks on datasets like CelebA-HQ and LSUN-Church demonstrate superior accuracy (S_dir up to 0.1982) and rapid runtimes, affirming its practical effectiveness.

RemEdit is a diffusion-based image editing framework designed to reconcile the trade-off between semantic fidelity and inference speed in controllable generative AI. It achieves this through a synergistic integration of Riemannian geometry for latent space navigation, dual spherical interpolation (SLERP), grounded prompt enrichment, and task-specific attention pruning. RemEdit demonstrates state-of-the-art editing accuracy and real-time performance benchmarks across various datasets, substantiating its practicality and robustness for high-fidelity image manipulation (Adhikarla et al., 25 Jan 2026).

1. Riemannian Manifold Navigation in Latent Space

RemEdit models the U-Net bottleneck feature space ("h-space") as a Riemannian manifold (M,g)(\mathcal M, g) of dimension N=C×H×WN = C \times H \times W. The metric tensor gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle induces inner products on the tangent space ThMT_h\mathcal M, enabling the computation of geodesics—paths following the data distribution instead of mere straight Euclidean offsets. The framework learns a non-trivial affine connection ∇\nabla parameterized by Christoffel symbols Γijk(h)\Gamma^k_{ij}(h), which underlie geodesic calculations in feature space.

A lightweight "Mamba" network fθf_\theta takes the concatenated feature vector y0=concat(h,  timestep)y_0 = \mathrm{concat}(h,\;\mathrm{timestep}) and predicts the Christoffel symbols. In parallel, a tangent-vector predictor vϕv_\phi outputs an initial velocity v0=vϕ(y0)v_0 = v_\phi(y_0), constrained to the unit ball via tanh-based retraction. Geodesic endpoints are obtained by integrating the following ODE system over N=C×H×WN = C \times H \times W0: N=C×H×WN = C \times H \times W1 yielding N=C×H×WN = C \times H \times W2 and the geodesic edit N=C×H×WN = C \times H \times W3. The model is trained to match example edits N=C×H×WN = C \times H \times W4 using the objective

N=C×H×WN = C \times H \times W5

This Riemannian approach preserves the semantics of edits by respecting the intrinsic structure of the latent distribution.

2. Dual-SLERP Blending for Edit and Identity Control

RemEdit employs a dual-SLERP approach, hierarchically interpolating between the original and edited states in both feature and noise spaces. Given two unit vectors N=C×H×WN = C \times H \times W6 and N=C×H×WN = C \times H \times W7 with angle N=C×H×WN = C \times H \times W8, the SLERP is defined as

N=C×H×WN = C \times H \times W9

The inner SLERP applies this to gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle0 (original latent) and gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle1 (geodesic endpoint), facilitating continuous modulation of edit strength: gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle2 The outer SLERP operates in noise space after U-Net forward passes. It separates semantic change (identity-orthogonal component) from fidelity, fusing predictions as: gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle3 This dual mechanism enables explicit, fine-grained control over semantic transformation and identity retention.

3. Goal-Aware Semantic Prompt Enrichment

Textual edit prompts are typically under-specified for targeted manipulations (e.g., "face → face with makeup"). RemEdit enriches prompts by extracting a fine-grained caption gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle4 from the input image gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle5 using a pretrained Vision-LLM (Qwen2-VL): gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle6 Semantic edit direction gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle7 is then computed in text-embedding space (e.g., via CLIP): gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle8 This approach grounds edits in actual image content, enhancing consistency and specificity without the need for additional training—prompt enrichment is a single forward pass through the VLM.

4. Task-Specific Attention Pruning

RemEdit introduces a novel attention pruning mechanism to accelerate inference while maintaining semantic fidelity. Each self-attention block's feature map gij(h)=⟨∂i,∂j⟩g_{ij}(h) = \langle \partial_i, \partial_j \rangle9 is reshaped into tokens ThMT_h\mathcal M0, with ThMT_h\mathcal M1. An MLP pruner ThMT_h\mathcal M2 processes ThMT_h\mathcal M3 to produce soft importance scores ThMT_h\mathcal M4. During inference, a pruning rate ThMT_h\mathcal M5 determines the top ThMT_h\mathcal M6 tokens to retain; attention is computed only over these, yielding: ThMT_h\mathcal M7 The pruner is trained to optimize

ThMT_h\mathcal M8

enabling effective acceleration without degrading semantic edit quality.

5. Empirical Performance and Benchmarks

RemEdit demonstrates state-of-the-art results on several benchmark datasets. Quantitative metrics on CelebA-HQ (ThMT_h\mathcal M9) include directionality score (∇\nabla0), segmentation consistency, Fréchet Inception Distance (FID), and runtime:

Method ∇\nabla1 Seg. Consistency ∇\nabla2 FID∇\nabla3 Time (s) ∇\nabla4
Asyrp (h-space) 0.1900 87.9% 24.3 28.9
LEdits++ 0.1820 89.7% 22.5 20.1
RemEdit (full) 0.1982 92.4% 19.8 2.8

(∇\nabla5 FID evaluated on 250 edited samples vs. held-out real images.)

On LSUN-Church, RemEdit similarly outperforms Diffusion-CLIP and BoundaryDiffusion in both ∇\nabla6 and segmentation consistency. Ablation experiments (CelebA-HQ, "smiling" edit) further show that geodesic navigation and dual-SLERP deliver superior semantic accuracy, and that up to 50% attention pruning (∇\nabla7) preserves most fidelity while reducing runtime to approximately 2.31 s.

6. End-to-End Algorithmic Structure

RemEdit’s computational workflow integrates the above components into a coherent editing pipeline:

Γijk(h)\Gamma^k_{ij}(h)0

This architecture demonstrates the synthesis of geometric, linguistic, and computational efficiency innovations, each contributing distinctly to controllable, high-speed image editing.

7. Significance and Implications

RemEdit's unified approach—learned Riemannian manifold navigation, dual-stage SLERP blending, prompt enrichment via VLM, and semantic attention pruning—addresses core limitations of prior diffusion editors. It sets new standards in edit fidelity (∇\nabla8 up to 0.198 on CelebA-HQ) and runtime (∇\nabla9 s at 50% pruning), evidencing not only superior semantic accuracy but also practical deployment feasibility. A plausible implication is that future generative editing frameworks may increasingly rely on data-driven geometric structures and task-aware resource allocation to reconcile interpretability, controllability, and efficiency (Adhikarla et al., 25 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RemEdit Diffusion-Based Framework.