Bi-VLM: Diverse Bi-directional Vision–Language Models
- Bi-VLM is a term that covers diverse bi-directional vision–language models aimed at restoring missing modalities and aligning multimodal representations.
- Key techniques include diffusion-based feature restoration, asymmetric twin-brain architectures for embodied control, and bi-directional prompt learning within frozen backbones.
- Additional formulations extend Bi-VLM to ultra-low precision quantization, bimanual manipulation, and graph matching for specialized applications such as medical segmentation.
to=arxiv_search.search qq彩票 大发快三是不是 下载彩神争霸াদি 平台直属 _影音先锋 code 手机上天天中彩票json {"4query4 OR ti:\4"Bi-VLM\" OR abs:\4"Bi-VLM\"","max_results":4all:\4query4 to=arxiv_search.search _一本道 微信里的天天中彩票json {"4query4 OR id:(&&&4all:\4&&&) OR id:(&&&4 OR ti:\4&&&) OR id:(&&&4 OR abs:\4&&&) OR id:(Zhou et al., 26 Sep 2025) OR id:(Wenting et al., 2023)","max_results":4all:\4query4,"sort_by":"submittedDate","sort_order":"descending"} Bi-VLM is a label used in recent arXiv literature for several distinct vision–language research programs rather than a single standardized model family. In one line of work, it denotes bi-directional vision–language modeling through missing-modality feature restoration and alignment in a frozen CLIP latent space; in another, it denotes a two-brain VLM architecture for embodied control; in another, bi-directional modality interaction prompt learning inside frozen CLIP; and in another, an ultra-low-precision post-training quantization pipeline for VLMs. Related formulations extend comparable ideas to vision-language-anchored bimanual manipulation and bi-level vision-language graph matching for text-guided medical segmentation (&&&4query4&&&, &&&4all:\4&&&, &&&4 OR ti:\4&&&, &&&4 OR abs:\4&&&, Zhou et al., 26 Sep 2025, Wenting et al., 2023).
4all:\4. Terminological scope and recurring design motifs
A common source of confusion is that the label does not denote one canonical architecture in the cited literature. Instead, recent uses cluster around several technical motifs: bidirectional restoration between image and text feature spaces, asymmetric coordination between two VLM streams, prompt-level bi-directional interaction inside frozen backbones, and distribution-aware quantization for efficient deployment.
| Usage | Core mechanism | Setting |
|---|---|---|
| Missing-modality Bi-VLM | Conditional diffusion, DMG, CMML | Foundation VLM robustness |
| TwinBrain Bi-VLM | Frozen Left Brain, trainable Right Brain, AsyMoT | Embodied VLA control |
| BMIP-style bi-directional interaction | Layered prompt replacement with attention-derived weights | Few-shot CLIP adaptation |
| Quantization Bi-VLM | Gaussian-quantile saliency-aware hybrid PTQ | Ultra-low-bit VLM deployment |
Across these uses, several structural themes recur. One is the preservation of a frozen semantic backbone while adding a lightweight or specialized mechanism around it: frozen CLIP encoders in missing-modality restoration and BMIP, a frozen generalist VLM in TwinBrainVLA, and post-training rather than retraining-oriented compression in ultra-low-bit Bi-VLM. Another is explicit asymmetry: the available modality conditions restoration of the missing one, the Left Brain semantically anchors the Right Brain without receiving gradients, and saliency-aware quantization treats outlier and inlier weights differently. This suggests that, in current usage, “Bi-VLM” often marks an attempt to preserve a stable semantic prior while introducing a second pathway, direction, or precision regime for adaptation.
4 OR ti:\4. Bi-directional feature restoration for missing-modality inference
In "Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration," Bi-VLM is realized as a mid-stage training module inserted between frozen CLIP dual encoders and downstream heads. The visual encoder PRESERVED_PLACEHOLDER_4query4^ and text encoder PRESERVED_PLACEHOLDER_4all:\4^ remain frozen, while an enhanced diffusion transformer operates in the VLM feature space to restore the missing feature PRESERVED_PLACEHOLDER_4 OR ti:\4^ from the available conditional feature PRESERVED_PLACEHOLDER_4 OR abs:\4. The forward noising process is
and the reverse conditional process is
The corresponding clean estimate is
This formulation supports both text-to-vision and vision-to-text restoration, with denoising trajectories anchored on the frozen CLIP manifold.
Two mechanisms define the paper’s bi-directional character. Dynamic Modality Gating (DMG) adaptively fuses conditional and restored content, conceptually written as
and implemented inside DiT blocks through learnable attention-pooled gate vectors and layer-wise gates and . Cross-Modal Mutual Learning (CMML) then forces restored features to be valid conditions in the opposite direction. The total loss is
PRESERVED_PLACEHOLDER_4all:\4query4^
During inference, if text is missing, the model restores PRESERVED_PLACEHOLDER_4all:\4all:\4^ from PRESERVED_PLACEHOLDER_4all:\4 OR ti:\4; if image is missing, it restores PRESERVED_PLACEHOLDER_4all:\4 OR abs:\4^ from PRESERVED_PLACEHOLDER_4all:\44; the available and restored features are then concatenated and fed to a lightweight decoder.
The reported evaluation frames this Bi-VLM as a robustness mechanism for missing-modality deployment rather than a replacement for the underlying foundation model. Benchmarks are MM-IMDb (F4all:\4-Macro), N4 OR ti:\44News (Accuracy), MMHS4all:\4all:\4K (Accuracy), and Food4all:\4query4all:\4^ (Accuracy). At 74query4% missing, the paper reports MM-IMDb (missing image) F4all:\4-M +7.4query47% over the best baseline, MMHS4all:\4all:\4K (missing image) ACC +8.4all:\46%, and N4 OR ti:\44News ACC +4 OR abs:\4.88% for image missing, with Food4all:\4query4all:\4^ showing small but consistent gains. Performance degrades only ~4 OR abs:\4% from 4all:\4query4%→94query4 missing in MM-IMDb, whereas baselines drop >4all:\4query4%. Ablations further report that replacing DMG with AdaLN or simple concat degrades performance, removing CMML weakens robustness and increases semantic drift, performance scales positively from Food4all:\4query4all:\4^ (68K) to COCO (445K) to CC4 OR abs:\4M (4 OR abs:\4M), gains rise from 4all:\46→4 OR ti:\4query4^ DiT layers with diminishing returns beyond 4 OR ti:\4query4, and accuracy plateaus at ~54query4^ DDIM steps, with latency ~4 OR ti:\4 OR ti:\4.94 ms under complete modalities and ~678–74query47 ms under missing modalities for 54query4^ DDIM steps (&&&4query4&&&).
4 OR abs:\4. Dual-VLM asymmetry for embodied action
In "TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers," Bi-VLM denotes a two coordinated VLM design. The architecture pairs a frozen “Left Brain” VLM that retains robust general visual reasoning with a fully trainable “Right Brain” VLM specialized for embodied perception and proprioceptive grounding. Both streams are isomorphic stacks initialized from the same checkpoint, using Qwen4 OR ti:\4.5-VL-4 OR abs:\4B-Instruct or Qwen4 OR abs:\4-VL-4B-Instruct, but they receive different token sequences: the Left Brain processes visual and text tokens only, while the Right Brain additionally ingests state tokens PRESERVED_PLACEHOLDER_4all:\45 derived from proprioceptive state PRESERVED_PLACEHOLDER_4all:\46.
Coordination is implemented through the Asymmetric Mixture-of-Transformers (AsyMoT). At each layer PRESERVED_PLACEHOLDER_4all:\47, the Right Brain forms joint keys and values by concatenating its own projections with stop-gradient copies from the Left Brain,
PRESERVED_PLACEHOLDER_4all:\48
and computes
PRESERVED_PLACEHOLDER_4all:\49
The asymmetry is strict: PRESERVED_PLACEHOLDER_4 OR ti:\4query4^ is frozen, PRESERVED_PLACEHOLDER_4 OR ti:\4all:\4, and no explicit scalar gating network is introduced. This architecture structurally isolates the semantic generalist from action gradients, while allowing the specialist stream to 4query4^ stable semantic representations.
Action generation is delegated to a Diffusion Transformer policy head trained with flow matching. The policy models
PRESERVED_PLACEHOLDER_4 OR ti:\4 OR ti:\4^
with straight-path interpolation PRESERVED_PLACEHOLDER_4 OR ti:\4 OR abs:\4^ and loss
PRESERVED_PLACEHOLDER_4 OR ti:\44^
Training freezes the Left Brain and optimizes only the Right Brain VLM, the state encoder, and the DiT policy head, using 44query4k steps, 4all:\46× NVIDIA H4all:\4query4query4^ GPUs, batch size 4all:\46 per device, AdamW, lr = 4all:\4e−5, cosine annealing, gradient clipping (norm 4all:\4.4query4 and DeepSpeed ZeRO-4 OR ti:\4.
Empirically, TwinBrainVLA is evaluated on SimplerEnv and RoboCasa. On SimplerEnv, TwinBrainVLA + Qwen4 OR ti:\4.5-VL-4 OR abs:\4B-Instruct reaches 58.4% average success, while TwinBrainVLA + Qwen4 OR abs:\4-VL-4B-Instruct reaches 64 OR ti:\4.4query4%, surpassing Isaac-GR4query4query4T-N4all:\4 (57.4all:\4%) by +4.9%. On RoboCasa GR4all:\4^ Tabletop, the corresponding averages are 54 OR abs:\4.5% and 54.6%, with the Qwen4 OR abs:\4-VL-4B-Instruct version outperforming Isaac-GR4query4query4T-N4all:\4 (47.6%) by +7.4query4%, QwenGR4query4query4T (47.8%) by +6.8%, and QwenPI (44 OR abs:\4.9%) by +4all:\4query4.7%. The paper argues semantic preservation by architectural design rather than by post-training VQA or captioning benchmarks; explicit external VLM evaluations are not reported. Limitations include the same-architecture constraint for Left/Right pairing, simulation-first evaluation, and the higher inference cost of running both VLMs plus DiT ODE steps (&&&4all:\4&&&).
4. Bi-directional modality interaction in prompt learning
"BMIP: Bi-directional Modality Interaction Prompt Learning for VLM" uses the idea of Bi-VLM at the prompt-learning level inside a frozen CLIP backbone. The model keeps the image encoder PRESERVED_PLACEHOLDER_4 OR ti:\45 and text encoder PRESERVED_PLACEHOLDER_4 OR ti:\46 frozen and learns layered prompts, projection heads, and attention-derived weighting layers. Deep language prompts PRESERVED_PLACEHOLDER_4 OR ti:\47 are inserted into the first PRESERVED_PLACEHOLDER_4 OR ti:\48 layers of the text encoder, and deep vision prompts PRESERVED_PLACEHOLDER_4 OR ti:\49 are inserted into the first PRESERVED_PLACEHOLDER_4 OR abs:\4query4^ layers of the image encoder. BMIP does not add new cross-attention modules; instead, it reads the native self-attention outputs of each branch and uses them to derive prompt substitution weights.
The central interaction block defines
PRESERVED_PLACEHOLDER_4 OR abs:\4all:\4^
and then performs bi-directional prompt replacement:
PRESERVED_PLACEHOLDER_4 OR abs:\4 OR ti:\4^
PRESERVED_PLACEHOLDER_4 OR abs:\4 OR abs:\4^
Thus the vision prompt absorbs transformed language information and the language prompt absorbs transformed vision information, with the balance learned from attention outputs. The paper positions this against single-modal prompt learning and uni-directional prompt transfer, arguing that simple aggregation or one-way transfer neglects the alignment effects resulting from two-way interaction.
The training objective remains the standard VLM classification objective under few-shot finetuning with frozen encoders,
PRESERVED_PLACEHOLDER_4 OR abs:\44^
with gradients flowing through the prompts, projection heads, and attention-weight generators. BMIP also proposes open-world generalization, defined as simultaneous evaluation on an unknown distribution composed of both base and novel classes, alongside cross-dataset transfer and domain generalization.
Quantitatively, for open-world generalization on ViT-B/4all:\46, BMIP reports average HM/Acc of 79.4query44^ / 74 OR ti:\4.4all:\47, compared with MaPLe 78.4 OR ti:\4 OR ti:\4^ / 74all:\4.76, CoCoOp 74.74 OR ti:\4^ / 67.67, CoOp 74 OR ti:\4.4all:\44^ / 65.57, and CLIP 74query4.84 / 64 OR abs:\4.94 OR ti:\4. Representative dataset-level results include EuroSAT HM 86.4all:\4query4^ ± 4all:\4.58 for BMIP versus 84all:\4.44 OR abs:\4^ ± 4query4.54 OR abs:\4^ for MaPLe, Flowers4all:\4query4 OR ti:\4^ HM 84 OR abs:\4.86 ± 4all:\4.74query4^ versus 84 OR ti:\4.78 ± 4query4.69, FGVC-Aircraft HM 4 OR abs:\47.4 OR ti:\45 ± 4query4.94 OR abs:\4^ versus 4 OR abs:\45.4 OR ti:\49 ± 4query4.58, and SUN4 OR abs:\497 HM 79.4query4 OR ti:\4^ ± 4query4.4 OR ti:\44^ versus 79.58 ± 4query4.4all:\4 OR abs:\4. In cross-dataset transfer, BMIP reports average accuracy 66.86 versus MaPLe 66.4 OR abs:\4query4, and in domain generalization an overall average 64 OR ti:\4.54query4^ versus 64 OR ti:\4.4 OR abs:\46 and OOD average 64query4.44query4^ versus 64query4.4 OR ti:\48. Aggregation ablations report BMIP at 79.4query44^ / 74 OR ti:\4.4all:\47, above IVLP, CoCoOp, MaPLe, PRESERVED_PLACEHOLDER_4 OR abs:\45, Addition, Attention, and Joint, supporting the paper’s claim that attention-guided learnable substitution is more effective than simple aggregation. The method is also composable: MaPLe+BMIP improves HM/Acc from 78.4 OR ti:\4 OR ti:\4^ / 74all:\4.76 to 79.4query4 OR abs:\4^ / 74 OR ti:\4.54, PromptSRC+BMIP from 79.67 / 74 OR abs:\4.44 OR abs:\4^ to 84query4.4query4 OR abs:\4^ / 74 OR abs:\4.97, and CoPrompt+BMIP from 78.99 / 74all:\4.48 to 79.54 / 74 OR ti:\4.4 OR abs:\45 (&&&4 OR ti:\4&&&).
5. Bi-VLM as ultra-low-precision post-training quantization
In "Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-LLMs," the term denotes a post-training quantization method for VLM deployment under hardware constraints. The method assumes approximate Gaussian layerwise weight statistics and partitions each weight matrix into an outlier subset and multiple inlier subsets using Gaussian quantiles. For a layer PRESERVED_PLACEHOLDER_4 OR abs:\46, the standardized magnitude variable is
PRESERVED_PLACEHOLDER_4 OR abs:\47
with thresholds derived from PRESERVED_PLACEHOLDER_4 OR abs:\48 and salient percentile PRESERVED_PLACEHOLDER_4 OR abs:\49. The saliency metric is
4query4^
which labels weights as salient when they exceed the learned tail threshold.
Quantization is hybrid and saliency-aware. Salient entries receive 4 OR ti:\4-bit non-uniform quantization with row-wise scalers,
4all:\4^
whereas each inlier subset is strictly binarized,
4 OR ti:\4^
The layerwise reconstruction objective is
4 OR abs:\4^
and the optimal scalar for an inlier subset is
4
Layerwise saliency search uses Brent’s method under bounded salient proportions, with empirical caps of ≈5% in vision and ≈4all:\4% in the LM. After quantization, the paper further prunes image tokens based on attention mass and reports 94query4%–99% image token redundancy in quantized models.
The method is evaluated on Llama 4 OR abs:\4.4 OR ti:\4-Vision instruction 4all:\4all:\4B, LLaVA-One-Vision 7B, and Qwen4 OR ti:\4.5-VL-7B-Instruct over MME, MMMU, ScienceQA-IMG, and VizWiz-VQA, using PTQ with 64 calibration samples. The abstract reports that, for the LLM part of the VLM, Bi-VLM outperforms the SOTA by 4 OR abs:\4%–47% on the visual question answering task in terms of four different benchmarks and three different models, and for the overall VLM it outperforms the SOTA by 4%–45%. The detailed results state, for example, that on Llama 4 OR abs:\4.4 OR ti:\4-Vision 4all:\4all:\4B the language part improves over SOTA by 4%–47% and the whole VLM by 8%–45%; on LLaVA-One-Vision 7B the gains are +4 OR abs:\4%–4 OR ti:\4query4% for the language part and +4%–4all:\49% for the whole model; and on Qwen4 OR ti:\4.5-VL-7B-Instruct they are +5%–4all:\4query4% and +4%–4all:\4 OR ti:\4%, respectively. Storage analysis gives
5
with average storage ≈ 4all:\4.4query4all:\44^ bits/weight for 6 and 7 in the LM, implying 8 for 4all:\46-bit FP weights and 9 for 4 OR abs:\4 OR ti:\4-bit FP weights. The paper identifies the vision encoder as highly sensitive, the adaptor as low sensitivity, and the LM as considerably sensitive in the 4all:\4–4all:\4 bit regime, and notes that text-token pruning harms accuracy much more than image-token pruning (&&&4 OR abs:\4&&&).
6. Adjacent formulations: bimanual manipulation and bi-level graph matching
Related work broadens the technical neighborhood around the label. "VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Robotic Bimanual Manipulation" defines a framework for bimanual manipulation in which a single precisely labeled human demonstration is decomposed into invariant anchors and adjustable components, with adaptation driven by Florence-4 OR ti:\4^ and SAM4 OR ti:\4^ segmentation, stereo back-projection, and geometric feasibility constraints. The demonstration trajectory 4query4^ is partitioned into invariant primitives 4all:\4^ and adjustable components 4 OR ti:\4^ using the bind indicator and object geometry tolerance, while scene adaptation is computed from
4 OR abs:\4^
Execution uses progressive IK refinement,
4
and dynamic collision compensation
5
On six primary tasks with new placements and same objects, VLBiMan reports average 85.4query4%, versus ReKep+ 65.4query4%, ReKep 44 OR abs:\4.4 OR abs:\4%, Robot-ABC 4 OR abs:\46.7%, MAGIC 45.8%, and Mechanisms 4 OR ti:\48.4 OR abs:\4%; with novel instances the figure is 78.4 OR abs:\4%. Under dynamic interference, the averages are 74query4.4query4 for same objects and 59.4 OR ti:\4% for novel instances. Long-horizon tasks reach 54 OR ti:\4.5% and 44all:\4.4 OR abs:\4% without interference, and cross-embodiment transfer on a humanoid dual-arm robot reaches 84 OR abs:\4.8% and 76.4 OR abs:\4% without interference. Ablations report that removing IK refinement drops the six-task novel-instance-with-interference result from 59.4 OR ti:\4% to 4 OR ti:\49.4 OR ti:\4%, removing collision avoidance to 4 OR abs:\44.4 OR ti:\4%, replacing the paper’s grasp adaptation with AnyGrasp to 4 OR abs:\4all:\4.7%, and downgrading VLMs to SAM+DINOv4 OR ti:\4^ to 4 OR abs:\45.8% (Zhou et al., 26 Sep 2025).
A second adjacent line is "Bi-VLGM : Bi-Level Class-Severity-Aware Vision-Language Graph Matching for Text Guided Medical Image Segmentation," which is not titled Bi-VLM but is explicitly a bi-level vision-language matching framework. It augments HRNetV4 OR ti:\4^ with a word-level VLGM module for local-class alignment and a sentence-level VLGM module for global-severity alignment. Soft correspondence matrices are produced by AIS over GCN-embedded visual and textual graphs, with losses
6
7
encoder objective
8
and segmentation objective
9
Severity-aware prompts are constructed from lesion area ratios using thresholds 4query4^ and 4all:\4. On IDRiD, Bi-VLGM reports mAUPR 68.4all:\4 OR ti:\4%, mF 66.74all:\4%, and mIoU 54all:\4.4 OR abs:\47%; on DDR it reports mIoU 4 OR abs:\4all:\4.78% and mF 47.4 OR ti:\49%. Ablations show mIoU rising from 47.74 for HRNetV4 OR ti:\4^ to 54query4.4query4query4^ with word-level only, 54query4.54query4^ with sentence-level only, and 54all:\4.4 OR abs:\47 with both. A comparison to contrastive loss reports VLGM mIoU 54all:\4.4 OR abs:\47 versus 49.46. These neighboring formulations show that the broader “bi-” vision-language design space also includes explicit decomposition into two alignment levels, two arms, or two control regimes, not only two modalities or two model branches (Wenting et al., 2023).