Papers
Topics
Authors
Recent
Search
2000 character limit reached

PETAR-4B: Dual-Context MLLM and Collider Physics

Updated 3 July 2026
  • PETAR-4B is a dual-use concept, defining both a 4B-parameter auto-thinking multimodal LLM and a 4b quark final state in collider experiments.
  • The model employs bi-mode annealing and policy optimization to dynamically switch between reasoning and direct-answer modes, reducing computational cost.
  • In collider physics, PETAR-4B denotes the 4b final state used in resonant di-Higgs analyses, enhancing signal discrimination with advanced techniques.

PETAR-4B appears in two distinct research contexts: as a tag for general-purpose auto-thinking multimodal LLMs (“R-4B” and associated PETAR-4B nomenclature) and as a label for the $4b$ (“four bottom”) quark final state in resonant di-Higgs collider analyses. The following provides a comprehensive technical overview of both usages, drawing on (Jiang et al., 28 Aug 2025) for the R-4B/PETAR-4B MLLM and (Li et al., 2019) for $4b$-channel collider physics.

1. General-Purpose Auto-Thinking in Multimodal LLMs: PETAR-4B

PETAR-4B is fundamentally associated with the R-4B model, a 4B parameter multimodal LLM (MLLM) developed for adaptive, task-sensitive reasoning (“auto-thinking”). In this paradigm, the model learns—through data-driven and reinforcement learning procedures—to autonomously allocate reasoning effort, switching between “thinking mode” (explicit multistep deliberation) and “direct-answer mode” per input complexity, without external control signals or user specification (Jiang et al., 28 Aug 2025).

The design addresses a critical inefficiency of prior MLLMs: always generating step-by-step reasoning traces, regardless of task difficulty, increases latency and compute cost and can introduce unnecessary verbosity for trivial inputs. PETAR-4B’s key innovation is a two-stage training pipeline:

  1. Bi-mode Annealing: Instills both reasoning and direct-answering capacities by mixing “reasoning-intensive” and “non-reasoning” samples—marked with unified > tags—during fine-tuning. > > 2. Bi-mode Policy Optimization (BPO): An RL phase in which the policy is forced to generate both modes per query and is optimized to choose the mode conferring highest utility (measured via a simple, rule-based reward, especially from the math domain), refining the model’s auto-thinking judgment. > > After these stages, R-4B (PETAR-4B) achieves robust state-of-the-art or competitive performance across 25 vision-language benchmarks, including math, science, VQA, document understanding, and OCR tasks, outperforming Qwen2.5-VL-7B and often matching or exceeding the much larger (16B) Kimi-VL-A3B-Thinking-2506 on reasoning benchmarks with lower compute cost. Notably, token usage in auto-thinking mode closely tracks that of non-thinking mode on simple tasks, but scales up for hard examples, verifying successful policy adaptivity. > > ## 2. Adaptive Mode Selection: Decision Process and Policy Learning > > PETAR-4B (as R-4B) employs a learned policy, not a rule-based controller, for mode selection. During evaluation, control tokens (<think>/blank or <think>…</think>) contextualize the prompt. The model operates in three evaluation modes: > > - N-T (Non-thinking): Always uses a direct-answer token context. > > - T (Thinking): Always includes reasoning by prompting with <think>. > > - A-T (Auto-thinking): Allows the model to decide which trajectory to pursue. > > The learned policy is optimized via BPO, which prevents “mode drift” or “thinking atrophy” (loss of reasoning preference) by always requiring explicit comparison between both modalities per query and then updating the policy with a clipped RL objective under KL regularization constraints—a necessity due to the asymmetry and potential bias remaining after bi-mode annealing. > > ## 3. Dataset Construction and Bi-Mode Annealing > > For bi-mode annealing, the PETAR-4B approach curates a 16.37M-sample dataset partitioned into "reasoning-intensive" (5.50M) and "non-reasoning" (10.87M) subsets. The labeling employs two heuristics: > > - Difficulty-based heuristic: Prompting Qwen2.5-32B-VL to assess subjectivity/open-ended complexity. > > - Performance-based hard mining: Repeated demonstration generation (N=8); if no run is correct, tag as reasoning-intensive; otherwise, as non-reasoning. > > All samples, regardless of tag, use a structurally aligned format—either <think>reasoning</think>answer or <think>answer. This consistency mitigates format-induced distributional shift and supports mode-compositional generalization during both fine-tuning and BPO.

Ablations across data strategies demonstrate that mixed-mode annealing ("Mixed-R") yields the highest benchmark average (69.5%), outperforming single-mode and two-stage-curriculum alternatives.

4. Reinforcement Learning with Bi-mode Policy Optimization

During the BPO stage, the policy is trained to output both thinking and non-thinking responses for each input, with group sizes matched by design. The objective is: JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right) where Rk(θ)R_k(\theta) is the policy ratio for the kkth response, AkA_k is the advantage, ϵ\epsilon controls clipping, and β\beta the KL penalty. This objective—adapted from Group Relative Policy Optimization (GRPO)—ensures balanced exploration and prevents preference collapse towards a single mode.

The reward is a simple, rule-based metric aligned primarily with correct math domain completion, eschewing elaborate manual reward shaping or human complexity annotation. A plausible implication is that the policy’s generalization capacity across non-math domains is anchored in the broad-domain mix of the bi-mode annealing dataset.

5. Empirical Evaluation and Efficiency Characteristics

PETAR-4B (R-4B) achieves strong performance across a suite including MMMUval, MMStar, MMVet, CharXiv, LogicVista, DynaMath, MathVerse, OlympiadBench, and WeMath, as well as visual and OCR benchmarks. Specifically, R-4B outperforms Qwen2.5-VL-7B on most tasks and is competitive with Kimi-VL-A3B-Thinking-2506, particularly in domains such as chart reasoning, logic, and math, despite a much smaller parameter count.

A token analysis demonstrates that the auto-thinking policy sharply reduces output length (and thus inference cost) on tasks like OCRBench (66 tokens vs. 394 in thinking mode), while incurring longer traces where complex reasoning is required. This suggests substantial gains in throughput and latency for real-world applications deploying PETAR-4B-class models.

6. PETAR-4B as the $4b$ Final State in Collider Physics

In high-energy physics, PETAR-4B refers to the $4b$ ("four bottom quark") channel central to studying resonant di-Higgs production in scalar singlet extensions of the Standard Model (Li et al., 2019). Specifically, the $4b$0 final state arises from the process: $4b$1 where $4b$2 is a heavy singlet-like scalar, and $4b$3 the $4b$4 GeV Higgs. This process probes parameter space regions leading to a strong first-order electroweak phase transition (SFOEWPT)—a requirement for electroweak baryogenesis.

The study employs an ATLAS-inspired resolved $4b$5 selection strategy, enhanced by a boosted decision tree (BDT) for improved signal-background discrimination at the HL-LHC. The selection strategy involves identifying two dijet pairs with invariant masses consistent with $4b$6, using $4b$7 and mass-window criteria optimized per benchmark, and then analyzing kinematic variables (including $4b$8, $4b$9, JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)0, JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)1) as BDT inputs.

7. JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)2 Channel Results and Phenomenological Relevance

The JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)3 analysis at 14 TeV HL-LHC (3 abJBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)4) shows that, for scenarios with maximal singlet-Higgs mixing that produce a SFOEWPT, a discovery reach up to JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)5 GeV is achievable; exclusion up to JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)6 GeV is possible if no resonance is seen. For higher JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)7, the JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)8 channel surpasses JBPO(θ)=EqP(Q)12gk=12gmin(Rk(θ)Ak, clip(Rk(θ),1ϵ,1+ϵ)Ak)βDKL(πθ(q)πref(q))J_{\mathrm{BPO}}(\theta) = \mathbb{E}_{q\sim P(Q)}\frac{1}{2g}\sum_{k=1}^{2g} \min\Big(R_k(\theta)A_k,\ \mathrm{clip}(R_k(\theta),1-\epsilon,1+\epsilon)A_k\Big) -\beta\,D_{\mathrm{KL}}\left(\pi_\theta(\cdot|q)\,\|\,\pi_{\mathrm{ref}}(\cdot|q)\right)9 and Rk(θ)R_k(\theta)0 in sensitivity, though it is slightly outperformed by Rk(θ)R_k(\theta)1 under optimistic assumptions. Notably, the Rk(θ)R_k(\theta)2 channel cannot fully exclude all SFOEWPT-allowed parameter space, thereby motivating a multi-channel resonant di-Higgs search program.

Across both machine learning and collider physics domains, PETAR-4B denotes rigorously defined and experimentally validated methods pivotal to their respective fields. In MLLMs, it encapsulates the principle of policy-driven, resource-efficient generalized reasoning; in electroweak baryogenesis studies, it identifies the critical experimental channel for probing singlet-extended phase transition dynamics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PETAR-4B.