---
title: 'Bielik 11B v3: LLM & Fusion Reactor Advances'
url: https://www.emergentmind.com/topics/bielik-11b-v3
type: topic
---

# Bielik 11B v3: LLM & Fusion Reactor Advances

Bielik 11B v3 denotes two separate high-impact research developments: (1) a state-of-the-art large language model (LLM) architecture optimized for Polish and multilingual tasks in the Bielik LLM series and (2) a compact, spherical-torus p–¹¹B fusion plasma reactor design concept. Both represent significant advances in their domains, with the former introducing innovations in tokenizer adaptation, embedding initialization, and training pipelines for language modeling, and the latter implementing multi-magnetofluid equilibrium with enhanced suprathermal populations and sophisticated transport physics for aneutronic fusion. Below, both aspects are presented in detail.

## 1. Transformer Language Model: Design and Architecture

Bielik 11B v3 in language modeling is an 11.2-billion-parameter Transformer designed explicitly for Polish and European languages [2604.10799], [2601.11579]. It leverages the Mistral 7B backbone, expanded from 32 to 50 layers using depth up-scaling (DUS), as follows:

- **Parameterization:** 50 Transformer layers, ≈8k hidden dimension, 32 attention heads, and 32k × 4k feed-forward width. Derived directly from Mistral-7B v0.2 layout (GQA, RoPE).
- **Attention and Position Encoding:** Implements Grouped-Query Attention (GQA) with shared key/value projections and Rotary Positional Embeddings (RoPE), supporting context lengths natively up to 32,768 tokens.
- **Scaling Mechanism:** Depth up-scaling is performed by duplicating and splicing layer blocks of the parent network, preserving width while scaling the depth from 32 to 50 layers.
- **Compression Variant:** The Minitron-7B is obtained by structured hybrid pruning (removing entire Transformer layers and reducing MLP width) and logit-based knowledge distillation, enabling ≈33% parameter count reduction and ≈50% faster inference while recovering ≈90% of the teacher’s performance [2603.11881].
- **Deployment and Quantization:** Supports post-training quantization via GPTQ, AWQ, and HQQ, allowing 8-, 5-, and 4-bit deployment. Fits within 22 GB fp16, 11 GB 8-bit, or 5 GB 4-bit weight footprints, with <10% task performance loss at 4-bit [2601.11579].

## 2. Tokenizer and Embedding Optimization for Polish

Bielik 11B v3 addresses fundamental inefficiencies of universal tokenizers by introducing a Polish-optimized tokenizer (APT4) [2604.10799]:

- **Vocabulary and Fertility Metrics:** APT4 vocabulary set to 32,000, replacing the Mistral-based 32,128-token multilingual vocabulary. It is specifically designed to reflect Polish morphosyntactic structure, including morpheme and diacritic resolution, and optimized digit/punctuation handling.
- **Quantitative Improvement:** Reduces tokens/word from ≈3.22 to ≈1.62 and fertility ratio from ≈0.21 to ≈0.209 tokens/character. This nearly halves the token count required for Polish text encoding, enables longer effective context (from ≈10,176 to ≈20,224 Polish words per 32k-token context), and yields up to ≈50% reduction in inference cost for Polish inputs.
- **Embedding Initialization (FOCUS):** Transition to the new tokenizer uses FOCUS (Fast Overlapping Token Combinations Using Sparsemax), which initializes new token embeddings as sparse linear combinations of the original vocabulary’s embeddings, guided by semantic similarity and sparsemax normalization. This approach mitigates catastrophic forgetting by preserving semantic relationships [2604.10799].

## 3. Multi-Stage Pretraining and Alignment Pipeline

Bielik 11B v3 adopts a two-stage continued pretraining curriculum after tokenizer transition:

- **Stage 1: Partial Freezing**  
  Input and output embedding matrices, along with a small number of extremal Transformer layers, are unfrozen for 4B tokens to adapt boundary representations to the new tokenizer. The remainder of the model is frozen.
- **Stage 2: Full Model Adaptation**  
  All parameters are unfrozen for a further 16B tokens. Autoregressive cross-entropy loss is optimized throughout, allowing a full global retuning.
- **Post-Training Alignment:**  
  Alignment to user preferences is achieved by:
  - **Supervised Fine-Tuning (SFT):** 20M Polish and English instruction–response pairs, up to 32,768 token sequences, 3 epochs.
  - **Direct Preference Optimization (DPO-P):** 114k human-labeled preference pairs, using a positive-only DPO loss, explicitly optimizing user-preferred generations.
  - **Reinforcement Learning via GRPO:** 143k domain-specialized examples in mathematics and STEM, optimized using group-relative reward maximization constrained by a KL-divergence term to balance exploration and prior fidelity.

## 4. Performance Benchmarks and Efficiency

Empirical evaluation demonstrates leading performance of Bielik 11B v3 on Polish and multilingual tasks [2604.10799], [2601.11579]. Selected results include:

| Benchmark                                      | Bielik-11B-v3.0-Instruct | Bielik-PL-11B-v3.0-Instruct |
|------------------------------------------------|--------------------------|-----------------------------|
| Open PL LLM (Polish, 5-shot avg)               | 65.93                    | 64.11                       |
| Polish EQ-Bench                                | 71.20                    | 71.15                       |
| CPTUB Complex Polish Text Understanding        | 3.73                     | 3.80                        |
| Medical PL (PES 2018–22, 5-shot %)             | 50.21%                   | 48.42%                      |
| Open LLM Leaderboard (English, core tasks)     | 72.45                    | 71.49                       |
| INCLUDE-base-44 (20 EU langs, avg/PL)          | 64.8 / 69.0              | 53.92 / 64.23               |
| Belebele Multilingual RC (avg/PL)              | 82.98 / 82.11            |                             |
| FLORES Translation (BLEU, avg/PL)              | 19.22 / 18.54            | 17.82 / 17.58               |

With full 32k context, the model supports batch deployment on 24 GB GPUs. For real Polish text, end-to-end inference latency can match or outperform the compressed 7B variant due to reduced input length from lower token fertility.

## 5. Spherical-Torus p–¹¹B Fusion Reactor: Bielik 11B v3 Design Concept

Bielik 11B v3 also refers to a compact, multi-magnetofluid equilibrium concept for p–¹¹B fusion plasmas [2604.04002].

- **Multi-Species Equilibrium:** Six-fluid axisymmetric model (thermal/suprathermal protons, ¹¹B, electrons) with continuity and momentum conservation under common electric and magnetic fields. Differential rotation and pressure anisotropies are fully self-consistent.
- **p–¹¹B Double-Peak Cross-Section:** Exploits two resonances at E₁ ≈ 160–165 keV and E₂ ≈ 650–675 keV, enhancing fusion rates in plasmas with significant suprathermal components.
- **Species Parameters (on axis):**
  - Thermal p: n = 0.43 × 10²⁰ m⁻³, T = 132 keV
  - Suprathermal p: n = 0.0046 × 10²⁰ m⁻³, T = 892 keV
  - Thermal B: n = 0.12 × 10²⁰ m⁻³, T = 132 keV
  - Suprathermal e: n = 0.011 × 10²⁰ m⁻³, T ≈ 0.5–1 MeV
- **Rotation:** NBI-induced toroidal rotation u_φ up to ±2,700 km/s for suprathermal ions, creating large differential shear (Δu_φ > 2 × 10⁶ m/s).
- **Outboard Magnetic Well and Omnigeneity:** The design features an outboard |B| well (minimum 2.75 T, maximum 3.14 T, ΔR ≈ 0.10 m beyond LCFS) that reduces neoclassical diffusion (“orbit squeezing”) and supports high E×B shear rates (ω_E ∼ 2 × 10⁶ s⁻¹), conditions favorable for turbulence suppression.
- **Orbit Physics:** Suprathermal protons exhibit orbit excursions up to 0.15–0.25 m, with orbit loss and recycling shaping edge plasma and pedestal conditions. Non-local orbit integration shows up to 15–20% of P_fusion is contributed by suprathermal protons.
- **Scaling:** R₀ = 1.4 m, a = 0.86 m, Iₚ = 13 MA, Bₜ = 3.0 T, V ≈ 30 m³, P_fusion ≈ 300 MW (Q ≳ 5 if P_aux ≲ 60 MW).
- **Challenges:** Sustained burn requires controlling edge recycling, minimizing plasma-wall interactions from strong suprathermal orbit losses, managing current drive with suprathermal electron populations carrying ≈21% of Iₚ, and mitigating power drain from relativistic bremsstrahlung.

## 6. Experimental and Evaluation Considerations

- **Cross-Section Standardization:** For nuclear and fusion physics, Bielik 11B v3 relies on the updated ¹¹B(p,α₀) cross section from 0.5–3.5 MeV, as established by recent high-precision experimental datasets resolving ambiguity in normalization and energy scale [1908.04064]. Linear interpolation between tabulated measurements is the recommended standard.
- **Supernova Neutrino Constraints:** In nucleosynthesis, the ¹¹B yield linkage to helium-burning rate uncertainties and supernova neutrino spectra/oscillations has been quantified [1102.4858]. Model uncertainties in λ₃α and λ_{12C α} dominate over oscillation effects, but improvements could make ¹¹B a probe for supernova ν-parameters.

## 7. Outlook and Future Research Directions

- **LLM Pathways:** The tokenizer specialization and FOCUS-based embedding transfer enable efficient adaptation to morphologically complex, under-resourced languages. The post-training alignment pipeline (SFT, DPO-P, GRPO) represents the state-of-the-art, and continued benchmarking against larger and specialized multilingual LLMs is in progress [2604.10799], [2601.11579].
- **Fusion Research Focus:** Priorities include validating multi-magnetofluid and suprathermal orbit predictions on EXL-50U/EHL-2, refining non-local ⟨σv⟩ modeling, integrated core-edge-wall simulations, and experimental investigation of advanced techniques (alpha-channeling, avalanche enhancement, spin-polarization) for power balance improvement [2604.04002].

Bielik 11B v3, both as a language technology and as a fusion reactor concept, stands at the intersection of algorithmic optimization, domain-specific adaptation, and empirical rigor, marking significant milestones in their respective fields.

Source: https://www.emergentmind.com/topics/bielik-11b-v3