---
title: Binary Input Symmetric Output (BISO) Channels
url: https://www.emergentmind.com/topics/binary-input-symmetric-output-biso-channels
type: topic
---

# Binary Input Symmetric Output (BISO) Channels

Binary input symmetric output (BISO) channels are discrete memoryless channels with binary input \(X\in\{0,1\}\) whose output alphabet can be arranged in symmetric pairs so that the conditional law under input \(0\) is the mirror image of the conditional law under input \(1\). In one standard formulation, the output alphabet is \(\{k:-l\le k\le l\}\) and the symmetry is \(P(Y=k\mid X=0)=P(Y=-k\mid X=1)\); in another, there exists an involutive permutation \(\pi\) such that \(W(y\mid 1)=W(\pi(y)\mid 0)\) and \(\pi^{-1}=\pi\). This symmetry implies that the uniform input distribution is capacity-achieving, so the symmetric capacity \(I(W)\) coincides with Shannon capacity. BISO channels therefore form a tractable but still expressive class that includes the binary symmetric channel (BSC) and binary erasure channel (BEC), and they recur in polarization, broadcast-channel ordering, secrecy, and low-complexity code design [1001.2062] [0807.3917].

## 1. Formal model and equivalent symmetry descriptions

A BISO channel is a binary-input discrete memoryless channel whose output alphabet admits a reflection symmetry. One explicit representation uses
\[
\mathcal Y=\{0,\pm 1,\pm 2,\dots,\pm l\},
\]
with
\[
P_{Y|X}(y|0)=P_{Y|X}(-y|1)\coloneqq p_y.
\]
A singleton output \(0\) can be split into \(0_+\) and \(0_-\) to ensure an even-sized output alphabet when convenient. In the broader symmetric-channel notation used in polar coding, the same structure is expressed through an involution \(\pi\) on outputs satisfying \(W(y|1)=W(\pi(y)|0)\) and \(\pi^{-1}=\pi\). For symmetric binary-input channels in this sense, uniform input is capacity-achieving [2504.16726] [0807.3917].

Related terminologies in the cited literature are binary-input memoryless output-symmetric (BMS), binary-input output-symmetric (BIOS), and binary-input symmetric-output memoryless (BISOM). One BMS formulation states that a memoryless channel with binary input \(X\) and output \(Y\) is BMS if there exists a sufficient statistic
\[
T(Y)=(X\oplus Z_A,A),
\]
where \((A,Z_A)\) are statistically independent of \(X\), and \(Z_A\) is binary with \(\Pr(Z_A=1\mid A=a)=a\). This places BISO channels inside a larger family of binary-input symmetric channels whose outputs can be represented as a random state plus a state-dependent binary symmetric corruption [2401.14710].

The significance of these equivalent descriptions is structural. The involutive-output view is natural for coding theorems and polarization; the paired-output view is natural for Lorenz-curve and partial-order arguments; and the sufficient-statistic view is natural for extremality results involving BSC and BEC comparisons. This suggests that the usefulness of BISO channels comes less from a single canonical parameterization than from the persistence of symmetry across multiple analytical frameworks.

## 2. Capacity, Bhattacharyya parameter, and basic extremes

For a binary-input channel \(W\), the symmetric capacity is the mutual information under uniform input. For BISO channels, this is the ordinary channel capacity because the uniform input distribution is optimal. A second fundamental quantity is the Bhattacharyya parameter
\[
Z(W)=\sum_y \sqrt{W(y\mid 0)W(y\mid 1)},
\]
which measures reliability and is central in both coding bounds and polarization [0807.3917].

A general capacity characterization in terms of \(Z(W)\) holds for all binary-input memoryless channels and therefore specializes directly to BISO channels:
\[
1-Z(W)\le C(W)\le 1-\mathcal H_b\!\left(\frac{1-\sqrt{1-Z^2(W)}}{2}\right),
\]
where \(\mathcal H_b\) is the binary entropy function. In the BISO case these inequalities apply directly to the Shannon capacity because \(C(W)=I(W;1/2)\). The bounds are sharp at the endpoints: \(Z(W)=0\) gives \(C(W)=1\), and \(Z(W)=1\) gives \(C(W)=0\) [1710.06908].

These reliability descriptors are complemented by contraction-type quantities that become unusually explicit on BISO channels. For example, one recent closed form for the KL contraction coefficient is
\[
\eta_{KL}(P)=\sum_{y>0}\frac{(p_y-p_{-y})^2}{p_y+p_{-y}},
\]
written in terms of the symmetric output pairs. This is analytically valuable because quantities that are usually defined by variational optimization become simple finite sums in the BISO class [2504.16726].

Taken together, \(I(W)\), \(Z(W)\), and the contraction coefficients provide several non-equivalent but tightly linked ways to quantify information retention. A plausible implication is that BISO channels are often used as a testbed not because they trivialize information theory, but because they make distinct information measures simultaneously computable.

## 3. Lorenz curves, order relations, and extremal channels

For BISO channels, the transition law can be reorganized into a geometric object called the BISO curve, and its integral
\[
F(t)=\int_0^t f(\tau)\,d\tau
\]
is the Lorenz curve. In this representation, the capacity is
\[
C=1-F(1).
\]
If two BISO channels with the same capacity have Lorenz curves \(F(t)\) and \(G(t)\) such that \(F(t)\le G(t)\) for all \(t\in[0,1]\), then the first channel is more capable than the second. This leads to the extremal sandwich
\[
BEC(C)\gg F(C)\gg BSC(C),
\]
meaning that among BISO channels of a fixed capacity, the BEC is most capable and the BSC is least capable. For output alphabet size at most \(3\), two equal-capacity BISO channels are always more-capable comparable, but the paper also gives counterexamples showing that comparability fails in general [1001.2062].

An especially striking property is that, for equal-capacity BISO channels, the more-capable and essentially-less-noisy orders reverse each other:
\[
F_1\gg F_2 \iff F_2\succeq F_1.
\]
Within broadcast-channel theory, this reversal yields a sharp criterion: for two BISO channels with the same capacity, superposition coding is optimal if and only if the channels are more-capable comparable. The same work also shows that if either receiver channel is a BSC or a BEC, then the superposition coding region is the capacity region [1001.2062].

More recent order-theoretic work extends this extremal picture beyond Shannon mutual information. For BISO channels with the same KL contraction coefficient, the BEC and BSC are extremal with respect to the less noisy order:
\[
BEC(C_\eta)\preceq_{\mathrm{ln}} F(C_\eta)\preceq_{\mathrm{ln}} BSC(C_\eta).
\]
For fixed Dobrushin coefficient, equivalently fixed Doeblin coefficient or maximum leakage, they are extremal with respect to degradability:
\[
BEC(C_\alpha)\preceq_{\mathrm{deg}} F(C_\alpha)\preceq_{\mathrm{deg}} BSC(C_\alpha).
\]
In the same binary-input setting,
\[
\eta_{TV}(P)=1-\alpha(P)=\alpha_{\max}(P)-1=e^{(X\rightarrow Y)}-1.
\]
These identities connect total-variation contraction, Doeblin mass, and leakage in a single algebraic chain [2504.16726].

A further extension replaces Shannon mutual information by Sibson Rényi mutual information and introduces \(\alpha\)-Lorenz curves. The cited theorem states that if two BISO channels have the same \(\alpha\)-capacity and \(F_{\alpha,1}(t)\le F_{\alpha,2}(t)\), then for \(\alpha>1\), \(\frac12\le\alpha<1\), and \(0<\alpha\le\frac13\), one has \(W_1\succeq_\alpha W_2\), whereas for \(\frac13\le\alpha\le\frac12\), the order reverses. At \(\alpha=\frac13\) and \(\alpha=\frac12\), all equal-\(\alpha\)-capacity BISO channels are equivalent under the \(\alpha\)-more-capable order. The BEC and BSC remain the two extremal channels throughout this framework [2508.19951].

## 4. Polarization and explicit coding constructions

Because every BISO channel is a symmetric binary-input discrete memoryless channel, Arıkan’s channel-polarization framework applies directly. Starting from \(N=2^n\) independent copies of \(W\), the first combining step is
\[
W_2(y_1,y_2\mid u_1,u_2)=W(y_1\mid u_1\oplus u_2)\,W(y_2\mid u_2),
\]
and the full transform is
\[
x_1^N=u_1^N G_N,\qquad G_N=B_NF^{\otimes n},\qquad
F=\begin{bmatrix}1&0\\1&1\end{bmatrix}.
\]
This produces synthesized subchannels \(W_N^{(i)}\) corresponding to the successive-cancellation decisions. The polarization theorem states that for any fixed \(\delta\in(0,1)\), the fraction of indices with \(I(W_N^{(i)})\in(1-\delta,1]\) tends to \(I(W)\), while the fraction with \(I(W_N^{(i)})\in[0,\delta)\) tends to \(1-I(W)\) [0807.3917].

This asymptotic dichotomy leads to polar codes. For any rate \(R<I(W)\), there exist information sets \(\mathcal A_N\subset\{1,\dots,N\}\) with \(|\mathcal A_N|\ge NR\) such that
\[
Z\!\left(W_N^{(i)}\right)\le O(N^{-5/4})\qquad \text{for all } i\in\mathcal A_N.
\]
Information bits are placed on \(\mathcal A_N\), the complement is frozen, and the resulting code has block length \(N=2^n\), rate at least \(R\), block error probability under successive cancellation bounded by
\[
P_e(N,R)=O(N^{-1/4}),
\]
and encoding and decoding complexity \(O(N\log N)\). For symmetric channels, including BISO channels, the performance is independent of the specific frozen vector [0807.3917].

Single-step recursions clarify the mechanism:
\[
I(W')+I(W'')=2I(W),\qquad Z(W'')=Z(W)^2,\qquad Z(W')\le 2Z(W)-Z(W)^2.
\]
For the BEC these become equalities, making the good/bad split completely explicit. This conservation-plus-separation behavior is the core of channel polarization [0807.3917].

The kernel viewpoint also generalizes. For an invertible binary \(\ell\times \ell\) matrix \(G\), repeated application of \(G^{\otimes n}\) polarizes symmetric binary-input memoryless channels if and only if \(G\) is not upper triangular; if \(G\) is upper triangular, every synthesized channel remains equivalent to the original channel. The same work retains \(O(N\log N)\) complexity for block length \(N=\ell^n\) [0811.1770].

A recent algebraic refinement shows that for symmetric underlying channels the synthetic channels generated by Arıkan transformations can be characterized as random switching channels of BSCs. In particular, Arıkan transforms preserve symmetry, and the likelihood-ratio profile of a synthetic channel can be tracked through explicit \(\star\) and \(\diamond\) operations on BSC crossover parameters. This suggests a compressed representation of polarized BISO channels in terms of BSC mixtures rather than full output alphabets [2510.22896].

## 5. Broadcast, wiretap, and feedback-assisted secrecy

In binary-input broadcast channels, a central inequality states that for any \((U,V,X,Y,Z)\) satisfying \((U,V)\to X\to (Y,Z)\),
\[
I(U;Y)+I(V;Z)-I(U;V)\le \max\{I(X;Y),I(X;Z)\}.
\]
Combined with cardinality reduction, this implies that for every binary-input broadcast channel, the maximum sum-rate of Marton’s inner bound equals the randomized time-division sum-rate. Since BISO broadcast channels are a subclass of binary-input broadcast channels, this simplification applies to them directly [1001.1468].

Within the more specialized BISO broadcast setting, equal-capacity comparability controls coding optimality. If two BISO receiver channels are more-capable comparable, superposition coding is optimal; if they are not more-capable comparable, the cited theorem states the equivalence of several strict inclusions, including \(TD\subset MIB\) and \(MIB\subset OB\). This isolates BISO subclasses where the best known inner and outer bounds differ [1001.2062].

The same symmetry is decisive in secrecy. For a binary-input symmetric physically degraded wiretap channel, polar coding achieves the secrecy capacity
\[
C_s=C(G_{Y|X})-C(Q_{Z|X}),
\]
and, more strongly, the entire rate-equivocation region
\[
\left\{(R,R_e):
0\le R\le C(G_{Y|X}),\;
0\le R_e\le R,\;
R_e\le C(G_{Y|X})-C(Q_{Z|X})
\right\}
\]
under weak secrecy. The construction places random bits on the subchannels that are good for the eavesdropper and secret bits on the subchannels that remain good for the legitimate receiver [1005.2759].

A BSC specialization shows how feedback can be turned into a secrecy resource even when the forward wiretap asymmetry is unfavorable. In the cited feedback scheme, the induced equivalent wiretap channel is again binary symmetric, with unscaled secrecy rate
\[
R_{s,u}=h(\epsilon_b+\delta_b-2\epsilon_b\delta_b)-h(\epsilon_b).
\]
The paper proves that this can yield a strictly positive secrecy rate even when the eavesdropper’s forward channel is less noisy than the legitimate receiver’s forward channel. Its appendix also identifies the relevant subproblem as a binary-input symmetric-output broadcast channel, and in the binary-\(U\) case the optimal auxiliary channel \(U\to X\) is itself a BSC with uniform \(U\) [0909.5120].

## 6. Contemporary extensions and structured ensembles

Recent work uses the BISO class to sharpen extremal mutual-information comparisons. For inputs uniform on shifted linear codes over BMS channels, the mutual information through a BSC of capacity \(t\) is lower bounded by a constant fraction of the mutual information through a BEC of the same capacity:
\[
\alpha_t\cdot I_{\mathrm{BEC}^{(t)}}(X^n;Y^n)\le I_{\mathrm{BSC}^{(t)}}(X^n;Y^n)\le I_{\mathrm{BEC}^{(t)}}(X^n;Y^n),
\]
with
\[
\alpha_t=\frac{t}{\eta_t},\qquad \eta_t=(1-2h^{-1}(1-t))^2,
\]
and the paper notes the uniform lower bound \(\alpha_t>\log_2(e)/2\) for all \(0<t\le 1\). The same paper also derives a general information-combining lower bound
\[
I(X;Y^n)\ge \frac{I(P_X,W)}{\eta(P_X,W)}\left(1-(1-\eta(P_X,W))^n\right)
\]
for arbitrary \(P_X\) and channels \(W\) [2401.14710].

Capacity-approaching sparse-graph constructions have also been developed specifically for symmetric binary-input channels. Irregular LDGM-LDPC ensembles are analyzed as subcodes of punctured LDPC codes and are shown to achieve rates arbitrarily close to the capacity of BISOM channels with bounded complexity, where complexity is measured by the average check-node degree per information bit. A key near-capacity condition is puncturing of the form
\[
p=1-\kappa\epsilon,
\]
which keeps the complexity lower bound finite as the gap to capacity \(\epsilon\) tends to zero [1003.2454].

A different line studies repetition and superposition (RaS) codes over BIOS channels. The cited theorem proves that block RaS codes are capacity-achieving over BIOS channels in frame-error rate, extends the same framework to source coding and joint source and channel coding, and shows that the associated enlarged QC-LDPC ensemble can also achieve capacity. The convolutional extension, Conv-RaS, is proved capacity-achieving in first error event probability and is described as a universal JSCC scheme with flexible rate [2402.13603].

A separate analytical perspective models discrete symmetric channels thermodynamically. For the BSC, a 4-symbol symmetric channel, and general discrete memoryless symmetric channels with equiprobable symbols, the mutual information is derived from a generalized second law and reduces to
\[
I=-\gamma U(\gamma)\big|_0^\beta
\]
because the correction term involving the temperature-dependent Hamiltonian vanishes identically in that class. In the BSC case this recovers \(I=1-H(\delta)\) [0807.4322].

Across these developments, BISO channels remain valuable for the same reason: symmetry forces enough regularity to make exact formulas, extremal comparisons, and explicit code constructions possible, while still leaving room for nontrivial phenomena such as non-comparable broadcast receivers, kernel-dependent polarization, and structured capacity-achieving ensembles.

Source: https://www.emergentmind.com/topics/binary-input-symmetric-output-biso-channels