Papers
Topics
Authors
Recent
Search
2000 character limit reached

Forward Stagewise Multiview Boosting

Updated 25 March 2026
  • Forward Stagewise Additive Multiview Boosting is an ensemble method that trains weak learners on distinct feature views to address multiclass classification challenges.
  • The approach introduces a novel exponential loss function with 1/V normalization, ensuring that errors from multiple views are collaboratively penalized to improve convergence.
  • Empirical results show that SAMA-AdaBoost achieves faster convergence, higher margins, and better generalization compared to traditional boosting methods.

Forward Stagewise Additive Multiview Boosting refers to a mathematically grounded ensemble learning approach in which weak learners are trained collaboratively across multiple feature subsets ("views") for multiclass classification. The SAMA-AdaBoost algorithm, the canonical representative of this class, extends traditional forward stagewise (additive) boosting to a multiview setting by minimizing a novel exponential loss tailored to collaborative, multiclass prediction. This approach is characterized by a rigorous mathematical framework, explicit convergence and margin bounds, and an emphasis on collaborative regularization among weak learners from different views (Lahiri et al., 2016).

1. Problem Setup and Mathematical Notation

The multiview boosting scenario considers a labeled training set:

S={(xi1,xi2,…,xiV,yi)}i=1NS = \{(x_i^1, x_i^2, \dots, x_i^V, y_i)\}_{i=1}^N

where VV is the number of views (feature subsets), each xiv∈Xvx_i^v \in \mathcal{X}^v is the vv-th view of instance ii, and yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\} is the multiclass label encoded as a one-hot vector Yi∈{0,1}KY_i \in \{0,1\}^K, with (Yi)k=1(Y_i)_k = 1 iff yi=ky_i = k.

On each view vv, a series of weak learners VV0 are trained. Each weak learner's output is mapped to a signed one-hot encoded VV1-vector VV2: VV3

2. Forward Stagewise Additive Model

The SAMA-AdaBoost algorithm generalizes the forward additive model to multiview multiclass learning: VV4 where VV5 are the stagewise weights. The single-view case corresponds to classical boosting such as SAMME.

A novel exponential loss function that reflects the aggregate margin over all views is defined: VV6 The overall loss to minimize is

VV7

This loss function upweights examples that are misclassified by a greater number of views.

3. Stagewise Optimization and Weight Updates

Boosting proceeds in rounds. At each round VV8:

  • Instance weights are defined by the current ensemble margin:

VV9

  • The optimal set of weak learners xiv∈Xvx_i^v \in \mathcal{X}^v0 and shared step-size xiv∈Xvx_i^v \in \mathcal{X}^v1 are selected to minimize

xiv∈Xvx_i^v \in \mathcal{X}^v2

where xiv∈Xvx_i^v \in \mathcal{X}^v3 denotes the count of views which misclassify xiv∈Xvx_i^v \in \mathcal{X}^v4.

  • xiv∈Xvx_i^v \in \mathcal{X}^v5 is determined numerically as the minimum of a strictly convex function:

xiv∈Xvx_i^v \in \mathcal{X}^v6

  • The weights are updated:

xiv∈Xvx_i^v \in \mathcal{X}^v7

with renormalization to ensure xiv∈Xvx_i^v \in \mathcal{X}^v8.

  • At prediction time, the ensemble outputs class xiv∈Xvx_i^v \in \mathcal{X}^v9 maximizing the weighted vote:

vv0

This stagewise process ensures that no single view can dominate, and examples misclassified by more views receive stronger weight adjustments.

4. Regularization and View Collaboration

The exponential loss contains a vv1 normalization in the exponent, regularizing the influence of individual views. An example incorrectly classified by only a small fraction of views receives a moderate upweight in loss, while broad disagreement among views leads to more significant penalization. This mechanism restricts overlearning by weak learners in any single view and addresses overfitting by fostering collaboration among views.

5. Convergence Analysis and Margin Bounds

Two principal theoretical guarantees underpin SAMA-AdaBoost:

  • Training Error Upper Bound: Let vv2 at round vv3.

vv4

As each vv5, the bound decays to zero as vv6.

  • Margin-Based Generalization Bound: Define a normalized classifier

vv7

For margin vv8,

vv9

This quantifies the decay in low-margin examples during training and establishes superior convergence properties over previous multiview and classical boosting formulations.

6. Comparative Performance and Empirical Findings

SAMA-AdaBoost exhibits several advantages compared to prior methods:

  • Versus traditional AdaBoost and SAMME, SAMA-AdaBoost demonstrates faster convergence in (theoretical) training-error bounds and produces higher margins, indicative of improved generalization.
  • Relative to earlier heuristic models such as MA-AdaBoost, SAMA-AdaBoost employs an exact convex minimization for the step-size, resulting in solutions closer to the global minimum for the exponential loss.
  • Compared to other multiview algorithms (e.g., Mumbo, 2-Boost, Co-AdaBoost, AdaBoost.Group), SAMA-AdaBoost:
    • Scales to multiclass and ii0 views settings
    • Relies on forward-stagewise optimization, not heuristic weight-transfer strategies
    • Achieves lower test error at equivalent boosting rounds on benchmarks such as 100-Leaves, eye/non-eye, MNIST, and standard UCI datasets
    • Produces accurate, diverse ensemble members (as evidenced in kappa-error diagrams) and demonstrates robustness under label noise
    • Is computationally faster per round than SAMME or Mumbo, leveraged by training multiple weak learners in parallel within low-dimensional feature spaces

7. Algorithmic Summary

A summary of the SAMA-AdaBoost algorithmic workflow is as follows:

  • Input: Training dataset ii1 for ii2, rounds ii3
  • Initialize weights ii4
  • For ii5 to ii6:

    1. Train each view's weak learner ii7 using ii8
    2. For each instance, compute ii9
    3. (Optional) Select a subset of high-performing views
    4. Solve for yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}0 minimizing yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}1
    5. Update yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}2, then renormalize
  • Output: Classifier yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}3 assigns the class yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}4 maximizing the summed, weighted votes across rounds and views

  • Termination: When yi∈{1,2,…,K}y_i \in \{1,2,\dots,K\}5 rounds are reached or the error bound falls below a specified threshold

SAMA-AdaBoost provides a principled, collaborative multiview boosting methodology that robustly optimizes a loss function tailored to multiclass, multiview settings, yielding both theoretical and empirical improvements over prior boosting frameworks (Lahiri et al., 2016).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Forward Stagewise Additive Multiview Boosting.