Papers
Topics
Authors
Recent
Search
2000 character limit reached

Asymmetric Tile Buffering (ATB) in GEMM

Updated 27 November 2025
  • Asymmetric Tile Buffering (ATB) is a tiling strategy for GEMM that decouples input and output tile dimensions to enhance arithmetic intensity and throughput.
  • Performance models demonstrate that ATB can achieve up to 4.54× speedup over symmetric tiling, as evidenced by case studies on AMD’s XDNA2 AI Engine.
  • ATB optimizes buffer utilization by balancing the trade-offs between higher arithmetic intensity and increased kernel switching overhead.

Asymmetric Tile Buffering (ATB) is a tiling strategy for general matrix multiplication (GEMM) that decouples the dimension of buffered input AA tiles along the MM axis from the dimension of output CC tiles, permitting significant enhancements in arithmetic intensity and overall throughput. Introduced in the context of accelerating AI workloads, ATB provides a methodical means to optimize buffer utilization in manycore and array architectures by exploiting a previously overlooked asymmetry in input/output operand buffering. Systematic performance modeling reveals that ATB can offer substantial real-world speedup over conventional symmetric tiling, as established in a detailed case study on AMD’s XDNA2 AI Engine (Wang et al., 20 Nov 2025).

1. Definition and Formulation

In symmetric GEMM tiling, a single set of tile dimensions (TM,TK,TN)(T_M,T_K,T_N) governs the size of buffered blocks for AA, BB, and CC:

  • AA-tiles are TM×TKT_M \times T_K,
  • BB-tiles are MM0,
  • MM1-tiles are MM2.

Asymmetric Tile Buffering introduces four distinct tile parameters:

  • MM3: rows of MM4 buffered,
  • MM5: rows of MM6 buffered,
  • MM7: reduction-dimension tile size,
  • MM8: columns of MM9/CC0 buffered,

with the constraint CC1. The asymmetry ratio is defined as

CC2

which quantifies how many output rows are accumulated per input row loaded. Figure 1 in (Wang et al., 20 Nov 2025) illustrates the structural difference between symmetric tiling and ATB.

2. Performance Model and Analytical Framework

The mathematical performance model comprises several core metrics:

Arithmetic Intensity (CC3):

Let CC4, CC5, CC6 denote per-element bytes of CC7, CC8, CC9 (e.g., for BF16, (TM,TK,TN)(T_M,T_K,T_N)0), and (TM,TK,TN)(T_M,T_K,T_N)1 the global reduction length. For output-stationary scheduling: (TM,TK,TN)(T_M,T_K,T_N)2

(TM,TK,TN)(T_M,T_K,T_N)3

Substituting (TM,TK,TN)(T_M,T_K,T_N)4 and simplifying gives: (TM,TK,TN)(T_M,T_K,T_N)5

This is subject to the L1 buffer constraint: (TM,TK,TN)(T_M,T_K,T_N)6

Kernel-Switching Overhead:

With a microkernel switching cost (TM,TK,TN)(T_M,T_K,T_N)7,

(TM,TK,TN)(T_M,T_K,T_N)8

Combined Throughput:

Let (TM,TK,TN)(T_M,T_K,T_N)9 be the peak per-core MAC rate and AA0 the microkernel efficiency.

AA1

AA2

32-core throughput is bounded by

AA3

3. Trade-offs and Design Space

ATB’s principal trade-off is between maximized arithmetic intensity and increased kernel switching overhead:

  • High AA4, small AA5: Favors greater arithmetic intensity and memory reuse but exacerbates switching overhead and reduces microkernel efficiency due to shorter steady phases.
  • Low AA6, large AA7: Improves core efficiency by yielding longer microkernel chains and fewer invocations, but at the expense of attainable arithmetic intensity.

The optimal configuration lies where buffer, compute, and switching costs are jointly minimized. This point is quantitatively determined by jointly satisfying constraints and optimizations in equations (1), (9), (11), and (12) from the analytical model.

4. Parameter Selection Guidelines

For effective deployment of ATB, the following process is recommended:

  1. Choose AA8: Select a AA9 large enough that microkernel efficiency BB0–BB1 (see Table 1).
  2. Increase BB2: Grow the asymmetry ratio until switching overhead or total buffer capacity becomes the limiting factor.
  3. Buffer Allocation: Allocate available buffer to maximize BB3, thus enlarging the “memory reuse volume” and boosting BB4.
  4. Full Array Evaluation: Simulate array-level performance; if memory-bound, consider higher BB5; if compute-bound, adjust by increasing BB6 or reducing BB7.

Table 1. Microkernel/core performance under Config 1, BB8, BB9:

CC0 CC1 CC2 (TF) CC3
8 1 0.36 0.156
8 4 0.36 0.134
32 4 0.75 0.312
64 4 1.16 0.511

5. Practical Implementation and Architectural Case Study

ATB’s effectiveness was demonstrated on AMD’s XDNA2 AI Engine comprising 32 compute cores (4 × 8), each core with 64 KB L1 and two input/output streams. The studied GEMM used mixed-precision (BFP16/BF16) with the following configuration (Config 1, Table 3):

  • Problem Size: CC4
  • L1 Tile: CC5, CC6
  • Buffer Used: CC7 KB (vs. 91 KB if symmetric)
  • Measured Throughput: CC8 TFLOPS/core × 32 cores = CC9 TFLOPS (compute limit)
  • Arithmetic Intensity (array): AA0 op/B (memory limit: AA1 TFLOPS)
  • Final Throughput: AA2 TFLOPS
  • Speedup: AA3 (baseline MLIR-AIE symmetric tiling achieves AA4 TFLOPS)

Table 3. Impact of ATB on throughput:

L1 tile AA5 AA6 AA7 (TF) Speedup
AA8 (symmetric) 1 4.8 1.00×
AA9 1 17.3 3.61×
TM×TKT_M \times T_K0 4 24.3 4.54×

Table 3 and additional configurations (see (Wang et al., 20 Nov 2025), Table 3) confirm 2–3× throughput gains across other precisions. ATB enlarges the feasible TM×TKT_M \times T_K1 memory reuse volume, often doubling or tripling arithmetic intensity in fixed scratchpad resources.

6. Practical Considerations and Implementation

ATB is a minor extension to standard tiling loops for GEMM: only TM×TKT_M \times T_K2 rows of TM×TKT_M \times T_K3 are buffered while accumulation to TM×TKT_M \times T_K4 rows of TM×TKT_M \times T_K5 proceeds. This increased reuse of the TM×TKT_M \times T_K6 tile directly leverages buffer capacity for enhanced output accumulation, enabling larger TM×TKT_M \times T_K7 without surpassing scratchpad constraints. ATB is particularly beneficial under tight buffer budgets or on architectures that can tolerate moderate kernel switching overheads.

A plausible implication is that architectural features such as hardware support for fast context switching and flexible buffer management can further amplify the benefits of asymmetric tile buffering in practice. However, performance gains are contingent on careful tuning of buffer, reduction, and output tile parameters as prescribed by the performance model (Wang et al., 20 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Asymmetric Tile Buffering (ATB).