- The paper introduces bridge-alpha curves and area, absolute-error, squared-error, and sup-norm tests to diagnose whether a factor model prices all portfolios generated by a prespecified characteristic axis.
- Using U.S. equities from 1967–2024, the study finds sign reversals across value, profitability, investment, and momentum axes, with HML and CMA-related models often overcorrecting while RMW- and UMD-containing models perform better on their corresponding axes.
- The results show that maximum-Sharpe improvements and axis-level pricing accuracy are largely distinct, so factor selection should evaluate both span expansion and localized pricing errors rather than relying on one criterion.
Motivation and contribution
Factor models are conventionally evaluated either by alpha tests on a selected set of test assets (GRS, characteristic-sorted deciles) or by spanning and maximum-Sharpe comparisons of factor spans. The first family depends on how portfolios are binned; the second does not reveal where pricing errors remain. This paper, by Useong Shin, fills the gap between these two approaches by generalizing an earlier cap-axis diagnostic (Shin, 2 Jul 2026) to arbitrary characteristics: value, operating profitability (OP), investment, and momentum. The core object is the bridge-alpha curve p↦αmx(p): for each cutoff p along a pre-specified characteristic rank order, a zero-investment "bridge" portfolio is formed as the prefix return minus an equal-exposure aggregate return, and its alpha under factor model m is estimated. The null is not pointwise significance on chosen deciles but a zero-curve restriction on the closed subspace Vx generated by the axis.
Theoretical structure
The construction is mechanical once the characteristic is fixed ex ante. Sorting stocks in descending order of characteristic x, with wealth shares as the measure, the bridge is
Dtx(p)=∫0prtx(u)du−pRtA,x,
which is closed at both endpoints, so a single prefix curve records body–tail offsets at every cutoff. Under finite-second-moment regularity, a proposition establishes that model m prices every return in Vx if and only if it prices the axis aggregate and has a zero bridge-alpha curve at all cutoffs — a complete test within the axis, explicitly not a full SDF test. Four functionals summarize each curve: signed area (SA), integrated absolute error (IAE), integrated squared error (p0), and sup norm (p1). Notably, p2 is implementable as the alpha of a single rank-area portfolio with linear weights p3, giving a clean HAC p4-test; p5, p6, and p7 are tested via simulated null distributions under a finite-grid Gaussian approximation. An important caveat is the aggregate gate: when the valid-characteristic subuniverse (e.g., CRSP∩Compustat) itself carries an alpha under the model, curve results are interpreted conditionally on that gate.
Data and implementation
The sample covers NYSE/AMEX/NASDAQ common stocks from 1967–2024, screened for investibility via market capitalization and liquidity hysteresis rules; the universe retains 99.7% of market capitalization while dropping extreme microcaps. Accounting axes use annual June formation with July–June holding; momentum uses monthly formation. Valid-universe coverage averages 79–81% for accounting axes and 97.5% for momentum. A sanity check confirms the internally constructed market return correlates at 0.9998 with published factors (daily RMSE ≈ 2 bp), reducing concern that results are artifacts of market-factor implementation. Replication checks of HML, RMW, and CMA show component-portfolio correlations of 0.98–0.998, with residual intercepts too small to explain the sign patterns.
Empirical findings: sign reversal and overcorrection
The central empirical pattern across all four axes is sign reversal. Models lacking the counterpart factor leave positive curves (undercorrection); adding the counterpart factor shifts the curve downward — but not always to zero. Key monthly results:
| Model |
Value |
OP |
Investment |
Momentum |
| CAPM |
+24.8 / NR |
+30.4 / B+ |
+43.0 / R+ |
+72.2 / R+ |
| FF3 |
−27.5 / R− |
+28.1 / B+ |
+16.9 / NR |
+92.6 / R+ |
| Carhart |
−27.6 / R− |
+25.2 / B+ |
+8.9 / NR |
−7.1 / NR |
| FF5 |
−28.5 / R− |
−5.7 / NR |
−20.1 / R− |
+91.2 / R+ |
| FF6 |
−27.9 / R− |
−6.9 / NR |
−22.1 / R− |
+7.0 / NR |
| q5 |
−19.8 / NR |
−17.1 / NR |
−26.8 / R− |
+3.3 / NR |
(Entries are annualized rank-area alphas in basis points, with verdict codes: R+/R− = significant rejection, B+ = borderline, NR = no rejection.)
Three contrasts stand out. On the value axis, HML-based models are rejected for negative overcorrection (FF5: −28.5 bp, p8) despite high rank-area p9 of roughly 0.75, while q5 — which contains no explicit value factor — passes cleanly. On the investment axis, the reversal is largest: CAPM's +43 bp becomes −20 to −27 bp in FF5, FF6, and q5, all with clean aggregate gates, while Carhart (no investment factor) is flattest, though conditionally due to gate failure. On the momentum axis, the diagnostic acts as a positive control: UMD-containing models (Carhart, FF6) move roughly 100 bp of distortion to within 7 bp of zero without rejection, and q5 passes without UMD through its ROE/EG block. On the OP axis, FF5 and FF6 achieve the ideal combination of high m0 (~0.53) and insignificant alphas, though daily diagnostics reveal a weak q5 overcorrection signal (m1 bp, m2).
A striking corollary is that FF3 and FF5 leave larger momentum-axis distortions than CAPM (+92.6 and +91.2 vs. +72.2 bp): adding non-momentum factor blocks amplifies rather than neutralizes momentum-axis pricing errors.
Sharpe gains versus axis distortion
The paper demonstrates that maximum-Sharpe improvement (m3) and axis-level pricing error are nearly separate coordinates. Across 155 candidate factors added one at a time to CAPM, the Spearman correlation between m4 and axis m5 is essentially zero on the value axis (0.041) and only weakly negative on OP (−0.357) and investment (−0.331). The q_EG factor is the clearest case: it delivers the largest m6 (1.270) yet leaves among the worst value-axis m7 (60.6 bp). Conversely, GFD OPE/BE achieves m8 bp on the OP axis with m9. More broadly, factors whose construction aligns closely with the sorting variable — OPE/BE for profitability, NOA_GR1A and INV_GR1 for investment, intermediate 12–7 momentum for momentum — consistently flatten their axes better than the canonical 2×3 or triple-sort counterpart factors. This suggests overcorrection stems from size-neutralized sort construction rather than from the underlying characteristic information, yielding a testable prediction (a size-split-free CMA should overcorrect less) that the author explicitly defers to future work.
Limitations and scope
The paper is candid about three restrictions. First, accounting-based axes are statements about the CRSP∩Compustat valid subuniverse, not the full market; gate-failing models (FF3, Carhart on several axes) admit only conditional interpretations. Second, verdicts are strictly axis-specific: no global ranking emerges, and the same model can pass one coordinate while being rejected on another. Third, a pass certifies pricing only of prefix, tail, interval, and step-function portfolios generated by that single characteristic order — not industry, other-anomaly, or idiosyncratic directions. Inference treats each axis separately rather than jointly, and nonlinear functionals rely on simulated Gaussian nulls rather than analytic distributions. The interpretation of overcorrection as a construction mismatch remains an interpretation, supported by the factor-scan gradient but not formally established.
Conclusion
This paper converts factor-model diagnosis into a functional problem: pricing errors along a fixed characteristic coordinate become a curve whose signed area, integral norms, and sup norm capture direction, magnitude, and local concentration. Applied to four canonical axes over 1967–2024, the diagnostic shows that containing a counterpart factor determines the direction of correction, while factor construction determines whether the axis is actually priced — with HML and CMA/IA overcorrecting into significant rejections and RMW and UMD flattening their axes within noise. The near-orthogonality of axis distortion to maximum-Sharpe gains indicates that span expansion and restricted zero-alpha pricing are distinct evaluation criteria, and the resulting model-by-axis fingerprints offer a complementary lens to GRS, spanning, and mean–variance comparisons.