---
title: Monetary Policy Expectations (MPE) Index
url: https://www.emergentmind.com/topics/monetary-policy-expectations-mpe-index
type: topic
---

# Monetary Policy Expectations (MPE) Index

Searching arXiv for the cited papers to verify metadata and ensure current grounding.
The **Monetary Policy Expectations (MPE) Index** denotes a text-derived measure of monetary-policy expectations, but the term is used for two distinct constructions in the recent literature. In one formulation, derived from Federal Open Market Committee (FOMC) statements and contemporaneous *New York Times* coverage, the index is the fitted informational-content component of a high-frequency monetary policy shock, identified as the projection of market movements onto the difference between FOMC and private-sector expectations [2111.06365]. In another formulation, developed for crypto-financial applications, the index is a weekly, engagement-weighted average of Large Language Model (LLM)-classified hawkish and dovish narratives extracted from StockTwits messages tagged with `$FED` and `$MACRO`, designed to capture ex-ante market expectations rather than realized policy implementation [2604.08825]. Across both uses, the core object is forward-looking monetary-policy information embedded in text, but the sampling frequency, unit of measurement, data source, and econometric role differ materially.

## 1. Conceptual scope and definitional variants

The term **MPE Index** is not attached to a single canonical formula across the two papers. In the FOMC-announcement framework, the paper itself refers to the **“information/news component”** rather than a named index, and the index can be defined from the method as
$$
MPE_t^{R} := \hat{\theta} D_t,
$$
where $D_t$ is the difference between text-implied FOMC expectations and text-implied private-sector expectations, and $\hat{\theta} D_t$ is the fitted component of a high-frequency shock variable such as FF4 or PNS explained by that difference [2111.06365]. In the StockTwits-based framework, the MPE index is explicitly a **high-frequency, text-derived measure of ex-ante market expectations about monetary policy**, aggregated at **weekly resolution** from LLM-classified investor communications [2604.08825].

These two constructions share a common interpretive axis: both attempt to isolate what market participants expect monetary policy to mean, rather than merely recording realized changes in the policy rate. The overlap is conceptual rather than operational. One series is event-based and denominated in **basis points (bps)** of a shock variable explained by information disclosure; the other is a weekly weighted average of ordinal hawkish/dovish scores. A plausible implication is that “MPE Index” is best treated as a family of text-based expectation measures rather than a single standardized macro-financial indicator.

| Variant | Core definition | Frequency / units |
|---|---|---|
| FOMC informational-content measure | Fitted value from projecting $\Delta R_t$ onto $D_t$ | Scheduled FOMC meetings / bps |
| StockTwits narrative measure | Engagement-weighted average of LLM-assigned scores $s \in \{-2,-1,0,+1,+2\}$ | Weekly / index level |

The distinction matters substantively. The FOMC-based measure is an identification device for decomposing high-frequency policy surprises into **information** and **monetary** components. The StockTwits-based measure is a narrative indicator used as an explanatory and predictive variable in VAR, VMD, and LSTM-SHAP analyses.

## 2. Event-based MPE from FOMC and private-sector text

In the 2021 framework, the informational content of FOMC announcements is identified by modeling the expectations of the FOMC and private-sector agents using computational linguistic tools on **FOMC statements** and *New York Times* articles, then projecting high-frequency financial-market movements onto differences in those expectations [2111.06365]. The sample covers **January 2000 to March 2014**, includes **106 scheduled meetings**, and excludes **unscheduled meetings**.

The text pipeline uses **BERT (bert-base-uncased; Hugging Face implementation)** to generate document embeddings. Preprocessing removes **numbers, dates, and stop-words (gensim list)**, tokenizes with the **Hugging Face default**, splits documents into **256-token windows with 10-token overlap**, and computes the document embedding as the mean of the BERT **[CLS] token** across windows. The output dimension is **768**. The description explicitly states that **no topic models or sentiment scorers are used**; the framework relies on **contextual embeddings and linear mapping**.

These embeddings are mapped to expected FFR decisions through two **linear elastic-net models**:
- $f^{FOMC}(X_t^{FOMC}) \approx E_t^{FOMC}(FFR)$
- $f^{NYT}(X_t^{NYT}) \approx E_t^{PS}(FFR)$

The target variable is the FOMC’s **target policy rate decision** $FFR_t$ at meeting $t$ in **bps**. The elastic net is selected for transparency and regularization, with tuning parameters $\lambda$ and $\eta$, and the models are chosen to maximize the **$R^2$ in Stage 3**, emphasizing predictive ability and interpretability over black-box complexity. The framework assumes the conditional mean restrictions
$$
E[\epsilon_t^{NYT}\mid X_t^{NYT}] = 0,\qquad E[\epsilon_t^{FOMC}\mid X_t^{FOMC}] = 0,
$$
and linearity in embeddings for tractability.

The expectations difference is
$$
D_t := E_t^{FOMC} - E_t^{PS}.
$$
This quantity is the driver of the identification strategy. It is not itself yet the index; rather, it is the text-based gap between what the FOMC statement implies and what private-sector reporting implied before or at the meeting.

## 3. Formal identification and decomposition in the FOMC framework

The high-frequency shock variable is denoted $\Delta R_t$ and may be a **Fed Funds Futures surprise (FF4)** or the **Policy News Shocks (PNS)** composite. Identification proceeds via the Stage 3 regression
$$
\Delta R_t = \zeta + \theta D_t + \nu_t,\qquad E[\nu_t\mid D_t]=0.
$$
The informational-content component is the fitted value from projecting $\Delta R_t$ onto $D_t$, written in vector form as
$$
\operatorname{proj}_D(\Delta R)=D(D'D)^{-1}D'\Delta R.
$$
In the scalar-$D_t$ case, this is equivalent to $\hat{\theta}D_t$ up to an intercept [2111.06365].

On that basis, the MPE Index for shock $R$ at meeting $t$ is defined as
$$
MPE_t^{R}:=\hat{\theta}D_t.
$$
The residual
$$
Monetary_t^{R}:=\Delta R_t-\hat{\theta}D_t
$$
is the **policy action residual**, yielding the decomposition
- **News\(_t\)** or information component: $\hat{\theta}D_t$
- **Monetary\(_t\)** or action component: $\Delta R_t-\hat{\theta}D_t$

The sign convention is explicit. A **positive MPE** indicates that the FOMC is more hawkish than the private sector, so markets learn **“tightening” or stronger inflation outlook from the announcement**, moving short rates up. A **negative MPE** indicates dovish information and lower short rates. No additional controls are required in Stage 3 because the theory posits that the information effect is proportional to $D_t$; the description further states that omitted-variable bias arises if one regresses on $f^{FOMC}(X_t^{FOMC})$ alone without $f^{NYT}(X_t^{NYT})$.

The high-frequency identification details are also tightly specified. **PNS** is computed over **30-minute windows around the statement release** and is the **first principal component of five unanticipated rate changes** involving current and near-term Fed funds expectations and **2–4 quarter eurodollar rates**. **FF4** follows the published **Gertler–Karadi** construction adopted by the paper. Validation uses daily changes in **nominal Treasury yields (3m, 1y, 2y, 5y, 10y, 20y)** and **TIPS real yields (2y, 5y, 10y, 20y)**. Following **Nakamura–Steinsson**, some displays omit **July 2008 to July 2009** to avoid crisis-period confounds.

## 4. Weekly LLM-based MPE from market messages

The 2026 construction defines the MPE index as a **weekly, engagement-weighted** measure of hawkish and dovish monetary-policy expectations extracted from **118,479 StockTwits messages** over **2014-09 to 2025-02** [2604.08825]. The corpus consists solely of messages tagged **`$FED`** and **`$MACRO`**, with no addition of news wires, central-bank speeches, analyst notes, or multilingual sources.

Classification is performed using **Mistral-7B (Jiang et al., 2023)**, run via the **ollama Python stack**. Each message is assigned one of five ordered categories:
- **Very Hawkish** $(−2)$
- **Hawkish** $(−1)$
- **Neutral** $(0)$
- **Dovish** $(+1)$
- **Very Dovish** $(+2)$

The sign convention is therefore the reverse of the event-based bps interpretation: here, **hawkishness is negative** and **dovishness is positive**. Ambiguous or descriptive messages default to **Neutral (0)**, intensity is encoded through the five-point scale, and **no continuous $[-1,1]$ scoring** is used. The paper does not report **a human-labeled evaluation set, accuracy/precision/recall, confusion matrices, or inter-annotator agreement**.

Let $s_i \in \{-2,-1,0,+1,+2\}$ denote the LLM-assigned score for message $i$, and let $w_i \ge 0$ denote engagement-based weight derived from **likes and reshares**. For week $t$, with message set $M_t$, the index is
$$
MPE_t=\frac{\sum_{i\in M_t} w_i s_i}{\sum_{i\in M_t} w_i}.
$$
The series is aligned to **week-ending Friday**. No additional normalization, such as **z-scoring**, and no smoothing, such as **exponential moving averages**, is applied.

The descriptive statistics reported for **2014–2025** are:
- mean **−0.03**
- standard deviation **0.09**
- skewness **−1.37**
- kurtosis **13.56**

The paper also notes that **Neutral comprises 95.96% of messages**, while extreme views attract more engagement, with **Very Hawkish averaging 1.32 interactions** against **0.31 for Neutral**. Beyond engagement-weighting, there are **no additional credibility filters or source-level reweighting**. The paper does not report **explicit deduplication or spam filtering** for index construction.

Although the empirical analysis uses the aggregate index, the description provides a conceptual decomposition into hawkish and dovish contributions:
$$
MPE_t^H=\frac{\sum_{i\in M_t,\ s_i<0} w_i s_i}{\sum_{i\in M_t} w_i},\qquad
MPE_t^D=\frac{\sum_{i\in M_t,\ s_i>0} w_i s_i}{\sum_{i\in M_t} w_i}.
$$

## 5. Empirical behavior and econometric roles

In the FOMC-based framework, the central empirical object is the Stage 3 regression of $\Delta R_t$ on $D_t$. The reported **pre-ZLB** results are:
- **PNS:** $R^2 = 0.079$; $\hat{\theta} = 0.083$; $p = 0.011$
- **FF4:** $R^2 = 0.085$; $\hat{\theta} = 0.183$; $p = 0.006$
- **FFR:** $R^2 = 0.038$; $\hat{\theta} = 0.051$; $p = 0.254$

For the **full sample (Jan 2000–Mar 2014)**, the reported result for **FF4** is
- **FF4:** $R^2 = 0.073$; $\hat{\theta} = 0.020$; $p = 0.017$

The description states that **PNS and FFR coefficients are positive but less significant** in the full sample [2111.06365]. The interpretation is direct: positive $\hat{\theta}$ implies that when FOMC information points to a higher optimal policy rate than markets expected, FF4 or PNS rises. The recovered series is therefore treated as economically meaningful and statistically significant for **FF4**, and for **PNS pre-ZLB**.

The time-series behavior is illustrated with specific episodes. **September 18, 2007** is described as a **large negative MPE**, reflecting markets inferring a deteriorating macro outlook from an aggressive cut. **March 18, 2008** is a **large positive MPE**, associated with inflation concerns. **January 30, 2002** is also a **positive MPE**, interpreted as communication of stabilization prospects.

The decomposition into **Monetary\(_t\)** and **News\(_t\)** is then used to study yield effects. For **PNS**, **News\(_t\)** significantly raises short-term nominal yields, including **3m: 1.208 bps per unit of News\(_t\); highly significant**, while longer real rates show limited or negative responses, including **TIPS 5–10y negative and significant**. For **FF4**, **News\(_t\)** significantly affects **3m nominal yields**, whereas **Monetary\(_t\)** explains most **mid- to long-term nominal yield** movements. The stated economic meaning is that the information effect **primarily moves short maturities**, consistent with revelation about near-term macro conditions and policy stance.

In the StockTwits-based framework, the MPE index is not a decomposition term but a predictive explanatory variable. It is designed to distinguish **ex-ante expectations** from **ex-post policy moves**, and the reported correlations are:
- $\operatorname{Corr}(MPE, FFR) = -0.11$
- $\operatorname{Corr}(MPE, SP500) = 0.11$
- $\operatorname{Corr}(MPE, VIX) = -0.03$
- $\operatorname{Corr}(MPE, 5y\ inflation\ expectations) = -0.29$
- $\operatorname{Corr}(MPE, Google\ “inflation”\ searches) = -0.41$

The econometric specification is a **bivariate VAR** for Bitcoin returns with lag order chosen by **AIC**, testing the null that lagged MPE terms do not enter the return equation. The reported result is that the MPE index **Granger-causes Bitcoin returns at lags 3–5**, with **p-values 0.069, 0.092, 0.089**. In the frequency domain, **Variational Mode Decomposition (VMD)** splits variables into **three intrinsic mode functions (IMFs)**: **long-term (IMF1)**, **medium-term (IMF2)**, and **high-frequency (IMF3)**. The reported scale-specific result is that **MPE IMF2 significantly predicts BTC IMF3 at lags 2 and 3**, with **$F = 6.495, p = 0.002$** and **$F = 3.991, p = 0.008$** [2604.08825].

## 6. Modeling extensions, interpretation, and limitations

The 2026 study supplements linear causality tests with an **LSTM** that predicts **one-week-ahead BTC return** using **18 macro/market predictors plus lagged BTC**, trained in an **expanding-window walk-forward scheme** with **4 folds**, **Optuna** tuning, **MSE** loss, and **RMSE** and **MAE** evaluation metrics [2604.08825]. Reported performance is:
- **LSTM overall RMSE 0.0896 (MAE 0.0709)**
- **ARIMA overall RMSE 0.0873 (MAE 0.0666)**

**Diebold–Mariano tests** show **no significant difference in 1-step predictive accuracy across folds ($p > 0.05$)**, and the LSTM is retained for its ability to capture **non-linear, regime-dependent interactions**. **SHAP** then ranks **MPE** as the **third most important predictor on average**, behind **Google “inflation”** and **FFR**, with global importances **0.008**, **0.006**, and **0.005**, respectively. Temporal attribution assigns notable SHAP mass at **$t$, $t-4$, $t-5$, $t-6$**, described as **“predictive memory.”** Stratification by **FFR regime (Rising/Flat/Falling)** shows a **robust negative slope** between hawkish MPE readings and SHAP contributions to Bitcoin returns, including in the **Flat regime**, which the paper interprets as narrative effects independent of realized policy adjustments.

The FOMC-based measure is explicitly compared with the high-frequency shock literature. It is said to **refine standard high-frequency shocks (FF4, PNS)** by isolating the central bank information effect via **text-based expectation differences**, and to complement prior decompositions such as **Jarociński–Karadi’s monetary vs information shocks**, but by directly modeling **private vs FOMC expectations from text rather than imposing sign/moment restrictions** [2111.06365]. The incremental explanatory power is summarized as **nontrivial $R^2$ for FF4**, reaching **8.5% pre-ZLB** and **7.3% full sample**.

Both constructions also carry clear caveats. In the FOMC case, the limitations include the **rational expectations conditional on text embeddings** assumption, **linearity** in Stage 2 and Stage 3, possible loss of information from removing **numbers/dates**, the use of **general-corpus BERT**, the **small sample** of **106 scheduled meetings**, exclusion of **press conference vs statement effects**, exclusion of **unscheduled meetings**, and possible post-crisis underperformance at the **zero lower bound**. The description adds that the simple linear specification is interpreted as producing **lower bounds**.

In the StockTwits case, the stated limitations include **classification risk** due to the absence of a human-labeled validation set, **coverage bias** because StockTwits skews toward active retail or semiprofessional investors focused on U.S. policy, the dominance of **Neutral** messages, the possibility that engagement-weighting amplifies **echo-chamber effects**, sensitivity to **weekly aggregation** and **Friday-close alignment**, and the possibility that **LLM updates or platform behavior changes** may shift classification boundaries. External validity beyond **U.S. monetary policy narratives** and beyond the studied crypto context is not assessed.

A common misconception is that an MPE index simply restates realized policy-rate changes. Neither construction is designed that way. The FOMC-based measure isolates the fitted **information disclosure** component of a high-frequency surprise, while the StockTwits-based measure is explicitly distinguished from realized **EFFR** changes by construction and by weak correlation. Another potential misconception is that the two indices are directly interchangeable. They are not: one is an **event-study projection in basis points**, the other a **weekly narrative score**. What they share is the objective of quantifying expectations embedded in monetary-policy communication and market discourse rather than treating observed rate changes as sufficient statistics for policy transmission.

Source: https://www.emergentmind.com/topics/monetary-policy-expectations-mpe-index