Papers
Topics
Authors
Recent
Search
2000 character limit reached

Reddit Inflation Scores (RIS)

Updated 5 July 2026
  • Reddit Inflation Scores (RIS) are a family of measures derived from Reddit data that quantify vote-induced score changes, narrative signals, and excess activity anomalies.
  • RIS formulations vary from experimental causal estimates to LLM-based inflation indicators and subreddit activity benchmarks, each targeting a specific unit of analysis.
  • The diverse methods offer actionable insights for adjusting ranking biases, nowcasting macroeconomic trends, and understanding community dynamics in online platforms.

Searching arXiv for the cited RIS-related Reddit papers and adjacent formulations to ground the article. arxiv_search.query({"search_query":"all:\"Reddit Inflation Scores\" OR ti:\"Random Voting Effects in Social-Digital Spaces\" OR ti:\"Nowcasting the euro area with social media data\" OR ti:\"Super-Linear Growth and Rising Inequality in Online Social Communities\" OR ti:\"GuessTheKarma\" OR ti:\"Learning Inflation Narratives from Reddit\" OR ti:\"Consumers and Curators\"","max_results":10,"sort_by":"submittedDate","sort_order":"descending"}) arxiv_search.query returned 6 relevant papers matching the supplied ids and titles, including (Saito et al., 23 Mar 2026, Boss et al., 12 Jun 2025, Machado et al., 4 Mar 2025, Glenski et al., 2018, Glenski et al., 2017), and (Glenski et al., 2015). Reddit Inflation Scores (RIS) denotes a family of Reddit-derived measures rather than a single standardized statistic. In the literature, the term is used or operationalized in at least four distinct ways: as a causal estimate of vote-induced score inflation on posts, as a daily or monthly text-derived signal of inflation expectations or narratives, as a subreddit-level measure of activity inflation relative to super-linear scaling, and as a calibration error or adjusted score intended to correct popularity metrics for ranking, exposure, and browsing biases (Glenski et al., 2015, Boss et al., 12 Jun 2025, Machado et al., 4 Mar 2025, Glenski et al., 2018, Saito et al., 23 Mar 2026, Glenski et al., 2017).

1. Terminological scope and principal variants

Across the cited work, RIS is paper-specific and tied to the object being measured. In some cases the acronym is explicit; in others it is an operational mapping introduced to summarize a paper’s core construct. The most important distinction is between score inflation on Reddit itself and inflation measurement extracted from Reddit text for macroeconomic analysis (Boss et al., 12 Jun 2025, Saito et al., 23 Mar 2026).

Paper RIS object Core formulation
(Glenski et al., 2015) Post-level causal score inflation/deflation Percent change in final score from a random +1+1 or −1-1 treatment
(Boss et al., 12 Jun 2025) Daily euro-area Reddit inflation indicator LLM labels plus social interaction, then aggregation, normalization, and smoothing
(Saito et al., 23 Mar 2026) Monthly U.S. Reddit inflation score CPI-linked post classification aggregated by month and smoothed
(Machado et al., 4 Mar 2025) Subreddit activity inflation Deviation from the super-linear baseline C=N+kNβC = N + kN^\beta
(Glenski et al., 2018) Pairwise popularity miscalibration Difference between Reddit-implied preference and independent preference
(Glenski et al., 2017) Adjusted quality-aligned score Raw score discounted for title-only voting, fatigue, and exposure bias

Two interpretive cautions follow directly from this diversity. First, RIS is not comparable across studies unless the unit of analysis is fixed: post, pair, day, month, or subreddit. Second, the causal status differs sharply across formulations. The field experiment on random voting identifies downstream voting effects through randomization, whereas the macroeconomic and calibration-oriented variants are measurement frameworks built from text classification, aggregation, or behavioral adjustment rather than randomized intervention (Glenski et al., 2015, Boss et al., 12 Jun 2025).

2. Post-level RIS as causal vote-induced score inflation

In the field experiment on Reddit post submissions, RIS is grounded in an in vivo randomized design run continuously from September 1, 2013 to January 31, 2014 on N=93,019N = 93{,}019 posts. Every 2 minutes, an automated program identified the most recent post and randomly assigned it with equal probability to one of three arms: up-treated, down-treated, or control. The treatment consisted of a single injected upvote or downvote, applied after a random delay of $0$, $0.5$, $1$, $5$, $10$, $30$, or −1-10 minutes, each chosen with equal likelihood. Treated posts were re-sampled 4 days later to obtain final score, and the injected treatment vote was removed before analysis so that estimated effects reflected downstream organic voting changes rather than the mechanically added vote (Glenski et al., 2015).

The RIS-compatible post-level definitions are

−1-11

and

−1-12

The reported values are −1-13 and −1-14. Positive treatment also increased the probability of reaching a final score of at least −1-15 by −1-16 relative to control and at least −1-17 by −1-18. The corresponding high-score RIS is expressed as a probability lift,

−1-19

with reported lifts of C=N+kNβC = N + kN^\beta0 at C=N+kNβC = N + kN^\beta1 and C=N+kNβC = N + kN^\beta2 at C=N+kNβC = N + kN^\beta3 (Glenski et al., 2015).

The score distribution was extremely heavy-tailed and positively skewed: full-distribution skewness was C=N+kNβC = N + kN^\beta4 and kurtosis was C=N+kNβC = N + kN^\beta5; after removing the top C=N+kNβC = N + kN^\beta6 of scores, skewness remained C=N+kNβC = N + kN^\beta7 and kurtosis C=N+kNβC = N + kN^\beta8. Most posts had low final scores, with median approximately C=N+kNβC = N + kN^\beta9 or less, and down-treatment lowered the median from N=93,019N = 93{,}0190 to N=93,019N = 93{,}0191. For global distributional comparisons, the Kolmogorov-Smirnov statistics were N=93,019N = 93{,}0192 for up-treated versus control and N=93,019N = 93{,}0193 for control versus down-treated, with N=93,019N = 93{,}0194 in both cases. Student’s N=93,019N = 93{,}0195-tests on N=93,019N = 93{,}0196, excluding scores N=93,019N = 93{,}0197, gave N=93,019N = 93{,}0198 for up-treated N=93,019N = 93{,}0199 control and $0$0 for down-treated $0$1 control (Glenski et al., 2015).

Mechanistically, the paper attributes these effects to herding, visibility feedback, and ranking dynamics. Reddit ranked posts dynamically using score-based ranking with time dynamics affecting frontpage visibility, although the exact ranking formula was not provided. Early votes could improve visibility, which increased the chance of being seen and voted on, producing path dependence. A notable asymmetry qualifies the interpretation: at the upper tail, negative treatments did not reduce the probability of reaching very high scores, even though the average negative treatment effect remained statistically significant. The paper explicitly notes that this contrasts with Muchnik et al. (2013), who found little effect for negative treatments on comments in a different platform (Glenski et al., 2015).

3. Daily RIS as an LLM-derived euro-area inflation indicator

In the euro-area nowcasting framework, RIS refers operationally to the paper’s daily Reddit inflation signal that incorporates social interaction. The paper does not explicitly define a “Reddit Inflation Score” or the acronym RIS; instead, it constructs daily indicators $0$2 and their social-interaction-enhanced counterparts $0$3. For operational purposes, $0$4 is mapped to the inflation signal $0$5 after aggregation, normalization, and smoothing (Boss et al., 12 Jun 2025).

The data source is Reddit’s r/europe subreddit over 2012m1–2023m12. The corpus contains $0$6 submissions and $0$7 total comments for r/europe as a whole. The inflation-related subset contains $0$8 submissions, $0$9 first-level comments, and $0.5$0 keyword-filtered comments. Submissions are filtered by inflation-related keywords including “inflation, deflation, hyperinflation, price.” For each submission or comment, the net score, defined as upvotes minus downvotes, is recorded and can be used as a weight (Boss et al., 12 Jun 2025).

The classification model is LLaMa-3-70b-instruct, run on European Commission local servers. Each submission or comment is classified as UP, DOWN, or NEUTRAL with respect to the future direction of the “inflation rate” in Europe. The paper reports no fine-tuning. Against human-labeled subsets comprising $0.5$1 of submissions, or $0.5$2 inflation cases, the LLM achieves median $0.5$3 for inflation, compared with dictionary baselines of $0.5$4 and $0.5$5. Temperature was varied in $0.5$6, and the chosen temperature is $0.5$7 (Boss et al., 12 Jun 2025).

At the submission level, the raw label is $0.5$8 and comment labels are $0.5$9. The equal-weight interaction score is

$1$0

Threshold-based reclassification is then

$1$1

The paper also evaluates an optional vote-weighted variant,

$1$2

Daily RIS is then aggregated as

$1$3

with $1$4 if $1$5. Standardization and smoothing are

$1$6

For headline inflation, the best-performing specification uses first-level comments only, no upvote weighting, $1$7, and a moving-average window of $1$8 days (Boss et al., 12 Jun 2025).

The daily Reddit indicator enters a MIDAS-AR nowcasting model for euro-area HICP year-over-year inflation and unemployment. The monthly target $1$9 is linked to daily predictors through

$5$0

where the weights are a normalized exponential of a second-degree Almon polynomial. In recursive out-of-sample evaluation from 2018m1 to 2023m12, the best Reddit inflation signal yields RMSFE $5$1, MAFE $5$2, and CRPS $5$3 for HICP headline relative to an AR(1) baseline; for HICP food, the best Reddit specification yields RMSFE $5$4, MAFE $5$5, and CRPS $5$6. The paper reports that social interaction matters: reclassification after comment voting occurs in $5$7 of inflation submissions, or $5$8, and upward revisions outnumber downward revisions by approximately $5$9 (Boss et al., 12 Jun 2025).

Three nuances are central. First, the best inflation model does not use vote weights, even though weighting schemes were tested. Second, smoothing is critical: for headline inflation, windows around $10$0–$10$1 days are optimal, whereas energy prefers shorter windows and core, services, and food benefit from longer windows up to $10$2 days. Third, information available only up to $10$3–$10$4 days before month-end still yields similar improvements to full-month availability, which the paper interprets as robust real-time usefulness (Boss et al., 12 Jun 2025).

4. Monthly RIS as CPI-linked inflation narratives in U.S. Reddit data

A distinct use of RIS appears in the U.S. narrative-measurement framework, where monthly Reddit inflation scores are constructed from posts and comments discussing prices of goods and services aligned to U.S. CPI components. The corpus is drawn from The-Eye archives, with observational analysis from January 2012 to December 2022 and training, validation, and test data from January 2010 to December 2011. CPI-aligned subreddits include r/food, r/Frugal, r/cars, r/travel, and r/RealEstate, selected when they had several hundred monthly keyword hits and clear thematic correspondence to CPI components (Saito et al., 23 Mar 2026).

The keyword filter uses “price,” “cost,” “inflation,” “deflation,” “expensive,” “cheap,” “purchase,” and “sale,” with U.S.-specific filters for r/travel. To reduce subreddit-specific volume bias, the preprocessing pipeline applies sampling caps of $10$5 submissions and $10$6 comments per month and subreddit. Exact or near-duplicate posts are removed. Posts are predominantly English, and for lexical analysis URLs are removed and scraping terms are appended to stopwords (Saito et al., 23 Mar 2026).

The annotation scheme is tri-polar: deflation, neither, or inflation. The labeled dataset contains $10$7 posts after de-duplication, with majority vote among three annotators; full agreement is $10$8, two-out-of-three agreement is $10$9, and the remaining $30$0 are set to “neither” due to lack of majority. Krippendorff’s $30$1 is $30$2 overall. Of these instances, $30$3 are used for model development and $30$4 are reserved as held-out test data (Saito et al., 23 Mar 2026).

The selected inference model is Gemini 2.0 Flash Lite after supervised fine-tuning. Test-set performance is $30$5 accuracy and $30$6 $30$7 for Gemini, $30$8 for Llama 3.2, $30$9 for Phi 2.7, and approximately −1-100 for DeBERTaV3-Large and RoBERTa-Large. The paper’s argument is that lightweight LLMs are data-efficient: high zero-shot baselines for API models and competitive supervised performance for small open-source models on a labeled dataset of approximately −1-101k examples (Saito et al., 23 Mar 2026).

At post level, if class probabilities are available, the expected inflation score is

−1-102

If hard labels are used, −1-103 with −1-104. The paper reports that the hard-label mapping is used without additional calibration. Aggregate monthly RIS is

−1-105

and the reported smoothing step is a 3-month moving average,

−1-106

The resulting series correlates strongly with realized and survey-based inflation indicators. For March 2012 to December 2022, the 3-month-smoothed RIS has Pearson correlation −1-107 and Spearman −1-108 with CPI, both with −1-109. Against University of Michigan Inflation Expectation (MICH), the correlations are Pearson −1-110 with −1-111 and Spearman −1-112 with −1-113. Granger causality tests with lags of −1-114, −1-115, and −1-116 months show significant RIS −1-117 CPI and RIS −1-118 MICH effects, while CPI −1-119 RIS and MICH −1-120 RIS are not significant. For example, RIS −1-121 CPI gives −1-122 with −1-123, −1-124 with −1-125, and −1-126 with −1-127 (Saito et al., 23 Mar 2026).

The narrative dimension is a defining feature of this RIS variant. Change-point detection with PELT identifies multiple robust structural breaks, with pronounced RIS increases during 2020–2022 across all communities. Bigram TF-IDF shift analysis shows sector-specific narrative changes: food moves from quality and preference terms toward affordability and price salience; cars shift from consumer utility to supply constraints and costs; housing moves from transaction language toward “home prices,” “rising rates,” and “mortgage rates”; travel shifts toward pandemic and cost constraints; and r/Frugal shifts toward substitution and coping strategies such as “dollar tree” and “month groceries” (Saito et al., 23 Mar 2026).

The paper also states an important limitation directly: RIS is not a measure of realized inflation levels and is unscaled to CPI units. This distinguishes it from the nowcasting-oriented euro-area series, which is evaluated in a direct predictive framework, even though both rely on Reddit text and LLM-based directional classification (Boss et al., 12 Jun 2025, Saito et al., 23 Mar 2026).

5. Subreddit activity inflation under super-linear growth and rising inequality

A different RIS formulation treats inflation as excess community activity relative to an empirically estimated scaling law. In this setting, subreddit size is the number of active users, denoted here by −1-128, and activity is the total number of comments −1-129 in a monthly snapshot. Excess comments are defined as −1-130, which captures activity beyond the baseline of one comment per active user. Using one-month snapshots over 2021, the paper fits the super-linear relation

−1-131

equivalently −1-132, with −1-133, −1-134, and −1-135 for subreddits with −1-136 (Machado et al., 4 Mar 2025).

The paper also reports heavy-tailed individual activity. The global distribution of comments-per-user in a subreddit follows a power law −1-137 with −1-138 and −1-139. As subreddits grow in size, the complementary cumulative distributions of comments-per-user widen with −1-140, indicating higher probabilities of very large user-level activity. Two null models—random sampling from global distributions and shuffling users across subreddits while preserving size constraints—collapse to the global curve, which the paper uses to argue that the observed broadening is not explained by finite size or sampling artifacts (Machado et al., 4 Mar 2025).

The core inequality measure is the Gini coefficient. For sample data −1-141 with mean −1-142, the computable form reported is

−1-143

The paper states that −1-144 increases monotonically with −1-145 and then slows or plateaus at very large sizes, indicating rising centralization of activity in larger subreddits (Machado et al., 4 Mar 2025).

RIS is then defined as activity inflation relative to the expected super-linear baseline. The paper gives

−1-146

with the interpretation that −1-147 means activity matches the expected baseline, −1-148 indicates inflated activity, and −1-149 indicates deflated activity. An excess-only version is

−1-150

The framework also introduces an inequality-adjusted score,

−1-151

or, alternatively, a linear adjustment around −1-152 (Machado et al., 4 Mar 2025).

This RIS variant is not about votes, ranking, or price expectations. It quantifies whether a subreddit is more active than expected given its size and the observed Reddit-wide super-linear scaling of excess comments. A plausible implication is that it is best interpreted as a size-normalized activity anomaly statistic rather than a signal of content quality or economic inflation.

6. RIS as popularity miscalibration and as a browsing-bias correction

Two additional strands treat RIS as a correction to the meaning of Reddit popularity. In the GuessTheKarma framework, the issue is miscalibration between Reddit scores and independent preference judgments. In the browsing-logs framework, the issue is inflation arising from title-only voting, exposure bias, and cognitive fatigue (Glenski et al., 2018, Glenski et al., 2017).

GuessTheKarma collected independent judgments for −1-153 image pairs drawn from eight image-focused subreddits, using −1-154 players and −1-155 total preferences. The majority preference across players served as a path-independent proxy for true population preference. Overall, the higher-scored Reddit item matched the majority preference only −1-156 of the time, while Imgur view counts achieved −1-157. Accuracy improved sharply only in extreme score-imbalance conditions: for Very High–Low pairs, Reddit reached −1-158 accuracy, and for High–Low pairs, both Reddit and Imgur reached −1-159. The paper reports the percentile-difference calibration

−1-160

with −1-161 and −1-162 for Reddit. On this basis, the RIS definition proposed in the synthesis is pair-level inflation,

−1-163

where −1-164 is the fraction of independent raters preferring −1-165 to −1-166. Positive RIS means Reddit scores overstate preference for −1-167; negative RIS means they understate it. The paper further reports that predictive accuracy declines with subreddit subscriber count, with −1-168 and −1-169 for Reddit, which it interprets as evidence that feedback loops and crowding weaken score-preference alignment in large communities (Glenski et al., 2018).

The browsing-logs study focuses on how users actually vote. It instruments a cohort of −1-170 consenting Reddit users over approximately one year and finds that −1-171 of posts were rated without first viewing the content. Sessions are defined with a −1-172-minute inactivity threshold. For link posts, content viewing requires clicking through to the external URL; for self posts, it requires expanding the text body or visiting the comments or permalink page. The paper reports position and ranking bias in voting likelihood, and it shows evidence of cognitive fatigue in the browsing sessions of users most likely to vote (Glenski et al., 2017).

On top of these findings, the RIS synthesis defines a family of adjusted scores. With raw score

−1-173

and estimated fraction of votes with prior content view

−1-174

the simplest adjustment is

−1-175

A fatigue-adjusted version uses the predicted probability −1-176 of voting without a content view,

−1-177

and an exposure-adjusted version uses inverse propensity weighting,

−1-178

where −1-179 is the estimated exposure propensity for vote −1-180 and −1-181 denotes upvote or downvote. The paper summary is explicit that these formulas are proposed extensions grounded in the observed behaviors rather than formulas reported directly in the original paper (Glenski et al., 2017).

Taken together, these two RIS formulations address a common misconception: a Reddit score is not automatically an unbiased proxy for content quality or independent preference. GuessTheKarma shows that score differences are informative mainly when they are very large, while browsing telemetry shows that a substantial share of votes do not follow content inspection at all (Glenski et al., 2018, Glenski et al., 2017).

7. Comparability, limitations, and methodological implications

RIS formulations differ along four axes: target construct, unit of analysis, causal status, and aggregation rule. The randomized-vote RIS targets downstream score amplification on individual posts; the euro-area and U.S. RIS variants target inflation expectations or narratives in daily or monthly text streams; the subreddit-scaling RIS targets excess activity relative to −1-182; and the preference-calibration or browsing-adjusted RIS targets deviations between popularity and quality-aligned judgment (Glenski et al., 2015, Boss et al., 12 Jun 2025, Machado et al., 4 Mar 2025, Glenski et al., 2018, Saito et al., 23 Mar 2026, Glenski et al., 2017).

Several paper-specific limitations matter for interpretation. The post-voting experiment measures posts, not comments, and Reddit’s exact ranking formula is not provided (Glenski et al., 2015). The euro-area nowcasting study is confined to r/europe, uses English keyword filters in a multilingual community, and does not explicitly apply spam or bot filtering (Boss et al., 12 Jun 2025). The U.S. narrative study is limited to CPI-aligned subreddits, public English-language posts, and moderate annotation reliability with −1-183 (Saito et al., 23 Mar 2026). The super-linear activity study covers comments in one-month subreddit snapshots and does not include subscribers, lurkers, posts, or view dynamics (Machado et al., 4 Mar 2025). GuessTheKarma is limited to image posts in eight subreddits, with posts from 2008–2015 judged in 2017 (Glenski et al., 2018). The browsing-log study is desktop-based and therefore does not resolve mobile inline-preview behavior (Glenski et al., 2017).

Despite these differences, recurring methodological themes are visible. Heavy tails recur in post scores and comment activity; visibility and ranking feedback recur in both voting experiments and popularity-calibration work; and social interaction can be either a confounder to be corrected or a signal to be exploited, depending on whether the objective is de-biasing Reddit scores or extracting macroeconomic information from Reddit discourse (Glenski et al., 2015, Boss et al., 12 Jun 2025, Machado et al., 4 Mar 2025, Glenski et al., 2018, Glenski et al., 2017).

A plausible implication is that RIS should be treated as a paper-specific operationalization rather than as a universal Reddit metric. For causal inference on platform dynamics, the post-level experimental RIS is the most directly identified. For macroeconomic nowcasting, the daily and monthly LLM-based RIS series are more relevant. For community science, the scaling-based RIS captures anomalous activity concentration. For platform evaluation and recommender calibration, the pairwise and browsing-adjusted RIS variants quantify how far visible popularity departs from independent preference or content-inspection-based judgment.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Reddit Inflation Scores (RIS).