---
title: Mutual Impact Analysis of OSS (MIAO)
url: https://www.emergentmind.com/topics/mutual-impact-analysis-of-oss-miao
type: topic
---

# Mutual Impact Analysis of OSS (MIAO)

Searching arXiv for the referenced papers and closely related work to ground the article.
arxiv_search(query="2602.17131 OR \"Mutual Impact Analysis of OSS\" OR \"Quantifying Competitive Relationships Among Open-Source Software Projects\"", max_results=10)
arxiv_search(query="1504.04971 OR \"Impact assessment for vulnerabilities in open-source software libraries\"", max_results=10)
arxiv_search(query="2411.06027 OR \"A Toolkit for Measuring the Impacts of Public Funding on Open Source Software Development\"", max_results=10)
arxiv_search(query="2304.02490 OR \"Opening the random forest black box by the analysis of the mutual impact of features\"", max_results=10)
Mutual Impact Analysis of OSS (MIAO) denotes an automated method for quantifying competitive relationships among open-source software projects through time-series econometrics, specifically structural vector autoregression (SVAR) and impulse response functions (IRFs). In this usage, a project’s trajectory is not treated as a purely internal phenomenon; rather, activity in one project may propagate to another over time, and asymmetric influence can be associated with competition-driven cessation. In a broader sense, the mutual-impact perspective in OSS research also appears in work on vulnerability assessment and on public funding evaluation, where impact is likewise treated as contextual, directional, and ecosystem-dependent rather than reducible to dependency presence or project-internal output alone [2602.17131] [1504.04971] [2411.06027].

## 1. Conceptual basis and motivation

MIAO was proposed to address a gap in OSS survivability research: prior studies were described as focusing mainly on internal factors such as developer activity, community structure, and code quality, while largely ignoring external competitive pressure [2602.17131]. The motivating claim is that OSS history contains recurrent cases in which one project overtakes another and the weaker project fades away. The paper uses Chainer versus PyTorch as its canonical example, noting that Chainer’s activity drops around 2019 while PyTorch continues to grow, and that Preferred Networks stated it would stop new Chainer development and move to PyTorch. The authors term this competition-driven cessation a Rising Event (REV) [2602.17131].

The core intuition is temporal and directional. If two OSS projects compete, then a shock to one project’s activity should affect the other over subsequent periods. Simple static comparisons or correlation coefficients do not capture that directionality: they do not answer whether a competitor suppresses the target, whether the target still affects the competitor, or how that relationship changes across stages of development. MIAO therefore treats OSS projects analogously to interacting economic variables, each with its own activity time series, possible contemporaneous and lagged effects, and responses to shocks that may be delayed and persistent [2602.17131].

This framing makes “impact” narrower than generic ecosystem influence but more operational than descriptive accounts of rivalry. The paper’s emphasis is not merely that projects coexist in the same domain, but that their observed activity histories can be used to estimate directional competitive influence. This suggests a specific interpretation of mutual impact in OSS: not co-presence or similarity, but measurable propagation of activity shocks across projects over time.

## 2. Formal model and score construction

The formal core of MIAO is an SVAR model:
$$
B_0 y_t = c + B_1 y_{t-1} + B_2 y_{t-2} + \ldots + B_p y_{t-p} + u_t
$$
where \(y_t\) is the vector of observed variables at time \(t\), here the activity time series of multiple OSS projects; \(B_0\) captures instantaneous or contemporaneous structure among variables; \(c\) is a constant vector; \(B_1,\dots,B_p\) are lag coefficient matrices; \(u_t\) is the structural shock vector, assumed to be white noise; and \(p\) is the lag order [2602.17131]. In the OSS setting, each element of \(y_t\) can be a project’s commit activity at time \(t\). Current activity is thus modeled as depending on contemporaneous mutual effects together with past activity.

The dynamic effect of a shock is quantified through the impulse response function:
$$
\text{IRF}_{ij}(k) = [\Psi_k B_0^{-1}]_{ij}
$$
with \(i\) the receiving project, \(j\) the shocked project, and \(k\) the number of periods after the shock. \(\text{IRF}_{ij}(k)\) expresses how much project \(i\) changes after \(k\) periods when project \(j\) receives a one-unit shock at time \(t\) [2602.17131].

MIAO then compresses the IRF profile into a scalar directional score. It first computes the Shock Cumulative Effect:
$$
\mathrm{SCE}_{ij} = \sum_{k=1}^\infty \mathrm{IRF}_{ij}(k)
$$
and then defines the MIAO score:
$$
\mathrm{MS}_{ij} = (-1) \cdot \sum_{k=1}^m \mathrm{SCE}_{ji}^k
$$
The paper specifies two important implementation details. First, the sign is flipped so that larger negative impact corresponds to a larger positive score. Second, the indices are reversed so that the score is intuitively read as influence from \(i\) to \(j\). When \(\mathrm{MS}_{ij}\) is near 0, positive and negative effects largely cancel [2602.17131].

The resulting pairwise relations are categorized by sign pattern:

| Pattern | Condition | Interpretation |
|---|---|---|
| M1 | \(\mathrm{MS}_{ij} > 0 \land \mathrm{MS}_{ji} \le 0\) | \(A_i\) negatively impacts \(A_j\) |
| M2 | \(\mathrm{MS}_{ij} \le 0 \land \mathrm{MS}_{ji} > 0\) | \(A_j\) negatively impacts \(A_i\) |
| M3 | \(\mathrm{MS}_{ij} \ge 0 \land \mathrm{MS}_{ji} \ge 0\) | mutual negative impact |
| M4 | \(\mathrm{MS}_{ij} \le 0 \land \mathrm{MS}_{ji} \le 0\) | mutual positive impact |

To quantify asymmetry, the paper uses the distance
$$
\mathrm{D}(a,b)=|a-b|
$$
A large gap between \(\mathrm{MS}_{ij}\) and \(\mathrm{MS}_{ji}\) is interpreted as evidence of unidirectional influence, which the study treats as a key clue for competition and possible decline [2602.17131].

## 3. Dataset construction and analytical workflow

The empirical study mined GitHub and constructed 187 OSS project groups. It began from the August 2024 GitHub Search dataset and retained projects with at least 2,538 stars, at least 23 contributors, and at least 500 commits. These thresholds were chosen to focus on relatively mature, active projects, making internal decline less likely and external competition more plausible. Bot-generated commits were excluded by filtering out authors containing “[bot]” [2602.17131].

A project was considered a REV candidate if it had ceased activity under at least one of three criteria: the repository was archived, there had been no commits for over a year, or the average commits over the last 12 months were less than 1.5. To determine whether cessation was due to competition, README files and related documentation were inspected for evidence that the project recommended migration to another OSS project, such as a deprecation notice in favor of another package. This yielded 234 REV candidates initially [2602.17131].

To make the SVAR models comparable and stable, the study fixed the dimensionality to a 3-variable model consisting of one target project and two competitors. If a project had two competitors, they were used directly; if it had three or more, two were chosen with the largest temporal overlap; if it had only one, LLM-assisted competitor discovery was used and one additional plausible competitor in the same domain was manually validated. Exclusions covered ambiguous competitors, direct repository transfers, official or native replacements, overlap shorter than one year, data inconsistencies, and fork-based competition. Forks were excluded because shared pre-fork histories create perfect multicollinearity, making VAR coefficients unidentified. After filtering, the corpus consisted of 87 REV groups and 100 non-REV groups [2602.17131].

The analysis was carried out over multiple windows rather than with a single global model. For each group, the start time was the earliest date common to all series. For REV cases, the end date was the first date when the 12-month moving average commits dropped below 1.5; for non-REV cases, the end date was the last common date. Nested time windows \(T_1, T_2, \dots, T_m\) were then created by expanding the window by one year at a time, up to a maximum of four years, because longer windows often induced trends and serial correlation. To increase robustness, the window was also shifted by 0, 1, 2, and 3 months and the resulting scores were averaged; the paper likens this to bagging in ensemble learning [2602.17131].

The preprocessing and model-selection pipeline is explicitly statistical. Stationarity was checked using the Augmented Dickey–Fuller (ADF) test; if the p-value exceeded 0.05, the series was treated as non-stationary and transformed using fractional differencing rather than ordinary integer differencing. Lag order was selected from candidate VAR models using information criteria such as AIC, BIC, or HQIC, together with a whiteness requirement on residuals. Residual autocorrelation was tested with the Ljung–Box test, whose null hypothesis is that residuals are white noise [2602.17131].

Recursive SVAR identification depends on variable order, and the study fixed the ordering differently by class. For the REV class, the order was Competitor 1 \(\rightarrow\) Competitor 2 \(\rightarrow\) Target, allowing the target to receive contemporaneous influence from both competitors. For the non-REV class, all \(3! = 6\) permutations were tested because the true ordering was unknown [2602.17131].

## 4. Empirical results and interpretation

The paper reports that over 90% of series were already stationary by ADF, while the remainder were made stationary by fractional differencing, typically with very small differencing orders. Selected VAR lags averaged about 10.93, and residuals generally passed the Ljung–Box test, which the authors interpret as support for the white-noise assumption and for the statistical workability of the SVAR or VAR machinery on these OSS data [2602.17131].

The central evaluation had two tasks. In EVAL1, the full available history was used to classify whether a target project was REV or non-REV. Across permutations, the model achieved up to 0.81 accuracy, with average accuracy about 0.75. The paper notes that 0.81 accuracy means that in the best setup about 81% of all 187 groups were correctly classified as REV or non-REV; it does not mean that 81% of all REV projects were found, since that would be recall. The paper also reports class-wise precision, recall, and F1-score; for the best retrospective setting, REV F1 was around 0.79 and non-REV F1 around 0.83 [2602.17131].

In EVAL2, the last year of data was masked and the task was to predict cessation from a point one year earlier. The best accuracy was up to 0.77, with average around 0.73. In the best one-year-ahead setting, REV F1 was around 0.72, though with greater variability than in the retrospective task. The authors interpret this as evidence that predicting decline earlier is harder than explaining it retrospectively, but still sufficiently informative for early warning [2602.17131].

The decision-tree analysis suggests that the most informative signals are unidirectional influence patterns. In the retrospective setting, classification was driven mainly by the target’s influence on a competitor, especially \(t \to c2\), which the paper describes as somewhat counterintuitive. In the one-year-ahead setting, the more intuitive competitor-to-target pattern appeared, especially \(c2 \to t\). The paper interprets this sequence as an evolving decline process: warning signs appear one year earlier as the competitor starts influencing the target, and later the target’s own influence on the competitor becomes more pronounced or the balance shifts. The key signal is therefore not merely that one project is larger, but asymmetric influence over time [2602.17131].

The practical implications follow directly from that interpretation. For OSS maintainers, MIAO is presented as an early-warning dashboard for monitoring increasing negative influence from competitors, detecting asymmetry in influence flows, anticipating loss of the development race, and informing roadmap changes, partnerships, or repositioning. For organizations choosing OSS, it is described as a risk assessment tool for identifying projects likely to stagnate and preferring more resilient ecosystems. For researchers, the methodological contribution lies in shifting survivability analysis from internal characteristics to ecosystem-level dynamic competition [2602.17131].

## 5. Relation to other OSS impact-assessment frameworks

Although the acronym MIAO is attached specifically to competitive analysis in [2602.17131], the broader mutual-impact logic has precedents in OSS research. A notable example is the vulnerability-impact work of Ponta, Plate, and Sabetta, which does not study project-versus-project competition but instead asks whether the vulnerable part of a bundled OSS library is actually exercised in the application’s real usage context, and therefore whether urgent patching is warranted [1504.04971]. The key assumption is stated as \(A1\): whenever an application that includes a vulnerable library executes a fragment of the library that would be updated in a security patch, there exists a significant risk that the vulnerability can be exploited.

That framework defines a change-list \(C_{ij}\) as the set of programming constructs in OSS component \(i\) modified, added, or deleted by the security patch for vulnerability \(j\); a trace-list \(T_a\) as the set of programming constructs executed at least once during runtime of application \(a\), including constructs in bundled libraries; and an application construct set \(S_a\) as the set of all programming constructs belonging to the application itself. The core decision signal is the intersection \(C_{ij} \cap T_a\). A non-empty intersection is taken as evidence that code touched by the security patch was executed in the application context, while \(T_a \cap S_a\) serves as a coverage indicator for how much of the application’s own code was observed at runtime [1504.04971]. This is a different operationalization of mutual impact: the effect of an OSS vulnerability depends on actual application usage, not merely on the presence of the dependency.

The architecture in that work consists of a Patch Analyzer, a Runtime Tracer, and a Source Code Analyzer feeding an Assessment Engine. The Java proof of concept integrates with Maven and uses SAP HANA with a web frontend for the Assessment Engine, JGit, SVNKit, and ANTLR for the Patch Analyzer, Javassist for instrumentation in the Runtime Tracer, and ANTLR in a Maven plugin for the Source Code Analyzer. Dynamic instrumentation is described as suitable for unit tests and CI, while static instrumentation is described as suitable for integration and end-user tests, particularly where startup overhead matters, such as in Tomcat containers or PaaS environments. The case study on CVE-2014-0050 in Apache Commons FileUpload concludes that the vulnerability is exploitable in the evaluated application context because a known vulnerable version was in use, execution of patch-modified code was observed, and the result was manually verified via an exploit from Exploit-DB [1504.04971].

A second complementary line of work concerns the impacts of public funding on OSS development. The 2024 toolkit paper is not about competition-driven cessation, but it systematizes OSS impact as positive or negative, direct or indirect, internal or external, and short-, medium-, or long-term. It further argues that measurement design must account for funding objectives, project life stage and social structure, and regional and organizational cost factors; and it reviews qualitative, quantitative, and mixed-methods approaches, alongside caveats about causal attribution, metric gaming, and oversimplification [2411.06027]. A plausible implication is that MIAO’s competitive score can be interpreted as one technological and ecosystem-level indicator within a larger socio-technical measurement framework rather than as a complete account of OSS impact by itself.

## 6. Scope, limitations, and adjacent usages

The 2026 MIAO paper is explicit that causality is not fully proven. It identifies statistical competitive influence, but not definitive proof that competition alone caused cessation. The labeling of REV cases is also noisy because the evidence comes from README migration hints and related materials, so some projects may be false positives or may have mixed causes of decline. Commit count is only one activity proxy and may not capture all dimensions of project health; the paper suggests that issue activity, releases, or dependency updates may improve the model. Further limitations include the exclusion of fork-based competition, the fixed 3-variable model, and limited external validity because the dataset covers 187 GitHub projects and may not generalize to other platforms or to commercial software [2602.17131].

These caveats are substantive rather than peripheral. They delimit what MIAO currently measures: dynamic competitive influence among independent OSS projects under a particular empirical design. It is not a complete theory of OSS decline, and it does not establish that every detected negative influence is causal in a strong sense. This suggests that MIAO is best understood as a quantitative signal extraction framework for competitive ecosystems, not as a standalone explanation of software evolution.

The term “mutual impact” also has an adjacent usage outside OSS ecosystems. In random-forest research, Hornung and Wright propose Mutual Forest Impact (MFI) as a relation measure and Mutual Impurity Reduction (MIR) as an importance measure for analyzing pairwise feature relations and joint impact on the outcome. MFI is defined through surrogate split agreement corrected by permutation, while MIR combines that relation parameter with actual impurity reduction to rank features in high-dimensional settings [2304.02490]. That work is terminologically related but substantively distinct: its objects are features within a predictive model, not OSS projects within a software ecosystem.

Taken together, these strands indicate that “mutual impact” has become a recurring analytical motif across different research problems. In OSS proper, however, MIAO most specifically refers to the SVAR- and IRF-based quantification of competitive relationships among projects, with empirical evidence that asymmetric influence can identify competition-driven cessation retrospectively and one year ahead [2602.17131].

Source: https://www.emergentmind.com/topics/mutual-impact-analysis-of-oss-miao