---
title: 'DataSTORM: Dual-System Research Framework'
url: https://www.emergentmind.com/topics/datastorm
type: topic
---

# DataSTORM: Dual-System Research Framework

DataSTORM encompasses two distinct systems at the intersection of reasoning and exploration over large, multifaceted data assets: (1) an LLM-agentic framework for deep thesis-driven analysis over structured databases and web corpora [2604.06474]; and (2) a platform for simulation ensemble management and alternative timeline exploration in complex dynamical domains [2407.14571]. Both employ advanced methodologies to drive insight discovery, decision support, and narrative synthesis in the presence of massive, heterogeneous information sources.

## 1. LLM-Agentic Deep Research: System Overview

DataSTORM [2604.06474] is an LLM-based agentic system structured as a three-stage pipeline targeting both structured databases and internet sources, culminating in the automated generation of evidence-grounded analytical narratives. The architecture tightly integrates multi-agent exploration, exploratory data analysis (EDA), and iterative thesis refinement, reframing deep research as a thesis-driven process.

The pipeline operates as follows:

- **Warm-Start Module (Co-STORM):** Ingests an internet corpus $\mathcal I$ and user query $q$, initiating a lightweight report $r_0$ and seed insight bank $B_0$ through multi-agent discourse over web evidence.
- **Multi-Agent Exploration Module:** Consists of a Planner agent that generates exploration questions tagged by destination (database or internet); an Executor agent which handles SQL synthesis and execution (implementing a ReAct-style reasoning loop and outputting both SQL $s_{i,j}$ and answer $a_{i,j}$); a Query Consistency Module ensuring predicate alignment and context-aware rectification; Inductive Surfacing via automatic computation of summary statistics after SQL execution; and a Thesis Generation & Refinement agent, invoked every $p$ layers.
- **Final Report Generation:** Comprised of staged narrative assembly with outline generation (Stage A), evidence-grounded section drafting (Stage B), citation grounding and revision (Stages C & D), and linguistic polishing (Stage E).

A unified insight bank $B_i$ maintains evidence from both database (quantitative) and web (qualitative) sources, annotated with provenance and merged into the evolving narrative arc via the system’s dynamic thesis $t_i$.

## 2. Exploratory Data Analysis and Data Storytelling Integration

The system encodes EDA principles both in its control flow and content surfacing:

- The Planner acts deductively by posing targeted “what-if” or “why” questions; the Executor acts inductively by surfacing bottom-up patterns and summary statistics after each SQL execution (including $\text{distinct\_pct}(C)$, $\text{top5}(C)$, and for numeric columns, $\min(C), \max(C), \text{median}(C), \text{mean}(C)$).
- Query Consistency ensures alignment of filter predicates across questions and enforces normalization by issuing follow-ups in the face of detected semantic drift.
- The system's Data Storytelling module generates candidate theses $T$ from $B_p$ (the insight bank at layer $p$), attaches research strategies for each, and invokes periodic refinement to maintain a coherent “through-line.”
- The report planning module (OutlineGen) programmatically links thesis, evidence, and narrative section structure.

## 3. Formal Thesis-Driven Analytical Process

The thesis-driven analytical workflow consists of:

- **Candidate Thesis Discovery:** Given $B_p$, the thesis generation LLM $f_1$ generates up to three candidate theses $T = \{t_1, ..., t_n\}$. Each $t_k$ is scored for coherence with $B_p$ using a function $\sigma$, yielding $t_p = \arg\max_{t\in T} \sigma(t,B_p)$.
- **Iterative Validation via Cross-Source Hypotheses:** For each $t_i$, corresponding evidence sets $E_i^{\mathcal D}$ (quantitative/database) and $E_i^{\mathcal I}$ (qualitative/web) are extracted, and a consistency score $v(t_i) = \mathrm{Sim}(\mathrm{Embed}(E_i^{\mathcal D}), \mathrm{Embed}(E_i^{\mathcal I}))$ is computed via embedding similarity.
- **Convergence Criteria:** Stopping is based on narrative and evidence bank stability: if the symmetric difference $\Delta B_i = |B_i \Delta B_{i-1}| \leq \epsilon_B$ and $1 - \mathrm{Jaccard}(\mathrm{kw}(t_i), \mathrm{kw}(t_{i-1})) \leq \epsilon_t$, or a fixed number of $m$ layers is reached.

This process ensures research output converges to a supported, focused analytical narrative.

## 4. Cross-Source Querying and Evidence Integration

In each layer, the Planner emits up to $n$ questions $q_{i,j} = (\text{text}, \text{destination})$, dispatched to either SQL execution or web search. For database destinations, the Executor uses a ReAct reasoning loop:

- Thought: Choose relevant tables/columns
- Actions: $\text{get\_tables()} \rightarrow \text{retrieve\_tables\_details()} \rightarrow \text{execute\_sql(sql)}$
- Loop continues until a stopping criterion is satisfied

All answers $a_{i,j}$, enriched with summary statistics and destination metadata, are merged into $B_i$. The InsightFilter retains the top $K$ insights per iteration by relevance to the current thesis ($t_{i-1}$), applying LLM-based re-ranking. This produces a parallel evidence structure where both qualitative and quantitative results are treated equivalently in downstream thesis and report generation.

## 5. Empirical Evaluation and Benchmarking

### 5.1 Quantitative Metrics on InsightBench

Performance is assessed using:

- **Insight-level recall:** $ \text{recall} = \frac{1}{|G|} \sum_{k=1}^{|G|} \max_{j} \text{match}(g_k, s_j) $, with match scored by LLM similarity.
- **Summary-level score:** $ \rho(\text{ref\_summary}, \text{sys\_summary}) $, judged by LLM.

Results:

| System         | Judge      | Insight Recall | Summary Score |
|----------------|------------|---------------|---------------|
| AgentPoirot    | Qwen-3-30B |      49.9%    |    51.5%      |
| DataSTORM      | Qwen-3-30B |  **69.3%**    | **58.7%**     |
| AgentPoirot    | GPT-4o     |      47.1%    |    46.6%      |
| DataSTORM      | GPT-4o     |  **61.9%**    | **52.5%**     |

DataSTORM demonstrated a +19.4% (Qwen) and +14.8% (GPT-4o) absolute improvement in insight-level recall and substantial gains in summary score [2604.06474].

### 5.2 ACLED Benchmark and Human Evaluation

Key metrics: reference-induced matching, RACE framework (Comprehensiveness, Depth, Instruction-following, Readability, 50 = parity), and database use ratio.

| System                 | Ref-Match | RACE   | DB Use Ratio |
|------------------------|-----------|--------|--------------|
| OpenAI DR (CSV/MCP)    | 51.2/48.5 | 46.8/46.1 | 23.3/30.4%   |
| DataSTORM              |   61.8    |  52.6  |   66.4%      |

Human expert assessment over 20 topics found DataSTORM outperforming the baseline on 6/7 rubric dimensions, with significant improvement in Originality (+0.84 pp, $p < 0.05$), and winning 57.5% of pairwise preferences versus the baseline.

*This suggests that the thesis-driven and EDA-anchored design enables DataSTORM to surface both more relevant and original insights, with a higher proportion of claims explicitly grounded in structured data compared to proprietary LLM research systems. A plausible implication is that this architecture can generalize to other settings requiring rigorous multi-source synthesis and reasoning.*

## 6. Algorithmic Structure and Reproducibility

The system’s operational core is summarized as follows:

```python
Algorithm DataSTORM(𝓓, 𝓘, q, m, p):
  (r₀, B₀) ← CoSTORM(𝓘, q)
  t₀ ← null
  for i in 1..m:
    Q_i ← Planner.generate(B_{i−1}, t_{i−1})
    new_insights ← ∅
    for each (q, dest) ∈ Q_i:
      if dest == "database":
        (s, a) ← Executor.run(q, 𝓓)
        s' ← QueryConsistency.standardize(s, B_{i−1})
        a' ← Executor.run(s', 𝓓).answer_with_stats
      else:
        a' ← WebSearcher.query(q, 𝓘)
      new_insights ← new_insights ∪ extract_insights(a')
    B_i ← InsightFilter.select_top_k(B_{i−1} ∪ new_insights, K)
    if i mod p == 0:
      t_i ← ThesisRefiner.refine(t_{i−1}, B_i)
    else:
      t_i ← t_{i−1}
    if Converged(B_i, B_{i−1}, t_i, t_{i−1}):
      break
  plan ← OutlineGenerator.plan(t_i, B_i)
  sections ← for each section_spec in plan:
                SectionWriter.draft(section_spec)
                SectionWriter.fact_check_and_revise()
  r ← Polisher.finalize(sections)
  return r
```

Key subroutines include LLM-driven question generation, ReAct SQL execution, LLM-based filter and thesis refinement, and convergence/stopping checks. This formal structure supports reproduction and extension in novel data-centric research contexts [2604.06474].

## 7. Relation to Simulation Ensembles: DataStorm-EM

DataStorm also labels a platform for managing and analyzing large ensembles of simulation instances—DataStorm-EM [2407.14571]—designed for high-complexity domains requiring the exploration of alternative system trajectories under uncertainty.

DataStorm-EM is architectured over four modules:

- **Ensemble Generation:** Directed acyclic model graphs (DS-Actors), sampling of parameter vectors, and distributed execution orchestration.
- **Parameter Management & Optimization:** Solves a 0–1 quadratic program to optimize the selection of simulation instances for diversity and informativeness under budget $B$.
- **Timeline Extraction & Analysis:** Hierarchical agglomerative clustering of timelines based on weighted time-series distances, selection of medoid representatives, and optional 2D projection for exploratory analysis.
- **Visualization & Exploration Tools:** Rich web-based interfaces (React + D3) for timeline plotting, cluster mapping, parallel coordinates of parameter vectors, interactive drill-down, and scenario summaries.

In practice, DataStorm-EM achieves efficient workflow orchestration, scalable clustering, and interactive exploration for ensembles up to 5,000 simulation instances on cloud/HPC clusters, with visualization overhead maintained below 5 minutes for $K\leq20$ timelines. Prototypes in pandemic response and urban sustainability further demonstrate its domain-general applicability. The system leverages a modern stack (Python, Ansible, Docker, Apache Kafka, Neo4j, Parquet, SciPy, scikit-learn) and supports extension via new model wrappers and analytic plug-ins [2407.14571].

Source: https://www.emergentmind.com/topics/datastorm