Experiversum: Reproducible Data-Driven Science
- Experiversum is a lakehouse-based ecosystem that integrates raw data, experimental processes, and team context using a structured metadata model.
- It employs ELT pipelines and a three-layered metadata schema to capture data acquisition, processing, and iterative experimental reviews.
- The framework supports transparent workflows and collaborative decision-making across diverse domains such as life, Earth, and political sciences.
Experiversum is a lakehouse-based ecosystem for curating, documenting, and reproducing data-driven experimental science, especially in exploratory and iterative settings. It links raw data, processing steps, and experimental context through a structured metadata model, ELT-style pipelines, and exploration interfaces, with the explicit aim of bridging exploratory and reproducible research while supporting transparent workflows, collaborative decision capture, and multi-perspective result interpretation (Vargas-Solar et al., 30 Sep 2025). A broader interpretation is suggested by adjacent work on living labs, active experimentation, and experiential learning: in that sense, Experiversum denotes an experimental environment in which interventions, observations, and learning are embedded in the operational substrate itself rather than isolated in offline evaluation (Schaer, 2017, Chandra et al., 28 Jun 2026, Yang et al., 27 Nov 2025).
1. Conceptual scope and research problem
Experiversum addresses a recurrent difficulty in empirical research: raw data may be available in abundance, yet the conditions that produced publishable results are often incompletely recorded. The ecosystem is designed around a definition of a data-driven experiment as comprising three elements: raw data from empirical sources, the research team responsible for data selection and validation, and contextual metadata describing collection, processing, and analysis conditions (Vargas-Solar et al., 30 Sep 2025). On this account, the central technical challenge is twofold: first, to design metadata models that represent not only data and workflows but also team and context; second, to implement ELT pipelines that support experiment curation while tracking how decisions affect outcomes, with metadata serving as an execution guide for reproducibility (Vargas-Solar et al., 30 Sep 2025).
This formulation places Experiversum at the intersection of data management, provenance engineering, and scientific workflow systems. Its distinctive emphasis is not merely storage, but the preservation of experimental structure across successive iterations. The system is intended for settings in which hypotheses, inclusion criteria, transformation logic, validation protocols, and interpretive judgments evolve during the research process rather than being fixed in advance (Vargas-Solar et al., 30 Sep 2025).
A broader conceptual lineage can be seen in earlier work on living labs. There, experimentation moves from in vitro evaluation toward in vivo settings, where real operational systems such as wikis, search engines, repositories, online shops, or social platforms are modified while ordinary users continue to use them (Schaer, 2017). This suggests a general family resemblance: Experiversum is not only a platform for preserving experiments after the fact, but also a way of treating experimentation as a first-class property of a live research or decision environment.
2. Architectural organization and metadata stratification
The core architecture is lakehouse-based. Like a data lake, it can ingest raw data in original formats with schema-on-read flexibility; like a data warehouse, it adds structured layers and metadata to enable querying, analytics, and governance (Vargas-Solar et al., 30 Sep 2025). In the prototype, the storage layer is backed by SQLite3 and is coupled to a metadata layer, orchestration logic, curation pipelines, and an exploration and analytics layer (Vargas-Solar et al., 30 Sep 2025).
The architecture is organized into extraction and loading, metadata management, experiment curation, Experiversum management, and multiple pipelines. Extraction and loading ingest raw data such as SAC files, CSVs, images, or social-media-derived records; datasets are grouped into catalogues according to ingestion date, size, format, and quantitative traits, with basic profiling at ingestion, including license, size, and provenance (Vargas-Solar et al., 30 Sep 2025). Metadata management implements the metamodel and normalizes metadata into a repository that supports navigation and comparison across datasets and experiments (Vargas-Solar et al., 30 Sep 2025). Experiment curation records data selection criteria, team roles and composition, research questions, performance constraints, evaluation criteria, and validation protocols, making it possible to inspect experiments on similar topics conducted under different methods or conditions (Vargas-Solar et al., 30 Sep 2025).
The metamodel is explicitly three-layered. Level 1: Raw Content covers datasets or releases, files or records, profiles, and annotations, together with attributes such as licensing, size, provenance, structure, and quantitative summaries. Level 2: Experimental Specifications covers actions, artefacts and models, parameters for actions, evaluation criteria, and validation protocols, capturing the structure, execution, and provenance of each action. Level 3: Experiment Context covers research teams, composition by role and seniority, responsibilities, and guiding research questions and objectives, thereby recording decision-making context and supporting inter-experiment comparison (Vargas-Solar et al., 30 Sep 2025).
This layered organization is significant because it formalizes a separation often blurred in ordinary research practice. Raw content, procedural transformations, and the social organization of inquiry are represented as linked but distinct objects. A plausible implication is that Experiversum functions simultaneously as a data repository, a workflow memory, and a contextual record of scientific judgment.
3. Iterative data cycles, pipelines, and reproducibility mechanisms
Experiversum is organized around iterative data cycles rather than linear pipelines. The cycle consists of data acquisition, preprocessing or transformation, analysis, interpretation and decisions, and iteration, with hypotheses, selection criteria, methods, or parameters being refined and the cycle repeated as needed (Vargas-Solar et al., 30 Sep 2025). All artefacts produced at each step are stored, and the metadata model records the processes and context that generated them (Vargas-Solar et al., 30 Sep 2025).
The operational pipeline structure is correspondingly explicit. The Extraction and Loading Pipeline performs data extraction from APIs, sensors, databases, or unstructured files; data cleaning to remove duplicates, errors, and inconsistencies; data enrichment with contextual details such as timestamps or geolocation reliability; and data loading into structured collections with associated metadata (Vargas-Solar et al., 30 Sep 2025). The Tagging Experimental Processes Pipeline specifies experiment identity, tags datasets and processes using algorithmic or user-defined tags, and stores those tags for reuse and reference (Vargas-Solar et al., 30 Sep 2025). The Transformation Pipeline structures text, signals, and media according to metadata entities, adds contextual enrichment that reflects experimental settings, and prepares data for analytics, machine learning, or further experimentation (Vargas-Solar et al., 30 Sep 2025). The exploration and analytics layer supports querying and retrieval by parameters such as time or location, filtering and aggregation, descriptive and predictive analytics, data visualization, and collaboration and sharing (Vargas-Solar et al., 30 Sep 2025).
Reproducibility is supported through several linked mechanisms. Each ingestion and transformation yields a dated, profiled dataset or release; provenance is tracked from sources through transformations and pipeline actions; parameters, evaluation criteria, validation protocols, and performance thresholds are recorded; and Level 3 metadata captures the team roles and research questions under which a result was produced (Vargas-Solar et al., 30 Sep 2025). The system’s stated position is that metadata should act as an execution guide for ELT processes rather than as passive documentation alone (Vargas-Solar et al., 30 Sep 2025).
Transparency is also framed as multi-perspective. Alternative analysis paths can be preserved and compared, such as human versus machine classifications or the outputs of different annotators. In the seismic monitoring case, the visualizations show about 90% agreement between human and machine classifications, with discrepancies flagged for expert review, while in species classification low-confidence cases below 0.6 identify items requiring manual checks (Vargas-Solar et al., 30 Sep 2025). This makes disagreement a stored analytic object rather than an informal afterthought.
4. Case studies and implementation domains
The ecosystem is demonstrated in Life Sciences, Earth Sciences, and Political Sciences. These case studies are methodologically important because they show that the same architecture is intended to span signals, text, images, qualitative annotation, and machine learning outputs within a common metadata framework (Vargas-Solar et al., 30 Sep 2025).
In the Life Sciences case, the goal is to classify sightings of Physalia physalis along the Brazilian coast. Instagram posts tagged with relevant hashtags such as #aguaviva and #caravelaportuguesa are retrieved, converted to CSV with fields including ID, source, location, and media URL, and ingested into the system (Vargas-Solar et al., 30 Sep 2025). CSV headers are then mapped to metadata entities such as experiment, media, content, and tags, and geo-temporal information is cleaned when descriptions are vague or imprecise (Vargas-Solar et al., 30 Sep 2025). Two research teams participate: data scientists use ML models to classify posts, while biologists manually tag posts and define classification categories; inclusion criteria, attributes such as gender of the affected person, model calibration, and performance thresholds are treated as experimental settings (Vargas-Solar et al., 30 Sep 2025). Querying supports analysis by time and region, comparison of human and ML classifications, and inspection of low-confidence cases (Vargas-Solar et al., 30 Sep 2025).
In the Earth Sciences case, the objective is to curate seismic data to distinguish natural from anthropogenic events and to produce validated bulletins. Raw SAC files are uploaded and validated, amplitude values on X, Y, and Z channels are extracted, and metadata such as station_id, channel_id, and timestamps are stored (Vargas-Solar et al., 30 Sep 2025). Junior analysts plot waveforms, detect events, and tag them; each event is annotated by station, year, and magnitude; analysts identify P and S wave arrivals; and a triangulated event list is validated by a senior analyst as the basis of the official bulletin (Vargas-Solar et al., 30 Sep 2025). The metadata model separates raw waveforms, detected events, analyst annotations, and validation decisions, thereby preserving the provenance of the bulletin (Vargas-Solar et al., 30 Sep 2025).
In the Political Sciences case, the goal is to determine whether political messages can be traced through graffiti across a city. A two-member team, junior and senior, refines the central research question over two cycles, defines inclusion criteria and political indicators, and builds a corpus in which the junior researcher photographs 1,050 graffiti images across districts; after review, 546 are validated and shared on Instagram (Vargas-Solar et al., 30 Sep 2025). Manual classification tagsets are then applied, and unsupervised ML using k-means and hierarchical clustering via Orange is used to cluster the images (Vargas-Solar et al., 30 Sep 2025). Narratives and metadata, including tags, cluster assignments, and review decisions, are compiled through successive review rounds (Vargas-Solar et al., 30 Sep 2025).
The prototype implementation uses Python for pipelines, Flask and Bootstrap for web interfaces, SQLite3 as the backend store, and executable Jupyter notebooks for programmable analysis and documentation (Vargas-Solar et al., 30 Sep 2025). The implementation is explicitly demonstrative rather than performance-oriented, and no detailed performance benchmarks or complexity analyses are reported (Vargas-Solar et al., 30 Sep 2025).
5. Related paradigms: living laboratories, experimentalist agents, and experiential benchmarks
Experiversum has strong affinities with the living labs paradigm in information retrieval. Living labs shift evaluation from static test collections toward operational systems in which real users issue queries, click, buy, edit, or otherwise interact while experimental switches alter ranking, interfaces, or content selection (Schaer, 2017). The contrast is often drawn against the Cranfield/TREC paradigm, which relies on a static document collection, a set of topics, and human relevance judgments for offline evaluation (Schaer, 2017). Within this in-vivo framework, the SOFISwiki experiment manipulated MediaWiki software to present an additional dashboard immediately after login, randomly assigned users to one of 17 user groups, and found that content attribution had a positive motivational effect on user contributions (Schaer, 2017). The REGIO JÁTÉK living lab in CLEF LL4IR 2015 evaluated a product-search ranking that combined a standard keyword-based text-relevance score with historical click data, using Team Draft Interleaving rather than simple A/B testing (Schaer, 2017). SSOAR, in turn, integrated a Living Labs API from LL4IR/TREC OpenSearch so that external researchers could submit queries and ranked lists for display to real repository users (Schaer, 2017). Taken together, these cases suggest a broader understanding of Experiversum as a live experimental substrate in which evaluation arises from usage data rather than precomputed labels.
A second related development appears in work on experimentalist LLM agents. "Hierarchical Experimentalist Agents" introduces HExA as a training-free, in-context self-improvement framework in which a tool-augmented actor runs experiments, an evolver converts trajectories into a persistent natural-language skill bank, and a retriever injects top-ranked skills and mistakes back into future episodes (Chandra et al., 28 Jun 2026). The associated Interphyre benchmark is built on the PHYRE 2D procedural physics environment, exposes simulation and analysis APIs such as simulate_action, predict_first_contact, and simulate_with_trace, and supports counterfactual branching via snapshot and restore (Chandra et al., 28 Jun 2026). Reported results show that Claude Sonnet 4.6 achieves only 2% success on the hardest levels in a direct setting, while HExA improves the same model to up to 77% success and reaches 44% success when only transferred skills from easier levels are used without active experimentation on the target level (Chandra et al., 28 Jun 2026). This indicates that an Experiversum-like setting can be instantiated not only for human-centered experimental science but also for agents that learn by structured intervention.
A third adjacent paradigm is BELA, the "Benchmark for Experiential Learning and Active Exploration" in repeated product recommendation. BELA uses 71K Amazon products organized into 2K choice sets, personas drawn from a set of up to 1M persona specifications, and an LLM user simulator that generates both answers and feedback from latent user preferences (Yang et al., 27 Nov 2025). Episodes are defined by customer, choice set, and latent factor, and evaluation is based on regret, , where is the best attainable product score and is the score of the recommended product (Yang et al., 27 Nov 2025). The central empirical result is that current frontier models perform better than random and popularity baselines but show no statistically significant downward trend in regret across episodes, even when given summaries of prior episodes or exact numerical regret (Yang et al., 27 Nov 2025). In this line of work, Experiversum can be understood as an environment of repeated, partially observed interaction in which experiential adaptation is itself the object of study.
6. Ethical, methodological, and prospective issues
The strongest explicit ethical critique associated with Experiversum-like systems comes from living labs research. Once experimentation is embedded in a running platform, problems of informed consent, privacy, harm, power asymmetry, and accountability become unavoidable (Schaer, 2017). In the SOFISwiki case, the terms of use were updated and users saw a notification pop-up, but they were not explicitly informed about the specific experiment or about being automatically assigned to one of 17 test conditions; there was no granular opt-out, because avoiding participation required ceasing use of the platform altogether (Schaer, 2017). In open systems based on MediaWiki, users often do not employ pseudonyms, making it possible to link projects, user accounts, and real persons, even when much of the underlying activity data is in principle publicly visible (Schaer, 2017). In ranking experiments, knowingly deploying sub-optimal or biased orderings may degrade user experience, while popularity-based signals may induce a Matthew effect in which already-visible items accumulate more clicks and low-visibility items remain suppressed (Schaer, 2017). In SSOAR’s living-lab configuration, external researchers could influence what users saw, yet the platform operator remained publicly accountable and had, by the authors’ account, no real possibility to validate the bias or validity of external rankings beyond relying on good will (Schaer, 2017).
Methodologically, these issues reinforce one of Experiversum’s central theses: sound experimental infrastructure must link data, processes, and context explicitly, including who decided what, under which assumptions, and with which consequences (Vargas-Solar et al., 30 Sep 2025). The architecture’s attention to team roles, responsibilities, validation protocols, and multi-perspective comparison can be read as a partial technical response to the opacity of ordinary exploratory work. Even so, the system’s own limitations are clearly noted. The SQLite3-based prototype may not scale to very large datasets or high-concurrency environments; tagging unstructured data such as images and social-media posts is difficult and requires more advanced NLP or related techniques; and interoperability across disciplines and data types remains challenging (Vargas-Solar et al., 30 Sep 2025).
The agentic literature adds a further methodological tension. HExA shows that active experimentation plus an external skill bank can produce substantial gains in a simulator (Chandra et al., 28 Jun 2026), whereas BELA shows that current frontier LLMs, when given past trajectories as ordinary prompt context, do not reliably improve across episodes and may even ask fewer questions while remaining poorly calibrated (Yang et al., 27 Nov 2025). This suggests that an Experiversum is not exhausted by storage and provenance. A plausible implication is that its long-term significance will depend equally on memory design, retrieval, uncertainty estimation, and governance: preserving experiments is one problem, but enabling responsible learning from them is another.
Across these literatures, Experiversum names a shift in the organization of empirical inquiry. Experiments are no longer treated as isolated runs whose outputs are summarized after completion. Instead, data, interventions, interpretive judgments, and sometimes even agent policies become part of a persistent, queryable, and ethically consequential environment.