AQUAH: Autonomous Hydrologic Modeling Agent
- AQUAH is an automated hydrologic modeling framework that converts free-form text into complete simulation outputs and expert reports.
- It integrates four core modules and eight specialized subagents using vision-enabled LLMs and physics-based CREST modeling to configure and run simulations.
- The system autonomously retrieves data, initializes parameters, and executes models, producing clear, analyst-ready results while noting areas for calibration.
AQUAH, short for Automatic Quantification and Unified Agent in Hydrology, is described as the first end-to-end language-based agent designed specifically for hydrologic modeling. It converts a free-form prompt such as “simulate floods for the Little Bighorn basin from 2020 to 2022” into a completed hydrologic simulation and a self-contained PDF report by autonomously retrieving terrain, forcing, and gauge data, configuring a hydrologic model, running the simulation, and generating documentation. Its workflow is LLM-driven, vision-enabled, and agentic, with initial experiments across contiguous-U.S. basins indicating that it can complete cold-start simulations and produce analyst-ready documentation without manual intervention; the reported outputs were judged by hydrologists as clear, transparent, and physically plausible, while still requiring further calibration and validation for operational deployment (Yan et al., 4 Aug 2025).
1. System definition and scope
AQUAH is an LLM-driven, vision-enabled agentic framework whose central abstraction is the transformation of natural-language intent into a hydrologic modeling workflow. The system begins from a prompt that specifies a basin or spatial envelope, a time window, and optional requests such as “include bias analysis,” and it ends with simulation outputs, diagnostics, maps, hydrographs, parameter tables, and narrative commentary in Markdown → PDF form (Yan et al., 4 Aug 2025).
The system is organized around four major modules, described in the paper as four “pillars,” and these are coordinated by a crewAI task-manager that passes intermediate artifacts among eight specialized subagents. Vision is explicitly integrated into three of those subagents—Perceptor, OutletSelector, and ParamInitializer—which interpret maps and rasters on the fly, emit structured descriptors, and justify their choices in plain language. This architecture places map interpretation, outlet choice, parameter seeding, model execution, and reporting inside a single prompt-driven pipeline rather than in separate analyst-managed stages (Yan et al., 4 Aug 2025).
| Module | Function | Key output |
|---|---|---|
| Context Parser Agent () | Extracts structured metadata from user text | Basin name or latitude/longitude polygon, dates, specific requests |
| Dataset Retriever Agent () | Downloads and clips terrain, forcing, and gauge data | DEM, drainage direction, flow accumulation, precipitation, PET, discharge, metadata |
| Operator Agent () | Configures and runs CREST | Simulation outputs and run diagnostics |
| Report Writer Agent () | Assembles publication-quality report | Maps, hydrographs, parameter tables, narrative commentary |
A plausible implication is that AQUAH is intended to lower the barrier between Earth observation data, physics-based tools, and decision makers by shifting the user interface from dataset and model configuration toward prompt specification.
2. Agentic workflow and orchestration
The Context Parser Agent ingests free-form text and extracts structured metadata, including a latitude/longitude polygon or basin name, start and end dates, and specific requests. Once that spatial–temporal envelope is known, the Dataset Retriever Agent downloads raw GeoTIFFs for DEM, drainage direction, and flow accumulation from HydroSHEDS; reprojects each layer to a common grid using GDAL and Rasterio; clips to the basin polygon via Shapely; and writes the minimal GeoTIFF footprint required by CREST. It also fetches MRMS precipitation and FEWS NET PET by bounding-box queries, resamples these to the DEM grid, stores them in CF-compliant NetCDF, and queries USGS NWIS web services for discharge time series and gauge metadata such as drainage area, elevation, and record length (Yan et al., 4 Aug 2025).
Any missing layers trigger fallback defaults, including global PRISM rainfall climatology or a uniform PET field, together with a warning in the final report. This fallback behavior is part of the automation boundary: AQUAH is designed to preserve end-to-end execution even when data coverage is incomplete, while surfacing the incompleteness explicitly in the resulting documentation (Yan et al., 4 Aug 2025).
The Operator Agent configures and runs the CREST distributed water-balance plus kinematic-wave routing model over the requested period and records run diagnostics including NSE, RMSE, bias, and KGE. The Report Writer Agent assembles run metadata, maps, hydrographs, parameter tables, and narrative commentary into a publication-quality report. In the paper’s cold-start pseudocode, the sequence is: parse the user prompt, retrieve and clip rasters and forcings, describe the basin with a vision LLM, scan CREST manuals using RAG, initialize parameters, run CREST, compute discharge metrics, and assemble the report (Yan et al., 4 Aug 2025).
3. Hydrologic modeling core and parameter initialization
AQUAH employs the distributed CREST model, which solves at each grid cell a lumped water-balance coupled to kinematic-wave routing downstream. The local water-balance is written as
where is the soil-water storage, is precipitation input, is surface plus subsurface runoff exiting the cell, and is evapotranspiration up to potential (Yan et al., 4 Aug 2025).
Infiltration follows a variable-infiltration curve parameterized by exponent , maximum storage 0, and impervious fraction 1. Surface and subsurface flows are routed via a kinematic-wave scheme that, in CREST, reduces to a Muskingum–Cunge form with parameters 2. The modeling core is therefore physics-based, but the initial configuration is delegated to LLM-mediated perception and retrieval components (Yan et al., 4 Aug 2025).
Parameter initialization is handled by the ParamInitializer Agent. A vision-augmented LLM reads the DEM and flow-accumulation rasters to produce a short basin description containing drainage area, relief statistics, and land-cover hints. In parallel, retrieval-augmented generation scans the CREST user manual PDFs and web documentation to extract each parameter’s physical meaning and admissible range. The agent then combines those two knowledge sources to emit a JSON namespace of CREST arguments, exemplified in the paper by
9
This vector is explicitly characterized as a “first-guess” seed for a cold-start simulation. The paper therefore positions AQUAH not as an automated final calibration loop, but as an autonomous mechanism for physically informed initialization and execution (Yan et al., 4 Aug 2025).
4. Vision-grounded reasoning and uncertainty handling
Three agents leverage vision-grounded LLM calls. The Perceptor consumes quick-look PNGs of the clipped DEM, flow-accumulation raster, and basin outline, and extracts morphological descriptors including relief range, drainage density, candidate gauge locations, and their raster elevations. The OutletSelector then receives candidate gauges and applies an ordered rule set: exclude reservoir-controlled gauges; pick the lowest-elevation pour point; maximize drainage area; maximize record length; and sanity-check upstream regulation. The output can take the form Selected gauge: 02473000 together with a brief textual justification. The ParamInitializer uses vision cues to adjust infiltration and routing parameters based on slope, impervious-area inference, and land-cover signals (Yan et al., 4 Aug 2025).
This use of multimodal inference is central to AQUAH’s claim of end-to-end autonomy. Rather than treating rasters solely as inputs to numerical code, the framework uses LLM vision to generate intermediate symbolic descriptions that steer downstream choices. This suggests a hybrid control logic in which perception, rule-based selection, RAG, and numerical simulation are interleaved rather than separated into purely statistical and purely physical subsystems.
AQUAH also supports ensemble and sensitivity workflows by iterating runs over perturbed parameter sets. For each ensemble member 3, one obtains a discharge time series 4, from which one can compute, for example, a 90% confidence interval at each time 5. The report embeds statistical measures including Nash–Sutcliffe Efficiency, Kling–Gupta Efficiency, and root-mean-square error. Ensemble uncertainty is summarized by the median run plus the 5%–95% band, and parameter sensitivity is sometimes illustrated via tornado plots on one or two key parameters, such as 6 versus NSE (Yan et al., 4 Aug 2025).
5. Implementation, datasets, and evaluation
The reported implementation tested GPT-4o (OpenAI), Claude-4-Sonnet (Anthropic), and Gemini-2.5-Flash (Google), with production defaulting to GPT-4o for its balance of vision accuracy and reasoning consistency. Task orchestration uses crewAI v0.75.01. Geospatial operations rely on GDAL, Rasterio, Shapely, and Folium, while the CREST model is wrapped in Python with C extensions for the kinematic-wave solver. Vision prompts use temperature 0.3 and a 5 MB per-image limit with automatic down-scaling and JPEG compression; text-only prompts run at temperature 0 (Yan et al., 4 Aug 2025).
AQUAH was trialed on contiguous-U.S. basins spanning physiographic provinces: Little Bighorn, Maine Coastal, Mad-Redwood, and Upper Leaf River. Simulations ran from 2020-01-01 to 2022-12-31. Terrain came from HydroSHEDS 3″ (90 m), forcing sources were MRMS precipitation and USGS FEWS NET PET, and discharge records came from USGS NWIS. Evaluation combined objective metrics over the full period—NSE, KGE, CC, RMSE, and bias—with expert hydrologist scores on Model Completeness, Simulation Results, Reasonableness, and Clarity, as well as LLM co-evaluation via gpt-o3 using the same rubric (Yan et al., 4 Aug 2025).
The results section reports that all candidates produced physically plausible hydrographs without any manual pre-configuration. For parameter initialization, Figure 1 shows boxplots of CC and NSCE from ten independent first-guess runs per LLM, with GPT-4o occasionally achieving an NSCE 7 on the first try, Claude-Sonnet-4 delivering the most consistent positive-median NSCE, and Gemini showing larger scatter and lower medians. For outlet selection, Figure 2 reports that GPT-4o and Claude-Sonnet-4 correctly identified the single natural outlet more than 90% of the time in simple basins, whereas Gemini sometimes split choices among interior stations. When a reservoir lay just upstream of the nominal outlet, all models needed an explicit reservoir-mask hint to avoid regulated gauges. For report quality, aggregated human plus LLM scores show that Claude-4-Opus reached the highest average, 7.01/10, leading on Model Completeness and Clarity, while GPT-4o scored best on Simulation Results but lagged on readability (Yan et al., 4 Aug 2025).
6. Limitations, future directions, and nomenclature
The paper identifies several limitations. Prompt portability is incomplete because a single prompt set, tuned for OpenAI, was reused across all LLM back ends; model-specific prompt engineering may improve performance. Reservoir awareness remains limited because the current vision agents lack an explicit reservoir mask layer, allowing downstream regulated gauges to slip through unless disqualified by wording. Data accessibility is a dependency because full automation relies on always-available public APIs, including USGS, MRMS, and HydroSHEDS; offline or regional deployments would require local caching. Calibration loop functionality is also incomplete: AQUAH can generate first-guess parameters, but a tighter RAG-guided iterative calibration subagent feeding back NSE and KGE diagnostics was still to be implemented (Yan et al., 4 Aug 2025).
The stated future directions are technically specific: embedding a differentiable surrogate of CREST for gradient-based calibration; extending the vision toolkit to detect land-cover classes, wetlands, and snowpack; wrapping additional hydrologic models such as SWAT and HEC-HMS in the same agentic framework; and deploying a locally hosted, open-source LLM+CV stack for global basins outside CONUS. These directions indicate that AQUAH is framed as a generalizable orchestration layer over hydrologic model execution rather than as a single-model endpoint (Yan et al., 4 Aug 2025).
A potential source of confusion is the similarity of acronyms in adjacent literatures. On arXiv, “AQUA: A Collection of H8O Equations of State for Planetary Models” concerns planetary interior modeling and a multiphase water equation of state (Haldemann et al., 2020), while “q-AQUA: a many-body CCSD(T) water potential, including 4-body interactions” concerns a water potential energy surface for cluster and liquid-phase simulations (Yu et al., 2022). AQUAH, by contrast, designates an end-to-end, language-based agent for hydrologic modeling (Yan et al., 4 Aug 2025).