Globus Flows: Automated Research Workflows
- Globus Flows is a cloud-hosted automation service that orchestrates multi-step research workflows, linking instruments, compute resources, and storage endpoints securely.
- It employs a state-machine model with a declarative JSON/Python DSL to define actions, choices, and error handling, ensuring reproducible and scalable processes.
- The system integrates AWS services, OAuth2 for secure access, and event-driven triggers to efficiently manage tasks from data transfer to ML training.
Globus Flows is a cloud-hosted research automation service central to the reproducible, efficient, and secure orchestration of multi-step workflows involving disparate scientific instruments, compute, storage, and human-in-the-loop tasks. At its core, Globus Flows enables declarative specification and execution of heterogeneous research processes—spanning spatial and temporal scales—by linking instruments, HPC clusters, storage endpoints, and external services in robust, parameterizable pipelines (Chard et al., 2022, Bicer et al., 2021).
1. Flow Specification and Execution Model
A "flow" in the Globus platform is a state-machine–style, multi-step process defined in JSON or a Python-oriented domain-specific language, extending the Amazon States Language (ASL). Each flow comprises:
- A
"StartAt"field marking the initial state. - A
"States"mapping of state names to definitions.
Supported state types include:
- Action: Invokes asynchronous external action providers (e.g., data transfers, compute functions).
- Choice: Conditional branching.
- Wait: Time-based delays.
- Pass: No-op, context reshaping.
- Fail: Flow termination with failure.
Each Action state specifies:
- An
ActionUrl(registered REST endpoint of the action provider), Parameters(input, mapped via JSONPath from flow context),ResultPath(output location in context),WaitTime(timeout),NextorEndfor control flow.
Flows declare an input schema (JSON Schema) both to validate initial context and to drive GUI or SDK front ends (Chard et al., 2022).
2. Architecture and System Integration
Globus Flows and companion services (Queues, Triggers, Timers) are fully cloud-native, leveraging AWS managed services for reliability, scaling, and resilience. Key components include:
- Front-end API: Implemented as AWS ECS Docker containers behind load balancers.
- Persistent State: AWS RDS for flow/run metadata and DynamoDB for OAuth tokens.
- Execution Engine: AWS Step Functions (ASF), driving each flow instance as a state machine, with ASF Task states enqueueing on Amazon SQS. AWS Lambda functions dequeue and invoke action providers; exponential backoff polling assures liveness.
- Queues Service: REST layer over SQS, with Globus Auth–protected roles to guarantee in-order, at-least-once message delivery.
- Triggers: Back-end workers that poll queues, evaluate user-defined predicates, transform event payloads, and post flow invocations.
- Timers: Priority-queued schedulers for periodic or time-triggered actions that persist through outages (Chard et al., 2022).
Integration with resources is federation-based:
- At instrument (edge), data is produced into directories or storage endpoints.
- Compute clusters (e.g., ALCF's ThetaGPU) host endpoints for Globus Transfer (data staging) and funcX (remote function execution).
- Globus Auth enables passwordless, single-consent access, with delegated OAuth2 tokens for flows to call Transfer/funcX on users’ behalf (Bicer et al., 2021).
3. Action Provider API and Authorization Delegation
Every action provider is a registered RESTful service supporting:
GET <action_url>/: Introspect provider, report scope/schema.POST <action_url>/run: Execute action, returning action_id and status.GET <action_url>/<action_id>/status: Poll state and output.POST <action_url>/<action_id>/cancel: Request cancellation.POST <action_url>/<action_id>/release: Free provider-side state.
OAuth2 scopes are issued and registered via Globus Auth per provider and per flow, with scope dependencies such that user consent to a flow’s run scope transits to all underlying actions. On invocation, the Flows front-end acquires needed tokens and caches them securely; Lambda backends retrieve and use these tokens as needed. "RunAs" directives allow flow states to execute as group-managed or alternate identities (Chard et al., 2022).
4. Event-Driven and Parameterized Workflow Automation
Globus Flows orchestrates automation using event decoupling:
- Event Generators (file watchers, instrument controllers) emit JSON payloads to named queues.
- Triggers poll queues, evaluate predicates, map events to flow inputs, and invoke parameterized flows.
- Timers periodically enqueue flow/action invocations based on schedule/count/end-time, ensuring orchestration over long timescales and resilience to service interruption.
Flows often parameterize actions per experiment (e.g., {scan_id} for beamline data), enabling massive concurrent sub-flow launches—e.g., one per sample, scan, or view (Bicer et al., 2021).
5. Use Cases and End-to-End Orchestration
Globus Flows underpins diverse, distributed research automations:
- Synchrotron Imaging: At APS, a driver detects new ptychographic scan directories, launching parameterized flows that (1) transfer raw data (bulk directory) from the beamline to ThetaGPU via Globus Transfer, (2) invoke GPU-accelerated reconstruction via funcX, and (3) stream results back to the beamline, with retries and dynamic job scheduling (Bicer et al., 2021). Concurrent flows (up to 64 GPUs) maximize pipeline overlap of transfers and compute for hundreds of views.
- Real-Time X-ray Crystallography: Multi-step per-image and per-batch flows automate image transfer, processing, visualization, catalog submission, and merging (Chard et al., 2022).
- ML Training and Inference: Flows tie together image staging, HPC-based MIDAS analysis, neural net training, and model distribution to edge devices.
- Data Publication and Human-in-the-loop: Flows coordinate transfers, interactive metadata entry, curator approval, DOI minting, and Search indexing in the Materials Data Facility.
- Analysis-as-a-Service: e.g., AlphaFold job orchestration, storage allocation, compute, publication, and user notification (Chard et al., 2022).
A characteristic flow for ptychographic reconstruction includes chained states with error and retry logic:
0
States are linked via the next field; each action has retry and failure-handling paths (Bicer et al., 2021).
6. Performance, Scaling, and Mathematical Foundations
End-to-end throughput depends on transfer, compute, and orchestration overheads. In experimental deployments:
- Single-node (8 A100 GPUs): For 1.8K frames, Catalyst dataset (~20 GB), ~1.0 s/iteration on one GPU, with best scaling at 1–2 GPUs due to communication overhead. Larger datasets (Coin 8K/16K, Siemens 32K) achieve >90% scaling to 8 GPUs.
- Multi-node: 8 nodes (64 GPUs) process 168 views in ~85 s end-to-end (vs. ~350 s serial/1 node, ~2500 s per GPU workstation).
- Flow run overhead (per step): For simple cases, mean overhead ~2.88 s; action provider (Transfer, funcX) median latencies are 3–4 s for trivial cases; authentication costs ~200–400 ms (Chard et al., 2022, Bicer et al., 2021).
The only explicit mathematical relation is for the ptychography iterative solver:
where is raw diffraction data, is the probe, and implements the update via Fourier transforms and projections (Bicer et al., 2021).
The composite makespan per view is:
for GPUs and iterations. For views processed with GPUs in parallel, the wall-clock time is approximately
modulo pipeline overlaps (Bicer et al., 2021).
7. Design Insights, Limitations, and Future Directions
Adapting ASL for declarative state-machine logic yields familiarity but can result in verbose flow specifications; higher-level SDKs and GUIs are in development (Chard et al., 2022). The action provider API is general, covering transfers, compute, DOI generation, search, notification, and manual tasks, though building robust providers requires custom implementation, token management, and scaling constructs.
Cloud-hosted, managed-service implementation eliminates user-side installations and provides reliable execution for long-lived flows. Throughput, not sub-second latency, is prioritized; overheads (1–4 s/step) limit applicability to real-time control loops but are acceptable for most distributed scientific automation use cases.
The OAuth2 authorization delegation model offers robust cross-resource security but can engender consent-management challenges for long-running, multi-domain workflows. Group-based access mitigates some user-experience limitations.
Early adopters report sufficient throughput and reliability for campaigns ranging from sub-second instrument triggers to multi-day HPC processing (Chard et al., 2022). Enhanced provenance capture, debugging, and analytics tools are recognized as desirable for future system evolution.
References:
- (Chard et al., 2022) "Globus Automation Services: Research process automation across the space-time continuum"
- (Bicer et al., 2021) "High-Performance Ptychographic Reconstruction with Federated Facilities"