ActionQuery Features
- ActionQuery Features are a set of direct-manipulation operations that enable scalable exploration of large multidimensional datasets using aggregate transformations.
- They employ the P⁶ operators—Pivot, Partition, Peek, Pile, Project, and Prune—to incrementally sculpt and analyze data with efficient, reversible interactions.
- The design emphasizes performance and scalability, ensuring efficient processing and fluid user interactions even when handling billions of rows.
Aggregate Query Sculpting (AQS) introduces a direct-manipulation interaction model for visual exploration of large-scale multidimensional datasets. The approach is fundamentally “born scalable”, commencing analysis with a single highly-aggregated mark representing the entire dataset and enabling users to incrementally “sculpt” the data through a fixed sequence of aggregation-preserving operations. These operations, collectively termed the P⁶ operators—Pivot, Partition, Peek, Pile, Project, and Prune—enable fluid traversal from global views to finely faceted slices and aggregations, maintaining tractability over datasets spanning from tens of thousands to billions of rows. The Dataopsy prototype provides a validated implementation of this paradigm across desktop and mobile interaction modalities (Hoque et al., 2023).
1. P⁶ Operations: Formal Definitions and Semantics
The six primary operations—Pivot, Partition, Peek, Pile, Project, and Prune—operate on the notion of “supernodes”: aggregates over arbitrary subsets within the dataset .
- Pivot splits a supernode along an attribute into disjoint subsets for all values of (or bins in the quantitative case).
- Partition is a purely visual operation that spatially organizes supernodes along a chosen axis without altering their underlying records or summary statistics. For supernodes, cells are assigned dimensions proportional to or equally partitioned along axis length .
- Peek substitutes the basic aggregate glyph with a detailed small-multiples glyph 0, such as a histogram, pie chart, or bar chart of a selected secondary attribute 1: 2.
- Pile merges a selected set of sibling supernodes 3 into a single aggregated supernode 4, optionally relabeling the new node. The parent’s child list updates accordingly.
- Project removes a supernode 5 (or user-unioned group) from its current substrate and instantiates it as the root of a new, disjoint canvas or card. The substrate set transitions from 6 to 7 and 8.
- Prune excises a supernode 9 (or selected value slices within 0) from its parent’s child list, recursively eliminating all descendant records from subsequent traversals.
Each operator is defined to operate solely on aggregates, avoiding direct manipulation of individual records, thus ensuring computational and perceptual scalability.
2. User Interaction Models and Parameterization
Each P⁶ operator corresponds to distinct UI triggers and parameter sets engineered for rapid and reversible exploration:
- Pivot: Initiated via pivot/partition icons; parameters include the pivot attribute, optional bin boundaries (quantitative), and bin labels. Operates on the currently selected supernode.
- Partition: Activated by partition icons, with axis (horizontal/vertical) and sizing mode (equal/proportional) as parameters; governs layout for all children under a parent node.
- Peek: Accessed via an “eye” icon; parameters are the peek attribute and glyph type (pie/bar), with scope over selected or all supernodes.
- Pile: Supports multi-selection (e.g., shift-click/drag) followed by pile icon; parameter is an optional new label.
- Project: Triggered using a “bulb” icon; parameters include canvas selection and an optional card name. Operates on one/multiple nodes.
- Prune: Uses a scissors icon, with confirmation and undo support; acts on any selected supernodes or value slices.
These direct-manipulation affordances are optimized for both desktop and touch-based environments, affording fluid, reversible workflows.
3. Data Transformations and State Changes
P⁶ operations correspond to well-defined transformations within the underlying aggregation tree:
| Operation | Input | Output |
|---|---|---|
| Pivot | 1 | 2 (split by 3) |
| Partition | 4 (siblings) | Visual grid coordinates |
| Peek | 5 | 6 (new glyphs) |
| Pile | 7 | 8 |
| Project | 9/0 | New canvas/card with root node 1 |
| Prune | 2value slices | 3 excised from aggregation tree |
Supernodes maintain references to their constituent data; state transitions do not require data duplication, and data transformations are always conducted at the level of aggregates, never on raw rows.
4. Performance and Scalability Guarantees
AQS’s “born scalable” design ensures computational tractability for datasets up to and beyond the billion-row scale. Key considerations:
- Pivot: Group-by aggregation is 4; for large 5, the grouping step is parallelizable (e.g., with Dask). Memory complexity scales with number of groups rather than 6.
- Partition: Layout cost is 7 for 8 cells. For 9, glyphs are automatically resized; scrollbars added as needed.
- Peek: Local histograms are 0; parallel or sampled group-by is employed on large aggregates.
- Pile: Union is 1 in worst case, but frequently implemented as merged count/summaries.
- Project: No data duplication; only view pointers re-assigned.
- Prune: Excising a node is 2 in tree structure and summary count updates, with undo efficiently supported via reference restoration.
Empirical benchmarks in Dataopsy show even the partitioning of 1.7 billion NYC taxi trips by Year can execute in under a minute on a parallel backend (Hoque et al., 2023).
5. Illustrative Applications and Workflow Patterns
AQS and Dataopsy have been validated on multiple case studies, reflecting common workflow idioms in large-scale exploratory data analysis.
- Adult Income dataset: Starting from a single supernode, users Pivot on “Race,” Partition horizontally, then Partition within each race by “Gender”—yielding a 5×2 grid. Peeking by “Accuracy Flag” within these cells as pie charts immediately exposes subgroup disparities.
- Cars dataset: Users Pile low fuel economy bins [0,10), [10,20) into a “poor economy” supernode, simplifying exploration. Prune is used to remove irrelevant categories, such as the “Other” race.
- Screenplay dataset: Prune eliminates low-frequency characters, enabling focus on the core ensemble.
These workflows exemplify interleaving of P⁶ operations: Pivot (slicing), Partition (layout), Peek (distribution probing), Pile/Prune (aggregation structure refinement), and Project (canvas reorganization), with every step operating exclusively on aggregates.
6. Constraints, Reversibility, and Interaction Fluidity
The AQS interaction vocabulary is expressly engineered to constrain visual and cognitive complexity:
- No individual P⁶ operation increases visible glyph count beyond manageable quantities; Partition always strictly respects the current supernode grouping.
- Project distributes complexity across canvases, mitigating combinatorial explosion from deep nesting.
- All operations are fully reversible via an undo stack that maintains parent/child references, with minimal recomputation required for restoration.
At each phase, users retain precise control over aggregation granularity and subset selection, avoiding information overload or accidental proliferation of irrelevant branches. This ensures analyses remain fluid, tractable, and fully under user direction, even at extreme data scales (Hoque et al., 2023).