---
title: ActionQuery Features
url: https://www.emergentmind.com/topics/actionquery-features
type: topic
---

# ActionQuery Features

Aggregate Query Sculpting (AQS) introduces a direct-manipulation interaction model for visual exploration of large-scale multidimensional datasets. The approach is fundamentally “born scalable”, commencing analysis with a single highly-aggregated mark representing the entire dataset and enabling users to incrementally “sculpt” the data through a fixed sequence of aggregation-preserving operations. These operations, collectively termed the P⁶ operators—Pivot, Partition, Peek, Pile, Project, and Prune—enable fluid traversal from global views to finely faceted slices and aggregations, maintaining tractability over datasets spanning from tens of thousands to billions of rows. The Dataopsy prototype provides a validated implementation of this paradigm across desktop and mobile interaction modalities [2308.02764].

## 1. P⁶ Operations: Formal Definitions and Semantics

The six primary operations—Pivot, Partition, Peek, Pile, Project, and Prune—operate on the notion of “supernodes”: aggregates over arbitrary subsets $D' \subseteq D$ within the dataset $D$.

- **Pivot** splits a supernode along an attribute $A$ into disjoint subsets $D'_i = \{ d \in D' \mid A(d) = v_i \}$ for all values $v_i$ of $A$ (or bins in the quantitative case).
- **Partition** is a purely visual operation that spatially organizes supernodes along a chosen axis without altering their underlying records or summary statistics. For $n$ supernodes, cells are assigned dimensions proportional to $|D'_i|$ or equally partitioned along axis length $L$.
- **Peek** substitutes the basic aggregate glyph $g(D')$ with a detailed small-multiples glyph $g_A(D')$, such as a histogram, pie chart, or bar chart of a selected secondary attribute $A$: $g_A(D') \gets \text{Histogram}(\{A(d) ~|~ d \in D'\})$.
- **Pile** merges a selected set of sibling supernodes $\{D'_{i},\, i \in I\}$ into a single aggregated supernode $D'' = \cup_{i \in I} D'_{i}$, optionally relabeling the new node. The parent’s child list updates accordingly.
- **Project** removes a supernode $D'$ (or user-unioned group) from its current substrate and instantiates it as the root of a new, disjoint canvas or card. The substrate set transitions from $S$ to $S_1 = S \setminus \{D'\}$ and $S_2 = \{D'\}$.
- **Prune** excises a supernode $D'$ (or selected value slices within $D'$) from its parent’s child list, recursively eliminating all descendant records from subsequent traversals.

Each operator is defined to operate solely on aggregates, avoiding direct manipulation of individual records, thus ensuring computational and perceptual scalability.

## 2. User Interaction Models and Parameterization

Each P⁶ operator corresponds to distinct UI triggers and parameter sets engineered for rapid and reversible exploration:

- **Pivot:** Initiated via pivot/partition icons; parameters include the pivot attribute, optional bin boundaries (quantitative), and bin labels. Operates on the currently selected supernode.
- **Partition:** Activated by partition icons, with axis (horizontal/vertical) and sizing mode (equal/proportional) as parameters; governs layout for all children under a parent node.
- **Peek:** Accessed via an “eye” icon; parameters are the peek attribute and glyph type (pie/bar), with scope over selected or all supernodes.
- **Pile:** Supports multi-selection (e.g., shift-click/drag) followed by pile icon; parameter is an optional new label.
- **Project:** Triggered using a “bulb” icon; parameters include canvas selection and an optional card name. Operates on one/multiple nodes.
- **Prune:** Uses a scissors icon, with confirmation and undo support; acts on any selected supernodes or value slices.

These direct-manipulation affordances are optimized for both desktop and touch-based environments, affording fluid, reversible workflows.

## 3. Data Transformations and State Changes

P⁶ operations correspond to well-defined transformations within the underlying aggregation tree:

| Operation | Input                     | Output                                   |
|-----------|---------------------------|------------------------------------------|
| Pivot     | $D'$                      | $\{D'_1, ..., D'_n\}$ (split by $A$)     |
| Partition | $\{D'_i\}$ (siblings)     | Visual grid coordinates                  |
| Peek      | $D'_i$                    | $g_A(D'_i)$ (new glyphs)                 |
| Pile      | $\{D'_{i}, i \in I\}$     | $D'' = \cup_{i \in I} D'_{i}$            |
| Project   | $D'$/$\{D'_{i}\}$         | New canvas/card with root node $D'$      |
| Prune     | $D',\,$value slices       | $D'$ excised from aggregation tree       |

Supernodes maintain references to their constituent data; state transitions do not require data duplication, and data transformations are always conducted at the level of aggregates, never on raw rows.

## 4. Performance and Scalability Guarantees

AQS’s “born scalable” design ensures computational tractability for datasets up to and beyond the billion-row scale. Key considerations:

- **Pivot:** Group-by aggregation is $O(|D'|)$; for large $D'$, the grouping step is parallelizable (e.g., with Dask). Memory complexity scales with number of groups rather than $|D'|$.
- **Partition:** Layout cost is $O(n)$ for $n$ cells. For $n \gg 1$, glyphs are automatically resized; scrollbars added as needed.
- **Peek:** Local histograms are $O(|D'_i|)$; parallel or sampled group-by is employed on large aggregates.
- **Pile:** Union is $O(\sum_{i \in I} |D'_i|)$ in worst case, but frequently implemented as merged count/summaries.
- **Project:** No data duplication; only view pointers re-assigned.
- **Prune:** Excising a node is $O(1)$ in tree structure and summary count updates, with undo efficiently supported via reference restoration.

Empirical benchmarks in Dataopsy show even the partitioning of 1.7 billion NYC taxi trips by Year can execute in under a minute on a parallel backend [2308.02764].

## 5. Illustrative Applications and Workflow Patterns

AQS and Dataopsy have been validated on multiple case studies, reflecting common workflow idioms in large-scale exploratory data analysis.

- **Adult Income dataset:** Starting from a single supernode, users Pivot on “Race,” Partition horizontally, then Partition within each race by “Gender”—yielding a 5×2 grid. Peeking by “Accuracy Flag” within these cells as pie charts immediately exposes subgroup disparities.
- **Cars dataset:** Users Pile low fuel economy bins [0,10), [10,20) into a “poor economy” supernode, simplifying exploration. Prune is used to remove irrelevant categories, such as the “Other” race.
- **Screenplay dataset:** Prune eliminates low-frequency characters, enabling focus on the core ensemble.

These workflows exemplify interleaving of P⁶ operations: Pivot (slicing), Partition (layout), Peek (distribution probing), Pile/Prune (aggregation structure refinement), and Project (canvas reorganization), with every step operating exclusively on aggregates.

## 6. Constraints, Reversibility, and Interaction Fluidity

The AQS interaction vocabulary is expressly engineered to constrain visual and cognitive complexity:

- No individual P⁶ operation increases visible glyph count beyond manageable quantities; Partition always strictly respects the current supernode grouping.
- Project distributes complexity across canvases, mitigating combinatorial explosion from deep nesting.
- All operations are fully reversible via an undo stack that maintains parent/child references, with minimal recomputation required for restoration.

At each phase, users retain precise control over aggregation granularity and subset selection, avoiding information overload or accidental proliferation of irrelevant branches. This ensures analyses remain fluid, tractable, and fully under user direction, even at extreme data scales [2308.02764].

Source: https://www.emergentmind.com/topics/actionquery-features