Papers
Topics
Authors
Recent
Search
2000 character limit reached

DataSway: Geotech Data & Visual Animation

Updated 7 July 2026
  • DataSway is a term representing two independent artifacts: a geotechnical dataset for soil expansion analysis and an SVG-based human–AI visualization system.
  • The geotechnical dataset comprises 395 one-dimensional swelling tests, detailing variables like swell percent, Atterberg limits, and compaction properties for regression and prediction.
  • The visualization system employs a three-stage human–AI workflow—clip generation, weight-based coordination, and timeline arrangement—to create dynamic, metaphoric animations.

DataSway is a name used in the arXiv literature for two distinct research artifacts in different domains. In geotechnical engineering, it denotes a compiled dataset of one-dimensional vertical free swelling tests and related soil properties intended for correlation analysis, data analytics, and machine-learning prediction of soil expansion during preliminary soil investigation and foundation design (U et al., 2021). In information visualization, it denotes a human–AI co-creation system for animating SVG-based metaphoric visualizations through vision–language-model clip generation, weight-based coordination, and timeline arrangement (Xie et al., 29 Jul 2025). The shared name can obscure the fact that these works address unrelated technical problems: expansive-soil characterization in one case, and bespoke animation authoring for data-rich visual metaphors in the other.

1. Dual usage of the term

In the available arXiv record, “DataSway” does not identify a single research program. One usage comes from documentation extracted and synthesized from Uyo and Onyekpe (2022) for a soil-swelling dataset (U et al., 2021). The other comes from a 2025 visualization system for animation clip generation and coordination in metaphoric graphics (Xie et al., 29 Jul 2025). A common source of confusion is therefore terminological rather than methodological.

DataSway usage Domain Core object
Data on one-dimensional vertical free swelling potential of soils Geotechnical engineering Dataset of swelling tests and soil properties
Vivifying metaphoric visualization with animation clip generation and coordination Information visualization / human–AI co-creation SVG-centric animation authoring system

This suggests that the most accurate encyclopedic treatment is disambiguating: the two DataSway artifacts share a name but differ in subject matter, methods, evaluation criteria, and intended users.

2. Geotechnical DataSway: dataset composition and scope

The geotechnical DataSway dataset reports standardized one-dimensional soil swelling on laboratory engineered and natural soils, together with corresponding index and compaction properties (U et al., 2021). It contains 395 one-dimensional free swelling tests. These tests include naturally occurring soils, both disturbed and undisturbed, as well as laboratory-engineered soils such as mixtures of bentonite–kaolinite, silica sand, and artificial clays. The dataset is organized into 13 groups, DS-1 through DS-13. Engineered soils appear in groups DS-3 and DS-11, while the remaining groups are predominantly natural soils.

The documented variables span swelling response, consistency limits, grain-size fractions, density-related quantities, and classification fields. The counts of available records differ by property, so the dataset is not a complete rectangular matrix.

Property Abbreviation Records
Swell percent SP 395
Plasticity index PI 395
Liquid limit LL 347
Plastic limit PL 347
USCS classification USCS 347
Clay content CC 339
Initial moisture content Mc 321
Dry unit weight γd\gamma_d 273
Specific gravity GsG_s 219
Activity index AA 209
Optimum moisture content OMC 246
Maximum dry density MDD 228
Void ratio ee 163
Silt content Sc 174

The dataset documentation also summarizes nominal value ranges to be computed from the raw Excel file: swell percent is described as ranging from approximately 0.5%0.5\% to 30%30\% or more, with a mean on the order of $10$–15%15\% and standard deviation approximately $5$–8%8\%; Atterberg limits are summarized as LL approximately GsG_s0–GsG_s1, PL approximately GsG_s2–GsG_s3, and PI approximately GsG_s4–GsG_s5; compaction parameters are summarized as OMC approximately GsG_s6–GsG_s7 and MDD approximately GsG_s8–GsG_s9 (U et al., 2021). Because the documentation explicitly states that exact descriptive statistics should be computed from the raw dataset, these figures function as approximate orientation rather than final benchmark values.

3. Geotechnical methodology, formulas, and modeling uses

The swelling measurements are anchored in one-dimensional free-swelling procedures centered on ASTM D4546, with additional references to AASHTO T258 and Turkish TS1900 laboratories (U et al., 2021). The apparatus is described as a standard oedometer or consolidometer fitted with stainless-steel or brass rings, porous stones, filter papers, and a surcharge loading cap. Undisturbed or remoulded soil is placed in the ring at its initial water content; compacted soils are prepared to target density and water content using Standard or Modified Proctor energy. The test sequence applies an initial surcharge, floods the sample by admitting water from the bottom, records vertical swell AA0 over time, and computes free swell percent.

Several definitions are explicit in the documentation:

AA1

AA2

AA3

AA4

Here, AA5 and AA6 are the swollen and dry sample volumes under zero osmotic stress; AA7 is gravimetric water content; and AA8 is the percentage of clay-sized particles. Associated index-property procedures include LL by Casagrande cup or fall-cone per ASTM D4318, PL by the hand-rolling method, moisture content by oven-drying at AA9 per ASTM D2216, specific gravity by pycnometer or volumetric flask per ASTM D854 or TS1900, and grain-size distribution by sieve and hydrometer or percussion per ASTM D422 and TS1900. Compaction characteristics are drawn from ASTM D698-07, ASTM D1557-07, AASHTO T99, AASHTO T180, and TS1900, with the documentation listing Standard Proctor energy as ee0 and Modified Proctor energy as ee1 (U et al., 2021).

The intended uses are explicitly analytical. The documentation proposes empirical correlations such as

ee2

and

ee3

with coefficients fit by ordinary least squares or advanced regression. It also identifies preliminary soil investigation and risk screening as direct applications, including rapid mapping of potential expansion hazards using PI, CC, and activity, and identification of “high-risk” soils such as those with ee4 and ee5. For machine learning, the raw .xlsx tables are described as suitable for supervised algorithms including ANNs, SVMs, and random forests, with soft-computing methods such as GA-optimized ANNs and ensemble classifiers noted as prior usage contexts (U et al., 2021).

A recurrent limitation is incompleteness and heterogeneity rather than lack of breadth. Some rows lack properties such as ee6 or ee7; the documentation recommends either imputation, including mean or ee8-NN, or removal of incomplete records. It also notes scatter induced by protocol variation in surcharge loads, ring sizes, and compaction energy, and recommends normalization by surcharge or tagging by group DS-1 through DS-13 during modeling. Unit consistency is treated as essential: moisture contents should remain in percent, densities in ee9, and length changes should be referenced to the same ring height.

4. Visualization DataSway: motivation and three-stage workflow

The visualization-system DataSway addresses the authoring of animations for SVG-based metaphoric visualizations such as flowers, trees, and playful glyphs (Xie et al., 29 Jul 2025). The underlying problem formulation is that animating such visualizations is expensive because creators must translate metaphorical intent into low-level keyframes, preserve data mappings during motion, and integrate animation with interactivity. The 2025 work positions this problem against both general motion-graphics tools, such as After Effects, Rive, and SVGator, and chart-animation toolkits that either privilege standard templates or lack semantic alignment with metaphoric imagery. It also situates metaphoric visualization in a lineage that includes Lupi and Posavec’s Dear Data and the “Four Seasons” tree example (Xie et al., 29 Jul 2025).

The system is structured as a three-stage human–AI co-creation workflow. Stage I, element-wise clip generation, takes an SVG snapshot of a static metaphoric visualization and a natural-language description of the desired motion. A VLM, specifically GPT-4o-mini, returns structured JSON keyframe specifications: one uniform clip per data-element class. These clips describe affine transforms, including translate, rotate, and scale, and style attributes such as opacity and fill. A validator prompt then checks whether the animated channels conflict with existing data encodings, warning, for example, when animating fill color could blur a categorical hue mapping.

Stage II, group-wise clip coordination, addresses the flatness that results when a single clip is applied identically to all instances of a class. Each element receives a normalized weight 0.5%0.5\%0 and an overall offset parameter 0.5%0.5\%1, giving a per-element delay

0.5%0.5\%2

where 0.5%0.5\%3 is the base delay. The system implements four coordination modes: data-centric, layout-centric, layer-centric, and random. Stage III, global timeline arrangement, presents the coordinated clips as blocks on a multi-track timeline, allowing users to drag and resize blocks to adjust start times and durations. The decomposition into clip generation, weight-based offsetting, and timeline editing is presented as a way to support rapid prototyping while retaining control (Xie et al., 29 Jul 2025).

The design rationale emerged from a formative study with eight experienced practitioners—five designers, two media artists, and one visualization researcher—using corpora of 20 live animated metaphoric examples and 30 static designs. That study yielded three design considerations: engagement across interaction phases, metaphor alignment, and complexity balance; it also identified the challenges of keyframe specification, multi-element coordination, and rapid iteration, which in turn motivated five design goals labeled DG1 through DG5 (Xie et al., 29 Jul 2025).

5. Coordination model, implementation architecture, and evaluation

The coordination strategies constitute the most formalized part of the visualization system (Xie et al., 29 Jul 2025). In the data-centric mode, weights are derived from data-proxy geometry:

0.5%0.5\%4

Here, 0.5%0.5\%5 approximates a glyph’s data value, allowing animations to run from largest to smallest or the reverse. Layout-centric coordination includes radius-based, projection-based, and sketch-based methods. Radius-based coordination uses a clicked center 0.5%0.5\%6 and Euclidean distance 0.5%0.5\%7; projection-based coordination projects each anchor point 0.5%0.5\%8 onto a drawn line 0.5%0.5\%9 and uses signed distance; sketch-based coordination records a polyline path 30%30\%0, finds the closest point 30%30\%1 on that path to each anchor, and uses the normalized arc-length of 30%30\%2. Layer-centric coordination uses the SVG DOM order, assigning

30%30\%3

while random coordination draws 30%30\%4.

The implementation is a React + TypeScript web application with five major components: an SVG Preview Panel, a Chat Panel powered by GPT-4o-mini via REST, a Keyframe Editor, a Coordination Panel, and a Timeline Panel. For large or complex SVGs that exceed the model’s context window, the system randomly samples repetitive element classes and strips non-essential attributes, falling back to core shapes. Export produces JavaScript functions using anime.js, and those functions recalculate weights at runtime based on the live SVG viewport so that animations adapt to responsive layouts and updated data (Xie et al., 29 Jul 2025).

A counterbalanced, within-subjects lab study compared DataSway with a chat-only baseline that prompted raw code from the VLM. Fourteen participants were recruited, with a mix of visualization and motion-design experience and with four overlapping the formative study.

Measure DataSway Baseline
Task A correct group coordination 78% 14%
Task B correct group coordination 71% 14%
Clip generation success 100% 100%
Creativity Support Index 66.64 (±12.25) 51.47 (±19.43)

The reported CSI difference was significant at 30%30\%5 using Wilcoxon. The highest gains were in Exploration, Expressiveness, and “Results Worth Effort.” Qualitative feedback emphasized layout-centric coordination, versioning support, and rapid previewing, while also identifying occasional hallucinations in clip generation, uncertainty about prompt phrasing such as “swing” versus “sway,” and a learning curve associated with the blended code/timeline interface (Xie et al., 29 Jul 2025). These results indicate not fully automated animation design, but rather a mediated workflow in which AI assistance is constrained by visualization semantics, authorial prompting, and explicit timeline editing.

6. Exemplars, implications, and comparative significance

The visualization paper demonstrates six gallery examples and an OECD Better Life Index case (Xie et al., 29 Jul 2025). The examples include disk glyphs rotating along a sketched path in “CD Sales,” hotel icons and stars sequenced by property size in “Hotel Ratings,” petal motion and glow in “Dreams,” legend-linked opacity and scale loops in “Eurovision,” random tick ripples and mountaintop fading in “Seal Collection,” and projection-based bamboo growth followed by swaying in “Bamboo Poems.” The OECD case combines projection-based reveal, randomized glints during resorting, and data-centric pulsing of individual petals while preserving filtering and reordering. These examples show that the system’s contribution lies less in inventing new animation primitives than in coordinating them in data- and layout-aware ways within web-based hypermedia.

The geotechnical dataset and the visualization system therefore occupy opposite ends of the data-workflow spectrum. One is a compiled empirical resource for expansive-soil analysis, with utility in regression, screening, and supervised prediction (U et al., 2021). The other is an authoring environment for expressive motion design in metaphoric visualization, evaluated through creativity support, usability, and coordination success rather than predictive performance (Xie et al., 29 Jul 2025). A plausible implication is that the shared name reflects a broad interest in “sway” as a metaphor for dynamic behavior, but the two artifacts are methodologically independent.

Future directions are explicit only for the visualization system. It raises questions about objective animation metrics, multimodal intent expression beyond text prompts, richer data-driven animation including physical simulation, and extension beyond SVG to canvas, WebGL, D3, and plot-grammar tools such as Vega-Lite (Xie et al., 29 Jul 2025). For the geotechnical dataset, the documentation is more application-oriented than prospective, but its emphasis on missing-data treatment, protocol heterogeneity, and model fitting suggests an ongoing role in statistical correlation studies and soft-computing prediction of soil expansion (U et al., 2021). Taken together, the two DataSway usages illustrate how the same label can designate either a scientific dataset for civil-infrastructure risk analysis or a human–AI interface for animating metaphor-rich data graphics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DataSway.