TouchThinker-1M: Multi-Source Tactile Commonsense
- TouchThinker-1M is a unified tactile reasoning corpus that standardizes 9 source datasets into a common visuotactile video format with over 1M frames.
- It unifies diverse tactile annotations into four key attributes—hardness, protrusion, elasticity, and friction—using action-aware, 6-8 second video clips to capture temporal dynamics.
- The dataset supports heterogeneous tactile-language tasks including attribute understanding, chain-of-thought reasoning, and open-ended QA, yielding significant performance gains over prior models.
Searching arXiv for the cited TouchThinker and closely related tactile reasoning papers. TouchThinker-1M is a million-scale, multi-source visuotactile dataset introduced to support open-world tactile commonsense reasoning in tactile-language systems. It is presented as the central data contribution of “TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation” (Lyu et al., 10 Jun 2026), where it is positioned as a response to three recurrent limitations in earlier tactile reasoning resources: insufficient scale, narrow task formats, and limited sensor diversity. In the reported construction, TouchThinker-1M is a unified tactile reasoning corpus assembled from 9 source datasets and standardized into a common annotation and video format, with 1,001,344 tactile frames, 415+ deduplicated objects, 8 scenarios, 7 tactile sensing platforms, and 4 major acquisition actions (Lyu et al., 10 Jun 2026). The dataset is explicitly designed as a multi-source, cross-sensor, cross-action, cross-object corpus so that models can learn transferable tactile semantics rather than sensor-specific shortcuts (Lyu et al., 10 Jun 2026).
1. Scope and defining characteristics
TouchThinker-1M is defined as a unified tactile reasoning corpus for open-world tactile commonsense reasoning (Lyu et al., 10 Jun 2026). The paper frames tactile commonsense reasoning as inference from tactile observations to physical properties and then to higher-level everyday commonsense, rather than as simple attribute classification alone (Lyu et al., 10 Jun 2026). Within that framing, TouchThinker-1M is intended to provide broader coverage in objects, sensors, interaction actions, and task formats than prior tactile datasets (Lyu et al., 10 Jun 2026).
The reported dataset statistics are central to its identity. According to the appendix summary reproduced in the source material, TouchThinker-1M contains 1,001,344 tactile frames, 415+ deduplicated objects, 8 scenarios or scene categories, 7 tactile sensing platforms in the main dataset construction, and 4 major acquisition actions (Lyu et al., 10 Jun 2026). The data are organized as standardized tactile video clips, typically 6–8 seconds long, after contact-region cropping, frame-rate resampling, interpolation, temporal truncation, and length normalization (Lyu et al., 10 Jun 2026). Multiple task formulations are included: tactile attribute understanding, tactile feature analysis, surface feature distinction, surface optimality identification, object sensation correlation, tactile scenario analysis, chain-of-thought tactile reasoning, and open-ended tactile commonsense QA (Lyu et al., 10 Jun 2026).
This suggests that TouchThinker-1M is not only a scale expansion but also a format expansion. A plausible implication is that its design attempts to shift tactile-language training away from narrowly templated supervision toward heterogeneous reasoning tasks that require grounding in temporal tactile evidence.
2. Dataset composition and source integration
The construction of TouchThinker-1M is based on the integration of nine source datasets (Lyu et al., 10 Jun 2026). These are VTV-150K, PhysiCLEAR, Touch and Go, TacQuad, Touch-Slide, YCB-Slide, FeelSight-Real, ObjectFolder-Real, and HaTT (Lyu et al., 10 Jun 2026). The paper emphasizes that the merged resource is deduplicated and normalized at the object-category level, yielding 415+ deduplicated objects and 7 tactile sensor platforms spanning diverse materials, household objects, tools, surfaces, foods, and textures (Lyu et al., 10 Jun 2026).
The source-level details reported in the appendix illustrate the heterogeneity of the merged corpus:
| Source dataset | Reported coverage |
|---|---|
| VTV-150K | 150,000 frames, 100 common objects, GelSight Mini / DIGIT / Tac3D, pressing + rotation + sliding |
| PhysiCLEAR | 29,211 frames, 48 everyday objects, GelSight var.1 |
| Touch and Go | 28,300 frames, 18 material categories, GelSight17 var.1, egocentric probing |
| TacQuad | 37,955 frames, 92 objects, GelSight Mini + DIGIT, pressing + twisting |
| Touch-Slide | 12,150 frames, 9 toy-kitchen objects, DIGIT sliding |
| YCB-Slide | 9,450 frames, 7 YCB objects, DIGIT sliding |
| FeelSight-Real | 21,000 frames, 5 objects, DIGIT sensors on Allegro hand |
| ObjectFolder-Real | 686,880 frames, 70 household object instances, GelSight17 var.2 |
| HaTT | 26,398 frames, 66 material textures, GelSight var.2 |
The dataset’s multi-source assembly is one of its main distinguishing features (Lyu et al., 10 Jun 2026). Earlier tactile reasoning datasets are characterized in the source material as being too small, too template-bound, and usually limited to 1–3 sensor types (Lyu et al., 10 Jun 2026). By contrast, TouchThinker-1M explicitly aggregates data across sensors, actions, and object classes, with the stated objective of supporting open-world generalization (Lyu et al., 10 Jun 2026).
A comparison with the contemporaneous TouchReason-1M resource in “Touch-R1: Reinforcing Touch Reasoning in MLLMs” is informative because it reveals a different design emphasis (Lai et al., 26 May 2026). TouchReason-1M is described as a multimodal tactile reasoning dataset containing over 1M synchronized tactile data pairs across four optical tactile sensors, collected under a force-guided two-stage acquisition protocol and paired with compensated 3D contact force signals (Lai et al., 26 May 2026). TouchThinker-1M, by contrast, is reported as a standardized multi-source corpus built from pre-existing datasets, with emphasis on cross-source semantic unification and action-aware temporal representation (Lyu et al., 10 Jun 2026). This suggests a methodological distinction between controlled synchronized acquisition and large-scale dataset unification.
3. Semantic schema, preprocessing, and supervision formats
A core design decision in TouchThinker-1M is annotation unification into a shared 4-dimensional tactile attribute space:
(Lyu et al., 10 Jun 2026). The paper states that existing labels from source datasets are mapped into this unified schema, while unannotated data are manually labeled using object appearance, tactile observations, and deformation patterns (Lyu et al., 10 Jun 2026). Each sample is labeled by multiple annotators; disagreements are cross-checked and adjudicated; and samples with insufficient tactile evidence are excluded (Lyu et al., 10 Jun 2026).
This unified semantic space is significant because the source datasets originally use different annotation styles (Lyu et al., 10 Jun 2026). Instead of mixing heterogeneous labels, the dataset imposes a common semantic layer across sources and sensors (Lyu et al., 10 Jun 2026). A plausible implication is that such normalization is intended to reduce fragmentation of supervision and make cross-dataset training more stable.
The preprocessing pipeline converts all source data into a unified tactile video format (Lyu et al., 10 Jun 2026). The reported steps are retaining only valid contact intervals, removing non-contact, noisy, and redundant segments, turning static tactile images into short video sequences when needed, contact-region cropping, frame-rate resampling, interpolation, temporal truncation, and length normalization (Lyu et al., 10 Jun 2026). The resulting clips are typically 6–8 seconds long (Lyu et al., 10 Jun 2026). The paper motivates this by arguing that tactile cues are localized, temporally redundant, and action-dependent, so dynamic clips preserve interaction structure more faithfully than still images alone (Lyu et al., 10 Jun 2026).
The supervision layer is also expanded beyond fixed-answer templates (Lyu et al., 10 Jun 2026). Three formats are reported. The first is template-based QA, standardized into the unified attribute/task space (Lyu et al., 10 Jun 2026). The second is chain-of-thought tactile reasoning, formatted as:
1 |
<think>...</think><answer>...</answer> |
where > contains tactile evidence and reasoning steps and <answer> provides the final attribute judgment or commonsense conclusion (Lyu et al., 10 Jun 2026). The third is open-ended tactile QA, including free-form description, comparative reasoning, attribute explanation, interaction prediction, and open-world decision making (Lyu et al., 10 Jun 2026). The appendix states that the open-ended instruction synthesis yields 5,000 touch-language instruction-following samples (Lyu et al., 10 Jun 2026).
4. Open-world reasoning objectives and benchmark coupling
TouchThinker-1M is closely tied to TouchThinker-Bench, the benchmark introduced in the same paper to evaluate open-world tactile commonsense reasoning (Lyu et al., 10 Jun 2026). The dataset is explicitly designed to support evaluation on unseen objects and unseen sensors, and the benchmark operationalizes that objective (Lyu et al., 10 Jun 2026).
TouchThinker-Bench contains three task families (Lyu et al., 10 Jun 2026). The first is basic tactile property understanding, covering hardness, roughness, elasticity, and friction (Lyu et al., 10 Jun 2026). The second is basic tactile reasoning, including SFD (Surface Feature Distinction), SOI (Surface Optimality Identification), OSC (Object Sensation Correlation), and TSA (Tactile Scenario Analysis) (Lyu et al., 10 Jun 2026). The third is open-ended tactile commonsense reasoning, including TAU (Touch Attribute Understanding), TIU (Touch Interaction Understanding), and TKU (Touch Knowledge Reasoning) (Lyu et al., 10 Jun 2026). The benchmark is built from held-out test splits of TouchThinker-1M plus additional unseen-sensor data from TacQuad, VisGel, and a self-collected GelSight Mini dataset (Lyu et al., 10 Jun 2026). It contains 10 tactile sensors total and 82 test objects, with both cross-object and cross-sensor splits (Lyu et al., 10 Jun 2026).
The paper’s problem formulation emphasizes that tactile understanding depends strongly on the acquisition action: pressing may reveal hardness, sliding may reveal friction, and rotation may reveal texture (Lyu et al., 10 Jun 2026). This action dependence motivates both the dataset design and the associated model architecture. The reported question-guided tactile representation is:
$F_{\mathrm{qa} = \mathrm{SelfAttn} \left( \mathrm{CrossAttn} \left( F,\tilde{Q}_w,\tilde{Q}_w \right) \right)$
and the Gaussian temporal MoE is:
$f_{\mathrm{moe} = \sum_{k=1}^{K}\pi_k \sum_{t=1}^{T} \alpha_{k,t} E_k(f_{\mathrm{qa},t})$
(Lyu et al., 10 Jun 2026). In the paper, these mechanisms are introduced to exploit the fact that tactile videos are redundant and action-specific (Lyu et al., 10 Jun 2026). This suggests that the dataset’s temporal clip standardization and action diversity are not incidental collection choices but are structurally aligned with the representation strategy.
A related but distinct benchmarking philosophy appears in TouchReason-Bench from Touch-R1 (Lai et al., 26 May 2026). That benchmark evaluates tactile perception and visual-tactile conflict resolution, measuring H-Acc, R-Acc, P-Acc, Mat-Acc, OMAE, L2-EM, SFD, SOI-, and CSC over 4,800 QA pairs and 200 held-out objects (Lai et al., 26 May 2026). The comparison indicates that TouchThinker-1M and TouchReason-1M occupy overlapping but non-identical problem spaces: the former emphasizes open-world tactile commonsense reasoning with multi-source generalization, while the latter emphasizes tactile-grounded reasoning with cross-sensor physical consistency and visual-prior revision.
5. Empirical role in model performance
The reported experiments attribute substantial gains to the TouchThinker-1M dataset combined with the TouchThinker modeling framework (Lyu et al., 10 Jun 2026). On VTV-150K, TouchThinker-7B outperforms VTV-LLM-7B on tactile property prediction by +5.2 on hardness, +5.8 on protrusion, +8.0 on elasticity, +6.9 on friction, and +5.1 combined (Lyu et al., 10 Jun 2026). On reasoning subtasks it improves SFD by +7.6, SOI by +6.6, OSC by +7.5, and TSA by +10.0, for an overall average improvement of +7.0 over VTV-LLM-7B (Lyu et al., 10 Jun 2026). The paper further states that TouchThinker-7B also beats VTV-LLM-14B in overall performance despite using fewer parameters (Lyu et al., 10 Jun 2026).
On TouchThinker-Bench open-ended tasks, the source material reports that TouchThinker-7B consistently beats Octopi and VTV-LLM on METEOR and GPT-5 / DeepSeek-V4 scores (Lyu et al., 10 Jun 2026). The examples given are TAU METEOR 34.06 vs 27.93, TIU METEOR 28.71 vs 27.45, and TKU METEOR 27.43 vs 22.17 when compared with VTV-LLM-7B (Lyu et al., 10 Jun 2026). The paper defines the open-ended evaluation score as:
where the five dimensions are semantic correctness, tactile consistency, commonsense and reasoning plausibility, information completeness, and language quality (Lyu et al., 10 Jun 2026).
The open-world generalization result is especially central to the dataset’s intended purpose. On the unseen-sensor/unseen-object benchmark, TouchThinker-7B averages 58.6, compared with 49.3 for VTV-LLM-7B and 38.0 for Octopi-13B (Lyu et al., 10 Jun 2026). Within the logic of the paper, this is one of the clearest demonstrations that TouchThinker-1M improves generalization beyond the training distribution (Lyu et al., 10 Jun 2026).
The ablations likewise support the claim that both dataset design and representation matter (Lyu et al., 10 Jun 2026). Without action-aware modeling, the average drops to 61.6; without stage I, to 58.3; without stage II, to 53.3; and the full model reaches 67.0 (Lyu et al., 10 Jun 2026). Additional ablations show that question-guided fusion alone is helpful, Gaussian temporal MoE alone is helpful, and combining both is best (Lyu et al., 10 Jun 2026). This suggests that the dataset’s multi-action temporal structure is functional rather than merely descriptive.
6. Relation to adjacent tactile datasets and methodological debates
TouchThinker-1M is part of a broader 2026 expansion of tactile-language resources, but it occupies a distinct niche. The paper explicitly positions prior tactile reasoning datasets as limited in scale, format, and sensor diversity (Lyu et al., 10 Jun 2026). Its response is a multi-source corpus with unified semantics and open-world evaluation, whereas TouchReason-1M in Touch-R1 pursues large-scale synchronized tactile-force records with force-guided standardized acquisition and reinforcement learning for tactile-grounded reasoning (Lai et al., 26 May 2026).
The contrast is technically meaningful. TouchReason-1M reports 1,323,000 synchronized tactile data pairs and 14,700 valid tactile sequences after filtering, spanning 1000+ objects, 9 material categories, and 4 heterogeneous optical tactile sensors: GelSight Mini, Xense, Tac3D, and DM-Tac X (Lai et al., 26 May 2026). Each tactile data pair includes raw tactile observation, deformation field, shear field, depth field, force measurement, and associated metadata, all synchronized with compensated 3D contact force signals and stored as lossless frame stacks (Lai et al., 26 May 2026). That design prioritizes physical grounding and controlled cross-sensor consistency. TouchThinker-1M instead prioritizes broad multi-source coverage, semantic unification, and open-world transfer across 9 datasets and 7 sensor platforms (Lyu et al., 10 Jun 2026).
A common misconception would be to treat these resources as interchangeable because both are million-scale tactile datasets. The available descriptions do not support that simplification. TouchThinker-1M is framed around open-world tactile commonsense reasoning, unified 4-attribute semantics, and action-aware temporal modeling (Lyu et al., 10 Jun 2026), while TouchReason-1M is framed around tactile-grounded R1-style reasoning, ordinal-aware rewards, and counterfactual tactile-use objectives (Lai et al., 26 May 2026). This suggests complementary rather than redundant roles in the tactile MLLM ecosystem.
Another methodological issue concerns whether tactile reasoning should be modeled from static contact images or from temporal interaction sequences. TouchThinker-1M takes a clear position by converting source data into tactile video clips and by emphasizing that tactile signals are redundant and action-specific (Lyu et al., 10 Jun 2026). The argument is that only a few temporal segments are informative, and that different actions expose different physical properties (Lyu et al., 10 Jun 2026). A plausible implication is that datasets lacking temporal or action diversity may underrepresent the structure of tactile evidence required for commonsense reasoning.
7. Limitations, applications, and significance
The paper lists three principal limitations relevant to TouchThinker-1M (Lyu et al., 10 Jun 2026). First, attribute coverage is incomplete: the current schema includes hardness, protrusion, elasticity, and friction, but real tactile perception also involves properties such as malleability and prickliness (Lyu et al., 10 Jun 2026). Second, the interactions are mostly short-horizon; TouchThinker-1M clips are typically 6–7 seconds, and long-horizon manipulation remains a future direction (Lyu et al., 10 Jun 2026). Third, the associated large LLM backbones are computationally heavy; the reported experiments use 7B and 14B models, and deployment on resource-constrained robots remains challenging (Lyu et al., 10 Jun 2026).
These limitations clarify the current scope of the dataset. Although it is designed for open-world generalization, its semantic coverage remains selective and its temporal horizon is comparatively short (Lyu et al., 10 Jun 2026). The dataset is therefore broad in source diversity and task variety, but not exhaustive in the full space of tactile phenomena.
The intended applications reported in the paper include embodied AI, robotic manipulation, tactile understanding, quality inspection, assistive technology, cross-sensor tactile generalization, and open-world physical commonsense reasoning (Lyu et al., 10 Jun 2026). In that sense, TouchThinker-1M is presented as infrastructure for tactile-language systems that must operate on unseen objects and unseen sensors rather than only on narrowly controlled in-distribution benchmarks (Lyu et al., 10 Jun 2026).
The larger significance of TouchThinker-1M lies in its attempt to make tactile reasoning scalable along both the data axis and the representation axis (Lyu et al., 10 Jun 2026). Its million-scale size, multi-source construction, unified tactile labels, temporal contact video format, and richer QA and CoT supervision collectively define a dataset intended to support transferable tactile semantics in open-world settings (Lyu et al., 10 Jun 2026). This suggests a broader shift in tactile-language research: from small, template-bound tactile benchmarks toward heterogeneous corpora in which sensor variation, action dependence, and free-form reasoning are treated as first-class design constraints.