SymSkill: Symbol & Skill Co-Invention
- SymSkill is a framework that jointly learns predicates, operators, and low-level skills from unsegmented demonstrations for effective long-horizon robotic manipulation.
- It integrates symbolic planning with compliant, impedance-controlled motion to provide compositional generalization and real-time failure recovery.
- The approach is validated on tasks ranging from pick-and-place to composite procedures, demonstrating improved success rates and flexible skill recomposition.
SymSkill is a symbol-and-skill co-invention framework for long-horizon robot manipulation that jointly learns predicates, operators, and low-level skills from unlabeled and unsegmented demonstrations, then uses a symbolic planner to compose those learned skills at execution time while performing recovery at both the motion and symbolic levels in real time (Shao et al., 2 Oct 2025). It is explicitly designed to bridge two failure modes: imitation learning is reactive but lacks compositional generalization, whereas classical task-and-motion planning offers compositionality but has prohibitive planning latency for real-time failure recovery (Shao et al., 2 Oct 2025). In its original formulation, SymSkill targets deterministic, fully observed manipulation domains with multiple objects and couples symbolic reasoning to stable motion policies executed through a passive impedance controller.
1. Problem formulation and design objective
The original SymSkill formulation considers a manipulation state
$\state_t = \Big(\pose_{ee},\, \{\pose_{(o)}\}_{o\in\mathcal O},\, \{\type(o)\}_{o\in\mathcal O}\Big),$
where $\pose_{ee}$ is the end-effector pose, $\pose_{(o)}$ are object poses, and each object has a type from a finite set (Shao et al., 2 Oct 2025). Training data are unsegmented demonstrations
$\mathcal D = \{\tau_i\}_{i=1}^N, \qquad \tau_i = \{\state_t\}_{t=0}^{T_i},$
with no labels, no segmentation boundaries, and no task-specific skill annotations (Shao et al., 2 Oct 2025). At test time, an initial state $\state_0$ and a symbolic goal state are given, and the system must execute a sequence of learned policies so that the goal is achieved while monitoring for failures and recovering online (Shao et al., 2 Oct 2025).
SymSkill’s central design objective is to obtain IL-like data efficiency and low-level motion learning together with TAMP-like compositional planning and symbolic reasoning, but with real-time execution and recovery (Shao et al., 2 Oct 2025). A common misconception is to treat it as either a conventional symbolic planner or a monolithic imitation learner. The paper instead positions it as a hybrid system: symbols and skills are both learned from demonstrations, and online behavior is governed by symbolic planning over those learned components rather than by a single reactive policy (Shao et al., 2 Oct 2025).
The execution substrate is a passive impedance controller,
$F_{ee} = G - D(\dot{\pose}_{ee} - f(\cdot)),$
where is gravity compensation and $\pose_{ee}$0 is a damping matrix (Shao et al., 2 Oct 2025). This controller-level choice is operationally important because the paper ties real-time recovery to compliant, energy-dissipating execution under disturbances (Shao et al., 2 Oct 2025).
2. Offline co-invention of symbols and skills
The offline pipeline comprises segmentation and reference-frame selection, predicate learning, operator learning, and skill learning (Shao et al., 2 Oct 2025). Demonstrations are assumed to alternate between premotion and motion: in premotion the gripper moves toward an object while only the gripper moves, whereas in motion the gripper and one manipulated object move together or in coordinated fashion (Shao et al., 2 Oct 2025). For each trajectory, SymSkill detects motion changes using velocity thresholds and identifies the motion object $\pose_{ee}$1; for motion segments it also identifies a stationary reference object $\pose_{ee}$2 that gives semantic meaning to the motion (Shao et al., 2 Oct 2025).
For premotion segments, the data are represented in the motion object’s frame,
$\pose_{ee}$3
and for motion segments the retained representation is
$\pose_{ee}$4
In the described experiments, a VLM, Gemini-2.5-Pro, identifies the semantically meaningful stationary reference object from sampled frames, constrained to known scene objects by structured output (Shao et al., 2 Oct 2025). The VLM is used offline for reference selection rather than for online planning or policy generation (Shao et al., 2 Oct 2025).
Predicate learning is based on typed Boolean functions over frames,
$\pose_{ee}$5
with $\pose_{ee}$6 and $\pose_{ee}$7 (Shao et al., 2 Oct 2025). SymSkill learns relative-pose predicates by fitting Gaussians to relative-pose distributions observed in demonstrations. For premotion, the object-to-gripper relation is modeled by translation and orientation Gaussians, and a predicate holds when the associated Mahalanobis distances are below thresholds (Shao et al., 2 Oct 2025). Motion predicates are learned analogously for object-object relations, including a short post-motion window to stabilize end-pose estimates (Shao et al., 2 Oct 2025). The resulting Gaussian ellipsoids are later used not only as classifiers but also as samplers for goal resampling during recovery (Shao et al., 2 Oct 2025).
Operator learning converts demonstrations into abstract state sequences and defines a typed template
$\pose_{ee}$8
Here, pre are preconditions, eff are add/delete effects, maintain are predicates that must remain true throughout execution, and skill is the low-level controller realizing the operator (Shao et al., 2 Oct 2025). Given symbolic states $\pose_{ee}$9 and $\pose_{(o)}$0 before and after a transition, add and delete effects are computed from set differences, preconditions are obtained from the intersection of predicates holding in $\pose_{(o)}$1, and maintenance conditions are the predicates holding continuously during the transition (Shao et al., 2 Oct 2025). This construction is the formal bridge between learned symbolic abstractions and monitored continuous execution.
3. Skill representation, planning, and recovery
For each learned operator, SymSkill learns a low-level skill $\pose_{(o)}$2, where $\pose_{(o)}$3 is a pose controller and $\pose_{(o)}$4 is the gripper action (Shao et al., 2 Oct 2025). The motion policy is an $\pose_{(o)}$5 LPV-DS, with position and orientation controlled separately:
$\pose_{(o)}$6
The position field is expressed as a mixture of locally linear systems,
$\pose_{(o)}$7
where $\pose_{(o)}$8 are derived from a GMM over demonstration trajectories (Shao et al., 2 Oct 2025). The local linear dynamics are learned under a stability-constrained optimization, which the paper states ensures global asymptotic stability (Shao et al., 2 Oct 2025). For premotion operators, skills are learned in the motion object’s frame; for motion operators, learning in the reference object’s frame works better than using object-to-object trajectories directly, especially for non-prehensile motion (Shao et al., 2 Oct 2025).
At execution time, the user specifies a symbolic goal state $\pose_{(o)}$9, either directly as a conjunction of learned predicates or by abstracting a goal state 0 (Shao et al., 2 Oct 2025). The current continuous state is abstracted to a symbolic state 1, and planning proceeds via A* search over the learned operator set 2:
3
Because operators are typed and symbolic, the system can reuse skills across tasks, generalize to different numbers of objects, and reorder skills when a goal can be achieved by different action sequences (Shao et al., 2 Oct 2025).
Recovery operates at both the motion and symbolic levels. Motion-level recovery modulates the DS with a local obstacle-avoidance matrix,
4
with objects modeled as ellipsoids (Shao et al., 2 Oct 2025). Symbolic-level recovery monitors whether maintain conditions remain true and whether expected effects are satisfied when a skill ends; if either check fails, the current plan is abandoned and a new symbolic plan is computed (Shao et al., 2 Oct 2025). A further recovery mechanism resamples target poses from the learned predicate distributions: for premotion or grasp failures the target is resampled from 5, and for motion-attractor failures from 6 (Shao et al., 2 Oct 2025). This means failure handling is not only replanning over discrete operators but also respecifying continuous attractors.
4. Empirical performance
In RoboCasa simulation, SymSkill executes 12 single-step tasks with an average success rate of 85.0%, compared with 65.0% for a no-monitoring ablation and 3.3% for a variant that replaces the DS skill with Diffusion Policy (Shao et al., 2 Oct 2025). The evaluated single-step tasks are OpenSingleDoor, CloseSingleDoor, PnPCounterToCab, PnPCabToCounter, PnPStoveToCounter, PnPCounterToStove, OpenDrawer, CloseDrawer, TurnOnStove, TurnOffStove, TurnOnSinkFaucet, and TurnOffSinkFaucet (Shao et al., 2 Oct 2025). Several door, drawer, and switch-like tasks reach 100% in the reported runs, whereas pick-and-place tasks are harder, particularly when container geometry is awkward or induces collisions (Shao et al., 2 Oct 2025).
The paper introduces a composite task, StoreCheese, to test skill recomposition without collecting additional data (Shao et al., 2 Oct 2025). SymSkill solves this task by composing previously learned operators from OpenSingleDoor, PnPCabToCounter, and CloseSingleDoor, executing a 6-skill sequence and recovering from symbolic errors multiple times (Shao et al., 2 Oct 2025). This suggests that the learned symbolic operators are not merely descriptive abstractions over the training trajectories but can support recomposition into longer procedures not explicitly demonstrated as monolithic tasks.
Real-robot experiments use a Franka arm with about 5 minutes of unsegmented play data, pose tracking from a motion capture system, visual data from a webcam, and human demonstration via a UMI gripper (Shao et al., 2 Oct 2025). In this setting, SymSkill learns semantically meaningful operators such as picking a lid from a cabinet, placing a lid onto cookware, picking a thing from a drawer, and placing a thing into cookware or a container (Shao et al., 2 Oct 2025). The paper reports that certain pick actions from cookware require the lid to be removed first, indicating that logically structured preconditions can be inferred from limited household interaction data (Shao et al., 2 Oct 2025).
5. Position within typed-composition research
Later work explicitly places SymSkill in a family of typed-composition methods alongside BLADE and Generative Skill Chaining (Qin et al., 29 Apr 2026). In that framing, long-horizon execution is represented as an ordered phase sequence
7
with phase-specific controllers selected from typed candidate pools, and SymSkill is cited as a method that jointly learns predicates, operators, and skills from demonstrations with symbolic recovery (Qin et al., 29 Apr 2026). The shared abstraction is that composition depends on typed module selection rather than a single end-to-end policy.
A major qualification introduced by the governance literature is that typed composition has often been studied under a frozen-library assumption. The skill-update paper states that existing typed-composition methods, including SymSkill, treat the library as static after construction and do not analyze how composition outcomes change when a skill is replaced (Qin et al., 29 Apr 2026). It therefore reframes a deployment question that is orthogonal to the original SymSkill paper: not only whether a skill composes, but whether an updated version preserves composition reliability (Qin et al., 29 Apr 2026). A plausible implication is that SymSkill’s symbolic and typed structure is well suited to update-aware governance, even though the original paper did not study versioned skill replacement directly.
The term also appears by analogy in later non-robotic work. One 2026 semantic communication paper describes its decomposition into semantic abstraction, transmission, repair, and execution as modular and SymSkill-like because the stages are explicit, reusable skills connected through typed interfaces (Fu et al., 4 May 2026). A separate LLM-agent paper characterizes test-time adaptive skill synthesis as directly about a “SymSkill”-style idea, although it uses the name SkillTTA and synthesizes a temporary, task-specific textual skill rather than learning symbolic robot operators (Wang et al., 16 May 2026). This suggests that “SymSkill” developed a broader interpretive role as a reference point for explicit skill structure, even when the methods in question differ substantially from the original robotic framework.
6. Limitations, misconceptions, and significance
The original paper states several limitations. SymSkill currently assumes access to full object state or reliable pose tracking, relies on a VLM to identify reference objects offline, is demonstrated mainly in single-arm manipulation with relatively structured scenes, and still depends on meaningful segmentation into premotion and motion episodes (Shao et al., 2 Oct 2025). Performance also degrades in difficult geometry cases, especially pick-and-place tasks with awkward containers (Shao et al., 2 Oct 2025). These constraints delimit the scope of the claimed real-time and data-efficient behavior.
A second misconception is to equate symbolic abstraction with brittle hand engineering. SymSkill’s contribution is precisely that predicates, operators, and skills are learned automatically from play-like data rather than manually specified (Shao et al., 2 Oct 2025). Conversely, it should not be conflated with unconstrained end-to-end reactive policies: the empirical gap between full SymSkill and the no-monitoring or Diffusion Policy ablations shows that symbolic monitoring, typed operators, and stable DS skills are not incidental components but central to the reported performance profile (Shao et al., 2 Oct 2025).
The broader significance of SymSkill lies in its demonstration that learned relative-frame predicates and stable DS skills can support both compositional generalization and reactive recovery in long-horizon manipulation (Shao et al., 2 Oct 2025). Subsequent governance work adds that typed composition alone is not sufficient for deployment once libraries become versioned and updateable (Qin et al., 29 Apr 2026). Taken together, these results situate SymSkill as both a concrete robotic system and a reference architecture: it is a learned-symbolic manipulation framework whose main claim is that data-efficient skill acquisition, symbolic recomposition, and real-time recovery can coexist in one pipeline.