ComBodied Agents: Human-Centered Agentic AI
- ComBodied Agents are human-centered AI systems that perceive, model, predict, and support a person’s evolving cognitive, physical, emotional, behavioral, and social states over time.
- They combine multimodal sensing, correctable longitudinal memory, and Personal World Models to predict how user choices, interventions, and contexts may affect outcomes such as wellbeing, safety, capability, and goal progress.
- Their interventions must remain consent-based, proportional, explainable, reversible, and agency-preserving, with evaluation focused on human outcomes, privacy, autonomy, dependence, relationship effects, and critical failures rather than engagement alone.
ComBodied Agents are a human-centered paradigm of agentic artificial intelligence that perceives, models, predicts, and supports an individual’s evolving human state over time. Unlike Digital Agents, which primarily transform software states, and conventional Embodied Agents, which primarily transform physical or simulated physical states, ComBodied Agents treat human state and agency as the principal objects of modeling, intervention, and evaluation. They may use software tools, sensors, wearables, robots, caregivers, clinicians, and institutional services as action channels, but these channels are subordinate to sustained human benefit, autonomy, safety, and capability development (Ding et al., 11 Aug 2026).
1. Conceptual foundations and definition
The term ComBodied combines “Companion” and “Body,” but denotes more than a conversational or emotional companion. A ComBodied Agent is defined as a human-centered intelligent agent that perceives, models, and influences the evolving state of a person through continuous multimodal sensing and longitudinal interaction (Ding et al., 11 Aug 2026).
The paradigm distinguishes three overlapping action substrates:
- Digital states and artifacts: webpages, files, code, APIs, calendars, databases, and software workflows.
- Physical or simulated physical states: locations, objects, bodies, environments, and manipulable resources.
- Evolving human states and agency: physiological, cognitive, emotional, behavioral, social, relational, and goal-directed conditions.
The distinction is functional rather than hardware-dependent. A ComBodied Agent may operate through a chatbot, browser, wearable, robot, caregiver, or clinical service. It is classified by whether the person’s evolving state and long-term outcome organize its perception, modeling, intervention, and evaluation.
The motivating medication-adherence example illustrates the distinction. A Digital Agent can send another reminder, while an Embodied Agent can physically bring medication. Neither action necessarily explains whether the person forgot, is confused, is experiencing side effects, lacks access, or deliberately refused. A ComBodied Agent instead treats the reason for non-adherence, the person’s agency, and the proportionality of support as central modeling and policy questions.
The paradigm incorporates five properties:
- Human-centric state modeling: the person, rather than an external task alone, is the primary object of modeling.
- Longitudinality: the agent reasons across repeated interactions, changing contexts, and extended horizons.
- Intervention: it may inform, remind, recommend, coach, nudge, reflect, coordinate, protect, execute, or escalate.
- Co-agency: it works with the person and adapts the division of labor rather than assuming replacement.
- Agency preservation: it should preserve or strengthen autonomy, control, dignity, relationships, and long-term capability.
A chatbot that remembers a name, or a wearable that measures heart rate, is not thereby a ComBodied Agent. The defining condition is integration of sensing, longitudinal memory, personalized state modeling, action-conditioned prediction, intervention policy, and feedback around an ongoing and correctable representation of the person.
2. Human-state perception and longitudinal memory
The general agent loop is represented as
where is an observation, an internal belief or task-relevant state, a current goal, an optional plan, a selected action, the resulting external or target-domain state, and the subsequent observation. In ComBodied Agents, the primary target state is the evolving human state
Because is only partially observable, the architecture treats observations as evidence rather than direct access to truth.
Event-based multimodal perception
ComBodied Agents use event-based personal data perception rather than indiscriminate collection of continuous personal data. The system filters, aligns, and interprets fragments of multimodal information to reconstruct meaningful events relevant to human state, goals, safety, or future support.
Potential modalities include language and text, speech and audio, vision, physiological and biochemical signals, motion and behavior, social and relational data, environmental and contextual data, and clinical, institutional, and structured records. Examples include self-reported goals, speech hesitation, posture, activity episodes, heart rate, sleep, glucose, electrodermal activity, falls, routine changes, calendar metadata, temperature, noise, medication lists, appointments, treatment plans, and laboratory results.
An event-evidence record should preserve what was observed, when it occurred and was acquired, the acquisition source and modality, device or sensor details, signal quality, observed versus inferred fields, uncertainty and alternative explanations, personal baseline, provenance, sensitivity, consent basis, third-party involvement, and retention, visibility, correction, and deletion rules.
Multimodal fusion involves more than combining embeddings. It requires temporal alignment of asynchronous evidence, determination of whether fragments belong to the same episode, comparison of corroborating and contradictory observations, comparison with personal baselines, and assessment of whether an event is sufficiently relevant and authorized for memory or action. Elevated heart rate, for example, does not establish stress or illness; exercise, environmental conditions, or delayed recovery may provide alternative explanations. Missing data are likewise ambiguous: a wearable gap may reflect a dead battery, device removal, connectivity failure, refusal, or a behavioral change.
Longitudinal and correctable memory
The memory system links events, goals, relationships, interventions, outcomes, and corrections. Its components include episodic memory, semantic person memory, trajectory memory, goal and commitment memory, relationship memory, intervention-response memory, and user-control memory.
Memory items may contain content, timestamp, source, modality, confidence, sensitivity, relevant goals, intervention and outcome links, and retention policy. The system is intended to be inspectable, provenance-preserving, uncertainty-aware, user-correctable, deletable, and scoped by relationship and purpose.
A key distinction is between explicit statements and inferred states. Corrections should alter later behavior without erasing the provenance of the original claim. Memory therefore records not merely that a reminder was sent, but whether it was accepted, ignored, helpful, annoying, harmful, or rejected because the person intentionally chose not to act.
The memory update is represented as
0
where 1 is current memory, 2 new observations, 3 the selected agent action, 4 the observed outcome, and 5 explicit or implicit feedback.
3. Personal World Models
A Personal World Model (PWM) is a purpose-bounded, individual-specific event-dynamics model that assimilates governed longitudinal evidence and current context, then predicts calibrated distributions over future human states, observable personal events, and scenario-relevant outcomes under alternative decisions, interventions, and environmental changes (Ding et al., 11 Aug 2026).
A PWM is not merely a user profile, memory store, personalized response generator, next-message predictor, or generic simulation of human behavior. Its functional role is to model how a particular person’s state-event trajectory may change under alternative actions and contexts.
A candidate scenario is represented as
6
where 7 represents possible user decisions, 8 agent interventions, 9 environmental or social changes, and 0 the prediction horizon.
The PWM estimates
1
where 2 are future latent human states, 3 future observable events, 4 future outcomes, 5 governed history, 6 the current inferred state, 7 context, and 8 current goals.
Possible predicted outcomes include medication adherence, goal progress, wellbeing, capability, safety, relationship quality, autonomy, and agency preservation. Unresolved user responses and environmental events should be varied or marginalized rather than treated as known. Long-horizon predictions are scenarios with widening uncertainty, not precise forecasts of a person’s life.
State inference
Let 9 denote governed event-evidence records, 0 longitudinal memory, 1 current context, 2 current user goals, 3 the uncertainty-bearing internal representation of current state, and 4 retrieved decision-relevant evidence. The architecture specifies
5
6
The resulting progression is
7
An inference is not automatically permission to act. Elevated heart rate may produce a hypothesis about fatigue, but only the intervention policy determines whether recommending rest, asking a question, or taking no action is admissible.
Causal extension
A stronger causal PWM may estimate
8
This is an interventional distribution, not automatically an identified individual counterfactual. Claims about what would have happened to the same person under alternative actions require potential outcomes such as 9 and 0, or a structural causal model with abduction, action, and prediction.
Writing 1 does not remove confounding. Motivation, hidden context, health status, prior interactions, and selective engagement may affect both intervention and outcome. Identification requires consistency, positivity, appropriate control of time-varying confounding, a defensible causal graph, and preferably randomized, micro-randomized, N-of-1, or carefully validated observational designs. High-risk systems should not perform unconstrained experimentation merely to improve a PWM.
4. Admissible intervention and co-agency
The PWM predicts; it does not authorize. Authorization belongs to an admissible intervention policy. The proposed policy is
2
where 3 is the admissible action set, 4 identifies nondominated actions, 5 is the predicted distribution under action 6, and 7 is a vector-valued objective.
The objective preserves trade-offs among immediate benefit, capability, autonomy, safety, relationships, wellbeing, and goal progress. The admissible set constrains consent, scope, safety, uncertainty thresholds, reversibility, proportionality, escalation requirements, and user-defined boundaries. It includes non-intervention, silence, clarification, confirmation, and referral, not only active intervention.
Consent should specify what the agent may sense, remember, infer, share, recommend, execute, or communicate to caregivers and institutions. It should be revocable and context-sensitive. Permissions must be relationship-scoped because family, intimate, workplace, clinical, and institutional contexts involve different authorities and third-party rights.
As uncertainty increases, the policy may remain silent, ask for clarification, request confirmation, reduce intervention intensity, defer action, or escalate to a qualified human. Safety and consent cannot be exchanged for higher predicted utility or engagement. Low-risk, routine, reversible actions may be delegated more fully, whereas high-stakes, intimate, health-related, or irreversible actions require stronger evidence, confirmation, explanation, and human oversight.
Action channels
Possible actions include:
- Inform: explain or educate.
- Remind: support attention and execution.
- Recommend: suggest options or plans.
- Coach: train and provide feedback.
- Nudge: gently steer behavior toward a goal.
- Reflect: support self-understanding.
- Coordinate: contact caregivers, clinicians, or services.
- Protect: detect scams or warn about danger.
- Escalate: refer to a human expert or emergency service.
- Execute: perform delegated external actions such as booking an appointment or operating software.
These actions may be delivered through software tools, phones, computers, sensors, ambient systems, wearables, robots, smart-home devices, caregivers, clinicians, institutional services, or emergency channels. The device or service is an action channel rather than the final objective.
Human-state transition
The person is not a passive object changed deterministically by the agent. The next state depends on agent action, user action, and exogenous conditions:
8
The next observation is
9
where 0 is the human-state transition process and 1 the observation process. Feedback may be explicit, such as “that reminder was unhelpful,” or implicit, such as accepting, ignoring, undoing, or repeatedly refusing an intervention.
5. Design space and deployment models
ComBodied Agents are organized along axes of human-state targets, relational contexts, and agent roles.
Human-state targets
Representative targets include cognitive and learning states, behavioral and habit trajectories, health and care, emotional and relational conditions, life management and goal alignment, protective and advocacy functions, and identity, reflection, and meaning.
Each target changes the necessary evidence, intervention authority, safety threshold, and evaluation criteria. A health-oriented agent may require clinical escalation rules, whereas a learning-oriented agent may prioritize capability growth and reduced scaffolding.
Relational contexts
Relevant contexts include self-relation, family, intimate or romantic relationships, friendship and peer relations, collaborative or professional contexts, institutions, and adversarial or asymmetric relationships.
Relationship mode affects memory scope, persona and tone, intervention rights, represented interests, privacy boundaries, escalation rules, and evaluation criteria. A family-care agent should not automatically expose intimate or workplace memories; a workplace agent should not disclose health or emotional inferences to an employer merely because it has access to relevant sensors.
Agent roles
Possible roles include tool, coach, mediator, caregiver, companion, intimate partner-like role, advocate, and guardian. Role transitions should be explicit. A scheduling tool should not silently become a behavioral coach, a companion should not present itself as a therapist, and a caregiver should not become a surveillance system.
Additional design dimensions include memory scope, initiative and authority, intervention intensity, evaluation target, edge or cloud deployment, time horizon, evidence quality, and user vulnerability.
Edge-native personal models
Because ComBodied Agents handle sensitive longitudinal information, the framework proposes edge-native personal models. In such systems, longitudinal memory, PWM, preferences, intervention policy, and safety boundaries reside primarily on trusted user-side devices. Cloud services may provide external knowledge, specialist models, broad reasoning, generation, and tools, but the edge controls disclosure and interprets returned results before they affect memory or action.
Three deployment stages are distinguished:
- Cloud-centric: the cloud performs most reasoning, memory retrieval, tool selection, and response generation.
- Hybrid edge–cloud: the edge performs privacy-sensitive perception, filtering, memory mediation, safety checks, task routing, and action authorization, while cloud models receive purpose-limited summaries.
- Edge-native: the edge maintains authoritative personal state, memory, PWM, policy, and safety boundaries, with selective cloud assistance.
Local adaptation may update memory, perception calibration, dynamics estimation, and intervention policy through retrieval updates, calibration layers, lightweight adapters, local reward models, parameter-efficient learning, model compression, or federated learning. Such adaptation should be inspectable, reversible, safety-bounded, and resettable. Local storage alone does not guarantee privacy or agency preservation.
6. Evaluation, governance, and limitations
ComBodied Agents require evaluation beyond task completion, engagement, or physical safety. The proposed evaluation dimensions include system reliability; perception and memory; PWM quality; decision and intervention quality; human outcomes; agency preservation; relationship and privacy effects; and critical failures.
Scenario-centered evaluation
The basic evaluation unit is a longitudinal scenario episode containing prior trajectory, current context, evidence available to the agent, hidden or partially observed state, permissible action space, agent authority, expected explanation, delayed outcome window, and unacceptable failure modes.
Evaluation questions include:
- What human state was targeted?
- What authority did the agent possess?
- Was the intervention appropriate?
- Did it preserve choice and understanding?
- Did it improve or weaken capability?
- Did it respect refusal and correction?
- Did it cause delayed harm or dependence?
- Did it escalate appropriately?
The proposed CombodiedBench spans Human State Perception, Memory Continuity, Goal Negotiation, Intervention Appropriateness, Agency Preservation, Relationship Boundaries, Escalation, and Longitudinal Outcomes. Critical failures are non-compensatory: unauthorized high-impact action, privacy violation, refusal failure, manipulation, harmful dependency, irreversible action without consent, persistent false memory, or unsafe failure to escalate should not be offset by strong average performance elsewhere.
PWM evaluation includes next-event prediction, next-state prediction, multi-step trajectory prediction, calibration, alternative-scenario discrimination, intervention-response accuracy, individual adaptation, drift detection, out-of-distribution abstention, and delayed-outcome validity. Relevant methods include field studies, N-of-1 studies, randomized or micro-randomized trials, simulated users, expert audits, and mixed-method evaluations.
Agency-preservation metrics
Agency preservation concerns the person’s capacity to understand, choose, act, refuse, correct, recover, and develop capability. Proposed dimensions include:
- Autonomy preservation: availability of alternatives, confirmation for consequential actions, refusal acceptance, adjustable autonomy, and perceived control.
- Contestability and correction: ability to inspect, challenge, correct, restrict, and delete memories, inferences, goals, and actions.
- Informed decision-making: understanding of reasons, uncertainty, risks, alternatives, and consequences.
- Capability preservation and growth: independent post-support performance, retention, reduced scaffolding, self-efficacy, and transfer.
- Over-reliance and dependence risk: declining independent attempts, distress when the agent is unavailable, unnecessary delegation, or preference for the agent over appropriate human support.
- Reversibility and accountability: action logs, confirmation, rollback, responsibility boundaries, and recovery procedures.
- Boundary and consent respect: compliance with do-not-remember, do-not-infer, forbidden-topic, role, age, and privacy restrictions.
- Relationship and social-world preservation: encouragement of appropriate human contact, avoidance of social isolation, and prevention of agent substitution for healthy relationships.
Governance
Governance must specify who owns personal memory and model state, who may inspect or correct it, what the agent may infer, how third-party data are handled, how consent is recorded and revoked, which actions require confirmation, when human escalation is mandatory, how model drift is detected, how memories are deleted or migrated, how cloud disclosure is audited, and how responsibility is assigned after failure.
Recommended safeguards include local-first processing where feasible, raw-data minimization, provenance-preserving event records, scoped memories, user-visible controls, audit logs, versioned and reversible local adaptation, explicit role and authority boundaries, stronger protections for minors and vulnerable users, and clinical, legal, or institutional escalation rather than fabricated authority. Engagement, retention, and dependence should not serve as success proxies in intimate, emotional, or health-related systems.
Failure modes and limitations
Potential technical failures include false state inference, sensor non-wear interpreted as behavior, correlated sensor failures mistaken for corroboration, persistent false memories, incorrect identity or third-party attribution, temporal misalignment, overfitting to temporary states, catastrophic forgetting after local updates, drift, poor rare-event calibration, confounded intervention-response learning, unjustified causal claims, unsafe exploration, cloud–edge inconsistency, failure under device or network loss, hidden role transitions, and failure to escalate.
Ethical risks include manipulation, emotional or practical dependency, sycophancy, social displacement, surveillance, privacy leakage, unauthorized inference, commercial exploitation, paternalism, identity manipulation, relationship-boundary violations, workplace monitoring, clinical overreach, false reassurance, delayed care, unsafe crisis handling, and disproportionate effects on minors, older adults, disabled people, patients, and people in distress.
The framework is a conceptual and architectural paradigm rather than a validated universal algorithm. PWMs are an organizing abstraction, human states are latent and partially observable, longitudinal data are sparse and confounded, long-horizon outcomes are difficult to measure, agency and wellbeing are multidimensional, causal intervention effects are difficult to identify, and the proposed benchmarks and metrics require empirical development.
ComBodied Agents therefore represent a shift from external task completion toward sustained human benefit. Their defining loop is
2
The central design principle is that an agent should act with a person in ways that preserve and strengthen long-term human agency.