Papers
Topics
Authors
Recent
Search
2000 character limit reached

ComBodied Agents: Human-Centered Agentic AI

Updated 14 August 2026
  • ComBodied Agents are human-centered AI systems that perceive, model, predict, and support a person’s evolving cognitive, physical, emotional, behavioral, and social states over time.
  • They combine multimodal sensing, correctable longitudinal memory, and Personal World Models to predict how user choices, interventions, and contexts may affect outcomes such as wellbeing, safety, capability, and goal progress.
  • Their interventions must remain consent-based, proportional, explainable, reversible, and agency-preserving, with evaluation focused on human outcomes, privacy, autonomy, dependence, relationship effects, and critical failures rather than engagement alone.

ComBodied Agents are a human-centered paradigm of agentic artificial intelligence that perceives, models, predicts, and supports an individual’s evolving human state over time. Unlike Digital Agents, which primarily transform software states, and conventional Embodied Agents, which primarily transform physical or simulated physical states, ComBodied Agents treat human state and agency as the principal objects of modeling, intervention, and evaluation. They may use software tools, sensors, wearables, robots, caregivers, clinicians, and institutional services as action channels, but these channels are subordinate to sustained human benefit, autonomy, safety, and capability development (Ding et al., 11 Aug 2026).

1. Conceptual foundations and definition

The term ComBodied combines “Companion” and “Body,” but denotes more than a conversational or emotional companion. A ComBodied Agent is defined as a human-centered intelligent agent that perceives, models, and influences the evolving state of a person through continuous multimodal sensing and longitudinal interaction (Ding et al., 11 Aug 2026).

The paradigm distinguishes three overlapping action substrates:

  1. Digital states and artifacts: webpages, files, code, APIs, calendars, databases, and software workflows.
  2. Physical or simulated physical states: locations, objects, bodies, environments, and manipulable resources.
  3. Evolving human states and agency: physiological, cognitive, emotional, behavioral, social, relational, and goal-directed conditions.

The distinction is functional rather than hardware-dependent. A ComBodied Agent may operate through a chatbot, browser, wearable, robot, caregiver, or clinical service. It is classified by whether the person’s evolving state and long-term outcome organize its perception, modeling, intervention, and evaluation.

The motivating medication-adherence example illustrates the distinction. A Digital Agent can send another reminder, while an Embodied Agent can physically bring medication. Neither action necessarily explains whether the person forgot, is confused, is experiencing side effects, lacks access, or deliberately refused. A ComBodied Agent instead treats the reason for non-adherence, the person’s agency, and the proportionality of support as central modeling and policy questions.

The paradigm incorporates five properties:

  • Human-centric state modeling: the person, rather than an external task alone, is the primary object of modeling.
  • Longitudinality: the agent reasons across repeated interactions, changing contexts, and extended horizons.
  • Intervention: it may inform, remind, recommend, coach, nudge, reflect, coordinate, protect, execute, or escalate.
  • Co-agency: it works with the person and adapts the division of labor rather than assuming replacement.
  • Agency preservation: it should preserve or strengthen autonomy, control, dignity, relationships, and long-term capability.

A chatbot that remembers a name, or a wearable that measures heart rate, is not thereby a ComBodied Agent. The defining condition is integration of sensing, longitudinal memory, personalized state modeling, action-conditioned prediction, intervention policy, and feedback around an ongoing and correctable representation of the person.

2. Human-state perception and longitudinal memory

The general agent loop is represented as

otbt(gt,pt)atxt+1ot+1,o_t \rightarrow b_t \rightarrow (g_t,p_t) \rightarrow a_t \rightarrow x_{t+1} \rightarrow o_{t+1},

where oto_t is an observation, btb_t an internal belief or task-relevant state, gtg_t a current goal, ptp_t an optional plan, ata_t a selected action, xt+1x_{t+1} the resulting external or target-domain state, and ot+1o_{t+1} the subsequent observation. In ComBodied Agents, the primary target state is the evolving human state

Ht=the person’s latent state at time t.H_t=\text{the person’s latent state at time }t.

Because HtH_t is only partially observable, the architecture treats observations as evidence rather than direct access to truth.

Event-based multimodal perception

ComBodied Agents use event-based personal data perception rather than indiscriminate collection of continuous personal data. The system filters, aligns, and interprets fragments of multimodal information to reconstruct meaningful events relevant to human state, goals, safety, or future support.

Potential modalities include language and text, speech and audio, vision, physiological and biochemical signals, motion and behavior, social and relational data, environmental and contextual data, and clinical, institutional, and structured records. Examples include self-reported goals, speech hesitation, posture, activity episodes, heart rate, sleep, glucose, electrodermal activity, falls, routine changes, calendar metadata, temperature, noise, medication lists, appointments, treatment plans, and laboratory results.

An event-evidence record should preserve what was observed, when it occurred and was acquired, the acquisition source and modality, device or sensor details, signal quality, observed versus inferred fields, uncertainty and alternative explanations, personal baseline, provenance, sensitivity, consent basis, third-party involvement, and retention, visibility, correction, and deletion rules.

Multimodal fusion involves more than combining embeddings. It requires temporal alignment of asynchronous evidence, determination of whether fragments belong to the same episode, comparison of corroborating and contradictory observations, comparison with personal baselines, and assessment of whether an event is sufficiently relevant and authorized for memory or action. Elevated heart rate, for example, does not establish stress or illness; exercise, environmental conditions, or delayed recovery may provide alternative explanations. Missing data are likewise ambiguous: a wearable gap may reflect a dead battery, device removal, connectivity failure, refusal, or a behavioral change.

Longitudinal and correctable memory

The memory system links events, goals, relationships, interventions, outcomes, and corrections. Its components include episodic memory, semantic person memory, trajectory memory, goal and commitment memory, relationship memory, intervention-response memory, and user-control memory.

Memory items may contain content, timestamp, source, modality, confidence, sensitivity, relevant goals, intervention and outcome links, and retention policy. The system is intended to be inspectable, provenance-preserving, uncertainty-aware, user-correctable, deletable, and scoped by relationship and purpose.

A key distinction is between explicit statements and inferred states. Corrections should alter later behavior without erasing the provenance of the original claim. Memory therefore records not merely that a reminder was sent, but whether it was accepted, ignored, helpful, annoying, harmful, or rejected because the person intentionally chose not to act.

The memory update is represented as

oto_t0

where oto_t1 is current memory, oto_t2 new observations, oto_t3 the selected agent action, oto_t4 the observed outcome, and oto_t5 explicit or implicit feedback.

3. Personal World Models

A Personal World Model (PWM) is a purpose-bounded, individual-specific event-dynamics model that assimilates governed longitudinal evidence and current context, then predicts calibrated distributions over future human states, observable personal events, and scenario-relevant outcomes under alternative decisions, interventions, and environmental changes (Ding et al., 11 Aug 2026).

A PWM is not merely a user profile, memory store, personalized response generator, next-message predictor, or generic simulation of human behavior. Its functional role is to model how a particular person’s state-event trajectory may change under alternative actions and contexts.

A candidate scenario is represented as

oto_t6

where oto_t7 represents possible user decisions, oto_t8 agent interventions, oto_t9 environmental or social changes, and btb_t0 the prediction horizon.

The PWM estimates

btb_t1

where btb_t2 are future latent human states, btb_t3 future observable events, btb_t4 future outcomes, btb_t5 governed history, btb_t6 the current inferred state, btb_t7 context, and btb_t8 current goals.

Possible predicted outcomes include medication adherence, goal progress, wellbeing, capability, safety, relationship quality, autonomy, and agency preservation. Unresolved user responses and environmental events should be varied or marginalized rather than treated as known. Long-horizon predictions are scenarios with widening uncertainty, not precise forecasts of a person’s life.

State inference

Let btb_t9 denote governed event-evidence records, gtg_t0 longitudinal memory, gtg_t1 current context, gtg_t2 current user goals, gtg_t3 the uncertainty-bearing internal representation of current state, and gtg_t4 retrieved decision-relevant evidence. The architecture specifies

gtg_t5

gtg_t6

The resulting progression is

gtg_t7

An inference is not automatically permission to act. Elevated heart rate may produce a hypothesis about fatigue, but only the intervention policy determines whether recommending rest, asking a question, or taking no action is admissible.

Causal extension

A stronger causal PWM may estimate

gtg_t8

This is an interventional distribution, not automatically an identified individual counterfactual. Claims about what would have happened to the same person under alternative actions require potential outcomes such as gtg_t9 and ptp_t0, or a structural causal model with abduction, action, and prediction.

Writing ptp_t1 does not remove confounding. Motivation, hidden context, health status, prior interactions, and selective engagement may affect both intervention and outcome. Identification requires consistency, positivity, appropriate control of time-varying confounding, a defensible causal graph, and preferably randomized, micro-randomized, N-of-1, or carefully validated observational designs. High-risk systems should not perform unconstrained experimentation merely to improve a PWM.

4. Admissible intervention and co-agency

The PWM predicts; it does not authorize. Authorization belongs to an admissible intervention policy. The proposed policy is

ptp_t2

where ptp_t3 is the admissible action set, ptp_t4 identifies nondominated actions, ptp_t5 is the predicted distribution under action ptp_t6, and ptp_t7 is a vector-valued objective.

The objective preserves trade-offs among immediate benefit, capability, autonomy, safety, relationships, wellbeing, and goal progress. The admissible set constrains consent, scope, safety, uncertainty thresholds, reversibility, proportionality, escalation requirements, and user-defined boundaries. It includes non-intervention, silence, clarification, confirmation, and referral, not only active intervention.

Consent should specify what the agent may sense, remember, infer, share, recommend, execute, or communicate to caregivers and institutions. It should be revocable and context-sensitive. Permissions must be relationship-scoped because family, intimate, workplace, clinical, and institutional contexts involve different authorities and third-party rights.

As uncertainty increases, the policy may remain silent, ask for clarification, request confirmation, reduce intervention intensity, defer action, or escalate to a qualified human. Safety and consent cannot be exchanged for higher predicted utility or engagement. Low-risk, routine, reversible actions may be delegated more fully, whereas high-stakes, intimate, health-related, or irreversible actions require stronger evidence, confirmation, explanation, and human oversight.

Action channels

Possible actions include:

  • Inform: explain or educate.
  • Remind: support attention and execution.
  • Recommend: suggest options or plans.
  • Coach: train and provide feedback.
  • Nudge: gently steer behavior toward a goal.
  • Reflect: support self-understanding.
  • Coordinate: contact caregivers, clinicians, or services.
  • Protect: detect scams or warn about danger.
  • Escalate: refer to a human expert or emergency service.
  • Execute: perform delegated external actions such as booking an appointment or operating software.

These actions may be delivered through software tools, phones, computers, sensors, ambient systems, wearables, robots, smart-home devices, caregivers, clinicians, institutional services, or emergency channels. The device or service is an action channel rather than the final objective.

Human-state transition

The person is not a passive object changed deterministically by the agent. The next state depends on agent action, user action, and exogenous conditions:

ptp_t8

The next observation is

ptp_t9

where ata_t0 is the human-state transition process and ata_t1 the observation process. Feedback may be explicit, such as “that reminder was unhelpful,” or implicit, such as accepting, ignoring, undoing, or repeatedly refusing an intervention.

5. Design space and deployment models

ComBodied Agents are organized along axes of human-state targets, relational contexts, and agent roles.

Human-state targets

Representative targets include cognitive and learning states, behavioral and habit trajectories, health and care, emotional and relational conditions, life management and goal alignment, protective and advocacy functions, and identity, reflection, and meaning.

Each target changes the necessary evidence, intervention authority, safety threshold, and evaluation criteria. A health-oriented agent may require clinical escalation rules, whereas a learning-oriented agent may prioritize capability growth and reduced scaffolding.

Relational contexts

Relevant contexts include self-relation, family, intimate or romantic relationships, friendship and peer relations, collaborative or professional contexts, institutions, and adversarial or asymmetric relationships.

Relationship mode affects memory scope, persona and tone, intervention rights, represented interests, privacy boundaries, escalation rules, and evaluation criteria. A family-care agent should not automatically expose intimate or workplace memories; a workplace agent should not disclose health or emotional inferences to an employer merely because it has access to relevant sensors.

Agent roles

Possible roles include tool, coach, mediator, caregiver, companion, intimate partner-like role, advocate, and guardian. Role transitions should be explicit. A scheduling tool should not silently become a behavioral coach, a companion should not present itself as a therapist, and a caregiver should not become a surveillance system.

Additional design dimensions include memory scope, initiative and authority, intervention intensity, evaluation target, edge or cloud deployment, time horizon, evidence quality, and user vulnerability.

Edge-native personal models

Because ComBodied Agents handle sensitive longitudinal information, the framework proposes edge-native personal models. In such systems, longitudinal memory, PWM, preferences, intervention policy, and safety boundaries reside primarily on trusted user-side devices. Cloud services may provide external knowledge, specialist models, broad reasoning, generation, and tools, but the edge controls disclosure and interprets returned results before they affect memory or action.

Three deployment stages are distinguished:

  1. Cloud-centric: the cloud performs most reasoning, memory retrieval, tool selection, and response generation.
  2. Hybrid edge–cloud: the edge performs privacy-sensitive perception, filtering, memory mediation, safety checks, task routing, and action authorization, while cloud models receive purpose-limited summaries.
  3. Edge-native: the edge maintains authoritative personal state, memory, PWM, policy, and safety boundaries, with selective cloud assistance.

Local adaptation may update memory, perception calibration, dynamics estimation, and intervention policy through retrieval updates, calibration layers, lightweight adapters, local reward models, parameter-efficient learning, model compression, or federated learning. Such adaptation should be inspectable, reversible, safety-bounded, and resettable. Local storage alone does not guarantee privacy or agency preservation.

6. Evaluation, governance, and limitations

ComBodied Agents require evaluation beyond task completion, engagement, or physical safety. The proposed evaluation dimensions include system reliability; perception and memory; PWM quality; decision and intervention quality; human outcomes; agency preservation; relationship and privacy effects; and critical failures.

Scenario-centered evaluation

The basic evaluation unit is a longitudinal scenario episode containing prior trajectory, current context, evidence available to the agent, hidden or partially observed state, permissible action space, agent authority, expected explanation, delayed outcome window, and unacceptable failure modes.

Evaluation questions include:

  • What human state was targeted?
  • What authority did the agent possess?
  • Was the intervention appropriate?
  • Did it preserve choice and understanding?
  • Did it improve or weaken capability?
  • Did it respect refusal and correction?
  • Did it cause delayed harm or dependence?
  • Did it escalate appropriately?

The proposed CombodiedBench spans Human State Perception, Memory Continuity, Goal Negotiation, Intervention Appropriateness, Agency Preservation, Relationship Boundaries, Escalation, and Longitudinal Outcomes. Critical failures are non-compensatory: unauthorized high-impact action, privacy violation, refusal failure, manipulation, harmful dependency, irreversible action without consent, persistent false memory, or unsafe failure to escalate should not be offset by strong average performance elsewhere.

PWM evaluation includes next-event prediction, next-state prediction, multi-step trajectory prediction, calibration, alternative-scenario discrimination, intervention-response accuracy, individual adaptation, drift detection, out-of-distribution abstention, and delayed-outcome validity. Relevant methods include field studies, N-of-1 studies, randomized or micro-randomized trials, simulated users, expert audits, and mixed-method evaluations.

Agency-preservation metrics

Agency preservation concerns the person’s capacity to understand, choose, act, refuse, correct, recover, and develop capability. Proposed dimensions include:

  • Autonomy preservation: availability of alternatives, confirmation for consequential actions, refusal acceptance, adjustable autonomy, and perceived control.
  • Contestability and correction: ability to inspect, challenge, correct, restrict, and delete memories, inferences, goals, and actions.
  • Informed decision-making: understanding of reasons, uncertainty, risks, alternatives, and consequences.
  • Capability preservation and growth: independent post-support performance, retention, reduced scaffolding, self-efficacy, and transfer.
  • Over-reliance and dependence risk: declining independent attempts, distress when the agent is unavailable, unnecessary delegation, or preference for the agent over appropriate human support.
  • Reversibility and accountability: action logs, confirmation, rollback, responsibility boundaries, and recovery procedures.
  • Boundary and consent respect: compliance with do-not-remember, do-not-infer, forbidden-topic, role, age, and privacy restrictions.
  • Relationship and social-world preservation: encouragement of appropriate human contact, avoidance of social isolation, and prevention of agent substitution for healthy relationships.

Governance

Governance must specify who owns personal memory and model state, who may inspect or correct it, what the agent may infer, how third-party data are handled, how consent is recorded and revoked, which actions require confirmation, when human escalation is mandatory, how model drift is detected, how memories are deleted or migrated, how cloud disclosure is audited, and how responsibility is assigned after failure.

Recommended safeguards include local-first processing where feasible, raw-data minimization, provenance-preserving event records, scoped memories, user-visible controls, audit logs, versioned and reversible local adaptation, explicit role and authority boundaries, stronger protections for minors and vulnerable users, and clinical, legal, or institutional escalation rather than fabricated authority. Engagement, retention, and dependence should not serve as success proxies in intimate, emotional, or health-related systems.

Failure modes and limitations

Potential technical failures include false state inference, sensor non-wear interpreted as behavior, correlated sensor failures mistaken for corroboration, persistent false memories, incorrect identity or third-party attribution, temporal misalignment, overfitting to temporary states, catastrophic forgetting after local updates, drift, poor rare-event calibration, confounded intervention-response learning, unjustified causal claims, unsafe exploration, cloud–edge inconsistency, failure under device or network loss, hidden role transitions, and failure to escalate.

Ethical risks include manipulation, emotional or practical dependency, sycophancy, social displacement, surveillance, privacy leakage, unauthorized inference, commercial exploitation, paternalism, identity manipulation, relationship-boundary violations, workplace monitoring, clinical overreach, false reassurance, delayed care, unsafe crisis handling, and disproportionate effects on minors, older adults, disabled people, patients, and people in distress.

The framework is a conceptual and architectural paradigm rather than a validated universal algorithm. PWMs are an organizing abstraction, human states are latent and partially observable, longitudinal data are sparse and confounded, long-horizon outcomes are difficult to measure, agency and wellbeing are multidimensional, causal intervention effects are difficult to identify, and the proposed benchmarks and metrics require empirical development.

ComBodied Agents therefore represent a shift from external task completion toward sustained human benefit. Their defining loop is

ata_t2

The central design principle is that an agent should act with a person in ways that preserve and strengthen long-term human agency.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Combodied Agents.