ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
First 10 authors:
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. Overview
This paper introduces a new idea called Combodied Agents.
Most AI agents today are designed to complete tasks. For example, a digital agent might fill out a form or write computer code, while a robot might move objects or bring medicine. The paper argues that these systems often focus too much on the task and not enough on the person affected by the task.
A Combodied Agent is an AI system designed to understand and support a person over time. It might consider the person’s health, goals, emotions, habits, abilities, relationships, and personal choices. Its main goal is not simply to do things for people, but to help people stay safe, capable, informed, and in control.
The name combines ideas from companion and body. However, the system does not have to be a talking companion or a robot. It could work through a phone, wearable device, computer, robot, or human care service.
2. Main Questions and Objectives
The paper mainly asks:
- How can AI systems focus on the changing needs and abilities of people, rather than only completing outside tasks?
- How can an AI remember important events and experiences over a long period?
- How can it tell whether a person needs information, a reminder, coaching, protection, or help from another human?
- How can AI support people without taking away their independence or encouraging unhealthy dependence?
- How should these systems be tested and judged?
The authors want to create a general framework for building AI that improves a person’s long-term wellbeing and abilities. For example, a learning agent should not only help a student get the right answer. It should also help the student understand the subject and solve similar problems independently later.
3. Research Approach
This is mainly a conceptual and framework paper. That means the authors do not mainly report one experiment with a group of participants. Instead, they bring together ideas from many areas of AI and human-centered technology and propose a new way to organize them.
The proposed Combodied Agent works like a repeating loop:
- Notice what is happening
- Remember useful information
- Estimate what might happen next
- Choose an appropriate kind of support
- Observe the result and learn from feedback
This is similar to how a thoughtful helper might act. Suppose an older person misses a medication dose. A simple system might just send another reminder. A Combodied Agent would try to understand the situation first. Perhaps the person forgot, felt confused, experienced side effects, or intentionally chose not to take the medicine. Each situation would require a different response.
Human-state perception
The system collects clues about the person from different sources, such as:
- Spoken or written messages
- Wearable sensors
- Movement and activity
- Facial expressions or posture
- Medical or care records, when authorized
- Information about the person’s surroundings
The paper calls this multimodal perception. “Multimodal” simply means using several kinds of information, much like understanding a situation by using sight, sound, words, and context together.
The system should focus on meaningful events, rather than saving every piece of data. For example, it might record that a person has repeatedly been sleeping poorly or has struggled with a daily task. It should also record how certain it is, where the information came from, and whether the person agrees with the interpretation.
Longitudinal memory
“Longitudinal” means looking across a long period of time.
The agent’s memory would not only store simple facts, such as a person’s name or favorite food. It could also remember:
- Important events and experiences
- Goals and commitments
- Changes in health, mood, habits, or skills
- Relationships with family members, caregivers, teachers, or doctors
- Which kinds of help worked or did not work
- The person’s corrections, privacy rules, and requests to delete information
This memory should be correctable. If the AI misunderstands someone, the person should be able to fix or remove the mistake.
Personal World Models
The paper proposes Personal World Models, or PWMs.
A Personal World Model is like a changing, incomplete map of how a particular person’s situation may develop. It could help answer questions such as:
- What might happen if the person receives a reminder?
- Would coaching be more useful than simply completing the task?
- Is the person becoming tired or overwhelmed?
- Could a proposed action improve the person’s ability, or make them more dependent?
The authors do not propose creating a perfect digital copy of a human. Such a complete copy would be unrealistic because people’s bodies, minds, relationships, and surroundings are extremely complicated. Instead, the system should build a limited model for a specific purpose, such as supporting learning, health, or daily independence.
Intervention policies
An intervention is something the AI does to help or influence a situation.
Possible interventions include:
| Type of support | Simple example |
|---|---|
| Inform | Explain a difficult idea |
| Remind | Give a medication or appointment reminder |
| Recommend | Suggest several possible next steps |
| Coach | Help someone build a skill or habit |
| Reflect | Help a person think about feelings or choices |
| Coordinate | Contact an approved caregiver or service |
| Protect | Warn about a scam or unsafe action |
| Escalate | Ask a doctor, caregiver, or emergency service for help |
| Execute | Book an appointment or complete a delegated computer task |
The AI should not always act. Sometimes it should stay quiet, ask a question, request permission, or refer the person to a human expert.
The paper says that decisions should consider consent, uncertainty, safety, reversibility, and user control. For example, an AI could automatically set a low-risk reminder, but it should probably ask for permission before contacting a doctor or sharing sensitive information.
4. Main Findings and Contributions
Because this is a framework paper, its main results are proposed ideas rather than numerical experimental findings.
The authors’ central conclusion is that existing AI agents are often organized around the wrong main target:
- Digital Agents mainly change digital information, files, websites, or software.
- Embodied Agents mainly change physical environments through movement and robotics.
- Combodied Agents would focus on the changing state and agency of a person.
The paper argues that many current systems have useful abilities, but these abilities are separated. For example, one system may have memory, another may track health, and another may provide emotional support. Combodied Agents would connect these abilities into one long-term system.
The paper identifies five important characteristics:
- Human-centered modeling: The person is the main focus, not just an outside task.
- Long-term understanding: The system considers changes over days, months, or years.
- Intervention: It can provide different kinds of support, not just answers.
- Co-agency: The AI works together with the person instead of automatically replacing them.
- Agency preservation: The system should protect independence, dignity, choice, relationships, and skill development.
The authors also say that success should not be measured only by speed or task completion. A system should also be judged by whether it helps people:
- Understand more
- Make better decisions
- Develop useful skills
- Stay independent
- Maintain healthy relationships
- Reach their own goals
- Remain able to disagree with or correct the AI
These points are important because an AI might complete a task successfully while making the person less knowledgeable or less independent. For example, an AI that always writes a student’s essays may produce good essays but prevent the student from learning to write.
5. Implications and Potential Impact
If developed responsibly, Combodied Agents could improve many areas of life.
In healthcare, such an agent might notice changes in a person’s routine, help them understand treatment choices, and contact a professional when necessary. It should not make unsupported medical guesses or secretly monitor people.
In education, it could adjust explanations to a student’s needs while gradually helping the student solve problems alone.
For older adults or people needing assistance, it could offer reminders and practical help while still respecting their choices and dignity.
In daily life, it could help people plan, understand information, manage commitments, or avoid dangers such as scams.
However, the idea also creates serious risks. A system with detailed information about a person could become too powerful. It might make incorrect assumptions, watch people too closely, manipulate their behavior, or encourage them to depend on the AI instead of human relationships. It could also make decisions based on private information without clear permission.
The paper therefore argues that these systems need strong rules for:
- Privacy and data ownership
- Consent and permission
- Clear explanations
- User correction and deletion
- Human oversight
- Safe limits on automatic action
- Protection from emotional manipulation
- Measuring independence and wellbeing, not just engagement
Conclusion
The paper proposes a major change in how we think about AI agents. Instead of asking only, “Can the AI complete this task?”, we should also ask, “Does this AI help the person become safer, more capable, more informed, and more in control over time?”
Combodied Agents are meant to act as supportive partners. They may use software, sensors, robots, or human services, but their true purpose is to improve human lives while protecting human choice. The idea could lead to more helpful and respectful AI, provided that people remain able to understand, correct, refuse, and control what the system does.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- No empirical prototype or deployment study demonstrates the full closed loop. The paper proposes perception, longitudinal memory, personal world models, intervention policies, and feedback adaptation, but does not show whether these components can operate reliably together in a real system.
- The concept of a “human state” remains underspecified. The framework lists physiological, cognitive, behavioral, emotional, social, and goal-related dimensions, but does not define a standardized state representation, measurement protocol, or method for resolving conflicts between these dimensions.
- The causal validity of Personal World Models is unresolved. The proposed models are expected to predict how an individual will respond to alternative interventions, but the paper does not explain how they will distinguish causal intervention effects from correlation, confounding, regression to the mean, or simultaneous life changes.
- Methods for learning from sparse, noisy, and missing personal data are not established. The paper recognizes discontinuous sensing and incomplete observations, but does not specify how models should handle missing-not-at-random data, sensor failure, selective self-reporting, or changes in device usage.
- Ground truth for latent human states is unavailable or highly subjective. The paper does not identify reliable labels or validation procedures for constructs such as agency, dignity, wellbeing, autonomy, capability growth, emotional state, or relationship safety.
- The framework lacks operational definitions for agency preservation. Agency is treated as a central success criterion, but the paper does not establish how to measure autonomy, control, understanding, refusal ability, responsibility, or long-term independence quantitatively.
- Trade-offs among human-centered objectives remain unresolved. The paper does not specify how an intervention policy should prioritize competing outcomes—for example, immediate safety versus autonomy, adherence versus informed refusal, efficiency versus learning, or emotional support versus social independence.
- The intervention policy lacks a validated decision rule for proportionality. Although consent, uncertainty, reversibility, safety, and escalation are identified as constraints, no formal thresholds determine when an agent should remain silent, ask for confirmation, intervene, execute an action, or contact another person.
- The appropriate level of autonomy for different domains is not established. The paper suggests greater human oversight for health, learning, emotional support, and high-stakes decisions, but does not provide domain-specific criteria for allocating authority between the agent and the user.
- Longitudinal adaptation may create feedback loops that are not analyzed. Repeated interventions can alter the user's behavior, data-generation process, expectations, and dependence on the agent, potentially invalidating the model that selected those interventions.
- No strategy is provided for distribution shift in personal life contexts. Major changes such as illness, relocation, bereavement, unemployment, changes in relationships, or altered routines may make historical memories and learned baselines misleading, but procedures for detecting and managing such shifts are not specified.
- Memory correction and deletion remain conceptually rather than technically defined. The paper calls for inspectable, correctable, and deletable memory but does not explain how corrections propagate through derived embeddings, inferred traits, causal models, cached predictions, backups, or information shared with external services.
- The framework does not resolve conflicts between user corrections and externally verified records. It remains unclear how the system should represent disagreements between a user's account, sensor data, caregiver reports, clinical records, and institutional data without privileging one source automatically.
- Consent models for continuous and multimodal sensing are insufficiently specified. The paper does not define how consent should be obtained, renewed, scoped by modality and context, withdrawn in real time, or managed when bystanders and other data subjects are captured.
- Bystander privacy and relational consent are underdeveloped. Ambient audio, video, institutional records, and caregiver communication may involve people who are not users of the system, but the framework does not establish their rights, notification mechanisms, or exclusion procedures.
- The boundary between support and unauthorized influence remains unclear. Actions such as nudging, coaching, protecting, and emotional reflection can shape preferences and behavior; the paper does not provide criteria for distinguishing beneficial assistance from manipulation or paternalism.
- Risks of emotional attachment and dependency lack measurable detection criteria. Relationship safety is identified as a requirement, but the paper does not define indicators, intervention thresholds, or validated safeguards for dependence, social withdrawal, sycophancy, or substitution for human relationships.
- Escalation and responsibility are not operationalized. The paper mentions referral to caregivers, clinicians, experts, or emergency services, but does not specify who is responsible for decisions, how false positives and false negatives are handled, or what happens when no authorized human is available.
- Safety guarantees for high-stakes interventions are absent. The framework does not provide evidence that agents can safely support medication, mental health, elder care, crisis response, financial protection, or other settings where incorrect inferences may cause serious harm.
- The effects of interventions on vulnerable populations are unexplored. Children, older adults, people with cognitive impairments, disabled users, people with mental-health conditions, and users with limited digital literacy may have different consent, agency, and interpretability requirements that are not concretely addressed.
- Cultural, linguistic, and socioeconomic variation is not empirically examined. The proposed signals and concepts—such as emotion, wellbeing, autonomy, appropriate support, and relationship boundaries—may vary substantially across populations, but no cross-cultural validation or adaptation method is presented.
- Potential demographic and sensor biases in human-state inference are not evaluated. The discussion does not report whether speech, vision, wearable, behavioral, or language-based models perform equitably across skin tones, genders, ages, disabilities, accents, languages, or device conditions.
- The paper does not establish whether multimodal sensing improves decisions enough to justify its privacy and burden costs. No ablation or cost–benefit analysis compares additional modalities against simpler, user-reported, or less intrusive alternatives.
- User burden and intervention fatigue are not modeled. Frequent prompts, confirmations, explanations, corrections, and consent requests may reduce adherence and increase cognitive load, but the framework does not specify how to optimize interaction burden over time.
- The notion of “beneficial human trajectories” is not sufficiently personalized or contestable. The paper assumes that capability, wellbeing, autonomy, and goal pursuit can jointly define benefit, but does not explain how users set priorities when these values conflict or change.
- Evaluation methods for long-term human gains are missing. The proposed scenario-centered evaluation is not accompanied by validated longitudinal designs capable of measuring retention, independent judgment, self-efficacy, health outcomes, relationship effects, or changes in capability after months or years.
- There are no benchmark datasets or task definitions for intervention-response prediction. Future researchers still need privacy-preserving benchmarks containing multimodal longitudinal evidence, user goals, interventions, outcomes, uncertainty, corrections, and adverse effects.
- The framework lacks baseline systems and comparative experiments. It does not specify how Combodied Agents should be compared with ordinary assistants, reminders, human support, domain-specific agents, or no-intervention controls.
- The benefits of edge-native personal models are asserted but not quantified. The paper does not evaluate the trade-offs among on-device privacy, computational cost, model capacity, personalization quality, synchronization, reliability, and selective cloud access.
- Security threats to personal models are not fully analyzed. Longitudinal memories and inferred vulnerabilities could enable profiling, coercion, identity theft, prompt injection, unauthorized inference, or malicious intervention, yet threat models and defenses are not specified.
- Ownership and governance of personal representations remain unresolved. The paper calls for user control but does not determine who owns memories, latent state representations, learned preferences, intervention-response data, or models trained from them.
- The framework does not explain how agents should behave when users request harmful, contradictory, or autonomy-reducing support. Procedures for negotiating goals, detecting coercion, handling impaired decision-making, and preserving the user's right to refuse are left open.
- The proposed paradigm’s boundaries relative to existing systems remain partly normative. The paper distinguishes Combodied Agents by their target state and success criteria, but does not provide an operational test for determining when an existing assistant, health agent, robot, or adaptive system qualifies as Combodied.
- The scalability of individualized modeling is unexamined. It remains unclear whether maintaining uncertainty-aware, correctable, multimodal, and longitudinal models is computationally and economically feasible for large populations without relying on centralized surveillance infrastructure.
- The supplied paper text does not present results validating the proposed claims. The available content is primarily conceptual and architectural; consequently, claims about improved human benefit, agency, safety, and long-term outcomes remain hypotheses requiring controlled empirical testing.
Practical Applications
Immediate Applications
The paper’s framework can be applied immediately by adapting existing software agents, wearables, personal assistants, care platforms, and educational tools. These deployments should use purpose-bounded data, explicit consent, user correction, and human escalation rather than attempting to construct a complete digital replica of a person.
- Medication and chronic-care support — Healthcare / eldercare
- Deploy an agent that combines medication schedules, user reports, wearable data, calendar context, and prior intervention outcomes to distinguish a missed dose caused by forgetfulness from confusion, side effects, access problems, or deliberate refusal.
- The workflow could begin with a low-intensity reminder, ask a clarifying question, provide education, notify an authorized caregiver, or escalate to a clinician depending on uncertainty and risk.
- Assumptions and dependencies: reliable medication records, user consent, clinically validated escalation rules, integration with healthcare systems, and clear limits preventing the agent from making unauthorized medical decisions.
- Personalized health and wellbeing monitoring — Healthcare / consumer health
- Extend existing health applications with longitudinal memory of symptoms, sleep, activity, mood reports, routines, and responses to previous recommendations.
- Potential products include a “personal health trajectory” dashboard, intervention-response logs, and a clinician-facing summary that separates user-reported facts, sensor observations, and model inferences.
- Assumptions and dependencies: sensor accuracy, individual baseline calibration, protection against false alarms, medical oversight for high-stakes recommendations, and compliance with health-data regulations.
- Adaptive reminders and routine assistance — Daily life / productivity
- Replace repetitive, context-insensitive reminders with proportionate support. For example, an agent could delay a low-priority reminder when the user is driving, ask whether a task remains relevant, or recommend a planning adjustment after repeated failures.
- Applications include reminders for appointments, hydration, rest, household tasks, deadlines, and travel preparation.
- Assumptions and dependencies: access to calendars and contextual signals, user-configured priorities, reversible actions, and safeguards against excessive notification or behavioral pressure.
- User-controlled personal memory — Software / personal information management
- Build memory assistants that store events, preferences, goals, relationships, interventions, and corrections with provenance, confidence, sensitivity labels, and retention rules.
- A practical product could provide controls such as “show why you remember this,” “correct this inference,” “do not remember similar information,” and “delete all memories related to this event.”
- Assumptions and dependencies: dependable source attribution, robust retrieval, secure storage, effective deletion, resistance to memory contamination, and interfaces understandable to nontechnical users.
- Agency-preserving workplace assistants — Enterprise software
- Adapt workflow, coding, data-analysis, and computer-use agents to assess cognitive load, responsibility boundaries, user understanding, and the reversibility of proposed actions before automating work.
- Examples include agents that draft but do not silently submit high-consequence documents, explain code changes before applying them, or return decision-relevant reasoning to the employee instead of optimizing only for speed.
- Assumptions and dependencies: organizational policies defining delegated authority, audit logs, user confirmation for consequential actions, and evaluation metrics beyond task completion and productivity.
- Learning assistants that strengthen independent capability — Education
- Use the proposed “coach,” “inform,” and “reflect” action modes to provide hints, diagnostic feedback, retrieval practice, and metacognitive prompts rather than immediately supplying answers.
- The system could track whether a learner later solves similar problems independently, thereby evaluating capability growth and calibrated reliance.
- Assumptions and dependencies: valid measures of learning and transfer, teacher involvement, age-appropriate safeguards, protection of student data, and avoidance of covert monitoring or excessive personalization.
- Scam, fraud, and unsafe-action protection — Finance / consumer security
- A personal agent could detect suspicious messages, unusual payment requests, or risky software actions and respond with explanation, warning, confirmation, or referral to a trusted person.
- Unlike a conventional fraud filter, the agent could account for the user’s known preferences, financial routines, vulnerability indicators, and prior corrections while preserving the user’s final control.
- Assumptions and dependencies: secure financial integrations, low false-positive rates, transparent explanations, strict authorization boundaries, and escalation protocols that do not expose private information unnecessarily.
- Assistive technology for older adults and people with disabilities — Accessibility / care
- Combine environmental sensors, calendars, voice interaction, and caregiver coordination to support daily routines, navigation, communication, and appointment management.
- The agent could adjust intervention intensity based on the person’s preferences and demonstrated capability, aiming to increase independence rather than replace it.
- Assumptions and dependencies: accessible interfaces, reliable consent from the user or authorized representative, protection against caregiver over-surveillance, robust performance across impairments, and preservation of dignity.
- Relationship-aware care coordination — Healthcare / social services
- Create a shared but permissioned coordination layer connecting a person, caregivers, clinicians, and service providers.
- The system could summarize relevant events, identify missed follow-ups, record who has been informed, and recommend escalation while maintaining separate access controls for sensitive memories.
- Assumptions and dependencies: interoperable records, role-based permissions, consent management, legal accountability, and safeguards against conflicting instructions from different parties.
- Scenario-centered evaluation of existing AI systems — Academia / industry
- Apply the paper’s evaluation principles to current assistants, tutors, health coaches, and workplace agents by measuring not only immediate task success but also comprehension, autonomy, retention, wellbeing, overreliance, and user ability to correct the system.
- A deployable evaluation workflow would compare intervention types—silence, clarification, reminder, recommendation, coaching, execution, and escalation—across realistic longitudinal scenarios.
- Assumptions and dependencies: validated agency-preservation metrics, longitudinal study designs, representative participants, privacy-preserving data collection, and agreement on what constitutes beneficial human change.
- Governance and product requirements for personal agents — Policy / standards
- Translate the paper’s principles into product requirements: explicit consent, visible sensing status, provenance of memories, user correction and deletion, uncertainty disclosure, reversible actions, intervention-rate limits, escalation rules, and independent audits.
- These requirements could inform procurement standards for schools, hospitals, employers, eldercare providers, and public agencies.
- Assumptions and dependencies: enforceable regulation, sector-specific definitions of harm, technical auditability, and mechanisms for users to challenge automated decisions.
Long-Term Applications
The following applications require advances in causal modeling, multimodal sensing, privacy-preserving infrastructure, clinical and educational validation, or governance. They are plausible extensions of the paper’s proposed Personal World Models and closed-loop intervention policies, not capabilities established by the paper itself.
- Personal World Models for intervention planning — Cross-sector AI
- Develop models that estimate how a particular person may respond to alternative actions, contextual changes, or non-intervention across health, learning, work, relationships, and goal pursuit.
- Such models could simulate options such as “remind now,” “ask for clarification,” “suggest a break,” “provide a lesson,” or “contact a caregiver,” while representing uncertainty and possible side effects.
- Assumptions and dependencies: sufficiently long and high-quality longitudinal data, causal rather than merely correlational learning, calibrated uncertainty, protection against feedback loops, and rigorous validation before use in high-stakes settings.
- Edge-native personal intelligence — Consumer hardware / privacy technology
- Move personal memory, state estimation, and intervention policies onto phones, wearables, home hubs, and robots, using cloud services only for explicitly authorized tasks.
- Potential tools include a user-owned personal model, local event detection, encrypted memory stores, and on-device policy enforcement for sharing and intervention.
- Assumptions and dependencies: efficient multimodal hardware, secure model updates, robust local inference, interoperability, energy efficiency, recovery from device loss, and verifiable privacy guarantees.
- Integrated home-care and rehabilitation robots — Robotics / healthcare
- Combine physical assistance with human-state modeling so that robots adapt support to fatigue, confidence, capability, and rehabilitation progress.
- A rehabilitation robot might reduce physical assistance as the user improves, provide coaching instead of taking over, or escalate to a therapist when progress or safety deteriorates.
- Assumptions and dependencies: reliable perception of human state, safe physical control, clinically validated outcome measures, liability frameworks, affordable hardware, and strict limits on autonomous physical intervention.
- Adaptive mental-health and emotional-support systems — Mental healthcare
- Build systems that detect changes in mood, stress, routine, or social connection and select among reflection, coping education, grounding exercises, human referral, or silence.
- The system could maintain relationship boundaries and detect possible dependency, social withdrawal, or escalating risk rather than optimizing conversation duration or engagement.
- Assumptions and dependencies: clinically meaningful detection across cultures and populations, validated intervention efficacy, crisis-response integration, professional oversight, privacy protection, and prevention of anthropomorphic manipulation.
- Longitudinal education and lifelong learning companions — Education / workforce development
- Model a learner’s evolving knowledge, confidence, misconceptions, goals, and independent judgment over months or years.
- The agent could coordinate tutoring, practice, project work, peer interaction, and human instruction while deliberately returning responsibility to the learner as competence develops.
- Assumptions and dependencies: reliable cross-context learning records, safeguards against educational profiling, teacher and learner control, culturally valid measures of capability, and evidence that personalization improves—not narrows—learning opportunities.
- Human-centered decision support in high-stakes institutions — Public policy / law / finance / healthcare
- Use Combodied Agents to help people understand complex choices while preserving responsibility and contestability.
- Examples include patient decision aids, benefits-navigation assistants, financial planning tools, legal-information systems, and public-service agents that explain options, identify missing information, and support informed human decisions rather than automatically choosing outcomes.
- Assumptions and dependencies: explainability appropriate to the decision, nondiscrimination testing, human accountability, accessible alternatives for people who opt out, legal review, and safeguards against institutional pressure to accept automated recommendations.
- Agency-aware automation allocation — Enterprise / robotics / software
- Create policies that dynamically decide whether the agent should execute independently, request confirmation, teach the user, provide a recommendation, or defer to a human.
- The allocation could depend on risk, reversibility, user expertise, fatigue, current goals, and whether repeated delegation is reducing the user’s capability or understanding.
- Assumptions and dependencies: measurable models of agency and capability, reliable risk estimation, organizational acceptance of slower but safer workflows, and mechanisms to prevent systems from manipulating users into granting more authority.
- Population-level prevention and public-health planning — Public health / policy
- Aggregate privacy-preserving trajectory signals to identify emerging barriers to medication adherence, access to care, isolation, or unhealthy routines and target voluntary support services.
- The framework could help policymakers compare interventions by long-term human benefit rather than engagement or short-term compliance.
- Assumptions and dependencies: strong anonymization, representative data, safeguards against surveillance and discriminatory profiling, transparent governance, and evidence that population-level patterns do not justify harmful individual-level inference.
- Formal benchmarks for human gain and agency preservation — Academia / standards bodies
- Establish benchmarks that evaluate whether an agent improves understanding, independent performance, self-efficacy, judgment, relationships, and wellbeing while reducing overreliance, dependence, and unwanted automation.
- Benchmark tasks should include longitudinal scenarios, intervention-response histories, uncertain observations, user corrections, refusal, non-intervention, and escalation.
- Assumptions and dependencies: consensus on constructs and metrics, costly longitudinal testing, culturally diverse samples, protection of participants, and resistance to reducing complex human outcomes to a single score.
- Human Digital Twin integration without full-person replication — Healthcare / industrial and research systems
- Connect domain-specific digital twins—for example, physiological, rehabilitation, or learning models—to a Combodied Agent that governs memory, consent, intervention, and user control.
- This would allow specialized predictive models to contribute evidence without granting them authority over the person or treating them as complete representations of human identity.
- Assumptions and dependencies: interoperable model interfaces, clear separation between prediction and decision authority, uncertainty calibration, continuous validation, and governance preventing domain models from being reused outside their agreed purpose.
- Personal agents as long-term capability partners — Daily life / social infrastructure
- Mature systems could support a person across changing life stages, helping coordinate health, work, learning, relationships, and personal goals while adapting their role over time.
- The desired outcome would be sustained human benefit—greater understanding, capability, autonomy, and wellbeing—not maximum automation, engagement, or dependence.
- Assumptions and dependencies: trustworthy longitudinal modeling, user ownership of data and goals, continuity across providers and devices, strong relationship-safety mechanisms, equitable access, and evidence that long-term use strengthens rather than displaces human relationships and judgment.
Glossary
- Admissible intervention policy: A decision policy that selects only interventions permitted by consent, safety, uncertainty, reversibility, and user-control constraints. “an admissible intervention policy selects proportionate support under constraints of consent, uncertainty, safety, reversibility, and user control.”
- Agentic AI: Artificial intelligence designed to interpret situations, pursue goals, plan, act, and adapt using feedback. “This ambiguity exposes a structural gap in Agentic AI.”
- Action substrate: The class of states that primarily organizes an agent’s modeling, decision-making, intervention, and evaluation. “We call this class the agent's action substrate.”
- Agency preservation: Designing support so that a person’s autonomy, control, dignity, relationships, and long-term capabilities are maintained or strengthened. “Agency preservation: support must preserve or strengthen autonomy, control, dignity, relationships, and long-term human capability.”
- Ambient acquisition: Collection of information from a surrounding environment without requiring direct user interaction. “We refer to these configurations as user-reported, device-mediated, on-body or contact, ambient or contactless, and institutional acquisition.”
- Autonomous laboratory: A laboratory system that combines automated digital planning with physical experimentation or execution. “Autonomous laboratories, for example, couple digital planning with physical execution.”
- Calibration: The degree to which a model’s confidence corresponds to the actual reliability of its predictions. “Personal World Models transform longitudinal event evidence and current context into calibrated distributions over future personal states.”
- Causal intervention modeling: Modeling how a particular action or intervention causes changes in a person or system over time. “Lack causal intervention modeling.”
- Closed loop: A system architecture in which actions produce feedback that updates later perception, memory, prediction, and action. “We synthesize these capabilities into a unified closed loop.”
- Co-agency: A collaborative arrangement in which an agent and a human share responsibilities rather than treating human replacement as the default. “Co-agency: the agent works with the human and adapts the division of labor instead of treating human replacement as the default.”
- Context-triggered acquisition: Data collection initiated when a specified environmental or situational context occurs. “Their sampling patterns range from continuous and periodic sensing to episodic, event-triggered, context-triggered, user-initiated, and query-on-demand collection.”
- Correctable memory: A memory system whose stored information can be inspected, revised, or removed by the user. “Longitudinal and correctable memory provides temporal context.”
- Digital state: The software representations, files, interfaces, or other computational conditions on which a Digital Agent operates. “Digital Agents are primarily organized around transformations of digital states.”
- Embodied Agent: An agent that primarily perceives and transforms physical or simulated physical environments through movement and manipulation. “Embodied Agents connect language and multimodal perception to navigation, manipulation, and physical control.”
- Episodic memory: Memory for specific events, interactions, or experiences rather than general facts. “Episodic memory & Records concrete events, interactions, and user experiences.”
- Event-evidence record: A structured representation of an observed event, including its timing, source, interpretation, uncertainty, and relevance. “only then can they become event-evidence records describing what was observed, when and how it was acquired.”
- Event-based multimodal perception: The process of filtering and interpreting data from multiple modalities to identify meaningful personal events. “Event-based multimodal perception reconstructs evidence about meaningful personal events.”
- Exogenous influence: An external factor that affects a person’s state but is not directly controlled by the agent or user. “It also depends on the person's own actions, contextual changes, and exogenous influences.”
- Human Digital Twin (HDT): A dynamically updated digital representation of a person or a specific aspect of that person, synchronized with ongoing data. “Combodied Agents share with Human Digital Twins (HDTs) an individual-centered and longitudinal perspective.”
- Human-state perception: Estimation of a person’s current condition from incomplete, noisy, multimodal evidence. “Human-State Perception estimates current personal state from incomplete, noisy, and multimodal evidence.”
- Intervention-response memory: Memory that links an intervention to the person’s subsequent acceptance, rejection, benefit, or harm. “Intervention-response memory supplies evidence for adapting future support rather than merely repeating a previously generated action.”
- Latent human state: The unobservable underlying condition of a person inferred from available observations. “Let denote the unobservable human state at time .”
- Longitudinal memory: Memory that organizes information about a person across extended periods and repeated interactions. “Longitudinal Memory organizes events, goals, relationships, interventions, outcomes, and corrections across time.”
- Multimodal foundation model: A broadly trained model capable of processing and relating multiple data modalities, such as text, images, audio, or video. “LLMs and multimodal foundation models have expanded this pattern into general-purpose agents.”
- Multimodal perception: Interpretation of information obtained through multiple sensory or data modalities. “Combodied Agents require multimodal perception.”
- On-device inference: Performing model computation locally on a device rather than sending data to a remote server. “Edge AI Agents” list “On-device inference.”
- Paralinguistic information: Non-lexical features of speech, such as pitch, rhythm, pauses, intensity, and voice quality. “Speech restores temporal and paralinguistic information through pitch, intensity, rhythm, pauses, hesitation, articulation, fluency, and voice quality.”
- Personal World Model (PWM): A person-specific model that predicts future personal states, events, and outcomes under alternative actions or interventions. “Personal World Models transform longitudinal event evidence and current context into calibrated distributions over future personal states.”
- Posterior representation: A probability-based representation of a hidden state after incorporating observed evidence. “The agent maintains an uncertainty-bearing posterior representation and retrieves decision-relevant evidence.”
- Provenance: Information documenting the origin, acquisition method, and history of a piece of data or memory. “Memory writing, consolidation, retrieval, and correction must preserve provenance and uncertainty.”
- Reversibility: The extent to which an intervention or action can be undone or withdrawn without lasting consequences. “the policy chooses only among actions admissible under consent, stated boundaries, safety constraints, uncertainty thresholds, reversibility, and escalation requirements.”
- Sampling pattern: The temporal manner in which data are collected, such as continuously, periodically, episodically, or on demand. “Their sampling patterns range from continuous and periodic sensing to episodic, event-triggered, context-triggered, user-initiated, and query-on-demand collection.”
- Semantic person memory: Memory containing relatively stable facts, preferences, values, and boundaries about an individual. “Semantic person memory & Stores relatively stable facts, preferences, values, and boundaries.”
- Speaker diarization: Identification and segmentation of which speaker is talking during different portions of an audio recording. “Speech research has traditionally separated voice activity detection, speaker diarization, speech recognition, acoustic feature extraction, and environmental sound detection.”
- State estimator: A computational component that infers an underlying state from observations and evidence. “The new evidence updates memory and may subsequently revise the state estimator, PWM, or Intervention Policy.”
- Trajectory memory: Memory that records changes in health, behavior, emotion, cognition, or routines over time. “Trajectory memory & Tracks changes in health, behavior, emotion, cognition, and routines.”
- Uncertainty-bearing representation: A representation that explicitly records doubt or confidence about its inferred content. “The agent maintains an uncertainty-bearing posterior representation and retrieves decision-relevant evidence.”
- User-contestable adaptation: Personalization that the user can challenge, reject, or correct. “Limited user-contestable adaptation.”
- Voice activity detection: Identification of portions of an audio signal that contain human speech. “Speech research has traditionally separated voice activity detection, speaker diarization, speech recognition, acoustic feature extraction, and environmental sound detection.”




