Papers
Topics
Authors
Recent
Search
2000 character limit reached

A2A Agent Cards: Protocol and Benchmarking

Updated 8 July 2026
  • A2A Agent Cards are JSON manifests that serve as machine-readable business cards, detailing agent identity, network endpoints, and invocation protocols for secure discovery and delegation.
  • They span various schema families—from baseline capability declarations to extended formats with security, trust, and economic attributes—catering to different operational needs.
  • Empirical studies demonstrate that card-driven routing and benchmark evaluations enhance multi-agent task accuracy while balancing latency, cost, and performance trade-offs.

Searching arXiv for papers on A2A Agent Cards and related A2A protocol work. A2A Agent Cards are machine-readable descriptors through which software agents advertise identity, capabilities, endpoints, and invocation requirements to other agents. In the A2A literature, they are typically JSON manifests published at a well-known URI for discovery, authentication, and task delegation; in a distinct but related usage, "AgentCARD" denotes a role-aware benchmark suite for evaluating multi-role LLM agent teams across model assignments and deployment modes (Ehtesham et al., 4 May 2025, Jiang et al., 28 May 2026).

1. Definition and publication model

In implementation-oriented accounts of A2A, an Agent Card is a JSON metadata document that functions as a machine-readable "business card" for an agent. It describes the agent’s identity, network endpoint, supported RPC methods, authentication schemes, and capability catalog, allowing a client agent to discover a remote agent, determine how to authenticate, and dynamically construct task payloads for delegation (Habler et al., 23 Apr 2025). A closely related description in A2A v0.2 characterizes the Agent Card as a capability descriptor that advertises identity, metadata, supported modalities through inputModes and outputModes, services or skills, and any structured data APIs; routers and orchestration layers fetch the card at runtime to decide how to encode and route each message part (Srinivasan, 14 Apr 2026).

The literature is not uniform about the publication path. Several sources place the manifest at /.well-known/agent.json, while the modality-native routing work for A2A v0.2 uses GET /.well-known/agent-card.json (Ehtesham et al., 4 May 2025, Srinivasan, 14 Apr 2026). This suggests that the well-known publication pattern is stable, while the exact filename is still evolving across drafts, implementations, and protocol-adjacent reconstructions.

Some systems present the card not only as a passive manifest but also as part of active registration. In AgentMaster, the so-called A2A agent card is reconstructed as a standardized metadata manifest published on startup or registration, delivered in a JSON-RPC register_agent call or as part of an A2A "Hello" handshake. In that formulation, the card supports both discovery and registration by the orchestrator and runtime routing through capability matching (Liao et al., 8 Jul 2025).

2. Schema families and representational variants

Across papers, A2A Agent Cards share a common structural core: agent identity, descriptive metadata, endpoint location, declared capabilities or skills, and invocation constraints such as authentication, modality support, or schema contracts. The main differences arise from the degree of specialization: some schemas remain minimal and interface-oriented, while others add trust, payment, reasoning-profile, or lineage fields.

Line of work Representative fields Emphasis
Baseline A2A manifests id, name, description, inputModes, outputModes, services, skills, metadata capability declaration and modality support
Security- and API-oriented variants a2aEndpointUrl, securitySchemes, capabilities, pushNotifications, publicKey, signature authentication, verification, invocation contract
Rich identity or economic extensions trust_domain, reasoning_profile, cost_profile, x402PaymentInfo, identity, lineage_support routing quality, trust, payments, provenance

The A2A v0.2-style schema emphasizes media handling and service discovery. Its representative fields are id, name, description, inputModes, outputModes, services, skills, and metadata, where inputModes and outputModes list MIME types and services declare callable methods such as tasks/send and tasks/sendSubscribe (Srinivasan, 14 Apr 2026).

Security- and interface-oriented formulations are more expansive. One secure-deployment account defines fields such as id, name, version, provider, a2aEndpointUrl, securitySchemes, capabilities, capabilitiesTags, pushNotifications, hostedAt, metadata, publicKey, and signature, explicitly tying the card to OpenAPI SecuritySchemes, JSON-Schema parameter blocks, and digital signatures (Habler et al., 23 Apr 2025). A survey-oriented A2A schema similarly uses id, name, version, skills, endpoints, authentication, metadata, and signature, with skill-level inputSchema, outputSchema, and scope for capability-based access control (Ehtesham et al., 4 May 2025).

AgentMaster reconstructs an operational registration schema with agent_id, agent_name, version, description, capabilities, input_schema, output_schema, access_rights, endpoint, heartbeat_interval, and metadata, reflecting orchestration and failover needs rather than a canonical protocol specification (Liao et al., 8 Jul 2025).

A more ambitious extension appears in LDP, which introduces rich delegate identity cards with first-class model-level properties: agent_id, principal_id, model_family, model_name, model_version, runtime_version, weights_fingerprint, endpoint_address, trust_domain, public_key, jurisdiction, data_handling_policy, context_window, modalities_supported, languages_supported, tokenizer_fingerprint, capability objects with quality_hint, latency_hint_ms_p50, and cost_hint, plus reasoning_profile and cost_profile (Prakash, 9 Mar 2026). Another extension adds x402PaymentInfo with asset, network, amount, payTo, nonce, and validUntil for micropayments (Vaziry et al., 24 Jul 2025). For regulated environments, an augmented identity block can embed agent_id, public_key, identity_proof, and lineage_support flags for Merkle proof generation and DPoP binding (Malkapuram et al., 22 Sep 2025).

3. Runtime semantics: discovery, routing, and delegation

At runtime, Agent Cards are consulted to decide whether a peer should be contacted and how a payload should be transformed before transmission. In baseline A2A operation, client agents bootstrap peer-to-peer task delegation by fetching candidate remote agents’ cards, inspecting skills and schemas, obtaining the required credentials, and then invoking JSON-RPC methods over HTTPS; long-running tasks may use Server-Sent Events for status streaming and optional push-notification endpoints for callbacks (Ehtesham et al., 4 May 2025).

AgentMaster presents a card-driven orchestration loop in which the orchestrator consults a registry of agent cards at task decomposition time, selects an agent whose capabilities array matches a sub-task label, opens a JSON-RPC call to the agent’s endpoint, and, if an agent fails to reply within heartbeat_interval or returns an error, fails over to another agent with overlapping capabilities based on metadata.load_weight. Returned partial outputs are expected to conform to the agent’s output_schema before synthesis into a final response (Liao et al., 8 Jul 2025).

In multimodal A2A, the card becomes a routing primitive. MMA2A inserts a Modality-Aware Router between the orchestrator and destination agents; it fetches or caches the destination card, inspects inputModes, and applies the routing rule

route⁡(p,Aj)={pif m∈cap⁡(Aj) transcode⁡(p,t)otherwise.\operatorname{route}(p,\mathcal{A}_j)= \begin{cases} p & \text{if } m \in \operatorname{cap}(\mathcal{A}_j)\ \operatorname{transcode}(p,t) & \text{otherwise.} \end{cases}

Here m∈{v,i,t}m \in \{v,i,t\} denotes voice, image, or text, and the alternative path transcodes non-supported parts to text through speech-to-text or image captioning (Srinivasan, 14 Apr 2026).

The empirical effect of card-driven modality routing is substantial when downstream reasoning is capable of exploiting the preserved signal. On the 50-task CrossModal-CS benchmark, MMA2A achieves 52.0%52.0\% task completion accuracy versus 32.0%32.0\% for a text-bottleneck baseline, with a native routing rate of 81.7%81.7\% versus 50.5%50.5\%. The improvement is accompanied by higher end-to-end latency: 13.04 s13.04\,\mathrm{s} versus 7.19 s7.19\,\mathrm{s}, or approximately 1.8×1.8\times slower (Srinivasan, 14 Apr 2026). The same study reports that an ablation replacing LLM-backed reasoning with keyword matching removes the accuracy gap entirely (36%36\% versus m∈{v,i,t}m \in \{v,i,t\}0), indicating that protocol-level routing and agent-level reasoning are jointly necessary (Srinivasan, 14 Apr 2026).

A related extension generalizes card-driven routing from static capabilities to identity-aware specialization. LDP’s rich delegate identity cards add quality hints, reasoning profiles, cost characteristics, progressive payload negotiation, governed sessions, structured provenance, and trust domains. In that setting, identity-aware routing achieves approximately m∈{v,i,t}m \in \{v,i,t\}1 lower latency on easy tasks through delegate specialization; semantic frame payloads reduce token count by m∈{v,i,t}m \in \{v,i,t\}2 with m∈{v,i,t}m \in \{v,i,t\}3 and no observed quality loss; and governed sessions eliminate m∈{v,i,t}m \in \{v,i,t\}4 token overhead at m∈{v,i,t}m \in \{v,i,t\}5 rounds (Prakash, 9 Mar 2026).

4. Security model, trust assumptions, and failure modes

The security literature treats Agent Cards as both an interface contract and a security surface. One secure A2A account states that Agent Cards must be signed by the agent’s private key to guarantee authenticity and integrity, with signature generation formalized as

m∈{v,i,t}m \in \{v,i,t\}6

and verification performed client-side against the corresponding public key. The same account recommends ECDSA P-256 (ES256) or RSA-2048/3072, SHA-256 or better, HTTPS with TLS 1.3, and optionally mTLS (Habler et al., 23 Apr 2025).

Operationally, A2A deployments frequently layer OAuth 2.0 bearer tokens, JWT validation, JSON-RPC 2.0 over HTTPS, and SSE keep-alives for status streaming. A protocol-centric threat analysis identifies risks specific to the creation, operation, and update phases, including long-lived JWT replay, insufficiently granular token scopes, unsigned Agent Cards, lack of nonce or timestamp replay protection in JSON-RPC calls, injection into JSON payloads, cross-vendor trust-boundary exploitation, shadowing attacks, collusion or free-riding in sub-workflows, context explosion through unbounded streaming, semantic drift, supply-chain compromise, and downgrade under protocol fragmentation. Its qualitative risk formula is

m∈{v,i,t}m \in \{v,i,t\}7

where m∈{v,i,t}m \in \{v,i,t\}8 is likelihood and m∈{v,i,t}m \in \{v,i,t\}9 is impact (Anbiaee et al., 11 Feb 2026).

Threat-modeling work centered on A2A emphasizes two card-specific attack classes. Agent Card Spoofing occurs when an attacker hosts a fake card at a lookalike domain, inducing clients to bind to a malicious agent; recommended mitigations include HTTPS with certificate validation and pinning, a trusted Agent Card registry or DNSSEC, and signature validation against a known CA hierarchy or DID registry. Poisoned Agent Card refers to malicious prompts or shell instructions embedded in capability descriptions; recommended mitigations are sanitization of all textual fields, strict JSON-Schema validation, and content-security controls when rendering in a browser (Habler et al., 23 Apr 2025).

A recurring misconception is that integrity of the manifest implies truthfulness of the manifest. Comparative trust-model work explicitly argues otherwise: AgentCards embody the Claim model. Fields such as capabilities and contact are self-described claims; TLS or OAuth can ensure the JSON came from the claimed endpoint, but do not guarantee that the advertised capabilities are accurate. No collateral, stake, or in-protocol reputation is required at issuance, and the design is therefore suitable for enterprise and known-partner settings but brittle in fully open networks (Hu et al., 5 Nov 2025).

5. Decentralized identity, micropayments, and lineage assurance

Several recent extensions attempt to move Agent Cards beyond a lightweight discovery layer toward stronger guarantees about identity, payment, and provenance. One architecture anchors AgentCards on-chain as smart contracts, storing fields such as agentID, publicKey, endpoint, capabilities, skills, defaultInputModes, defaultOutputModes, metadata, and optional x402PaymentInfo. Integrity is modeled through a stored hash

52.0%52.0\%0

an off-chain signature 52.0%52.0\%1, and verification through ecrecover(σ, H) == agentID (Vaziry et al., 24 Jul 2025).

That architecture also specifies the on-chain lifecycle. Registration uses registerAgent(...), update uses updateAgent(...), and revocation uses revokeAgent(), with reported gas usage of approximately 52.0%52.0\%2 gas for registry deployment, 52.0%52.0\%3 gas for full registerAgent, 52.0%52.0\%4 gas for partial updateAgent, and 52.0%52.0\%5 gas for revokeAgent. On Ethereum Sepolia, the paper reports average block time of approximately 52.0%52.0\%6, roughly 52.0%52.0\%7 tx/sec network throughput, about 52.0%52.0\%8 for one confirmation, and about 52.0%52.0\%9 for 32.0%32.0\%0-block finality (Vaziry et al., 24 Jul 2025).

The same work extends A2A with x402 micropayments. A client initially issues an A2A POST without X-PAYMENT, receives 402 Payment Required together with payment metadata, signs an EIP-3009 transferWithAuthorization, and retries with X-PAYMENT: base64(authorization). The reported overhead is about one extra HTTP round trip, while an on-chain ERC-20 transferWithAuthorization costs approximately 32.0%32.0\%1 gas, or about ~ $0.05 at 50 gwei</code>; off-chain incremental payment channels are proposed to batch many micropayments before final settlement (<a href="/papers/2507.19550" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">Vaziry et al., 24 Jul 2025</a>).</p> <p>For critical multi-agent systems, another line of work augments the Agent Card with a cryptographically grounded identity block and append-only lineage support. The identity block includes <code>&quot;agent_id&quot;: &quot;aid://sha256(pubkey∥domain∥timestamp)&quot;</code>, <code>&quot;public_key&quot;: &quot;ed25519:&lt;base64-pk&gt;&quot;</code>, <code>&quot;identity_proof&quot;: &quot;ed25519:&lt;base64-sig&gt;&quot;</code>, and <code>lineage_support</code> flags. The semantics are explicit: <code>agent_id = aid://SHA-256(public_key ∥ provider.domain ∥ timestamp)</code> and <code>identity_proof = \mathrm{Sign}_{priv}(agent_id ∥ skills)</code> (<a href="/papers/2509.18415" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">Malkapuram et al., 22 Sep 2025</a>).</p> <p>Lineage is then anchored in Merkle trees modeled after Certificate <a href="https://www.emergentmind.com/topics/transparency-logs" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">Transparency logs</a>. Leaves are hashed as $32.0\%$2, internal nodes as $32.0\%$3, and the Lineage Store emits a Signed Tree Head

$32.0\%$4

Inclusion proofs have $32.0\%$5 audit paths, and multiproof size is reported as $32.0\%$6 for $32.0\%$7 leaves in a tree of size $32.0\%$8 (Malkapuram et al., 22 Sep 2025). A federated Proof Server can aggregate inclusion and consistency proofs into a signed attestation, allowing external verifiers to validate multi-hop provenance without access to the full execution trace (Malkapuram et al., 22 Sep 2025). This suggests an evolution from point-to-point authenticity toward verifiable chain-of-custody for non-human identities.

6. AgentCARD as a role-aware benchmark interpretation

A distinct usage of the term appears in the paper "Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams," which introduces AgentCARD as a benchmark suite for multi-role LLM agent teams rather than as a manifest format (Jiang et al., 28 May 2026). The framework models a benchmark domain as $32.0\%$9, roles as $81.7\%$0, a pool of candidate LLMs as $81.7\%$1, and deployment mode as $81.7\%$2. A team configuration is

$81.7\%$3

and the harness executes a planner, an optional verifier, and an executor. Task accuracy is binary,

$81.7\%$4

with aggregate accuracy

$81.7\%$5

AgentCARD couples this role-decomposed harness to a unified API/self-hosted cost model. Per-task cost is

$81.7\%$6

For API-deployed roles,

$81.7\%$7

and for self-hosted roles,

$81.7\%$8</p> <p>with an amortized GPU-hour price of `\$81.7\%9(Ctask(θ),Acc(θ))9(C_{\text{task}}(\theta), \mathrm{Acc}(\theta)) in cost-accuracy space and extracts the Pareto frontier by dominance testing.

To identify bottlenecks, AgentCARD uses Shapley-value diagnostics over weak-to-strong planner and executor upgrades:

50.5%50.5\%0

The larger value identifies the more critical role (Jiang et al., 28 May 2026).

The reported empirical findings are highly specific. On SWEBench-Lite, pairing GPT-5.4 as planner with MiniMax-M2.7 as executor yields Acc=76.2 % versus a budget-mixed homogeneous baseline of approximately 32.2 %, a real synergy of Δ_acc=+44.0 %; reversing the same pair loses 8.1 %. On IMO-AnswerBench, the hybrid team Claude-Opus-4.6→Qwen3.5-27B matches Claude-only accuracy at 80.6 % but at \$50.5\%$11.283/task, a 12× reduction. The bottleneck structure is domain-dependent: MedAgentBench is planner-bottlenecked with φ_P≈29 % and φ_E≈3 %, whereas FinanceBench, IMO-AnswerBench, and MCP-Atlas are executor-bottlenecked with φ_E−φ_P ranging from +12 % to +34 %. In a three-role extension on MCP-Atlas, adding a verifier raises accuracy from 80.4 % to 83.8 % with a heterogeneous team (GLM-5.1 planner, GPT-5.4 verifier, MiniMax-M2.7 executor) at slightly lower cost, and the verifier helps only when it is "stronger" at the planning role than the original planner (Jiang et al., 28 May 2026).

Taken together, the two literatures place "A2A Agent Cards" at different layers of the agent stack. In the A2A protocol literature, the card is chiefly a discovery and invocation artifact; in AgentCARD, the term denotes a bottom-up evaluation framework centered on role assignment, deployment mode, cost-accuracy frontiers, and role decomposition (Jiang et al., 28 May 2026). The coexistence of these usages underscores that the "card" abstraction in agent systems now spans interface description, routing, security, economic coordination, and performance evaluation.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to A2A Agent Cards.