---
title: 'A2A Agent Cards: Protocol and Benchmarking'
url: https://www.emergentmind.com/topics/a2a-agent-cards
type: topic
---

# A2A Agent Cards: Protocol and Benchmarking

Searching arXiv for recent papers on A2A Agent Cards and related A2A protocol work.
A2A Agent Cards are machine-readable descriptors through which software agents advertise identity, capabilities, endpoints, and invocation requirements to other agents. In the A2A literature, they are typically JSON manifests published at a well-known URI for discovery, authentication, and task delegation; in a distinct but related usage, "AgentCARD" denotes a role-aware benchmark suite for evaluating multi-role LLM agent teams across model assignments and deployment modes [2505.02279] [2606.20629].

## 1. Definition and publication model

In implementation-oriented accounts of A2A, an Agent Card is a JSON metadata document that functions as a machine-readable "business card" for an agent. It describes the agent’s identity, network endpoint, supported RPC methods, authentication schemes, and capability catalog, allowing a client agent to discover a remote agent, determine how to authenticate, and dynamically construct task payloads for delegation [2504.16902]. A closely related description in A2A v0.2 characterizes the Agent Card as a capability descriptor that advertises identity, metadata, supported modalities through `inputModes` and `outputModes`, services or skills, and any structured data APIs; routers and orchestration layers fetch the card at runtime to decide how to encode and route each message part [2604.12213].

The literature is not uniform about the publication path. Several sources place the manifest at `/.well-known/agent.json`, while the modality-native routing work for A2A v0.2 uses `GET /.well-known/agent-card.json` [2505.02279] [2604.12213]. This suggests that the well-known publication pattern is stable, while the exact filename is still evolving across drafts, implementations, and protocol-adjacent reconstructions.

Some systems present the card not only as a passive manifest but also as part of active registration. In AgentMaster, the so-called A2A agent card is reconstructed as a standardized metadata manifest published on startup or registration, delivered in a JSON-RPC `register_agent` call or as part of an A2A "Hello" handshake. In that formulation, the card supports both discovery and registration by the orchestrator and runtime routing through capability matching [2507.21105].

## 2. Schema families and representational variants

Across papers, A2A Agent Cards share a common structural core: agent identity, descriptive metadata, endpoint location, declared capabilities or skills, and invocation constraints such as authentication, modality support, or schema contracts. The main differences arise from the degree of specialization: some schemas remain minimal and interface-oriented, while others add trust, payment, reasoning-profile, or lineage fields.

| Line of work | Representative fields | Emphasis |
|---|---|---|
| Baseline A2A manifests | `id`, `name`, `description`, `inputModes`, `outputModes`, `services`, `skills`, `metadata` | capability declaration and modality support |
| Security- and API-oriented variants | `a2aEndpointUrl`, `securitySchemes`, `capabilities`, `pushNotifications`, `publicKey`, `signature` | authentication, verification, invocation contract |
| Rich identity or economic extensions | `trust_domain`, `reasoning_profile`, `cost_profile`, `x402PaymentInfo`, `identity`, `lineage_support` | routing quality, trust, payments, provenance |

The A2A v0.2-style schema emphasizes media handling and service discovery. Its representative fields are `id`, `name`, `description`, `inputModes`, `outputModes`, `services`, `skills`, and `metadata`, where `inputModes` and `outputModes` list MIME types and `services` declare callable methods such as `tasks/send` and `tasks/sendSubscribe` [2604.12213].

Security- and interface-oriented formulations are more expansive. One secure-deployment account defines fields such as `id`, `name`, `version`, `provider`, `a2aEndpointUrl`, `securitySchemes`, `capabilities`, `capabilitiesTags`, `pushNotifications`, `hostedAt`, `metadata`, `publicKey`, and `signature`, explicitly tying the card to OpenAPI SecuritySchemes, JSON-Schema parameter blocks, and digital signatures [2504.16902]. A survey-oriented A2A schema similarly uses `id`, `name`, `version`, `skills`, `endpoints`, `authentication`, `metadata`, and `signature`, with skill-level `inputSchema`, `outputSchema`, and `scope` for capability-based access control [2505.02279].

AgentMaster reconstructs an operational registration schema with `agent_id`, `agent_name`, `version`, `description`, `capabilities`, `input_schema`, `output_schema`, `access_rights`, `endpoint`, `heartbeat_interval`, and `metadata`, reflecting orchestration and failover needs rather than a canonical protocol specification [2507.21105].

A more ambitious extension appears in LDP, which introduces rich delegate identity cards with first-class model-level properties: `agent_id`, `principal_id`, `model_family`, `model_name`, `model_version`, `runtime_version`, `weights_fingerprint`, `endpoint_address`, `trust_domain`, `public_key`, `jurisdiction`, `data_handling_policy`, `context_window`, `modalities_supported`, `languages_supported`, `tokenizer_fingerprint`, capability objects with `quality_hint`, `latency_hint_ms_p50`, and `cost_hint`, plus `reasoning_profile` and `cost_profile` [2603.08852]. Another extension adds `x402PaymentInfo` with `asset`, `network`, `amount`, `payTo`, `nonce`, and `validUntil` for micropayments [2507.19550]. For regulated environments, an augmented identity block can embed `agent_id`, `public_key`, `identity_proof`, and `lineage_support` flags for Merkle proof generation and DPoP binding [2509.18415].

## 3. Runtime semantics: discovery, routing, and delegation

At runtime, Agent Cards are consulted to decide whether a peer should be contacted and how a payload should be transformed before transmission. In baseline A2A operation, client agents bootstrap peer-to-peer task delegation by fetching candidate remote agents’ cards, inspecting skills and schemas, obtaining the required credentials, and then invoking JSON-RPC methods over HTTPS; long-running tasks may use Server-Sent Events for status streaming and optional push-notification endpoints for callbacks [2505.02279].

AgentMaster presents a card-driven orchestration loop in which the orchestrator consults a registry of agent cards at task decomposition time, selects an agent whose `capabilities` array matches a sub-task label, opens a JSON-RPC call to the agent’s `endpoint`, and, if an agent fails to reply within `heartbeat_interval` or returns an error, fails over to another agent with overlapping capabilities based on `metadata.load_weight`. Returned partial outputs are expected to conform to the agent’s `output_schema` before synthesis into a final response [2507.21105].

In multimodal A2A, the card becomes a routing primitive. MMA2A inserts a Modality-Aware Router between the orchestrator and destination agents; it fetches or caches the destination card, inspects `inputModes`, and applies the routing rule
$$
\operatorname{route}(p,\mathcal{A}_j)=
\begin{cases}
p & \text{if } m \in \operatorname{cap}(\mathcal{A}_j)\\
\operatorname{transcode}(p,t) & \text{otherwise.}
\end{cases}
$$
Here $m \in \{v,i,t\}$ denotes voice, image, or text, and the alternative path transcodes non-supported parts to text through speech-to-text or image captioning [2604.12213].

The empirical effect of card-driven modality routing is substantial when downstream reasoning is capable of exploiting the preserved signal. On the 50-task CrossModal-CS benchmark, MMA2A achieves $52.0\%$ task completion accuracy versus $32.0\%$ for a text-bottleneck baseline, with a native routing rate of $81.7\%$ versus $50.5\%$. The improvement is accompanied by higher end-to-end latency: $13.04\,\mathrm{s}$ versus $7.19\,\mathrm{s}$, or approximately $1.8\times$ slower [2604.12213]. The same study reports that an ablation replacing LLM-backed reasoning with keyword matching removes the accuracy gap entirely ($36\%$ versus $36\%$), indicating that protocol-level routing and agent-level reasoning are jointly necessary [2604.12213].

A related extension generalizes card-driven routing from static capabilities to identity-aware specialization. LDP’s rich delegate identity cards add quality hints, reasoning profiles, cost characteristics, progressive payload negotiation, governed sessions, structured provenance, and trust domains. In that setting, identity-aware routing achieves approximately $12\times$ lower latency on easy tasks through delegate specialization; semantic frame payloads reduce token count by $37\%$ with $p=0.031$ and no observed quality loss; and governed sessions eliminate $39\%$ token overhead at $10$ rounds [2603.08852].

## 4. Security model, trust assumptions, and failure modes

The security literature treats Agent Cards as both an interface contract and a security surface. One secure A2A account states that Agent Cards **must** be signed by the agent’s private key to guarantee authenticity and integrity, with signature generation formalized as
$$
\sigma = \mathrm{Sign}_{sk}\bigl(H(\mathrm{AgentCard}_{\text{content}})\bigr),
$$
and verification performed client-side against the corresponding public key. The same account recommends ECDSA P-256 (ES256) or RSA-2048/3072, SHA-256 or better, HTTPS with TLS 1.3, and optionally mTLS [2504.16902].

Operationally, A2A deployments frequently layer OAuth 2.0 bearer tokens, JWT validation, JSON-RPC 2.0 over HTTPS, and SSE keep-alives for status streaming. A protocol-centric threat analysis identifies risks specific to the creation, operation, and update phases, including long-lived JWT replay, insufficiently granular token scopes, unsigned Agent Cards, lack of nonce or timestamp replay protection in JSON-RPC calls, injection into JSON payloads, cross-vendor trust-boundary exploitation, shadowing attacks, collusion or free-riding in sub-workflows, context explosion through unbounded streaming, semantic drift, supply-chain compromise, and downgrade under protocol fragmentation. Its qualitative risk formula is
$$
R = L \times I,
$$
where $L$ is likelihood and $I$ is impact [2602.11327].

Threat-modeling work centered on A2A emphasizes two card-specific attack classes. **Agent Card Spoofing** occurs when an attacker hosts a fake card at a lookalike domain, inducing clients to bind to a malicious agent; recommended mitigations include HTTPS with certificate validation and pinning, a trusted Agent Card registry or DNSSEC, and signature validation against a known CA hierarchy or DID registry. **Poisoned Agent Card** refers to malicious prompts or shell instructions embedded in capability descriptions; recommended mitigations are sanitization of all textual fields, strict JSON-Schema validation, and content-security controls when rendering in a browser [2504.16902].

A recurring misconception is that integrity of the manifest implies truthfulness of the manifest. Comparative trust-model work explicitly argues otherwise: AgentCards embody the **Claim** model. Fields such as capabilities and contact are self-described claims; TLS or OAuth can ensure the JSON came from the claimed endpoint, but do not guarantee that the advertised capabilities are accurate. No collateral, stake, or in-protocol reputation is required at issuance, and the design is therefore suitable for enterprise and known-partner settings but brittle in fully open networks [2511.03434].

## 5. Decentralized identity, micropayments, and lineage assurance

Several recent extensions attempt to move Agent Cards beyond a lightweight discovery layer toward stronger guarantees about identity, payment, and provenance. One architecture anchors AgentCards on-chain as smart contracts, storing fields such as `agentID`, `publicKey`, `endpoint`, `capabilities`, `skills`, `defaultInputModes`, `defaultOutputModes`, `metadata`, and optional `x402PaymentInfo`. Integrity is modeled through a stored hash
$$
H = \mathrm{sha256}(\mathrm{serialize}(\mathrm{AgentCard})),
$$
an off-chain signature $\sigma = \mathrm{Sign}_{SK_{agent}}(H)$, and verification through `ecrecover(σ, H) == agentID` [2507.19550].

That architecture also specifies the on-chain lifecycle. Registration uses `registerAgent(...)`, update uses `updateAgent(...)`, and revocation uses `revokeAgent()`, with reported gas usage of approximately $350\,000$ gas for registry deployment, $200\,000$ gas for full `registerAgent`, $50\,000$ gas for partial `updateAgent`, and $30\,000$ gas for `revokeAgent`. On Ethereum Sepolia, the paper reports average block time of approximately $12\,\mathrm{s}$, roughly $15$ tx/sec network throughput, about $12\,\mathrm{s}$ for one confirmation, and about $2.4\,\mathrm{min}$ for $12$-block finality [2507.19550].

The same work extends A2A with x402 micropayments. A client initially issues an A2A POST without `X-PAYMENT`, receives `402 Payment Required` together with payment metadata, signs an EIP-3009 `transferWithAuthorization`, and retries with `X-PAYMENT: base64(authorization)`. The reported overhead is about one extra HTTP round trip, while an on-chain ERC-20 `transferWithAuthorization` costs approximately $100\,\mathrm{k}$ gas, or about `~ $0.05 at 50 gwei`; off-chain incremental payment channels are proposed to batch many micropayments before final settlement [2507.19550].

For critical multi-agent systems, another line of work augments the Agent Card with a cryptographically grounded identity block and append-only lineage support. The identity block includes
`"agent_id": "aid://sha256(pubkey∥domain∥timestamp)"`,
`"public_key": "ed25519:<base64-pk>"`,
`"identity_proof": "ed25519:<base64-sig>"`,
and `lineage_support` flags. The semantics are explicit: `agent_id = aid://SHA-256(public_key ∥ provider.domain ∥ timestamp)` and `identity_proof = \mathrm{Sign}_{priv}(agent_id ∥ skills)` [2509.18415].

Lineage is then anchored in Merkle trees modeled after Certificate Transparency logs. Leaves are hashed as $L_i = H(0x00 \parallel c_{e_i})$, internal nodes as $N = H(0x01 \parallel L \parallel R)$, and the Lineage Store emits a Signed Tree Head
$$
\mathrm{STH}_n = \mathrm{Sign}_{LS}(n, R_n, wallclock_t, monotonic_ctr, log_id).
$$
Inclusion proofs have $O(\log n)$ audit paths, and multiproof size is reported as $O(\log n + k \cdot \log(n/k))$ for $k$ leaves in a tree of size $n$ [2509.18415]. A federated Proof Server can aggregate inclusion and consistency proofs into a signed attestation, allowing external verifiers to validate multi-hop provenance without access to the full execution trace [2509.18415]. This suggests an evolution from point-to-point authenticity toward verifiable chain-of-custody for non-human identities.

## 6. AgentCARD as a role-aware benchmark interpretation

A distinct usage of the term appears in the paper "Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams," which introduces AgentCARD as a benchmark suite for multi-role LLM agent teams rather than as a manifest format [2606.20629]. The framework models a benchmark domain as $\mathcal{D}=\{1,\ldots,N\}$, roles as $\mathcal{R}=\{P,E,V,\ldots\}$, a pool of candidate LLMs as $M$, and deployment mode as $d \in \{\mathrm{API},\mathrm{SH}\}$. A team configuration is
$$
\theta = ((m_r,d_r): r \in \mathcal{R}),
$$
and the harness executes a planner, an optional verifier, and an executor. Task accuracy is binary,
$$
a_j(\theta)=1 \text{ if } \hat{y}_j \text{ matches gold answer, else } 0,
$$
with aggregate accuracy
$$
\mathrm{Acc}(\theta)=\frac{1}{|\mathcal{D}|}\sum_{j \in \mathcal{D}} a_j(\theta).
$$

AgentCARD couples this role-decomposed harness to a unified API/self-hosted cost model. Per-task cost is
$$
C_{\text{task}}(\theta)=\sum_{r \in \mathcal{R}} C_r(\theta).
$$
For API-deployed roles,
$$
C_r = p_{in}\cdot N_{in}^{fresh} + p_{cache}\cdot N_{in}^{hit} + p_{out}\cdot N_{out},
$$
and for self-hosted roles,
$$
C_r = \frac{N_{GPU,r}\cdot p_{GPU/hr}\cdot T_{wall}}{3600 \cdot N_{tasks}},
$$
with an amortized GPU-hour price of `\$1.87/hr for H100` [2606.20629]. The framework then places each configuration as a point $(C_{\text{task}}(\theta), \mathrm{Acc}(\theta))$ in cost-accuracy space and extracts the Pareto frontier by dominance testing.

To identify bottlenecks, AgentCARD uses Shapley-value diagnostics over weak-to-strong planner and executor upgrades:
$$
\phi_P = \tfrac{1}{2}[(a_{SW}-a_{WW}) + (a_{SS}-a_{WS})], \qquad
\phi_E = \tfrac{1}{2}[(a_{WS}-a_{WW}) + (a_{SS}-a_{SW})].
$$
The larger value identifies the more critical role [2606.20629].

The reported empirical findings are highly specific. On SWEBench-Lite, pairing GPT-5.4 as planner with MiniMax-M2.7 as executor yields `Acc=76.2 %` versus a budget-mixed homogeneous baseline of approximately `32.2 %`, a real synergy of `Δ_acc=+44.0 %`; reversing the same pair loses `8.1 %`. On IMO-AnswerBench, the hybrid team Claude-Opus-4.6→Qwen3.5-27B matches Claude-only accuracy at `80.6 %` but at `\$0.108/task` versus `\$1.283/task`, a `12×` reduction. The bottleneck structure is domain-dependent: MedAgentBench is planner-bottlenecked with `φ_P≈29 %` and `φ_E≈3 %`, whereas FinanceBench, IMO-AnswerBench, and MCP-Atlas are executor-bottlenecked with `φ_E−φ_P` ranging from `+12 %` to `+34 %`. In a three-role extension on MCP-Atlas, adding a verifier raises accuracy from `80.4 %` to `83.8 %` with a heterogeneous team `(GLM-5.1 planner, GPT-5.4 verifier, MiniMax-M2.7 executor)` at slightly lower cost, and the verifier helps only when it is "stronger" at the planning role than the original planner [2606.20629].

Taken together, the two literatures place "A2A Agent Cards" at different layers of the agent stack. In the A2A protocol literature, the card is chiefly a discovery and invocation artifact; in AgentCARD, the term denotes a bottom-up evaluation framework centered on role assignment, deployment mode, cost-accuracy frontiers, and role decomposition [2606.20629]. The coexistence of these usages underscores that the "card" abstraction in agent systems now spans interface description, routing, security, economic coordination, and performance evaluation.

Source: https://www.emergentmind.com/topics/a2a-agent-cards