---
title: 'Perceived AI Capacity: Understanding Human Beliefs'
url: https://www.emergentmind.com/topics/perceived-ai-capacity
type: topic
---

# Perceived AI Capacity: Understanding Human Beliefs

Searching arXiv for the cited works and closely related papers on perceived AI capacity.
arXiv search query: "Perceived AI capacity mental models trust calibration AIQ AI literacy human AI collaboration"
Perceived AI capacity denotes the beliefs humans hold about what AI systems can do, how reliably they can do it, and how far those capacities extend across domains, time horizons, and social settings. Recent research approaches the construct through mental models of AI “capabilities and limitations,” through scenario-based expectations about whether specific AI outcomes will occur “in the next 10 years,” and through judgments about competence, trustworthiness, social intelligence, and usefulness in concrete workflows [2503.16438][2412.01459]. Taken together, these studies suggest that perceived AI capacity is not a single scalar variable, but a family of representations linking objective system performance, human expectations, and the quality of human–AI collaboration.

## 1. Conceptual foundations

A central distinction in the literature is between **objective AI capacity** and **perceived AI capacity**. In the AIQ framework, objective AI capacity is the “actual functional performance of the AI in a given domain and context,” whereas perceived AI capacity is the user’s “internal model of that performance—how powerful, accurate, flexible, risky, or constrained they believe the AI to be” [2503.16438]. This distinction is also used to separate perceived capacity from both traditional human IQ and standalone AI benchmarks such as accuracy, BLEU, and MMLU, which assess the system itself but “say nothing about whether humans can accurately perceive those capacities” or about joint human–AI performance [2503.16438].

A second foundation is the claim that humans often infer AI capacity through models originally developed for evaluating humans. “Human Learning about AI” formalizes this as an Ability Model combined with **Difficulty Projection**, in which perceived AI task difficulty is partly anchored to human difficulty:  
$$
\tilde \delta^A(t)=\lambda \delta^H(t)+(1-\lambda)\delta^A(t), \quad \lambda \in [0,1].
$$
Under this model, people treat AI failures on human-easy tasks and successes on human-difficult tasks as unusually informative about AI’s overall ability, even when AI performance is in fact only weakly related to human task difficulty [2406.05408].

A third foundation is that perceived capacity is often inseparable from broader judgments about AI as a social or moral actor. Public ratings of “mind” and “morality” in AI distinguish **agency**, **experience**, **moral agency**, and **moral patiency**, showing that perceived capacity includes planning, acting, feeling, and responsibility, not merely technical competence [2502.18683]. Related work on Social-AI treats “social intelligence” as a perceived, holistic construct rather than only a bundle of technical benchmarks, emphasizing that ordinary users infer capacity from observable behavior more than from explicit beliefs about internal mechanism or intent [2605.29938].

## 2. Principal dimensions and domain-specific manifestations

One of the most explicit multidimensional formulations is the AIQ proposal, which defines human–AI collaborative intelligence through eight dimensions: **Strategic AI Understanding**, **Prompt Engineering Intelligence**, **Critical Evaluation Capability**, **Integration Intelligence**, **Adaptive Learning Capability**, **Ethical Judgment in AI Utilization**, **Context Sensitivity**, and **Creative Synthesis** [2503.16438]. Within this framework, perceived AI capacity is most explicit in Strategic AI Understanding, but it also enters Critical Evaluation, Integration Intelligence, and Context Sensitivity, because all of them require judgments about when AI is likely to be reliable or fallible.

In software engineering, perceived AI capacity appears as a problem of **mental models** and **role attribution**. Developers in interviews describe AI-powered development tools either as an **inanimate tool**—a “machine,” “software,” “tool,” “search engine,” or “pattern recognizer”—or as a **human-like teammate**, such as a “colleague,” “assistant,” “teacher,” “junior engineer,” or “senior colleague.” Factor analysis then groups concrete roles into **Support Roles** and **Expert Roles**, with “assistant,” “reference guide,” and “tool” loading on the former, and “problem solver,” “advisor,” and “reviewer” on the latter [2504.20329]. Perceived capacity here is role-structured: bounded assistance differs qualitatively from advisory or evaluative authority.

Other work emphasizes that perceived capacity can be social, moral, or experiential rather than purely functional. Across 26 entities, AIs are rated as having **low-to-moderate agency** and **low experience**, but more varied **moral agency**; the highest-rated AI, a Tesla Full Self-Driving car, is rated as morally responsible for harm “as a chimpanzee” [2502.18683]. In studies of socially intelligent agents, participants associate social intelligence with “recognize and respond appropriately to human emotion,” “build rapport and trust with humans,” “take initiative during conversations and offer useful assistance,” and “understand what is ethically right and wrong” [2605.29938].

A further strand shifts the focus from machine properties to human experience with AI. “AI usage patterns are shaped by perceived gains in human agency” treats perceived AI capacity primarily through **perceived gains in human agency**, described as a “holistic phenomenological sensation of capacity.” It codes these gains into five dimensions: **instrumental**, **cognitive**, **affective**, **relational**, and **structural** agency [2607.02313]. This suggests that in practice many users do not evaluate AI capacity only as machine competence, but as the extent to which AI makes them feel more capable, effective, or future-oriented.

## 3. Measurement strategies

The literature uses several distinct measurement paradigms. The AIQ proposal is deliberately conceptual rather than psychometric: it does not provide a numeric scoring algorithm or factor model, but it does specify an **AIQ Assessment Process Flow** with **Initial Assessment**, then **Problem Decomposition tasks**, **Critical Evaluation tasks**, **Integration Tasks**, **Practical Application tasks**, and **Continuous Feedback**, followed by **Scoring and Analysis** through **Quantitative Metrics**, **Qualitative Assessment**, a **Performance Profile**, and a **Development Plan** [2503.16438].

A second strategy is **scenario-based expectancy measurement**. In the 71-scenario studies of public and expert perceptions, participants rate “How likely is the projection to occur in the next 10 years?” on a 6-point semantic differential rescaled to \([-100\%, +100\%]\), alongside risk, benefit, and valence. Valence is then modeled as  
$$
S_i=\alpha+\beta_R R_i+\beta_B B_i+\varepsilon_i,
$$
which makes it possible to separate perceived capacity from perceived desirability [2412.01459].

A third strategy models perceived capacity as a latent psychometric construct. “Capturing Humans’ Mental Models of AI” uses an Item Response Theory approach in which perceived teammate ability and perceived task difficulty jointly determine estimated performance:
$$
\theta_{ij}=a_i-d_j.
$$
The framework is extended to multidimensional IRT across four trivia topics, allowing perceived AI ability to be estimated separately for History of Art, Video Games, Cities, and Math, as well as through the covariance structure across those dimensions [2305.09064].

A fourth strategy uses **indirect elicitation** rather than direct self-report. In the academic policy–practice study, the logic of implementation is expressed as  
$$
\text{Policy in effect at practice}\iff P \land A \land R,
$$
where \(R\) includes practitioner capacity. Perceived AI capacity is then inferred from a ten-item instrument and three filtered indicators: **AI-integrated assessment capacity (proxy)**, **sector-level necessity (proxy)**, and **ontological stance** [2511.02875]. This approach treats perceived capacity as something revealed by coherent response patterns rather than by a single explicit question.

A fifth strategy aggregates local perceptions into organizational climate. In vocational education, **teacher-perceived AI capability** is measured through four items about whether generative AI has “information retrieval, analysis, and decision-making capabilities similar to (or even beyond) those of humans,” “multi-turn dialogue, contextual awareness, and self-correction capabilities,” and content generation “at or even beyond human-level quality,” then aggregated to the school level as a Level-2 mediator [2603.20056].

## 4. Empirical patterns and miscalibration

A consistent empirical result is that perceived capacity is often miscalibrated. In the expert–public comparison over 71 scenarios, experts report higher average **Expectancy** than the public: \(+25.2\%\) versus \(+12.7\%\). They also report lower **Perceived risk** (\(+19.3\%\) versus \(+34.7\%\)), higher **Benefit** (\(+15.3\%\) versus \(-5.2\%\)), and less negative **Valence** (\(-4.0\%\) versus \(-19.7\%\)) [2412.01459]. Yet the two groups still show a strong scenario-level correlation in expectancy, \(r=0.734\), indicating agreement on which developments are more plausible, but disagreement in degree [2412.01459].

The projection problem is especially clear in experimental work. On 414 standardized math tasks, ChatGPT 3.5 achieves average success of 82%, while human success averages 67%, and ChatGPT’s performance is almost uncorrelated with human difficulty: coefficient \(-0.001\), \(R^2=0.002\). Nonetheless, people’s perceived AI success falls sharply with human difficulty; the corresponding coefficient is \(-0.32\), with \(R^2=0.12\) [2406.05408]. The same work shows downstream adoption errors: in an environment where AI is better on human-hard tasks but worse on human-easy tasks, **Full Adoption** occurs for 34% of participants under an anthropomorphic framing but only 15% under a Black Box framing [2406.05408].

Perceptual bias also appears in judgment of AI-generated writing. In the logical argumentation study, when participants compared one AI-generated and one human-written text on the same topic, 50.0% rated the AI text lower, 18.14% rated the two texts equal, and 31.86% rated the AI text higher. In the regression on score classification, only the **logic** factor had the strongest and most reliable association, \(\beta=-0.3423\), \(t=-3.398\), \(p=0.001\) [2511.22151]. The study interprets this as a bias in which preconceptions about AI reasoning shape evaluation of the same textual evidence.

Human–AI teamwork studies further show that higher perceived capacity does not guarantee better joint performance. In the embodied control task, an AI teammate is rated as less helpful and less leader-like than a human teammate in Session 1, but ratings of helpfulness, leadership, and trust increase across sessions, with trust showing \(F(2,22)=10.405, P=0.0007\). Despite this, human-only teams significantly outperform human–AI teams overall, \(F(1,20)=23.278, P=0.0001\), especially at intermediate and hard difficulty [2501.15332]. The result is a dissociation between rising perceived AI capacity and persistently worse team outcomes.

Comparable dissociations appear in social adoption. In the Social-AI survey, 89% report having interacted with an AI agent they perceived as socially intelligent, and the composite Perceived AI Social Intelligence score averages 3.51 on a 1–5 scale. Yet across 12 scenarios there is a consistent **support–adoption gap**: “glad this service existed for others” exceeds “actively seek out this service for myself” by about 0.20 on a normalized 0–1 scale [2605.29938]. Perceived capacity therefore supports approval without necessarily producing personal adoption.

## 5. Consequences for adoption, work, and institutions

In education and organizational development, perceived AI capacity is increasingly treated as an institutional variable rather than only an individual attitude. The AIQ proposal explicitly recommends applications in **Educational development**, **Professional development**, and **Organizational strategy**, arguing that assessments could diagnose overreliance, underreliance, and deficits in strategic understanding, prompting, evaluation, and integration [2503.16438]. In vocational education, **overall school AI readiness** predicts **aggregated teacher-perceived AI capability** with \(\beta=0.165\), \(SE=0.033\), \(t=5.003\), \(p<0.001\), and teacher-perceived capability predicts student AI literacy with \(\beta=0.009\), \(SE=0.004\), \(t=2.333\), \(p<0.001\); the indirect effect is \(ab=0.0015\), \(z=2.114\), \(p=0.034\) [2603.20056]. This positions perceived capacity as a transmission mechanism from institutional readiness to learning outcomes.

In software engineering, adoption is closely tied to role-based perceptions of capacity. The total number of roles assigned to AI-powered development tools correlates with **Perceived Usefulness** at \(r=0.59, p<0.001\) and with **Perceived Ease of Use** at \(r=0.56, p<0.001\); **Expert Roles** correlate with \(PU\) at \(r=0.41, p<0.001\) and with \(PEU\) at \(r=0.42, p<0.001\) [2504.20329]. This indicates that richer and more capable role attributions are associated with higher acceptance and more experimentation with AI tools.

In the workplace, perceived capacity reorganizes expectations about decency, status, and meaningfulness. Employees in IT and healthcare often anticipate increased satisfaction with decency aspects such as working hours, but decreased satisfaction with meaningfulness aspects such as social image “due to misconceptions about AI handling most of their tasks,” whereas service workers anticipate little improvement in working hours but a higher social standing from the status of working with AI [2605.28680]. A related vignette study shows that **low AI competency** or **low AI proactivity** generally improve ownership, meaningfulness, satisfaction, and role dynamics, while highly competent and highly proactive AI can make the system appear as a “superior,” reducing human accountability and job identity [2606.00182].

At the level of everyday use, perceived capacity can dominate more conventional trust variables. Daily AI chatbot users often continue using systems not because they see them as reliably “good,” but because they experience **perceived gains in individual agency** that “often outweigh concerns about accuracy, reliability, and consistency” [2607.02313]. This suggests that adoption may be driven as much by perceived augmentation of human capability as by judgments about objective system accuracy.

Workforce-level framing pushes the same point further. The AI Pyramid argues that individuals, organizations, and governments often misread AI readiness by relying on titles, credentials, or exposure to tools, when what matters is the distribution of **AI Native capability**, **AI Foundation capability**, and **AI Deep capability** across the system [2601.06500]. Here, perceived AI capacity becomes a policy problem: institutions can overestimate readiness if they mistake a small apex of technical experts for broad AI Nativity.

## 6. Limitations, controversies, and open questions

A recurring limitation is that many current treatments are still conceptual or proxy-based. The AIQ paper offers no explicit \( \text{AIQ}=f(C_1,\dots,C_n) \) formula, no reliabilities, and no factor model; it is a conceptual proposal rather than an empirical psychometric instrument [2503.16438]. The academic policy–practice instrument likewise presents its indicators as “filtered, conservative proxies” and a “reusable scaffold, not a diagnostic device,” while also warning about “Goodhart-type distortions” once such indicators enter governance or procurement [2511.02875].

A second limitation is that perceptions move faster or slower than actual capabilities. Several papers emphasize that AI capabilities evolve quickly, so mental models can lag behind both improvements and newly revealed failure modes [2503.16438]. The scenario studies are explicitly cross-sectional and geographically bounded; one is Germany-only for the public sample, and another concentrates on U.S. adults evaluating Social-AI, making cross-cultural variation an open issue [2412.01459][2605.29938].

A third controversy concerns what perceived capacity is actually tracking. In some settings it tracks expected objective performance; in others it tracks perceived social intelligence, perceived human agency gains, ontological stance, or role identity. This suggests that “perceived AI capacity” may be partly an umbrella label for several constructs that are related but not interchangeable. A plausible implication is that future work will need to separate **objective AI knowledge**, **subjective perceived AI capacity**, **trust calibration**, and **realized human–AI performance**, then model their interactions directly rather than treating them as substitutes [2503.16438].

Finally, the literature increasingly converges on a calibration agenda. Proposed directions include longitudinal tracking of how perceptions and behaviors co-evolve with AI upgrades, cross-cultural comparisons, richer typologies of AI users, integration of self-report with performance-based tasks, and interfaces or training regimes that make AI strengths, weaknesses, and failure modes legible enough for humans to form accurate mental models [2406.05408][2607.02313]. Across these approaches, the central problem remains stable: effective human–AI systems depend not only on what AI can do, but on how accurately people perceive, update, and act on those capacities.

Source: https://www.emergentmind.com/topics/perceived-ai-capacity