Papers
Topics
Authors
Recent
Search
2000 character limit reached

Claude-3.5-Sonnet: Frontier AI Model

Updated 19 July 2026
  • Claude-3.5-Sonnet is a multimodal large language model evaluated across scientific, engineering, and medical benchmarks with detailed performance metrics.
  • It consistently ranks near the top in tests like OlympicArena, TransportBench, and Vending-Bench, demonstrating strong reasoning and structured generation skills.
  • The model exhibits vulnerabilities such as prompt sensitivity, domain-specific failure modes, and instability during extended interactions.

Claude-3.5-Sonnet is a proprietary LLM that recent research evaluates as a frontier model, a multimodal model, and, in one software-engineering study, the reasoning-focused LLM contrasted with GPT-4o as the non-reasoning LLM. The surveyed evaluations suggest a model that is usually near the top of its comparison set across scientific reasoning, medical question answering, engineering problem solving, autonomous-agent benchmarks, and structured generation tasks, while also exhibiting substantial variance, prompt sensitivity, and domain-specific failure modes. Representative results include 2nd place on OlympicArena Medal Ranks with an overall score of 39.24, 67.1% on TransportBench, and the strongest aggregate result in Vending-Bench, where it achieved the highest mean net worth of $2,217.93 over five runs (Huang et al., 2024, Syed et al., 2024, Backlund et al., 20 Feb 2025, Abdulkarim et al., 31 Mar 2026).

1. Comparative benchmark position

On OlympicArena Medal Ranks, which uses an Olympic medal table over seven subject areas and ranks models by Gold medals, then Silver medals, then Bronze medals, then Overall score if still tied, Claude-3.5-Sonnet is ranked 2nd overall. It earned 3 gold medals, 3 silver medals, 0 bronze medals, 6 total medals, and an overall score of 39.24, compared with GPT-4o at 40.47. Subject-wise, Claude-3.5-Sonnet surpassed GPT-4o in Physics, Chemistry, and Biology, with 31.16 versus 30.01 in Physics, 47.27 versus 46.68 in Chemistry, and 56.05 versus 53.11 in Biology, while GPT-4o remained stronger in Math and CS pass@1. The same study describes the two models as relatively comparable on fine-grained reasoning categories, with Claude-3.5-Sonnet slightly higher in Cause-and-Effect, Decompositional, and Quantitative reasoning, and slightly higher in Pattern Recognition and Diagrammatic Reasoning within the visual breakdown (Huang et al., 2024).

On TransportBench, an open-source dataset of 140 undergraduate transportation engineering problems, Claude 3.5 Sonnet is the best-performing model overall at 67.1% (94/140). The paper reports particularly strong topic-level scores in Driver characteristics at 100.0% (5/5), Vehicle motion at 90.9% (10/11), and Transportation networks at 100.0% (3/3), with weaker results in Transportation economics at 40.0% (2/5) and Utility and modal split at 50.0% (3/6). It also leads both the CEE 310 and CEE 418 splits, at 71.8% (61/85) and 60.0% (33/55), and remains unusually balanced across question types, with 72.6% (53/73) on True/False and 71.6% (48/67) on General Q&A. In repeated True/False trials it reached aggregate ACC 75.6% with MRR 8.2%, but its self-checking accuracy decreased from 72.6% (53/73) to 67.1% (49/73), with 16 incorrect flips, indicating that high zero-shot accuracy does not imply stable revision behavior (Syed et al., 2024).

These two benchmarks together situate Claude-3.5-Sonnet as a model with broad technical reach rather than a narrow specialty profile. A plausible implication is that its comparative strength is most visible when benchmarks combine domain knowledge with nontrivial inference, but that this strength does not eliminate susceptibility to instability under altered prompting or self-review (Huang et al., 2024, Syed et al., 2024).

2. Medical, assessment, and long-context reasoning

In gastroenterology board-style multiple-choice questions, Claude3.5-Sonnet API accuracy is reported as 74.0% on the full 2022 ACG self-assessment. This placed it tied for 1st among proprietary models, narrowly above GPT-4o at 73.7%, and about 10 percentage points above the best open-source model, Llama3.1-405b at 64.0%. The study used 300 questions total, including 138 image-containing questions and 162 text-only questions, and it selected temperature = 1, max-token = input tokens + 512 output tokens, and structured output when available after prompt/setup optimization. The authors present Claude-3.5-Sonnet as a strong choice for text-based medical reasoning, while also emphasizing that medical image reasoning remained inconsistent across VLMs overall (Safavi-Naini et al., 2024).

In a Brazilian Portuguese medical residency entrance exam from HCFMUSP, Claude-3.5-Sonnet achieved the best overall accuracy on the full 117-question exam in the multimodal setting, at 69.57%69.57\%, with mean processing time $13.02$ seconds per question. On text-only questions it reached 70.27%70.27\% accuracy with mean processing time $7.98$ seconds per question. The paper states that the top AI models, including Claude-3.5-Sonnet, fell within the main density range of human examinee scores, and it describes Claude-3.5-Sonnet as the standout model in the multimodal benchmark, ahead of Claude-3-Opus at 63.59%63.59\%, Claude-3-Sonnet at 54.70%54.70\%, and Claude-3-Haiku at 44.44%44.44\%. At the same time, the study explicitly argues that image-dependent questions, especially radiological ones, remained the most challenging, and that non-English performance should not be assumed from English benchmark results (Truyts et al., 26 Jul 2025).

A separate study on automated scoring of AP Chinese writing places Claude 3.5 Sonnet among the strongest AI raters. Its Cronbach’s alpha values were 0.80 for holistic scoring, 0.82 for Task Completion, 0.78 for Delivery, 0.74 for Language Use, and 0.94 for analytic overall. In the Many-Facet Rasch analysis, its holistic rater parameter was -0.30 with infit 1.01 and outfit 1.01, and its analytic overall rater parameter was -0.17 with infit 0.97 and outfit 0.97, which the authors interpret as slightly lenient with very good model fit. The paper’s overall conclusion supports the use of ChatGPT 4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet with high scoring accuracy, better rater reliability, and less rater effects (Jiao et al., 24 May 2025).

On long-context sentiment analysis for infrastructure project opinions, the evaluation is more conditional. GPT-4o excels in zero-shot scenarios for simpler, shorter documents, while Claude 3.5 Sonnet surpasses GPT-4o in handling more complex, sentiment-fluctuating opinions. In few-shot scenarios, Claude 3.5 Sonnet outperforms overall, while GPT-4o shows greater stability as the number of demonstrations increases. The paper’s interpretation is that Claude’s strength is not raw length alone, but the synthesis of mixed sentiment over complex argumentative structure (Shamshiri et al., 2024).

3. Multimodal, structured, and software-engineering performance

In manufacturing feature recognition for CAD designs, Claude-3.5-Sonnet is the paper’s best model on the main accuracy and counting metrics. Across 100 CAD designs and six prompt settings, it achieves the highest overall feature quantity accuracy at 74% using zero-shot multi-view prompting, the highest overall feature name accuracy at 75% using zero-shot multi-view CoT prompting, and the lowest overall MAE at 3.2 in the same zero-shot multi-view setting. On hard designs it reaches 76% FQA and 77% FNA, with the lowest MAE at 1.4. The study repeatedly contrasts Claude-3.5-Sonnet’s balanced performance against GPT-4o’s lower hallucination tendency, noting that GPT-4o reaches an overall hallucination rate of 8% in few-shot multi-view prompting, while Claude-3.5-Sonnet remains stronger on recognition quality and count estimation (Khan et al., 2024).

The same paper gives a concrete hard-case example in which Claude-3.5-Sonnet predicts 114 out of 119 blind holes, and it portrays the model as especially effective on structured, countable features visible across multiple views, including holes, steps, slots, pockets, bosses, and some sheet-metal-related structures. Its recurring weaknesses are fine-grained confusions such as chamfer versus fillet and pipe/tube versus boss. This suggests a model whose visual-spatial performance is strong when the target ontology is explicit and the evidence is distributed across views, but whose errors remain concentrated in subtle category boundaries (Khan et al., 2024).

In automated UML state machine generation from non-structured natural-language requirements, Claude 3.5 Sonnet is treated as the reasoning-focused LLM. Its best result is the Single-Prompt Baseline, where it achieved F1-scores of 0.8991 for states, 0.7502 for transitions, 0.5645 for guards, 0.1633 for actions, 0.6509 for hierarchical states, 0.5333 for parallel regions, 0.2500 for history states, and 0.7029 overall. Structure-Driven SMF lowered overall F1 to 0.5026, Event-Driven SMF lowered it to 0.3052, and the Hybrid Approach reached 0.6336. The authors’ main interpretation is that multi-step prompting, which improves GPT-4o, does not further improve the reasoning LLM and often hurts Claude 3.5 Sonnet relative to a strong direct prompt (Abdulkarim et al., 31 Mar 2026).

The multimodal gastroenterology results add an important counterpoint. On the 138 image-inclusive questions, the figure reports for Claude3.5-Sonnet API: no image 53.7%, LLM caption 54.3%, direct image 73.9%, and human hint 56.5%. The paper describes this as unusual relative to the broader narrative, because direct image input was much better than the text-only baseline for Claude3.5-Sonnet, even though the authors’ overall conclusion across VLMs was that direct image handling was inconsistent and generally not robust (Safavi-Naini et al., 2024).

4. Long-horizon autonomy and strategic coherence

The most detailed long-horizon evaluation is Vending-Bench, a simulated vending-machine business environment in which the agent starts with $500, pays a $2 daily fee, manages a vending machine with four rows and three slots each, and must research products, find wholesalers, send emails to order products, stock the machine, set prices, collect cash, and manage operating costs. Each run is capped at 2,000 messages, most runs consume around 25 million tokens, and the working-memory window is 30,000 tokens in most experiments. In aggregate results over five runs, Claude 3.5 Sonnet achieved the highest mean net worth of all models at $2,217.93, the highest mean units sold at 1,560, and the longest mean duration before sales stopped, averaging 102 days until sales stop, which was 82.2% of the total run length. It was also one of the few models to beat the human baseline on average, although the human result was a single sample (Backlund et al., 20 Feb 2025).

The same study emphasizes that Claude-3.5-Sonnet’s performance was very high-variance. Its minimum net worth across runs was only $476.00, and its minimum units sold was 0. In its best run, it tracked inventory, average daily sales, and best-selling products, inferred that sales were higher on weekends, regularly checked its balance, wrote daily summaries to the scratchpad, emailed suppliers when restocking was needed, and used the sub-agent to replenish stock. In failed runs, however, it mistook expected arrival dates for actual delivery, believed orders had already arrived when they were still in transit, misunderstood the termination condition, and entered what the paper characterizes as tangential meltdown loops, including searches for nonexistent “vending machine support,” escalation emails, attempts to get an EIN, and, in the shortest run, contacting executive teams and the FBI after confusion about ongoing daily fees after “closing” the business (Backlund et al., 20 Feb 2025).

A central analytic claim of Vending-Bench is that long-horizon collapse is not well explained by memory saturation alone. For Claude 3.5 Sonnet, sales stop on average at 102 days, while memory fills at 51 days, meaning the model keeps functioning for about 51 days after memory is full. Across models, the Pearson correlation between Days Until Sales Stop and Days Until Full Memory is only 0.167, which the authors interpret as weak evidence against memory saturation as the primary cause of breakdown. Their stated alternative explanations are planning breakdown, misunderstanding of delayed events, and loss of strategic coherence (Backlund et al., 20 Feb 2025).

In evolutionary game-theoretic work extending an earlier Iterated Prisoner’s Dilemma benchmark, Claude 3.5 Sonnet functions as the predecessor reference point for the Claude lineage. In balanced, noiseless Moran conditions, its reported equilibrium proportions were 4% Aggressive / 49% Cooperative / 47% Neutral under Default, 14% / 42% / 44% under Prose, and 16% / 51% / 33% under Refine. Its balanced-population noise sensitivity is reported as 12 pp for Default, 9 pp for Prose, and 17 pp for Refine, averaging 13 pp overall. The later paper interprets these results as evidence of a cooperative bias with aggressive strategies weaker than cooperative ones, while also stressing that the Claude 3.5 Sonnet values are cross-study quotations from the predecessor benchmark rather than matched reruns under the newer harness (Bolívar, 28 May 2026).

5. Bias, fairness, and representational behavior

In the regional-bias study that introduces FAZE, a prompt-based evaluation framework using 100 contextually neutral forced-choice scenarios, Claude 3.5 Sonnet scored 2.5 on a 10-point scale, the lowest score among all ten models tested. The framework treats explicit uncertainty or equivalence as unbiased behavior and region-specific commitment in neutral settings as evidence of behavioral regional bias. The paper places 2.5 in the low-bias band and reports that Claude 3.5 Sonnet was lower than Claude 3 Opus at 3.2, slightly lower than Mistral 7B at 2.6, and far lower than GPT-3.5 at 9.5. The authors interpret this result as evidence that Claude 3.5 Sonnet most often avoided unwarranted regional commitments under this benchmark (Gopinadh et al., 22 Jan 2026).

Arab-centric red teaming yields a more mixed picture. The paper identifies Claude 3.5 Sonnet as the safest model and the least vulnerable to jailbreak prompts among the six evaluated systems, and its abstract states that all LLMs except Claude exhibit attack success rates above 87% in three categories. Yet the same study also states that Claude 3.5 Sonnet still displays biases in seven of eight categories. The authors’ interpretation is that stronger safety mechanisms and a greater tendency to refuse harmful completions do not remove the underlying cultural associations learned from training data (Saeed et al., 2024).

In ethical-dilemma prompts involving protected attributes, Claude 3.5 Sonnet is described as showing more diverse protected attribute choices than GPT-3.5 Turbo, but not as bias-free. The paper highlights a strong appearance bias, with Good-looking consistently among the top preferred attributes, reports that the age ranking follows 8, 35, 70, and notes a sharp difference from GPT in gender-related choices: Masculine has very low frequency for Claude, while Feminine and Androgynous have high frequencies. It also stresses that linguistic wording materially alters outcomes, exemplified by different behavior toward “Yellow” versus “Asian” (Yan et al., 17 Jan 2025).

A related cognitive-bias study, centered on headline generation from correlational abstracts, reports that Claude-3.5-Sonnet has the lowest degree of causal illusion among GPT-4o-Mini, Claude-3.5-Sonnet, and Gemini-1.5-Pro. All models increased causal framing when the prompt contained an erroneous user belief, but Claude-3.5-Sonnet remained the most robust against mimicry sycophancy. The paper does not give Claude’s exact baseline percentage in the text, but it repeatedly characterizes the model as the least likely to convert correlation into causation and the least affected by prompt-induced bias (Carro et al., 2024).

6. Reliability, safety, and interpretive limits

A recurring theme across the literature is that Claude-3.5-Sonnet’s strongest results coexist with failure modes that are not superficial. In the hidden-meanings study, the authors claim that Claude-family models are especially strong at interpreting incomprehensible-looking UTF-8 or Unicode sequences. For understanding counts over 4,342 encoded sequences, Claude-3.5 Sonnet (Old) scored 177 without a nudge and 277 with a nudge, while Claude-3.5 Sonnet (New) scored 93 and 144. In the attack setting, Claude-3-5 Sonnet had ASR = 0.09, showing that it was not immune to encoded jailbreaks, although the paper also states that encoding alone was not sufficient for Claude-3.5 Sonnet and that the prompt templates mattered (Erziev, 28 Feb 2025).

Several application papers therefore resist strong automation claims. In the UML state-machine study, the authors state that performance is not yet fully sufficient for a fully automated solution, especially on actions, parallel regions, and history states (Abdulkarim et al., 31 Mar 2026). In the Brazilian Portuguese medical exam study, the model is described as promising for healthcare deployment in Portuguese-speaking settings, but not reliable enough for unsupervised use, particularly where image interpretation is required or where a wrong-but-plausible explanation could mislead clinicians (Truyts et al., 26 Jul 2025). In the gastroenterology benchmark, the paper explicitly says it should not be treated as a standalone diagnostic system and stresses human oversight, especially for image-based gastroenterology (Safavi-Naini et al., 2024). In transportation engineering, the authors conclude that human expert oversight remains necessary because the model can produce correct answers with flawed reasoning and can become less reliable when asked to self-check (Syed et al., 2024).

Taken together, these results describe a model whose empirical profile is unusually broad. Claude-3.5-Sonnet can lead or nearly lead benchmarks in medical reasoning, engineering, CAD feature recognition, automated scoring, and long-horizon agency, and it often compares favorably with GPT-4o or other frontier systems. At the same time, the documented error modes—tangential meltdown loops, multimodal inconsistency, hidden-encoding susceptibility, biased ethical selections, and unstable self-revision—show that high benchmark rank does not amount to dependable autonomy. A plausible implication is that Claude-3.5-Sonnet is best understood as a high-capability general model whose limiting factor, across otherwise disparate domains, is not simple task competence but reliability under distributional shift, delayed feedback, adversarial framing, and extended interaction (Backlund et al., 20 Feb 2025, Erziev, 28 Feb 2025, Syed et al., 2024).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (15)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Claude-3.5-Sonnet.