Contact Support
    Anthropic/Claude-3 Sonnet

    Model Card

    Claude 3 Sonnet has been evaluated across various benchmarks, demonstrating significant advancements in intelligence and performance compared to other models in its class, including Claude 3 Opus, Claude 3 Haiku, and GPT-4, etc.

    Here’s a comparison of Claude 3 Sonnet's performance against other models: (for more details read [here(https://www.anthropic.com/news/claude-3-family))

    BenchmarkClaude 3 OpusClaude 3 SonnetClaude 3 HaikuGPT-4GPT-3.5Gemini 1.0 UltraGemini 1.0 Pro
    Undergraduate level knowledge (MMLU)86.8% 5-shot79.0% 5-shot75.2% 5-shot86.4% 5-shot70.0% 5-shot83.7% 5-shot71.8% 5-shot
    Graduate level reasoning (GPQA, Diamond)50.4% 0-shot CoT40.4% 0-shot CoT33.3% 0-shot CoT35.7% 0-shot CoT28.1% 0-shot CoT
    Grade school math (GSM8K)95.0% 0-shot CoT92.3% 0-shot CoT88.9% 0-shot CoT92.0% 5-shot CoT57.1% 5-shot CoT94.4% MajI@3286.5% MajI@32
    Math problem-solving (MATH)60.1% 0-shot CoT43.1% 0-shot CoT38.9% 0-shot CoT52.9% 4-shot CoT34.1% 4-shot CoT53.2% 4-shot CoT32.6% 4-shot CoT
    Multilingual math (MGSM)90.7% 0-shot83.5% 0-shot75.1% 0-shot74.5% 8-shot79.0% 8-shot63.5% 8-shot
    Code (HumanEval)84.9% 0-shot73.0% 0-shot75.9% 0-shot67.0% 0-shot48.1% 0-shot74.4% 0-shot67.7% 0-shot
    Reasoning over text (DROP, F1 score)83.1% 3-shot78.9% 3-shot78.4% 3-shot80.9% 3-shot64.1% 3-shot82.4% Variable shots74.1% Variable shots
    Mixed evaluations (BIG-Bench-Hard)86.8% 3-shot CoT82.9% 3-shot CoT73.7% 3-shot CoT83.1% 3-shot CoT66.6% 3-shot CoT83.6% 3-shot CoT75.0% 3-shot CoT
    Knowledge Q&A (ARC-Challenge)96.4% 25-shot93.2% 25-shot89.2% 25-shot96.3% 25-shot85.2% 25-shot
    Common Knowledge (HellaSwag)95.4% 10-shot89.0% 10-shot85.9% 10-shot95.3% 10-shot85.5% 10-shot87.8% 10-shot84.7% 10-shot

    The Claude 3 models have sophisticated vision capabilities on par with other leading models. Here’s a compiled table comparing the models’ performance:

    Vision Capabilities Benchmarks

    BenchmarkClaude 3 OpusClaude 3 SonnetClaude 3 HaikuGPT-4VGemini 1.0 UltraGemini 1.0 Pro
    Math & reasoning (MMMU val)59.40%53.10%50.20%56.80%59.40%47.90%
    Document visual Q&A (ANLS score, test)89.30%89.50%89%88.40%90.90%88.10%
    Math (MathVista testmini)50.5% CoT47.9% CoT46.4% CoT49.90%53.00%45.20%
    Science diagrams (AI2D, test)88.10%88.70%87%78.20%79.50%73.90%
    Chart Q&A (Relaxed accuracy test)80.8% 0-shot CoT81.1% 0-shot CoT81.7% 0-shot CoT78.5% 4-shot CoT80.80%74.10%

    Claude Family Model Comparison

    To help you choose the right model for your needs, here’s a compiled table comparing the key features and capabilities of each model in the Claude family:

    Claude ModelClaude 3.5 SonnetClaude 3 OpusClaude 3 SonnetClaude 3 Haiku
    DescriptionMost intelligent modelPowerful model for highly complex tasksBalance of intelligence and speedFastest and most compact model for near-instant responsiveness
    StrengthsHighest level of intelligence and capabilityTop-level performance, intelligence, fluency, and understandingStrong utility, balanced for scaled deploymentsQuick and accurate targeted performance
    MultilingualYesYesYesYes
    VisionYesYesYesYes
    API model nameclaude-3-5-sonnet-20240620claude-3-opus-20240229claude-3-sonnet-20240229claude-3-haiku-20240307
    API formatMessages APIMessages APIMessages APIMessages API
    Comparative latencyFastModerately fastFastFastest
    Context window200K200K200K200K
    Max output8192 tokens4096 tokens4096 tokens4096 tokens
    Cost (Input / Output per MTok)$3.00 / $15.00$15.00 / $75.00$3.00 / $15.00$0.25 / $1.25
    Training data cut-offApr 2024Aug 2023Aug 2023Aug 2023

    Benchmark Metric Glossary

    • MMLU: Tests general knowledge across subjects (e.g., history, math).
    • GPQA: Graduate-level reasoning and knowledge tasks.
    • GSM8K: Grade-school math problems.
    • MATH: Advanced high school/college-level math.
    • MGSM: Multilingual version of GSM8K math.
    • HumanEval: Tests coding ability via programming tasks.
    • DROP: Complex reading comprehension with reasoning.
    • BIG-Bench-Hard: Difficult reasoning and problem-solving tasks.
    • ARC-Challenge: Complex reasoning and common-sense questions.
    • HellaSwag: Common-sense story continuation.
    • MMMU: Tests understanding across text, math, and visuals.
    • ANLS: Measures similarity between answers.
    • MathVista: Challenging math requiring reasoning.
    • AI2D: Understanding scientific diagrams.
    • Relaxed Accuracy (Chart Q&A): Answering questions based on charts/graphs.

    Meta data

    200,000 tokens
    $3 per million
    $15 per million
    Aug 2023
    Create an agent Pipe