Claude 3 Sonnet has been evaluated across various benchmarks, demonstrating significant advancements in intelligence and performance compared to other models in its class, including Claude 3 Opus, Claude 3 Haiku, and GPT-4, etc.
Here’s a comparison of Claude 3 Sonnet's performance against other models: (for more details read [here(https://www.anthropic.com/news/claude-3-family))
| Benchmark | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| Undergraduate level knowledge (MMLU) | 86.8% 5-shot | 79.0% 5-shot | 75.2% 5-shot | 86.4% 5-shot | 70.0% 5-shot | 83.7% 5-shot | 71.8% 5-shot |
| Graduate level reasoning (GPQA, Diamond) | 50.4% 0-shot CoT | 40.4% 0-shot CoT | 33.3% 0-shot CoT | 35.7% 0-shot CoT | 28.1% 0-shot CoT | — | — |
| Grade school math (GSM8K) | 95.0% 0-shot CoT | 92.3% 0-shot CoT | 88.9% 0-shot CoT | 92.0% 5-shot CoT | 57.1% 5-shot CoT | 94.4% MajI@32 | 86.5% MajI@32 |
| Math problem-solving (MATH) | 60.1% 0-shot CoT | 43.1% 0-shot CoT | 38.9% 0-shot CoT | 52.9% 4-shot CoT | 34.1% 4-shot CoT | 53.2% 4-shot CoT | 32.6% 4-shot CoT |
| Multilingual math (MGSM) | 90.7% 0-shot | 83.5% 0-shot | 75.1% 0-shot | 74.5% 8-shot | — | 79.0% 8-shot | 63.5% 8-shot |
| Code (HumanEval) | 84.9% 0-shot | 73.0% 0-shot | 75.9% 0-shot | 67.0% 0-shot | 48.1% 0-shot | 74.4% 0-shot | 67.7% 0-shot |
| Reasoning over text (DROP, F1 score) | 83.1% 3-shot | 78.9% 3-shot | 78.4% 3-shot | 80.9% 3-shot | 64.1% 3-shot | 82.4% Variable shots | 74.1% Variable shots |
| Mixed evaluations (BIG-Bench-Hard) | 86.8% 3-shot CoT | 82.9% 3-shot CoT | 73.7% 3-shot CoT | 83.1% 3-shot CoT | 66.6% 3-shot CoT | 83.6% 3-shot CoT | 75.0% 3-shot CoT |
| Knowledge Q&A (ARC-Challenge) | 96.4% 25-shot | 93.2% 25-shot | 89.2% 25-shot | 96.3% 25-shot | 85.2% 25-shot | — | — |
| Common Knowledge (HellaSwag) | 95.4% 10-shot | 89.0% 10-shot | 85.9% 10-shot | 95.3% 10-shot | 85.5% 10-shot | 87.8% 10-shot | 84.7% 10-shot |
The Claude 3 models have sophisticated vision capabilities on par with other leading models. Here’s a compiled table comparing the models’ performance:
Vision Capabilities Benchmarks
| Benchmark | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4V | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|
| Math & reasoning (MMMU val) | 59.40% | 53.10% | 50.20% | 56.80% | 59.40% | 47.90% |
| Document visual Q&A (ANLS score, test) | 89.30% | 89.50% | 89% | 88.40% | 90.90% | 88.10% |
| Math (MathVista testmini) | 50.5% CoT | 47.9% CoT | 46.4% CoT | 49.90% | 53.00% | 45.20% |
| Science diagrams (AI2D, test) | 88.10% | 88.70% | 87% | 78.20% | 79.50% | 73.90% |
| Chart Q&A (Relaxed accuracy test) | 80.8% 0-shot CoT | 81.1% 0-shot CoT | 81.7% 0-shot CoT | 78.5% 4-shot CoT | 80.80% | 74.10% |
Claude Family Model Comparison
To help you choose the right model for your needs, here’s a compiled table comparing the key features and capabilities of each model in the Claude family:
| Claude Model | Claude 3.5 Sonnet | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku |
|---|---|---|---|---|
| Description | Most intelligent model | Powerful model for highly complex tasks | Balance of intelligence and speed | Fastest and most compact model for near-instant responsiveness |
| Strengths | Highest level of intelligence and capability | Top-level performance, intelligence, fluency, and understanding | Strong utility, balanced for scaled deployments | Quick and accurate targeted performance |
| Multilingual | Yes | Yes | Yes | Yes |
| Vision | Yes | Yes | Yes | Yes |
| API model name | claude-3-5-sonnet-20240620 | claude-3-opus-20240229 | claude-3-sonnet-20240229 | claude-3-haiku-20240307 |
| API format | Messages API | Messages API | Messages API | Messages API |
| Comparative latency | Fast | Moderately fast | Fast | Fastest |
| Context window | 200K | 200K | 200K | 200K |
| Max output | 8192 tokens | 4096 tokens | 4096 tokens | 4096 tokens |
| Cost (Input / Output per MTok) | $3.00 / $15.00 | $15.00 / $75.00 | $3.00 / $15.00 | $0.25 / $1.25 |
| Training data cut-off | Apr 2024 | Aug 2023 | Aug 2023 | Aug 2023 |
Benchmark Metric Glossary
- MMLU: Tests general knowledge across subjects (e.g., history, math).
- GPQA: Graduate-level reasoning and knowledge tasks.
- GSM8K: Grade-school math problems.
- MATH: Advanced high school/college-level math.
- MGSM: Multilingual version of GSM8K math.
- HumanEval: Tests coding ability via programming tasks.
- DROP: Complex reading comprehension with reasoning.
- BIG-Bench-Hard: Difficult reasoning and problem-solving tasks.
- ARC-Challenge: Complex reasoning and common-sense questions.
- HellaSwag: Common-sense story continuation.
- MMMU: Tests understanding across text, math, and visuals.
- ANLS: Measures similarity between answers.
- MathVista: Challenging math requiring reasoning.
- AI2D: Understanding scientific diagrams.
- Relaxed Accuracy (Chart Q&A): Answering questions based on charts/graphs.