Mistral Large 2 has been tested across a wide range of benchmarks, demonstrating state-of-the-art performance in multiple domains. Here are the results: (for more details read here)
Base Pretrained Benchmarks
| Benchmark | Score |
|---|---|
| MMLU | 84.00% |
Base Pretrained Multilingual Benchmarks (MMLU)
| Benchmark | Score |
|---|---|
| French | 82.80% |
| German | 81.60% |
| Spanish | 82.70% |
| Italian | 82.70% |
| Dutch | 80.70% |
| Portuguese | 81.60% |
| Russian | 79.00% |
| Korean | 60.10% |
| Japanese | 78.80% |
| Chinese | 74.80% |
Instruction Benchmarks
| Benchmark | Score |
|---|---|
| MT Bench | 8.63 |
| Wild Bench | 56.3 |
| Arena Hard | 73.2 |
Code & Reasoning Benchmarks
| Benchmark | Score |
|---|---|
| Human Eval | 92% |
| Human Eval Plus | 87% |
| MBPP Base | 80% |
| MBPP Plus | 69.00% |
Math Benchmarks
| Benchmark | Score |
|---|---|
| GSM8K | 93% |
| Math Instruct (0-shot, no CoT) | 70% |
| Math Instruct (0-shot, CoT) | 72% |
Benchmarks Glossary
- MMLU: Measures performance across multiple language understanding tasks.
- MT Bench: Measures multi-task performance on a variety of tasks.
- Wild Bench: Benchmark for handling open-ended or unpredictable tasks.
- Arena Hard: Performance on difficult adversarial tasks.
- Human Eval: Tests code generation capabilities.
- Human Eval Plus: Enhanced version of Human Eval with additional challenges.
- MBPP Base: Code benchmark focusing on programming tasks.
- MBPP Plus: Advanced programming task benchmark.
- GSM8K: Benchmark for mathematical problem solving.
- Math Instruct (0-shot, no CoT): Measures math performance without chain-of-thought reasoning.
- Math Instruct (0-shot, CoT): Measures math performance with chain-of-thought reasoning.
