OpenAI-o1-preview has been tested across a wide range of benchmarks, demonstrating state-of-the-art performance in multiple domains. Here are the results: (more details here)
| Benchmark | GPT-4o | o1-preview | o1 | Expert Human |
|---|---|---|---|---|
| Competition Math (AIME 2024) | 13.40% | 56.70% | 83.30% | - |
| Competition Code (Codeforces) | 11.00% | 62.00% | 89.00% | - |
| PhD-Level Science Questions (GPQA Diamond) | 56.10% | 78.30% | 78.00% | 69.70% |
ML Benchmarks
| Benchmark | GPT-4o | o1 improvement |
|---|---|---|
| MATH | 60.3 | 94.8 |
| MathVista (testmini) | 63.8 | 73.2 |
| MMMU (val) | 69.1 | 78.1 |
| MMLU | 88 | 92.3 |
PhD-Level Science Questions (GPQA Diamond)
| Benchmark | GPT-4o | o1 improvement |
|---|---|---|
| Chemistry | 40.2 | 64.7 |
| Physics | 59.5 | 92.8 |
| Biology | 61.6 | 69.2 |
Benchmarks Glossary
- Competition Math (AIME 2024): Measures accuracy in advanced math problems.
- Competition Code (Codeforces): Evaluates programming skills using Elo ratings.
- PhD-Level Science Questions (GPQA Diamond): Assesses performance on complex science questions.
- MATH: Benchmark for solving mathematical problems.
- MathVista (testmini): Tests performance on mathematical reasoning.
- MMMU (val): Evaluates understanding across various multi-modal tasks.
- MMLU: Measures general language understanding across multiple languages.
- Chemistry/Physics/Biology: PhD-level science problem-solving ability.