Llama 3.1-405B-Instruct outperforms many of the available open-source and closed chat models on common industry benchmarks.
Here’s a table summarizing the comparison between different Llama models against various benchmarks: (for more details read here)
Base Pretrained Models
Category: General
| Benchmark | # Shots | Metric | Llama 3 8B | Llama 3.1 8B | Llama 3 70B | Llama 3.1 70B | Llama 3.1 405B |
|---|---|---|---|---|---|---|---|
| MMLU | 5 | macro_avg/acc_char | 66.7 | 66.7 | 79.5 | 79.3 | 85.2 |
| MMLU-Pro (CoT) | 5 | macro_avg/acc_char | 36.2 | 37.1 | 55.0 | 53.8 | 61.6 |
| AGIEval English | 3-5 | average/acc_char | 47.1 | 47.8 | 63.0 | 64.6 | 71.6 |
| CommonSenseQA | 7 | acc_char | 72.6 | 75.0 | 83.8 | 84.1 | 85.8 |
| Winogrande | 5 | acc_char | - | 60.5 | - | 83.3 | 86.7 |
| BIG-Bench Hard (CoT) | 3 | average/em | 61.1 | 64.2 | 81.3 | 81.6 | 85.9 |
| ARC-Challenge | 25 | acc_char | 79.4 | 79.7 | 93.1 | 92.9 | 96.1 |
Category: Knowledge Reasoning
| Benchmark | # Shots | Metric | Llama 3 8B | Llama 3.1 8B | Llama 3 70B | Llama 3.1 70B | Llama 3.1 405B |
|---|---|---|---|---|---|---|---|
| TriviaQA-Wiki | 5 | em | 78.5 | 77.6 | 89.7 | 89.8 | 91.8 |
Category: Reading Comprehension
| Benchmark | # Shots | Metric | Llama 3 8B | Llama 3.1 8B | Llama 3 70B | Llama 3.1 70B | Llama 3.1 405B |
|---|---|---|---|---|---|---|---|
| SQuAD | 1 | em | 76.4 | 77.0 | 85.6 | 81.8 | 89.3 |
| QuAC (F1) | 1 | f1 | 44.4 | 44.9 | 51.1 | 51.1 | 53.6 |
| BoolQ | 0 | acc_char | 75.7 | 75.0 | 79.0 | 79.4 | 80.0 |
| DROP (F1) | 3 | f1 | 58.4 | 59.5 | 79.7 | 79.6 | 84.8 |
Instruction-Tuned Models
Category: General
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| MMLU | 5 | macro_avg/acc | 68.5 | 69.4 | 82.0 | 83.6 | 87.3 |
| MMLU (CoT) | 0 | macro_avg/acc | 65.3 | 73.0 | 80.9 | 86.0 | 88.6 |
| MMLU-Pro (CoT) | 5 | micro_avg/acc_char | 45.5 | 48.3 | 63.4 | 66.4 | 73.3 |
| IFEval | - | - | 76.8 | 80.4 | 82.9 | 87.5 | 88.6 |
Category: Reasoning
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| ARC-C | 0 | acc | 82.4 | 83.4 | 94.4 | 94.8 | 96.9 |
| GPQA | 0 | em | 34.6 | 30.4 | 39.5 | 41.7 | 50.7 |
Category: Code
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| HumanEval | 0 | pass@1 | 60.4 | 72.6 | 81.7 | 80.5 | 89.0 |
| MBPP ++ base version | 0 | pass@1 | 70.6 | 72.8 | 82.5 | 86.0 | 88.6 |
| Multipl-E HumanEval | 0 | pass@1 | - | 50.8 | - | 65.5 | 75.2 |
| Multipl-E MBPP | 0 | pass@1 | - | 52.4 | - | 62.0 | 65.7 |
Category: Math
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| GSM-8K (CoT) | 8 | em_maj1@1 | 80.6 | 84.5 | 93.0 | 95.1 | 96.8 |
| MATH (CoT) | 0 | final_em | 29.1 | 51.9 | 51.0 | 68.0 | 73.8 |
Category: Tool Use
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| API-Bank | 0 | acc | 48.3 | 82.6 | 85.1 | 90.0 | 92.0 |
| BFCL | 0 | acc | 60.3 | 76.1 | 83.0 | 84.8 | 88.5 |
| Gorilla Benchmark API | 0 | acc | 1.7 | 8.2 | 14.7 | 29.7 | 35.3 |
| Nexus (0-shot) | 0 | macro_avg/acc | 18.1 | 38.5 | 47.8 | 56.7 | 58.7 |
Category: Multilingual
| Benchmark | # Shots | Metric | Llama 3 8B Instruct | Llama 3.1 8B Instruct | Llama 3 70B Instruct | Llama 3.1 70B Instruct | Llama 3.1 405B Instruct |
|---|---|---|---|---|---|---|---|
| Multilingual MGSM (CoT) | 0 | em | - | 68.9 | - | 86.9 | 91.6 |
Multilingual Benchmarks
Category: General
| Benchmark | Language | Llama 3.1 8B | Llama 3.1 70B | Llama 3.1 405B |
|---|---|---|---|---|
| MMLU (5-shot, macro_avg/acc) | Portuguese | 62.12 | 80.13 | 84.95 |
| MMLU (5-shot, macro_avg/acc) | Spanish | 62.45 | 80.05 | 85.08 |
| MMLU (5-shot, macro_avg/acc) | Italian | 61.63 | 80.40 | 85.04 |
| MMLU (5-shot, macro_avg/acc) | German | 60.59 | 79.27 | 84.36 |
| MMLU (5-shot, macro_avg/acc) | French | 62.34 | 79.82 | 84.66 |
| MMLU (5-shot, macro_avg/acc) | Hindi | 50.88 | 74.52 | 80.31 |
| MMLU (5-shot, macro_avg/acc) | Thai | 50.32 | 72.95 | 78.21 |
Benchmarks and Metrics Glossary
- MMLU (Massive Multitask Language Understanding): A benchmark designed to measure a model's performance across a wide range of tasks, focusing on its ability to handle diverse and complex language tasks.
- MMLU-Pro (CoT) (Massive Multitask Language Understanding - Chain of Thought): A variant of MMLU where the model is expected to generate reasoning chains (step-by-step explanations) for solving tasks.
- AGIEval: A benchmark designed to measure a model's performance on English language tasks.
- CommonSenseQA: A benchmark that tests a model's ability to answer questions requiring common sense reasoning.
- Winogrande: A commonsense reasoning benchmark focusing on resolving ambiguity in sentences.
- BIG-Bench (Beyond the Imitation Game Benchmark): A comprehensive benchmark covering a wide array of language tasks. The "Hard" subset focuses on more challenging tasks.
- ARC (AI2 Reasoning Challenge) - Challenge: A benchmark designed to measure a model's ability to answer challenging science questions.
- TriviaQA-Wiki: A benchmark for evaluating the model's ability to answer trivia questions, using Wikipedia as the source of truth.
- SQuAD (Stanford Question Answering Dataset): A reading comprehension benchmark where models must answer questions based on a given text passage.
- QuAC (Question Answering in Context): A benchmark for evaluating a model's ability to answer questions based on a dialogue context.
- BoolQ (Boolean Questions): A yes/no question-answering benchmark, testing the model’s ability to provide accurate binary answers.
- DROP (Discrete Reasoning Over Paragraphs): A reading comprehension benchmark focusing on discrete reasoning over paragraphs.
- IFEval: A benchmark designed to test a model’s understanding and performance across various English tasks.
- GPQA (Generalized Professional QA): A benchmark that tests a model's ability to answer professional-level questions across various domains.
- HumanEval: A benchmark designed to evaluate the model’s ability to write correct code based on problem descriptions.
- MBPP (MultiPL-E Benchmark for Programming Proficiency): A benchmark that tests a model's proficiency in writing code across multiple languages.
- GSM-8K (Grade School Math 8K): A benchmark focusing on the model’s ability to solve grade-school-level math problems.
- MATH: A benchmark evaluating a model’s ability to solve complex mathematical problems.
- API-Bank: A benchmark testing the model's accuracy in using APIs to solve tasks.
- BFCL (Benchmark for Commonsense Language): A benchmark designed to measure a model’s ability to understand and use commonsense knowledge.
- Gorilla Benchmark API Bench: A benchmark testing the model’s ability to work with APIs, focusing on more complex or rare tasks.
- Nexus: A benchmark testing the model's macro-average accuracy in performing zero-shot tasks.
- Multilingual MGSM (CoT): A multilingual version of the GSM-8K benchmark, testing the model’s ability to solve math problems in multiple languages.
Metrics:
- macro_avg/acc (Macro Average Accuracy): A metric that averages the accuracy of a model across all classes or tasks.
- acc_char (Accuracy Character): A measure of how accurately a model performs at the character level, typically used in text-based benchmarks.
- em (Exact Match): A metric that measures how often the model's output exactly matches the correct answer.
- f1 (F1 Score): A metric that considers both precision and recall to measure a model's accuracy, particularly useful for imbalanced classes.
- pass@1: A coding metric that measures the model's ability to generate correct code on the first attempt.
- em_maj1@1: A measure of exact match accuracy, focusing on the first major attempt at a task, particularly in complex reasoning or math problems.
- final_em: Final exact match score, often used in more challenging benchmarks like MATH.