Command R+ has demonstrated exceptional performance across several key benchmarks that are critical for enterprise use cases. Here are the results (for more details read here):
LLMS Performance on Azure
| Capability (in %) | Command R+ | Mistral-Large | GPT-4 Turbo |
|---|---|---|---|
| Multilingual (BLEU) | 35.9 | 31.4 | 36.6 |
| RAG (Average Accuracy) | 71.5 | 60.7 | 76.7 |
| Tool Use (Average Success Rate) | 74.5 | 63.1 | 73.7 |
LLMS Pricing on Azure
| Token Type | Command R+ | Mistral-Large | GPT-4 Turbo |
|---|---|---|---|
| Input ($/M tokens) | $3.00 | $8.00 | $10.00 |
| Output ($/M tokens) | $15.00 | $24.00 | $30.00 |
Human Preference Evaluation for Summarization with Citations
| Model Comparison | Command R+ Win Rate | Other Model Win Rate |
|---|---|---|
| vs. Claude 3 Sonnet | 74% | 26% |
| vs. GPT-4 Turbo | 53% | 47% |
Multi-Step Reasoning with Search Tools
| Model | Bamboogle Accuracy | HotpotQA Accuracy | StrategyQA Accuracy |
|---|---|---|---|
| Command R+ | 76.20% | 75.60% | 71.20% |
| Claude 3 Sonnet | 64.20% | 67.20% | 69.90% |
| Mistral-Large | 59.00% | 61.60% | 60.10% |
| GPT-4 Turbo | 75.20% | 63.10% | 79.50% |
Conversational Agent Evaluation (ToolTalk Hard Benchmark)
| Model | ToolTalk Soft Success Rate |
|---|---|
| Command R+ | 71.10% |
| Claude 3 Sonnet | 56% |
| Mistral-Large | 57% |
| GPT-4 Turbo | 69.80% |
Function Calling Evaluation (Berkeley Function Calling Leaderboard)
| Model | Function Pass Rate |
|---|---|
| Command R+ | 78.00% |
| Claude 3 Sonnet | 77% |
| Mistral-Large | 69% |
| GPT-4 Turbo | 77.60% |
Multilingual Evaluations (Translation Quality - BLEU Score)
| Model | FLORES (en->L2) | FLORES (L2->en) | WMT23 (en->L2) | WMT23 (L2->en) |
|---|---|---|---|---|
| Command R+ | 37.7 | 39.6 | 36.1 | 30.3 |
| Claude 3 Sonnet | 35.4 | 37.1 | 33.2 | 31.3 |
| Mistral-Large | 32.3 | 34.2 | 29.7 | 29.2 |
| GPT-4 Turbo | 38.2 | 37.4 | 38.4 | 32.5 |
Multilingual Token Cost (Relative to Cohere Tokenizer)
| Language | Cohere Tokenizer | Mistral Tokenizer | OpenAI Tokenizer |
|---|---|---|---|
| French | 1 | 1.12 | 1.27 |
| Spanish | 1 | 1.14 | 1.32 |
| German | 1 | 1.14 | 1.27 |
| Italian | 1 | 1.15 | 1.28 |
| Portuguese | 1 | 1.18 | 1.38 |
| Chinese | 1 | 1.18 | 1.5 |
| Korean | 1 | 1.67 | 1.92 |
| Japanese | 1 | 1.67 | 1.77 |
| Arabic | 1 | 1.85 | 2.33 |
Benchmarks Glossary
- Bamboogle Accuracy: Measures the accuracy of a model in retrieving and using information from the web in a multi-hop reasoning setup.
- HotpotQA Accuracy: Assesses the model's ability to answer questions requiring multi-hop reasoning across different documents.
- StrategyQA Accuracy: Evaluates the model's accuracy in answering yes/no questions that require implicit multi-hop reasoning.
- ToolTalk Soft Success Rate: Percentage of successful tool use conversations, where the model correctly identifies and uses the appropriate tools without unintended side effects.
- Function Pass Rate: The success rate of function-calling tasks, where the model correctly executes functions without errors.
- BLEU Score: A metric used to evaluate the quality of machine-generated text, particularly in translation, by comparing it to reference translations.
- FLORES (en->L2): BLEU score measuring translation quality from English to a target language (L2).
- FLORES (L2->en): BLEU score measuring translation quality from a target language (L2) to English.
- WMT23 (en->L2): BLEU score evaluating translation quality from English to a target language (L2) in the WMT23 translation task.
- WMT23 (L2->en): BLEU score evaluating translation quality from a target language (L2) to English in the WMT23 translation task.