Contact Support
    Cohere/Command-R-plus
    License

    Model Card

    Command R+ has demonstrated exceptional performance across several key benchmarks that are critical for enterprise use cases. Here are the results (for more details read here):

    LLMS Performance on Azure

    Capability (in %)Command R+Mistral-LargeGPT-4 Turbo
    Multilingual (BLEU)35.931.436.6
    RAG (Average Accuracy)71.560.776.7
    Tool Use (Average Success Rate)74.563.173.7

    LLMS Pricing on Azure

    Token TypeCommand R+Mistral-LargeGPT-4 Turbo
    Input ($/M tokens)$3.00$8.00$10.00
    Output ($/M tokens)$15.00$24.00$30.00

    Human Preference Evaluation for Summarization with Citations

    Model ComparisonCommand R+ Win RateOther Model Win Rate
    vs. Claude 3 Sonnet74%26%
    vs. GPT-4 Turbo53%47%

    Multi-Step Reasoning with Search Tools

    ModelBamboogle AccuracyHotpotQA AccuracyStrategyQA Accuracy
    Command R+76.20%75.60%71.20%
    Claude 3 Sonnet64.20%67.20%69.90%
    Mistral-Large59.00%61.60%60.10%
    GPT-4 Turbo75.20%63.10%79.50%

    Conversational Agent Evaluation (ToolTalk Hard Benchmark)

    ModelToolTalk Soft Success Rate
    Command R+71.10%
    Claude 3 Sonnet56%
    Mistral-Large57%
    GPT-4 Turbo69.80%

    Function Calling Evaluation (Berkeley Function Calling Leaderboard)

    ModelFunction Pass Rate
    Command R+78.00%
    Claude 3 Sonnet77%
    Mistral-Large69%
    GPT-4 Turbo77.60%

    Multilingual Evaluations (Translation Quality - BLEU Score)

    ModelFLORES (en->L2)FLORES (L2->en)WMT23 (en->L2)WMT23 (L2->en)
    Command R+37.739.636.130.3
    Claude 3 Sonnet35.437.133.231.3
    Mistral-Large32.334.229.729.2
    GPT-4 Turbo38.237.438.432.5

    Multilingual Token Cost (Relative to Cohere Tokenizer)

    LanguageCohere TokenizerMistral TokenizerOpenAI Tokenizer
    French11.121.27
    Spanish11.141.32
    German11.141.27
    Italian11.151.28
    Portuguese11.181.38
    Chinese11.181.5
    Korean11.671.92
    Japanese11.671.77
    Arabic11.852.33

    Benchmarks Glossary

    • Bamboogle Accuracy: Measures the accuracy of a model in retrieving and using information from the web in a multi-hop reasoning setup.
    • HotpotQA Accuracy: Assesses the model's ability to answer questions requiring multi-hop reasoning across different documents.
    • StrategyQA Accuracy: Evaluates the model's accuracy in answering yes/no questions that require implicit multi-hop reasoning.
    • ToolTalk Soft Success Rate: Percentage of successful tool use conversations, where the model correctly identifies and uses the appropriate tools without unintended side effects.
    • Function Pass Rate: The success rate of function-calling tasks, where the model correctly executes functions without errors.
    • BLEU Score: A metric used to evaluate the quality of machine-generated text, particularly in translation, by comparing it to reference translations.
    • FLORES (en->L2): BLEU score measuring translation quality from English to a target language (L2).
    • FLORES (L2->en): BLEU score measuring translation quality from a target language (L2) to English.
    • WMT23 (en->L2): BLEU score evaluating translation quality from English to a target language (L2) in the WMT23 translation task.
    • WMT23 (L2->en): BLEU score evaluating translation quality from a target language (L2) to English in the WMT23 translation task.

    Meta data

    128,000 tokens
    $3 per million
    $15 per million
    Create an agent Pipe