Contact Support
    Mistral/Mistral Large 2
    License

    Model Card

    Mistral Large 2 has been tested across a wide range of benchmarks, demonstrating state-of-the-art performance in multiple domains. Here are the results: (for more details read here)

    Base Pretrained Benchmarks

    BenchmarkScore
    MMLU84.00%

    Base Pretrained Multilingual Benchmarks (MMLU)

    BenchmarkScore
    French82.80%
    German81.60%
    Spanish82.70%
    Italian82.70%
    Dutch80.70%
    Portuguese81.60%
    Russian79.00%
    Korean60.10%
    Japanese78.80%
    Chinese74.80%

    Instruction Benchmarks

    BenchmarkScore
    MT Bench8.63
    Wild Bench56.3
    Arena Hard73.2

    Code & Reasoning Benchmarks

    BenchmarkScore
    Human Eval92%
    Human Eval Plus87%
    MBPP Base80%
    MBPP Plus69.00%

    Math Benchmarks

    BenchmarkScore
    GSM8K93%
    Math Instruct (0-shot, no CoT)70%
    Math Instruct (0-shot, CoT)72%

    Benchmarks Glossary

    • MMLU: Measures performance across multiple language understanding tasks.
    • MT Bench: Measures multi-task performance on a variety of tasks.
    • Wild Bench: Benchmark for handling open-ended or unpredictable tasks.
    • Arena Hard: Performance on difficult adversarial tasks.
    • Human Eval: Tests code generation capabilities.
    • Human Eval Plus: Enhanced version of Human Eval with additional challenges.
    • MBPP Base: Code benchmark focusing on programming tasks.
    • MBPP Plus: Advanced programming task benchmark.
    • GSM8K: Benchmark for mathematical problem solving.
    • Math Instruct (0-shot, no CoT): Measures math performance without chain-of-thought reasoning.
    • Math Instruct (0-shot, CoT): Measures math performance with chain-of-thought reasoning.

    Meta data

    128,000 tokens
    $3 per million
    $9 per million
    Create an agent Pipe