Contact Support
    OpenAI/o1-preview

    Model Card

    OpenAI-o1-preview has been tested across a wide range of benchmarks, demonstrating state-of-the-art performance in multiple domains. Here are the results: (more details here)

    BenchmarkGPT-4oo1-previewo1Expert Human
    Competition Math (AIME 2024)13.40%56.70%83.30%-
    Competition Code (Codeforces)11.00%62.00%89.00%-
    PhD-Level Science Questions (GPQA Diamond)56.10%78.30%78.00%69.70%

    ML Benchmarks

    BenchmarkGPT-4oo1 improvement
    MATH60.394.8
    MathVista (testmini)63.873.2
    MMMU (val)69.178.1
    MMLU8892.3

    PhD-Level Science Questions (GPQA Diamond)

    BenchmarkGPT-4oo1 improvement
    Chemistry40.264.7
    Physics59.592.8
    Biology61.669.2

    Benchmarks Glossary

    • Competition Math (AIME 2024): Measures accuracy in advanced math problems.
    • Competition Code (Codeforces): Evaluates programming skills using Elo ratings.
    • PhD-Level Science Questions (GPQA Diamond): Assesses performance on complex science questions.
    • MATH: Benchmark for solving mathematical problems.
    • MathVista (testmini): Tests performance on mathematical reasoning.
    • MMMU (val): Evaluates understanding across various multi-modal tasks.
    • MMLU: Measures general language understanding across multiple languages.
    • Chemistry/Physics/Biology: PhD-level science problem-solving ability.

    Meta data

    128,000 tokens
    $15 per million
    $60 per million
    Oct 2023
    Create an agent Pipe