Contact Support
    Meta/
    Llama-3.1-8B-Instruct
    License

    Model Card

    The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction-tuned generative models in 8B, 70B, and 405B sizes (text in/text out). The Llama 3.1 instruction-tuned text-only models (8B, 70B, 405B) are optimized for multilingual dialogue use cases and outperform many of the available open-source and closed chat models on common industry benchmarks.

    Model Architecture: Llama 3.1 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety.

    Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.

    Llama 3.1 family of models: Token counts refer to pretraining data only. All model versions use Grouped-Query Attention (GQA) for improved inference scalability.

    Intended Use

    Intended Use Cases: Llama 3.1 is intended for commercial and research use in multiple languages. Instruction-tuned text-only models are intended for assistant-like chat, whereas pretrained models can be adapted for a variety of natural language generation tasks. The Llama 3.1 model collection also supports the ability to leverage the outputs of its models to improve other models, including synthetic data generation and distillation. The Llama 3.1 Community License allows for these use cases.

    Key Features

    • Multilingual capabilities: Supports English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai, optimized for multilingual dialogue use cases.
    • Efficient architecture: Built on an optimized transformer architecture with 8 billion parameters for high performance in natural language generation tasks.
    • Instruction tuned: Uses Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) to align with human preferences for helpfulness and safety.
    • High context capacity: Handles a context length of up to 131k tokens, enabling the processing of more complex and lengthy inputs.
    • Scalable inference: Employs Grouped-Query Attention (GQA) for improved inference scalability, enhancing processing efficiency.
    • Environmental responsibility: Developed with a focus on sustainability, contributing to Meta's net-zero greenhouse gas emissions goal in global operations.
    • Tool integration: Supports tool use and integration with third-party services, allowing for versatile application in various domains.

    Meta data

    131,072 tokens
    $0.2 per million
    $0.2 per million
    Dec 2023
    Create an agent Pipe