Groq logoGroq

    llama-3.1-8b-instant

    groq/llama-3.1-8b-instant

    ChatToolsSystem prompt

    Context

    131K

    Max output

    8K

    Input / 1M

    $0.05

    Output / 1M

    $0.08

    Modality

    Chat

    Cutoff

    Dec 2023

    About this model

    Llama 3.1 8B on Groq provides low-latency, high-quality responses suitable for real-time conversational interfaces, content filtering systems, and data analysis applications. This model offers a balance of speed and performance with significant cost savings compared to larger models. Technical capabilities include native function calling support, JSON mode for structured output generation, and a 128K token context window for handling large documents.

    Best suited for

    • Low-latency chat interfaces, edge or device deployments, mobile applications, simple content generation, classification, and high-throughput, cost-efficient workloads.

    Built-in tools

    Gtwy web search

    Hosted by the gateway — enable them per request without wiring your own endpoint.

    Capabilities

    Tools

    Native function calling, so agents can invoke your endpoints.

    System prompt

    Honours a dedicated system role, separate from the user turn.

    Supported parameters

    creativity_level

    Controls the creativity of responses. Higher values (e.g., 0.7) increase creativity; lower values (e.g., 0.2) make responses more predictable.

    max_tokensMax Tokens Limit

    Specifies the maximum number of text units (tokens) allowed in a response, limiting its length.

    probability_cutoffProbability Cutoff (Top P)

    Focuses on the most likely words based on a percentage of probability.

    log_probability

    If true, returns the log probabilities of each output token returned in the content of message.

    repetition_penalty

    The `frequency_penalty` controls how often the model repeats itself, with higher positive values reducing repetition and negative values encouraging it.

    novelty_penalty

    Discourages responses that are too similar to previous ones.

    stop

    This parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)

    tools

    Lists tool definitions or capabilities available to the model.

    tool_choice

    Decides whether to use tools or just the model for generating responses.

    response_type

    Defines the format or type of the generated response.

    stream

    Sends the response in real-time as it's being generated.