llama-3.1-8b-instruct

Text Generation • Meta • Hosted

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models. The Llama 3.1 instruction tuned text only models are optimized for multilingual dialogue use cases and outperform many of the available open source and closed chat models on common industry benchmarks.

Model Info
Planned Deprecation	5/30/2026
Context Window ↗	7,968 tokens
Terms and License	link ↗
Unit Pricing	$0.28 per M input tokens, $0.83 per M output tokens

Playground

Try out this model with Workers AI LLM Playground. It does not require any setup or authentication and an instant way to preview and test a model directly in the browser.

Launch the LLM Playground

Usage

export interface Env {
  AI: Ai;
}

export default {
  async fetch(request, env): Promise<Response> {

    const messages = [
      { role: "system", content: "You are a friendly assistant" },
      {
        role: "user",
        content: "What is the origin of the phrase Hello, World",
      },
    ];

    const stream = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", {
      messages,
      stream: true,
    });

    return new Response(stream, {
      headers: { "content-type": "text/event-stream" },
    });
  },
} satisfies ExportedHandler<Env>;

export interface Env {
  AI: Ai;
}

export default {
  async fetch(request, env): Promise<Response> {

    const messages = [
      { role: "system", content: "You are a friendly assistant" },
      {
        role: "user",
        content: "What is the origin of the phrase Hello, World",
      },
    ];
    const response = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { messages });

    return Response.json(response);
  },
} satisfies ExportedHandler<Env>;

import os
import requests

ACCOUNT_ID = "your-account-id"
AUTH_TOKEN = os.environ.get("CLOUDFLARE_AUTH_TOKEN")

prompt = "Tell me all about PEP-8"
response = requests.post(
  f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/meta/llama-3.1-8b-instruct",
    headers={"Authorization": f"Bearer {AUTH_TOKEN}"},
    json={
      "messages": [
        {"role": "system", "content": "You are a friendly assistant"},
        {"role": "user", "content": prompt}
      ]
    }
)
result = response.json()
print(result)

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/meta/llama-3.1-8b-instruct \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{ "messages": [{ "role": "system", "content": "You are a friendly assistant" }, { "role": "user", "content": "Why is pizza so good" }]}'

Parameters

Input

prompt

stringrequiredminLength: 1The input text prompt for the model to generate a response.

lora

stringName of the LoRA (Low-Rank Adaptation) model to fine-tune the base model.

▶response_format{}

object

raw

booleandefault: falseIf true, a chat template is not applied and you must adhere to the specific model's expected formatting.

stream

booleandefault: falseIf true, the response will be streamed back incrementally using SSE, Server Sent Events.

max_tokens

integerdefault: 256The maximum number of tokens to generate in the response.

temperature

numberdefault: 0.6minimum: 0maximum: 5Controls the randomness of the output; higher values produce more random results.

top_p

numberminimum: 0maximum: 2Adjusts the creativity of the AI's responses by controlling how many possible words it considers. Lower values make outputs more predictable; higher values allow for more varied and creative responses.

top_k

integerminimum: 1maximum: 50Limits the AI to choose from the top 'k' most probable words. Lower values make responses more focused; higher values introduce more variety and potential surprises.

seed

integerminimum: 1maximum: 9999999999Random seed for reproducibility of the generation.

repetition_penalty

numberminimum: 0maximum: 2Penalty for repeated tokens; higher values discourage repetition.

frequency_penalty

numberminimum: 0maximum: 2Decreases the likelihood of the model repeating the same lines verbatim.

presence_penalty

numberminimum: 0maximum: 2Increases the likelihood of the model introducing new topics.

Output

Synchronous — Send a request and receive a complete response

response

stringThe generated text response from the model

▶usage{}

objectUsage statistics for the inference request

▶tool_calls[]

arrayAn array of tool calls requests made during the response generation

Streaming — Send a request with `stream: true` and receive server-sent events

type

string

format

binary

API Schemas (Raw)

Synchronous Input

Synchronous Output

Streaming Input

Streaming Output