Skip to content
Dashboard

Llama 4 Maverick 17B 128E Instruct FP8

Llama 4 Maverick 17B 128E Instruct FP8 is Meta's natively multimodal Mixture of Experts (MoE) model with 17B active parameters across 128 experts. Published benchmarks span image and text tasks, and the MoE activates a fraction of the parameters that comparable dense models use.

View API reference
Input and output price
Prices from: Input $0.20, Output $0.80, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'meta/llama-4-maverick',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingPlayground

Try out Llama 4 Maverick 17B 128E Instruct FP8 by Meta. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

meta logo
meta logo

Llama 4 Maverick 17B 128E Instruct FP8

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Free Tier
Release Date
131K8K0.2 s78 tps
$0.20/M
$0.80/M
04/05/2025
128K8K0.2 s
$0.24/M
$0.97/M
04/05/2025

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Getting started

Call Llama 4 Maverick 17B 128E Instruct FP8 through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.

Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.

index.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'meta/llama-4-maverick',
prompt: 'Why is the sky blue?',
});
console.log(result.text);
}
main().catch(console.error);

Top-level parameters

The same Llama 4 Maverick 17B 128E Instruct FP8 request in each API format AI Gateway supports.

top-level-params.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'meta/llama-4-maverick',
system: 'You are a concise technical assistant.',
prompt: 'Summarize the tradeoffs between static generation and SSR.',
maxOutputTokens: 1024,
temperature: 0.5,
});
console.log(result.text);
}
main().catch(console.error);

Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.

ParameterTypeRequiredDescription
modelstringYesModel ID in the form creator/model, e.g. meta/llama-4-maverick. AI Gateway routes the request to an available provider.
maxOutputTokensnumberNoHard cap on generated tokens. Llama 4 Maverick 17B 128E Instruct FP8 supports up to 8,192 output tokens.
providerOptionsRecord<string, JSONValue>NoAI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below.

Input limits

InputFormatsSourcesMax countMax sizeLimits
TextPrompt and response share the 131K-token context window
ImageURL, base64, Uint8ArraySent as image parts in messages; counts as input tokens

Provider options

Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.

Learn more in the AI SDK provider docs.

provider-options.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'meta/llama-4-maverick',
prompt: 'Why is the sky blue?',
providerOptions: {
gateway: {
only: ['deepinfra', 'bedrock'],
},
},
});
console.log(result.text);
}
main().catch(console.error);

These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.

ParameterTypeRequiredDescription
providerOptions.gateway.onlystring[]NoRestrict routing to these provider slugs. Requests fail over only within the listed providers.
providerOptions.gateway.orderstring[]NoPreferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks.
providerOptions.gateway.sort'cost' | 'ttft' | 'tps'NoRank candidate providers by price, time to first token, or tokens per second instead of the default routing order.
providerOptions.gateway.zeroDataRetentionbooleanNoRoute only to providers with a zero-data-retention policy for this model.

Routing across providers

AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.

Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.

Image input

Send images alongside text as message parts. Images count as input tokens.

image-input.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'meta/llama-4-maverick',
messages: [
{
role: 'user',
content: [
{ type: 'text', text: 'Describe this image.' },
{ type: 'image', image: 'https://example.com/photo.jpg' },
],
},
],
});
console.log(result.text);
}
main().catch(console.error);

Tool calling

Expose tools the model can call. Define each tool’s inputs with a Zod schema.

tool-calling.ts
import { generateText, tool } from 'ai';
import { z } from 'zod';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'meta/llama-4-maverick',
prompt: 'What is the weather in San Francisco?',
tools: {
getWeather: tool({
description: 'Get the current weather for a location',
inputSchema: z.object({ location: z.string() }),
execute: async ({ location }) => ({ location, temperatureC: 18 }),
}),
},
});
console.log(result.text);
}
main().catch(console.error);

Copy link to headingMore models by Meta

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
$0.01/img
meta logo
08/26/2026
1M5.1 s130 tps
$0.10/M
$0.20/M
Read$0.002/M
+2
meta logo
08/05/2026
1M4.0 s149 tps
$1.25/M
$4.25/M
Read$0.15/M
+2
meta logo
08/05/2026
1M2.8 s169 tps
$1.25/M
$4.25/M
Read$0.15/M
+2
meta logo
07/09/2026
1M4.1 s91 tps
$0.10/M
$0.20/M
Read$0.002/M
+2
meta logo
1M4.7 s94 tps
$1.25/M
$4.25/M
Read$0.15/M
+2
meta logo

Copy link to headingAbout Llama 4 Maverick 17B 128E Instruct FP8

Meta released Llama 4 Maverick 17B 128E Instruct FP8 on April 5, 2025 as one of the first two models in the Llama 4 generation. The collection is built around two architectural advances: native multimodality through early fusion, and Mixture of Experts (MoE). Llama 4 Maverick 17B 128E Instruct FP8 is the larger and more capable of the two initial releases, with 17 billion active parameters, 128 routed experts plus one shared expert, and 400 billion total parameters. Each token activates only 17B of those 400B parameters (the shared expert plus one routed expert). This makes inference substantially more efficient than a dense 400B model while preserving the quality benefits of the larger total parameter budget.

Llama 4's native multimodality represents a different architectural approach from the adapter-based vision in Llama 3.2. Rather than adding image understanding to an existing text backbone, Llama 4 treats text and vision tokens together from the beginning in a unified backbone. This enables more coherent cross-modal reasoning.

On the LMArena leaderboard, an experimental chat version of Llama 4 Maverick 17B 128E Instruct FP8 scored an Elo of 1417. Llama 4 Maverick 17B 128E Instruct FP8 exceeds comparable frontier models on coding, reasoning, multilingual, long-context, and image benchmarks. It achieves results comparable to other open-weight models on reasoning and coding at less than half the active parameters.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: For workloads that mix images and long text, Llama 4 Maverick 17B 128E Instruct FP8's efficiency advantage over dense models shows most at scale. Validate throughput at your expected concurrency level before you pick a provider tier. Compare $0.24 and $0.97.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Llama 4 Maverick 17B 128E Instruct FP8

Best for

  • Production multimodal applications: Pairing image understanding with long-form text generation for product catalog processing and document analysis with mixed visual and textual content
  • Creative and coding workloads: Multilingual applications where the MoE architecture reaches dense-model scores on published benchmarks at lower active-parameter cost
  • Cost-per-quality sensitive workloads: Comparable-capability dense models are significantly more expensive to serve
  • Long-context multimodal tasks: Image and text reasoning must be maintained coherently across extended conversations
  • General assistant and chat: Meta designates Llama 4 Maverick 17B 128E Instruct FP8 as the intended product workhorse

Consider alternatives when

  • Extreme long documents: Llama 4 Scout's 10M token context window is purpose-built for that use case
  • Text-only workload: The MoE overhead of loading all experts into memory is not offset by quality gains over a dense model at similar cost
  • Maximum reasoning depth: Llama 4 Behemoth (when available) or other frontier reasoning models may be appropriate

Llama 4 Maverick 17B 128E Instruct FP8 combines native multimodality, a 128-expert MoE architecture, and strong benchmark results on image and text tasks at a fraction of the active-parameter cost of dense alternatives. For teams building multimodal production applications on open models, Llama 4 Maverick 17B 128E Instruct FP8 is the more capable of the two initial Llama 4 releases.

Copy link to headingFrequently Asked Questions

  • What is Mixture of Experts (MoE) and how does it work in Llama 4 Maverick 17B 128E Instruct FP8?

    Each input token activates only a subset of the total parameters. Llama 4 Maverick 17B 128E Instruct FP8 uses alternating dense and MoE layers. MoE layers route each token to a shared expert plus one of 128 routed experts. Only 17B of the 400B total parameters are active per token, reducing inference cost while the full parameter budget contributes to model quality.

  • What does "natively multimodal" mean compared to the adapter-based vision in Llama 3.2?

    Llama 3.2 added vision to an existing text backbone via cross-attention adapters, keeping language model weights frozen. Llama 4 Maverick 17B 128E Instruct FP8 processes text and vision tokens together in a unified backbone. This enables deeper cross-modal reasoning because the model was never strictly text-only.

  • What Elo score did Llama 4 Maverick 17B 128E Instruct FP8 achieve on LMArena?

    An experimental chat version of Llama 4 Maverick 17B 128E Instruct FP8 scored an Elo of 1417 on LMArena.

  • What languages does Llama 4 support?

    Llama 4 supports 200 languages, including over 100 with more than 1 billion tokens each, representing 10x more multilingual coverage than Llama 3.

Your use is subject to Meta's Terms & Privacy Policies.