Llama 4 Maverick 17B 128E Instruct FP8
Llama 4 Maverick 17B 128E Instruct FP8 is Meta's natively multimodal Mixture of Experts (MoE) model with 17B active parameters across 128 experts. Published benchmarks span image and text tasks, and the MoE activates a fraction of the parameters that comparable dense models use.
View API reference- Input and output price
- Prices from: Input $0.20, Output $0.80, Per 1M tokens
- 24h uptime
- Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({ model: 'meta/llama-4-maverick', prompt: 'Why is the sky blue?'})Copy link to headingPlayground
Try out Llama 4 Maverick 17B 128E Instruct FP8 by Meta. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Llama 4 Maverick 17B 128E Instruct FP8
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Copy link to headingUptime24 hours
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
Copy link to headingLatency24 hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Getting started
Call Llama 4 Maverick 17B 128E Instruct FP8 through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.
Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'meta/llama-4-maverick', prompt: 'Why is the sky blue?', });
console.log(result.text);}
main().catch(console.error);Top-level parameters
The same Llama 4 Maverick 17B 128E Instruct FP8 request in each API format AI Gateway supports.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'meta/llama-4-maverick', system: 'You are a concise technical assistant.', prompt: 'Summarize the tradeoffs between static generation and SSR.', maxOutputTokens: 1024, temperature: 0.5, });
console.log(result.text);}
main().catch(console.error);Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID in the form creator/model, e.g. meta/llama-4-maverick. AI Gateway routes the request to an available provider. |
maxOutputTokens | number | No | Hard cap on generated tokens. Llama 4 Maverick 17B 128E Instruct FP8 supports up to 8,192 output tokens. |
providerOptions | Record<string, JSONValue> | No | AI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below. |
Input limits
| Input | Formats | Sources | Max count | Max size | Limits |
|---|---|---|---|---|---|
| Text | — | — | — | — | Prompt and response share the 131K-token context window |
| Image | — | URL, base64, Uint8Array | — | — | Sent as image parts in messages; counts as input tokens |
Provider options
Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.
Learn more in the AI SDK provider docs.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'meta/llama-4-maverick', prompt: 'Why is the sky blue?', providerOptions: { gateway: { only: ['deepinfra', 'bedrock'], }, }, });
console.log(result.text);}
main().catch(console.error);These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.
| Parameter | Type | Required | Description |
|---|---|---|---|
providerOptions.gateway.only | string[] | No | Restrict routing to these provider slugs. Requests fail over only within the listed providers. |
providerOptions.gateway.order | string[] | No | Preferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks. |
providerOptions.gateway.sort | 'cost' | 'ttft' | 'tps' | No | Rank candidate providers by price, time to first token, or tokens per second instead of the default routing order. |
providerOptions.gateway.zeroDataRetention | boolean | No | Route only to providers with a zero-data-retention policy for this model. |
Routing across providers
AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.
Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.
Image input
Send images alongside text as message parts. Images count as input tokens.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'meta/llama-4-maverick', messages: [ { role: 'user', content: [ { type: 'text', text: 'Describe this image.' }, { type: 'image', image: 'https://example.com/photo.jpg' }, ], }, ], });
console.log(result.text);}
main().catch(console.error);Tool calling
Expose tools the model can call. Define each tool’s inputs with a Zod schema.
import { generateText, tool } from 'ai';import { z } from 'zod';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'meta/llama-4-maverick', prompt: 'What is the weather in San Francisco?', tools: { getWeather: tool({ description: 'Get the current weather for a location', inputSchema: z.object({ location: z.string() }), execute: async ({ location }) => ({ location, temperatureC: 18 }), }), }, });
console.log(result.text);}
main().catch(console.error);Copy link to headingAbout Llama 4 Maverick 17B 128E Instruct FP8
Meta released Llama 4 Maverick 17B 128E Instruct FP8 on April 5, 2025 as one of the first two models in the Llama 4 generation. The collection is built around two architectural advances: native multimodality through early fusion, and Mixture of Experts (MoE). Llama 4 Maverick 17B 128E Instruct FP8 is the larger and more capable of the two initial releases, with 17 billion active parameters, 128 routed experts plus one shared expert, and 400 billion total parameters. Each token activates only 17B of those 400B parameters (the shared expert plus one routed expert). This makes inference substantially more efficient than a dense 400B model while preserving the quality benefits of the larger total parameter budget.
Llama 4's native multimodality represents a different architectural approach from the adapter-based vision in Llama 3.2. Rather than adding image understanding to an existing text backbone, Llama 4 treats text and vision tokens together from the beginning in a unified backbone. This enables more coherent cross-modal reasoning.
On the LMArena leaderboard, an experimental chat version of Llama 4 Maverick 17B 128E Instruct FP8 scored an Elo of 1417. Llama 4 Maverick 17B 128E Instruct FP8 exceeds comparable frontier models on coding, reasoning, multilingual, long-context, and image benchmarks. It achieves results comparable to other open-weight models on reasoning and coding at less than half the active parameters.
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: For workloads that mix images and long text, Llama 4 Maverick 17B 128E Instruct FP8's efficiency advantage over dense models shows most at scale. Validate throughput at your expected concurrency level before you pick a provider tier. Compare $0.24 and $0.97.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use Llama 4 Maverick 17B 128E Instruct FP8
Best for
- Production multimodal applications: Pairing image understanding with long-form text generation for product catalog processing and document analysis with mixed visual and textual content
- Creative and coding workloads: Multilingual applications where the MoE architecture reaches dense-model scores on published benchmarks at lower active-parameter cost
- Cost-per-quality sensitive workloads: Comparable-capability dense models are significantly more expensive to serve
- Long-context multimodal tasks: Image and text reasoning must be maintained coherently across extended conversations
- General assistant and chat: Meta designates Llama 4 Maverick 17B 128E Instruct FP8 as the intended product workhorse
Consider alternatives when
- Extreme long documents: Llama 4 Scout's 10M token context window is purpose-built for that use case
- Text-only workload: The MoE overhead of loading all experts into memory is not offset by quality gains over a dense model at similar cost
- Maximum reasoning depth: Llama 4 Behemoth (when available) or other frontier reasoning models may be appropriate
Copy link to headingConclusion
Llama 4 Maverick 17B 128E Instruct FP8 combines native multimodality, a 128-expert MoE architecture, and strong benchmark results on image and text tasks at a fraction of the active-parameter cost of dense alternatives. For teams building multimodal production applications on open models, Llama 4 Maverick 17B 128E Instruct FP8 is the more capable of the two initial Llama 4 releases.
Copy link to headingFrequently Asked Questions
What is Mixture of Experts (MoE) and how does it work in Llama 4 Maverick 17B 128E Instruct FP8?
Each input token activates only a subset of the total parameters. Llama 4 Maverick 17B 128E Instruct FP8 uses alternating dense and MoE layers. MoE layers route each token to a shared expert plus one of 128 routed experts. Only 17B of the 400B total parameters are active per token, reducing inference cost while the full parameter budget contributes to model quality.
What does "natively multimodal" mean compared to the adapter-based vision in Llama 3.2?
Llama 3.2 added vision to an existing text backbone via cross-attention adapters, keeping language model weights frozen. Llama 4 Maverick 17B 128E Instruct FP8 processes text and vision tokens together in a unified backbone. This enables deeper cross-modal reasoning because the model was never strictly text-only.
What Elo score did Llama 4 Maverick 17B 128E Instruct FP8 achieve on LMArena?
An experimental chat version of Llama 4 Maverick 17B 128E Instruct FP8 scored an Elo of 1417 on LMArena.
What languages does Llama 4 support?
Llama 4 supports 200 languages, including over 100 with more than 1 billion tokens each, representing 10x more multilingual coverage than Llama 3.