Kimi K2 Instruct
Kimi K2 Instruct is Moonshot AI's Mixture-of-Experts (MoE) language model with one trillion total parameters and 32 billion active per forward pass, a context window of 131.1K tokens, available through AI Gateway via Novita AI.
View API reference- Input and output price
- Input $0.57, Output $2.30, Per 1M tokens
- 24h uptime
- Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({ model: 'moonshotai/kimi-k2', prompt: 'Why is the sky blue?'})Copy link to headingPlayground
Try out Kimi K2 Instruct by Moonshot AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Kimi K2 Instruct
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Copy link to headingUptime24 hours
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
Copy link to headingLatency24 hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Getting started
Call Kimi K2 Instruct through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.
Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'moonshotai/kimi-k2', prompt: 'Why is the sky blue?', });
console.log(result.text);}
main().catch(console.error);Top-level parameters
The same Kimi K2 Instruct request in each API format AI Gateway supports.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'moonshotai/kimi-k2', system: 'You are a concise technical assistant.', prompt: 'Summarize the tradeoffs between static generation and SSR.', maxOutputTokens: 1024, temperature: 0.5, });
console.log(result.text);}
main().catch(console.error);Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID in the form creator/model, e.g. moonshotai/kimi-k2. AI Gateway routes the request to an available provider. |
maxOutputTokens | number | No | Hard cap on generated tokens. Kimi K2 Instruct supports up to 131,072 output tokens. |
providerOptions | Record<string, JSONValue> | No | AI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below. |
Input limits
| Input | Formats | Sources | Max count | Max size | Limits |
|---|---|---|---|---|---|
| Text | — | — | — | — | Prompt and response share the 131K-token context window |
Provider options
Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.
Learn more in the AI SDK moonshotai provider docs.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'moonshotai/kimi-k2', prompt: 'Why is the sky blue?', providerOptions: { gateway: { only: ['novita'], }, }, });
console.log(result.text);}
main().catch(console.error);These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.
| Parameter | Type | Required | Description |
|---|---|---|---|
providerOptions.gateway.only | string[] | No | Restrict routing to these provider slugs. Requests fail over only within the listed providers. |
providerOptions.gateway.order | string[] | No | Preferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks. |
providerOptions.gateway.sort | 'cost' | 'ttft' | 'tps' | No | Rank candidate providers by price, time to first token, or tokens per second instead of the default routing order. |
providerOptions.gateway.zeroDataRetention | boolean | No | Route only to providers with a zero-data-retention policy for this model. |
Routing across providers
AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.
Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.
Tool calling
Expose tools the model can call. Define each tool’s inputs with a Zod schema.
import { generateText, tool } from 'ai';import { z } from 'zod';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'moonshotai/kimi-k2', prompt: 'What is the weather in San Francisco?', tools: { getWeather: tool({ description: 'Get the current weather for a location', inputSchema: z.object({ location: z.string() }), execute: async ({ location }) => ({ location, temperatureC: 18 }), }), }, });
console.log(result.text);}
main().catch(console.error);Copy link to headingAbout Kimi K2 Instruct
Kimi K2 Instruct, released July 11, 2025, is a Mixture-of-Experts (MoE) language model from Moonshot AI.
Sparse expert routing at 32B activation. The full trillion parameters encode broad knowledge: programming languages, API conventions, domain facts, and tool-use patterns. At inference time, a routing mechanism selects roughly 32 billion parameters per token. Latency and compute cost stay comparable to a dense 32B model, while the knowledge base spans the entire trillion-parameter budget.
With 32B active parameters for reasoning depth and a full 1T parameter budget encoding broad tool-use and coding knowledge, K2 handles structured sequences of API calls, multi-step planning, and code synthesis.
Kimi K2 Instruct is available through AI Gateway at $0.57 per million input tokens and $2.3 per million output tokens.
AI Gateway routes K2 across Novita AI, giving you automatic failover across multiple providers.
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: K2 routes across Novita AI. Choose it when uptime and provider redundancy matter most.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use Kimi K2 Instruct
Best for
- Agentic pipelines: Structured sequences of API calls, data processing, and code synthesis
- Provider redundancy: Deployments where failover across multiple providers matters most
- K2 architecture baseline: Teams evaluating the K2 architecture for the first time who want the original release
- Broad knowledge at low cost: Workloads that benefit from trillion-parameter knowledge breadth at 32B-dense inference economics
Consider alternatives when
- Chain-of-thought traces: Kimi K2 Thinking layers extended reasoning on top of this foundation
- Minimum latency: Kimi K2 Turbo is the speed-optimized variant
- September 2025 checkpoint: Use Kimi K2-0905 for expanded context and refined agentic training
- Multimodal inputs: K2 processes text only, so reach for a vision-capable model
Copy link to headingConclusion
Kimi K2 Instruct established that sparse expert routing can deliver dense-model responsiveness at trillion-parameter scale. Its architecture anchors the entire K2 family of specialized variants. Routing across Novita AI gives you automatic failover for high-availability production.
Copy link to headingFrequently Asked Questions
How does sparse routing translate to cost savings?
The full 1T parameters store broad knowledge, but only ~32B activate per token via the expert router. You pay compute proportional to a 32B dense model while drawing on knowledge encoded across the entire trillion-parameter budget.
Why does base K2 list many providers on AI Gateway?
It was the first K2 variant adopted across providers, so routing across Novita AI reflects earlier integration. Later checkpoints and variants can have narrower provider sets.
Is K2 text-only?
Yes. Kimi K2 Instruct accepts and produces text. Multimodal capabilities are not part of this release.
What agentic patterns does K2 handle well?
Structured multi-step sequences: invoke an API, parse the response, branch on results, call a second API, and synthesize a final output. The function-calling interface in AI Gateway maps directly to these workflows.
Can I bring my own provider credentials?
Yes. AI Gateway supports Bring Your Own Key for providers where you hold a direct account. BYOK requests are excluded from ZDR coverage.
Your use is subject to Moonshot AI's Terms & Privacy Policies.