Skip to content
Dashboard

Kimi K2 Instruct

Kimi K2 Instruct is Moonshot AI's Mixture-of-Experts (MoE) language model with one trillion total parameters and 32 billion active per forward pass, a context window of 131.1K tokens, available through AI Gateway via Novita AI.

View API reference
Input and output price
Input $0.57, Output $2.30, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'moonshotai/kimi-k2',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingPlayground

Try out Kimi K2 Instruct by Moonshot AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

moonshotai logo
moonshotai logo

Kimi K2 Instruct

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Free Tier
Release Date
131K131K1.2 s40 tps
$0.57/M
$2.30/M
07/11/2025

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Getting started

Call Kimi K2 Instruct through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.

Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.

index.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'moonshotai/kimi-k2',
prompt: 'Why is the sky blue?',
});
console.log(result.text);
}
main().catch(console.error);

Top-level parameters

The same Kimi K2 Instruct request in each API format AI Gateway supports.

top-level-params.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'moonshotai/kimi-k2',
system: 'You are a concise technical assistant.',
prompt: 'Summarize the tradeoffs between static generation and SSR.',
maxOutputTokens: 1024,
temperature: 0.5,
});
console.log(result.text);
}
main().catch(console.error);

Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.

ParameterTypeRequiredDescription
modelstringYesModel ID in the form creator/model, e.g. moonshotai/kimi-k2. AI Gateway routes the request to an available provider.
maxOutputTokensnumberNoHard cap on generated tokens. Kimi K2 Instruct supports up to 131,072 output tokens.
providerOptionsRecord<string, JSONValue>NoAI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below.

Input limits

InputFormatsSourcesMax countMax sizeLimits
TextPrompt and response share the 131K-token context window

Provider options

Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.

Learn more in the AI SDK moonshotai provider docs.

provider-options.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'moonshotai/kimi-k2',
prompt: 'Why is the sky blue?',
providerOptions: {
gateway: {
only: ['novita'],
},
},
});
console.log(result.text);
}
main().catch(console.error);

These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.

ParameterTypeRequiredDescription
providerOptions.gateway.onlystring[]NoRestrict routing to these provider slugs. Requests fail over only within the listed providers.
providerOptions.gateway.orderstring[]NoPreferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks.
providerOptions.gateway.sort'cost' | 'ttft' | 'tps'NoRank candidate providers by price, time to first token, or tokens per second instead of the default routing order.
providerOptions.gateway.zeroDataRetentionbooleanNoRoute only to providers with a zero-data-retention policy for this model.

Routing across providers

AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.

Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.

Tool calling

Expose tools the model can call. Define each tool’s inputs with a Zod schema.

tool-calling.ts
import { generateText, tool } from 'ai';
import { z } from 'zod';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'moonshotai/kimi-k2',
prompt: 'What is the weather in San Francisco?',
tools: {
getWeather: tool({
description: 'Get the current weather for a location',
inputSchema: z.object({ location: z.string() }),
execute: async ({ location }) => ({ location, temperatureC: 18 }),
}),
},
});
console.log(result.text);
}
main().catch(console.error);

Copy link to headingMore models by Moonshot AI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M1.6 s122 tps
$4.50/M
$22.50/M
Read$0.45/M
+2
fireworks logo
morph logo
07/27/2026
1M0.6 s163 tps
$2.50/M+1 more
$12.75/M+1 more
Read$0.29/M
+2
alibaba logo
baseten logo
blackbox logo
+11
07/16/2026
262K2.1 s181 tps
$1.90/M
$8/M
Read$0.38/M
+2
moonshotai logo
06/15/2026
262K0.8 s90 tps
$0.74/M+1 more
$3.50/M+1 more
Read$0.15/M
+2
baseten logo
deepinfra logo
fireworks logo
+1
06/12/2026
262K0.4 s131 tps
$0.95/M
$4/M
Read$0.16/M
+1
baseten logo
fireworks logo
moonshotai logo
+1
04/20/2026
262K0.7 s52 tps
$0.60/M
$3/M
Read$0.10/M
+1
bedrock logo
moonshotai logo
novita logo
01/26/2026

Copy link to headingAbout Kimi K2 Instruct

Kimi K2 Instruct, released July 11, 2025, is a Mixture-of-Experts (MoE) language model from Moonshot AI.

Sparse expert routing at 32B activation. The full trillion parameters encode broad knowledge: programming languages, API conventions, domain facts, and tool-use patterns. At inference time, a routing mechanism selects roughly 32 billion parameters per token. Latency and compute cost stay comparable to a dense 32B model, while the knowledge base spans the entire trillion-parameter budget.

With 32B active parameters for reasoning depth and a full 1T parameter budget encoding broad tool-use and coding knowledge, K2 handles structured sequences of API calls, multi-step planning, and code synthesis.

Kimi K2 Instruct is available through AI Gateway at $0.57 per million input tokens and $2.3 per million output tokens.

AI Gateway routes K2 across Novita AI, giving you automatic failover across multiple providers.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: K2 routes across Novita AI. Choose it when uptime and provider redundancy matter most.
  • Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Kimi K2 Instruct

Best for

  • Agentic pipelines: Structured sequences of API calls, data processing, and code synthesis
  • Provider redundancy: Deployments where failover across multiple providers matters most
  • K2 architecture baseline: Teams evaluating the K2 architecture for the first time who want the original release
  • Broad knowledge at low cost: Workloads that benefit from trillion-parameter knowledge breadth at 32B-dense inference economics

Consider alternatives when

  • Chain-of-thought traces: Kimi K2 Thinking layers extended reasoning on top of this foundation
  • Minimum latency: Kimi K2 Turbo is the speed-optimized variant
  • September 2025 checkpoint: Use Kimi K2-0905 for expanded context and refined agentic training
  • Multimodal inputs: K2 processes text only, so reach for a vision-capable model

Kimi K2 Instruct established that sparse expert routing can deliver dense-model responsiveness at trillion-parameter scale. Its architecture anchors the entire K2 family of specialized variants. Routing across Novita AI gives you automatic failover for high-availability production.

Copy link to headingFrequently Asked Questions

  • How does sparse routing translate to cost savings?

    The full 1T parameters store broad knowledge, but only ~32B activate per token via the expert router. You pay compute proportional to a 32B dense model while drawing on knowledge encoded across the entire trillion-parameter budget.

  • Why does base K2 list many providers on AI Gateway?

    It was the first K2 variant adopted across providers, so routing across Novita AI reflects earlier integration. Later checkpoints and variants can have narrower provider sets.

  • Is K2 text-only?

    Yes. Kimi K2 Instruct accepts and produces text. Multimodal capabilities are not part of this release.

  • What agentic patterns does K2 handle well?

    Structured multi-step sequences: invoke an API, parse the response, branch on results, call a second API, and synthesize a final output. The function-calling interface in AI Gateway maps directly to these workflows.

  • Can I bring my own provider credentials?

    Yes. AI Gateway supports Bring Your Own Key for providers where you hold a direct account. BYOK requests are excluded from ZDR coverage.

Your use is subject to Moonshot AI's Terms & Privacy Policies.