Skip to content
Dashboard

Qwen3-30B-A3B

Qwen3-30B-A3B is a mixture-of-experts model from Alibaba Cloud that activates only 3 billion of its 30 billion parameters per inference, outperforming QwQ-32B while running at a fraction of the compute cost.

View API reference
Input and output price
Input $0.12, Output $0.50, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'alibaba/qwen-3-30b',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingPlayground

Try out Qwen3-30B-A3B by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

alibaba logo
alibaba logo

Qwen3-30B-A3B

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Free Tier
Release Date
41K16K0.2 s
$0.12/M
$0.50/M
04/28/2025

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Getting started

Call Qwen3-30B-A3B through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.

Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.

index.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'alibaba/qwen-3-30b',
prompt: 'Why is the sky blue?',
});
console.log(result.text);
}
main().catch(console.error);

Top-level parameters

The same Qwen3-30B-A3B request in each API format AI Gateway supports.

top-level-params.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'alibaba/qwen-3-30b',
system: 'You are a concise technical assistant.',
prompt: 'Summarize the tradeoffs between static generation and SSR.',
maxOutputTokens: 1024,
});
console.log(result.text);
}
main().catch(console.error);

Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.

ParameterTypeRequiredDescription
modelstringYesModel ID in the form creator/model, e.g. alibaba/qwen-3-30b. AI Gateway routes the request to an available provider.
maxOutputTokensnumberNoHard cap on generated tokens. Qwen3-30B-A3B supports up to 16,384 output tokens. Reasoning tokens count toward this limit.
reasoning'provider-default' | 'none' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh'NoProvider-agnostic reasoning effort, available in AI SDK 7 or later. Maps to the provider’s native reasoning configuration; reasoning settings under providerOptions take precedence when both are set. See the Reasoning section below.
providerOptionsRecord<string, JSONValue>NoAI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below.

Input limits

InputFormatsSourcesMax countMax sizeLimits
TextPrompt and response share the 41K-token context window

Provider options

Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.

Learn more in the AI SDK alibaba provider docs.

provider-options.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'alibaba/qwen-3-30b',
prompt: 'Why is the sky blue?',
providerOptions: {
gateway: {
only: ['deepinfra'],
},
},
});
console.log(result.text);
}
main().catch(console.error);

These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.

ParameterTypeRequiredDescription
providerOptions.gateway.onlystring[]NoRestrict routing to these provider slugs. Requests fail over only within the listed providers.
providerOptions.gateway.orderstring[]NoPreferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks.
providerOptions.gateway.sort'cost' | 'ttft' | 'tps'NoRank candidate providers by price, time to first token, or tokens per second instead of the default routing order.
providerOptions.gateway.zeroDataRetentionbooleanNoRoute only to providers with a zero-data-retention policy for this model.

Routing across providers

AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.

Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.

Reasoning

AI Gateway bridges reasoning across every API format. The AI SDK exposes a provider-agnostic top-level reasoning level (none, minimal, low, medium, high, or xhigh); the Chat Completions and Responses formats take the same effort under reasoning.effort; and the Anthropic Messages format uses a native thinking token budget. Whichever you send, the gateway maps it to the target model’s native configuration, converting between effort levels and token budgets as needed. Reasoning-related settings under providerOptions take full precedence over the top-level reasoning value and are never merged. Reasoning tokens typically count toward your output-token usage, though how they’re reported and billed varies by provider.

Learn more in the AI Gateway reasoning guide.

reasoning.ts
import { generateText } from 'ai';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'alibaba/qwen-3-30b',
prompt: 'Explain the Monty Hall problem step by step.',
reasoning: 'high',
});
console.log(result.text);
}
main().catch(console.error);

Tool calling

Expose tools the model can call. Define each tool’s inputs with a Zod schema.

tool-calling.ts
import { generateText, tool } from 'ai';
import { z } from 'zod';
import 'dotenv/config';
async function main() {
const result = await generateText({
model: 'alibaba/qwen-3-30b',
prompt: 'What is the weather in San Francisco?',
tools: {
getWeather: tool({
description: 'Get the current weather for a location',
inputSchema: z.object({ location: z.string() }),
execute: async ({ location }) => ({ location, temperatureC: 18 }),
}),
},
});
console.log(result.text);
}
main().catch(console.error);

Copy link to headingMore models by Alibaba Cloud

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
991K4.0 s50 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+2
alibaba logo
09/01/2026
991K1.1 s96 tps
$0.16/M
$0.47/M
Read$0.02/M
Write$0.20/M
+2
alibaba logo
08/26/2026
1M0.6 s133 tps
$0.10/M
$0.40/M
Read$0.01/M
Write$0.63/M
+1
alibaba logo
deepinfra logo
morph logo
+3
08/14/2026
1M2.5 s160 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+1
alibaba logo
fireworks logo
08/02/2026
991K2.5 s140 tps
$0.03/M+2 more
$0.13/M+2 more
Read$0.006/M
Write$0.04/M
+2
alibaba logo
07/28/2026
1M2.0 s58 tps
$0.40/M+1 more
$1.60/M+1 more
Read$0.08/M
Write$0.50/M
+2
alibaba logo
06/02/2026

Copy link to headingAbout Qwen3-30B-A3B

Qwen3-30B-A3B occupies a distinctive position in the Qwen3 lineup: it's the smaller of two MoE models in the family, but its efficiency story is the more striking one. Inference activates only 3 billion parameters, comparable to serving a small model, yet the full 30 billion parameter capacity gives it a much larger representational space than a genuinely 3B model would have.

Alibaba Cloud's benchmarks position this model above QwQ-32B, which was previously one of the stronger open reasoning models. QwQ-32B is a dense model that activates all 32 billion of its parameters on every token, meaning Qwen3-30B-A3B achieves superior results at roughly one-tenth the active parameter count. For teams running inference at volume, this ratio has direct cost implications.

Like the rest of the Qwen3 family, the model supports hybrid thinking modes. The enable_thinking parameter switches between step-by-step chain-of-thought reasoning and direct-response mode. The thinking budget can be configured per request, so applications can use extended reasoning for genuinely complex queries while defaulting to fast responses for routine ones.

The 30B-A3B supports 119 languages and dialects, the same multilingual coverage as the rest of the Qwen3 family, and includes tool calling, agentic workflow, and MCP support, giving the model strong instruction-following and coding capabilities relative to its inference cost.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: Provider selection is most consequential when your application has region-specific latency requirements or data handling policies that point to particular infrastructure.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Qwen3-30B-A3B

Best for

  • High-volume inference where token cost matters: Activating only 3B parameters per request makes this model economical to run at scale. Applications processing thousands of requests per hour benefit from the efficiency gap over fully-dense alternatives at similar quality levels
  • Reasoning tasks that previously required larger models: If your workload was reaching for Qwen2.5-32B or QwQ-32B, the 30B-A3B delivers comparable or better results with significantly lower serving costs
  • Applications with variable complexity: The hybrid thinking mode is particularly useful here; route complex queries through thinking mode and simpler ones through non-thinking mode, keeping costs proportional to actual task difficulty
  • Production deployments requiring predictable throughput: MoE models with small active parameter counts tend to be faster to serve than dense models of comparable benchmark performance, which helps when maintaining consistent response latency under load

Consider alternatives when

  • You need maximum reasoning headroom: For the most demanding tasks, the Qwen3-235B-A22B offers a higher capability ceiling. The 30B-A3B is efficient; the 235B MoE variant offers more reasoning headroom when needed
  • You're comparing against genuinely tiny models for simple tasks: If your application primarily handles simple classification, short-form generation, or keyword extraction, even smaller models may provide adequate quality at lower cost
  • Multimodal input processing is required: Qwen3-30B-A3B handles text only

Qwen3-30B-A3B delivers strong reasoning performance without the serving costs of large dense models. It outperforms QwQ-32B while activating one-tenth the parameters per token, fitting well into high-throughput applications where quality and efficiency need to coexist. AI Gateway adds reliable failover across DeepInfra and a single billing integration.

Copy link to headingFrequently Asked Questions

  • How is it possible for a 3B-active-parameter model to outperform QwQ-32B?

    The mixture-of-experts architecture separates total parameter count from inference compute. At inference, routing selects the most relevant 3 billion parameters for each token. The model benefits from the broad capacity of its 30 billion total parameters while keeping serving costs proportional to the 3B active count. QwQ-32B activates all 32 billion parameters but has less total representational capacity.

  • What does "A3B" mean in the model name?

    "A3B" indicates that 3 billion parameters are activated during inference (A = activated, 3B = 3 billion). The "30B" is the total parameter count across all expert layers.

  • How does the 30B-A3B architecture affect serving cost?

    At inference, only 3 billion parameters activate per token, so per-token compute is comparable to a 3B dense model even though the full MoE has 30 billion parameters. This is the source of the cost advantage over dense 32B-class models at similar quality.

  • How does the thinking budget control work in practice?

    You set a token budget for the thinking trace via the API. Higher budgets allow the model to explore more reasoning steps before producing its answer. Lower budgets constrain the reasoning phase, producing faster responses, useful when a question is straightforward and extended reasoning wouldn't add value.

  • Does Qwen3-30B-A3B support the same 119 languages as other Qwen3 models?

    Yes. The 119-language coverage applies across the Qwen3 family, including this model.

  • What agentic use cases is this model suited for?

    The model supports tool calling and MCP (Model Context Protocol). It fits automated workflows where the model needs to select and invoke tools across multiple steps, particularly in cost-sensitive deployments where running a larger model per agent step would be prohibitive.

  • How does AI Gateway route requests for this model?

    AI Gateway selects among DeepInfra based on availability and performance. If a provider returns an error or is slow to respond, requests automatically retry with another provider in the pool, so your application doesn't need to implement retry logic independently.

Your use is subject to Alibaba Cloud's Terms & Privacy Policies.