Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 307 models · Page 2 of 6.
Search and filter all models →- GoogleGemini Omni Flash PreviewGemini Omni Flash (Preview) is a multimodal model designed for video, image, and text tasks. It is optimized for video generation, offering video output alongside text responses in a single model.
- GoogleGemma 4 26B A4B ITGemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
- GoogleGemma 4 31B ITGemma 4 31B is engineered to tackle the most demanding enterprise workloads and complex reasoning tasks. With an expansive 256K-token context window, the 31B model can effortlessly ingest entire codebases, and massive sets of images in a single prompt.
- Z.AIGLM 4.5GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters.
- Z.AIGLM 4.5 AirGLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters.
- Z.AIGLM 4.5VBuilt on the GLM-4.5-Air base model, GLM-4.5V inherits proven techniques from GLM-4.1V-Thinking while achieving effective scaling through a powerful 106B-parameter MoE architecture.
- Z.AIGLM 4.6As the latest iteration in the GLM series, GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications.
- Z.AIGLM 4.7GLM-4.7 is Z.ai’s latest flagship model, with major upgrades focused on two key areas: stronger coding capabilities and more stable multi-step reasoning and execution.
- Z.AIGLM 4.7 FlashGLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option. Beyond coding, it is also recommended for creative writing, translation, long-context tasks, and roleplay.
- Z.AIGLM 4.7 FlashX GLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option.
- Z.AIGLM 5GLM-5 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5.
- Z.AIGLM 5 TurboGLM 5 Turbo is a foundation model deeply optimized for the OpenClaw scenario. It has been specifically optimized for the core requirements of OpenClaw tasks since the training phase, enhancing key capabilities such as tool invocation, command following, timed and persistent tasks, and long-chain execution.
- Z.AIGLM 5.1GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours—autonomously planning, executing, and improving itself throughout the process—ultimately delivering complete, engineering-grade results.
- Z.AIGLM 5.2GLM-5.2 delivers powerful coding capabilities, usable 1M-context support, and continued strengths in long-horizon tasks.
- Z.AIGLM 5.2 FastFast version of GLM 5.2 with 120-250 TPS.
- Z.AIGLM 5V TurboGLM-5V-Turbo is Z.AI’s first multimodal coding foundation model, built for vision-based coding tasks. It can natively process multimodal inputs such as images, video, and text, while also excelling at long-horizon planning, complex coding, and action execution. Deeply optimized for agent workflows, it works seamlessly with agents such as Claude Code and OpenClaw to complete the full loop of “understand the environment → plan actions → execute tasks”.
- Z.AIGLM-4.6VGLM-4.6V series are Z.ai’s iterations in a multimodal large language model. GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales.
- Z.AIGLM-4.6V-FlashFor local deployment and low-latency applications. GLM-4.6V series are Z.ai’s iterations in a multimodal large language model. GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales.
- OpenAIGPT 4o Mini Search PreviewGPT-4o mini Search Preview is a specialized model trained to understand and execute web search queries with the Chat Completions API. In addition to token fees, web search queries have a fee per tool call.
- OpenAIGPT 5 ChatGPT-5 Chat points to the GPT-5 snapshot currently used in ChatGPT.
- OpenAIGPT 5.1 Codex MaxGPT‑5.1-Codex-Max is purpose-built for agentic coding.
- OpenAIGPT 5.1 Codex MiniGPT-5.1 Codex mini is a smaller, faster, and cheaper version of GPT-5.1 Codex.
- OpenAIGPT 5.1 ThinkingAn upgraded version of GPT-5 that adapts thinking time more precisely to the question to spend more time on complex questions and respond more quickly to simpler tasks.
- OpenAIGPT 5.2GPT-5.2 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- OpenAIGPT 5.2 Version of GPT-5.2 that produces smarter and more precise responses.
- OpenAIGPT 5.2 ChatThe model powering ChatGPT is gpt-5.2-chat-latest: this is OpenAI's best general-purpose model, part of the GPT-5 flagship model family.
- OpenAIGPT 5.2 CodexGPT‑5.2-Codex is a version of GPT‑5.2 further optimized for agentic coding in Codex, including improvements on long-horizon work through context compaction, stronger performance on large code changes like refactors and migrations, improved performance in Windows environments, and significantly stronger cybersecurity capabilities.
- OpenAIGPT 5.3 CodexGPT-5.3-Codex advances both the frontier coding performance of GPT‑5.2-Codex and the reasoning and professional knowledge capabilities of GPT‑5.2, together in one model, which is also 25% faster. This enables it to take on long-running tasks that involve research, tool use, and complex execution.
- OpenAIGPT 5.4GPT-5.4 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- OpenAIGPT 5.4 MiniGPT-5.4 Mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
- OpenAIGPT 5.4 NanoGPT-5.4 Nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.
- OpenAIGPT 5.4 ProGPT-5.4 Pro uses more compute to think harder and provide consistently better answers. It's designed to tackle tough problems.
- OpenAIGPT 5.5GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.
- OpenAIGPT 5.5 Pro
- OpenAIGPT 5.6 LunaGPT-5.6 Luna is a fast, affordable GPT-5.6 model that brings strong capability at the lowest cost in the series.
- OpenAIGPT 5.6 SolGPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 series, its most capable model for long-horizon agentic work across coding, biology, and cybersecurity.
- OpenAIGPT 5.6 TerraGPT-5.6 Terra is a balanced GPT-5.6 model for everyday work, with performance comparable to the previous generation at half the cost.
- OpenAIGPT Image 1GPT Image 1 is OpenAI's new state-of-the-art image generation model. It is a natively multimodal language model that accepts both text and image inputs, and produces image outputs.
- OpenAIGPT Image 1 MiniA cost-efficient version of GPT Image 1. It is a natively multimodal language model that accepts both text and image inputs, and produces image outputs.
- OpenAIGPT Image 1.5GPT Image 1.5 is OpenAI's latest image generation model, with better instruction following and adherence to prompts.
- OpenAIGPT Image 2GPT Image 2 is OpenAI's state-of-the-art image generation model for fast, high-quality image generation and editing. It supports flexible image sizes and high-fidelity image inputs.
- OpenAIGPT OSS 120BExtremely capable general-purpose LLM with strong, controllable reasoning capabilities
- OpenAIGPT OSS 20BA compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments.
- OpenAIGPT OSS Safeguard 20BOpenAI's first open weight reasoning model specifically trained for safety classification tasks. Fine-tuned from GPT-OSS, this model helps classify text content based on customizable policies, enabling bring-your-own-policy Trust & Safety AI where your own taxonomy, definitions, and thresholds guide classification decisions.
- OpenAIGPT-3.5 TurboOpenAI's most capable and cost effective model in the GPT-3.5 family optimized for chat purposes, but also works well for traditional completions tasks.
- OpenAIGPT-4 Turbogpt-4-turbo from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It has a knowledge cutoff of April 2023 and a 128,000 token context window.
- OpenAIGPT-4.1GPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
- OpenAIGPT-4.1 miniGPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases.
- OpenAIGPT-4.1 nanoGPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model.
- OpenAIGPT-4oGPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API.
- OpenAIGPT-4o miniGPT-4o mini from OpenAI is their most advanced and cost-efficient small model. It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast.
- OpenAIGPT-4o mini TranscribeGPT-4o mini Transcribe is a speech-to-text model that uses GPT-4o mini to transcribe audio. It offers improvements to word error rate and better language recognition and accuracy compared to original Whisper models. Use it for more accurate transcripts.
- OpenAIGPT-4o TranscribeGPT-4o Transcribe is a speech-to-text model that uses GPT-4o to transcribe audio. It offers improvements to word error rate and better language recognition and accuracy compared to original Whisper models. Use it for more accurate transcripts.
- OpenAIGPT-5GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.
- OpenAIGPT-5 miniGPT-5 mini is a cost optimized model that excels at reasoning/chat tasks. It offers an optimal balance between speed, cost, and capability.
- OpenAIGPT-5 nanoGPT-5 nano is a high throughput model that excels at simple instruction or classification tasks.
- OpenAIGPT-5 proGPT-5 pro uses more compute to think harder and provide consistently better answers. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish.
- OpenAIGPT-5-CodexGPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments.
- OpenAIGPT-5.1 InstantGPT-5.1 Instant (or GPT-5.1 chat) is a warmer and more conversational version of GPT-5-chat, with improved instruction following and adaptive reasoning for deciding when to think before responding.
- OpenAIGPT-5.1-CodexGPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments.