gpt-realtime-whisper
GPT Realtime Whisper is a streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. It is designed for realtime use cases where developers need to tune latency and accuracy. GPT Realtime Whisper is priced by audio duration rather than text tokens.
Websockets
index.ts
import { experimental_streamTranscribe as streamTranscribe } from 'ai';import { createGateway, gateway } from '@ai-sdk/gateway';import { readFile } from 'node:fs/promises';
const modelId = 'openai/gpt-realtime-whisper';
// Mint this on your server, then send only the short-lived token to the client.const { token } = await gateway.experimental_transcription.getToken({ model: modelId,});const clientGateway = createGateway({ apiKey: token });
// Raw 24 kHz, 16-bit signed little-endian mono PCM audio.const bytes = await readFile('audio.pcm');const audio = new ReadableStream<Uint8Array>({ start(controller) { controller.enqueue(new Uint8Array(bytes)); controller.close(); },});
const result = streamTranscribe({ model: clientGateway.transcription(modelId), audio, inputAudioFormat: { type: 'audio/pcm', rate: 24000 },});
for await (const part of result.fullStream) { if (part.type === 'transcript-delta') { process.stdout.write(part.delta); }}
console.log('\nFinal:', await result.text);Providers
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|