Skip to main content
Transcribe audio data to text using on-device speech recognition models.

Basic Transcription

Transcription with Options

Get detailed output including confidence scores and timestamps:

STTOutput

The detailed output object:

Example: Transcribe Audio File

Example: Record and Transcribe

Model Management

Supported Audio Formats

For best transcription accuracy: - Use 16kHz mono PCM audio - Keep audio clips under 30 seconds for optimal performance - Use a smaller model (whisper-tiny) for faster results, larger models (whisper-base) for better accuracy