Skip to main content

Fields

Language

Leave language unset to auto-detect. Set it when you know the language: detection costs time and picks wrong on short clips.

Word timestamps

Use these for captions, or to seek a player to a spoken word. On a desktop app transcribing long recordings, this is what makes the transcript navigable.

Diarization

diarization: true labels speakers inside the transcription. For a standalone “who spoke when” answer, use the diarization namespace.

Translation

translateToEnglish: true returns English regardless of source language.

Long recordings

A desktop is the one platform where transcribing an hour of audio is reasonable. Word timestamps and diarization both cost extra time per minute of audio, so turn on only what the interface will actually show.