Fields
Language
Leavelanguage unset to auto-detect. Set it when you know the language: detection costs time
and picks wrong on short clips, which matters more in a browser where the model is small.
Word timestamps
<audio> element to a spoken word.
Diarization
diarization: true labels speakers inside the transcription. For a standalone “who spoke
when” answer:
Translation
translateToEnglish: true returns English regardless of source language.