Skip to main content
RunAnywhere.voice assembles speech-to-text, a language model, and text-to-speech into one session that listens, thinks, and answers out loud. It owns the microphone while it runs.

Signature

Named arguments, three of them required.
Leaving vad unset ensures the catalogued default Silero VAD.

VoiceSession

say speaks a line without waiting for the user, which is how you greet someone. interrupt cuts off whatever is being spoken. Always close() the session. It holds the microphone.

Turn handling

Raise interruption.minDurationMs in a noisy room, or a cough cuts the model off.

Steering the replies

Spoken replies want to be short. Cap maxOutputTokens lower than for a chat interface.

A complete widget

Closing on lifecycle change and in dispose is what stops the microphone staying live behind a backgrounded app.

Permission

Request the microphone permission before start().

Three models at once

A voice session holds a speech model, a language model, and a synthesis model in memory together. On a phone that is a real constraint. Prefer the smallest usable model for each.

Diarization

To answer “who spoke when” over a recording rather than run a live conversation: