Skip to main content
synthesizeStream emits AudioChunk values as they are produced, so playback starts on the first chunk instead of the last.

Signature

When it is worth it

For a sentence, speak is simpler and the latency difference is invisible. For a paragraph, or for a model reply being streamed as it generates, streaming synthesis is what keeps the gap between “the model finished” and “the user hears something” short.

Pairing it with a streaming completion

Speak each sentence as the model finishes it:
For a full conversation rather than one reply, use the voice session, which handles turn-taking and interruption for you.

Feeding AudioTrack

Release the track when you are done. It holds an audio output.