Skip to main content
synthesizeStream emits AudioChunk values as they are produced, so playback starts on the first chunk instead of the last.

Signature

When it is worth it

For a sentence, speak is simpler and the latency difference is invisible. For a paragraph, or for a model reply being streamed as it generates, streaming synthesis is what keeps the gap between “the model finished” and “the user hears something” short.

Pairing it with a streaming completion

Speak each sentence as the model finishes it:
For a full conversation rather than one reply, use the voice session, which handles turn-taking and interruption.

Forwarding chunks to the renderer

When the audio graph lives in the window, send chunks as they arrive rather than buffering the whole utterance in the main process:
Audio chunks are much larger than text deltas, so batching matters here in a way it does not for tokens. Send each chunk once and let the renderer schedule them against a running clock.

Stopping