Local Siri Voices in macOS 27: What’s Inside
SirimacOS 27локальный TTS
What was found under the hood
The new Siri voices appear to run on a local generative stack rather than a conventional speech engine with a limited set of intonations. The combination of an LLM and a streaming audio decoder is especially notable: it can account for text meaning, shape delivery, and begin playback before the entire response has been generated.
Apple’s public Siri AI (Beta) documentation confirms that an advanced on-device model is used, but does not disclose the internal architecture. At the time of testing macOS 27, the stated requirements were a Mac with an M3 chip or newer and at least 12 GB of unified memory. Users can adjust speech rate and expressivity.
An independent teardown adds unofficial details. The system contains a file called instruct_3b.voice, linked to the voice model. During the research, it supported English, German, and Japanese, understood text content, and accepted some emotion tags, although Apple does not document that mechanism.
The expressivity parameter seems to affect more than a simple emotional setting. Observations suggest that the model connects expressiveness with sentence semantics, so the result cannot be reduced to changing pitch or speaking speed. Its quality is not yet on par with newer ElevenLabs or Gemini voice models, but it is clearly beyond the familiar monotone delivery.
The main limitation appeared during long generations. The maximum tested output was 163 seconds, yet after roughly one minute the model began to degrade and could loop on a single word. Splitting text into chunks noticeably reduced voice degradation, while working with float representations harmed the result more severely.
Why developers should care
This is already a useful local TTS building block, but not yet a production-ready component. A streaming decoder suits low-latency interfaces, while semantic and emotional processing provides more control than traditional speech synthesis.
The most intriguing part of the teardown is the possibility of supplying custom voice embeddings. Apple provides no public API or documentation for this, so format stability cannot be assumed: the internal file, parameters, or behavior may change with any system build.
I would first test chunk boundaries, decoder error accumulation, and intonation consistency between segments. A 163-second limit looks impressive only on paper if real stability ends after one minute. A strong local voice engine is clearly taking shape, but the central engineering question is not voice beauty yet—it is predictable long-form generation.