[ VOICE API ]

Voice API

A voice that answers as fast as you speak, carries a direction you wrote into the line, and never leaves your machine.

Real-time response

A character feels present when it can answer as quickly as you speak.

How long before the first word

US0.3s
CLOUD TTS A0.9s
CLOUD TTS B1.4s
CLOUD TTS C2.2s
Sample data · same scale

What the same line costs

US
CLOUD TTS A
CLOUD TTS B14×
CLOUD TTS C22×
Sample data · same scale
Placeholder data

Both charts are drawn the same way, aligned the same way, with the same accent, so reading them side by side says one thing: faster, and cheaper. The numbers are placeholders until the measured set lands.

Quality needn't cost more.

Every other engine trades quality for compute along today's frontier. Ours sits outside it.

← BETTER UP AND LEFT LISTENING QUALITY ↑ COMPUTE / COST → TODAY'S FRONTIER Fast small model Legacy on-device Open flow model Cloud LLM-TTS A Cloud LLM-TTS B OUTSIDE IT OURS
OURSOTHER ENGINESTODAY'S FRONTIER
Placeholder data

The dashed line is the frontier every engine alive today has to trade along. Ours sits outside it. The scales are deliberately left off, so what you read is the shape and not a number. The plotted positions are placeholders until the measured set lands.

A reference voice in. A new line out.

A voice you've never heard, saying something it never said.

01
REFERENCE · the line it was given
02
CLONED · a new line, same voice
Live audio

01 is the reference we handed it. 02 is the same voice saying a line it was never given. Press either one and check it for yourself.

The same voice, in another language.

Japanese and English, even mid sentence.

EN
ENGLISH · the cloned voice
JA
JAPANESE · the same voice, another language
Live audio

One voice, English and Japanese, and it can switch inside a sentence. Those two are what we have today. Cross language cloning beyond them has no measured sample yet.

Add the direction to the line and hear the difference.

[natural] The engineering team is deploying the new model this week.

Click to hear · placeholder audio

Pick a tag and hear that take. The small wave in front of each label is that clip's own shape, not a drawn icon. The audio here is a placeholder, produced by signal processing on a real output rather than by the tag path itself.

It breathes.

8 kHz4 kHz0
It stops here
Bandwidth limited engine · simulated
8 kHz4 kHz0
We keep going
Ours · real output

The brightest part of breath and sibilance sits above that line. Cut it and the voice goes muffled.

Nothing leaves the room.

The model runs on your machine. Your text and your voices never travel.

your device voicetext remote server REQUEST LEAVESAUDIO RETURNS two crossings · every line
Every line crosses the edge of your machine. Twice.
PRIVATE LOCAL LOOP your devicenothing crosses the edge
Request out, audio back, all inside the box. Nothing to send, nothing to wait for.
Diagram

Both sides are the same round trip: the request goes out, the audio comes back. The only difference is that the top one crosses the edge of your device to reach someone else's machine and back, with your text, your reference audio and your character settings on the road the whole way. The bottom one finishes inside the box. This describes the on-device and self-hosted build. The drawing is a diagram, not a measured topology.