A voice that answers as fast as you speak, carries a direction you wrote into the line, and never leaves your machine.
A character feels present when it can answer as quickly as you speak.
Both charts are drawn the same way, aligned the same way, with the same accent, so reading them side by side says one thing: faster, and cheaper. The numbers are placeholders until the measured set lands.
Every other engine trades quality for compute along today's frontier. Ours sits outside it.
The dashed line is the frontier every engine alive today has to trade along. Ours sits outside it. The scales are deliberately left off, so what you read is the shape and not a number. The plotted positions are placeholders until the measured set lands.
A voice you've never heard, saying something it never said.
01 is the reference we handed it. 02 is the same voice saying a line it was never given. Press either one and check it for yourself.
Japanese and English, even mid sentence.
One voice, English and Japanese, and it can switch inside a sentence. Those two are what we have today. Cross language cloning beyond them has no measured sample yet.
Pick a tag and hear that take. The small wave in front of each label is that clip's own shape, not a drawn icon. The audio here is a placeholder, produced by signal processing on a real output rather than by the tag path itself.
The brightest part of breath and sibilance sits above that line. Cut it and the voice goes muffled.
The model runs on your machine. Your text and your voices never travel.
Both sides are the same round trip: the request goes out, the audio comes back. The only difference is that the top one crosses the edge of your device to reach someone else's machine and back, with your text, your reference audio and your character settings on the road the whole way. The bottom one finishes inside the box. This describes the on-device and self-hosted build. The drawing is a diagram, not a measured topology.