Bring a VRM or FBX. We rig her. Give her text, yours or your LLM's, and she answers with her voice, her face and her whole body.
A file goes in. What comes out looks up and says the first line, in one take.
Drop a VRM or FBX. Rigging, retargeting and the viseme set are handled for you. Nothing to wire up. Auto rigging can still fail on unusual proportions, so there is a manual path in the docs.
A file goes in, we rig it, you hand her text. Open step three and cut in whenever you like.
A VRM or an FBX, straight out of your DCC. No naming convention to learn and no template to match. That is the whole ask.
Skeleton, expressions and visemes get wired automatically. Humanoid bones retargeted, face blendshapes matched to our expression set, five visemes timed to the voice.
Yours or your LLM's. She answers with her voice, her face and her whole body. Cut in whenever you like. Her sentence can wait.
Three steps named after whoever is doing the work, not after a stage in a pipeline. Step three is live: press Cut in and she stops on the half word she is on, then comes back from the same place. This panel is a simulation of the behaviour, not a live session.
A laugh is not a mouth shape. It starts in the diaphragm, travels through the shoulders and the neck, and reaches the face last.
Same laugh, both sides. We generate the whole body, so when she laughs, she laughs. Lip sync alone can only move the mouth. The drawing stands in for the split screen take until the footage lands.
When people talk, the gesture starts before the voice does. Motion generated after the audio can only chase it.
Voice and motion come out of one pass, so the hand can lead. That is an architecture difference, not a setting. The offset shown here is a placeholder until the measured pair lands.
Interrupt her any time. She stops on half a word, hears you out, then picks the sentence back up.
I was thinking we could take the long way back, past the station, and see if the shop is still open.
No queue to drain and no sentence to finish first. She stops where she is and comes back from the same place. The live version is step three above.
Streaming in and out with no seam between sentences. Nothing restarts, nothing buffers.
One unbroken track against a stitched one. At every seam the other track resets the mouth and jumps the motion. Ours has no seams to reset.
Model, generation and rendering all sit on your machine. Your character, your text, your footage. None of it takes a step outside.
This holds for the local and self hosted build. Nothing crosses the edge of the box. The drawing is a diagram, not a measured topology.