[ CHARACTER API ]

She was a file this morning. She hasn't stopped talking since.

Bring a VRM or FBX. We rig her. Give her text, yours or your LLM's, and she answers with her voice, her face and her whole body.

From a file to someone who answers

A file goes in. What comes out looks up and says the first line, in one take.

A file
Rigged for you
Someone who answers
Diagram

Drop a VRM or FBX. Rigging, retargeting and the viseme set are handled for you. Nothing to wire up. Auto rigging can still fail on unusual proportions, so there is a manual path in the docs.

Three steps. None of them is a pipeline.

A file goes in, we rig it, you hand her text. Open step three and cut in whenever you like.

  • A VRM or an FBX, straight out of your DCC. No naming convention to learn and no template to match. That is the whole ask.

  • Skeleton, expressions and visemes get wired automatically. Humanoid bones retargeted, face blendshapes matched to our expression set, five visemes timed to the voice.

  • Yours or your LLM's. She answers with her voice, her face and her whole body. Cut in whenever you like. Her sentence can wait.

She stops on the half word she is on.
Try it

Three steps named after whoever is doing the work, not after a stage in a pipeline. Step three is live: press Cut in and she stops on the half word she is on, then comes back from the same place. This panel is a simulation of the behaviour, not a live session.

Laughter starts somewhere below the face.

A laugh is not a mouth shape. It starts in the diaphragm, travels through the shoulders and the neck, and reaches the face last.

Whole body
Mouth only
Diagram

Same laugh, both sides. We generate the whole body, so when she laughs, she laughs. Lip sync alone can only move the mouth. The drawing stands in for the split screen take until the footage lands.

Her hands know first.

When people talk, the gesture starts before the voice does. Motion generated after the audio can only chase it.

216 ms earlier Gesture starts0.000s First syllable0.216s
Gesture starts First syllable
Placeholder timing

Voice and motion come out of one pass, so the hand can lead. That is an architecture difference, not a setting. The offset shown here is a placeholder until the measured pair lands.

Her sentence can wait.

Interrupt her any time. She stops on half a word, hears you out, then picks the sentence back up.

I was thinking we could take the long way back, past the station, and see if the shop is still open.

Stops here Resumes from here
One frame

No queue to drain and no sentence to finish first. She stops where she is and comes back from the same place. The live version is step three above.

She is never not there.

Streaming in and out with no seam between sentences. Nothing restarts, nothing buffers.

Aniwaffle · continuous
Clip based · restarts
Diagram

One unbroken track against a stitched one. At every seam the other track resets the mouth and jumps the motion. Ours has no seams to reset.

Nobody else has met her.

Model, generation and rendering all sit on your machine. Your character, your text, your footage. None of it takes a step outside.

YOUR MACHINE MODEL GENERATION RENDER NO DATA LEAVES
Diagram

This holds for the local and self hosted build. Nothing crosses the edge of the box. The drawing is a diagram, not a measured topology.