00:00 / 00:00

Persona

Lip Sync

One microphone can drive Live2D mouths and the vowel shapes of VRM avatars. It does not depend on face tracking: it works with no camera or phone at all, and blends with face tracking when there is one

The settings live in the Lip Sync section of Tracking, below Webcam Tracking. They are application settings that stay the same across scenes, and the web console's Tracking page edits the same ones. Audio is processed on this computer only: it is never played through the speakers or sent over the plugin API — plugins read the level and vowel readings, not the sound itself

Setting It Up

  1. Turn Enable Lip Sync on in the Lip Sync section. It is off by default. macOS asks for microphone permission the first time
  2. Pick Default Microphone, or a specific input, under Microphone
  3. Speak into the microphone and check that Input Level and the A, I, U, E and O indicators follow you, then adjust the levels if needed
  4. Microphone Lip Sync, in each avatar's own Tracking section, is Always by default, so once the master switch is on the avatars in the scene start lip-syncing

Microphone

Default Microphone follows the system's default input; you can also pick a specific microphone. The list shows inputs from the computer running Persona — in the web console too, which never lists the microphones of the device the browser runs on. Virtual audio devices you have installed appear in the list like any other input. The row shows the name of the active input

A selected microphone that cannot be found is never silently swapped for another: the list shows Unavailable Microphone and the status turns unavailable until you choose another or Retry

Persona does not play the microphone back to you (there is no monitoring), and does not capture the computer's own audio output directly

Levels

ControlRangeDefaultEffect
Input Boost0 – 30 dB0 dBLifts a quiet microphone before the noise gate. Too much holds the mouth wide open
Noise Gate−60 – 0 dB−45 dBIgnores sounds below this level
Smoothing0 – 300 ms80 msSoftens mouth opening and vowel changes. Higher values add more delay

The noise gate uses the level after Input Boost is applied. Everything above it, from the gate up to full scale, becomes the mouth's travel from closed to fully open — so if the mouth hangs open, lower Input Boost or raise Noise Gate, and if it barely moves when you speak, do the reverse. All three apply to the running analysis at once; choosing a microphone or turning lip sync on restarts capture

Status and Meters

The status beside Enable Lip Sync reads:

StatusMeaning
offLip sync is switched off
starting…Opening the microphone and loading the audio analyzer
listeningCapturing and analyzing
unavailableCapture failed; the reason and a Retry button show below

While lip sync is on, Input Level and the five vowel indicators sit below the settings. The indicators are similarity weights, not probabilities: ambiguous speech can raise two at once, and sounds that match no vowel make them fall back while the mouth still opens with the volume

Capture stops and reads unavailable in three cases: microphone access was denied (allow Persona to use the microphone in system settings first), the selected microphone is missing or was unplugged, or the audio analyzer failed to start. Retry restarts capture on the same microphone

Calibrate Voice

While a microphone is capturing, the Voice Profile row reads Calibrated for ‹microphone› when that microphone has its own calibration, and Default Profile otherwise. Calibrate Voice… is available while the status reads listening

The dialog lists A, I, U, E, O and Background Noise, each with a prompt telling you what to say:

  1. At your usual speaking volume and distance from the microphone, press Record on a row and hold that sound steadily for two seconds; for Background Noise, stay quiet and breathe normally
  2. A recorded row gets a check mark and its button becomes Retry, so any single sound can be redone. Vowels count only above the Noise Gate; background noise has no volume requirement. A take that falls short is reported — "That sound was too quiet. Speak above the noise gate and retry." or "Not enough audio was captured. Keep the sound steady and retry." — and the previous successful calibration data is kept
  3. With all six recorded, Preview Profile applies the draft at once. Say A, I, U, E and O in turn: the matching indicator should lead, and the avatars' mouths in the scene follow along. Re-record any sound that gets confused, then preview again
  4. Save Profile keeps it. Cancel, or closing the dialog, discards the draft and restores the saved or bundled profile

If the microphone already has a calibration, the dialog loads its calibration data, so one vowel can be improved without repeating all six. Turning lip sync off, switching microphones, a capture failure or Retry discards unsaved calibration data

Profiles are saved per physical microphone, one each, and load automatically when that microphone starts capturing. Reset Profile appears only when the current microphone has a calibration; it deletes that calibration and returns to the bundled profile, leaving other microphones' profiles untouched. With Default Microphone selected, Persona works out which physical microphone it is where it can; when it cannot, choose a specific microphone before calibrating

Until you calibrate, the bundled starting profile applies — the one the row calls Default Profile. It may not suit your voice or microphone. If vowels are detected incorrectly, calibrate with your own voice. MotionSync profiles cannot be imported

Microphone Lip Sync

You can set microphone lip sync separately for each avatar. Select its layer, and the Tracking section of its detail holds Microphone Lip Sync, which uses the microphone configured in Tracking:

OptionEffect
OffThis avatar ignores the microphone
Always · defaultFollows the microphone at all times, blending with live face tracking by volume
When Face Tracking Is UnavailableLeaves the mouth to face tracking while a face is live, and uses the microphone only when none is — including a face dropout while the last pose is held

With Always and a live face, the blend follows the smoothed volume: in silence the mouth keeps everything tracking sends, silent yawns, jaw movement and lip shapes included; as you speak, the microphone's vowels gradually take over. Without live face tracking, the microphone fully controls the mouth; the last tracked pose does not affect it

New avatars and those in older scenes start at Always, but Enable Lip Sync is off by default, so nothing moves until you turn it on. The mode belongs to that one instance in the scene and is saved with it: two copies of one model can use different modes, while parameter bindings stay with the model

Driving the Model

Live2D

No new sidecar is needed. Microphone values use the standard mouth inputs — MouthOpen, MouthSmile, JawOpen, plus ARKitJawOpen, ARKitMouthClose, ARKitMouthFunnel and ARKitMouthPucker — so the mouth bindings a model already has keep working. Mouth opening blends the vowel shape with the volume continuously, so the shape does not switch abruptly at a consonant

The LipSync parameter group a model declares in its .model3.json — the parameters listed in the Lip Sync row of Model Info — receives the volume directly wherever no binding already drives that parameter, so a rig with non-standard mouth ids moves without a binding. For a rig with its own per-vowel mouth parameters, bind VoiceAVoiceO in Parameters

While a motion's own audio or speech played through the plugin API is playing, it drives the mouth instead of the microphone

VRM

A, I, U, E and O drive the aa, ih, ou, ee and oh expressions, each weighted by the volume times that vowel's share. A model with none of those presets but an ARKit jawOpen expression opens its jaw with the volume instead. How good the mouth looks depends on the model's own expressions

When blending with live face tracking, the vowel presets and the ARKit mouth expressions they conflict with hand over by volume. The microphone's weights are restored at the end of every frame, so they never stick to a held tracking pose

Voice Inputs

The Mouth group of the parameter binding editor also carries twelve voice inputs, usable in both Live2D and VRM bindings:

InputLabelRangeMeaning
VoiceVolumeVoice Volume0 – 1The smoothed volume envelope
VoiceA, VoiceI, VoiceU, VoiceE, VoiceOVoice A … Voice O0 – 1Each vowel's strength, weighted by volume and smoothed
VoiceSilenceVoice Silence0 – 1One minus the volume, so 1 in silence
VoiceFrequencyVoice Form0 – 1The vowels projected onto a conventional mouth-form axis, neutral at 0.5 — not pitch
VoiceVolumePlusMouthOpenVoice + Mouth Open0 – 2Volume plus the tracked mouth opening
VoiceFrequencyPlusMouthSmileVoice + Mouth Form0 – 1The tracked mouth form plus the voice's signed form offset
VoiceMouthOpenVoice Mouth Open0 – 1Vowel-weighted mouth opening times the volume
VoiceMouthSpreadVoice Mouth Spread−1 – 1Signed, vowel-weighted mouth spread times the volume

The two composites, VoiceVolumePlusMouthOpen and VoiceFrequencyPlusMouthSmile, can sum past their range before the binding's own range clamps them. Without a live face each uses its audio half alone, and a pose held after a face dropout is never added. When the microphone is not acting on the avatar — lip sync off, the mode set to Off, or When Face Tracking Is Unavailable leaving the mouth to a live face — they carry only the tracked half, so a binding built on a composite keeps following face tracking

Importing

A .vtube.json may name these inputs in any letter case

Persona analyzes the voice in its own way, so volume levels and vowel detection for the same audio may differ from VTube Studio

Automations

The Microphone Level trigger in Automations listens through the same microphone, Input Boost, Noise Gate and Smoothing. It works with Enable Lip Sync off, without driving the avatars' mouths

An automation's Play Audio action can drive a mouth from a file instead: set its Drive Lip Sync to an avatar, Live2D or VRM, and that avatar's mouth follows the file while it plays, in place of the microphone. The analysis uses the bundled voice profile rather than a microphone's calibration, reads the sound before its volume and fades are applied, and never opens the microphone. When several files drive one avatar, the one that started last wins, and once none is playing, the microphone and face tracking take over again

Last updated on September 19, 2026

Tech otakus destroy the world