- Learn
- Sound & Music
- Lip Sync
Sound & Music · Film language
Lip Sync
Also called: lip synchronization, lip-synching, mouth sync
Lip sync is the match between the mouth movements we see and the voice we hear. When it holds, nobody notices; when it slips by even a few frames, the performance feels fake. Filmmakers protect it on set, rebuild it in ADR and dubbing, and fake it deliberately in musicals and music videos.
- What it does
- Makes a voice feel like it comes out of the face on screen.
- Use it when
- Any on-screen speech or singing: production dialogue, ADR, dubbing, playback, animation and AI talking characters.
- Watch out
- Audio that arrives early is noticed faster than audio that lags; check sync on p, b and m, where the lips close.
- Try this prompt
close-up of a singer mouthing lyrics into a vintage microphone, lips mid-word, face lit by a warm stage spot, 85mm, eye level
What is lip sync?
Lip sync (lip synchronization, also written lip-synching) means two related things. Technically, sound and picture line up, so a word is heard on the frame where the lips form it. As performance, someone mimes to an existing recording, like a singer mouthing along to playback in a musical or a music video.
Both rely on the same cue. Viewers read speech from a few mouth shapes, and the easiest to check are the bilabials, p, b and m, where the lips press shut, plus wide open vowels. A close-up makes every mismatch larger, which is why dialogue close-ups get the most sync attention in the edit.
Where does lip sync matter in filmmaking?
Production dialogue. Feature films record sound separately from the camera, so the clapperboard gives the editor a sync point: the frame where the sticks meet lines up with the spike on the audio track.
ADR. When a line is replaced in the studio, the actor has to match their own lip movements from the shoot. In ADR a great reading that drifts off the mouth is unusable.
Dubbing. In dubbing a translated line must fit another actor's mouth. Adapters rewrite it so closed-lip sounds and long vowels land where the original has them.
Musicals and playback. Classic Hollywood musicals prerecorded most songs and had performers mime to playback on set; some stars were sung for by other voices, as Marni Nixon did for Audrey Hepburn in My Fair Lady (1964). Singin' in the Rain (1952) builds its whole plot around this, from a disastrous preview where the soundtrack slips out of sync to a hidden singer supplying a star's voice.
Animation. Studios record the dialogue first, and animators time mouth shapes to the track frame by frame. Here the picture follows the voice acting, not the other way round.
How to get and keep lip sync
- Slate every take, even with timecode. A visible clap saves hours when a recorder drifts or metadata goes missing.
- Keep frame rates consistent. Mismatched rates between camera and sound (for example 23.976 vs 24) cause slow drift that looks fine at the start of a take and is off by the end.
- Know the tolerance. Broadcast guidance puts the point where viewers notice at roughly 45 ms of audio leading, about one frame at 24 fps, and around 125 ms of audio lagging, about three frames. Early sound feels wrong faster because in life light always reaches us before sound.
- Check on the bilabials. When syncing by eye, scrub to a p, b or m and slide the audio until the lips close on it. Then check it at full speed.
- For playback, give performers a loud, clean track and shoot to the song's timecode so every take cuts together. For dialogue editing, over-the-shoulder angles and cutaways let the editor hide a line that won't sync.
How does AI lip sync work for talking characters?
AI lip sync usually happens in two stages. First you need a performance shot: a face whose mouth is clearly visible and moving. Then a voice track is matched to it, either by a dedicated lip-sync step that reshapes the mouth or by sliding the audio in the edit, as in ADR.
Direction matters in the first stage. Frontal or three-quarter faces sync far better than profiles; a hand, microphone or beard over the lips makes the mouth hard to read. In FlashBoards, keep the script, the recorded voice takes and the generated performance shots on one board, so each line sits next to the face that has to say it.
Lip sync vs. asynchronous sound
Lip sync is the default contract: we see a mouth, we hear its words at the same instant. Asynchronous sound breaks that contract on purpose: a voice from another moment or a line over a different image, for irony or subjectivity. The difference is intent. A voice that drifts two frames late is not asynchronous sound; it is a sync error, and it reads as sloppy rather than expressive. Both live in the wider world of sound in film.
How to create a lip sync with AI
3 promptsImage and video models can't hear your dialogue, so the job is to generate a performance shot that is easy to sync: a readable mouth, clear shapes, steady framing. Build a start frame and a second take, animate the first, and keep all three next to the voice track on one FlashBoards board.
Music video close-up, eye-level camera, 85mm lens. A singer in her late twenties with a teal-dyed buzz cut, a thin silver septum ring and freckles across her nose sings into a vintage chrome microphone held below her chin, mouth wide open on a long vowel, eyes half closed. Warm amber spotlight from above-left, deep blue haze behind her, soft bokeh from small stage lights. Her whole mouth and jaw clearly visible, face turned three-quarters toward camera.
"Microphone held below her chin" and "whole mouth and jaw clearly visible" keep the lips unobstructed; "mouth wide open on a long vowel" gives a clean, readable shape.
Open FlashBoardsCinematic dialogue close-up in 2.39:1, eye-level, 65mm lens. A fisherman in his seventies with a white horseshoe mustache trimmed above the lip line, a sun-cracked nose and a faded green wool cap speaks to someone off camera right, lips pressed together as if saying an M. Overcast daylight from a harbor window camera left, soft shadow on the far cheek, blurred ropes and a lantern in the dark cabin behind him.
"Lips pressed together as if saying an M" asks for a bilabial, the shape editors sync on; the trimmed mustache keeps hair off the mouth.
Open FlashBoardsCamera: slow, steady push-in from close-up toward her face, keeping the mouth centered in the lower third of the frame. Action: The singer is already singing as the shot begins, lips shaping each word clearly, jaw opening wide on the long notes. She closes her lips firmly between phrases, takes a breath, tilts her head slightly back and opens her mouth again on a held note. The spotlight stays warm and steady, blue haze drifting slowly behind her.
"Already singing as the shot begins" keeps the clip alive from frame one; firm closes and wide openings give clear anchor points for syncing.
Open FlashBoardsCommon mistakes
- Mumbled mouths. Video models default to small, vague lip motion; name the shapes: "lips pressed together," "jaw opening wide."
- Blocked lips. Say where the microphone or hands are, "below her chin," and keep facial hair trimmed.
- Profiles hide half the mouth; ask for frontal or three-quarter faces.
- Motion isn't sync. A generated mouth says no particular words; match the voice afterward or use it as a listening shot.
Keep the reference frames, prompts and every generated take side by side — images and video in one canvas.
FAQ
4 questionsWhat does lip sync mean in film?
In film, lip sync means that dialogue or singing on the soundtrack lines up with the mouth movements on screen, so the voice seems to come from the character. It covers syncing production sound, replacing lines in ADR, fitting translated dialogue in dubbing and miming to prerecorded music.
How many frames out of sync is noticeable?
Most viewers notice audio leading the picture by about one frame at 24 fps (roughly 45 ms) and lagging by about three frames (roughly 125 ms). Early sound is caught sooner because in life we always see an event before we hear it. Editors check sync on p, b and m, where the lips close.
How is lip sync done in dubbing?
Dubbing adapters rewrite the translated script so its closed-lip sounds, open vowels and line lengths fall close to the original actor's mouth movements. Voice actors record while watching the picture, and the editor slides each line into place. The aim is to match the most visible shapes, especially in close-ups.
How does AI lip sync video work?
AI lip sync starts with a shot of a face whose mouth is clearly visible, either filmed or generated, and a voice track. A lip-sync step then reshapes the mouth to match the audio, or the editor aligns the voice by hand. Frontal or three-quarter faces, clear mouth shapes and steady framing give the best results.
Related terms
4- Asynchronous SoundSound & Music
Asynchronous sound doesn't match what we see at that moment — a sound from another place or time, or deliberately out of step with the image — used for contrast, irony or subjectivity; accidental drift out of sync is simply an error.
- DubbingSound & Music
Dubbing replaces a film's original dialogue with a new recording — usually a translation performed by voice actors and fitted to the lip movements on screen.
- ADRSound & Music
ADR (automated dialogue replacement, once called looping) is re-recording an actor's lines in a studio after the shoot, in sync with picture, to fix noise, change a line or improve a performance.
- Voice ActingActing & Casting
Voice acting is performing a character with the voice alone — for animation, games, dubbing or narration — where timing, pitch and texture have to carry the whole performance.
Further reading
ITU-R Recommendation BT.1359, Relative timing of sound and vision for broadcasting.