
Most lists of AI lip sync tools rank ten products on one scale, as if they all did the same thing.
They don't.
Re-syncing a real person in footage you already shot, making a still photo talk, and generating a presenter from a script are three different jobs, and a tool that is excellent at one is often useless at the other two.
AI lip sync takes a face on video and a piece of audio, then redraws the mouth and lower face so the speech looks like it came from that person.
Everything else in the frame stays as it was.
People use it for three jobs:
Pick the job first. The tool follows from it, and the wrong tool for the right job is still the wrong tool.
If you are re-syncing real footage, use a dedicated lip sync model, not an avatar platform.
Sync Labs, the company started by the researchers behind the Wav2Lip paper, is the reference point here, and it sells by the second through an app and an API.
If you are translating a talking-head video into other languages, a translation product with lip sync built in saves you stitching voice, translation and sync together yourself.
For that, use HeyGen. Its video translation lists 175 languages and dialects on paid plans from $29 a month, and lets you edit the translated script before it renders.
Our AI video translation roundup compares the rest.
If you are animating a photo or an AI-generated character, use a video model with lip sync attached, such as Kling.
Our Kling review and the guide to making video from a photo cover that route.
If there is no footage and no character yet, you don't need a lip sync tool at all.
You need an avatar platform, and the lip sync comes with it.
Start with our AI avatar tools roundup and the D-ID vs HeyGen decision.
A common mistake runs the other way: a founder with good real footage pays for an avatar subscription to fix one line, and ends up re-shooting the whole video as an avatar that looks worse than the original.
Once the job and the tool are picked, three things predict most bad results: profiles, hands, and the letter P.
The model is redrawing part of a face, and some faces are much harder to redraw than others.
Side profiles and steep angles are the first.
Sync Labs' own model documentation says its standard models give suboptimal results on profile views, and that its top model, sync-3, is the one built for profiles, close-ups and obstructions.
Hands, microphones and coffee cups crossing the mouth are the second.
On the standard models occlusion detection is a setting you switch on, and it slows processing.
Then there are the sounds where lips have to close fully: b, p and m.
A model that keeps the mouth slightly open through them gives the uncanny look people notice without being able to name.
Test every model on the word "maybe". Bad sync keeps the mouth open through the m, and once you see it you can't stop seeing it.
Two people in one shot is less of a problem than it used to be.
Sync Labs lists active speaker detection on all its models, which applies the sync only to whoever is talking.
The cheapest fix for all of this happens before the shoot.
If you know a video will be dubbed, film the speaker facing the camera, keep hands and mics below the chin, and light the face evenly.
Ten minutes of planning on set is worth more than any model upgrade.
For experiments, yes. For a client or a paid ad, check the license first.
Wav2Lip, the 2020 ACM Multimedia model most free tools still wrap, says on its repository that any form of commercial use is strictly prohibited.
A website offering free Wav2Lip lip sync doesn't change that license.
ByteDance's LatentSync is released under Apache 2.0, which allows commercial use.
Version 1.6 works at 512 by 512 pixels and, by its own README, needs about 18 GB of GPU memory for inference, against 8 GB for version 1.5.
A rented cloud GPU, yes. Most laptops, no.
Less than people expect for re-syncing, and it is priced by the second.
Sync Labs lists lipsync-2 at $0.04 to $0.05 per second, roughly $2.40 to $3 for a minute of video, and sync-3 at $0.107 to $0.133 per second, roughly $6.40 to $8 a minute.
Its subscription plans start at $5 a month for clips up to one minute and $19 a month for up to five.
HeyGen puts translation with lip sync on its Creator plan at $29 a month and up.
Do the math for a two-minute launch video dubbed into three languages: 360 seconds of sync at $0.05 a second is $18.
That is a rounding error next to paying a native speaker to check each translation, and you should still pay for that check, because a perfectly synced mistranslation is worse than an unsynced correct one.
Prices move often, so check the vendor's pricing page on the day you buy.
Sometimes, and the line is about whose words they are.
YouTube's disclosure policy requires a label when AI makes "a real person appear to say or do something they didn't do", and it lists cloning your own voice for voiceovers or dubs as something that does not need one.
So dubbing your own talk into Spanish in your own cloned voice is fine without a label.
The harder case is the one founders don't think of.
Lip-syncing a cofounder or a customer to fix a fumbled line feels like editing, but if the new line says something they didn't say, it is the exact case the policy describes.
Get their sign-off on the new words in writing before it ships, and label it if the change is more than a slip of the tongue.
If you are shooting a launch video with Flowjam and a dub is in the plan, say so in the brief so the speaker is filmed for it, and our guide to adding AI voiceover covers the audio half.
Adam is the founder of Flowjam, where he helps startups turn ideas into launch videos, product demos, and ads with AI video. He writes about AI video production, creative workflows, and go-to-market for early-stage teams.
AI lip sync is a model that redraws a speaker's mouth and lower face so it matches a new audio track. It is used to dub videos into other languages, fix lines in existing footage, animate photos or characters, and power AI avatars.
It depends on the job. For re-syncing real footage, a dedicated model such as Sync Labs is the reference point. For translating talking-head videos, HeyGen's video translation bundles voice, translation and lip sync. For animating a photo or character, a video model such as Kling works. For script-to-video, use an avatar platform.
Yes. LatentSync from ByteDance is open source under Apache 2.0 and allows commercial use, but needs a GPU with around 18 GB of memory for version 1.6. Wav2Lip is free but its license prohibits commercial use. Most paid tools also offer limited free tiers.
Sync Labs lists its lipsync-2 model at $0.04 to $0.05 per second, about $2.40 to $3 per minute of video, with plans from $5 a month. HeyGen includes translation with lip sync from its $29 a month Creator plan.
The most common causes are side profiles, hands or microphones covering the mouth, and sounds like b, p and m where the lips should fully close. Filming the speaker facing camera with nothing in front of the mouth fixes most of it before generation.