Announcement

Introducing LipDub 2.0: Short-Form Localization Just Got a Fast Lane

Short-form teams work in loops: make, localize, review, repeat. The content is 30 seconds long, the deadline is tomorrow, and there are nine markets waiting. In that loop, every hour between submit and output is an hour the loop isn't spinning.

A round art piece with a large object.

LipDub Team

Quick Summary

LipDub 2.0 is live: 8x faster lip sync for short-form video, zero training, every existing model stays.

Abstracted UI elements showing the new lip sync model, LipDub 2.0, bolded and floating over the older, legacy models.

What zero training actually changes

Until now, every LipDub job followed the same path: preprocess the footage, train on the speaker, then render. That training step is what gives LipDub 1.0 its precision on detail-critical work, and it's exactly the right trade for a broadcast spot or an hour-long course. But for a 20-second ad, it meant most of your wait had nothing to do with generating your video.


LipDub 2.0 is an entirely different lip sync model. There's no preprocessing and no training step. You submit, it renders. That one change makes it 8x faster than LipDub 1.0, and it changes more than the clock:


  • No training fees. Where Premium and Ultra add a one-time training fee per speaker, LipDub 2.0 has none. You pay the same per-minute rate you always have, and nothing else.

  • No setup. New speaker, new campaign, new creator every week? Nothing to prepare. Upload and go.

  • No waiting to find out. The feedback loop between "submit" and "see it" collapses, which is where the real workflow change lives.


Handles the shots that used to break

Short-form footage is messy on purpose. Speakers turn mid-sentence. A hand brings the product right up to the face. There's a mic in front of the mouth for the whole take. The profile shot is the one your editor loved.


These are the frames where lip sync has traditionally fallen apart, and one broken frame is all it takes: a single flicker makes the whole video unusable. LipDub 2.0 handles them natively. Side profiles, three-quarter turns, occlusions, all by default, with no flickers, no artifacts, and no extra setup.


Fast enough to change how you work

LipDub 2.0 works on videos up to 2 minutes with a single speaker, rendered up to 1080p, in all 80+ languages.


That's the design. Social ads, UGC, influencer content, presentation clips, talking-head footage: the videos teams ship constantly, to every market, on deadlines that don't move. One upload becomes every market you sell in.


And when the render comes back in a fraction of the time, you don't just save minutes. You run the Spanish version, tweak a line in the translation editor, and run it again, before the review call, not after it. You test two openings in German instead of settling for one. Iteration stops being a luxury and starts being the workflow.


It's selected by default when you upload an eligible video. Run your next short and feel the difference.


Everything else stays exactly where it was

LipDub 2.0 is an addition, not a replacement.


Ultra, Premium, and Flash are training-based quality tiers of LipDub 1.0, and all of them remain available, unchanged, at the same pricing. For long-form video localization, multi-speaker scenes, and detail-critical work where texture carries the shot, LipDub 1.0 remains the right choice, and it isn't going anywhere.


The rule of thumb: under 2 minutes and one speaker, use LipDub 2.0. Anything else, use LipDub 1.0.


And short-form is where LipDub 2.0 starts, not where it ends. We're actively optimizing the model and expanding what it can handle, with support for longer content high on the list.


See it live

We're hosting a live walkthrough of LipDub 2.0: what it's for, how it handles the hard shots, and what's coming next. Bring your footage and your questions. Save your spot.


For API teams: benchmark it

LipDub 2.0 is live in the API with the same eligibility as the app. Run one representative job against your current pipeline and compare turnaround end to end. Read the LipDub 2.0 API docs.


FAQ


What is LipDub 2.0? A new, additional lip sync model built for short-form video. It's an entirely different technology: no preprocessing or training step, so jobs start rendering the moment you upload, and it handles motion natively, side profiles, turns, hands and objects in front of the face, without flickers or artifacts.


Is it replacing LipDub 1.0, Ultra, Premium or Flash? No. Ultra, Premium and Flash are training-based quality tiers of LipDub 1.0, and all of them remain available, unchanged. LipDub 2.0 is a new option alongside them.


What content works with LipDub 2.0? Videos up to 2 minutes with a single speaker, rendered up to 1080p, in any of our 80+ languages. Upload requirements are the same as every other LipDub model. The rule of thumb: under 2 minutes and one speaker, LipDub 2.0. Anything else, LipDub 1.0.


What if my video is longer than 2 minutes or has two speakers? Use LipDub 1.0, it seamlessly handles long-form and multi-speaker content today. And this launch is just the starting point for LipDub 2.0: we're actively optimizing the model and expanding what it can handle, with longer content high on the list.


How fast is it? 8x faster than LipDub 1.0. There's no training step, so rendering starts as soon as you submit.


Does it cost more? No. Pricing works the same as every other model: a per-minute rate that scales with your plan. And because there's no training step, there are no training fees, where Premium and Ultra add a one-time fee per speaker.


Where is LipDub 2.0 available? It's available in the API, with the same limits as the app: see the LipDub 2.0 API docs. In the app, it's in the Translation and Dialogue Replacement project types, and it's preselected by default for eligible videos.

What could your video do in 80 more languages?

Find out on your own footage in minutes.

No Credit Card required

What could your video do in 80 more languages?

Find out on your own footage in minutes.

No Credit Card required

What could your video do in 80 more languages?

Find out on your own footage in minutes.

No Credit Card required