Sync AI Avatar Lips to Spanish Audio with Wav2Lip
Job to be done: Lip-sync an AI avatar video with Spanish audio using Wav2Lip
🇳🇬 Ways to use this in Nigeria
Ideas to get you started, adapt to your situation.
- Entrepreneur
Create Spanish explainer videos for your product using an AI avatar speaking fluently.
- Student
Generate Spanish language practice videos for your exams using an AI avatar.
What you’ll get
You will create a video where an AI avatar’s lips move in sync with a Spanish audio track. This approach works by using a specialized AI model to map spoken words onto a video of a face, making it appear as though the avatar is speaking the audio.
Tools you need
- edge-tts (free): A Python library to generate speech from text using Microsoft Edge’s TTS engine.
- Wav2Lip (free): An AI model and associated code that synchronizes lip movements in a video with an audio file.
Steps
-
Generate Spanish audio: Use edge-tts to create an audio file from your Spanish text. You will need to install edge-tts first. The author doesn’t share the exact installation command, but a common way is
pip install edge-tts.# Example command to generate audio (replace with your desired text and voice) # The author used: edge-tts --text "La IA no espera. Tu negocio tampoco." --voice es-ES-AlvaroNeural --write-media avatar-es.mp3 # You will need to run this in your terminal or a Python script.You should get an MP3 audio file (e.g.,
avatar-es.mp3). -
Prepare your avatar video: Have a video file of your AI avatar ready. The author used a video named
miguel-face-small.mp4. -
Run Wav2Lip for lip-syncing: Execute the Wav2Lip inference script with specific parameters to combine your avatar video and the Spanish audio. You will need to have Wav2Lip installed and its pre-trained model (
wav2lip_gan.pth) downloaded. The author’s successful command is provided below. Note that running Wav2Lip often requires a Python environment and specific dependencies.OMP_NUM_THREADS=4 python inference.py \ --checkpoint_path wav2lip_gan.pth \ --face miguel-face-small.mp4 \ --audio avatar-es.mp3 \ --wav2lip_batch_size 4 \ --resize_factor 2 \ --outfile miguel-avatar-es.mp4You should see the script processing frames and eventually produce an output video file (e.g.,
miguel-avatar-es.mp4) where the avatar’s lips match the Spanish audio.
Original source
This workflow is based on a blog post by migbolivar titled “5 attempts, 2 SIGKILLs, 1 non-existent flag: how I got a 4-second avatar to speak Spanish” shared on DEV Community. The author details their troubleshooting process to achieve lip-syncing for an AI avatar.
Notes & variations
- Free-tier alternatives: Both edge-tts and Wav2Lip are free and open-source tools, making them accessible without cost. However, running Wav2Lip can be computationally intensive and may require a machine with sufficient RAM.
- Common mistake: A common pitfall is using incorrect command-line arguments. The author highlights how they mistakenly used
--batch_sizeinstead of the correct--wav2lip_batch_size. Always double-check the available arguments using the tool’s help command (e.g.,python inference.py --help). - Tip for better results: If you encounter memory errors (SIGKILL), try reducing the
--wav2lip_batch_sizeand/or increasing the--resize_factor(e.g., to 2 or higher) to lower memory usage. The author found that a--resize_factorof 2 produced a working result even if it meant slightly lower video resolution.