An AI talking avatar is a video in which a still portrait appears to speak a script: the voice is synthesized from your text and the mouth is animated to match every phoneme. On Aura AI you upload one front-facing photo, type the script and generate; the result is a lip-synced clip with no watermark, rendered by video models with native speech such as Kling 3.0, Seedance 2.5, Veo 3.1 and Grok Imagine Video 1.5. This tutorial covers the photo, the script, the model choice with credits, and the fixes for the usual lip-sync problems.
In This Article
How lip sync from a single photo works
Two systems run in sequence. First, a text-to-speech engine turns your script into a voice track with natural intonation, pauses and emphasis. Second, a video model animates the face in the photo: it breaks the generated speech into phonemes, maps each one to the matching mouth shape on that specific face, and adds subtle head motion, blinks and facial micro-movements so the result feels alive rather than pasted. There is no rigging, no keyframing and no source footage; the photo-to-talking-video conversion happens in a couple of minutes.
What changed in 2026 is that the animation step is no longer a dedicated lip-sync model bolted onto a still. Video models such as Kling 3.0, Seedance 2.5, Veo 3.1 and Grok Imagine Video 1.5 generate picture and speech in one pass, so mouth movement, expression and audio are produced together instead of being aligned afterwards. Kling 3.0, released in February 2026, went further with phoneme-level lip sync for multiple characters, so two faces in one frame can each speak their own lines.
Step-by-step: make a photo talk
- Pick the right portrait. Front-facing, evenly lit, face clearly visible, mouth not covered by hair, hands or a microphone. Headshots, selfies and studio portraits all work; sunglasses and heavy shadows do not. If you have no photo, generate one with text to image in the same session and use it as the source.
- Type the script. Write exactly what the avatar should say. Punctuation drives pacing: a full stop is a beat, a comma is a breath. Put dialogue in quotes if you also describe the scene. Plan on roughly two and a half words per second, so a 10-second clip is about 25 words.
- Choose the language and voice. Type the script in the target language and the engine generates a matching voice with native pronunciation. English, Spanish, French, German, Chinese, Arabic and dozens of others are supported.
- Generate, then review the mouth. Watch the clip once with sound off. If the lips lead or lag the audio, shorten the script or split it into two clips; long paragraphs are the most common cause of drift.
- Download. Every talking avatar video on Aura AI is delivered with no watermark, so it goes straight into an edit, a course or a social post.
Which model to use for a talking avatar
The talking avatar tool renders through video models with native speech and lip sync. Specs and credits below are the figures published on each model page on Aura AI.
| Model | Credits | Resolution | Duration | Audio and lip sync | Plan |
|---|---|---|---|---|---|
| Kling 3.0 | 3 credits for 5 s, +0.35 per extra second | 1080p | 3 to 15 s | Native audio, most accurate mouth movement | Every plan |
| Kling 3.0 Pro | 5 credits for 5 s, +0.5 per extra second | 1080p | 3 to 15 s | Native audio | Starter+ |
| Seedance 2.5 | 6 credits for 4 s at 720p, +0.35 per extra second; 1080p x2 | 480p to 1080p | 4 to 15 s | Single pass, lip sync in 10+ languages | Starter+ |
| Seedance 2.0 | 5 credits for 5 s at 720p, +0.3 per extra second; 1080p x2, 4K x4 | 720p to 4K | Up to 15 s | Stereo audio, character consistency across shots | Starter+ |
| Veo 3.1 | 7 credits for 4 s at 1080p, +1 per extra second; 720p x0.75, 4K x1.8 | 720p to 4K | 4, 6 or 8 s | Native dialogue, realistic faces | Starter+ |
| Veo 3.1 Lite | 2 credits flat | 720p or 1080p | 4, 6 or 8 s | Native audio | Every plan |
| Grok Imagine Video 1.5 | 4 credits for 5 s at 720p, +0.4 per extra second; 480p x0.7 | 480p or 720p | Up to 15 s | Native audio, image to video | Every plan |
Kling 3.0 for accuracy
The most precise mouth shapes on the platform and the cheapest entry point at 3 credits. Start here for spokesperson clips and anything where the viewer stares at the lips.
Veo 3.1 for realism
The most natural skin, eyes and lighting, with clean dialogue audio. Pick it when the face has to pass as a real recording, and use Veo 3.1 Lite at 2 credits for drafts.
Seedance 2.5 and 2.0 for longer scripts
Seedance keeps a character consistent across shots and generates speech in more than ten languages in one pass. Seedance 2.0 is also the route to a 4K master.
Sora 2 and Grok for delivery and speed
Sora 2 adds expressive, actor-like delivery at 1080p. Grok Imagine Video 1.5 is the fast, low-credit option for short social clips with sound.
Script tips for natural delivery
- Write for the ear. Short sentences, contractions, one idea per line. Read it aloud once; if you stumble, the voice will too.
- Spell out numbers and acronyms the way you want them spoken: "twenty-four seven", "A-P-I".
- Use punctuation as direction. An ellipsis adds a pause, an exclamation mark lifts the pitch, a question mark raises the ending.
- Describe the tone in the prompt when the model supports it: "warm, unhurried, slight smile" changes the performance, not just the words.
- Keep clips under 15 seconds. That is the maximum on Kling 3.0 and Seedance, and shorter clips sync more tightly. For a longer message, generate several clips and join them in the AI video editor, or read the tutorial on 10-minute AI speech videos.
Common problems and how to fix them
| Symptom | Usual cause | Fix |
|---|---|---|
| Mouth barely moves | Face too small in frame or turned away | Crop tighter on the face, use a front-facing shot |
| Lips drift out of sync near the end | Script too long for one clip | Split into 8 to 10 second clips |
| Teeth or jaw look wrong | Source photo has a closed, tight smile or low resolution | Use a neutral, relaxed expression at higher resolution |
| Voice sounds robotic | Flat punctuation, long sentences | Add commas and full stops, describe the tone |
| Face changes between clips | Different seeds or models per clip | Keep the same model and source photo; Seedance 2.0 for multi-shot consistency |
From one clip to a full ad
A single talking avatar covers a product intro, a tip video, a greeting or one line inside a longer edit. When the job needs several scenes, a written hook, b-roll and product shots cut for TikTok or Reels, move up to AIGC Studio. Its Director agent casts the avatar, writes the script, renders the voiceover with lip sync and assembles the multi-scene edit, using the same talking-avatar engine as one building block.
How to do it in Aura AI
- Open the talking avatar tool in the web app and upload a front-facing portrait, or generate one with text to image first.
- Paste the script in the language you want spoken, keeping each clip under about 25 words for 10 seconds.
- Pick the model: Kling 3.0 for lip accuracy, Veo 3.1 for realism, Seedance 2.5 for multilingual speech. The credit cost updates before you generate.
- Generate and download the clip with no watermark, then extend, upscale or drop it into an edit in the same workspace.
Make a photo talk
Upload a portrait, type the script, and get a lip-synced talking avatar in minutes, with no watermark.
Create a Talking AvatarFrequently Asked Questions
What is an AI talking avatar?
An AI talking avatar is a video generated from a still portrait in which the person appears to speak a script you typed. A text-to-speech engine produces the voice and a video model animates the mouth, head and expressions to match it, so no filming, rigging or 3D model is needed.
Which AI model has the best lip sync?
On Aura AI, Kling 3.0 gives the most accurate mouth movement and costs 3 credits for a 5-second clip on every plan. Veo 3.1 pairs the most realistic faces with clean dialogue audio, and Seedance 2.5 generates speech with lip sync in more than ten languages in a single pass.
What kind of photo works best?
A front-facing portrait with even lighting and the whole face visible. Avoid sunglasses, hands or hair over the mouth, side profiles and low-resolution crops. Headshots, casual selfies and AI-generated portraits all work.
How long can a talking avatar video be?
Single clips run up to 15 seconds on Kling 3.0, Seedance 2.5 and Grok Imagine Video 1.5, and up to 8 seconds on Veo 3.1. For a longer message, generate several clips from the same photo and join them in the AI video editor.
Do talking avatar videos have a watermark?
No. Talking avatar videos generated on Aura AI are delivered with no watermark and are ready for marketing, courses, presentations and social media without post-processing.