Kling AI Talking-Head Avatar Prompt
This prompt directs a believable AI talking-head presenter in Kling: medium close-up framing, soft office lighting, and explicit performance direction (natural lip-sync, subtle head movement, blinks, warm expression). Describe the presenter and background in the brackets โ the performance notes are what separate a convincing avatar from a mannequin reading cue cards.
Natural lip-sync talking heads in Kling: framing, lighting, and performance direction for believable AI presenters.
Talking-head presenter video, 8 seconds: [PERSON DESCRIPTION โ e.g. a friendly woman in her 30s, business casual] speaking directly to camera, natural conversational lip-sync, subtle head movement and occasional blinks, warm genuine expression. Framing: medium close-up, chest up, centered, eye-level camera. Setting: [BACKGROUND โ e.g. bright modern home office, softly blurred]. Lighting: soft key light from the left, gentle fill, no harsh shadows. Performance note: calm, confident delivery, slight smile. Photorealistic, no text overlays, no watermark.
See what this prompt creates
Talking-head avatars live or die on micro-motion. The uncanny valley isn't in the face โ modern models render faces beautifully โ it's in the stillness: a head that doesn't drift, eyes that don't blink, a mouth moving with metronomic precision. This prompt attacks exactly that by directing performance like a film director would: "subtle head movement," "occasional blinks," "slight smile." Those five words do more for believability than any resolution setting.
Kling is the current sweet spot for this format โ its lip-sync and natural motion handling lead the pack for presenter-style content โ but the prompt's real content is framing discipline. "Medium close-up, chest up, centered, eye-level" is the news-anchor framing every viewer subconsciously trusts. Deviate from it (low angles, extreme close-ups) and the avatar reads as dramatic or threatening regardless of the script.
Getting the best result
- Keep it to 8 seconds per generation. Lip-sync coherence degrades over longer clips. Generate in 8s chunks per script paragraph and stitch โ viewers never notice the cuts.
- Write the script for the mouth, not the page. Short sentences, conversational words, no tongue-twisters. The avatar can only look natural saying things a human would comfortably say.
- Blur the background. "Softly blurred" isn't just aesthetic โ it hides the background warping that betrays AI video fastest. Depth of field is camouflage.
- Match the voice to the face. If you're dubbing with ElevenLabs or similar, pick a voice whose age and energy match the visual. Mismatched dubbing breaks the illusion instantly.
3 variations to try
Customer testimonial: "casual setting, slightly off-center framing, enthusiastic genuine tone" โ the UGC variant (pair with the testimonial prompt below).
News-style update: "studio backdrop, formal attire, measured delivery" for company announcements and changelogs.
Multilingual: generate the same visual with different script dubs โ one avatar, five markets. Keep the performance neutral so no language looks mismatched.
How to use this prompt
- Copy it โ hit the copy button above to grab the raw prompt text.
- Fill the brackets โ replace anything in [BRACKETS] with your own details.
- Paste into your AI tool โ works with ChatGPT, Claude, Gemini for text prompts; Midjourney, DALL-E or Stable Diffusion for image prompts; Runway, Pika or Sora for video.
- Iterate โ tweak one bracket at a time until the output is exactly right.
Frequently asked questions
How do I add voice to the talking head?
Two routes: Kling's own lip-sync features (upload audio, it syncs the mouth), or generate silent and dub in editing with ElevenLabs, CapCut's text-to-speech, or your own recording. For the most natural result, record or generate the audio first, then sync the video to it โ not the reverse.
Why does my avatar look like a mannequin?
Missing micro-motion. Add explicit direction: blinks, subtle head drift, breathing-scale shoulder movement, varied mouth shapes. A perfectly still head with moving lips is the uncanny valley's home address โ this prompt includes the performance notes that fix it.
Can viewers tell it's AI?
At 8 seconds, usually not โ especially with motion and a dubbed voice. Over longer stretches, sharp-eyed viewers catch it. Disclose AI use where platform rules or ad policies require it; the tech is for scale, not deception.
What's the best aspect ratio?
9:16 for TikTok/Reels/Shorts (where talking heads live), 16:9 for YouTube explainers and course content. Generate in the delivery ratio โ cropping a 16:9 talking head to vertical butchers the framing.