๐ŸŽฌ Video Prompts
#kling#talking-head#avatar#explainer#ugc

Kling AI Talking-Head Avatar Prompt

This prompt directs a believable AI talking-head presenter in Kling: medium close-up framing, soft office lighting, and explicit performance direction (natural lip-sync, subtle head movement, blinks, warm expression). Describe the presenter and background in the brackets โ€” the performance notes are what separate a convincing avatar from a mannequin reading cue cards.

Natural lip-sync talking heads in Kling: framing, lighting, and performance direction for believable AI presenters.

Talking-head presenter video, 8 seconds: [PERSON DESCRIPTION โ€” e.g. a friendly woman in her 30s, business casual] speaking directly to camera, natural conversational lip-sync, subtle head movement and occasional blinks, warm genuine expression. Framing: medium close-up, chest up, centered, eye-level camera. Setting: [BACKGROUND โ€” e.g. bright modern home office, softly blurred]. Lighting: soft key light from the left, gentle fill, no harsh shadows. Performance note: calm, confident delivery, slight smile. Photorealistic, no text overlays, no watermark.
More video prompts prompts
Try it now ChatGPT โ†— Grok โ†—
Example output

See what this prompt creates

AI-generated example still frame for the 'Kling AI Talking-Head Avatar Prompt' prompt
Example frame generated by AI to show the look this prompt aims for. Your results will vary by tool and the details you fill in.

Talking-head avatars live or die on micro-motion. The uncanny valley isn't in the face โ€” modern models render faces beautifully โ€” it's in the stillness: a head that doesn't drift, eyes that don't blink, a mouth moving with metronomic precision. This prompt attacks exactly that by directing performance like a film director would: "subtle head movement," "occasional blinks," "slight smile." Those five words do more for believability than any resolution setting.

Kling is the current sweet spot for this format โ€” its lip-sync and natural motion handling lead the pack for presenter-style content โ€” but the prompt's real content is framing discipline. "Medium close-up, chest up, centered, eye-level" is the news-anchor framing every viewer subconsciously trusts. Deviate from it (low angles, extreme close-ups) and the avatar reads as dramatic or threatening regardless of the script.

Getting the best result

  • Keep it to 8 seconds per generation. Lip-sync coherence degrades over longer clips. Generate in 8s chunks per script paragraph and stitch โ€” viewers never notice the cuts.
  • Write the script for the mouth, not the page. Short sentences, conversational words, no tongue-twisters. The avatar can only look natural saying things a human would comfortably say.
  • Blur the background. "Softly blurred" isn't just aesthetic โ€” it hides the background warping that betrays AI video fastest. Depth of field is camouflage.
  • Match the voice to the face. If you're dubbing with ElevenLabs or similar, pick a voice whose age and energy match the visual. Mismatched dubbing breaks the illusion instantly.

3 variations to try

Customer testimonial: "casual setting, slightly off-center framing, enthusiastic genuine tone" โ€” the UGC variant (pair with the testimonial prompt below).

News-style update: "studio backdrop, formal attire, measured delivery" for company announcements and changelogs.

Multilingual: generate the same visual with different script dubs โ€” one avatar, five markets. Keep the performance neutral so no language looks mismatched.

How to use this prompt

  1. Copy it โ€” hit the copy button above to grab the raw prompt text.
  2. Fill the brackets โ€” replace anything in [BRACKETS] with your own details.
  3. Paste into your AI tool โ€” works with ChatGPT, Claude, Gemini for text prompts; Midjourney, DALL-E or Stable Diffusion for image prompts; Runway, Pika or Sora for video.
  4. Iterate โ€” tweak one bracket at a time until the output is exactly right.

Frequently asked questions

How do I add voice to the talking head?

Two routes: Kling's own lip-sync features (upload audio, it syncs the mouth), or generate silent and dub in editing with ElevenLabs, CapCut's text-to-speech, or your own recording. For the most natural result, record or generate the audio first, then sync the video to it โ€” not the reverse.

Why does my avatar look like a mannequin?

Missing micro-motion. Add explicit direction: blinks, subtle head drift, breathing-scale shoulder movement, varied mouth shapes. A perfectly still head with moving lips is the uncanny valley's home address โ€” this prompt includes the performance notes that fix it.

Can viewers tell it's AI?

At 8 seconds, usually not โ€” especially with motion and a dubbed voice. Over longer stretches, sharp-eyed viewers catch it. Disclose AI use where platform rules or ad policies require it; the tech is for scale, not deception.

What's the best aspect ratio?

9:16 for TikTok/Reels/Shorts (where talking heads live), 16:9 for YouTube explainers and course content. Generate in the delivery ratio โ€” cropping a 16:9 talking head to vertical butchers the framing.

Related prompts

Keep exploring the vault.