AI Adoption GuideMarketingCreate
Synthetic-avatar video ads
Text-to-video avatars produce localized spokesperson ads without shoots, using tools like HeyGen or Synthesia.
Marketing processResearchPlanCreateLaunchMeasureReport
By Don, DoneThat’s AI coach · updated
What synthetic-avatar video ads are
Synthetic-avatar video ads turn approved scripts and brand inputs into talking-head or presentational spots without booking talent, locations, or a crew. A text-to-video system generates a digital spokesperson (or a likeness you have rights to use), lip-syncs the script, and exports a cut ready for review. Common production stacks include HeyGen, Synthesia, and similar avatar platforms, often paired with your existing edit or ad-ops tooling.
The job is not “AI makes the campaign.” The job is cheaper, faster iteration on spokesperson creative when the message is clear and the brand already knows what it will and will not claim. Video drafts come from the model; brand still approves talent likeness, claims, and final cut before anything runs.
When this approach fits (and when it does not)
Use synthetic avatars when you need many variants of a similar spot: language versions, market-specific offers, seasonality, or A/B tests on hook and CTA. Cost and turnaround dominate. A shoot for every locale or every offer change rarely pencils out for mid-funnel or always-on performance work.
Skip or sharply limit the approach when the creative depends on documentary footage, product-in-hand demos that must match physical packaging, celebrity talent without likeness rights, or claims that legal will not clear without live talent and full production documentation. If brand guidelines or a cleared script are missing, output should stay empty. Do not invent spokesperson lines, product claims, or visual brand elements to “fill the gap.”
Cost outcome shows up as fewer shoots per market and shorter revision cycles, not as free media or guaranteed performance. You still pay for platform seats, rendering, editing polish, and human review time.
Inputs you need before generation
Treat generation as a pipeline with hard gates, not a chat box.
Required inputs
- Cleared script (or locked script template with approved claim slots)
- Brand guidelines: voice, visual system, prohibited claims, disclosure rules
- Likeness and talent policy: which avatars are allowed, which real-person likenesses need contracts, and who signs off
- Destination specs: aspect ratio, length, captions, end card, tracking pixels or UTM conventions
Optional but high-value
- Glossaries and locale notes for markets you localize into
- Reference cuts that define pacing and on-screen hierarchy
- Offer and legal footnotes that must appear on screen or in voiceover
If script or brand guidelines are missing, return no video draft. Surface a clear blocker instead of a speculative cut. That empty-output rule protects reviewers from “almost right” creative that still embeds unapproved claims or off-brand presentation.
How a practical workflow runs
- Brief and rights check. Confirm offer, audience, markets, and likeness rights before anyone prompts a platform.
- Script lock. Produce or adapt copy under brand and legal constraints. Prefer brand-grounded copy workflows so claims stay inside approved language.
- Avatar and scene selection. Choose only pre-approved avatars, outfits, backgrounds, and brand kits. Do not introduce a new face or wardrobe without the same review bar you would apply to talent casting.
- Generate drafts. Create a small set of cuts (for example primary plus one alternate hook), not dozens of unreviewed renders.
- Human review. Brand and legal check likeness, claims, captions, disclosures, and visual compliance. Marketing lead owns creative judgment; legal owns claim risk.
- Polish and package. Captions, lower thirds, end cards, loudness, and platform export settings.
- Launch and learn. Feed performance back into script and hook tests; keep avatar and claim changes on the approval path.
Human-in-the-loop is non-negotiable at likeness and claims. Models draft video; they do not clear talent usage or advertising law for you.
Localization without a full reshoot
Avatar ads earn most of their cost advantage in localization. One approved narrative structure can support multiple languages if you control translation quality and on-screen text. Keep claim equivalents reviewed per market. Do not assume a machine translation of a cleared English claim is cleared elsewhere.
Practical pattern: lock the English (or source) script and visual scaffold first, then generate locale variants with the same avatar family and layout. Review each locale for idioms, regulated phrasing, pricing displays, and mandatory disclosures. Where a market requires different proof or disclaimer language, treat that as a new cleared script, not a cosmetic text swap.
Brand safety, likeness, and disclosure
Likeness. Synthetic does not mean unrestricted. Using a face that resembles a real person, employee, or influencer without rights creates the same risk as unauthorized talent. Stick to platform stock avatars your policy allows, or to likenesses covered by contract. Brand approval is required before publish.
Claims. Avatar delivery does not soften substantiation requirements. If the script would be blocked in a live shoot, block it here. Prefer approved claim libraries and refuse generation when required claim inputs are absent.
Disclosure. Disclose synthetic media where law, platform policy, or brand policy requires it. Examples include labeling AI-generated or digitally created spokesperson content in captions, end cards, or platform fields when those rules apply. Build disclosure into templates so reviewers do not rely on memory.
Quality bar. Reject uncanny lip sync, wrong product names on screen, mismatched logos, or backgrounds that imply partnerships you do not have. Cost savings evaporate if takedowns or forced edits follow.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first