Design System
View
Social · Foundations

Voice & tone

The brand voice, adapted for a scroll. Warmer and plainer than the website, never casual to the point of undermining authority. This page sets how Pierce Law sounds on social and how the tone shifts between LinkedIn and YouTube.

Outline — to be filled. Writing-voice structure agreed; copy written in a later pass. Defers to Brand · Voice & Tone for the core attributes. The AI voiceover mastery section below is the canonical Pierce source used by production script.json records.

Social voice principles

The through-line across every channel: plain-language, human, authoritative without being cold. How we sound when we're explaining the law to a worried person on their phone.

Tone by channel

LinkedIn leans professional and thought-leadership; YouTube leans approachable and educational. The same idea, dialed for the room.

Writing rules

Hooks in the first line, one clear CTA, plain over jargon, the emoji policy, and how to handle sensitive subjects (death, divorce, estates) with care.

Mastering the AI voice

How to control pitch, pacing, volume, cadence, and speed when generating voiceover for these videos — whether through the TimDS edge-tts command (timds video voiceover <slug>) or an approved prompt-driven engine like ElevenLabs. This section is the canonical client source; TimDS supplies the generic orchestration.

The house voice: calm, confident, knowledgeable, reassuring. A trusted advisor who has explained this a thousand times and still means it. Light NC softness is welcome; announcer polish is not. The listener just lost someone or is worried about their family — the voice should lower the temperature of the room, not perform.

The five levers

LeverWhat it controlsToo littleToo much
PitchPerceived age, weight, authorityThin, salesyFake gravitas, "movie trailer"
Speed (rate)Words per minuteRushed, nervous, salesySomber, funereal, "AI slow"
Pacing (pauses)Where silence landsWall of words, no room to absorbMelodramatic, disjointed
CadenceThe rhythm between lines — varianceMetronomic, obviously syntheticErratic, hard to follow
Volume/energyIntensity, brightnessFlat, depressedForced brightness, customer-service cheer
The single most important lesson from tuning these videos: uniformity is what sounds like AI. A real person does not read five sentences at the same rate, pitch, and energy. Cadence — deliberate variance between lines — is the lever that breaks the synthetic read. Tune it before you touch anything else.

edge-tts pipeline (scripts/generate_voiceover.py)

edge-tts is a fixed neural TTS: it cannot add breath, hesitation, or texture. Your levers are voice choice, rate, pitch, and punctuation — applied per job and per line.

Config anatomy

// src/videos/<slug>/script.json
{
  "voice": "en-US-AndrewMultilingualNeural",  // job-wide voice
  "rate": "-10%",                             // job-wide default speed
  "pitch": "-2Hz",                            // job-wide default pitch
  "lines": [
    // per-line overrides beat the defaults — THIS is where cadence lives
    { "id": "01-hook", "tts": "…", "rate": "-13%", "pitch": "-3Hz" },
    { "id": "02-list", "tts": "…", "rate": "-6%",  "pitch": "+0Hz" }
  ]
}
Gotchas. Pitch must carry an explicit sign — +0Hz, -2Hz — bare 0Hz errors. Regenerate only the selected topic with timds video voiceover <slug> --force, and only after explicit approval: replacing the locked take re-times every edited word in that production.

Rate (speed)

SettingReads asUse for
+0% to -4%Normal conversationalLists, facts, mid-video info
-5% to -9%Measured, consideredDefault for most lines
-10% to -13%Weighty, deliberateHooks, empathy beats, CTA
below -15%Somber / "depressed"Avoid — this killed early drafts

Pitch

Small moves only. -1Hz to -3Hz adds groundedness; beyond -5Hz sounds processed. Drop pitch with rate on the lines that carry weight (hook, CTA), return to +0Hz on brisk informational lines.

Cadence recipe — the anti-AI pattern. Give every line its own rate/pitch, shaped like a real read:

hook        slow + low      (-13% / -3Hz)   let the opening land
list/facts  near-normal     (-6%  / +0Hz)   pick the energy up
empathy     slow again      (-11% / -2Hz)   soften, don't drag
proof       middle          (-8%  / -1Hz)   conversational confidence
CTA         slow + low      (-12% / -3Hz)   settled, final

The spread matters more than the exact values: 5–7 points of rate variance across the video is the difference between "read by a person" and "rendered by a machine."

Pacing & volume through text — edge-tts has no other knobs:

  • Ellipsis "…" → real pause with a falling lead-in. Best before a payoff: "the details… matter."
  • Em dash "—" → short catch-breath: "Every piece protected — so nothing's left to chance."
  • Short sentences. "Wills. Trusts. Probate." — each period is a beat.
  • Contractions (nothing's, who'll, doesn't) and small spoken-word filler ("…families now…") loosen the grammar-perfect AI read.
  • Volume: edge-tts supports a volume param but don't use it for emphasis — word emphasis comes from sentence position and punctuation. Put the important word at the end of the sentence, after a pause.

Voices that fit the brand

VoiceReads as
en-US-AndrewMultilingualNeuralWarm low-mid male — the house voice
en-US-AvaMultilingualNeuralWarmest of the female mid-registers
en-US-EmmaMultilingualNeuralMore neutral/calm female alternative

ElevenLabs (prompt-driven)

When edge-tts's ceiling is too low (no breath, no texture), ElevenLabs adds real disfluency. Two prompts are involved: the voice design prompt and the annotated script.

Voice design prompt — write it as: register → pace/energy → persona → texture → inflection → negative space. Example (the house male voice, light NC accent):

Middle-aged American male, light NC accent — soft southern warmth, not
heavy, professional. Warm low-mid register, steady, conversational, not
somber. Trusted-advisor tone — said this before, means it. Slight breath,
minor pacing imperfections, not studio-clean. Downward inflection, never
upward. Confident, easy, not brisk or flat. Warm on empathetic lines, never
saccharine. No vocal fry, no over-enunciation, no announcer polish.

Rules of thumb, learned the hard way:

  • Ask for imperfection explicitly ("slight breath, minor pacing imperfections, not studio-clean") — otherwise you get the polished AI read.
  • Every positive needs a cap: "warm, never saccharine"; "southern, not heavy"; "slow" drifts to somber unless you say "not somber".
  • Inflection direction matters: "downward inflection, never upward" is what makes statements sound settled instead of uncertain.
  • End with negatives — the "no vocal fry, no announcer polish" list is as load-bearing as the positives.
  • Generate 3–4 takes and pick the least "read"-sounding one.

Script annotation (v3 audio tags) — inline [tags] direct the read per phrase, eleven_v3 only (v2 reads the brackets aloud):

[warm, easy] When you're planning for your family's future [beat] ...the details matter.
Wills. Trusts. Probate. [brighter] Every piece protected — so nothing's left to chance.
[friendly] We've walked this road with over ten thousand North Carolina families [breath] ...
[sincere] That's not just a slogan. [beat] It's five stars, family after family...
[confident] Pierce Law Group. Your family's legacy — protected.

Tag vocabulary by lever:

  • Cadence/energy: [warm, easy] [brighter] [friendly] [confident] — vary them line to line, same principle as edge-tts rate variance. Avoid stacking [calm] [thoughtful] [settled] — that combination reads as slow/depressed.
  • Pacing: [beat] (short), [pause] (longer), [breath] (audible inhale). Use [beat] + ellipsis before payoffs; ration [pause].
  • Volume/emphasis: tags like [sincere] shift intensity; also CAPS sparingly for a single stressed word.

Voice settings (Studio / API)

SettingValueWhy
Stability0.35–0.45The real "irregularity" knob — high stability = flat robot
Similarity~0.75Keeps the designed voice consistent
Style0.2–0.3Empathetic softening without theater
Modeleleven_v3Required for [tags]; eleven_multilingual_v2 if untagged

Script-writing rules that serve the voice (any engine)

  1. Write for the ear, not the page. Read every line aloud once before generating. If you stumble, the TTS will too.
  2. One idea per line. Each script.json line is one breath-group; the pause between clips is free pacing.
  3. Front-load calm, end settled. Hook slow, middle conversational, CTA slow and low. Never end on an up-note — falling intonation = "you can trust this."
  4. Spell out anything the engine will mangle: "pierce law dot com" in the tts text, merged back for captions via mergeWords.
  5. Numbers as words when they carry weight: "ninety days", not "90".
  6. Empathy lines get the slowest rate and the simplest words. "You just lost someone. The paperwork shouldn't be the hard part."

QA checklist before rendering

  • Listen to the concatenated take end-to-end, eyes closed. Does any stretch sound metronomic? Add per-line variance there.
  • Does it sound sad? Raise mid-video rates toward -5%, keep only hook/CTA slow.
  • Does the CTA sound like an ad? Slow it down, drop pitch, add a beat before the URL.
  • Check captions.json timings — merged words (piercelaw.com) intact?
  • durationMs totals ≈ target runtime (30s spot ⇒ ~33–36s spoken is fine; video pacing adds gaps).

AI voiceover — source examples

The spoken voice for AI-animated videos, generated with ElevenLabs (eleven_v3), applying the mastery guide above. Two takes from the same prompt and settings — pick whichever reads least "read."

SettingValue
Modeleleven_v3
Speed100% (no time-stretch)
Stability50%
Similarity boost75%

Take 1

Take 2

Voice design prompt used — see Mastering the AI voice above for the full breakdown of why each phrase is there:

Middle-aged American male, light NC accent — soft southern warmth, not
heavy, professional. Warm low-mid register, steady, conversational, not
somber. Trusted-advisor tone — said this before, means it. Slight breath,
minor pacing imperfections, not studio-clean. Downward inflection, never
upward. Confident, easy, not brisk or flat. Warm on empathetic lines, never
saccharine. No vocal fry, no over-enunciation, no announcer polish.

Script / pacing prompt used:

[warm, easy] When you're planning for your family's future [beat] ...the details matter.
Wills. Trusts. Probate. [brighter] Every piece protected — so nothing's left to chance.
[friendly] We've walked this road with over ten thousand North Carolina families [breath] ...each guided by an attorney who'll call you back.
[sincere] That's not just a slogan. It's five stars, family after family [beat] ...because the guidance doesn't stop at the paperwork.
[confident] Pierce Law Group. Your family's legacy — protected. Free consultation, at piercelaw dot com.
Reuse this pair. Same prompt + settings, generate 3–4 takes, keep the ones that sound least synthetic. Swap in new script lines but keep the tag vocabulary and the negative-space rules (never saccharine / never somber / no announcer polish) — that's what keeps every future take on-voice.