# A voice that performs the story

The owner rejects the historical master as emotionless. It is not accepted under this
new gate. A prior model saying “emotional and deliberate” is insufficient: it did not
bind phrase-level expectations to audible evidence.

We listened through five bounded native-media audits: four reference excerpts plus the
historical opening. Full receipts and local clocks are in [vocal-reference](vocal-reference/source-alarm.json).
The model's adjectives remain interpretations, not acoustic measurements. The reference
payoff audit misidentifies the remote helper as a doctor; do not copy that role label.

| Reference evidence | Audible delivery function | Ukrainian transfer |
|---|---|---|
| [Opening, source 0–32 s](vocal-reference/source-alarm.mp4): rapid worried questions against measured answers | Contrast urgency with attempted reassurance; small hesitation can expose hidden fear. | The opening question needs concern; the mother's excuse should be lightly evasive; the narrator's discovery should lose that lightness. |
| [Explanation, source 77–102 s](vocal-reference/source-explanation.mp4): patient explanation interrupted by a quick question | The listener changes the pace; facts are delivered to someone who needs them. | Ingredient copy needs thought groups, emphatic distinctions and responsive curiosity; avoid a continuous label-reading tone. |
| [Skepticism, source 160–185 s](vocal-reference/source-skepticism.mp4): polite dismissal, pointed question, patient correction | Objection → answer is audible before the picture explains it. | “І чим він відрізняється…” gets a sharper rising challenge; the answer slows and grounds itself. |
| [Payoff, source 284–314 s](vocal-reference/source-payoff.mp4): softened gratitude opens into a more direct close | Relief changes timbre and rhythm; the CTA gains clarity while preserving trust. | “Сьогодні твоє місце — поруч” needs a smile and release. Ordering instructions need clear separate steps, then a warm final landing. |

[The complete vocal score](vocal-score.json) assigns all 21 original paragraphs a speaker,
context, intention, emotion, contour, pace, energy, pause, emphasis, negative direction,
source reference and listening criterion. Quoted speech changes the narrator's stance;
it must not become a cartoon imitation of age or gender. Emotion means fit to the moment,
not shouting or crying throughout.

The delivery ladder is **curiosity → realization → unease → vulnerable listening →
determination → skepticism → trust → cautious hope → informed choice → relief → joy →
gratitude → helpful invitation**. Reserve the largest release for the photo callback.
Quiet factual language still has an intention and rhythmic contrast.

New speech uses provider-compatible direction in the actual request, not a separate
document the provider never sees. The expressive runner supports the documented
[Kie Eleven v3 dialogue contract](https://docs.kie.ai/market/elevenlabs/text-to-dialogue-v3).
Tags are suggestions that must be auditioned; punctuation and tags can shape delivery.
[ElevenLabs prompting guide](https://elevenlabs.io/docs/best-practices/prompting).

For Ukrainian, every multisyllabic occurrence in the submitted script receives contextual
stress adjudication and U+0301. The ledger binds the exact script, direction score and
request bytes. Preserve plain spelling for captions by removing only U+0301. Do not use
generic Unicode accent stripping, which can corrupt Ukrainian letters.

**Acceptance requires listening, not metadata.** Each occurrence must have a local time,
heard stress, intact consonants and pass/fail/uncertain. Each vocal line also needs heard
delivery, context fit, no flat reading and no overacting. A correct ASR transcript, loudness,
pitch range, tag or written accent cannot approve it. Check the opening, repetitions,
quotes, product name and last CTA word. Repair a failed phrase in a new attempt, re-align,
then listen to the complete assembled voice because transitions can expose seams.

The short directed audition is an experiment, not a replacement for the full 5:49 voice.
Failed provider requests and listening verdicts remain visible in the review. New picture
cannot use it as a complete narration master.
