Video DNA passport — evidence before interpretation
Version 2 production contract. Read the callable meta-function,
the populated contract, the vocal performance passport
and the LevchykUGS study. New generation goes through
scripts/compile_narrative.py; a descriptive document alone cannot pass the factory gates.
Required passports: video craft; transferable persuasion function; each character;
product identity/use; product-driven change; wardrobe/props; attention; vocal performance;
Ukrainian pronunciation; regional visual references. Link them by stable IDs and exact
asset hashes. Describe what the provider must make and what the reviewer must actually see/hear.
A reusable way to reverse-engineer an ad and hand its structure to production. The unit of analysis is a change: someone learns, wants, feels, believes, acts, or sees something different. A shot list alone misses the causal story; a marketing summary alone misses the craft.
1. Source and provenance
Record asset ID, original path/URL, SHA-256, duration, dimensions, frame rate, audio tracks, languages, caption method, date analyzed, model and request receipts. Record campaign claims separately: who supplied spend, currency, date range, placements, attribution window, revenue, and whether raw account data was inspected. A large spend is a reason to study the creative, not proof that every feature caused success.
2. Evidence vocabulary
Every substantive annotation carries start_s, end_s, evidence, status, confidence, and reviewer. Use these statuses:
| Status |
Meaning |
Example |
| measured |
Derived from media or a reproducible tool |
Runtime, frame size, candidate cut at 18.4 s |
| observed |
Directly visible or audible |
A daughter grips a mug while asking for help |
| asserted-in-ad |
A character or overlay makes the claim |
A claimed blood-pressure improvement |
| seller-asserted |
Product page or business brief says it |
Return period, formula, delivery promise |
| interpreted |
A plausible account of what the craft does |
The mug and blanket signal safety |
| hypothesized |
A causal claim requiring audience testing |
Delaying the reveal may improve completion |
| unknown |
Evidence absent or contradictory |
Real conversion lift; whether performers are synthetic |
Confidence measures certainty about an annotation, never predicted sales. High = directly legible/audible and corroborated; medium = clear pattern with approximate timing; low = ambiguity, inference, or unverified source assertion. A transcription system must not label summaries “verbatim.” Preserve rejected provider output with reasons. Test a whole-video model’s clock against distinctive start/middle/end anchors. If it assigns an event to the wrong time, reject those timecodes and adjudicate short excerpts with explicit media hashes and known offsets; the model must not become the authoritative edit clock.
3. Source fact sheet
Capture category, format, aspect, visual realism, dialogue/narration mode, perspective, main viewer promise, apparent audience, destination/placement if known, first product appearance, first spoken brand, first benefit, first CTA, final CTA, total words, speaking time, silence, and normalized reveal position. Separate screen time from story time. Count speaking roles independently of visible roles and extras.
4. Hook microscope
Inspect the first frame and 1, 3, 5, 10, and 30 seconds. At each checkpoint record: visible subject, spoken proposition, information withheld, immediate stakes, emotional invitation, motion or edit interrupt, legible caption, and reason the next second is needed. Identify overlapping hooks: visual incongruity, danger, relationship, knowledge gap, contradiction, identity, and specificity. Do not reduce the hook to its first sentence.
5. Story and belief timeline
Use one row per meaningful turn, not a forced universal scene count:
time → event → cause → character goal → obstacle → decision → viewer question → emotional change → belief changed → evidence → next beat
Identify inciting incident, commitment to the quest, failed approaches, mentor arrival, explanation, product discovery, objections, adoption, elapsed-time proof, relational payoff, and commercial close when present. Mark absent stages rather than inventing them. A beat may perform several functions. Explain why the next beat follows from this one.
6. Character passport
For each recurring person: stable identity markers; screen and story role; relationship graph; what they know at each stage; external want; internal need; fear; misconception; agency; turning point; dialogue style; recurring props; wardrobe; body posture; gaze; brow; eyelids; mouth tension; gesture; spatial position; entrance and exit. Record the narrator separately from the protagonist. A child may be the plot's agent while the parent is the buyer avatar.
7. Product-concern continuity
Identity is independent of the concern. Bind affected person, exact body/object region,
separate observable markers, baseline, actual use, story days, intermediate states and
resolved endpoint. Keep age, face shape, body size, hair, lens, pose and lighting stable.
The cream example requires the problem clearly present in week one, visibly lower in
week two, lower again in week three, and visually absent at week four. “Less severe”
cannot pass as “resolved.” First use remains at baseline; it is not instant erasure.
Use four distinct matched references, hold each state long enough to inspect, and use a
short matched dissolve at labeled elapsed-time boundaries. The generator must not morph
a face mid-shot, change clothes or relight the neck to imply improvement. Review actual
pixels at state, board, clip and final-film levels. A smile, caption, removed scarf or
another actor's neck cannot stand in for the result. These are illustrative visual
targets, not clinical measurements or guaranteed results. Mentor and family members are
explicitly excluded from the target concern. Other categories use their own observable
and timescale: a cleared error queue or a controlled alert test, not cosmetic weeks.
Wardrobe and prop changes have passports too. Each garment has an owner, ID, neckline,
color, sleeve, material, fit and reference. Every shot assigns a wardrobe state to every
visible person. Preserve states across reverse angles and adjacent scenes. A change needs
an explicit event and visible action, such as mother putting on the high collar or
changing into the birthday blouse. The daughter cannot inherit the mother's concealment
garment. Bind narrated actor → action → object → expected visible result to the audio
words; inspect the actual frame interval and both sides of every cut. Reject a mother
already wearing the supposedly newly offered blouse and wrong-sleeve product inserts.
8. Emotional architecture
Map surface discomfort → functional loss → social consequence → identity wound → threatened relationship → desired future. For each beat record the emotion invited, valence, arousal, perceived control, intimacy, and uncertainty. Optional 0–4 curves are analyst ratings with written anchors, never audience measurement. Include recovery intervals, dignity/absolution, and whether persuasion relies on fear, shame, belonging, hope, anger, curiosity, relief, or agency.
9. Retention ledger
Give every open loop an ID, opening time, exact question, reminders/escalations, answer time, delay, and whether the answer creates another loop. Separate plot suspense from product curiosity. Track microhooks, specificity, character reversals, location changes, interruption, reveals, callbacks, pattern breaks, progressive repetition, surprise, payoff and residual uncertainty. Measure viewer drop-off only when actual platform data exists.
10. Persuasion ledger
Describe awareness and market sophistication as hypotheses. Store problem definition; mechanism of problem; mechanism of solution; comparison/failed alternative; authority source; discovery story; demonstration; proof claim; objection; answer; commitment; offer; risk reversal; urgency; CTA verb; destination; and transaction friction. For each mechanism ask whether it is literal, metaphorical, asserted, or independently supported. Distinguish a credential cue from a verified credential, a character result from clinical evidence, and a seller review from an audited sample.
11. Visual grammar
Annotate shot scale, camera angle, lens impression, movement, horizon, subject placement, eyeline, negative space, foreground/midground/background, occlusion, scene geography, focus/depth, color temperature, contrast, light direction, motivated practical light, wardrobe, production design, weather/time, props and motifs. Note visual continuity and probable artifacts without declaring the production method from appearance alone. Use contact sheets for identity and color arcs; moving excerpts for acting and continuity.
12. Editing grammar
Create a candidate cut list from scene detection, then adjudicate boundaries. Record hard cut, dissolve, fade, insert, reaction shot, shot/reverse shot, montage, match action, flashback, temporal ellipsis, repeat and hold. Measure shot durations and cadence by act, first 10-second cut count, close-up proportion, visual novelty interval and time spent on reaction vs explanation. A scene detector's threshold output is not a ground-truth shot count.
13. Audio and captions
Audio is the clock. Preserve the approved spoken script, voice IDs and language. Every
line requires context, speaker, intention, emotion, pitch contour, energy/timbre, pace,
pause, emphatic words, quoted-role switch, source delivery reference, negative direction
and listening criterion. Analyze actual reference speech, not just the words or face.
Questioning, hesitant reassurance, patient explanation, doubt, joyful release and the
CTA cannot share one generic warm reading. The actual provider request must carry the
directions its model supports. Listen for each intended turn; reject flat or overacted
delivery independently of pronunciation. Measure loudness, peak, silence and duration
as separate technical checks. For this format do not create or decompose music.
Ukrainian stress is an occurrence-level contract: contextual evidence for every inflected
word, U+0301 after the stressed vowel of every multisyllabic word in actual submitted
text, explicit foreign-brand pronunciation, and exact script/request/score hashes.
Listen to every occurrence in the real recording and record intended vs heard stress,
consonants, time and pass/fail/uncertain. ASR spelling and written accents do not prove
stress. Uncertainty blocks acceptance; repair in a versioned attempt, realign and recheck
the assembled voice. Strip only U+0301 for plain captions, preserving Ukrainian letters.
Captions carry the heard words, real alignment source, cue start/end, line breaks, position, face/product occlusion, contrast and phone legibility. Distinguish measured provider word alignment from manually verified timing and from estimates. Record where picture anticipates, illustrates, reacts to, or contradicts a spoken phrase. Reject orphaned words and cuts through meaning units.
14. Localization and product mapping
For each source beat, specify function kept → local event → product truth → visual → audio → reason for change. Translate idiom, social roles, settings, familiar objects, authority, currency, transaction model, calendar and CTA. Localization is not flag decoration. Keep dignity and recognizable daily life. Do not transplant a country's supposed biological superiority, disease claims, numerical outcomes, discounts, guarantees, or competitor allegations into a new product.
Preserve image attraction: saturated complementary accents, subject/background contrast,
layered space, tactile materials, appealing silhouettes and memorable environments. Attach
actual local visual references with source/rights, hashes, palette, local cue and a precise
usage role. Put these images into generation refs; prose alone is not a reference bank.
The original Ukrainian moodboard demonstrates
teal/coral street contrast, emerald/amber tea, cobalt/orange care and a plum/green courtyard.
It conditions setting/color, not cast identity or product application. Do not reproduce its grid.
For attention, annotate every retained second with a stimulus, meaningful change, small
payoff and next question. Keep a ledger of open loops and handoffs. Flag unexplained holds
and visual repetition; let a deliberate suspense hold pass with an explicit acting reason.
Do not confuse frequent edits with meaning or promise zero drop-off without audience data.
15. Generation handoff
Lock source analysis → claim ledger → adapted script → rendered and reviewed audio → word/phrase timeline → character and concern references → shot plan → video generation → edit → captions → validation. Every generated item stores prompt, model, tier, references/hashes, task ID, receipt, output path, retained range, review and dependency. A changed line invalidates its alignment and dependent shots, not the entire paid run. Poll existing task IDs before resubmitting.
16. Quality gates and score anchors
Use pass / revise / unknown with concrete evidence: story causality, preserved hook function, sustained questions, clear mechanism, claim traceability, role distinction, Ukrainian diction, identity continuity, concern continuity, picture/speech relationship, mobile legibility, complete CTA, no music, correct video model/tier, and actual playable output. Subjective 1–5 craft scores require anchors: 1 absent/broken; 2 inconsistent; 3 intelligible; 4 clear and deliberate; 5 unusually coherent with multiple observed receipts. Do not create a single “DNA score” pretending to predict performance.
17. Experiment plan and limitations
Specify one-variable ablations, hypotheses, primary metrics, guardrails, audience and placement matching, random assignment, spend/sample allocation and decision rules before a test. Useful comparisons include caregiver quest vs direct explanation; early vs delayed reveal; dialogue vs single narrator; technical vs simple mechanism; relational vs appearance-only payoff. Track attention, qualified traffic, conversion and adverse feedback separately. Never invent an expected percentage gain from one reference.
18. Report packaging
Provide an executive reading, watchable reference, synchronized timeline, expandable evidence, character/concern passports, emotional and belief maps, source-to-adaptation table, claim ledger, full script, audio master, generated film, editable project, provider receipts, test plan and explicit unknowns. Machine-readable JSON should describe the same evidence as the readable report. Link timestamps to the actual local asset; keep temporary provider URLs distinct from durable deliverables.