Entry 05 in the AI Task Playbook series. Scope: the full production path for explainer, voiceover and screen-recording video, from topic to publish. It does not cover on-camera performance, narrative shorts or commercials that need a real crew, and it is not a comparison of AI video generators — for that see The Definitive 2026 AI Video Generator Review.
The goal
When you are done, you should have: a three- to eight-minute explainer video ready to publish (or a 60–90 second short), with a complete script, a shot list, voiceover, proofread subtitles and a cover, with the source material filed tidily so the next one takes less time.
"The next one takes less time" is the operative phrase. How long the first video takes barely matters; what matters is whether you kept anything reusable — a script template, a shot list format, a subtitle checklist, an asset library, a cover template. Without those, video ten costs as much as video one, which is exactly why most people stop after three.
The tool stack
| Layer | Job | Recommended | Alternatives |
|---|---|---|---|
| Research | Topic, fact-checking, sourcing | Perplexity, Metaso | Felo |
| Script | Structure, argument order, information density | Claude | ChatGPT |
| Spoken voice | Turn written prose into speakable lines | Kimi, DeepSeek | Doubao |
| Voiceover | Narration, one or many languages | Moyin, iFlytek Zhizuo | ElevenLabs, Fish Audio |
| Visuals | Charts, diagrams, covers, a few generated images | Gaoding AI, Recraft | Tongyi Wanxiang, Napkin |
| Generated clips | A few seconds of atmosphere, transitions | Kling, Jimeng | Runway, Pika |
| Editing | Sync, auto subtitles, audio levels, export | CapCut Pro | Dujia Creator, Clipchamp |
| Transcription | Turn a recorded take into text when you speak first | Tingwu | iFlytek Tingjian |
Figure 1: The whole pipeline. The shot list is the hinge — it simultaneously decides how the voiceover is phrased, which assets you need, and how long each shot holds. Without it, all three decisions land on you at once inside the editor.
Why this stack
Why the weight sits on "the script carries the shot list." What actually eats time in video is not typing, it is dragging things around a timeline: what visual goes with this line, how long this beat holds, where the transition sits. If the script stage already records the visual, the duration and the on-screen text for every line, editing becomes execution, and the speed difference is several-fold. Walking into an editor with a plain text script means thinking and cutting at the same time — the slowest possible mode.
Why you must read the script aloud. Models default to written prose, and written prose is very hard to follow by ear: long qualifiers, nested clauses, stacked abstractions. It reads smoothly and listens terribly. There is exactly one test: read it out loud. Any sentence that needs two breaths gets split, and anywhere you instinctively rephrase, rewrite it the way you just said it.
Why not generate the whole video with AI. Today's video models are good at a few seconds of atmosphere, establishing shots and transitions, and bad at carrying specific information — they can't draw a correct interface, can't spell, and can't keep a subject looking the same across two shots. Filling five minutes this way is expensive, visually incoherent, and tiring to watch. The sensible use is seasoning: the opening seconds, transitions between chapters, an abstract visual when you reach an abstract idea.
Why subtitles need a human pass. Automatic subtitles are very accurate on ordinary speech and noticeably worse on proper nouns, acronyms, product names and numeric units. And a typo in a subtitle is the one thing viewers will screenshot and share.
The full steps
Step 1: Topic and the one-line conclusion
Write the one-line conclusion first: the sentence viewers take away. If you can't write it, the topic isn't clear yet, and starting the script now guarantees a pile of information with nowhere to land.
Step 2: Research and sources
Use search-based AI to assemble the key facts and record a source for each. Video is harder to annotate than text, so verification has to finish up front. For data, dates, prices and version numbers, use the official page — the model's recollection is not reliable.
Step 3: Write the voiceover script
Use the first prompt below. Put the hard requirements in the prompt: short sentences, spoken register, no written-prose connectives, a takeaway roughly every 30 seconds. Also require that the first 15 seconds give the viewer a reason to stay — not an introduction, but the problem this video solves for them.
Step 4: Read it aloud
Read the whole thing out loud with a pen in hand. This step typically cuts ten to twenty percent of the words, and everything it cuts is material nobody could have followed anyway. Note where you naturally pause; those are your shot boundaries.
Step 5: Build the shot list
Have the model convert the script into a table: number / line / visual / asset source / estimated duration / on-screen text. Classify the asset source explicitly: own screenshot, chart to be made, stock shot, clip to be generated. Only once this table exists do you know how much material to prepare.
Step 6: Record the voiceover
Two routes. Record it yourself: a phone and a headset mic are enough; closing the window, choosing a room with soft furnishings and staying a fist's distance from the mic matter more than the equipment. AI voiceover: prefer a service that lets you tune pace and pauses, and always audition how it reads numbers and unusual names. Either way, do not imitate a real person's voice.
Step 7: Prepare the assets
Work down the shot list; do not hunt for material while cutting. Keep screenshots at one resolution and charts on one palette, and hold generated clips to under a tenth of the runtime. Put everything in one folder named by shot number.
Step 8: Edit
Fixed order: lay the voiceover across the timeline → place visuals against the shot list → generate subtitles → proof the subtitles line by line → normalise levels and reduce noise → add transitions and music → watch it through once → export.
Push the music clearly below the voice. A lot of amateur feel comes from music that is simply too loud.
Step 9: Cover, headline, review
Use a fixed cover template and write a headline that names the problem solved. A few days after publishing, look at three numbers: three-second retention (a cover and opening problem), average view duration (a pacing problem), and engagement (a topic problem).
Figure 2: The nine steps and the single loop. "Does it read cleanly" sits before the voiceover because changing the script after recording means redoing everything.
Prompts
1. Write the voiceover script
You are a scriptwriter for explainer videos. Turn the material below into a
{duration} voiceover script.
Hard requirements:
1. This is written to be spoken and heard, not read. Average sentence under
20 words, one idea per sentence, no written-prose connectives like
"furthermore", "in conclusion" or "it is worth noting".
2. The first 15 seconds must answer "why should I keep watching" with a concrete
question or a counterintuitive fact. No self-introduction, no "today we're
going to talk about".
3. Give a takeaway roughly every 30 seconds.
4. The whole script serves one conclusion: {one-line conclusion}.
5. For numbers, dates and names, use only the material I supply. Do not add
anything from your own memory. Where the material is silent, write "{to verify}".
6. End on one concrete action. No uplift, no "please subscribe".
Material:
{paste your verified facts and sources}2. Convert the script into a shot list
Convert the voiceover script below into a shot list as a Markdown table with columns:
number / line / visual / asset source / estimated seconds / on-screen text
Rules:
1. Break at natural sense boundaries; one line of script per visual, no cell
longer than 12 seconds.
2. "Asset source" must be one of four values: own screenshot / chart to be made /
stock shot / clip to be generated. Generated clips must total under 10% of runtime.
3. "Visual" must be specific enough to execute, e.g. "settings page screenshot,
third item on the left highlighted". Never "relevant footage".
4. "On-screen text" holds only key words to burn into the frame, under 10 words
per cell; leave it empty where none is needed.
5. Estimate duration at a normal narration pace and give the total runtime at the end.
Script:
{paste the read-aloud-corrected script}3. Generate the asset preparation list
From the shot list below, produce an asset preparation list grouped by type so I can
gather everything in one pass.
Four groups:
1. Screenshots or footage I must capture myself: exactly what screen, which state,
what to highlight.
2. Charts to be made: chart type, the one-line conclusion it must convey, the data needed.
3. Stock shots: search keywords, two or three alternatives each.
4. Clips to be generated: a visual description and a duration. Also tell me whether a
still image with a slow push or pan could replace each one — if it can, prefer that.
Finish with counts per group and an estimated preparation time.
Shot list:
{paste the shot list}4. Subtitle proofing checklist
Below is the auto-generated subtitle text. Find everything that needs human
confirmation and output a list.
Focus on:
1. Proper nouns, product names, company names and acronyms rendered wrongly or as
homophones.
2. Numbers, units, dates and version numbers that disagree with my script.
3. Line breaks that create ambiguity, and long lines that should be split in two.
4. Filler words and repetitions that should be removed.
5. Whole lines missing or added compared with the original script.
Output: timecode / subtitle text / problem type / suggested correction.
Do not polish lines that have no problem.
Original script:
{paste the script}
Auto subtitles:
{paste the exported subtitle text}Worked example: turning a voiceover script into a shot list you can execute
Test setup: macOS with the Claude Code CLI 2.1.226,
--model sonnet, August 2026. The script was written for this demonstration; the output was actually run.
Before you start: install the CLI
The demos use the Claude Code command line. Three commands to install and verify:
# 1. Node.js 22 or newer is required
node --version
# 2. Install globally
npm install -g @anthropic-ai/claude-code
# 3. Verify
claude --version # should print something like 2.1.226 (Claude Code)
claude doctor # checks that the installation is healthyThe first run needs a login: type claude in a terminal to open the interactive interface and follow the prompts to authorise your account (claude auth manages the login state afterwards). Once that is done, the one-shot -p calls below run directly.
Every command here passes --model sonnet so that you reproduce against the same model tier used for these runs; without it you get your account's default model and the output will differ.
Step one: create the material (copy and paste)
A short voiceover script on why prompts stop working, how to notice, and how to fix them — already read aloud and split per step 4:
mkdir -p ~/demo/video && cd ~/demo/video
cat > script.txt <<'EOF'
Have you noticed that the same prompt works today and stops working two weeks later?
That is not your imagination.
The model gets updated. Your prompt stays where it was.
Three things today: why prompts break, how to notice, and how to fix them.
First, why they break. A prompt leans on the model's default behaviour, and an update
changes the defaults.
Maybe you used the word "brief" to hold the output short. The new version may not
respect that.
Second, how to notice. You need a fixed set of test questions you rerun after every
model update.
Five questions is enough, but it has to be the same five every time.
Third, how to fix. Do not rewrite the whole thing. Find which single constraint stopped
working.
Replace vague adjectives with checkable conditions: swap "brief" for "no more than
three sentences".
Next time the model updates, rerun those five questions and you will know whether to
touch anything.
EOF
cat > shotlist-prompt.txt <<'EOF'
Convert the voiceover script below into a shot list as a Markdown table with columns:
number / line / visual / asset source / estimated seconds / on-screen text
Rules:
1. Break at natural sense boundaries; one line of script per visual, no cell longer
than 12 seconds.
2. "Asset source" must be one of four values: own screenshot / chart to be made /
stock shot / clip to be generated. Generated clips must total under 10% of runtime.
3. "Visual" must be specific enough to execute. Never "relevant footage".
4. "On-screen text" holds only key words, under 10 words per cell; empty where none
is needed.
5. Estimate duration at a normal narration pace and give total runtime plus a count
per asset source at the end.
Output the table and the counts only.
EOFStep two: run it
{ cat shotlist-prompt.txt; echo; echo "Script:"; cat script.txt; } > input.txt
claude --model sonnet -p "$(cat input.txt)"Figure 3: The actual terminal output of the commands in this section. The counts at the end are the most practical part — they turn "prepare assets" into a task with a finish line.
Step three: read the result
The shot list came back with 17 rows. Four of them:
| # | Line | Visual | Asset source | Est. seconds | On-screen text |
| 4 | Your prompt stays where it was. | Screenshot of the original saved prompt text, unchanged, in a text editor | own screenshot | 2 | Prompt: unchanged |
| 8 | Maybe you used the word "brief" to hold the output short. | Screenshot of prompt text with the word "brief" highlighted | own screenshot | 4 | "brief" |
| 12 | Five questions is enough, but it has to be the same five every time. | Timeline graphic showing the same 5-question list repeated at three update dates | chart to be made | 6 | Same 5, every time |
| 16 | Replace vague adjectives with checkable conditions: swap "brief" for "no more than three sentences". | Before/after screenshot: "brief" crossed out, replaced with "no more than three sentences" | own screenshot | 6 | brief → ≤3 sentences |
Total runtime: ≈ 66 seconds
Asset source counts: own screenshot 9 · chart to be made 7 · stock shot 1
clip to be generated: 0 (0% of runtime)Three reasons this table is directly usable.
The visuals are specific enough to execute. Row 16 is not "relevant footage," it is "before/after screenshot: 'brief' crossed out, replaced with 'no more than three sentences'." You can go and capture that without thinking again. The prompt line forbidding "relevant footage" is what produced this.
Asset needs became a countable quantity. Nine screenshots, seven charts, one stock shot — that is the entire list of things to prepare tonight. Without the table, "prepare assets" is an activity with no end; with it, it is a list of 17 items.
Zero generated clips. The prompt capped generated clips at 10% of runtime and it scheduled none at all, because this video is entirely concrete information with nowhere that needs atmosphere. That is the cost point from the section above in practice: hold this line and the AI-video part of the bill essentially disappears.
One caveat: the durations are estimates, not measurements. It calculated 66 seconds from a nominal narration pace; the recorded voiceover will differ, and the voice track is what you sync against. The duration column exists to show whether the piece is unbalanced, not to cut against.
What you have to supply
- The one-line conclusion, and who this video is for
- A verified fact and source list (video is a poor place to annotate, so verify beforehand)
- Your own opinion or experience — the only thing separating a video from a literature review
- Your own assets: product screenshots, screen recordings, photos
- Recording conditions: a quiet room and a headset mic, or a usable voiceover service
- Brand visuals: cover template, subtitle style, opening and closing frames
- A pair of headphones for checking, and one pass watching the final cut on a phone
What you end up with
- A publishable cut where visuals, subtitles and levels all line up
- A voiceover script and a shot list, both reusable as templates
- A proofread subtitle file, useful for repurposing into text
- An asset folder organised by shot number
- A cover template and a headline pattern
- A review record: three-second retention, average view duration, engagement
Time
| Stage | First video (5-minute explainer) | Practised | 60–90 second short |
|---|---|---|---|
| Topic and conclusion | 30–45 minutes | 15 minutes | 10 minutes |
| Research and verification | 60–90 minutes | 40 minutes | 20 minutes |
| Writing the script | 60–90 minutes | 30 minutes | 15 minutes |
| Reading aloud and fixing | 20–30 minutes | 15 minutes | 5 minutes |
| Shot list | 30–45 minutes | 15 minutes | 10 minutes |
| Voiceover | 30–60 minutes | 20 minutes | 10 minutes |
| Preparing assets | 60–120 minutes | 40 minutes | 20 minutes |
| Editing and subtitle proofing | 120–180 minutes | 60–90 minutes | 30 minutes |
| Cover and publishing | 30 minutes | 10 minutes | 10 minutes |
| Total | roughly 6–10 hours | roughly 3–4 hours | roughly 1.5–2 hours |
The two big blocks are preparing assets and editing, and both scale almost entirely with how specific the shot list is. Specific means execution; vague means improvising on the timeline.
Cost
- A zero-cost route works: record yourself, use the free tier of an editor and its bundled stock, make the cover yourself. Total spend zero; the trade-off is that voice quality depends on your room and your throat.
- Voiceover is the main variable: AI narration is usually billed per character or per minute, with monthly tiers available. A five-minute explainer is not many characters, so per-video cost is low — but multiple language versions or frequent re-recording after script changes raise it noticeably.
- Editing and stock: an editor subscription mostly buys premium stock, watermark-free export and higher resolutions, billed monthly and genuinely optional. Commercial use of stock and music needs separate confirmation; free does not mean commercially licensed.
- Generated clips are the most expensive line: billed per second or per generation, and you often generate several before one is usable. Holding them under a tenth of the runtime makes the cost negligible; trying to fill the whole video with them moves it up an order of magnitude.
- Easy to forget: cloud storage (video eats space fast), font licences, and commercial images used on covers.
How it fails
Figure 4: Seven ways this goes wrong. The first and sixth decide whether anyone finishes watching, the second through fifth decide whether it feels professional, and the seventh decides whether there is a second video.
1. A written-prose script. The most insidious and the most damaging. The text reads perfectly and viewers cannot follow it, so average view duration is short and you cannot see why. Reading it aloud fixes ninety percent of this.
2. Subtitle errors. Automatic subtitles turn product names into homophones, spell out numbers oddly and break acronyms apart. Individually minor, collectively the one thing viewers screenshot. Proofing a five-minute video against the script takes about ten minutes.
3. Overusing generated clips. The symptom is a visual style that shifts between shots, a subject whose face changes, and a stack of unusable generations you paid for. Treat them as seasoning, not the meal.
4. Inconsistent audio. Especially obvious when your own recording is mixed with AI narration. Normalise every track to a similar loudness before export and push music clearly below the voice.
5. Infringing assets. Background music, film clips and sports footage are the three high-risk categories. Even material labelled "free" in a stock library may exclude commercial use — read the licence before you place it.
6. A wasted opening. Fifteen seconds of "hi everyone, today we're going to talk about" and the audience has already gone. Open with the problem, the conflict, or a fact that contradicts expectations.
7. Making the first video too heavy. Spend three weeks polishing one video and the feeling afterwards is usually "never again." The goal of the first one is to get the pipeline working and leave templates behind, not to produce your best work. Fix the runtime and cap the asset count first, then raise the bar.
Alternatives
If you're a poor writer but a good talker: record ten to twenty minutes of free speech first, transcribe it with Tingwu, and have AI shape it into a script and shot list. For many people this is much faster, and the register is naturally your own.
If the content is a how-to: record your screen and narrate as you go. The recording is its own shot list; post-production is only trimming pauses, adding subtitles and highlighting. That can halve the times in the table above.
If you just want to turn an article into video: use a text-to-video tool and keep expectations calibrated — that output works as supplementary material, not as your main content.
If editing is your bottleneck: get the script and shot list right and outsource the edit. The more specific the shot list, the cheaper the handoff — which is a side benefit of this whole process.
Other tasks in this series: Run a Content Channel with AI, Read and Organize PDFs with AI, Build a Knowledge Base with AI.