Script and Edit Video with AI: One Pipeline from Topic to Final Cut

1 viewsAI VideoVoiceover ScriptCapCutShot ListAI Task Playbook

Breaks "make a video with AI" into nine steps you can follow: settle the one-line conclusion and your sources, write a script that reads aloud cleanly, turn that script into a shot list with durations, then take it into the editor for sync, subtitle proofing and a cover. Includes copy-ready prompts, a material checklist, first-video versus practised timings, the cost structure, and the seven most common ways it goes wrong.

A video creator editing an explainer video beside a paper shot list in a small studio

A video creator editing an explainer video beside a paper shot list in a small studio

Entry 05 in the AI Task Playbook series. Scope: the full production path for explainer, voiceover and screen-recording video, from topic to publish. It does not cover on-camera performance, narrative shorts or commercials that need a real crew, and it is not a comparison of AI video generators — for that see The Definitive 2026 AI Video Generator Review.

The goal

When you are done, you should have: a three- to eight-minute explainer video ready to publish (or a 60–90 second short), with a complete script, a shot list, voiceover, proofread subtitles and a cover, with the source material filed tidily so the next one takes less time.

"The next one takes less time" is the operative phrase. How long the first video takes barely matters; what matters is whether you kept anything reusable — a script template, a shot list format, a subtitle checklist, an asset library, a cover template. Without those, video ten costs as much as video one, which is exactly why most people stop after three.

The tool stack

LayerJobRecommendedAlternatives
ResearchTopic, fact-checking, sourcingPerplexity, MetasoFelo
ScriptStructure, argument order, information densityClaudeChatGPT
Spoken voiceTurn written prose into speakable linesKimi, DeepSeekDoubao
VoiceoverNarration, one or many languagesMoyin, iFlytek ZhizuoElevenLabs, Fish Audio
VisualsCharts, diagrams, covers, a few generated imagesGaoding AI, RecraftTongyi Wanxiang, Napkin
Generated clipsA few seconds of atmosphere, transitionsKling, JimengRunway, Pika
EditingSync, auto subtitles, audio levels, exportCapCut ProDujia Creator, Clipchamp
TranscriptionTurn a recorded take into text when you speak firstTingwuiFlytek Tingjian
Topic + one-line conclusion Research and sources Voiceover scriptwritten to be spoken Read it aloudfix what stumbles Shot listvisual / duration / on-screen text Voiceover Assets: screenshots / charts / a few clips Edit: sync + subtitle proof + levels Cover and publish

Figure 1: The whole pipeline. The shot list is the hinge — it simultaneously decides how the voiceover is phrased, which assets you need, and how long each shot holds. Without it, all three decisions land on you at once inside the editor.

Why this stack

Why the weight sits on "the script carries the shot list." What actually eats time in video is not typing, it is dragging things around a timeline: what visual goes with this line, how long this beat holds, where the transition sits. If the script stage already records the visual, the duration and the on-screen text for every line, editing becomes execution, and the speed difference is several-fold. Walking into an editor with a plain text script means thinking and cutting at the same time — the slowest possible mode.

Why you must read the script aloud. Models default to written prose, and written prose is very hard to follow by ear: long qualifiers, nested clauses, stacked abstractions. It reads smoothly and listens terribly. There is exactly one test: read it out loud. Any sentence that needs two breaths gets split, and anywhere you instinctively rephrase, rewrite it the way you just said it.

Why not generate the whole video with AI. Today's video models are good at a few seconds of atmosphere, establishing shots and transitions, and bad at carrying specific information — they can't draw a correct interface, can't spell, and can't keep a subject looking the same across two shots. Filling five minutes this way is expensive, visually incoherent, and tiring to watch. The sensible use is seasoning: the opening seconds, transitions between chapters, an abstract visual when you reach an abstract idea.

Why subtitles need a human pass. Automatic subtitles are very accurate on ordinary speech and noticeably worse on proper nouns, acronyms, product names and numeric units. And a typo in a subtitle is the one thing viewers will screenshot and share.

The full steps

Step 1: Topic and the one-line conclusion

Write the one-line conclusion first: the sentence viewers take away. If you can't write it, the topic isn't clear yet, and starting the script now guarantees a pile of information with nowhere to land.

Step 2: Research and sources

Use search-based AI to assemble the key facts and record a source for each. Video is harder to annotate than text, so verification has to finish up front. For data, dates, prices and version numbers, use the official page — the model's recollection is not reliable.

Step 3: Write the voiceover script

Use the first prompt below. Put the hard requirements in the prompt: short sentences, spoken register, no written-prose connectives, a takeaway roughly every 30 seconds. Also require that the first 15 seconds give the viewer a reason to stay — not an introduction, but the problem this video solves for them.

Step 4: Read it aloud

Read the whole thing out loud with a pen in hand. This step typically cuts ten to twenty percent of the words, and everything it cuts is material nobody could have followed anyway. Note where you naturally pause; those are your shot boundaries.

Step 5: Build the shot list

Have the model convert the script into a table: number / line / visual / asset source / estimated duration / on-screen text. Classify the asset source explicitly: own screenshot, chart to be made, stock shot, clip to be generated. Only once this table exists do you know how much material to prepare.

Step 6: Record the voiceover

Two routes. Record it yourself: a phone and a headset mic are enough; closing the window, choosing a room with soft furnishings and staying a fist's distance from the mic matter more than the equipment. AI voiceover: prefer a service that lets you tune pace and pauses, and always audition how it reads numbers and unusual names. Either way, do not imitate a real person's voice.

Step 7: Prepare the assets

Work down the shot list; do not hunt for material while cutting. Keep screenshots at one resolution and charts on one palette, and hold generated clips to under a tenth of the runtime. Put everything in one folder named by shot number.

Step 8: Edit

Fixed order: lay the voiceover across the timeline → place visuals against the shot list → generate subtitles → proof the subtitles line by line → normalise levels and reduce noise → add transitions and music → watch it through once → export.

Push the music clearly below the voice. A lot of amateur feel comes from music that is simply too loud.

Step 9: Cover, headline, review

Use a fixed cover template and write a headline that names the problem solved. A few days after publishing, look at three numbers: three-second retention (a cover and opening problem), average view duration (a pacing problem), and engagement (a topic problem).

No Yes 1. One-line conclusion 2. Research and sources 3. Write the voiceover script 4. Read it aloud Reads cleanly? 5. Shot list 6. Voiceover 7. Prepare assets 8. Edit: sync / subtitles / levels 9. Cover / publish / review

Figure 2: The nine steps and the single loop. "Does it read cleanly" sits before the voiceover because changing the script after recording means redoing everything.

Prompts

1. Write the voiceover script

You are a scriptwriter for explainer videos. Turn the material below into a
{duration} voiceover script.

Hard requirements:
1. This is written to be spoken and heard, not read. Average sentence under
   20 words, one idea per sentence, no written-prose connectives like
   "furthermore", "in conclusion" or "it is worth noting".
2. The first 15 seconds must answer "why should I keep watching" with a concrete
   question or a counterintuitive fact. No self-introduction, no "today we're
   going to talk about".
3. Give a takeaway roughly every 30 seconds.
4. The whole script serves one conclusion: {one-line conclusion}.
5. For numbers, dates and names, use only the material I supply. Do not add
   anything from your own memory. Where the material is silent, write "{to verify}".
6. End on one concrete action. No uplift, no "please subscribe".

Material:
{paste your verified facts and sources}

2. Convert the script into a shot list

Convert the voiceover script below into a shot list as a Markdown table with columns:
number / line / visual / asset source / estimated seconds / on-screen text

Rules:
1. Break at natural sense boundaries; one line of script per visual, no cell
   longer than 12 seconds.
2. "Asset source" must be one of four values: own screenshot / chart to be made /
   stock shot / clip to be generated. Generated clips must total under 10% of runtime.
3. "Visual" must be specific enough to execute, e.g. "settings page screenshot,
   third item on the left highlighted". Never "relevant footage".
4. "On-screen text" holds only key words to burn into the frame, under 10 words
   per cell; leave it empty where none is needed.
5. Estimate duration at a normal narration pace and give the total runtime at the end.

Script:
{paste the read-aloud-corrected script}

3. Generate the asset preparation list

From the shot list below, produce an asset preparation list grouped by type so I can
gather everything in one pass.

Four groups:
1. Screenshots or footage I must capture myself: exactly what screen, which state,
   what to highlight.
2. Charts to be made: chart type, the one-line conclusion it must convey, the data needed.
3. Stock shots: search keywords, two or three alternatives each.
4. Clips to be generated: a visual description and a duration. Also tell me whether a
   still image with a slow push or pan could replace each one — if it can, prefer that.

Finish with counts per group and an estimated preparation time.

Shot list:
{paste the shot list}

4. Subtitle proofing checklist

Below is the auto-generated subtitle text. Find everything that needs human
confirmation and output a list.

Focus on:
1. Proper nouns, product names, company names and acronyms rendered wrongly or as
   homophones.
2. Numbers, units, dates and version numbers that disagree with my script.
3. Line breaks that create ambiguity, and long lines that should be split in two.
4. Filler words and repetitions that should be removed.
5. Whole lines missing or added compared with the original script.

Output: timecode / subtitle text / problem type / suggested correction.
Do not polish lines that have no problem.

Original script:
{paste the script}

Auto subtitles:
{paste the exported subtitle text}

Worked example: turning a voiceover script into a shot list you can execute

Test setup: macOS with the Claude Code CLI 2.1.226, --model sonnet, August 2026. The script was written for this demonstration; the output was actually run.

Before you start: install the CLI

The demos use the Claude Code command line. Three commands to install and verify:

# 1. Node.js 22 or newer is required
node --version

# 2. Install globally
npm install -g @anthropic-ai/claude-code

# 3. Verify
claude --version        # should print something like 2.1.226 (Claude Code)
claude doctor           # checks that the installation is healthy

The first run needs a login: type claude in a terminal to open the interactive interface and follow the prompts to authorise your account (claude auth manages the login state afterwards). Once that is done, the one-shot -p calls below run directly.

Every command here passes --model sonnet so that you reproduce against the same model tier used for these runs; without it you get your account's default model and the output will differ.

Step one: create the material (copy and paste)

A short voiceover script on why prompts stop working, how to notice, and how to fix them — already read aloud and split per step 4:

mkdir -p ~/demo/video && cd ~/demo/video

cat > script.txt <<'EOF'
Have you noticed that the same prompt works today and stops working two weeks later?
That is not your imagination.
The model gets updated. Your prompt stays where it was.
Three things today: why prompts break, how to notice, and how to fix them.
First, why they break. A prompt leans on the model's default behaviour, and an update
changes the defaults.
Maybe you used the word "brief" to hold the output short. The new version may not
respect that.
Second, how to notice. You need a fixed set of test questions you rerun after every
model update.
Five questions is enough, but it has to be the same five every time.
Third, how to fix. Do not rewrite the whole thing. Find which single constraint stopped
working.
Replace vague adjectives with checkable conditions: swap "brief" for "no more than
three sentences".
Next time the model updates, rerun those five questions and you will know whether to
touch anything.
EOF

cat > shotlist-prompt.txt <<'EOF'
Convert the voiceover script below into a shot list as a Markdown table with columns:
number / line / visual / asset source / estimated seconds / on-screen text

Rules:
1. Break at natural sense boundaries; one line of script per visual, no cell longer
   than 12 seconds.
2. "Asset source" must be one of four values: own screenshot / chart to be made /
   stock shot / clip to be generated. Generated clips must total under 10% of runtime.
3. "Visual" must be specific enough to execute. Never "relevant footage".
4. "On-screen text" holds only key words, under 10 words per cell; empty where none
   is needed.
5. Estimate duration at a normal narration pace and give total runtime plus a count
   per asset source at the end.
Output the table and the counts only.
EOF

Step two: run it

{ cat shotlist-prompt.txt; echo; echo "Script:"; cat script.txt; } > input.txt
claude --model sonnet -p "$(cat input.txt)"

Terminal window: entering the demo/video directory, joining the shot-list prompt and script into input.txt, calling claude, and printing a shot table with line, visual, asset source, duration and on-screen text, ending with asset counts

Terminal window: entering the demo/video directory, joining the shot-list prompt and script into input.txt, calling claude, and printing a shot table with line, visual, asset source, duration and on-screen text, ending with asset counts

Figure 3: The actual terminal output of the commands in this section. The counts at the end are the most practical part — they turn "prepare assets" into a task with a finish line.

Step three: read the result

The shot list came back with 17 rows. Four of them:

| # | Line | Visual | Asset source | Est. seconds | On-screen text |
| 4 | Your prompt stays where it was. | Screenshot of the original saved prompt text, unchanged, in a text editor | own screenshot | 2 | Prompt: unchanged |
| 8 | Maybe you used the word "brief" to hold the output short. | Screenshot of prompt text with the word "brief" highlighted | own screenshot | 4 | "brief" |
| 12 | Five questions is enough, but it has to be the same five every time. | Timeline graphic showing the same 5-question list repeated at three update dates | chart to be made | 6 | Same 5, every time |
| 16 | Replace vague adjectives with checkable conditions: swap "brief" for "no more than three sentences". | Before/after screenshot: "brief" crossed out, replaced with "no more than three sentences" | own screenshot | 6 | brief → ≤3 sentences |

Total runtime: ≈ 66 seconds
Asset source counts: own screenshot 9 · chart to be made 7 · stock shot 1
clip to be generated: 0 (0% of runtime)

Three reasons this table is directly usable.

The visuals are specific enough to execute. Row 16 is not "relevant footage," it is "before/after screenshot: 'brief' crossed out, replaced with 'no more than three sentences'." You can go and capture that without thinking again. The prompt line forbidding "relevant footage" is what produced this.

Asset needs became a countable quantity. Nine screenshots, seven charts, one stock shot — that is the entire list of things to prepare tonight. Without the table, "prepare assets" is an activity with no end; with it, it is a list of 17 items.

Zero generated clips. The prompt capped generated clips at 10% of runtime and it scheduled none at all, because this video is entirely concrete information with nowhere that needs atmosphere. That is the cost point from the section above in practice: hold this line and the AI-video part of the bill essentially disappears.

One caveat: the durations are estimates, not measurements. It calculated 66 seconds from a nominal narration pace; the recorded voiceover will differ, and the voice track is what you sync against. The duration column exists to show whether the piece is unbalanced, not to cut against.

What you have to supply

  • The one-line conclusion, and who this video is for
  • A verified fact and source list (video is a poor place to annotate, so verify beforehand)
  • Your own opinion or experience — the only thing separating a video from a literature review
  • Your own assets: product screenshots, screen recordings, photos
  • Recording conditions: a quiet room and a headset mic, or a usable voiceover service
  • Brand visuals: cover template, subtitle style, opening and closing frames
  • A pair of headphones for checking, and one pass watching the final cut on a phone

What you end up with

  • A publishable cut where visuals, subtitles and levels all line up
  • A voiceover script and a shot list, both reusable as templates
  • A proofread subtitle file, useful for repurposing into text
  • An asset folder organised by shot number
  • A cover template and a headline pattern
  • A review record: three-second retention, average view duration, engagement

Time

StageFirst video (5-minute explainer)Practised60–90 second short
Topic and conclusion30–45 minutes15 minutes10 minutes
Research and verification60–90 minutes40 minutes20 minutes
Writing the script60–90 minutes30 minutes15 minutes
Reading aloud and fixing20–30 minutes15 minutes5 minutes
Shot list30–45 minutes15 minutes10 minutes
Voiceover30–60 minutes20 minutes10 minutes
Preparing assets60–120 minutes40 minutes20 minutes
Editing and subtitle proofing120–180 minutes60–90 minutes30 minutes
Cover and publishing30 minutes10 minutes10 minutes
Totalroughly 6–10 hoursroughly 3–4 hoursroughly 1.5–2 hours

The two big blocks are preparing assets and editing, and both scale almost entirely with how specific the shot list is. Specific means execution; vague means improvising on the timeline.

Cost

  • A zero-cost route works: record yourself, use the free tier of an editor and its bundled stock, make the cover yourself. Total spend zero; the trade-off is that voice quality depends on your room and your throat.
  • Voiceover is the main variable: AI narration is usually billed per character or per minute, with monthly tiers available. A five-minute explainer is not many characters, so per-video cost is low — but multiple language versions or frequent re-recording after script changes raise it noticeably.
  • Editing and stock: an editor subscription mostly buys premium stock, watermark-free export and higher resolutions, billed monthly and genuinely optional. Commercial use of stock and music needs separate confirmation; free does not mean commercially licensed.
  • Generated clips are the most expensive line: billed per second or per generation, and you often generate several before one is usable. Holding them under a tenth of the runtime makes the cost negligible; trying to fill the whole video with them moves it up an order of magnitude.
  • Easy to forget: cloud storage (video eats space fast), font licences, and commercial images used on covers.

How it fails

Written-prose scripttiring to hear, views collapse Stop: read it aloudsplit anything needing two breaths Proper nouns wrong in subtitlesscreenshotted and shared Stop: proof line by line againstthe script, focus on names and numbers Generated clips overusedincoherent and expensive Stop: cap clips at 10%replace with a slow push on a still Levels jump aroundmusic buries the voice Stop: normalise loudnessmusic clearly under the voice Unclear asset sourcingmusic and footage infringement Stop: own or clearly licensed onlyfree is not commercial-use First 15 seconds is an introductionviewers leave immediately Stop: open with a problem ora counterintuitive fact Three weeks on one videonever want to make another Stop: fix runtime and asset caps firstraise standards once templated

Figure 4: Seven ways this goes wrong. The first and sixth decide whether anyone finishes watching, the second through fifth decide whether it feels professional, and the seventh decides whether there is a second video.

1. A written-prose script. The most insidious and the most damaging. The text reads perfectly and viewers cannot follow it, so average view duration is short and you cannot see why. Reading it aloud fixes ninety percent of this.

2. Subtitle errors. Automatic subtitles turn product names into homophones, spell out numbers oddly and break acronyms apart. Individually minor, collectively the one thing viewers screenshot. Proofing a five-minute video against the script takes about ten minutes.

3. Overusing generated clips. The symptom is a visual style that shifts between shots, a subject whose face changes, and a stack of unusable generations you paid for. Treat them as seasoning, not the meal.

4. Inconsistent audio. Especially obvious when your own recording is mixed with AI narration. Normalise every track to a similar loudness before export and push music clearly below the voice.

5. Infringing assets. Background music, film clips and sports footage are the three high-risk categories. Even material labelled "free" in a stock library may exclude commercial use — read the licence before you place it.

6. A wasted opening. Fifteen seconds of "hi everyone, today we're going to talk about" and the audience has already gone. Open with the problem, the conflict, or a fact that contradicts expectations.

7. Making the first video too heavy. Spend three weeks polishing one video and the feeling afterwards is usually "never again." The goal of the first one is to get the pipeline working and leave templates behind, not to produce your best work. Fix the runtime and cap the asset count first, then raise the bar.

Alternatives

If you're a poor writer but a good talker: record ten to twenty minutes of free speech first, transcribe it with Tingwu, and have AI shape it into a script and shot list. For many people this is much faster, and the register is naturally your own.

If the content is a how-to: record your screen and narrate as you go. The recording is its own shot list; post-production is only trimming pauses, adding subtitles and highlighting. That can halve the times in the table above.

If you just want to turn an article into video: use a text-to-video tool and keep expectations calibrated — that output works as supplementary material, not as your main content.

If editing is your bottleneck: get the script and shot list right and outsource the edit. The more specific the shot list, the cheaper the handoff — which is a side benefit of this whole process.


Other tasks in this series: Run a Content Channel with AI, Read and Organize PDFs with AI, Build a Knowledge Base with AI.