Free Chinese AI Workflow: Create Your First Short Video

AI content creationChinese AIfree AI toolsshort videocreator workflow

A beginner workflow using Chinese AI tools Kimi, DeepSeek, Doubao, Jimeng, and Jianying for research, scripting, visuals, editing, review, and publishing.

Cover image for the Chinese AI free edition
Cover image for the Chinese AI free edition

For people trying AI-assisted creation for the first time, with zero budget, who want to get their first piece of content out the door. This article does not chase "full automation." Instead, it builds a workflow that is repeatable, checkable, and can actually be finished.

Introduction: What You Lack Is Usually Not Tools, but a Process You Can Follow to the End

First-time creators usually blame their struggles on "I can't write," "I can't edit," or "I have no eye for shots." Only after actually starting do they discover the bigger problem: the work is chopped into fragments. Today you hunt for topics, tomorrow you write the script, the day after you make images; by the time you finally open the editing software, you find the footage aspect ratios don't match, the voiceover runs too long, and the title says something different from the body.

AI can ease these problems, but only if you stop treating it as a black box that hands in finished homework. The safer approach is to break creation into small steps and let each tool handle only what it is good at: one tool organizes research, one writes the draft, one makes the language sound more like real speech, one produces visual assets, and finally an editing tool assembles everything.

This article runs one concrete case study from start to finish: producing a 60–90 second knowledge short video, "Why AI Is Good at Organizing Meeting Notes." You will end up with a topic card, a voiceover script ready to record, a set of image prompts, a shot list, and a pre-publish review checklist.

Chinese AI free workflow
Chinese AI free workflow
graph LR
    A[Account positioning] --> B[Build topic pool]
    B --> C[Compile research cards]
    C --> D[Generate voiceover script]
    D --> E[Generate images and shots]
    E --> F[Editing and subtitles]
    F --> G[Human review]
    G --> H[Publish and retrospective]

1. Start With the Minimal Tool Stack

The free route does not require installing every popular product. The more tools you use, the higher the switching cost, and the easier it is for accounts, assets, and versions to descend into chaos. The combination below is already enough to finish your first video.

Division of labor among free Chinese AI tools
Division of labor among free Chinese AI tools
StageRecommended toolRole in this articleSubstitution principle
Research organizationKimiConsolidate material, extract key points, generate question listsAnything that can read long text and organize structure
Script writingDeepSeekGenerate outlines, voiceover scripts, and alternate versionsAnything with stable Chinese writing
Conversational rewritingDoubaoShorten sentences, simulate dialogue, check for awkward phrasingAny Chinese chat model works
Images and short shotsJimengText-to-image, image-to-video, local adjustmentsUse whatever free quota your account has
Editing and subtitlesJianyingAssemble footage, subtitles, voiceover, music, and exportMobile or desktop both work

Kimi's official page positions it as an AI assistant for knowledge work, documents, and creative tasks; DeepSeek offers a chat product and public API documentation; Doubao's official capability list includes Q&A, writing, translation, and creative assistance; Jimeng supports text-to-image, text-to-video, and image-to-video; Jianying (known internationally as CapCut) puts AI video generation, image design, voiceover, and multi-track editing into a single creation workflow. Specific free quotas, queue speeds, and available models may change, so this article teaches methods only and does not treat any quota as a fixed promise.

Many people's first message to an AI is: "What's trending today?" That question is far too broad. The model may return a dozen topics that look lively but have nothing to do with your account. A better approach is to first write an account positioning card that constrains your content scope.

Comparison of a bad question versus a good question
Comparison of a bad question versus a good question

Copy the following to any text model and replace the information in brackets:

You are my content planning assistant. Do not generate topics yet. First, use the information below to compile an account positioning card.

Account theme: How ordinary people can use AI to work more efficiently
Primary audience: Working professionals aged 22–40
Content format: 60–90 second vertical voiceover or image-and-text videos
Style: clear, restrained, minimal jargon, no exaggerated claims
Content we do not make: unverified news, income promises, complex programming tutorials
Publishing cadence: 3 pieces per week

Please output:
1. A one-sentence positioning statement
2. Three recurring content segments
3. Topic boundaries for each segment
4. What viewers should gain after watching
5. Five categories of content we are likely to drift into by mistake

A usable positioning result should be sufficiently narrow. For example:

In 90 seconds, explain one AI work method an ordinary person can try the same day, along with its use cases, limitations, and verification steps.

This sentence becomes the "master constraint" for every prompt that follows. No matter how many topics you gather later, check each one against this sentence first.

Account positioning card
Account positioning card

3. Build a Topic Pool That Does Not Depend on Inspiration

A topic pool is not a warehouse of titles; it is a continuously updated task table. Record at least the following fields:

FieldExample
Audience questionMeetings produce too much content — how do I organize it quickly?
Core conclusionAI is good at first drafts and classification, not at replacing fact-checking
Evidence sourcesProduct documentation, my own demos, public tutorials
Content format60-second voiceover plus interface demo
Risk pointsDo not upload sensitive meeting content; re-verify names and figures
StatusTo research / Drafting / Published / Awaiting retrospective
Topic pool example
Topic pool example

When asking AI to generate topics, require it to center on "problems" rather than "tool names." Viewers usually do not come for a specific model; they come to solve a problem.

Based on the account positioning below, generate 20 topics that can be used long-term, without chasing the day's trends.

Positioning: In 90 seconds, explain one AI work method an ordinary person can try the same day, along with its use cases, limitations, and verification steps.

Requirements:
- Every topic starts from a real problem
- Mark whether it suits voiceover, screen recording, or image-and-text
- Give a one-sentence core conclusion
- List the facts that need verification
- Exclude exaggerated income claims, AI worship, and anything that cannot be demonstrated
- Group into "easy to produce / needs preparation / not recommended yet"

Pick topics from the "easy to produce" group. This case study chooses: Why AI Is Good at Organizing Meeting Notes. It has a clear problem, a demonstrable process, and a privacy boundary worth flagging.

4. Build a Research Card First, Then Let AI Write the Script

Asking a model directly to "write a viral script" easily yields text that is correct but empty. The fix is to build a research card first, separating facts, opinions, and limitations.

Put your existing articles, product documentation, or your own operation logs into Kimi with the prompt below:

Using only the material I provide, compile a research card for a short video. Do not add facts beyond the material.

Topic: Why AI is good at organizing meeting notes

Split it into six parts:
1. The problem viewers are currently facing
2. Facts that can be confirmed
3. Work suited to AI
4. Work that must be checked by a human
5. Privacy and usage boundaries
6. Three steps usable for a demo

After every fact, mark which source it came from; put anything unconfirmable into "to verify" — do not guess.
Kimi research card output (illustrative)
Kimi research card output (illustrative)

A solid research card might look like this:

  • Problem: Meetings are information-dense; manually replaying and categorizing them takes time.
  • Suited to AI: Classifying the transcript by agenda item, extracting action items, generating a first-draft summary.
  • Must be human-checked: Owners, deadlines, numbers, negations, and final decisions.
  • Privacy boundary: When company secrets, client information, or personal sensitive data are involved, follow your organization's data rules and do not upload unauthorized raw material to external services.
  • Demo steps: Prepare redacted text; generate the summary and action items; check each item against the original.

The most important thing here is not getting more content from the AI, but making it state clearly "what cannot be confirmed." Trustworthy knowledge content often comes from spelling out the boundaries.

Facts, opinions, and items to verify
Facts, opinions, and items to verify

5. Use Three-Pass Writing to Get a Natural Voiceover Script

Three-pass writing process
Three-pass writing process

Pass One: Structure Only

Based on the research card, design the structure for a 60–90 second short video.

Requirements:
- The first 3 seconds state a concrete pain point without manufacturing anxiety
- The body makes only one core conclusion
- Demonstrate with a "three-step method"
- Must include one limitation reminder
- End with an actionable step, not an empty slogan
- Output only the structure and the goal of each segment for now; do not write the full copy

A reasonable structure is usually: pain point 8 seconds, conclusion 10 seconds, three-step demo 45 seconds, boundary reminder 12 seconds, closing 5 seconds.

Pass Two: Expand Into a Voiceover Script

Expand the structure above into a voiceover script.

Writing rules:
- Keep each sentence within about 22 Chinese characters (roughly 12–15 English words)
- Prefer concrete verbs; avoid abstract words like "empower," "disrupt," or "revolutionary"
- Do not claim AI can guarantee accuracy
- The first time a technical term appears, explain it in plain language
- Keep the total length within 260–330 Chinese characters (roughly 150–220 English words)
- Mark places that need on-screen visuals with [VISUAL]
Copyable script prompt card
Copyable script prompt card

Pass Three: Use Doubao for a "Read-Aloud Check"

Rewrite the voiceover script below so it sounds more like a real person speaking.

Requirements:
- Keep the facts and structure
- Remove repetition, boilerplate, and formal written connectors
- Break up any run of three or more long sentences
- Do not use hyperbole
- Flag the three sentences most likely to trip up a speaker, and provide replacements

Here is a finished draft ready to record:

After an hour-long meeting, the most tiring part is often not the discussion but the write-up. Who owns what, when it's due, and which items were only suggestions — all mixed into dozens of pages of notes, easy to miss. What AI is good at is not deciding for you, but organizing the mess into a first draft.

>

Step one: prepare a meeting transcript that has already been redacted. [VISUAL: removing names, client information, and internal data]

>

Step two: have the AI classify it into "conclusions, action items, owners, deadlines, open questions." [VISUAL: a five-column table appearing row by row]

>

Step three: check against the original. Numbers, names, dates, negations, and final decisions — skip none of them. [VISUAL: original and summary side by side]

>

What you save this way is organizing time, not the responsibility of judgment. When sensitive material is involved, follow your company's rules. Try it once with a fictional transcript and see whether the AI's first draft actually reduces your repetitive work.

This script contains no promises like "shocking" or "10x your productivity," but it delivers complete actions and states the limits — which makes it far better at building long-term trust.

Key sentence from the finished script
Key sentence from the finished script

6. Titles and Covers: Promise Only What the Body Actually Delivers

A title's job is to get the right people to click, not to trick everyone into clicking. Have the AI produce versions of different intensity, then choose yourself.

Based on this voiceover script, generate 12 short-video titles.

In three groups:
A. Problem-type: state the viewer's difficulty directly
B. Method-type: state what they will learn
C. Counterintuitive-type: point out what AI should NOT do

Constraints:
- No more than 20 Chinese characters (roughly 8–10 English words)
- Do not use words like "must-watch," "legendary," "insane," "kills it," or "guaranteed profit"
- Every title's promise must be redeemable within the 90-second body
- Attach one short cover line per title, no more than 10 Chinese characters (roughly 5 English words)

Candidate titles:

  • "Don't Let AI Hand In Your Meeting Notes Unreviewed"
  • "Organize Meetings With AI in Just Three Steps"
  • "AI Can Write Minutes, but These Five Things Need a Human Check"
Three title groups method
Three title groups method

Keep only one layer of information on the cover, such as "Three-Step Meeting Notes" or "Five Things to Re-Check." Do not cram the full title, tool logos, and ten selling points into one image.

Cover subtraction example
Cover subtraction example

7. Break the Voiceover Script Into 8 Shots

The hard part of image generation is not "writing a very long prompt" but keeping the visual logic consistent across shots. Have the AI output a shot list first, then move into Jimeng to generate the assets.

ShotDurationVisual taskAsset method
14 sMeeting just ended, desk piled with notesAI image or live footage
26 sPages of text scrolling rapidlyScreen recording or motion mockup
38 sCore conclusion: "AI drafts first"Subtitle card
410 sRedacting the textScreen recording demo
512 sFive-column table being generatedScreen recording demo
612 sOriginal and summary checked side by sideSplit screen
710 sPrivacy and human-check reminderInformation card
86 sThree-step process recapEnd card
Eight-shot table
Eight-shot table

Example Jimeng prompt for the opening shot:

Modern office scene just after a meeting ends, desk with a laptop, sticky notes, and meeting notes, only hands visible, no clear human faces, clean and restrained composition, realistic photography style, soft natural light, blue-gray tones, vertical 9:16, title space reserved at the top, no text, no brand logos
Jimeng generation interface (illustrative)
Jimeng generation interface (illustrative)

Example prompt for a process card:

Minimalist infographic, three-step process: redact notes, AI classification, human verification, white background, dark blue lines, consistent rounded cards, simple icons, vertical 9:16, space reserved for text layout, do not generate any text in the image

When generating visuals, fix five variables first: subject, scene, lighting, composition, and aspect ratio. Do not change the style words on every image, or the same video will look like it was stitched together by five different accounts. Any Chinese text in your visuals is best added later in Jianying rather than relying on the image model to render it directly.

Consistent-style nine-image grid
Consistent-style nine-image grid

8. Do Only Four Things While Editing

Beginners burn the most time inside editing software: trying dozens of transitions, constantly changing fonts, adding effects to every line. Knowledge videos need clarity, not spectacle.

  1. Lay down the voiceover or narration first. Let the audio determine total duration and pauses.
  2. Then add the supporting visuals. When the script says "redact," show redaction; when it says "verify," show the comparison.
  3. Generate subtitles, then check them line by line. Names, product names, numbers, and punctuation are the most error-prone.
  4. Add music last. Music must not drown out the voice, and transition sound effects must not interrupt comprehension.
Jianying timeline (illustrative)
Jianying timeline (illustrative)

Jianying currently offers AI video generation, AI image design, AI voiceover, and multi-track editing as an integrated feature set, but specific features may vary by version, device, and account. For this workflow, what matters most is still basic editing, subtitle proofreading, and reliable export.

Subtitle error correction
Subtitle error correction
Efficiency change illustration
Efficiency change illustration

9. Before Publishing, Pass the "Seven-Item Human Review Gate"

Pausing for five minutes before export saves more time than deleting the post afterward.

□ Does the title accurately reflect the body?
□ Can every fact be traced to source material or a demo you ran yourself?
□ Have names, numbers, dates, and negations been verified?
□ Has sensitive information and unlicensed material been removed?
□ Do the visuals sync with the voiceover instead of being mere decoration?
□ Are there any exaggerated promises or absolute claims?
□ Will viewers know what to do next after watching?

You can also have a model perform a reverse review:

You are now a strict content review editor. Examine this script. Do not polish it for me — only point out problems.

Output in these categories:
1. Facts that cannot be proven from the source material
2. Sentences that could be misread as performance guarantees
3. Privacy, copyright, or compliance risks
4. Mismatches between the title and the body
5. Terms the audience may not understand
6. Items the author must personally confirm

Quote the original sentence for every issue and suggest the smallest possible fix.

10. Turn Publishing Data Into Input for the Next Piece

Do not stare at view counts alone. For your first batch of content, what deserves more attention is: whether the opening seconds retain people, where viewers drop off, what questions keep recurring in the comments, and whether saves exceed likes. Real questions from the comments can go straight into the next round of the topic pool.

Publish-and-retrospective loop
Publish-and-retrospective loop
graph LR
    A[Publish] --> B[Record data]
    B --> C[Collect comment questions]
    C --> D[Find drop-off points]
    D --> E[Update title/structure rules]
    E --> F[Produce the next piece]
    F --> A

Record four sentences for every piece of content:

  • What problem was this piece most trying to solve?
  • Which sentence retained viewers best?
  • Which segment made people leave most easily?
  • What single variable will you change in the next piece?

Change only one variable at a time — only the opening, or only the cover. If title, length, visuals, music, and delivery all change at once, you cannot tell which one produced the result.

11. Twelve Prompts Worth Saving

1. Topic deduplication

Compare the following 30 topics and merge those expressing the same problem. Keep the most distinct version of each and explain how it differs from the others. Keep only 10 in the end.

2. Separating facts from opinions

Split the material into four columns: verifiable facts, author opinions, reasonable inferences, unconfirmable claims. Do not write inferences as facts.

3. Finding counterexamples

For my core conclusion, give three scenarios where it does not apply. If the conclusion overgeneralizes, rewrite it into a more accurate version.

4. Three-second openings

Write 8 three-second openings for the same topic: 2 each of problem-type, result-type, wrong-demo-type, and scenario-type. Do not manufacture anxiety.

5. Voiceover compression

Without removing key facts or the limitation reminder, compress the script by 25%. Prioritize cutting repetition and abstract boilerplate.

6. Jargon translation

Find the terms in the script that an ordinary office worker may not understand. Rewrite each term into one everyday sentence without sacrificing accuracy.

7. Shot list

Break the voiceover script into 8–10 shots. For each shot write: duration, voiceover line, visual purpose, asset type, and on-screen text.

8. Image prompts

Generate image prompts from the shot list. Keep the same color palette, lighting, and compositional language across all shots, and do not generate text inside the images.

9. Subtitle line breaks

Convert the voiceover script into subtitle format. Each line 8–16 Chinese characters (roughly 4–9 English words), broken at semantic boundaries, without changing the meaning.

10. Title delivery check

Check whether the title's promise is fully delivered in the body. If not, lower the title's promise first rather than force-fitting content.

11. Risk check

Check the script for absolute claims, performance guarantees, unlicensed material, sensitive information, and omitted conditions that could mislead.

12. Comments to topics

Classify the following comments into "doesn't know how, worried about risk, wants examples, doubts the results, advanced needs," and generate the next batch of topics.

12. The Ten Most Common Beginner Questions

1. Why does AI writing look obviously AI-written at a glance?

Usually it is not a lack of model capability but a prompt with no real material and no defined audience. Provide the research card, your own viewpoints, and a list of banned expressions first, then do the conversational rewrite.

2. Are free tools enough?

Enough to run the workflow end to end, but not necessarily suited to high-frequency batch production. Free quotas, queues, and features change, so the early priority is validating your content format, not generating in bulk every day.

3. How many images does one video need?

For a 60–90 second knowledge video, 8–12 effective visuals are usually enough. An effective visual is one that explains the current sentence, not one that merely adds variety.

4. Can I skip using a real human voice entirely?

You can use the voiceover provided by the tools, but check pronunciation, pauses, licensing scope, and platform requirements. For personal-expression accounts, keeping your own voice usually builds recognizability more easily.

5. Can AI-generated content be published directly?

Not recommended. At minimum, complete five checks: facts, privacy, copyright, title delivery, and subtitles.

6. Why is the generated image style inconsistent?

Because every image changed the style words and composition. Fix the color palette, lighting, camera distance, aspect ratio, and subject description, then fine-tune the action image by image.

7. Should I publish on many platforms at once?

Pick one primary platform first and get ten pieces of content through the workflow, then adapt versions based on each platform's aspect ratios and user habits. Copying the same version to every platform usually obscures what is actually going wrong.

No. A stable account usually needs "evergreen questions plus a small amount of trending topics." With only trends and no recurring segments, content is very hard to compound.

9. How do I judge whether a topic is worth making?

It should satisfy all of these: the audience genuinely asks it, you can offer concrete actions, there is material to verify against, and it can be explained clearly within the target duration.

10. How long until I see results?

There is no fixed answer. Your first ten to thirty pieces are more like a testing period; the goal is to find stable segments, a way of speaking, and a production rhythm — not to depend on one piece suddenly getting massive views.

13. The Real Cost of This Free Workflow

Software costs can approach zero, but time does not. Your first production may take two to three hours because you are building templates. Only after your account positioning, prompts, subtitle styles, and asset directories are fixed can a single piece shrink to tens of minutes.

Store your projects in the following directory structure:

AI-media/
├── 00-account-positioning/
├── 01-topic-pool/
├── 02-sources/
├── 03-script/
├── 04-images/
├── 05-video-clips/
├── 06-edit-project/
└── 07-published-and-review/

Keep file names consistent too, for example:

2026-07-20_meeting-notes_v01_source.md
2026-07-20_meeting-notes_v03_script.md
2026-07-20_meeting-notes_cover_02.png

It looks like a mere tidiness habit, but it prevents the classic mess of "final, final2, actually-final-this-time."

Closing: Finish Your First Piece Before Talking About Automation

Free Chinese AI tools already cover research organization, Chinese writing, images, short shots, voiceover, and editing. They are best at helping you validate three things: what you plan to say, whether the audience wants to watch it, and whether you can finish consistently.

In the first stage, do not aim for unattended automation, and do not hand all the work to a single chat window. Get the process running end to end, keep human judgment in the loop, and record every revision. Once you can publish steadily and know exactly which stage eats your time, move on to the professional route in the next article: using APIs, structured data, and n8n to turn copy-paste into controlled automation.

Source Verification Note

The product capabilities mentioned in this article were checked against the following official pages. Features, free quotas, and service availability may vary by account, region, and version; refer to the current account interface and latest terms before use.