Tencent Zhiying

A digital human video creation tool under Tencent, supporting AI broadcasting and video synthesis.

  • Popularity
  • Audio & Video
  • Free tier
No preview available

At a glance

  • Free tierPartial
  • Chinese supportYes
  • Works in ChinaYes
Pricing

A free tier is available, offering basic digital avatars along with limited monthly generation time and export counts.

Pricing changes over time; check the official site

There is a widely acknowledged inversion in the hierarchy of disdain within the content industry: video has far greater viral power than text and images, but its production cost is also significantly higher. Writing an article can be done by one person in an afternoon; producing a decent talking-head video requires an on-camera talent, lighting and audio setup, multiple retakes, editing, and subtitles—for organizations, this represents a continuous labor cost, and for individuals, the reluctance to show their face alone filters out a large portion of potential creators.

Digital human technology targets this wall of costs: it allows AI avatars to replace real people on camera, turning text scripts directly into broadcast videos, shifting video production from a "filming project" to a "typing task." Tencent Zhiying is one of the most representative products in this field in China.

What is Tencent Zhiying?

Tencent Zhiying (zenvideo.qq.com) is Tencent’s intelligent video creation platform. Its core capability is generating digital human broadcast videos, supplemented by a full suite of video production tools including AI voice synthesis, subtitle recognition, article-to-video conversion, and online editing. The entire platform is web-based; you can create videos directly in your browser without installing professional software.

It is backed by several of Tencent’s core technologies: AI speech synthesis, face generation and driving, Tencent Cloud’s video processing capabilities, and ecosystem synergy with its own distribution channels like WeChat Channels and Tencent Video. Its positioning is clear: to serve creators and organizations that need to produce video content at scale and low cost.

Core Features

Digital Human Broadcast Videos

This is the flagship feature, following a three-step process: input text script → select digital human avatar and voice → generate broadcast video.

The platform includes dozens of built-in digital human avatars covering different genders, ages, and styles (such as news anchor, approachable explainer, business professional, etc.). The avatar’s lip movements automatically sync with the voiceover, complemented by natural facial expressions and body language. You can freely add backgrounds, PPT-style graphics, subtitles, and brand logos to the screen, resulting in a complete talking-head video.

The value of this feature is most obvious in batch scenarios: if a news program needs to produce ten daily news briefs, the traditional method requires anchors and production teams to work around the clock, whereas the digital human solution only requires editors to paste in the scripts—the cost difference is an order of magnitude.

Custom Exclusive Digital Humans

Upload video footage of a real person to train an exclusive digital human that matches their appearance and voice. Afterward, all videos can feature this "digital twin" on camera, while the actual person only needs to provide the content.

For knowledge bloggers relying on personal IP or enterprises needing a fixed spokesperson image, this is the highest commercial value application of digital human technology—it transforms the non-replicable bottleneck resource of "personal on-camera presence" into an infinitely replicable digital asset. Of course, custom digital humans are a paid premium feature.

AI Voice Synthesis (TTS)

Offers multiple voice tones with adjustable speed, intonation, and emotional styles, delivering natural and fluent Chinese results. It can be used independently—for adding narration to existing videos or creating audio content—without necessarily requiring a digital human. Tencent’s accumulated expertise in speech synthesis ensures high quality in this area.

Subtitle Recognition and Script Matching

Automatically recognizes speech from uploaded videos to generate subtitles with high accuracy, saving the tedious work of manual alignment. It also supports "script matching"—provide the original script, and the system automatically aligns it to the timeline, ensuring error-free subtitles. Since subtitles are essential for video completion rates, this feature alone is worth using.

Article-to-Video

Paste a text-and-image article, and the AI will automatically extract key points, match visual assets, generate voiceovers and subtitles, and output a video. This is ideal for batch "video-ifying" existing content (such as WeChat Official Account articles or news drafts), achieving two goals with one effort. The generated results are somewhat template-driven, suitable for informational content but not for scenarios requiring high creativity.

Online Editing and Multi-Aspect Ratio Output

Lightweight editing capabilities on the web: cropping, splicing, transitions, stickers, and background music. It also offers one-click export in both landscape (16:9) and portrait (9:16) formats, adapting to the specifications of different platforms like WeChat Channels, Douyin, and Bilibili.

Comparison with Similar Products

vs HeyGen: HeyGen is the benchmark for digital human videos in the international market, leading the industry in avatar naturalness and boasting stunning video translation (language switching + lip-syncing) features; however, it requires overseas network access, USD subscriptions, and has average optimization for Chinese scenarios. For domestic users creating Chinese content, Zhiying is more practical in terms of accessibility, Chinese voice quality, and pricing; HeyGen is stronger for those pursuing ultimate avatar effects and multi-language global expansion.

vs iFlytek Zhizuo: The closest domestic competitor in positioning. iFlytek’s strength lies in speech technology, with a slight advantage in TTS tone richness and naturalness; Zhiying has a more complete video production toolchain (editing, subtitles, article-to-video) and benefits from Tencent’s distribution ecosystem. Both offer free quotas for trial; choose based on your primary needs.

vs Dujia Creative Tool (Baidu): Baidu’s similar layout, deeply integrated with the Baijiahao ecosystem. If you operate content on a specific platform, prioritizing that platform’s tools offers real implicit benefits from ecosystem synergy (traffic, format adaptation).

vs Jianying (CapCut): Jianying is an editing tool for those "with assets," strong in editing experience and template ecosystems, and is also adding digital human capabilities; Zhiying is a generation tool for those "without assets," going directly from text to video. They actually represent different stages of the workflow, and many creators use both.

vs Real People On-Camera: To be honest: digital humans still cannot replace the infectiousness and trustworthiness of real people. For live commerce, personal branding, and emotionally driven content, real on-camera presence still dominates; the digital human’s home turf is informational content like news broadcasts, training courseware, and product explanations—where the audience cares about the information itself, and who reads it matters less.

Who Should Use Tencent Zhiying?

News and Government/Enterprise Organizations: For news broadcasts, government propaganda, and notice interpretations, where content is standardized and production volume requirements are high, the cost-reduction effect of digital human solutions is most significant. This is currently the most mature scenario for digital human adoption.

Corporate Training and Marketing Departments: For training courseware videos, product introductions, and internal communications, using digital humans to unify image and batch-produce content eliminates reliance on specific employees’ schedules and camera presence.

Knowledge Creators Who Don’t Want to Show Their Face: If you have valuable content but are unwilling to show your face, digital humans offer the lowest barrier path from text/image to video.

E-commerce and Local Merchants: For batch production of product explanation and store promotion videos, it is a practical choice for cost-sensitive scenarios.

WeChat Channels Ecosystem Operators: Tencent’s products have smoother toolchain synergy in WeChat Channels content production.

Limitations

The "AI feel" of digital humans still exists—expressions and movements can appear mechanical upon close inspection, especially in long videos. While audience acceptance of digital human broadcasts is good in informational scenarios, there is still a significant gap in content requiring emotional connection.

The free version has limitations on avatar selection, duration quotas, and export resolution; serious usage essentially requires payment; the cost of custom exclusive digital humans is relatively high for individual users.

Limited creative freedom: It is an efficient "standardized video assembly line," not a creative tool. Creators seeking personalized visuals may find it restrictive.

Pricing

A free tier is available, offering basic digital avatars along with limited monthly generation time and export counts. Paid memberships are billed on a monthly or annual basis, unlocking additional avatar options, longer durations, and HD exports. Enterprise-grade capabilities, such as custom digital avatars and API integration, are priced separately. Refer to the official website for current pricing details.

The best way to evaluate it is straightforward: take one of your actual scripts, use the free allowance to generate a video, and publish it to your target platform to monitor performance. The effectiveness of digital avatar content depends heavily on the content type—informational content often performs surprisingly well, allowing you to validate its potential in under half an hour.