大模型中文能力怎么测?10 类真实任务清单
中文说得流利,不等于中文用得好。本文拆解中文大模型评测该看的维度——语言地道、文化常识、格式遵循、拒绝翻译腔、长文组织等,给出 10 类可复现的真实测试任务,帮你测出模型的真实中文水平。
中文说得流利,不等于中文用得好。本文拆解中文大模型评测该看的维度——语言地道、文化常识、格式遵循、拒绝翻译腔、长文组织等,给出 10 类可复现的真实测试任务,帮你测出模型的真实中文水平。
Use a realistic research-assistant project to understand goal planning, tool calls, memory management, result validation, and production boundaries for an AI agent that goes beyond chat to execute a workflow.
从文档清洗、切分、向量化、召回、重排到引用展示,完整讲清一个企业知识库问答系统怎么做,以及如何减少幻觉和无来源回答。
A practical guide to AI-assisted development: assign roles to Claude Code, Cursor, and Codex; manage context; review code; run tests; and keep AI from damaging an existing project.
用一个待办事项和文档查询案例讲清 MCP 的作用、服务端结构、工具定义、权限边界和调试方法,帮助开发者把 AI 从聊天框接到真实系统。
A practical local AI assistant guide for individuals and small teams, covering model selection, hardware, document Q&A, privacy boundaries, common performance problems, and cloud-model alternatives.
拆解实时语音 AI 助手的核心链路:收音、语音识别、对话模型、工具调用、语音合成、打断处理和客服场景上线注意事项。
A practical AI image generation workflow for design, operations, and content teams, covering requirements, prompt templates, reference-image control, inpainting, and repeatable batch production of consistent visual assets.
A team-operations example showing how to build an AI office automation workflow from meeting transcription and task extraction to spreadsheet updates and weekly reports, with guidance on permissions, review, and low-code alternatives.
A practical short-video workflow covering script breakdown, storyboard prompts, image-to-video generation, voice-over and captions, editing, quality checks, and alternative tools.
面向企业知识库和问答系统的 RAG 质量教程:如何建立测试集、评估召回、检查引用、处理拒答,并用人工反馈持续改进。
Agents are no longer limited to automated clicks in demo videos. They are redefining the interface to software across enterprise workflows, personal assistants, and development tools.
Showing 12 of 12 articles