Key Facts
Kimi K3 went live on July 16: a 2.8-trillion-total-parameter MoE architecture activating only 16 of 896 experts per token (about 1.8%), a 1-million-token context window, and native support for visual input. It had previously participated in blind testing as an anonymous checkpoint on Arena.ai's Frontend Code Arena, and after being unmasked ranked first with 1679 points, above Claude Fable 5 (1631) and GPT-5.6 Sol (1618), taking the top spot in 6 of 7 frontend task categories—a jump of 17 places from Kimi K2.6's 18th. API pricing is $3/M input ($0.3/M on cache hit) and $15/M output.
"Open" Is Still One Step Short
A distinction to draw: the Arena ranking comes from third-party platform voting, while the rest of the scores—like an 88.3 on Terminal-Bench 2.1—are currently self-reported by Moonshot, with some independent release-day evaluations giving a different ordering, pending verification by the technical report to be released with the weights on July 27. The license terms haven't been announced either—just how "open" the "open weights" will ultimately be remains unknown. Separately, per Artificial Analysis's independent composite evaluation, K3's intelligence level is on par with Opus 4.8 and GPT-5.5, still below Fable 5 and GPT-5.6 Sol.
Impact Assessment
If the weights are released on schedule, this will be the first time the open source camp has put a 3-trillion-parameter-class model on the table—a new option and a new threshold for teams running their own inference: the community estimates that even a quantized version needs roughly 650GB–1TB of memory. The more practical reference value is on the coding track—what K3 topped is precisely the subject where Gemini 3.5 Pro stumbled and slipped this time, and the gap between Chinese open source models and closed-source flagships on specific capabilities is narrowing month by month.
via: Moonshot AI official release, Tom's Hardware, Arena.ai leaderboard (2026-07-16)