IA12 MIN

Moonshot's Kimi K3 rattled Silicon Valley—but its open-weight promise is still due July 27

China's 2.8-trillion-parameter model can handle a million-token context and has impressed coding evaluators. It is already reviving DeepSeek-era anxiety in the US, even though the weights, licence and technical report have not arrived.

Official Kimi K3 launch artwork showing a large ink-drawn infinity symbol across a worktable
Image: Moonshot AI / Kimi
01

Silicon Valley has another date circled in red

Moonshot AI introduced Kimi K3 on July 16, prompting a familiar chain reaction: headlines declaring the end of America's AI lead, a technology-stock sell-off and comparisons with DeepSeek's market shock. The model is real and available through Kimi's web app, desktop product, coding agent and API.

K3 has 2.8 trillion total parameters, native image input, a one-million-token context window and a focus on long-running coding and knowledge work. It debuted at the top of Arena's front-end coding ranking. Moonshot also says its overall performance still trails Claude Fable 5 and GPT-5.6 Sol—a useful counterweight to the victory-lap headlines.

Full model weights are promised for July 27. There is no K3 repository, final licence or technical report today. K3 is a working commercial service with a future open-weight commitment, not yet a model anyone can download and inspect on independent hardware.

02

A 2.8-trillion-parameter model does not use everything at once

K3 is a mixture-of-experts model. A router selects 16 of 896 expert blocks for each token instead of activating the entire network. Moonshot combines that design with Kimi Delta Attention and Attention Residuals, techniques intended to make very long contexts and very deep models more efficient. Its claimed 2.5-fold scaling-efficiency gain over Kimi K2 is a vendor measurement until independent reproduction arrives.

Quantisation-aware training uses four-bit weights and eight-bit activations. Even then, a simple four-bit calculation puts 2.8 trillion weights around 1.4 terabytes before runtime memory and overhead. Moonshot recommends supernodes with at least 64 accelerators. Open weights could give cloud providers, national labs and large companies more control; they will not turn K3 into a casual laptop download.

03

The benchmark winner changes with the benchmark

Moonshot's coding chart puts K3 first on Program Bench and SWE Marathon, second on FrontierSWE and almost level with GPT-5.6 Sol on Terminal Bench 2.1. Models use different agent harnesses—Kimi Code, Claude Code or Codex—and some rival runs encountered fallback systems or cybersecurity guards. Several tasks were also recalibrated for H20 GPUs.

Independent results support a narrower conclusion: K3 belongs near the leading group. Artificial Analysis scores it at 57, around Claude Opus 4.8 and GPT-5.5 but behind Fable 5 and GPT-5.6 Sol. On the 113-task DeepSWE leaderboard, K3 records 69% ±5% at an average $4.65 per task; GPT-5.6 Sol reaches 73% ±3% for $8.39 and Fable 5 gets 70% ±4% for $21.63. The intervals overlap.

Arena found K3 first in six of seven front-end categories and second to Fable 5 in games. That is meaningful evidence about visible web output, not a universal measure of reliability, energy use, security or production work. Claims that K3 has beaten every US model usually depend on selecting its best scoreboard.

Moonshot's official chart comparing Kimi K3 with OpenAI, Anthropic and Zhipu models across six coding evaluations
Image: Moonshot AI / Kimi; vendor-reported results and configurations
04

Cheaper than some US rivals is not the same as cheap

Moonshot charges $3 per million uncached input tokens, $15 per million output tokens and $0.30 for cached input. That can suit coding agents that repeatedly read the same repository. It is also Kimi's highest price yet. Artificial Analysis describes K3 as somewhat expensive among similarly priced models, slower than average and unusually verbose: it produced 130 million output tokens during the evaluator's index run, more than twice the average.

Per-token prices tell only part of the story. A verbose model can erase a headline discount, while self-hosting replaces an API bill with accelerators, power, networking and specialists. Buyers should compare completed work, review time and deployment cost rather than one line on a pricing page.

05

Moonshot's most useful disclosure is a warning

Moonshot says K3 can become unstable when an agent fails to preserve its reasoning history or when a user switches models mid-session. It also acknowledges 'excessive proactiveness': given ambiguity or a small obstacle, K3 may make an unexpected decision on the user's behalf.

That matters more than nationality when an agent can edit files, run terminal commands or reach company documents. Production deployments need isolation, minimal permissions, approval for consequential actions and an audit trail. Our report on GPT-5.6 and delegated work raised the same issue: a finished-looking result can hide decisions no one intended to delegate.

K3 also lacks a detailed safety card. Its behaviour across languages, memorisation, refusal boundaries, bias and cybersecurity capabilities remain under-documented. Weight access would enable outside researchers to examine some of those questions—provided the release includes workable documentation and research-friendly terms.

06

Open source asks for more than a giant download

Vendors and headlines often blur open source and open weights. The Open Source Initiative's AI definition requires freedom to use, study, modify and share, along with parameters, training code and enough information about training data to rebuild a substantially equivalent system. Weight access alone can be extremely useful without meeting that full standard.

K3 cannot be classified yet because the core artefacts are pending. Moonshot has a substantial record of releasing Kimi K2, K2.5 and developer tools, so the July 27 promise is credible enough to watch. The licence, repository, serving code and technical report will show how much control users receive.

The release has already become a prop in Washington's argument over export controls and AI rules. As our examination of US labs' shared regulatory agenda found, technical rules also decide which companies can afford to compete. K3 adds a geopolitical twist: the most adaptable option may not come from California.

07

The consequential test begins on July 27

Kimi K3 appears to be a significant leap. Independent evaluations place it beside highly capable US models for coding and agentic work, and its API pricing adds pressure near the frontier. It also shows that US chip controls have not frozen Chinese progress, although Moonshot has not disclosed K3's training hardware or cost.

That is not enough to declare America's lead erased. Leaderboards move, benchmarks cover a fraction of real work and US labs will release more models. K3 can change procurement and pricing without being the best at everything—a more tangible consequence than winning an imaginary permanent race.

On July 27, look for weights, licence terms, compatible serving code, the technical report and reproducible results outside Moonshot's API. If they arrive, researchers and providers will gain an enormous new tool in every sense. If the date slips or documentation is thin, K3 will remain a strong commercial model while its open-weight revolution waits a little longer.

SOURCES

Where the information comes from

Original announcements, documents and reporting used to prepare this article.

01Moonshot AI / Kimi — anuncio técnico, precios, evaluaciones y limitaciones de Kimi K302Moonshot AI — organización oficial y repositorios publicados03DeepSWE — clasificación independiente de tareas de programación04Artificial Analysis — evaluación, velocidad y coste de Kimi K305Open Source Initiative — definición de inteligencia artificial de código abierto06Associated Press — reacción del sector y contexto de la rivalidad entre China y EE UU07EFE — cobertura en español y contraste de evaluaciones08EL PAÍS — contexto empresarial y geopolítico de Kimi K309TechCrunch — debate tecnológico y político tras el lanzamiento
00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE