Table of Contents
- 1. What Kimi K3 Is — Core Specs and the Company
- 2. What "Third Place" Actually Means — Which Ranking, and When
- 3. Who Ran the Benchmarks — Separating Vendor Numbers from Third-Party Ones
- 4. Pricing: The Verdict Flips Depending on What You Compare It To
- 5. Nobody Can Download It Yet — Where the Open Weights Stand
- 6. How Far Did US Stocks Fall — The Numbers Vary by Outlet
- 7. The Distillation Allegation, and the Rebuttal
- 8. Weaknesses — Speed, Hallucination, and Moonshot's Own Assessment
- 9. How to Try It Today
- 10. Sorting Claims by Confidence
- FAQ
On July 16, 2026, China's Moonshot AI released Kimi K3. At 2.8 trillion total parameters it is billed as the largest open model ever, and within days of launch US semiconductor stocks were being sold off and a White House official was publicly accusing the company of distilling Anthropic's models.
The trouble is that the numbers, and even the labels, disagree depending on who is reporting them. This article cross-checks the benchmark publishers, the financial press, Moonshot's own announcements and the actual Hugging Face page, and states who produced each number as it goes.
57 points, third place on Artificial Analysis at launch. First on Arena.ai's Frontend Code Arena. Overall, though, it does not reach Fable 5 or GPT-5.6 Sol.
At $3 in and $15 out it is clearly cheaper than Opus 4.8 or GPT-5.6 Sol. Yet it is the most expensive of the Chinese models — three times GLM-5.2.
The release is scheduled for July 27. As of writing, Hugging Face still shows a countdown and the license is unannounced.
1. What Kimi K3 Is — Core Specs and the Company
Kimi K3 is Moonshot AI's flagship model, and the API went live on July 16, 2026 (consistent across ITmedia NEWS, OpenRouter and others). The company was founded in 2023, is backed by Alibaba and Tencent, and according to Forbes its latest funding round put its valuation above $20 billion (up from $4.3 billion and then $10 billion), with ARR of $200 million as of April 2026.
Sources: Moonshot AI's announcement (as reported by ITmedia NEWS, quoting npaka's write-up) and Northflank's hands-on article. Context length and weight size are Northflank's measured figures.
Two architectural additions carry names of their own. Kimi Delta Attention (KDA) speeds up decoding at million-token context lengths; AINews reports it as up to 6.3x faster. And Attention Residuals (AttnRes) selectively pulls the representations it needs out of earlier layers; the same article puts the gain at roughly 25% better training efficiency for under 2% added cost. Weights are MXFP4 and activations MXFP8, and rather than bolting quantization on afterwards, training from SFT onward was run with quantization built in.
🟡 There is no official figure for active parameters. The notation "2.8T-A50B" (roughly 50 billion active) is circulating, but it is an estimate derived from the 16/896 ratio, not a number Moonshot published — Northflank explicitly notes that Moonshot has never stated the active parameter count. It is also worth remembering that neither OpenAI nor Anthropic discloses parameter counts for their own models, so "2.8 trillion, the world's largest" only holds among the models whose sizes are public.
2. What "Third Place" Actually Means — Which Ranking, and When
Saying it ranks "third, behind the latest models from OpenAI and Anthropic" is broadly correct. But without pinning down the source, the numbers stop adding up later.
The basis is the Intelligence Index from Artificial Analysis. In its July 17, 2026 write-up the firm scored Kimi K3 at 57 points, placing it third behind Claude Fable 5 and GPT-5.6 Sol, and rated its intelligence on par with Opus 4.8 and GPT-5.5. For reference, the same index puts GLM-5.2 at 51 and DeepSeek v4 Pro at 44.
⚠️ The same 57 points is reported as 3rd, 4th and 7th
Artificial Analysis's own announcement says third, Northflank's tally says fourth of 189 models, and the firm's live page shows seventh of 190. The score (57) never changed, so the gap comes from what gets included in the comparison — whether variants at different reasoning effort are counted separately — and when the tally was taken. Rankings move daily, and "third" is best read as a snapshot dated July 17, 2026.
On metrics closer to real work, the ordering is sharper. Here is Artificial Analysis's GDPval-AA v2, which scores practical tasks on an Elo scale with the human baseline set to 1,000.
| Model | GDPval-AA v2 (Elo) | Where it lands |
|---|---|---|
| Claude Fable 5 | 1760 | 1st |
| Kimi K3 | 1668 | Behind Fable 5 and Sol |
| Claude Opus 4.8 | 1600 | K3 beats it |
| GLM-5.2 | 1514 | Best of the previous Chinese generation |
| GPT-5.5 | 1494 | — |
| Kimi K2.6 (previous gen) | 1190 | K3 is 478 points higher |
Source: Artificial Analysis (the firm's article published July 17, 2026). On AA-Briefcase, its private agentic evaluation, K3 also comes second behind Fable 5 at Elo 1547, up 732 points from the previous-generation K2.6. On AutomationBench-AA it takes first place at 53%.
The number that really matters is not "third" but the size of the jump from the previous generation. From 1190 to 1668 on GDPval-AA v2, and +732 points on AA-Briefcase. What moved the market was the fact that a Chinese open model had come within striking distance of the US frontier in a single generation.
3. Who Ran the Benchmarks — Separating Vendor Numbers from Third-Party Ones
Benchmark figures carry very different weight depending on whether Moonshot measured them itself or an outside party did. ITmedia reported the two separately, and this article follows the same practice.
- Program Bench: K3 77.8% (GPT-5.6 Sol 77.6%, Fable 5 76.8%)
- SWE Marathon: K3 42.0% (second, Opus 4.8 at 40.0%)
- BrowseComp: K3 91.2% (GPT-5.6 Sol 90.4%)
Every win is narrow, and all of them come from the vendor's own testing — read them with that in mind.
- Frontend Code Arena (Arena.ai): K3 first at 1679 (Fable 5 1631, GPT-5.6 Sol 1618). It leads six of the seven subdomains, up 17 places from K2.6's 18th
- GDPval-AA v2: K3 third at 1668 (Artificial Analysis)
- FrontierSWE: Fable 5 86.6% > K3 81.2% ← K3 loses this one
- Vals Index: 74.70% (second of 38 models, per Northflank's tally)
The result worth noticing is that it trails Fable 5 by 5.4 points on FrontierSWE. Headlines about topping certain benchmarks are accurate, but it is not winning across the board. Moonshot itself concedes that it falls short of Fable 5 and GPT-5.6 Sol overall (more on that below).
🟡 The scoring method has been challenged. The author of Program Bench pointed out that Moonshot used an average implementation rate rather than the number of programs that ran to completion. AI entrepreneur Bindu Reddy has separately argued that verification on uncontaminated private evaluations such as LiveBench is needed. Narrow vendor-reported wins should be discounted accordingly.
4. Pricing: The Verdict Flips Depending on What You Compare It To
Kimi K3 costs $3.00 in and $15.00 out per million tokens ($0.30 for cached input). Because the assessment reverses completely depending on the comparison, looking at only one side leads you to the wrong conclusion.
- Kimi K3: $3 / $15
- Claude Opus 4.8: $5 / $25
- GPT-5.6 Sol: $5 / $30
About 40% cheaper on input and 40-50% cheaper on output. Frontier-class performance below Western pricing has been Kimi's consistent pitch, and on this point it holds up.
- Kimi K3: $3 / $15
- GLM-5.2: about one third of K3's cost per task
- DeepSeek V4 Pro: about one twenty-third of K3's cost per task
Simon Willison, calling it a steep increase over the company's earlier models, describes it as the most expensive model any Chinese AI lab has released.
One thing to be clear about: this is not another DeepSeek shock. What made DeepSeek in January 2025 so jarring was pricing an order of magnitude below the West. K3 is 40-50% cheaper than Western models, but in the same order of magnitude. This is not price disruption; it is performance catching up, at a modest discount.
And what you actually pay is not set by the per-token rate alone. Artificial Analysis measured the real cost of completing one task while running its Intelligence Index.
Cost per task as measured by Artificial Analysis (during its Intelligence Index run)
Source: Artificial Analysis. In its measurements Kimi K3 needed about 132 million output tokens to complete the Intelligence Index, fewer than the roughly 166 million the previous-generation K2.6 consumed. The rate is higher, but less verbose reasoning brings the total down.
Cheaper than Opus 4.8 on both list price and measured cost, and the most expensive of the Chinese models — that is the full picture on pricing. The "Chinese, therefore dirt cheap" framing does not apply, and dismissing it as expensive is equally wrong once you compare it with the West.
5. Nobody Can Download It Yet — Where the Open Weights Stand
Outlets do not even agree on what to call it. VentureBeat headlines it as the "largest open-source model ever" and SCMP as the "world's largest open-source AI model," while Reuters writes "open-weight". The latter is the technically accurate term, and here is why.
The moonshotai/Kimi-K3 repository on Hugging Face still shows a countdown, with not a single safetensors file listed. Artificial Analysis's model page also marks it proprietary. Checked on July 27, 2026 — that is the scheduled release date itself, so the situation has very likely changed by the time you read this. Follow the link above to see the current state.
As of July 17, Northflank states plainly that the license is unannounced. A Hugging Face community post likewise lists it as TBD until the weights land. The claim that it will be "Modified MIT" is secondhand inference from the K2 generation and is not settled.
What has been promised is the trained weights, not the training data or training code. The accurate name for that arrangement is open weights, which is a different thing from open source as OSI defines it.
Even once it ships, this is not something an individual can run. At 4-bit (MXFP4) the weights alone come to roughly 1.4TB, and with headroom for the KV cache and activations Northflank estimates you need something on the order of eight nodes of eight 80GB GPUs, while Moonshot recommends a supernode of 64 accelerators or more. It does not fit on a single H100, H200 or B200, a distributed cluster is mandatory, and production deployment is effectively Linux plus NVIDIA only. This is nothing like the personal setups covered in the hardware requirements for local LLMs.
That said, it still matters a great deal. Once frontier-class weights come down to earth, organizations that cannot let data leave their walls — regulated industries, government bodies — can run them in house. The mere existence of a public model on the same footing changes the bargaining power of closed APIs. The assumptions behind how to choose an open model need updating because of this release.
6. How Far Did US Stocks Fall — The Numbers Vary by Outlet
US semiconductor stocks really were sold off during launch week. But the reported percentage declines conflict between outlets, so they are given here as ranges rather than as settled facts.
| What moved | Reported move | Source and notes |
|---|---|---|
| Nasdaq | -1.5% on Friday (its worst session of the week) | Yahoo Finance |
| Taiwan market | Over -6% | Same |
| Japanese market | Closed -4% | Same |
| Semiconductor ETF (SMH) | Down more than 20% from its late-June high | Same |
| Philadelphia Semiconductor Index (SOX) | -9% to -12.5% on the week | ⚠️ The figure differs by outlet ("12.5%, worst week in 15 months", "over 9%", "about 10%, the steepest since April 2025") |
| Individual stocks (Jul 17) | AMD -5%, Intel -4%, NVIDIA -3% (all intraday, later recovered) | 24/7 Wall St. |
| Hong Kong-listed Chinese AI names | Zhipu -28.4%, MiniMax -15.6% (Jul 17) | Forbes. The Chinese names were sold off harder |
For comparison, in the DeepSeek shock of January 2025 that everyone reaches for, Yahoo Finance reported that NVIDIA alone lost roughly $590 billion of market value in a single session. NVIDIA's drop this time was 3% intraday — a completely different scale.
🟡 The decline has since been recovered
Yahoo Finance later reported that dip buyers came in and chip stocks pared their losses, and Bank of America took the calm view that this was merely a summer correction. Analysts are split, too. Bernstein's Robin Zhu framed K3 as confirmatory — one more data point in the already-known trend of Chinese labs closing the gap. Morgan Stanley's Gary Yu, by contrast, said K3 has drawn a globally favorable reception and shows that Chinese LLMs are catching up with the US leaders on model scale, performance and price alike.
A figure claiming that $470 billion of AI-related value evaporated in three days is also in circulation, but it traces back to a flash news item on a crypto outlet and could not be corroborated in the major financial press, so this article does not use it.
7. The Distillation Allegation, and the Rebuttal
On July 22-23, the story moved from technology to geopolitics. Michael Kratsios, director of the White House Office of Science and Technology Policy (OSTP), made two claims about Moonshot, as reported by CNBC, The Hill, Yahoo News and others.
That Moonshot obtained servers fitted with NVIDIA GB300s and also had access to GB300s in Thailand, possibly using them to train its own models. The GB300 is a Blackwell-generation part and cannot legally be sold into China.
That the company carried out industrial-scale covert distillation against Anthropic's Fable and even built a dedicated platform that rotated access methods to evade detection. Jacob Helberg, appearing alongside him, called it the theft of valuable American intellectual property.
Moonshot employee Randy Xian replied wryly on X that Fable went public on July 1 and K3 shipped on July 15, so we apparently trained a frontier model from scratch in 15 days — one for the record books. AI researcher Braden Hancock told TechCrunch that Fable had only been available since July 1 and that distilling that much data, training on it and shipping within two weeks is impossible.
🔴 No evidence has been made public. According to SCMP, several AI specialists called the accusation political and reckless, and noted that model outputs are not copyrightable in the first place. Reuters, citing diplomatic cables, reported that the State Department had instructed embassies back in April to brief host governments on IP infringement by Chinese AI firms, with Moonshot named among them. In other words, these accusations predate the K3 release, and no K3-specific evidence has been produced.
Nor is the US side united on this. NVIDIA publicly criticized Anthropic by name, saying of Anthropic's claims about export-control loopholes that American companies should focus on innovation and rise to the challenge rather than telling tall tales.
8. Weaknesses — Speed, Hallucination, and Moonshot's Own Assessment
Lost behind the performance headlines are three weaknesses that bite in production.
Artificial Analysis's live measurements show 33 tokens per second (145th of 190 models) and 177 seconds to first token. Northflank, meanwhile, reports 62 tokens per second and 1.99 seconds. Most of that gap is capacity strain from overwhelming demand — OpenRouter carries a warning on the model that upstream capacity is constrained and 429 errors may be frequent.
Artificial Analysis reports that while accuracy improved, the hallucination rate rose from 39% to 51%. A higher intelligence score is not the same as factuality, and this is not a model you can use without verification.
The company says there is a clear difference in user experience between its model and both Claude Fable 5 and GPT-5.6 Sol. Its own read is that narrow benchmark margins and how a model feels in daily use are two different things.
9. How to Try It Today
Before the weights land, there are three ways to try it.
| Route | Details | Who it suits |
|---|---|---|
| Official Kimi app / web | Via kimi.com. A free tier is available | Anyone who just wants to judge the quality |
| Moonshot API | $3 in, $15 out, $0.30 on cache hits (per million tokens) | Teams embedding it in their own product |
| Through OpenRouter | OpenRouter. Same endpoint you already call other models through | Anyone comparing several models |
| Self-hosting | After the weight release (scheduled for July 27). 64 accelerators or more is the recommendation, and you need MoE-aware scheduling in vLLM, TensorRT-LLM or SGLang | Organizations whose data cannot leave the building |
Bear in mind, though, that as noted above capacity is strained and 429s are common. Before moving a production workload over wholesale, the realistic approach — the same thinking as in choosing between cloud and local — is to route a slice of traffic to it with a fallback in place.
10. Sorting Claims by Confidence
- API released July 16, 2026; 2.8 trillion total parameters; 16 of 896 experts active; 1 million tokens
- API pricing of $3 in, $15 out, $0.30 cached
- 57 points on the Artificial Analysis Intelligence Index
- 1668 on GDPval-AA v2 (Fable 5 = 1760, Opus 4.8 = 1600)
- First place at 1679 on Frontend Code Arena
- Cost per task of $0.94, below Opus 4.8's $1.80
- Hallucination rate up from 39% to 51%
- Its rank on the Intelligence Index (3rd, 4th and 7th all in circulation)
- SOX's weekly decline (-9% to -12.5%)
- Output speed (33 tok/s versus 62 tok/s)
- Active parameters of "about 50 billion" (an estimate, not an official figure)
- A "Modified MIT" license (inferred from K2)
- The distillation allegation — no public evidence, and researchers reject it on the timeline
- Embargoed chip access — a US official's claim; neither Moonshot's formal response nor any evidence is public
- Weights and license — unreleased as of writing (scheduled for July 27)
- "$470 billion of AI value lost in three days" — could not be corroborated in the major financial press
In short, the real story of Kimi K3 is the fact that a Chinese open model has come within range of the frontier. It got there at 40-50% below Western frontier pricing, though not at an order-of-magnitude discount; it trails Fable 5 and GPT-5.6 Sol overall; it is slow; and it hallucinates more. Even so, it pushed GDPval-AA v2 up 478 points in a single generation and has committed to publishing its weights — and that is what the market reacted to.
FAQ
Q1. Is Kimi K3 open source?
No. What is planned is the release of the trained weights, meaning open weights, not the release of training data or training code. On top of that, as of writing the weights themselves are still not out — the moonshotai/Kimi-K3 page on Hugging Face still shows a countdown and no license has been announced. The release is scheduled for July 27, 2026.
Q2. Is it really third on the benchmarks?
The basis is 57 points and third place on Artificial Analysis's Intelligence Index, behind Claude Fable 5 and GPT-5.6 Sol, and as of July 17, 2026 that reading is correct. But tallies putting the same 57 points at fourth and seventh also exist, because the result shifts with how the comparison set is drawn and when the tally was taken. Read "third" as a snapshot. On the coding-focused Frontend Code Arena it takes first, while on FrontierSWE it loses to Fable 5.
Q3. Is it really cheap?
It flips depending on the comparison. At $3 in and $15 out per million tokens it is about 40% cheaper on input and 40-50% cheaper on output than Claude Opus 4.8 ($5/$25) or GPT-5.6 Sol ($5/$30), so against the Western frontier it is clearly inexpensive. Measured cost agrees: $0.94 per task versus $1.80 for Opus 4.8 (Artificial Analysis). Against Chinese rivals, however, it is the most expensive of the group — roughly three times GLM-5.2 and 23 times DeepSeek V4 Pro. Simon Willison calls it the most expensive model any Chinese AI lab has released. This is not the order-of-magnitude discount of the DeepSeek shock.
Q4. Why did US stocks fall?
Because of the worry that cheap Chinese models will reduce demand for expensive compute. During launch week the Nasdaq fell 1.5% on Friday, Taiwan more than 6%, Japan 4%, and the semiconductor ETF (SMH) was down over 20% from its late-June high. That said, chip stocks subsequently pared their losses as dip buyers stepped in, and Bank of America called it merely a summer correction. The scale is different from the DeepSeek shock, when NVIDIA alone lost roughly $590 billion in a single day in January 2025.
Q5. Is it true that they stole Anthropic's model?
No evidence has been made public. White House OSTP director Michael Kratsios alleged industrial-scale distillation of Anthropic's Fable, but Moonshot denies it, pointing to the 15-day window between Fable's public release on July 1 and K3's launch on July 15. AI researcher Braden Hancock likewise told TechCrunch that distilling, training and releasing within two weeks is impossible. According to SCMP, several specialists criticized the accusation as political. There is not enough material to reach a verdict at this point.
Q6. Can I run it on my own PC?
No. Even at 4-bit quantization (MXFP4) the weights alone are around 1.4TB, which will not fit on a single H100, H200 or B200. Northflank estimates something like eight nodes of eight 80GB GPUs, and Moonshot recommends 64 accelerators or more. macOS and Windows are out of scope for production use; Linux plus NVIDIA is the assumption. This is a completely different scale from running a local LLM yourself.
Q7. Should I use it in production?
Switching over entirely, straight away, is not advisable, for three reasons: (1) capacity is strained and 429 errors can be frequent (OpenRouter's own warning); (2) the hallucination rate rose from 39% to 51%; and (3) Moonshot itself admits a clear user-experience gap against Fable 5 and GPT-5.6 Sol. The realistic path is to keep a fallback in place and route part of the traffic through it in an area where it excels, such as front-end code generation, then compare.
Q8. How much better is it than the previous K2.6?
The size of the jump is the real story here. On Artificial Analysis's GDPval-AA v2 it went from 1190 to 1668, a gain of 478 points, and on the firm's private agentic evaluation AA-Briefcase it gained 732 points. On Arena.ai's Frontend Code Arena it climbed 17 places, from 18th to 1st. What the market responded to was less the absolute ranking than how much ground it closed in one generation.
Related Articles
- Claude Fable 5 Released: A Full Breakdown of Features, Benchmarks, Pricing, How It Differs from Mythos, and the New Safety Design
- GPT-5.6 Release Explained: Luna/Terra/Sol, Benchmarks, Pricing, and How It Compares to Claude
- Claude Opus 5 Released: How It Differs from Opus 4.8 and Fable 5, and What to Watch When Migrating
- The Best Local LLM Models Compared [2026]: How to Choose by Use Case and Size
- What PC Specs Do You Need for a Local LLM? A Quick Guide to VRAM, GPU and RAM [2026]
- The Complete Guide to Quantization Formats: GGUF, GPTQ, AWQ, and Which File to Pick