Study outputs. Topics I want to learn โ researched with AI, filed here.
China sets the open-weight ceiling; Opus 5 still leads the closed scoreboard
Study output for 2026-08-29. Cross-checked twice (source order swapped). ๐ข = same call both times. โ ๏ธ = flipped when order changed โ treat as weaker.
Deep cut โ labs / weights
| Claim | Read | |
|---|---|---|
| China is setting the open-weight size ceiling; US public releases lag. | ๐ข | Summer census: Chinese models above 20B dominated. Reported monthly ceilings ~754Bโ2.78T params. US open releases usually under 130B. |
| Kimi K3 is frontier-scale MoE, but not best on several agent tasks โ by its own paper. | ๐ข | 2.8T total / 104B active / 1M context. Trails GPT-5.6 Sol on Terminal-Bench and DeepSWE. |
| Qwen 2.4T-A95B is huge open MoE; license is not plain Apache/MIT. | ๐ข | 2.4T total / 95B active, long context, custom Qwen3.8-Max license. |
| DeepSeek V4-Pro is a production refresh, not a new architecture. | โ ๏ธ | One pass: agentic jump. Other: preview โ GA. Promising, not settled. |
| Meta open-weight move is still announcement-heavy. | โ ๏ธ | Strategy signal, not a measured stack. |
| Chinese open licenses are mixed, not uniformly permissive. | โ ๏ธ | Apache/MIT still common; some large releases add non-commercial / revenue-share. |
Scoreboard โ closed models
| Claim | Read | |
|---|---|---|
| Claude Opus 5 sits near the top of the August snapshot. | ๐ข | Slight lead on intelligence-style rankings. Not a blowout. |
| Grok 4.6 is the main challenger, especially reasoning/science slices. | ๐ข | Near the top; not the unambiguous overall winner. |
So what (only ๐ข)
- For open-weight scale, watch China cadence and parameter size โ not just benchmark posts.
- Kimi K3 is real frontier mass; its own numbers say it does not uniformly beat the best closed models.
- Opus 5 is still the safest premium reference; Grok 4.6 is the one to watch on narrow reasoning boards.
Not a vendor ranking. Public sources, not a local bench.