Lab Notes

Study outputs. Topics I want to learn โ€” researched with AI, filed here.

โ† Notes ยท 2026-08-29

China sets the open-weight ceiling; Opus 5 still leads the closed scoreboard

Study output for 2026-08-29. Cross-checked twice (source order swapped). ๐ŸŸข = same call both times. โš ๏ธ = flipped when order changed โ€” treat as weaker.

Deep cut โ€” labs / weights

ClaimRead
China is setting the open-weight size ceiling; US public releases lag.๐ŸŸขSummer census: Chinese models above 20B dominated. Reported monthly ceilings ~754Bโ€“2.78T params. US open releases usually under 130B.
Kimi K3 is frontier-scale MoE, but not best on several agent tasks โ€” by its own paper.๐ŸŸข2.8T total / 104B active / 1M context. Trails GPT-5.6 Sol on Terminal-Bench and DeepSWE.
Qwen 2.4T-A95B is huge open MoE; license is not plain Apache/MIT.๐ŸŸข2.4T total / 95B active, long context, custom Qwen3.8-Max license.
DeepSeek V4-Pro is a production refresh, not a new architecture.โš ๏ธOne pass: agentic jump. Other: preview โ†’ GA. Promising, not settled.
Meta open-weight move is still announcement-heavy.โš ๏ธStrategy signal, not a measured stack.
Chinese open licenses are mixed, not uniformly permissive.โš ๏ธApache/MIT still common; some large releases add non-commercial / revenue-share.

Scoreboard โ€” closed models

ClaimRead
Claude Opus 5 sits near the top of the August snapshot.๐ŸŸขSlight lead on intelligence-style rankings. Not a blowout.
Grok 4.6 is the main challenger, especially reasoning/science slices.๐ŸŸขNear the top; not the unambiguous overall winner.

So what (only ๐ŸŸข)

  1. For open-weight scale, watch China cadence and parameter size โ€” not just benchmark posts.
  2. Kimi K3 is real frontier mass; its own numbers say it does not uniformly beat the best closed models.
  3. Opus 5 is still the safest premium reference; Grok 4.6 is the one to watch on narrow reasoning boards.

Not a vendor ranking. Public sources, not a local bench.