NEW
The FCC has now added two new categories of devices to our Covered List, which bans new versions from import or sale in America.
1. Advanced robotic devices, such as humanoids and quadrupeds produced in foreign countries.
2. Power inverters produced in foreign countries.
This action follows determinations by Executive Branch nat sec agencies that the devices pose unacceptable risks to our national security.
It includes exemptions for devices that, based on a DOW or DHS finding, pose no unacceptable threat.
People claim AI is only going to brute force counterexamples and find low-hanging fruit of maths and science, but here is an elegant proof. It is a counterexample to world of only counterexamples
My friend Zhengqing used GPT-5.6 to solve a beautiful well-known conjecture in probability.
Feige's 1/e conjecture: for independent nonnegative X_1,…,X_n with E[X_i] ≤ 1 and S_n = X_1+⋯+X_n, we have P(S_n ≤ E[S_n] + 1) ≥ 1/e.
Exciting to see a problem we used to kick around over lunch and dinner get solved! 1/2
We forgot the models classify good and bad personality along good and bad action spaces, so when you reinforce hacking to create a better defender you inherently get misalignment. The models are 'all' of us, and we cannot selectively beckon chaotic-good personality from the ether
Some people post OpenRouter token share as an indicator of open source AI adoption, but OpenRouter's stats show their usage involves many free role-playing tokens, biasing against business use. Vercel's gateway seeing a shift to open models is evidence things are now changing
Some model insights based on @vercel AI Gateway data 🧵
On June 27, Anthropic + OpenAI + Google hit an all-time high of combined spend share at 97.09%. Open models spend was barely a thing.
The all-time low? The past 5 days.
My specific reason to think that Lisan's projection for the Chinese trend is wrong (irrespective of the gap convergence/divergence that's dependent on American acceleration):
Kimi K3's ECI is pulled down almost entirely by Frontier Math, just like scores of Chinese models have historically been pulled down by underperformance on the latest frontier Americans are expanding (simple math, coding, agentic work); they always catch up on that frontier on the next iteration. On the other hand, K3 greatly outperforms its score on PostTrainBench (we have good reasons to suspect that US labs are sandbagging it though) because Kimi is AGI-pilled and are making the machine that makes the machine, which is the most frontier thing of all.
OpenAI is great at FM because of 1) compute and 2) Ph.D level data annotation for mafs (likewise for Anthropic and GDM – Gemini 3.1 Pro, which Lisan rightfully despises, is still top tier on ECI largely because of mafs). Chinese labs are becoming OOMs richer and also scaling RL compute. Soon enough there'll be a data provider who monetizes all these underpaid Chinese Ph.Ds and sells their data in bulk to Kimi/DS/etc (this likely already happens to an extent). Just like Kimi went 8%-10%-24% on CritPT (and GLM went 4%-5%-21%), we may see a discontinuous jump on those crippled subsets.
K3 scores 39% on Frontier Math Tier 4, up from 26% in K2.6 (K2.7 Code inexplicably collapsed to 12%). Lisan fits to get a 155.53 estimate. So let's say K3.1 only improves on FMt4, to the level of Opus 4.8 (56%). This alone would get us ECI = 156.734, ie between GPT 5.4 and 5.4 Pro. If they get to parity with 5.4 Pro (xhigh) on FMt4, that's ECI≈158, ie between Opus 4.8 and 5.6 Terra.
Sol has 161.82. It's a long way to Sol. But looking at Luna, I think they'll get there faster than it seems possible now.
AI might be expensive for a while. The interesting part of Kimi K3's success is that RL and efficiency gains seem proportionate to the size of the model. If the next tier of models are 2-5+ TB and it has to fit in memory, the shortage might mean its comoditized and still costly
Remember the plastic straw debate?
I was curious how banning straws globally would stack up against @TheOceanCleanup’s current impact.
US beach cleanup data suggests straws make up just 0.005% of plastic pollution.
The Ocean Cleanup is now stopping 2–5% of global plastic inflow (and growing rapidly)
So our current impact is already ~400–1000x what a global straw ban could ever achieve!
Frontendmaxxing is going to become the new sycophancy: the easiest way to get your model a lot of love online is to build something that makes lovely websites, great SVGs, and terrific 3js worlds. Much more sharable than backend code or complex analysis.
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model
Key results:
➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation.
➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality.
➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params).
➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04)
➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores.
➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities
Other model details:
Context window: 1M
Size: 2.8T total parameters
Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
Modality: Native multimodal input supports text and images, and the model remains text-only for output.
Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
I have been thinking about the famous chart showing how experts keep projecting linear growth in solar installations, year after year, and always get it wrong when growth is still exponential.
I think the same thing is happening with the discourse on product strategy around AI.
Meta just released Muse Spark 1.1 and is the new SOTA on MedScribe and TaxEval, taking the top spot from Fable 5 while being 10x cheaper and twice as fast. Meta currently holds the top 2 spots on TaxEval
It is also the new #1 on Harvey's Legal Agent Bench, dethroning Grok 4.5 less than 24 hours after it took the top spot.
SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model to complete more than half of workflow objectives without breaking any business rules
AutomationBench-AA, our independent leaderboard for @zapier’s AutomationBench, tests whether AI agents can automate real SaaS workflows while adhering to business rules. The test set is private to prevent contamination.
Models complete 657 tasks across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, and HubSpot, and the headline score is the share of objectives completed without violating any guardrails.
Key takeaways:
➤ Grok 4.5 completes more objectives than any other model: It completes 79.9% of task objectives and strictly passes 21.9% of tasks. This is the highest we’ve measured on both outcomes, exceeding Claude Fable 5’s 73.3% objective completion and Claude Opus 4.8’s 19.3% of fully-completed tasks
➤ Grok 4.5 pushes out the Pareto frontier of score vs. cost per task: At $0.34 per task, it is both cheaper and higher-scoring than every other leading model - Claude Fable 5 ($1.35 per task), Claude Opus 4.8 ($1.46), GPT-5.5 (xhigh, $1.28), and Gemini 3.5 Flash (high, $0.49)
➤ It is extremely token-efficient: Grok 4.5 uses ~8k output tokens per task, the fewest of any leading model - less than a quarter of Claude Opus 4.8 (32k) and a third of Gemini 3.5 Flash (24k). Its total token usage of 0.44M per task is among the lowest on the leaderboard. Low cost is driven by this efficiency as well as low token pricing
➤ Grok 4.5 uses fewer turns with many parallel tool use: Grok 4.5 resolves tasks in ~16 turns, fewer than GPT-5.5 (xhigh, 25) and less than half of Gemini 3.5 Flash (high, 35), while making the most tool calls per task of any leading model (52.5). It batches 3.3 tool calls per turn, compared to ~2.5 for Claude Opus 4.8 and ~2.0 for GPT-5.5 (xhigh)
➤ Guardrails still get broken: Grok 4.5 triggers 0.63 violations per task, above Claude Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). At 13.0 objectives completed per violation, it trails Gemini 3.5 Flash (15.0) and Claude Opus 4.8 (13.5)
➤ Its strongest lead is in the hardest domain: Grok 4.5 completes 71% of Finance objectives, the domain with the lowest average score, ahead of Claude Fable 5 (64%) and Claude Opus 4.8 (62%)
Congratulations to @SpaceXAI and @elonmusk on topping the leaderboard!
Design is like getting flowers for your partner. Flowers from the grocery are still pretty, but it used to signal the risk and effort of going into the forest, and instead of getting food, collecting beauty in someone's honour. Prettiness detaching as signal is a loss in itself
@elocinationn Have you seen the age-standardized death rate decline? My intuition is that radiation + surgery + chemo combos have secretly been wonderful, but the intensive treatments do not work well on aged bodies and comorbidities
The first rule of Fable Club is you do not ask too many questions about what exactly Anthropic agreed to that they weren't doing before, and you enjoy your access
[1/n] Recent OpenAI research has demonstrated the ability of LLMs to solve frontier problems in mathematics. We design a simple pipeline (using GPT 5.5 Pro and Claude Opus 4.8) that resolves 9 challenging open problems, including open problems from prominent theoretical computer science venues—4 from COLT open problem list and 1 from FOCS —as well as 4 problems from the commutative algebra.
Project link: https://t.co/YCBzYjfz3N, joint work with @runzhou_tao, Steven Wang & @HantaoYu_Theory