@ArtificialAnlys 2.8T parameters and still behind Fable 5… but if they actually drop the weights this becomes the most important open model of the year.
Anyone else waiting to self-host this?
Still one of the strongest analyses out there
Kimi K3 hits 57 on the Intelligence Index — right next to Opus 4.8 and GPT-5.5.
Weights release would make it the open-weights king.
#KimiK3#OpenWeights#AIModels#MoonshotAI
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model
Key results:
➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation.
➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality.
➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params).
➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04)
➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores.
➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities
Other model details:
Context window: 1M
Size: 2.8T total parameters
Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
Modality: Native multimodal input supports text and images, and the model remains text-only for output.
Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
New open-weights flash model just dropped
Ant Group’s Ling 3.0 Flash scores 38 on the Intelligence Index and sits on the Pareto frontier.
124B total / only 5B active.
#OpenWeights#AIModels#Ling30#ArtificialAnalysis
Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models
@AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52.
Key results:
➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters.
➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545)
➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer
➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it.
Additional model details:
➤ Size: 124B total parameters, 5B active
➤ Context window: 262K
➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount
➤ License: MIT
➤ Providers: inclusionAI first-party API and DeepInfra third-party API
The next frontier model war is heating up. Fable 6 (Anthropic) is reportedly dropping mid-to-late August — same price as Fable 5, higher capability, and aimed straight at GPT-6. August could be absolute chaos.
#Fable6#Anthropic#GPT6#AI
Bro built one of the fastest-growing open-source AI agents ever… then disappeared into OpenAI.
Meanwhile OpenClaw keeps powering autonomous workflows for thousands.
The lobster era is wild.
#OpenClaw#OpenSource#AIAgents#AutonomousAgents#AgenticAI
@FabrizioRomano Arsenal won the league, and people are still saying 20+ years, bottlers, etc
You forget Arsenal won the league and reached the Champions League final.... What did your club achieve last season??
It’s rare to find an AI model that hits the sweet spot for speed, cost, and intelligence, but Grok 4.5 manages to do exactly that. It's quickly becoming a standout in the current landscape.
Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place amongst US labs
Muse Spark 1.2 (xhigh) lands at 54, up 3 points from Muse Spark 1.1 (51) and 11 points from Muse Spark 1.0 (43, April). It enters effectively tied with GPT-5.5 (xhigh, 55) and Grok 4.5 (high, 54), narrowly behind current frontier models Claude Opus 5 (max, 61), Claude Fable 5 (max w/ fallback, 60), GPT-5.6 Sol (max, 59), and Kimi K3 (max, 57)
Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release!
Key Takeaways:
➤ Muse Spark 1.2 gets closer to the frontier on agentic knowledge work. At Muse Spark 1.1's launch, we noted agentic knowledge work as its clearest gap; Muse Spark 1.2's gains help to close this. Its GDPval-AA v2 Elo rose 260 points to 1631, #5 among all models we have benchmarked and ahead of Claude Opus 4.8 (max, 1588). Terminal-Bench 2.1 gained 2 points (78% to 80%), and Tau3-Bench Banking rose 2 points (25% to 27%)
➤ Among the most cost-efficient models at its intelligence level. Muse Spark 1.2 costs $0.40 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing, with only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) cheaper in its intelligence cluster - GPT-5.6 Terra (max, $0.51), Kimi K3 (max, $0.86), and GPT-5.5 (xhigh, $1.18) all cost more per task. The cost increase over Muse Spark 1.1 ($0.29 per task) is driven by increased token usage per Intelligence Index task
➤ AA-Omniscience abstention rate increases. The score rose from 18 to 22 as the hallucination rate fell 10 points (38% to 28%) and the attempt rate dropped from 82% to 67%. This heavy abstention (not answering questions when unsure) now drives both the low hallucination rate and a lower accuracy (41% to 38%)
➤ Scientific Reasoning results remain largely unchanged. CritPt notably gained 3 points (15% to 18%), while SciCode fell 2 points (58% to 56%), and Humanity's Last Exam fell 1 point (45% to 44%)
Other model details:
➤ Context window: 1M tokens, unchanged from Muse Spark 1.1
➤ Pricing: unchanged from Muse Spark 1.1: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M
➤ Availability: Meta's first-party API at launch
AI model race right now:Claude Opus 5 topping charts
GPT-5.6 Sol right behind
DeepSeek V4-Flash crushing value
Meta Muse Spark 1.2 entering coding agentsClosed frontier is tight. Open models closing fast.Which one’s your daily driver?
Frontier-Bench scores are actually insane, and it's way less likely to hallucinate mid-project. Daily driver energy without the token anxiety. What are you building with it?
Claude Opus 5 is lowkey wild. Near-Fable 5 performance on coding & reasoning… at half the price? Anthropic really said, 'Why pay flagship money for flagship work?' Who's already switched their default to Opus 5? Drop your experience