Big news: Qwen3.8-Max-0902 by @Alibaba_Qwen just debuted at #1 overall in the Code Arena: WebDev with 1691 pts!
It scores 3 pts above Claude Opus 5 (Max), 17 pts above Kimi K3 (Max), and 22 pts above the previous Qwen3.8-Max.
Priced at a blended $5/MToken, Qwen3.8-Max-0902 also claims the highest-scoring position on the Pareto frontier! Stay tuned for a closer look at its Pareto positioning, and for Agent Arena scores coming soon.
Its strength carries across every Code Arena: WebDev category:
- #1 in Data & Analytics and Consumer Product
- #2 in Brand & Marketing, Gaming, and Simulations
- #3 in Content Creation Tools and Reference-Based Design
Congrats to the @Alibaba_Qwen team on this huge update!
We’re introducing Gemini 3.8 Flash ⚡️ built to tackle complex agentic and multi-step tasks with even greater diligence.
Our most intelligent workhorse model yet delivers significant improvements in reasoning, evolving to an AI partner that doesn’t just write code, but can also navigate complex projects.
While solving ambiguous and high-friction tasks is immensely helpful, it can also be expensive. Fortunately, 3.8 Flash features the usual effort controls, ensuring that the amount of thinking required to accomplish the task at hand is proportional to the tokens spent.
Watch Gemini 3.8 Flash combine native video understanding with advanced coding to autonomously build this 3D game, play it to find errors, and execute code changes in a seamless agentic loop in @antigravity.
This is pretty catastrophic and I'm not sure how either of these companies reverse the trend. It really appears the AI bubble is based on the whims of maybe a few hundred companies spending money on two companies to justify five companies spending money with one company
You can now run Cursor cloud agents on your infrastructure, including pools of machines that automatically scale with demand.
This lets you give agents access to internal services or specialized hardware, while the agent loop stays in Cursor.
New research: Training a Misaligned Reward Seeker
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.
In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
Read more: https://t.co/gs2ZjYkPan
In another case of it being quicker to build a tool than hunt around to find something that does what you need, here's a little vibe-coded thing for turning one or more GeoJSON shapes into a rendered PNG https://t.co/74K7U0hfkW
Copilot can approve the PR.
Every review now says whether it would approve. Actual approve is off by default. Flip it and the approve counts toward required reviews; new commits dismiss it like a human.
One MCP front door.
Kong AI Gateway 2.0 is GA on Konnect. One route, one catalog, ACL so tools you cannot use never show up in tools/list. That is the ship.
Mythos is a program, not a picker.
Life Sciences Verification is invite-only. Cyber Verification is near future. Claude Security already runs on Mythos 5.1. Same sticker price as Fable.
Pics sits on the slide.
https://t.co/bO54YqlNgL, Nano Banana underneath. Click an image in Docs or Slides and the editor stays on top. Upscale to 2K or 4K. Drive is later.
claude-fable-5-1 is live.
API, Bedrock, Google Cloud, Foundry, Claude Platform on AWS. Cache reads at $0.25, seventy-five percent off Fable 5. That is the SKU, not the bench table.