Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon!
It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback.
Here’s a look at the benchmarks:
"So NOW you care about performance??" is just more cope. I've always cared about performance, but I cared about human productivity and happiness more. Writing all web apps in Rust before agents would have been madness. Now it's trivial and cheap. Update your priors!
Sonnet 5.5 cost conflict — Anthropic “up to 30% cheaper per task” vs. Artificial Analysis ~50% *more* at max effort.
So, per-token price is almost meaningless;
Only cost-per-*passed* task mixed with *your* effort counts.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
"JEV-as-a-Judge: Accept When Confident, Escalate When Unsure"
This paper shows you can just use JEV for every evaluation instead of expensive LLM.
JEV basically acts as a cheap first-pass judge, returning both a verdict and how confident it is.
When confidence is high, keep the answer. When it’s low, escalate to a stronger LLM.
This simple routing keeps ~99% of GPT-6’s accuracy while reducing evaluation cost by a lot.
https://t.co/cP4EzOreaO
@mattpocockuk “Lock down your agents” is a strong one.
It’s like asking the agent to build a Minecraft world from scratch versus building a Minecraft world inside Minecraft.
What happens when agents produce code faster than CI can validate it?
Linear reports:
• Nearly 4× the tests since Jan
• PR wait: >6m → just over 5m
• Runner time/test: ~halved
Cut setup. Trim the critical path. Batch checks. Balance shards. Benchmark caching.
Agents were shipping code faster than our CI pipeline could keep up.
Our team optimized our pipeline, leading to a roughly 15% faster test suite and 50% reduction in runner-time spent per test, all while our test suite grew 4x in size.
@moofeez explains how:
https://t.co/gkEIqwrwjn
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Does your Claude model really need Claude Code…? 🤔
We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge:
1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost
2️⃣A simple harness can be competitive
3️⃣The native harness isn’t always the best.
Millions of people are using coding agents, but the impact of harness choice remains unclear.
(1/n) More details in the thread. 🧵
🎉New Paper Accepted !
Thrilled to share that the paper titled "A Survey on UAV-enabled Edge Computing: Resource Management Perspective” by Dr.Xiaoyu XIA, Dr.Sheik Mohammad Mostakim Fattah and Prof. Ali Babar has achieved acceptance in #ACMComputingSurveys. #CREST#EdgeComputing
CREST organized a year-end party for the members & #CRESTSummerProjects2021 students in #Adelaide 🎉. We gathered to share our stories throughout the year 2021 & the resolutions for the next year.
Happy holidays, everyone! 🎊
Looking forward to seeing you all refreshed in 2022.
Yesterday's a happy day 4 @crest_uofa 🎉🎊. We celebrated the #graduation of Dr. @_Chadni_, @Shagun_05@SkAnjitha Anupam, Nini Cui & many others working w/ us #crest_grad. Appreciate your great contributions. Wish you all the best for your next journey. @ecms_uofa @UniofAdelaide