your engineering team can write code 3-5x faster with AI agents.
but shipping? barely any faster.
coding agents didn't remove your delivery bottleneck.
they moved it downstream into review, testing, and context.
symptom 1: PRs pile up. bigger diffs, more branches. someone still has to manually understand every single one.
symptom 2: testing falls behind. agents produce changes faster than QA can verify. testing becomes the step everything waits on.
symptom 3: agents start cold. every task, they rebuild architecture, decisions, tribal knowledge from scratch. zero memory.
symptom 4: confidence lags the code. a task can look "done" while missing constraints a human would've caught instinctively.
net result: faster code. same slow ship cycle. bigger pile up.
that's why we built Alan (https://t.co/267w1GAgMy) the control plane for software delivery.
one context layer + orchestration across codex, cursor, claude code, github, slack, linear. verification built into the agent loop, not a queue at the end.
your agents code fast. now your entire SDLC can too.
As models become more capable, the risks associated with developing and testing them internally also grow.
We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage.
Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment.
https://t.co/ecbMMmVoox
New Stanford paper shows a leaderboard can tell you which agent scored highest on that benchmark, but it may not tell you which agent is actually better for your work.
Across 3 enterprise agent benchmarks, this paper finds that less than 3% of score variation comes from the agent itself.
A much larger 7–23% comes from the interaction between the agent and the specific task.
In plain English: an agent that looks better overall may simply fit the benchmark's task mix better.
The problem gets worse on hard tasks. On τ2-bench action checks, reliability falls from 0.752 overall to 0.000 on the hardest task quartile.
And across 50 train/test splits, projected reliability correlated with held-out reliability at r = -0.90.
So a leaderboard can look precise while being a weak guide for your actual workflow.
So for choosing agents, the better question is not "Which agent ranks 1st?" but "Which agent works reliably on my task types, difficulty level, and cost constraints?"
That also supports routing different task classes to different agents instead of forcing 1 model to win everything.
– arxiv. org/abs/2608.11323
Title: "Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations"
GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model
@Zai_org has just launched GLM-5.3, which ties Kimi K3 (60) for the most intelligent open weights models on the Artificial Analysis Intelligence Index (assuming weights are released as per Z AI’s guidance). Between Kimi K3 and GLM-5.3, the open weights frontier is closer than ever to the proprietary frontier.
GLM-5.3 keeps GLM-5.2's size at 753B total parameters and 40B active parameters. The Z AI team has shared that GLM-5.3’s weights are expected to be released within the next week.
Key Takeaways:
➤ GLM-5.3 sees the largest gain in agentic capabilities, placing it amongst frontier models. The model's Elo rating in GDPval-AA v2, our real-world agentic knowledge work evaluation, rises from 1524 to 1770, a 246-point jump. That places GLM-5.3 second among all models, behind only Claude Opus 5 (1855), and surpassing the previous open weights leader on this evaluation, Kimi K3 (1668), by more than 100 points.
➤ GLM-5.3 is less token efficient than its predecessor. Across the Artificial Analysis Intelligence Index v4.1, GLM-5.3 uses roughly 18,700 output tokens per task, up about 20% from GLM-5.2 (15,700) and 27% more than Kimi K3 (14,700). This would have implications for cost of deployment.
➤ At $0.68 per Intelligence Index task, GLM-5.3 costs 1.5x GLM-5.2's $0.44, but is still cheaper than other models in the same intelligence tier. GLM-5.3 is 19% cheaper per task compared to Kimi K3 ($0.84) and 45% cheaper than GPT-5.6 Sol ($1.23). GLM-5.3’s increase in cost is partially driven by a 20% token usage increase compared to its predecessor.
➤ GLM-5.3 makes an improvement in real-world knowledge accuracy, scoring 14 on AA-Omniscience up from 4 for GLM-5.2. GLM-5.3 is now the second-best open weights model on AA-Omniscience, behind only Kimi K3 (20). A more detailed breakdown of where improvements are made reflects a genuine gain in accuracy as opposed simply more cautious abstention, as accuracy rate (24% to 34%) and attempt rate (46% to 55%) both increased. However, GLM-5.3’s hallucination rate regressed up slightly, from 26% to 30%.
Additional model details:
➤ Size: 753B total parameters, 40B active (MoE), unchanged from GLM-5.2
➤ Context window: 1M tokens
➤ Pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens. An 81% cache hit discount is applied ($0.26 per 1M cached input tokens)
➤ License: MIT
➤ Accessibility: GLM-5.3 is currently accessibly through Z AI’s first party API. The Z AI team has shared they expect to release the weights shortly
Check out the full analysis of GLM-5.3 on Artificial Analysis: https://t.co/YI6GbwBUFD
A useful pattern is emerging for production agents: trace every run, make it replayable, and require approval at risky tool boundaries.
Observability explains what happened. Permissions limit what can happen. You need both.
It’s finally here! 🥁 Say hello to the new #TwitterAPI.
We’re rebuilding the Twitter API v2 from the ground up to better serve our developer community. And today’s launch is only the beginning.
https://t.co/32VrwpGaJw
The suspension of the H1B visa program is bad for the US, bad for innovation, and will shatter dreams and disrupt lives. As a former H1B visa holder, my heart goes out to all the families affected.
Trump just suspended the visa program that allowed me to move to the US to start @huggingface! Unfortunately, I won’t be able to vote in a few months but if you can, please vote him out, he's destroying what made America great in so many different ways! https://t.co/Y46hpawXdb
We have extra FDA-approved ventilators. Will ship to hospitals worldwide within Tesla delivery regions. Device & shipping cost are free. Only requirement is that the vents are needed immediately for patients, not stored in a warehouse. Please me or @Tesla know.
ER doc described 2 recent patients in enough respiratory distress to be admitted to the hospital, tested negative for flu and 20 common viruses, had CT scans consistent with Covid-19. State denied them both testing. Doc: “It made me realize that they weren’t testing anyone.”
@dan_s_becker@A_K_Nain Yes that is definitely a cause of concern but I think India is on a pretty good track with free testing facilities in most govt run hospitals. I am currently in the US and the condition here is far worse if anyone needs to get tested.
@A_K_Nain@dan_s_becker I agree the lack of testing facilities is a major problem but India has an advantage of numbers. We can still contain it and limit the numbers below 150, especially after recent travel restrictions.
@dan_s_becker I believe it's more about "What Indians are doing to contain this threat?". BJP govt has implemented some strict travel restrictions and India has pretty good healthcare infrastructure to accommodate masses if needed. But it mainly depends on how serious people are about all this
An example of how Incredible India really is. For millennia we have known how to harness the power of mind over matter. Veg, Non-Veg, what’s the difference? It’s all in the mind...😄