Just tried a new coding bootcamp 🚀 Turns out learning software feels like assembling a giant Lego set—one piece at a time, but the picture changes with every line of code. #CodeLife#SoftwareLove
Three years of war have taken away our children’s right to education. 💔
Now, school is finally returning in person, but my 3 children have no notebooks, pens, or school bags.
We need $200 to provide them with these basic school supplies. 🙏
Please donate, share, and help
Her objection was correct and it was the whole problem. She may only bill time she actually spent. Anything that makes her faster removes revenue. Every argument for adopting was an argument against her own invoice. https://t.co/8oOaU82cEw
After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 practice areas
Harvey LAB-AA tests models on a private set of 120 legal tasks built by the team at @harvey. The tasks span 24 practice areas from corporate M&A and capital markets to tax, litigation, and bankruptcy. Models work to create the legal outputs specified in the tasks, and each task is graded against a rubric of binary criteria. The primary metric we present is the all-pass rate: the share of tasks where all criteria in the rubric are satisfied, reflecting the high standard of real-world professional legal deliverables.
Claude Fable 5 (max, with fallback) from @AnthropicAI leads Harvey LAB-AA with a 14.2% all-pass rate, after falling back to Opus 4.8 in only 1 task. This is almost double the scores of the next best models Claude Opus 4.8 (max) and GLM-5.2 (max) from @Zai_org, which tie at 7.5%.
Key takeaways from Harvey LAB-AA:
➤ Frontier legal work is far from solved: At launch, most models pass a majority of individual criteria but very few fully satisfy the requirements of any given task. The best model, Claude Fable 5, fully satisfies rubrics on just 14.2% of tasks, leaving ~86% of professional legal deliverables incomplete. Claude Opus 4.8 (max) and GLM-5.2 (max) follow at 7.5%, MiniMax-M3 at 6.7%, and Claude Sonnet 5 at 5.0%, ahead of GPT-5.5 (xhigh) from @OpenAI and Claude Sonnet 4.6 (max), which both score 4.2%.
➤ Models can pass many requirements of legal tasks, but rarely all of them: the leading models pass >90% of individual rubric criteria, but 13 of the 28 evaluated models fully pass 0 tasks.
➤ The top open weights model scores just over half the frontier leader: GLM-5.2 (max) ties Claude Opus 4.8 for second with a 7.5% all-pass rate and criteria pass of 91.0% vs. 91.1% respectively, both now behind Claude Fable 5 (14.2%). GLM-5.2 reaches that at ~6% of Fable 5's cost per task (~$1 vs. ~$19).
➤ Cost per task spans ~950x: the most expensive model, Claude Fable 5, costs ~$19 per task, while Gemini 3.1 Flash-Lite passes 31.1% of criteria for ~$0.02 per task.
OpenAI just gave AI a new way to learn 🤯
With the latest Codex update, you can now “teach” AI by recording your screen.
Show Codex how you perform a task once — your clicks, steps, and workflow — then it can replay those actions later whenever you ask.
It’s like creating a digital teammate that learns by watching.
The future of AI interaction is moving from “tell it what to do” to “show it how it’s done.” 🚀