I’m building a free study tool for my classmates that turns lecture notes into quizzes and spaced repetition cards. Some of them can’t afford paid study apps, which is why I started it. I��d use Devin to help fix bugs and get the first version ready for them to try. Would love a shot at this.
I’m building a free study tool for my classmates that turns lecture notes into quizzes and spaced repetition cards. Some of them can’t afford paid study apps, which is why I started it. I’d use Devin to help fix bugs and get the first version ready for them to try. Would love a shot at this.
DeepSeek is climbing fast.
DeepSeek-V4.1-Flash just reached #6 overall for frontend design on Design Arena with an Elo of 1347, putting it in the same performance band as Claude Fable 5.1.
A 39-position jump for DeepSeek. That’s a serious comeback.
BREAKING: DeepSeek‑V4.1‑Flash takes 6th overall on Design Arena with an Elo of 1347!
This marks a 39-position jump over the next-highest DeepSeek model - and DeepSeek’s return to a top-10 placement on Design Arena.
The model ranks in the same performance band as Claude Fable 5.1 on real-world frontend design tasks.
Congratulations to the @deepseek_ai team!
SWE-2 is available today in Devin across Desktop and CLI. We’re making it free for all Pro, Max & Teams subscribers for the next month.
Read more about how we trained SWE-2:
https://t.co/s5L08Rmxgi
Introducing SWE-2, our closest model yet to the frontier.
On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.
We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
GOOGLE IS LOADING UP ON AI TALENT BEFORE ITS NEXT MODEL 👀
Mechanize co-founder Tamay Besiroglu has joined Google DeepMind
More than a dozen former Mechanize employees reportedly moved to Google
Many are working on AI “midtraining”
Mechanize specializes in improving AI coding capabilities
Google had reportedly discussed a $1.5B+ talent + technology deal
Google is clearly not slowing down.
Gemini 4 prep getting serious?
The AI talent war is getting ridiculous.
Andrew Tulloch is reportedly leaving Meta for Anthropic, less than a year after joining Meta.
He previously worked on GPT-4o, GPT-4.5 and o3, and co-founded Thinking Machines Lab.
Anthropic keeps collecting serious talent.
The AI safety debate just escalated.
Former OpenAI and Anthropic researcher Jacob Coxon quit Anthropic, warning frontier labs are “gambling with our lives.”
Anthropic’s Evan Hubinger then said he sees a >10% chance AI could kill all humans within a decade.
https://t.co/WsNYDLABT4
GLM-5.5 Leak: Mythos Level 🐉
>Team is reportedly preparing to launch soon, could land next month
>GLM-5.5 reportedly expected to beat Mythos 5.1 and GPT-6 Astra
>September release window reportedly targeted
>Rumored 3T+ parameter model
>Expected to remain open-weight
>1M-token context window reportedly carried over
Can GLM-5.5 actually deliver Mythos-level performance?
Introducing ApprenticeBench: computer use + continual learning on a real job.
We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness.
No FDEs. Agents deploy themselves into the job. 🧵
SWE-2 looks seriously impressive. Cognition reports 92.8% on Terminal-Bench 2.1, leading this comparison, with competitive results across the other benchmarks and up to 70% lower cost.
I wish I could try it on the free study tool I’m building for my classmates. I haven’t used it yet, but I’d love to see how it handles the coding tasks I’ve been working through.
Great work, @cognition.
Introducing Fusion in Devin CLI
The most efficient frontier harness for Fable & Astra; 39% cheaper across coding benchmarks.
Pick your favorite model for planning and a cost-effective model for execution.
$0 input tokens. $0 output tokens.
Novita AI is offering Ling 3.0 Flash VL API access free until September 23, 2026 at 02:30 UTC.
It handles text, images, and video, with reasoning, function calling, a 262K context window, and up to 32K output tokens.
To try it, create a Novita API key and use:
inclusionai/ling-3.0-flash-vl
OpenAI- and Anthropic-compatible endpoints are supported. The listed default limit is 30 requests/minute, so plan around that when testing.
Offer details
https://t.co/fyRy7TwtIm