Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with self-improving strategies.
It achieves state-of-the-art performance among open-source models of comparable size and delivers performance comparable to Claude Opus 4.8 across reasoning, agentic, and coding tasks:
✅Terminal-Bench 2.1 (86.1)
✅SWE-Bench (86 on verified, 65.1 on pro, 79.6 on Multilingual)
✅DeepSWE (56)
✅HLE (44.6)
✅ClawEval (81.4)
✅Tool Decathlon (71.2)
Ornith-1.5 takes a major step toward training foundation models through end-to-end self-improvement, extending the self-scaffolding strategies introduced in Ornith-1.0 into a more complete self-improvement loop: the model proposes new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning, continuously creating new learning experiences from which it can improve.
All models, along with their quantized versions (FP8, GGUF, MLX, and NVFP4), have been released under the MIT License, enabling unrestricted commercial and research use.
📘Tech Blog: https://t.co/OZ63scRWLB
🤗Huggingface: https://t.co/mGJLwhrQOM
Unpopular opinion: When I'm working on real stuff, I feel like current SOTA LLM is enough. I rarely run into hard problems, coding web apps and mobile apps is mostly routine. Opus 4.8, GPT sol or Kimi k3 is enough for this purpose. From this point, lower price brings more value than higher intelligence.
Hey builders, have you also been getting emails from people trying to blackmail you? They claim they've found a "security issue" and demand money in exchange for not making it public. Even if you reply, they often pretend they never received a response and keep asking for payment. 🙂
One of them is [email protected]. They appear to be running automated scans and are probably spamming half the internet. In my case, they sent me old DNS records that had already been changed before they contacted me and tried to present them as a "security issue."
By the way, isn't this kind of behavior illegal?
Introducing Claude Sonnet 5, our most agentic Sonnet yet.
It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.
As a result of a US government directive, we are suspending access to Claude Fable 5 for all users. You can continue to use all other Claude models.
Here’s what this means for you:
Across Claude products, new sessions will run on your selected default model or Opus 4.8, and existing Fable 5 sessions will end with an error.
On the Claude Platform, requests to Fable 5 will also return an error. Please update your integrations to other Claude models.
We know this is a disruption to your workflows; we appreciate your patience and support.
@TheoAugust8 Building https://t.co/Uu6LoNz7Wl, monitoring system for saas builders, because every builder should be notified about issues before their customers notice.
@JorgeMvrfil I’m working on mobile app for my monitoring system https://t.co/RspzalIeOO because not all busy people are sitting by the computer all day
@TTrimoreau Not sure which is the best, but I use Polar over Stripe because it’s merchant of record + it’s simple. But I’m eager to hear experience with other payment providers too.
@sflorimm I’m building https://t.co/RspzalIeOO - monitoring for busy saas builders. It watches your websites and other business critical systems and notifies you when anything needs your attention.
Introducing Gemini 3.5: our newest family of models combining frontier intelligence with real-world action.
The first release is 3.5 Flash, our strongest model yet for agents and coding 🧵