Fuck it, ban me if you want, but fuck you @OpenAI, @thsottiaux, @sama!
You're absolute hypocrites. Quietly taking our money while dumbing down the models and throttling us behind our backs.
You bunch of spineless clowns, do you actually have the balls to own up to this scam? Stop acting blind and posting those bullshit teaser tweets every week about "quota resets" like you're doing us a favor. Half the time every single prompt gets hit with "Selected model is at capacity. Please try a different model," and you've quietly slashed our limits by at least half.
If you don't have the compute, stop selling subscriptions, refund our money, compensate us, and issue a public apology! The whole company—especially Altman and that bunch of Chinese staff running the LLM teams—is full of absolute pussies.
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use.
Today we're open-sourcing CUA-S1-FORMS, the first in the family: https://t.co/J1frbEbZTQ
Jev, now with open weights + vision.
Classified 1,697 SF Tech Week events with Gemma 4 26B-A4B.
Zero labels. No fine-tuning.
Jev-ify any open-source model on SimpleJev.
Introducing Bespoke Nimble: an open data, open model, open recipe for an open Jev.
Code and info: https://t.co/aC7kPejrcj
Model: https://t.co/snwKGdhn1I
Data:
* A new data curation recipe called contrastive data curation.
* Slightly change facts to generate negative data. This pushes the model to discriminate better and become a better decision maker. The calibration is implicit.
* Didn't do ablations but I think this is a critical piece!
* This also means training data doesn't need probabilities.
* Data covered 10 categories, and is fully synthetic.
* This data is split into train and eval.
Training
* LoRA finetune of Qwen3.5-9B.
* Distillation-free: we use Jev to only evaluate.
* No RL yet!
Serving
* Parallel constrained decoding as suggested by @NielsRogge and @harshagundal.
Results:
* The post-trained Qwen (Nimble) became substantially better on our curated eval: 66% for Qwen to 90% for Nimble. Jev is at 93%.
* 100ms on H100 and free to use on your macbook! Feel the AGI for free.
* 2 days of building in public. :)
Big caveat is that there is no standard benchmark to measure performance, and it's possible Nimble is much worse on other benchmarks compared to Jev. But it should be better than Qwen!
We thank @typesafeai for making Jev and the inspiring discussions in the community. Hope this release lifts all the boats and encourages more research and activity in this space.
We're adding support for AGENTS.md to Claude Code.
Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md.
You can toggle this behavior in /config.
Breaking: Browser Use + Jev = Ultrafast ⚡
Findings flights took 7s and cost only $0.0039 🤯
> new action space every step
> DOM state space
> small LLM fallback to type
(this video is at 1x speed btw)
Built a tiny open source browser agent. try it below ↓
We open sourced BrowserSkill, a bridge between your agent and your actual browser.
most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back.
> login state is already there, it just works where you're signed in
> captchas and confirmation dialogs come back to you, then it continues
> it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes
one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around.
one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.
https://t.co/LBOjW8rxXp
We open sourced BrowserSkill, a bridge between your agent and your actual browser.
most tools give the agent a blank browser. We let it borrow a tab from yours, then hand it back.
> login state is already there, it just works where you're signed in
> captchas and confirmation dialogs come back to you, then it continues
> it's a CLI, not an MCP server => any agent that can run a shell can use it, and you see every call it makes
one thing that's easy to miss: the agent asks before borrowing a tab, and that switch lives in your browser settings, not in a flag, so it can't be talked around.
one line to install, works with Cursor, Claude Code, Codex, Hermes, Openclaw, CodeBuddy, WorkBuddy. Everything runs locally, MIT.
https://t.co/LBOjW8rxXp
Today we’re unveiling Odyssey-3, a big step forward for foundation world models.
It can control robots, power humanoids, drive cars (on the roads of India!), train AIs, pilot drones, and even play video games.
We can’t wait to see what intelligent systems it enables.