@SeanZCai Love the 3d scatter chart, would be helpful to see an index of all companies in the space beyond the obvious ones (there are prob 30+ still in stealth).
in some sense spacex buying cursor gave it permission to cannibalize its own moat: the task-level data that made the company so valuable.
now the deal is done. cursor can put the crown jewels to work in routers & reward models instead of preserving them as acquisition leverage.
Introducing Cursor Router, our intelligent model router that selects the right model for the task at hand.
Router delivers frontier-quality results at 60% lower cost.
If the task you’re working on is not repeatable, your prompt should be the transcript of a ~20 min unhinged ramble
you’ll probably get 10x better responses than your usual “you are a professional.. etc”
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
farewell to the Amp Editor, my daily driver for a solid few months up until 5.3-codex
hard to overstate how much has changed since then
agent companies must now have 0 attachment to their products and embrace change in order to meet the current pace
@DrorIvry@chrisbarber if that's true, then reviews would probably have to be deeply tied to specific tasks with verifiable outcomes instead of broader tools
@chrisbarber yea the proto version of that is the clawdhub skills store, which is mostly operated by agents.
thumbs-up scores haven't been very reliable because there's an incentive to self-promote / sybil, and agents are not using the comments section for some reason
most people think ideas come from:
- insight
- intelligence
- taste
- reading
- vibes
but in practice they actually come from:
- building the wrong thing
- hitting a constraint
- getting embarrassed by users
- realizing the obvious thing you missed
- noticing the second order effect you couldn’t see from the couch
a really great idea is the *output* of the work, not the input.
@a1zhang So the mental model is: REPL is the shared workspace, and sub-agents are invoked within that workspace like functions so LLM calls are small and don't carry the full context.
bro casually explains RL tuning for LLMs and the three critical components: training, inference, and environments. basically any RLVR algorithm such as GRPO comes down to this super simple concept.
Unexpected twist in the coding agent race:
GPT-5.2 (w/ the CodexCLI harness) is SOTA on Terminal-Bench. Opus-4.5 w/ Terminus is at a sweet spot of cost and accuracy, but it's surprising to see a ~10% gap between them.
Fascinating work by @Mike_A_Merrill, @alexgshaw and team:
We study frontier models across an array of agent harnesses. The best performing agent/harness combination in our experiments was GPT 5.2 with Codex CLI:
Day 2 of #PyTorchCon 🔥
What a ride. Talked with folks using #PyTorch to fine-tune models for drug discovery, cancer research, autonomous vehicles and, of course, customer support!
Thanks @PyTorch for having us!
Why did OpenAI train GPT-5 with less compute than GPT-4.5?
Due to the higher returns to post-training, they scaled post-training as much as possible on a smaller model
And since post-training started from a much lower base, this meant a decrease in total training FLOP 🧵
@JoshSeriesAI@iamtrask Since data needs to be priced competitively, we've built a platform that achieves that in the form of a marketplace. Check it out: https://t.co/kbY8gJEWSQ