@naoufal_elh Would be cool to see a writeup on this. It seems like one of the harder parts of routing.
The folks @cognition gave this explanation which seems to make sense.
Cursor Router keeps improving from millions of in-product user interactions each week.
We intelligently classify and route requests, lowering latency and reducing cost based on the task.
Frontier data is a research problem.
For example, some of the biggest academic CUA papers are solely focused on data synthesis for SFT (OpenCUA, from last August) and RL (CUA-Gym, from this June). Good data requires diligent researchers and engineers.
By next year, I'd expect data talent to look increasingly similar to the current research scientist talent. Already hear small stories of data talent poaching wars and such.
BREAKING: Bland Speech v3 by @usebland debuts as the most human-sounding TTS model on our newest evaluation: Audio Realism Bench.
Traditional speech benchmarks increasingly struggle to separate frontier systems that all sound clean, fluent, and intelligible.
Audio Realism Bench instead asks a harder question: when two voices read the same transcript, which one sounds more like a real person?
Vetted native speakers evaluate recordings blindly across phone agents, conversational assistants, and spoken explainers, judging the subtle qualities that define human-like speech.
We also place real human recordings directly into the evaluation pool, testing whether evaluators can distinguish TTS models from actual human speech.
Bland Speech v3 takes the top spot in our initial rankings, followed by MAI-Voice-2 by @MicrosoftAI and Grok TTS by @SpaceXAI, all demonstrating important steps toward voice systems capable of sustaining genuinely human-feeling interactions.
Congratulations to the @usebland team!
Made the top of our Internal "Yapper" Leaderboard at Corgi.
If you've been following awhile, Drop a comment and say hey!
Let me know what you want to see more of.
If you are interested in learning more about Opus's over-editing behaviour at high reasoning effort, here's a blog for you:
1) Reasoning models specifically have a default tendency to over-edit
2) It can be alleviated with prompting
The example is ~identical too :)
Tech has decided that regulation is the enemy of progress, but innovation requires capital
Regulation builds the trust necessary to inject billions into markets
To enable free innovation, Andera raised a $37M series A led by Lightspeed to scale financial oversight with agents
There’s nothing inherently accelerationist or decelerationist about open weights
Models are just information
Context and timing determine their impact, and sweeping claims otherwise deserve skepticism
we're hosting a dinner at ICML at one of the most sought after reservations in Seoul (chef was on a Netflix show) with General Catalyst!
if you like great food and a small curated group of friends and researchers, would love for you to join!
https://t.co/zTRswg3FzF
I trained an LLM from scratch on pre-1900 text to see if it could come up with quantum mechanics and relativity.
While the model is too small to do meaningful reasoning, it has glimpses of intuition.
When given observations from past landmark experiments, the model can declare that “light is made up of definite quantities of energy” and even suggest that gravity and acceleration are locally equivalent.
I’m releasing the dataset + models and leave this as an open problem to the research community.
I also include what this project has taught me about intelligence in a mini essay linked below.
🧵(1/n)
Excited to share @Standard_Kernel's seed round and some reflections on what we’ve learned about kernel generation and what we believe is next. Grateful to our amazing team, supporters, and the broader community pushing this space forward.