@giffmana Environment and harnesses is more of data scaling (in traditional sense) than compute
Unless ofc they meant that their environments are now more (computationally) complex worlds
It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future.
I asked Astra to help edit my math paper draft and it found a counterexample to a proposition in the literature, a proposition which I was planning to cite... I have now emailed the author... sign of the times
Wrote a short guide and Colab worksheet for anyone wanting to practice manual backpropagation. It breaks down the gradients step-by-step without relying on autograd, starting simple and building up to an MLP.
Check it out:
https://t.co/rqwlXlphrd
@aditi17goel Lots of alpha here for people getting into ML!
Nice balance of intuition + actual details instead of the usual surface-level distilled content.
Lately, I've been going deep into ML internals. If you're studying ML or just curious about how things work under the hood, a thread of my recent deep dives:
@finbarrtimbers Technically, it's allocated during the first optimizer.step(), not the second training step.
The first optimizer step just happens to be at the end of the first training (and beginning of the second) step.
🚀 Excited to share our work at Bytedance Seed!
Knapsack RL: Unlocking Exploration of LLMs via Budget Allocation 🎒
Exploration in LLM training is crucial but expensive.
Uniform rollout allocation is wasteful:
✅ Easy tasks → always solved → 0 gradient
❌ Hard tasks → always fail → 0 gradient
💡 Our idea: treat exploration as a knapsack problem → allocate rollouts where they matter most.
✨ Results:
🔼 +20–40% more non-zero gradients
🧮 Up to 93 rollouts for hard tasks (w/o extra compute)
📈 +2–4 avg points, +9 peak gains on math benchmarks
💰 ~2× cheaper than uniform allocation
📄 Paper: https://t.co/3HbOwV2tLL
Cohere intends to acquire Perplexity immediately after their acquisitions of TikTok and Google Chrome.
We will continue to monitor the progress of those deals closely so we can submit our term sheet upon completion.
Thoughts: While LLMs provide powerful priors for RL, many recent studies show that simply narrowing the model's output distribution can improve performance, but this also exhausts the model's potential for further exploration and improvement.
*Is it a blessing or a curse?* 🤔