yup, curriculum learning is probably as old as RL :D
People usually think of RL as "the agent collects its own data", but the robot can't really pick its own data to solve a hard exploration problem, and see useful gradient signal, if the chance of seeing success is super low (in a sparse reward setting). Which then leads to using denser rewards so that the gradient signal is not impossibly rare to stumble upon. Which then leads to tons of reward tuning and task-specific curriculums. Which then leads to suboptimal policies which are optimizing wrt a soup of mishmashed objectives. Which turns out leads to an inability to scale. So the details of SGS that matter are:
1. Expose a ton of "task configurations" (initial state, goal) to the robot, so that the difficulty of the set of tasks the robot can actually pick from is actually diverse.
2. Instantiate a very simple, automatic, adaptive curriculum based on success rate of each of these task configs (SGS) so that the robot doesn't waste experience in tasks that are neither too easy nor too hard.
So the surprising thing is that this simple instanciation lets us train policies with a sparse reward formulation, at much higher scale than previously possible, to achieve really hard robotic tasks :D
If you've ever worked at Google you'd see first hand just how much raw cognitive math-ish processing speed and working memory can decline for nationally ranked test takers and former top uni grads in their 30s and 40s, from decades of disuse and unambitious office work.
@elliotarledge Noob question, but why do we need an agent to hill climb this? Is it not possible to convert it into an optimization problem and solve that directly?
What GPT6-Astra tells us is that scaling spatial and 3D understanding rather than scaling action leads to better robotic manipulation.
We take a minimalist approach: what if we train a policy directly from a 3D understading model, not a VLM? Surprisingly, it outperforms most larger VLAs and WAMs.
Introducing Grounding Action Model—a brand new and SOTA policy train from a small 3D box detector WildDet3D(https://t.co/kpvWag0Scq)
Arxiv: https://t.co/Le5gdp6go3
Code: https://t.co/1YM093UgBc
After 2+ years in the robotics data space, we are shutting @Eidon_AI down.
The thesis was right. But the business is brutally hard.
We close this chapter by open-sourcing everything we built and sharing lessons for anyone venturing into the space. https://t.co/pyL9IkjQPT
“We found a steering vector for <concept> (usually using contrastive pairs)” is the paper idea that keeps giving…
It’s silly to be surprised by these results. Obviously capable LLMs have abstract representations of all concepts that appear in human texts—that’s how they work.
A carefully controlled look at looped transformers (arXiv 2609.19107):
1. Weight sharing is not compute-optimal on fresh data. It costs a constant ~1.06–1.16x compute at every scale.
2. The looped model's gains over vanilla come from elsewhere: the norm + input re-injection boundary operator, and growing depth mid-training. Both improve the scaling exponent. Neither needs weight sharing.
3. Looping is still a parameter-free way to get growth. Tied growth keeps the exponent gain at vanilla param count.
4. Weight sharing wins on multi-epoch, data-constrained training.
https://t.co/FYGexgqiWj
@marksaroufim This was a super cool competition. Although each of the individual operations in the target (squaring, modulo) are learnable with NNs, it seems nobody managed to train one that composes them. The winning approaches look like symbolic regression.