@yacineMTB Wasting flops on? You rather are more sample efficient since the critic just gets the info and doesn’t have to extract it from incomplete data
@creus_roger@mitrma Thanks! Not published because it isn't a result, it's the baseline for
something I'm working on. Runs on my own stack, PPO with a few PufferLib
(@jsuarez5341) tweaks, tuned on Craftax.
On Astra, agreed that it would be cool. 1% of 5B is 50M steps, where we're at about 10 return.
@creus_roger@mitrma For scale, this is a 1M parameter MinGRU. The 1B run takes 14 minutes on a single
A100 and the 5B run takes 1 hours, at about a million environment steps a
second.
@creus_roger@mitrma We get there with plain recurrent PPO too, running thousands of environments on one GPU. At 5B steps it reaches the vault in 19% of episodes and the troll mines in 4%. It scores 69 there and 47 at 1B, where the best leaderboard entry is 41. No trillion parameter LLM needed.
microdux 🦆 all 14 official Microduck tasks, in JAX.
Walk, turn, spin, skate. MuJoCo Playground native, so registry.load just works.
Rewards verified term by term against the original rather than eyeballed.
https://t.co/IGS7J8km8Z
Our paper on Streaming RL under Partial Observability won Best Paper Award at the Big Worlds Workshop at #RLC2026. We enable memory in streaming RL using recurrent trace units.
Project page: https://t.co/gXGExiEIIz
Collaborators: @__noahfarr__ *, @CarloDeramo , @Jan_R_Peters
@jxmnop reports of its death have been greatly exaggerated. built a browser extension that brings the crimson back with one click https://t.co/sL9qmBCqPs
@yacineMTB It really isn’t. It takes the elegant parts and slaps on stuff like clipping to make it more well behaved because that’s cheaper than a proper trust region.