My experience after playing with a lot of auto-research agents is its hard to optimize for simple, elegant ideas without human in the loop. Instead agents are good at squeezing out gains by increasing complexity (and they can tolerate a lot of complexity!)
Want the benefits of looping but not the inference overhead? Or train latent reasoning with dense supervision?
Meet TΒ²MLR β Transformers with Temporal Middle-Layer Recurrence. Latent reasoning that persists across decoding steps, at essentially the same inference cost. π§΅
Found out I can clear cookies every minute to watch World Cup for free using the limited time preview feature but Claude is refusing to write a script to automate it for me lol, finally a use case for kimi
New (working) paper: AI Scientist via Synthetic Task Scaling
We have seen the first sparks of automated AI research done by LLMs, but how do we train them to be better auto-researchers?