@din0s_ I think we’re quite fortunate that we learned how to code before the LLM era, and now we can reap the benefits of agents, but having the fundamentals solid.
The way I see it is that agentic programming is a new level of abstraction.
In era of pretraining, what mattered was internet text. You'd primarily want a large, diverse, high quality collection of internet documents to learn from.
In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit like what you'd see on Stack Overflow / Quora, or etc., but geared towards LLM use cases.
Neither of the two above are going away (imo), but in this era of reinforcement learning, it is now environments. Unlike the above, they give the LLM an opportunity to actually interact - take actions, see outcomes, etc. This means you can hope to do a lot better than statistical expert imitation. And they can be used both for model training and evaluation. But just like before, the core problem now is needing a large, diverse, high quality set of environments, as exercises for the LLM to practice against.
In some ways, I'm reminded of OpenAI's very first project (gym), which was exactly a framework hoping to build a large collection of environments in the same schema, but this was way before LLMs. So the environments were simple academic control tasks of the time, like cartpole, ATARI, etc. The @PrimeIntellect environments hub (and the `verifiers` repo on GitHub) builds the modernized version specifically targeting LLMs, and it's a great effort/idea. I pitched that someone build something like it earlier this year:
https://t.co/ANHhasxzD8
Environments have the property that once the skeleton of the framework is in place, in principle the community / industry can parallelize across many different domains, which is exciting.
Final thought - personally and long-term, I am bullish on environments and agentic interactions but I am bearish on reinforcement learning specifically. I think that reward functions are super sus, and I think humans don't use RL to learn (maybe they do for some motor tasks etc, but not intellectual problem solving tasks). Humans use different learning paradigms that are significantly more powerful and sample efficient and that haven't been properly invented and scaled yet, though early sketches and ideas exist (as just one example, the idea of "system prompt learning", moving the update to tokens/contexts not weights and optionally distilling to weights as a separate process a bit like sleep does).
Tired of downloading @Canvas_by_Inst files one by one? 😅
As a @UvA_Amsterdam student, I built a CLI tool that:
• Lists your courses
• Lets you pick files to download
• Organizes them neatly by course/module
📎 https://t.co/Px2xdJC8MC 🎥 demo
I'm noticing that due to (I think?) a lot of benchmarkmaxxing on long horizon tasks, LLMs are becoming a little too agentic by default, a little beyond my average use case.
For example in coding, the models now tend to reason for a fairly long time, they have an inclination to start listing and grepping files all across the entire repo, they do repeated web searchers, they over-analyze and over-think little rare edge cases even in code that is knowingly incomplete and under active development, and often come back ~minutes later even for simple queries.
This might make sense for long-running tasks but it's less of a good fit for more "in the loop" iterated development that I still do a lot of, or if I'm just looking for a quick spot check before running a script, just in case I got some indexing wrong or made some dumb error. So I find myself quite often stopping the LLMs with variations of "Stop, you're way overthinking this. Look at only this single file. Do not use any tools. Do not over-engineer", etc.
Basically as the default starts to slowly creep into the "ultrathink" super agentic mode, I feel a need for the reverse, and more generally good ways to indicate or communicate intent / stakes, from "just have a quick look" all the way to "go off for 30 minutes, come back when absolutely certain".
New #ReproducibilityCertification:
Reproducibility Study of "Robust Fair Clustering: A Novel Fairness Attack and Defense Framework"
Iason Skylitsis, Zheng Feng, Idries Nasim, Camille Niessink
https://t.co/dBaVap7CGE
#adversarial#fairness#clustering
Yoshua Bengio joined the party and started a blog.
His first blog post is about leveraging the advances in machine learning to help tackle climate change: https://t.co/NBHLIAcBql
Witnessed the most amazing thing on the train to Edinburgh yesterday. A guy boarded in Wigan & sat opposite me. He went to sleep for an hour.
When he woke up he bought a sandwich, ate it & went back to sleep. (This isn’t a maths test, you don’t need to know the distance/speed).
Most and least happy countries in the world according to the recent Word Happiness Report. Finland is most happy, along with Norway, Sweden, Australia, and Canada. The rest of the map is a humbling reminder of the hardship and suffering in the world today. https://t.co/ak2eLbH4ad