In era of pretraining, what mattered was internet text. You'd primarily want a large, diverse, high quality collection of internet documents to learn from.
In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit like what you'd see on Stack Overflow / Quora, or etc., but geared towards LLM use cases.
Neither of the two above are going away (imo), but in this era of reinforcement learning, it is now environments. Unlike the above, they give the LLM an opportunity to actually interact - take actions, see outcomes, etc. This means you can hope to do a lot better than statistical expert imitation. And they can be used both for model training and evaluation. But just like before, the core problem now is needing a large, diverse, high quality set of environments, as exercises for the LLM to practice against.
In some ways, I'm reminded of OpenAI's very first project (gym), which was exactly a framework hoping to build a large collection of environments in the same schema, but this was way before LLMs. So the environments were simple academic control tasks of the time, like cartpole, ATARI, etc. The @PrimeIntellect environments hub (and the `verifiers` repo on GitHub) builds the modernized version specifically targeting LLMs, and it's a great effort/idea. I pitched that someone build something like it earlier this year:
https://t.co/ANHhasxzD8
Environments have the property that once the skeleton of the framework is in place, in principle the community / industry can parallelize across many different domains, which is exciting.
Final thought - personally and long-term, I am bullish on environments and agentic interactions but I am bearish on reinforcement learning specifically. I think that reward functions are super sus, and I think humans don't use RL to learn (maybe they do for some motor tasks etc, but not intellectual problem solving tasks). Humans use different learning paradigms that are significantly more powerful and sample efficient and that haven't been properly invented and scaled yet, though early sketches and ideas exist (as just one example, the idea of "system prompt learning", moving the update to tokens/contexts not weights and optionally distilling to weights as a separate process a bit like sleep does).
I spent the last 2 years on the trail of the Zizians, a group of technically gifted young people who set out to save the world. Their ideas led them down a path littered with violence and death. Along the way 6 people have been killed, 6 are jail, others have gone underground.
@farhaj@diffuse_store@_buildspace@FarzaTV 😂 it doesn’t do avatars yet…. It did come up with this tho: “portrait of Farhaj, the king of cats, sitting on his royal throne surrounded by hundreds of cats, by davinci, renaissance”
I launched an AI Poster store: @diffuse_store 🥳
Do you want to hang this image in your house?
Or you can try endless variations with dogs🦮, iguanas🦎 or bananas🍌 instead 🤖
Weird, funny and cool wall art awaits! DM for a discount code 😊
@_buildspace@farzaTV@farhaj
@giangp_ @_buildspace i guess i’m not sure yet if "Generate posters from your favorite ideas or interests" is more interesting to people than just having AI create something for them. will just have to see how people end up using it.
@polixonrio @_buildspace for me I'd like to use this for things like a party playlist- when the music I want only exists on soundcloud or youtube - then I can play them back-to-back.
@Yaroslav1127@_buildspace my thought is that it doesn't need to be "personal"- this would be amazing for team meetings / daily scrum. It could auto post to slack after the daily scrum is over with a summary.