@robleclerc Yes. But the secret to a healthy and thriving subculture is to prevent it from being too useful and economically valuable, so it is not overrun with people who want fame, status and money. A-Life is way more fun than most other AI or bio conferences
Delighted to share our new work on automated open-ended capability discovery for frontier models!
Model eval is getting harder with more capable models and so we ask: can models self-explore and tell us their capabilities? 🔍🕵️♂️
Many thanks to @shengranhu and @jeffclune ❤️❤️❤️
Scaling laws in deep RL? Turns out that batch size, learning rate, and UTD (update-to-data) for getting the most efficient and scalable deep RL has predictable relationships. Checkout the analysis in new work by @_oleh & collaborators: https://t.co/QnX7EKX6yy
Great to see an extension of GLAM applied to larger LLMs and more complex environments ➡️ https://t.co/TPtHuqqsas 🔥
GLAM can be used to align and functionally ground LLMs in external environments' dynamics through Online RL (
https://t.co/zNjFu7XBsf)
TWOSOME is used on...
1/4
Language models trained with RL should do much more than just satisfy preferences. They can achieve conversational goals, do strategic reasoning, etc. But we need the right tasks to evaluate progress on this. LMRL gym aims to provide this: https://t.co/keM7nV2nSn
A thread 👇
Mistral 7B is out. It outperforms Llama 2 13B on every benchmark we tried. It is also superior to LLaMA 1 34B in code, math, and reasoning, and is released under the Apache 2.0 licence.
https://t.co/krGs0xwbLH
To pursue something for its own sake means not subordinating one’s epistemic drive to understand that thing to the thing as defined by a fixed conception that has been imposed upon you from without. Herein lies the delicate connection between autonomy and general intelligence.
One major reason why mathematics is considered difficult: proofs.
Reading and writing proofs are hard, but you cannot get away without them. The best way to learn is to do.
So, let's deconstruct the proof of the most famous mathematical result: the Pythagorean theorem.
The reason AI isn't more of a rigorous science is that the kind of person drawn to the idea of automating thinking does not actually have that much interest in thinking.
Excited to present our new paper on bridging the theory-practice gap in RL! For the first time, we give *provable* sample complexity bounds that closely align with *real deep RL algorithms'* performance in complex environments like Atari and Procgen.
Learn about designing simpler and more principled RL algorithms and much more with @ben_eysenbach from CMU!
Blog: https://t.co/MU39poLVJH
Video: https://t.co/EORozXGU7r
The hot mess theory of AI misalignment (+ an experiment!)
https://t.co/OukfipSkIJ
There are two ways an AI could be misaligned. It could monomaniacally pursue the wrong goal (supercoherence), or it could act in ways that don't pursue any consistent goal (hot mess/incoherent).
Toolformer: Language Models Can Teach Themselves to Use Tools
introduce Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction
abs: https://t.co/EimSc0SBXs
How long until we see language models grounded in APIs, enabling them to interact with real-world and trained with RL from human feedback? This could enable applications from robot control to user interface navigation.
I'll be thinking about this paper for a while: "Theory of Durable Dominance"
https://t.co/YlfHOKrosX
Upshot: "Contrary to economic assumption that increased competition displaces [dominant firms], competition instead entrenches dominants”
TLDR: Superstar effects in everything