Physics and computer science is almost a parent-child relationship. That's why many computer scientists have it in their DNA, whether they realise it or not
World models are all the rage these days, so it's worth reiterating a few points that are largely correct.
1. Yes, our agents need models. The primary use of these models is planning. Planning can be done in real time, to improve an immediate decision, or in the background when not much is going on, to improve future decisions.
2. Learning models that predict the next sensory percept, such as pixels, is insufficient. The models should predict agent state; agent state is a summary of the past observations.
3. Learning one-step models is insufficient. Models should be conditioned on sequences of actions (e.g., option models). Finding what sequences of actions they should be conditioned on is an unsolved problem.
'Empirical Design in Reinforcement Learning', by Andrew Patterson, Samuel Neumann, Martha White, Adam White.
https://t.co/HE6kB5jQLH
#reinforcement#experiments#benchmark
RLC will be held at the Univ. of Alberta, Edmonton, in 2025. I'm happy to say that we now have the conference's website out: https://t.co/ZjpvWi5jyV
We'll continue to update it, and the CFP will be out soon, but the relevant dates are already there.
@RL_Conference@UAlberta
Why I think most RL papers are probably wrong in 1 graph. This is a 100 experiment hyperparam sweep for 200M steps each. Didn't find anything useful until >10B samples. Showing anything at all about an algorithm requires extensive sweeps that are rarely done!
A year later and our work on Loss of Plasticity is finally published, in Nature no less! The Nature version is totally rewritten and has many new results:
https://t.co/QImypXpqQl
Congratulations to the authors:
@s_dohare@JFernandoHG
@LanceLan3
@rahman_parash@rupammahmood
Yesterday there was a completely-student-organized summit of the RLAI (Reinforcement Learning and Artificial Intelligence) research group at the University of Alberta, held at the lovely Amii headquarters. Nice folks and diverse new ideas!
Would AI researchers (students, profs and alumni) of Canadian universities be interested in having one unified Slack that we collectively own?
I've seen a lot of scientific and career value coming from the simple fact that we can communicate and share experience.
Ash Ketchum has been a part of millions of lives & while many stop watching the Pokémon anime Ash continued to be a role model trainer for new generations. He's lost every Pokémon League to teach children it's ok to lose & today he finally won. What a day to be a Pokémon fan :)