Reinforcement learning has libraries.
What it’s missing is workflow discipline.
Shipping RLCLI today.
1. Local-first
2. Reproducible
3. Built for real experiment iteration
GitHub: https://t.co/63AJqOPOhG
Install: https://t.co/OPFFBGSDD1
you can try playing as the player yourself against any of these levels. been trying to beat level 2 for the past hour buts its too fast. would love to see someone try beating it.
source code - https://t.co/Xi4hpUZRaR
command - python play_human.py --level 2 --seed 0 --fps 60
@kjcodes_ more chance of overfitting when changing billions of parameters in full ft, might shift the models pretrained structure, can run into catastrophic forgetting and might have waited hours for it. Full ft better for high quality data and if target task far from pre trained dist.