My book, Reinforcement Learning from Human Feedback is done!
This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.
Transferring as much of the intuitions of building Olmo as I possibly can in the book format.
The book is launching with an over 10 hour, full course with slidedecks, functional code for the training chapters, an example model completions library, and of course the free online web version.
Physical orders from Manning will ship in 1-2 weeks, and Amazon a week or so after. Thanks for your support!
With the new algo change, i want to show u guys my fav feature that we recently shipped. You can argue with the articles now!
(or just ask it normal questions if you're a more sincere, thoughtful, and curious person)
over the last 6 months I have been working on a clean room reimplementation of MW2 (2009) from scratch called Open Victor Zulu that runs entirely in the browser
finally have it in a demo-able state, lots to finish up on
Appreciate all the kind words over the last 48 hours
On our side, rapid response reliability + accelerated preplanned improvements have landed
More to come, but the goal is for an entire cloud to go offline without users noticing
@samuelcolvin@pydantic This is just super inaccurate.
Phoenix is fully OSS. It's literally free to set up and host. We don't charge for it. We don't charge by users, we don't charge per span, we don't charge period.
By that definition, Phoenix is the best ROI in the AI Observability market right now.