Something magical happened today - had issues with login to @cursor_ai and @bot ( the issue was that was logging in with GitHub which used my Gmail as login - so it was logging me using google instead ) , raised a ticket bot responded by email - asked me to try some simple …
Been working on this a while, and super happy to share the first few pre-release chapter of this book on post-training. It’s called:
Post-Training AI: A Practical Guide to Fine-Tuning and Reinforcement Learning.
The idea behind the book is to take you on a simple line through post-training, to build intuition about how agents are learning in SFT, GRPO, distillation, and environments. It definitely skips a few concepts, and prioritises clear explanation of depth, but I think that’ll help most people start post-training.
The first two chapters are live in draft. They’re pretty fun already, they show you a minimal post-training implementation in a few hundred lines of code.
Chapter 1 is an intro that defines post-training. What each stage changes, what none of them can fix, and why the techniques now work outside frontier labs.
Chapter 2 is a speed run. You implement SFT and GRPO from scratch in pure PyTorch. Each loop fits in under 100 lines and runs on a free notebook GPU.
The rest of the book follows one model from base checkpoint to post-trained system, using TRL, transformers, PEFT, and the Hugging Face Hub.
Early release means raw and unedited. You read while I write. If you find problems, tell me. That is the point of releasing early.
https://t.co/AjtnDe92n0
Asked it to create a promo walkthrough video of my app - kept recording a blank screen - half way through generating the promo it ran out of credit - what a waste of time and money - keep control!
Github plugin yet could not clone a repo - I logged in to gh on its computer and it kept telling me to to relogin - finally after 10% of my usage it was able to clone a repo
Getting ready to publish my complete guide to RL for LLMs tomorrow morning. Although the post contains many of my own thoughts / learnings, it is also a synthesis of so many great resources that have been published over the years:
- The RLHF Book (https://t.co/n3aUxqWtgS) by @natolambert
- Reinforcement Learning by Richard S. Sutton and Andrew G. Barto
- Spinning Up in Deep RL (https://t.co/EYcOOllvIy) from OpenAI
- Build an LLM (https://t.co/6HfXuofVwq) and Reasoning Model (https://t.co/l5uS247A04) from Scratch by @rasbt
- Various notes (https://t.co/hpsyAKOnRv) and papers (TRPO, PPO, etc.) from John Schulman
- Policy Gradient Algorithms (https://t.co/caNLakSjo5) by Lilian Weng
- A Vision Researcher’s Guide to RL (https://t.co/25awrjj5Bd) by @YugeTen
- From REINFORCE to Dr. GRPO (https://t.co/M23e3Fl3VZ) by @qingfeng_lan
- Async GRPO in the Wild (https://t.co/A5qeYMwcy4) by @yumo_xu
- Open RL infrastructure like TRL (https://t.co/TGrrnJ574t) and OpenInstruct (https://t.co/L9pObiUb1c)
I highly recommend reading all of them. They’ve truly helped me to learn so much.
BABYSIT your agent !
Loops are overrated , expensive and results in poor quality software - irrespective how strong a verification is in the loop.
deliver your feature in vertical slices - be the human in the "loop" for each slice - execute each slice in its own session ( cheaper ).
Why Human in the loop - as we dont know what we want until we see it.
the quant wars around qwen 3.8 27b dense are the best thing happening in local ai right now.
@atomic_chat_hq just dropped their AD ladder, 1bit at 8.5GB up to q8, and measured EVERY community gguf against BF16, including the spots where their rivals beat them, shaded in red on their own chart. that's how you publish quants.
the part that matters most is the low end. 9.9GB at 83.5% agreement, 13.8GB at 92.4%. a frontier-table 27b just became runnable on 12GB cards and 16GB macbooks.
their numbers, and i'm about to verify the ones that matter on my own metal, including one test nobody's thought of yet: whether the mtp speed flag survives their quantization. numbers soon.
Introducing fx, a tiny, open, native coding agent from Vercel Labs.
Originally an internal tool, fx is a harness and CLI written in Zig, optimized for research and embedding in larger systems. Today, we're open sourcing it.
fx is built on three principles:
1. Fast. A single native binary, no runtime to install. It cold starts in 10µs and does no unnecessary work or I/O before accepting input. fx is the answer to "how fast can a coding agent be?"
2. Light. The 6.3MiB binary uses single-digit megabytes of memory at baseline, made for instant installation and embedding in resource-constrained environments and agent sandboxes.
3. Open. Apache-2.0, model and provider agnostic, suitable for local and cloud inference. Its small core extends through skills, plugins, and MCP.
Minimalism is an obsession throughout the entire harness: system prompt, tools, features, binary. The goal was to keep context usage and time to first token low, and make fx optimal for model benchmarking, sandboxing, evals, and gyms.
You can use fx directly or embed it as infrastructure. The CLI feels more like a Unix shell than an IDE in the terminal: it preserves scroll history, produces minimal output, and uses complex TUI rendering very, very sparingly. Programmatically, 𝚏𝚡 𝚊𝚜𝚔 --𝚓𝚜𝚘𝚗 gives structured output, 𝚏𝚡 𝚊𝚌𝚙 connects to editors and other clients, and WebAssembly can even run the whole thing inside the browser (see: https://t.co/wf2Trg47sC).
Privacy is a design constraint: no product telemetry, sessions and usage stay local, and no source code or prompts are shared with any endpoint other than inference. With local inference and auto-updates off, fx is fully hermetic.
fx is experimental. Use at your own risk and expect frequent changes. Chat with us on X (https://t.co/A2AB2YythC) or file issues (https://t.co/GEjTHSoa1J).
𝚌𝚞𝚛𝚕 -𝚏𝚜𝚂𝙻 𝚏𝚡.𝚜𝚑/𝚜𝚎𝚝𝚞𝚙.𝚜𝚑 | 𝚋𝚊𝚜𝚑
https://t.co/g2uEuXhGnt
another incredible work for those who has ONE DGX-spark/GB10 you can now able to run full DeepSeekV4 Flash 0731 at 59 toks/sec you better follow @bleysg .
if you want to run GLM 5.2 on 2 DGX spark you also should check out his repo.
another incredible work for those who has ONE DGX-spark/GB10 you can now able to run full DeepSeekV4 Flash 0731 at 59 toks/sec you better follow @bleysg .
if you want to run GLM 5.2 on 2 DGX spark you also should check out his repo.