1/5 New rLLM Blog Post: Continual learning for agents 🧵
Can an AI agent keep learning after deployment—from real production traffic?
Each interaction yields one trajectory. Unlike build-time RL, you typically can't sample multiple rollouts for the same prompt, so GRPO's group-relative advantage no longer applies.
Our result shows a +16.2 points improvement on real SWE tasks.
Full post: https://t.co/P6Sog0P3Tg
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
Qwen Image 2.1 for character style infographics.
I didn't expect it at first, but after looking through these generated images, this model might actually also great at UI design.
Prompt (first image) ⬇️
🚀 Sol-H3 on DGX Spark: 768p in Under a Minute 🤩
Monday: 8×B300, 5s 768p in 1.65s — faster than playback.
Today: the same stack on one desktop Spark — about 56s hot E2E.
Five seconds of 1344×768 video at 24 FPS with stereo audio, on a single NVIDIA DGX Spark (GB10).
Two-stage, not the datacenter profile:
384p H3 draft → latent ×2 → H3-to-LTX VAE adapter → 768p LTX refine → VAE decode
No decode/re-encode between stages. Stage 2 is conditioned on the draft latent, so Gemma stays off the box. Quantized weights stay resident; sparse attention cuts the rest.
Stage 1 takes any MiniMax-H3 few-step LoRA. Timing is hot E2E (encode → both stages → video/audio VAE); cold start and MP4 mux are separate. Apache 2.0.
Server was realtime. Edge is one box, under a minute.
🔗 https://t.co/rfo0EVkJsI
Amazing team effort—full credits in the blog.
@haopengl33@lawrence_cjs@yitongli165665@shanasaimoe ,Jingyu Xin, @HaochengXiUCB@songhan_mit
This shot walks through seven rooms without a cut.
GPT-6 Astra mapped the space, the blocking never left the viewport, and Higgsfield Seedance 2.5 rendered the whole run.