@silennai@explabsai@OKfallah Congrats! Loved the world-model-as-a-harness, curious where prompt optimization tops out vs. training a world model like DreamGym does?
@samaneggs Love this, most setups just bury failed runs in a summary. Curious though, how do you tell when two attempts are the same dead end vs actually a new idea?
@spisallyouneed Love this. Resets have been quietly carrying RL benchmarks for years. What's the longest horizon an agent has to survive in the current suite? Will be playing around with the environment hub this week.
@gpuemi@wafer_ai Used it for a multi-hour coding session and TTFT stayed snappy even at long context, what’s keeping prefill off the decode path, disaggregated serving?
@proceduralia This matches everything I've hit building these systems. The apps are the easy part; the hard part is reliable resets, checking against ground truth, and enough variety to actually learn from. Love the PRISM framing!
@_xjdr So on nvfp4, nvidia's own numbers show it's already as good as fp8. it's actually ahead on long-context recall and tool-use, tied just about everywhere else, with only a tiny dip on one coding test and that's with no extra tuning at all.
https://t.co/MHzQ1Xg6Vv