@Abhinavsns@SamJWasserman@Apple Plus one for local video models (wan, ltx, hunyuan) working on Mac. So much hardware potential on Mac, so many open source models that aren’t as performant on Mac as on an Nvidia card.
Claude 4.7 leads PostTrainBench while managing time better than 4.6
On ArenaHard (a writing quality and instruction-following benchmark) it jumps from 6.7% to 24.2%.
From personal observations, 4.7 writes more, and some of it is richer, but some of it is the same point rephrased as if the new angle were the idea. We certainly need more interesting writing evals!
Model shaping is still a craft of a few. That's what AI agents are for: learning it and doing it for everyone else.
As a part of FrontierSWE benchmark we built a 20-hour post-training task on @tinkerapi and found the real bottleneck is research intuition.
We built a new task to test AI research capabilities! Agents asked to use @tinkerapi from @thinkymachines to train a model on logic games. That involves writing full training pipeline, running experiments across recipes, and submitting the best model.