So, starting on X after much internal debating. Show some love ❤️
About me currently:
AI engineer @ Oracle.
Cancer Early Detection ML @ Strand LS.
Robotics @ Cynlr.
Soulsborne Pro @ PS5.
More below 👇
Same noise predictor. Different sampling path.
Part 10 checks DDIM's variance budget and what one noise error does to the next step.
Full derivation + runnable checks below. https://t.co/gzmgXRDlpi
The model builds a game player, watches it lose and edits the code.
AAArena tests that loop against archived human programs. Six ladders topped, six still out of reach. https://t.co/t1yx3mIgKk
Your fine-tuning examples don't all teach something equally new.
PASS allocates a limited budget around weaker concept support. The results include a tie worth keeping in view. https://t.co/ZGhVUpCzQA
The agent noticed the conflict. Did it tell you?
This study tracks the gap between seeing contradictory evidence, using tools and writing the final answer. https://t.co/GLq0rgUjA3
Same model. Larger group. A different risk question.
This preprint models why testing a few agents may leave the population effect unmeasured. The assumptions matter. https://t.co/UlWzsSPNFc
A warning in your IDE can now open a Copilot fix.
The JetBrains update also changes default models, MCP startup and the oldest supported IDE. Details below. https://t.co/MSJAp7C827