Created an agent skill called “Visual Explainer” + set of complementary slash commands aimed to reduce my cognitive debt so the agent can explain complex things as rich HTML pages. The skill includes reference templates and a CSS pattern library so output stays consistently well-designed. Much easier for me to digest than squinting at walls of terminal text.
https://t.co/TsbtZwCtxg
Sakana AI developed a new coding agent, ALE-Agent, trained to solve NP-hard optimization problems.
Our agent participated in a live coding competition, the challenging AtCoder Heuristic Contest, and ranked #21 out of 1,000 human participants!
Learn more: https://t.co/KOYxix8oy0
Introducing ALE-Bench, ALE-Agent!
Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems
Blog: https://t.co/jjfLyIYQcY
Paper: https://t.co/PdHvigUCa3
ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company.
What makes ALE-Bench unique is its focus on hard optimization problems that demand long-horizon and creative reasoning. It’s open-ended, in the sense that true optima are out of reach (NP-hard) and scores can continuously improve. We believe this benchmark has the potential to become one of the key benchmarks for reasoning and coding in the next generation.
ALE-Agent is our end-to-end agent that we specifically designed for this challenging domain. In fact, our ALE-Agent has already built an impressive track record in the wild! In May 2025, our agent participated in a live AtCoder Heuristic Competition (AHC), alongside 1,000 other participants in real-time. AHC is considered to be one of the most challenging coding competitions in this domain.
Our ALE-Agent achieved an impressive ranking of 21st out of 1,000 human participants in the competition (top 2%), marking a turning point for AI discovery of solutions to hard optimization problems with a wide spectrum of important real world applications such as logistics, routing, packing, factory production planning, power-grid balancing.
We look forward to applying this technology to real industrial optimization opportunities. Building on the insights from this study, Sakana AI will continue to tackle the challenge of developing AI with even greater algorithm engineering capabilities.
ALE-Bench Dataset: https://t.co/tH7DuYk129
ALE-Bench Code: https://t.co/fbJhvKXlxQ
This research was conducted in collaboration with AtCoder Inc. (@atcoder). We are deeply grateful for their outstanding expertise and contributions in optimization and algorithms, which were invaluable in providing data, analyzing results, and enabling our AI agent’s participation in their contests.
Hot take: I rarely if ever do "git add *" or "git add ."
"git add -p" is super underrated. You basically do a mini code-review before making the commit. Essential part of my workflow
I just uploaded a 90 minute tutorial, which is designed to be the one place I point coders at when they ask "hey, tell me everything I need to know about LLMs!"
It starts at the basics: the 3-step pre-training / fine-tuning / classifier ULMFiT approach used in all modern LLMs.
Machine Learning Engineering for Production (MLOps) - DeepLearningAI
Building models is one thing. Making them useful is another thing. The first course of the MLOps specialization teaches the ML lifecycle & deploying models.
Course 1 is free on Youtube: https://t.co/IHNNAdgizs
Today, in collaboration with the @Harvard Lichtman Laboratory, we're releasing a novel resource to study the human brain — an imaging dataset covering a cubic mm of cortical tissue with traces of tens of thousands of neurons and 130M annotated synapses. https://t.co/FywrdxFWgt