a few months ago, Claude Fable beat Pokemon FireRed. it won using a very naive strategy, which it got away with because FireRed is a fairly simple game.
so i built a harder benchmark based on the ROM hack Pokemon Radical Red!
Interesting NYT story about the most popular Uber Eats delivery restaurant in the world - an Indian restaurant in a Houston strip mall. Some of the drivers only deliver from there,
yup very cool work, i was partially inspired by this as well!
the agent capability that i'm trying to isolate is specifically optimizing a team and move sequencing in a loop to achieve a single difficult goal (e.g. beating a gym leader). this leads to some really interesting qualitative results as agents search the large space of possible teams to find creative techs to specific problems.
and the best part is that you can arbitrarily scale difficulty up! if a battle is too easy then you can just lower the level cap, lower the team size, or disable certain Pokemon/items/moves.
although this benchmark is more of a fun project, i think it still tests important agent capabilities in an interesting way. an agent needs to:
- develop a strategy over a long-horizon under partial observability
- learn from its failures and adapt accordingly
- explore efficiently under a limited budget
i wrote up a longer blogpost about the benchmark and more detailed results here: https://t.co/sb3VUsBl3X
code is here if anyone wants to run evals themselves or contribute: https://t.co/CaHtuY3QvU
a few months ago, Claude Fable beat Pokemon FireRed. it won using a very naive strategy, which it got away with because FireRed is a fairly simple game.
so i built a harder benchmark based on the ROM hack Pokemon Radical Red!
the agent is decent at battling but struggles with team optimization. it always makes *local* updates to its team, constantly trying to patch 1-2 weaknesses without taking a step back and evaluating whether it should rework the team entirely.
over many attempts, it tends to forget changes it made early on and drop them, causing the same issues to come up again.
watch the agent try to beat Lt. Surge, the electric type gym leader, with a neat Cofagrigus tech:
For those wondering why an AI company needs a power trading desk, I wrote about this shift back in December: as compute scales, managing electricity exposure becomes increasingly tied to protecting the economics and utilization of the underlying GPU fleet. (link in comments)
@Dimsunlight an LLM needs the tools to teambuild the way a human would
in my setup the model battles for real and can look at its own history to update its team. should be sufficient context, but it's still a hard problem bc the search space is large and it needs to reason over long context
Our quant analyst @ash_locked latest research article:
There is no debate that AI infrastructure demand is accelerating, but the market is still struggling to understand when that demand actually becomes real grid-visible load.
I met @justoutquan on my very first day at Berkeley.
As my dorm neighbor in Blackwell, he quickly became the door I knocked on whenever I wanted to complain about a minor inconvenience, or chase the next random thought.
Then one day at Croads, sitting with Justin, @_shreya_s , and @armaanrgoel, and I pitched a random idea for a digital bumper sticker. By that night, we all were pulling an all-nighter for our first pitch competition. Somehow, that pivoted to food delivery, Good Neighbor, my first startup.
Justin taught me what hustle actually looked like: staying up all night for a vision, making memes about Mark Zuckerberg eating peas and cheese, driving around Berkeley making deliveries ourselves, failing, iterating, and waking up the next day ready to do it all again.
We never did build Tinder for hot dog delivery, but I have feeling @tomo is an even better idea.
Justin taught me hustle, but more than that, he showed me the lasting impact of kindness, conviction, and relentless energy. Beyond startups, he is one of the most thoughtful, and most supportive friends I know.
Iโm grateful I got to be a small part of the journey, and even more grateful that chasing ideas together led to a lifelong friendship. So excited to bet on Justin and @tomo โค๏ธ
overall i think if you see this in your VLM it's probably fine. the model learns to operate around it and inference-time patches will end up breaking things.
you can read the full blogpost here!
https://t.co/kihypKhaGi
first blogpost! i was training a VLM recently and noticed that the vision token norms were way higher than text token norms going into the LLM. i was worried this would break things because of the residual stream, but my model worked fineโฆ
i also swapped in the KV cache corresponding to a blank image during inference on a real image. the model still behaves closer to having received a real image than a blank one! this implies that image info is distributed across the entire prefill state, not just the KV cache