Nuclear scientist @Google building ML infra for Gemini, @GoogleDeepMind and @Waymo. Ph.D. @ImperialCollege. Steward of Gemini models. ๐ฌ๐ท๐จ๐ญ๐ฌ๐ง๐ท๐ด๐ถ๐ฆ
How do we know if an AI agent is ready for the complex, adversarial realities of the real economy?
Traditional benchmarks test LLMs on static games (like Chess or Poker). But in the wild, agents face fluid, unseen environments where static leaderboards saturate and data contamination runs rampant.
@ChrisGPT The simple solution here is to not believe anything until we officially release something. People posting fake benchmarks screams of trying to manipulate prediction markets.
My favorite AI joke:
Current models are getting close to PhD-level intelligence. After that, they're expected to achieve the intelligence of someone who decided not to do a PhD.
@suchenzang "Our effort began on September 1st after hearing a rumor...".
Brb, starting rumors about room temp superconductors, antimatter propulsion and curing cancer
Okay, this is the best way to explain why AI Slop is not what you think.
By using AI Denzel Washington to explain it all. This is actually really good. Well written, directed and generated.
It's about you not needing the studios anymore.
@davepl1968 Most people don't have a clear goal in mind. It's analogous to diffusion; without real specifications, token use in a project will expand to fill the available budget.