Bryce Mitchell and Mikey Musumeci debated whether wrestling or BJJ is gayer 💀
“If a man sits on his ass and lets another man on top of him, that is gay.”
Demetrious Johnson was losing it 🤣
Terry Tao's talking to ChatGPT about the Jacobian counterexample makes it clear how powerful domain expertise is for eliciting AI capabilities. I cannot believe his GPT 5.6 is the same as my GPT 5.6. I could never elicit that kind of capability out of 5.6 or Fable, even though we know they have it in them. I would fail because I couldn't follow what they are saying, or ask the right questions to push them towards a solution.
How much expertise do you need to elicit that capability? Could an undergrad math major elicit the intelligence we see in that conversation and that must have been present in the original Fable disproof? I would bet that if 100 math majors had Fable and tried to resolve the Jacobian conjecture, none of them would succeed. 100 math PhDs might have a higher success rate, but how much higher? Levent Alpoge seems like a prodigy who has done serious work in the space before. I imagine his conversation with Fable involved some pretty high-insight direction on his part, beyond what a typical PhD student could do.
As AI capabilities keep increasing, I expect the gap between their default performance and their maximum performance will persist or even grow. Put differently: I bet a layperson with access to a model 10x smarter than Fable could not elicit a finding like the Jacobian counterexample.
This is a special feature of scientific research, by the way. Most tasks have a ceiling on how well you can do them. There is no Shakespeare of customer service emails that is 100x better than the average customer service email. But scientific research is limited by the unbounded ceiling of everything we do not know. Perhaps as models get more intelligent, expert researchers will produce superhuman knowledge by eliciting genius capabilities, while the rest of us will be made equal, inputing the same prompts as everyone else and praying for rain.
Anthropic has a model that's too "dangerous"... Mythos..
Without access to the resources and tools Mythos has, I'm confident I can find more vulnerabilities than Mythos (with strict guidelines and rules of course)
I implemented the prorated pricing solution for Evolus and came up with the best approach for handling data as well as remainders. I'd write a blog on it but don't think I can.
None of our competitors have a solution as solid.