The danger of selling to big companies, if you're a startup, is that they don't say no outright. They have months of meetings with you first. Since you hate meetings, that seems to you a sign of commitment. But it's not. They love having meetings! It's almost all they do.
arXiv 2602.03061 reports 60 benchmark comparisons where its accuracy estimator beats plain averaging, while giving no error bar (i.e., uncertainty interval) for any of them.
What I observed from running their code without modification: For models whose answers vary a lot, their simulator works. But making the model more consistent adds noise instead (33% more, 18% to 50% across 250 runs).
The above is one out of five claims in the paper that I checked. The other four held up.
My reproduction supports a different claim: that their estimator helps on average across all 60 comparisons, but not enough to show in any single one.
https://t.co/tAiohgszOy
As pointed out by @sirbayes, this paper https://t.co/rqZn5gePU3 has a formal investigation into the observation from the first tweet -- that "my models were completely unable to learn to perform actual first-order logic -- despite the fact that this ability was definitely part of the representable function space. Instead, the models would inevitably latch onto statistical keyword associations to make their predictions"
The paper concludes: "while models can attain near-perfect test accuracy on training distributions, they fail catastrophically on other distributions; we demonstrate that they have learned to exploit statistical features rather than to emulate the correct reasoning function"
Basically, SGD will always latch onto statistical correlations as a shortcut for fitting the training distribution, which prevents it from finding the generalizable form of the target program (the one that would operate outside of the training distribution), despite that program being part of the search space.
Jack Dorsey on why building something for yourself can be better than trying to solve a problem
When a 29 year old Jack Dorsey joined the podcasting startup Odeo, he soon learned that no one there really cared about podcasting either:
“That was one of Odeo’s biggest failures. We were not building tools for us. We were building tools for other people.”
As Odeo struggled, the team began to look for new ideas, and Jack presented his idea for Twitter. Somewhat counterintuitively, the product would not solve a specific problem:
“Twitter solves no one’s problem at all. It was something we wanted to use. It was something we wanted to see in the world. It was something we wanted to use on a daily basis, and that’s all that drove us. That’s what got us up every single morning, and that’s what made it meaningful.”
It turns out there were more people like Jack and the Odeo team who wanted to use it to.
“I think that is one of the biggest lessons I learned—Twitter did not start as a company. Twitter started as a product within another company that was failing. And to me, this really emphasized the fact that entrepreneurship is not necessarily starting a new company. It’s actually just taking significant risk to build what you want to see in the world.”
Source: @Cal_Engineer (Oct 2013)
Prediction: we’ll see loads of conjectures and theorems proven and disproven by LLMs over the next 6 months and it will make basically zero difference to the world, and won’t even contribute meaningfully to advancing mathematics.
if your data is stored in a database that a company can freely read and access (i.e. not end-to-end encrypted), the company will eventually update their ToS so they can use your data for AI training — the incentives are too strong to resist
Paul Graham explains why you shouldn’t try to be a visionary
“Empirically, the way to do really big things seems to be to start with small things and grow them bigger. Want to dominate microcomputer software for decades? Start by writing a basic interpreter for a machine with a couple thousand users. Want to make the universal website and a giant vacuum for people’s time? Start by building a website where Harvard undergrads can stalk one another.”
Paul Graham continues:
“Neither Bill Gates nor Mark Zuckerberg knew how big their companies were going to get. All they knew was that they were onto something… Maybe it’s a bad idea to have really big ambitions initially, because the bigger your ambitions, the longer they’re going to take to realize and the long you’re projecting into the future, the more likely you’re going to be wrong.”
PG suggests starting with something small that works instead.
“I think the best way to do these big ideas is not to try and identify a precise point in the future and say, How do I get from here to there? Like the popular image of a visionary. I think a better model is Columbus who thought there was something to the West—I’ll sail westward. Start with something that works, that you know works, that’s small, and then when the opportunity comes to move, move westward. The popular image of a visionary is someone with a very precise view of the future, but empirically it’s probably better to have a blurry one.”
Hannah Mira Cairo (b. 2007, Nassau, Bahamas) specializes in harmonic analysis & Fourier restriction theory.
Homeschooled, she finished calculus by 11 via Khan Academy & self-taught advanced undergrad math by 14. After moving to California, she took UC Berkeley graduate courses.
At age 17 (2025), in Ruixiang Zhang’s class, she disproved the 40-year-old Mizohata–Takeuchi conjecture; a weighted L² inequality for Fourier extension operators on smooth hypersurfaces (open since 1980s); via a fractal counterexample of unexpected wave energy concentration. Now PhD student at the University of Maryland.
Are you a highly competent mid-career professional? Are you trying to break into AI safety/security? Apply to Lateral Workshop by Aug 9! I'll be speaking, alongside some awesome founders, researchers, and grantmakers.
https://t.co/xOt988Gt1h
Claude Code creator Boris Cherny ( @bcherny )
"For people who aren’t building agentic products but are using Claude Code, every 6 months, delete your claude.md file, delete your skills, and delete your hooks. Then see what the model does. It might surprise you.
For Opus 5, we strongly recommend trying to delete all of these things because the model may no longer need the extensive instructions that were necessary for previous models."
- at Y Combinator Startup School 2026.
----
From "Y Combinator" YouTube channel, (full video link in comment)
GDM AGI Safety is hiring! Open roles on all subteams, Lon/Bay/more
There's lots to do to reduce risks from AGI at Google, but we're bottlenecked on people. If you want to help, please apply!
I really love this team and my excellent colleagues. I'm excited to hire more!
Love that all FinTwit people making fun of Leopold still get mogged by his returns this year.
Leopold is a great guy and he's extremely in tune with where the world is headed.