@KaiyuYang4 Good post. Agree on most observations. Formal isn’t new and has been around since 1960’s.
My own experiences in the last decade+ creating products on neurosymbolic base at leading companies is inline with this post. Thanks for sharing
@hillelogram Modeling the properties of the model under verification, modeling the properties of the environment (constraints) remains a challenge in terms of accuracy - with/without LLM’s.
If the modeling is incorrect, results will be suspect. GIGO applies
It’s because of multiple reasons - learning a new language with strict semantics, modeling real world isn’t trivial, thinking in terms of what shouldn’t happen (alongside than what should happen), modeling the environment constraints (the hardest)….
Have had first hand experience and success with formal methods in hardware verification. Companies like Nvidia , Apple have large formal groups.
@SebastienBubeck 😑ok this is kinda weak because it’s focused on personal exchange and “intention” but doesn’t address the core allegation about data access at all…
🦔An NYU mathematician says OpenAI used his own progress against him to beat him to one of the biggest unsolved problems in mathematics. Tristan Buckmaster had been working toward a Millennium Prize proof using OpenAI's Codex when information about his progress reached OpenAI.
Days later, OpenAI published a full proof of the same problem using the same uncommon approach, after burning $22.5 million in compute to get there. When Buckmaster confronted them, exec Sébastien Bubeck allegedly said "Why would you ruin your career?" and "If you don't want me to be nice, then I don't have to be nice."
My Take
OpenAI spent $22.5 million to solve a problem with a $1 million prize. They didn't do this for the bounty. They needed a headline that says "our AI solved one of the hardest problems in mathematics" and they needed it before someone else got credit. They started days after they heard about Buckmaster's progress and took the same uncommon approach he'd pursued for months. That's hard to explain as coincidence.
Buckmaster did his work inside Codex. OpenAI reserves the right to train on Codex data. They admit they can't rule out that his usage helped improve their models. So a customer used their product, potentially handed them the roadmap, and then OpenAI outran him with $22.5 million in compute he could never match. I don't know if any of this was intentional. But if you're a researcher and you just watched this happen, I don't think you'd keep your best ideas inside someone else's product.
Hedgie🤗
https://t.co/1YnggU7xdI
@rynorhn Not surprised. Anthropic and OpenAI are definitely using sessions chats to train their models. Stop using these if you are working on innovative , differentiated applications