Thanks for the shoutout to TLA+ @bcherny !
As suggested by Markus Kuppe, let me showcase the brilliant work done by colleagues at Specula and the TLA+ community. While Specula (https://t.co/hoV6UJIybf) can effectively model system code in TLA+ (and then use model checkers to find hundreds of deep bugs), we can further prove the correctness of system specs using TLAPS. The TLAPS-Bench (https://t.co/hrYHc4DSow) project aims to offer a proof service for any TLA+ spec of important systems/protocols. Basically, Specula writes the spec; TLAS-Bench generates the proof.
AI writes amazing proofs! We had to retire all Proof-Completion tasks in TLAPS-Bench as they're too easy for frontier AI. We also proved classic protocols and algorithms like 2PC, Paxos, TCP, etc.
However, proving low-level specs of real-world system code _from scratch_ is still non-trivial. @qiancheng9788 ran a subset of TLAPS-Bench using Claude Code w/ Opus 5. The success rate is merely 51.4%. @Muse is great, but Muse Spark 1.3 can only achieve 16.7%.
Ruize has a bag of tricks to push AI agents. Once he pushed Codex to prove one of the nine invariants for the ZooKeeper implementation. Codex struggled 5 days, burned $2000, and did it correctly. We ran out of money to continue. And, @xu_dong_sun shows us that liveness proofs are still difficult for AI without guidance. We have an exciting avenue, but still a lot of work to be done.
If anyone wants to verify their protocols/systems, send us some tokens -- we can do it in TLAPS-Bench. @HacksonClark tells me that people are excited about long-horizon tasks -- TLAPS-Bench may be a perfect benchmark for that.
@bcherny Consider further tuning Claude for TLA+, you may find more bugs and/or prove the correctness of SDK code more efficiently.
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted.
Is formal verification the future of coding (or at least, bug finding)?
We released a beta version of Specula, a tool that automatically checks system code using TLA+ based formal methods. Specula instructs coding agents to fully automate formal specification (model + invariants) and mode-code conformance, which were major barriers to adopt formal methods in practice. We find Specula rather useful in finding deep bugs that require formal reasoning. (Specula found hundreds.)
If you need automatic formal reasoning of your code, give it a try. (The tool is push-button.)
Code: https://t.co/gcqFUHYGL6
Paper: https://t.co/4W2p1yX5ia
(While I'm posting it, the students, especially @qiancheng9788 and @SaadMRP, and Ruize Tang did all the hard technical work.)
CC @disalg_spec@Yiming_Su3
Last fall, we wrote a small benchmark called SysMoBench when developing Specula. SysMoBench asks AI to model current, distributed systems in TLA+. @qiancheng9788 and the team wrote about the experience and updated the results on newest LLMs. The abilities of frontier models today are rather impressive and our Specula agent can get 100% on the benchmark now.
@qiancheng9788 put up a nice website: https://t.co/U5HilIV354
New SIGOPS Blog -- "Can LLMs model real-world systems in TLA+?" by Qian Cheng, Ruize Tang, Emilie Ma, Finn Hackett, Peiyang He, Yiming Su, Ivan Beschastnikh, Yu Huang, Xiaoxing Ma, and Tianyin Xu.
https://t.co/nCEUg3zFV2
The main problem with modern-day universities are the lies. We have been lying so much that we cannot think clearly anymore. Here are a few lies.
- « Professor X is an an expert in ZYX » where ZYX is some socially relevant topic. For example, professor X is an expert in « disinformation » or « entrepreneurship » or « software engineering ». We have so abused the term « expert » that we no longer even know what it means. An expert is someone who has a lot of experience at solving a specific kind of problems. Very few professors become at expert at anything but: teaching, writing papers, sitting on committees. You simply cannot be an expert at entrepreneurship if you have never been an entrepreneur. At best, you can a scholar of entrepreneurship, in the sense that you know a lot about what has been written on the topic of entrepreneurship... but that's not the same.
- « If you stand in front of a class for 3 hours a week, talking to the class about topic X, you are transmitting your knowledge about topic X. » Learning is an action. It is something you do as you act. You learn to speak English by speaking English. You learn to code in Java by coding in Java. Want to become a philosopher? Then find someone who is eager to discuss deep issues and talk with them. Stop attending lectures, stop reading alone. You are not going to understand Socrates by reading Plato. You need to « work through the material ». Join a book club. Listening passively gives you the illusion of knowledge, but, when tested, you are found out. You can help students learn by giving them tasks, by providing examples that they can follow... but you cannot teach merely by talking to someone. I have spent hours, for years, « talking to » some graduate students, only to find out that they had understood none of what I said. The ones who did « get it » worked « with me » on projects.
- « Students are blank slates that you can build up by having them take courses. » So you take John, you have John pass 20 courses in computer science or philosophy, and John *is* then an expert in computer science or philosophy. If someone tells you that they have a degree in X, it tells you almost nothing about what they can do. My two sons learned to program software on their own, without my help, years before they ever had programming courses. They are not great basketball players nor musicians, but they can code circles around most people... Even if Jack had three degrees in computer science, my sons would probably, on average, beat Jack at a programming competition. Life is not fair, people are not blank slate.
- « Universities are independently-minded institutions. They cultivate critical thinking. » Modern universities are among the most conformists institutions you can find. They immediately bow to the dominant intellectual paradigms, whatever it is. The best way to succeed on campus is to immediately adopt the new dominant paradigms and to stay clear of new ideas that have not been vetted yet. Almost everything goes through some form of « peer review » which means that ideas must be accepted by your peers before they can be considered. So if you think differently, you are handicaped. That's by design.
The final and wort lie is that universities are dedicated to the pursuit of truth. There is no such dedication to truth. Universities will lie, misreport, misrepresent if it serves their goals.
But here is the cool thing about universities... there are bubbles, small groups, individuals here and there, that are dedicated to higher ideals. There is a long tradition going back to Thomas Aquinas, of crazy folks who insist on pursuing truth.
What a #kubecon, what a week!
In case you missed it, we announced a partnership with the @CloudNativeFdn to ensure that the open source software that the world depends on stays dependable. Watch the announcement below!