All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stability, order of accuracy, or physical consistency, even when given very detailed prompting and asked to formalize their implementations in Lean.
Even in cases where they succeed, the more reliable models (such as Fable 5) routinely consume >100x the tokens of a lightweight neurosymbolic model like Lanyon.
Our thesis: only a truly neurosymbolic model like Lanyon is able to produce ultra-reliable numerical solvers for complex scientific problems, with end-to-end correctness guarantees. And at least right, it's not even close. Read more in our latest @lanyon_ai benchmarking post below 👇