This is more interesting than the title suggests.
Lean4Agent treats agent workflows like programs. Each step has formal preconditions and postconditions, making workflows verifiable before execution and failures traceable to the step where they occurred.
Verified workflows outperform failing ones across SWE-Bench and paper-understanding tasks, and using formal feedback to improve workflows boosts performance further. For long-horizon agents, debugging the trajectory may matter as much as improving the model.
Paper: https://t.co/ZgUOHn8XDa
[5/n]
🔮 This work opens a new direction for how to uniformly model and verify the agent workflow and executions. It can be further applied in dynamic workflows for long-horizon agents, for debugging trajectories, for post-training enhancement, and even for building formal-aware model architectures.
🚀 Introducing the Lean4Agent framework. The first framework 🎉 that unifies the formal verification of agent workflow and executions. It views the agent workflow as “a computer program in natural language.”
🤖 The Lean4Agent provides FormalAgentLib that treats agent workflows like programs. We develop a predicate system that elegantly constrains the agent's execution step via preconditions and postconditions.
✅ This framework makes agent workflows’ self-consistency verifiable before execution, and enables the detection of which step of the workflow breaks the consistency for fine-grained fixes.
🔁 When a verified workflow fails on a specific problem, the predicate system, combined with the LeanEvolve proposed in the Lean4Agent framework. We can dynamically update the workflow to enhance its performance.
[1/n]
Paper: https://t.co/UQAYF7GY6S
Huggingface: https://t.co/WSDjPcq4u3
GitHub Repo: https://t.co/YgRJRZbROz
[4/n]
📝 When the verification provides no additional information, the LeanEvolve can also work, and outperforms the pure LLM-based workflow evolve baseline by 7.00% on average.
🚀 Introducing MA-LoT Theorem Framework: An open-source multi-agent framework utilizing the Long Chain-of-Thought to boost automated theorem-proving🎉
✅ Achieving 61.07% accuracy rate under pass@32 on MiniF2F-Test outperforming Goedel-Prover, Lean_STP and DeepSeek-Prover-V1.5
🔥 Proposing LoT-TL training-inference framework to train LLMs with field-specific Long CoT ability.
🤖 Composing the multi-agent system that combines the advantage of both whole-proof generation and tree-search.
[1/n]
Paper: https://t.co/J9m5i4lNf6
Website: https://t.co/OjQ9P0fYAN
Huggingface: https://t.co/tXsLlxnfvb
Github repo: https://t.co/PZZqk9wAo2
Wonderful collaborators: @rui4research@TheTallEric@shizhediao@RenjiePi@mircale2003@JunjieHu12
[2/n]
🚀 The Scaling Law study demonstrated the capability of our model can be further improved by considering data that are continuously produced by the community
🚀 Introducing Goedel-Prover: A 7B LLM achieving SOTA open-source performance in automated theorem proving! 🔥
✅ Improving +7% over previous open source SOTA on miniF2F
🏆 Ranking 1st on the PutnamBench Leaderboard
🤖 Solving 1.9X total problems compared to prior works on Lean Workbook
[1/n]
website: https://t.co/0lvSAyea9k
huggingface: https://t.co/PDZS9j0PXe
github: https://t.co/mxwOIfMROu
Amazing collaborators: @sangertang1999 (co-first author) @Lyubh22@wujiayun12@hongzhou__lin@KaiyuYang4@JiaLi52524397@xiamengzhou@danqi_chen@prfsanjeevarora@chijinML
[2/n]
🚀 The Scaling Law study demonstrated the capability of our model can be further improved by considering data that are continuously produced by the community
🚀 Excited to introduce TheoremLlama!
🎉 Our new framework transforms general-purpose LLMs into Lean4 experts. Achieving 36.48% and 33.61% on MiniF2F-Valid and Test, surpassing the GPT-4 baseline of 22.95% and 25.41%.
🌟 Check out our open-sourced Open Boostrapped Theorems (OBT) 📜🔍dataset and model checkpoints. Dive into the future of formal theorem proving!
#AI #MachineLearning #Lean4 #TheoremProving #Llama
📑Paper: https://t.co/nkoqGRkZdg
👉 Learn more:
GitRepo: https://t.co/53YH8dQFJT
Model Ckpt: https://t.co/jRTATXQ9F8
OBT Dataset: https://t.co/57vR5tn4ya