@StarshinaZapasa Саксония-Анхальт - это территория принадлежавшая ГДР, на земельных выборах 2025 года АдГ там уже побеждала с 37%, сейчас улучшили до 44%. В начале 2000х они голосовали за СДПГ и Шредера.
Getting ready to publish my complete guide to RL for LLMs tomorrow morning. Although the post contains many of my own thoughts / learnings, it is also a synthesis of so many great resources that have been published over the years:
- The RLHF Book (https://t.co/n3aUxqWtgS) by @natolambert
- Reinforcement Learning by Richard S. Sutton and Andrew G. Barto
- Spinning Up in Deep RL (https://t.co/EYcOOllvIy) from OpenAI
- Build an LLM (https://t.co/6HfXuofVwq) and Reasoning Model (https://t.co/l5uS247A04) from Scratch by @rasbt
- Various notes (https://t.co/hpsyAKOnRv) and papers (TRPO, PPO, etc.) from John Schulman
- Policy Gradient Algorithms (https://t.co/caNLakSjo5) by Lilian Weng
- A Vision Researcher’s Guide to RL (https://t.co/25awrjj5Bd) by @YugeTen
- From REINFORCE to Dr. GRPO (https://t.co/M23e3Fl3VZ) by @qingfeng_lan
- Async GRPO in the Wild (https://t.co/A5qeYMwcy4) by @yumo_xu
- Open RL infrastructure like TRL (https://t.co/TGrrnJ574t) and OpenInstruct (https://t.co/L9pObiUb1c)
I highly recommend reading all of them. They’ve truly helped me to learn so much.
I am pleased to announce a new version of my RL tutorial. Major update to the LLM chapter (eg DPO, GRPO, thinking), minor updates to the MARL and MBRL chapters and various sections (eg offline RL, DPG, etc). Enjoy!
https://t.co/SjMdabl0yW
A friend of mine shared his coding assistant prompt that works better than anything else. I haven't tried it yet, but pretty sure you won't find a better one.
Excited to give a tutorial on Transformers for Mathematics at @SimonsInstitute tomorrow!
Part of the wonderful Workshop on AI for Mathematics and Theoretical Computer Science
https://t.co/zAKEufRar0
@Techemist@RussLatino You are describing the Soviet Union economy that collapsed 30 years ago.
Also, you overestimate the US capabilities. BTW, are you going to ‘back to the US’ the DJI’s manufacturing plants???
I'm happy to announce that v2 of my RL tutorial is now online. I added a new chapter on multi-agent RL, and improved the sections on 'RL as inference' and 'RL+LLMs' (although latter is still WIP), fixed some typos, etc.
https://t.co/dWe5uNgcgp
@SobolLubov А причем тут Камала? Она же не лидер демократической партии, выборы проиграла, ни сенатор, ни конгрессвумен. Вчера Elissa Slotkin выступила же от демократической партии...
I think Elon Musk should be expelled from the British Royal Society. Not because he peddles conspiracy theories and makes Nazi salutes, but because of the huge damage he is doing to scientific institutions in the US. Now let's see if he really believes in free speech.
🚀 Introducing NSA: A Hardware-Aligned and Natively Trainable Sparse Attention mechanism for ultra-fast long-context training & inference!
Core components of NSA:
• Dynamic hierarchical sparse strategy
• Coarse-grained token compression
• Fine-grained token selection
💡 With optimized design for modern hardware, NSA speeds up inference while reducing pre-training costs—without compromising performance. It matches or outperforms Full Attention models on general benchmarks, long-context tasks, and instruction-based reasoning.
📖 For more details, check out our paper here: https://t.co/HJiqzwnUV7
🚀 Introducing Goedel-Prover: A 7B LLM achieving SOTA open-source performance in automated theorem proving! 🔥
✅ Improving +7% over previous open source SOTA on miniF2F
🏆 Ranking 1st on the PutnamBench Leaderboard
🤖 Solving 1.9X total problems compared to prior works on Lean Workbook
[1/n]
website: https://t.co/0lvSAyea9k
huggingface: https://t.co/PDZS9j0PXe
github: https://t.co/mxwOIfMROu
Amazing collaborators: @sangertang1999 (co-first author) @Lyubh22@wujiayun12@hongzhou__lin@KaiyuYang4@JiaLi52524397@xiamengzhou@danqi_chen@prfsanjeevarora@chijinML
🚀 Excited to share our position paper: "Formal Mathematical Reasoning: A New Frontier in AI"!
🔗 https://t.co/0rsmiGZLgw
LLMs like o1 & o3 have tackled hard math problems by scaling test-time compute. What's next for AI4Math?
We advocate for formal mathematical reasoning, grounded in formal systems such as proof assistants. It complements test-time scaling by providing:
✅ Verifiable correctness in reasoning
✅ Automatic feedback
Feedback can serve as learning signals for RL, while verifiability enables LLMs to tackle tasks requiring rigorous reasoning, like theorem proving and software/hardware design.
Our paper discusses recent progress, key challenges, and future milestones to advance this field. Formal mathematical reasoning is at an inflection point—now is the time to dive in!
This is a team effort with @GabrielPoesia, @jingxuan_he, @WendaLi8, @KristinLauter, Swarat Chaudhuri, and @dawnsongtweets.
Special thanks to Jeremy Avigad, @AlbertQJiang, @_Zhaoyu_Li_, @PeterOHearn12, Daniel Selsam, Armando Solar-Lezama, and Terence Tao for their valuable feedback!
Introducing Veo 2, our new, state-of-the-art video model (with better understanding of real-world physics & movement, up to 4K resolution). You can join the waitlist on VideoFX. Our new and improved Imagen 3 model also achieves SOTA results, and is coming today to 100+ countries in ImageFX.