FlexAttention is a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code @Boyuan_Feng & @__avik show how many existing attention variants can be implemented via FlexAttention & that we achieve competitive performance compared to handwritten kernels. 📺 Watch the full video on YouTube: https://t.co/oRi5myp9zu
OpenAI's o3 LLM to discover CVE-2025-37899, a remote zero-day use-after-free bug in the Linux kernel’s SMB (ksmbd) implementation.
o3 autonomously reasoned about 12k LoC and uncovered a bug missed by human review
-----
→ o3 found this bug in the SMB logoff handler, where sess->user is freed without synchronization, allowing other threads to still access the freed memory, leading to potential kernel code execution.
→ The test used no scaffolding or agents—just direct API calls. Code context was expanded to ~12k LoC (~100k tokens) for the test that discovered the new CVE.
→ In earlier tests, o3 identified another use-after-free (CVE-2025-37778) 8 out of 100 times with a 1:4.5 true-to-false positive ratio. Claude 3.7 only found it 3 times; Claude 3.5 missed entirely.
→ o3 outperformed Claude in clarity and structure of its bug reports, at times even outperforming its human operator in reasoning about patch sufficiency.
→ These results suggest LLMs, when used right, can dramatically boost effectiveness in vulnerability discovery—even in kernel-level codebases.
Tina: Tiny Reasoning Models via LoRA
"the best Tina model achieves a >20% reasoning performance increase and 43.33% Pass@1 accuracy on AIME24, at only $9 USD post-training and evaluation cost (i.e., an estimated 260x cost reduction). Our work reveals the surprising effectiveness of efficient RL reasoning via LoRA."
o3 and o4-mini are super good at coding, so we are releasing a new product, Codex CLI, to make them easier to use.
this is a coding agent that runs on your computer. it is fully open source and available today; we expect it to rapidly improve.
I've officially switched away from Cursor.
My new stack:
- Cline (w/ Sonnet 3.7) for frontend
- RepoPrompt (w/ o1 pro) for backend
Accuracy and quality are just so much higher when you're not trying to save $ by compressing to fewer tokens.
The AI world just got disrupted by Optimus Alpha, a SECOND mystery model that's outperforming almost everything.
Is it OpenAI's secret new release or something else entirely?
Let's investigate 🧵
Github 👨🔧: Awesome-GraphRAG: A curated list of resources (surveys, papers, benchmarks, and opensource projects) on graph-based retrieval-augmented generation.
-------------
→ Provides a categorized list of research papers on GraphRAG, structured according to a comprehensive survey.
→ Covers key areas within GraphRAG including knowledge organization, knowledge retrieval, and knowledge integration techniques.
→ Offers a comparison between traditional RAG and GraphRAG, clarifying the advantages of graph-based approaches.
Manus, the new AI product that everyone's talking about, is worth the hype.
This is the AI agent we were promised.
Deep Research+Operator+Computer Use+Lovable+memory.
Asked it to "Do a professional analysis of Tesla stock " and it did ~2wks of professional-level work in ~1hr!