Disclaimer: I had given early access to internal beta version of Grok 4.20
It found a new Bellman function for one of the problems I’d been working on with my student N. Alpay.
The problem reduces to identifying the pointwise maximal function U(p,q) under two constraints and understanding the behavior of U(p,0).
In our paper https://t.co/pgJw9MaEA1 we proved U(p,0)\geq I(p), where I(p) is the Gaussian isoperimetric profile, I(p) ~ p\sqrt{log(1/p)} as p ~ 0.
After ~5 minutes, Grok 4.20 produced an explicit formula U(p,q) = E \sqrt{q^2+\tau}, where \tau is the exit time of Brownian motion from (0,1) starting at p. This yields U(p,0)=E\sqrt{\tau} ~ p log(1/p) at p ~ 0, a square root improvement in the logarithmic factor.
Any significance of this result? It will not tell you how to change the world tomorrow. Rather, it gives a small step toward understanding what is going on with averages of stochastic analogs of derivatives (quadratic variation) of Boolean functions: how small can they be?
More precisely, this gives a sharp lower bound on the L1 norm of the dyadic square function applied to indicator functions 1_A of sets A \subset [0,1].
In my previous tweet about Takagi function, we saw that the sharp lower bound on ||S_1(1_A)||_1 miraculously coincides with Takagi function of |A| which (surprisingly to me) is related to the Riemann hypothesis. Here, we obtain a sharp lower bound on ||S_2(1_A)||_1 given by E \sqrt{\tau}, where Brownian motion starts at |A|. This function belongs to the family of isoperimetric-type profiles, but unlike the fractal Takagi function, it is smooth and does not coincide with the Gaussian isoperimetric profile.
Finally, in harmonic analysis it is known that the square function is not bounded in L^1. The question here was more about curiosity: how exactly does it blow up when tested on Boolean functions 1_A. Previously, the best known lower bound was |A|(1-|A|) (Burkholder—Davis—Gandy). In our paper, we obtained |A| (1-|A|)\sqrt{log(1/(|A|(1-|A|)))}. This new Grok’s Bellman function gives |A| (1-|A|) \log(1/(|A|(1-|A|))) and this bound is actually sharp.
You are missing a lot not by leveling up your skills in 2025
𝐈𝐁𝐌 𝐢𝐬 𝐨𝐟𝐟𝐞𝐫𝐢𝐧𝐠 𝐲𝐨𝐮𝐫 𝐟𝐫𝐞𝐞 𝐜𝐨𝐮𝐫𝐬𝐞𝐬 𝐰𝐢𝐭𝐡 𝐚 𝐜𝐞𝐫𝐭𝐢𝐟𝐢𝐜𝐚𝐭𝐞.
No payment is required!
[ Bookmark for future 🔖]🧵
🩺 Get started with MedGemma, a collection of Gemma 3 variants built for medical text and image comprehension. Choose between the 4B multimodal model or 27B text-only model to accelerate your healthcare AI projects ↓
https://t.co/iANwgjBLry
📣 Proud to share HealthBench, an open-source benchmark from our Health AI team at OpenAI, measuring LLM performance and safety across 5000 realistic health conversations. 🧵
Unlike previous narrow benchmarks, HealthBench enables meaningful open-ended evaluation through 48,562 unique physician-written rubric criteria spanning several health contexts (e.g., emergencies, global health) and behavioral dimensions (e.g., accuracy, instruction following, communication).
Blog, paper, code: https://t.co/NsSPeIoHZy
Scientific discovery with LLMs has so much potential yet is underexplored. Our new benchmark **LLM-SRBench** enable rigorous evaluations of equation discovery with LLMs!
🧠Key takeaway: Even SOTA discovery models with strong LLM backbones still fail to discover mathematical relations for a scientific problem. Introducing a new dataset designed to leverage LLM capabilities toward driving innovation instead of finding known solutions is a step forward to enhance these discovery agents.
🧵👇
The 𝕏 recommendation algorithm is being replaced with a lightweight version of @Grok, so will soon be dramatically better!
You should notice some improvement already.
🧵 1/ How well do LLMs actually do on Olympiad-level math?
We evaluated frontier models on 455 problems from the IMO Shortlist.
Unlike most benchmarks, we emphasize proof validity, not just final answer correctness.
Here’s what we found 👇
Can we build AI research agents to perform long-horizon tasks like ML engineering tasks e.g.Kaggle?
Introducing our new work MLAgentBench: Benchmarking Large Language Models as AI Research Agents!
Three components of Reasoning for AI:
1. Foundation (Pre-training)
2. Self-improvement (RL)
3. Test-time compute (planning).
@xai will soon have the best foundation in the world - Grok3. Join us to advance reasoning to the next-level! 🔥🔥
https://t.co/nIIlLjb1je
Many studies suggest AI has achieved human-like performance on various cognitive tasks. But what is “human-like” performance?
Our new paper conducted a human re-labeling of several popular AI benchmarks and found widespread biases and flaws in task and label designs. We make 5 concrete recommendations for future AI Benchmarks, drawing inspirations from decades of studies in cognitive science.
#AI #benchmark #cognitivescience #intelligence
BREAKING: Elon just announced that Grok 3 is FREE.
For now, at least..
No clue how long this will last, so if I were you, I’d jump on it ASAP.
I’ve created 500 solid Grok 3 prompts for you to play with.
Right now, it’s 100% FREE.
To get it, simply Like & Reply 'G3' and I'll send it to you via DM.
AI will determine the future of WORK, BUSINESS, and LIFE.
And there are huge questions:
Will AI unleash our creativity and help us to flourish?
Or will we lose even more power, agency, and ownership?
It all depends on the design of the tech infrastructure.
As @cdixon often says, “Architecture is destiny.”
And how AI systems are designed will be key to the future we all live in.
The team at @gensynai is building a protocol that connects every device in the world into an open network for machine intelligence, with no gatekeepers or artificial boundaries.
They are shaping a future of AI that is good for all of us.
At a recent Innoveum event, I chatted to @benfielding cofounder and CEO of @GensynAI about what they’re building, how it transforms AI, and why we should all be relieved, excited, and get involved.
The full conversation is below in Episode One of Innoveum | The Podcast. 👇
Are you struggling to pay a huge amount on paid courses?
I'm giving you access to 20+ FREE Courses
1. Artificial Intelligence
2. Machine Learning
3. Cloud Computing
4. Ethical Hacking
5. Data Analytics
6. AWS Certified
7. Data Science
8. BIG DATA
9. Python
10. MBA
To get it, just:
1. Like & Retweet
2. Comment "ALL"
3. MUST be Following (so that I can dm)