Rigorous evaluation of medical AI is good for everyone, and we welcome it. Counter to a half-dozen independent studies from institutions such as the Mayo Clinic that were highly positive on OpenEvidence—a lone paper now purports to show that generalized AI beats specialized clinical AI (@UpToDate, @EvidenceOpen). The paper has a massive undisclosed conflict of interest and irredeemable methodological flaws.
Behind the scenes: The study authors run a competing in-house medical AI at their hospital, and asked OpenEvidence for an API to power it — including rights to build a "competing product" with OpenEvidence's own API. OpenEvidence declined. Then, this paper coincidentally appeared.
Point-by-point, looking closely at the datasets used in the study, the disingenuous and fatal flaws become immediately apparent 🧵.
Ron Green (@rgreenjr), BA ’92, Life Member, took an AI class at @UTAustin in the ’90s. Today, he’s CTO & co-founder of https://t.co/vjShFyS5wJ, leading projects from Fortune 500 AI strategy to breast cancer risk prediction.
Listen now on Hello Longhorn → https://t.co/tOo4XSw5eZ
We are deeply saddened by the passing of Bill Atkinson. He was a true visionary whose creativity, heart, and groundbreaking work on the Mac will forever inspire us. Our thoughts are with his loved ones.
I don't have too too much to add on top of this earlier post on V3 and I think it applies to R1 too (which is the more recent, thinking equivalent).
I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed in AI. You may not always be utilizing it fully but I would never bet against compute as the upper bound for achievable intelligence in the long run. Not just for an individual final training run, but also for the entire innovation / experimentation engine that silently underlies all the algorithmic innovations.
Data has historically been seen as a separate category from compute, but even data is downstream of compute to a large extent - you can spend compute to create data. Tons of it. You've heard this called synthetic data generation, but less obviously, there is a very deep connection (equivalence even) between "synthetic data generation" and "reinforcement learning". In the trial-and-error learning process in RL, the "trial" is model generating (synthetic) data, which it then learns from based on the "error" (/reward). Conversely, when you generate synthetic data and then rank or filter it in any way, your filter is straight up equivalent to a 0-1 advantage function - congrats you're doing crappy RL.
Last thought. Not sure if this is obvious. There are two major types of learning, in both children and in deep learning. There is 1) imitation learning (watch and repeat, i.e. pretraining, supervised finetuning), and 2) trial-and-error learning (reinforcement learning). My favorite simple example is AlphaGo - 1) is learning by imitating expert players, 2) is reinforcement learning to win the game. Almost every single shocking result of deep learning, and the source of all *magic* is always 2. 2 is significantly significantly more powerful. 2 is what surprises you. 2 is when the paddle learns to hit the ball behind the blocks in Breakout. 2 is when AlphaGo beats even Lee Sedol. And 2 is the "aha moment" when the DeepSeek (or o1 etc.) discovers that it works well to re-evaluate your assumptions, backtrack, try something else, etc. It's the solving strategies you see this model use in its chain of thought. It's how it goes back and forth thinking to itself. These thoughts are *emergent* (!!!) and this is actually seriously incredible, impressive and new (as in publicly available and documented etc.). The model could never learn this with 1 (by imitation), because the cognition of the model and the cognition of the human labeler is different. The human would never know to correctly annotate these kinds of solving strategies and what they should even look like. They have to be discovered during reinforcement learning as empirically and statistically useful towards a final outcome.
(Last last thought/reference this time for real is that RL is powerful but RLHF is not. RLHF is not RL. I have a separate rant on that in an earlier tweet
https://t.co/RMIpFPVpuM)
Superintelligence is within reach.
Building safe superintelligence (SSI) is the most important technical problem of our time.
We've started the world’s first straight-shot SSI lab, with one goal and one product: a safe superintelligence.
It’s called Safe Superintelligence Inc.
SSI is our mission, our name, and our entire product roadmap, because it is our sole focus. Our team, investors, and business model are all aligned to achieve SSI.
We approach safety and capabilities in tandem, as technical problems to be solved through revolutionary engineering and scientific breakthroughs. We plan to advance capabilities as fast as possible while making sure our safety always remains ahead.
This way, we can scale in peace.
Our singular focus means no distraction by management overhead or product cycles, and our business model means safety, security, and progress are all insulated from short-term commercial pressures.
We are an American company with offices in Palo Alto and Tel Aviv, where we have deep roots and the ability to recruit top technical talent.
We are assembling a lean, cracked team of the world’s best engineers and researchers dedicated to focusing on SSI and nothing else.
If that’s you, we offer an opportunity to do your life’s work and help solve the most important technical challenge of our age.
Now is the time. Join us.
Ilya Sutskever, Daniel Gross, Daniel Levy
June 19, 2024
Our latest Hidden Layers episode is out now! 🎙️ Get the scoop on Mamba-2, KANs, the OpenAI Superalignment Team breakup, and more. Don’t miss it! 🚀 #AI#TechTalk#Podcast
https://t.co/Z99lFtZiKu
Thrilled to share my chat with the brilliant Scott Aaronson on his work at OpenAI and the challenges of aligning superintelligent AGI (ensuring advanced AI systems act safely and ethically). Scott is thoughtful and hilarious throughout.
https://t.co/sld04RDgXr
The *real* problem with The Force Awakens, The Last Jedi, and The Rise of Skywalker is that any of those three titles could have been used for any of the three films. Very unclear that the one they went with was the most fitting.
Breaking News: Former Gov. John Hickenlooper of Colorado, a Democrat, unseated Sen. Cory Gardner in a race seen as key to Democrats’ hopes of retaking the Senate.
https://t.co/q3VH7cZ5hX
My productivity algorithm is something like:
* Start to do something
* Remember higher-priority thing, switch to doing that
* Remember that high-priority thing is hard
* Do nothing
I’ve been thinking whether in American presidential history there’s ever been an act of cowardice and impotence as brazen as @realDonaldTrump's “I don’t take responsibility at all” comment, followed by his efforts—3+ years into his term—to blame his predecessors. 1/
Watch Fox News host Tucker Carlson call one of his guests a 'tiny brain...moron' during an interview. NowThis has obtained the full segment with historian Rutger Bregman that Fox News is refusing to air.