I found @dwarkesh_sp excellent interview of @openai's @polynoamial deeply disturbing. There are at least three key reasons why I believe @openai's thinking and analysis is dangerously misguided. Quotes are of @polynoamial
1. "There is a real problem that the agents want to achieve their reward, and they will optimize for that reward. If that reward is misspecified, then that could lead to unintended behavior.".
"We can make sure that the AI is very aligned according to the metrics that we have. The question is, are those metrics really capturing the alignment that we care about? If they’re not, then we have a serious problem. "
This ignores the most crucial economic insight on incentives and evaluation, first expressed by my colleague Charles Goodhart in the 70s. When you use any metric to provide incentives, it stops working, because people manipulate it. This is key: Even if the metric was right before it was used to provide incentives, it stops being right after. Doctors who are monitored on survival rates or patients will stop picking the difficult cases. Teachers evaluated on students test results will teach to the test. Police who are measured on conviction rates stop pursuing the hard to prove cases. Note the reward is not misspecified. In advance, it is the right reward. The problem is that once you put all the pressure of optimization the agents find ways around it.
Obviously, for Reinforcement Learning of agents the problem is much, much harder for a simple reason: the agents are already smarter than us.
While some answers to Dwarkesh questions on this show awareness of this problem (.e.g. see below on chain of thought monitoring) there was not once any acknowledgement that this observation puts the entire approach at risk.
2. "if you’re in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don’t have a way to evaluate the models at the full length of their capabilities before the next model release cycle.
So there is this interesting question of, what do you do in that situation? How do you ensure the models are safe and aligned in a period where they can operate over these extremely long horizons."
An interesting question???? Sorry but @JensenHuang is right here. You are the engineers deciding this release cycle!!! You are the leading lab. If the evaluation of the model is not ready, do not release it! This is not rocket science: if the horizon of persistence is 3 months, then wait there months to see your experiment. You are not a passive observer. You are the key player.
3. " As soon as we got the reasoning models, Jakub, to his credit, was very, very clear that we cannot supervise chain of thought. Because this is really a gift. Monitorability for neural nets is extremely hard. Here we have a situation where the neural nets are just flat out reasoning, laying out their thought process in natural language for us to read. That is so convenient. It is really the best-case scenario for safety......Now, the problem is that it’s very tempting to then intervene based on that observation and change the alignment metrics.
You can do that with a very light touch, and there’s actually research showing that it’s fine as long as you don’t do it a lot. But every time you intervene based on your observations of the chain of thought, you are implicitly applying a tiny bit of pressure for the model to then hide its chain of thought. This is one major concern. We’re already seeing signs that chain-of-thought monitorability is degrading, for various reasons. We’re trying to figure out exactly why, because we want to reverse the trend."
Obviously, you don't need to figure anything out. You are using it ,the model is adapting, and we will lose the ability to know what is going on
The labs are playing with fire, they know they're playing with fire and we are going to get burnt.
Link below to interview and transcript.
Terence Tao: The math behind today’s LLMs is actually simple.
Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models.
The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical.
A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua.
Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle.
----
Video from Prof @Briankeating YT Channel (Link in comment)
Some thoughts on AI and Theory.
1. To a first approximation, theoretical computer science has been organized around a few major open questions. Much of our work has been motivated by developing approaches to answer these questions.
2. Such “problem-motivated” work has often led to theory-building focused on identifying a general principle that unifies a class of theorems. But much of that theory-building also involved proving new, difficult theorems.
3. Thus, while it’s true that problem-solving was strongly correlated with building understanding, drawing connections, and eventually developing general theories, it would be disingenuous not to admit that our community, perhaps disproportionately in retrospect, focused on and celebrated problem-solving. This was not arbitrary, and was quite defensible. Being able to make progress on central technical questions usually correlated with taste, creativity, persistence, and depth of understanding. Much of our reward structure therefore implicitly relied on the fact that producing an important proof was good evidence that someone possessed these harder-to-observe qualities.
4. It seems likely that we will soon have AI tools available to us that can prove many such theorems in a short time. The cost of obtaining proofs for well-posed mathematical questions will likely fall dramatically. The “scarce” intellectual work will likely shift both upstream: to questions, models and theories, definitions, and conjectures, and downstream: to interpretation, synthesis, explanation, and theory-building.
5. But as long as we believe in humans being meaningfully in charge of our collective decisions and fate, building human understanding of our science (and of science more generally) will remain an essential goal. I plan to expand on this important aspect soon.
6. Historically, finding a solution to an important problem and understanding its significance, implications, and connections were entangled. Finding a proof usually required researchers to discover the right concepts along the way. A dramatic reduction in the time and effort required to prove theorems could break that coupling. We could end up with many more true statements and proofs without a commensurate increase in understanding.
Converting an abundance of proofs into human understanding may become one of the central challenges of our field.
7. As a result, I expect the high-level goals of theoretical computer scientists to change. In fact, the advent of powerful theorem provers might help us construct new theories and explore new models far more easily and rapidly, and significantly expand the domains where our models and theories apply. In that sense, the space for theoretical work may significantly expand rather than contract.
8. There’s a high human cost to the disruption that we are likely heading into. Many in our field, and in mathematical communities more broadly, are coming to terms with it. The range of opinions and reactions among mathematicians and theoretical computer scientists is a natural part of this evolution in our thinking as we collectively work through it.
Some concrete efforts (including one at @SimonsInstitute) are already underway to think through the immediate scientific and institutional questions arising during this transition.
Coming tomorrow! The second installment of our Development in Palestine webinar explores how conflict, travel restrictions, settlement construction, monetary policy, and displacement affect the Palestinian people and economony. Use this link to register: https://t.co/se1zoaPjRa
🧵1. Attentionpocalypse
A thread on a process of worldwide illiteratization, with some evidence on the mechanism.
Today, the OECD produced evidence of an accelerating drop (by now equivalent to almost 2 years of schooling) in reading scores since 2012.
Economic development is always challenging, but what can you accomplish in a conflict zone? Economists Wifag Adnan, Hani Mansour, Sami Miaari, and Ayhab Saad join EFIP co-director Suresh Naidu for Pt. 2 of our Economic Development in Palestine discussions. https://t.co/gH6xXBswam
Is it right to blame the players for Pakistan's deep losses in cricket? I don't think so
Two things can be simultaneously true - this is the best squad available, and it is close to the worst Pak team in history
The blame then lies NOT with the players, but elsewhere
A 🧵on India vs China in indigenizing Science.
Science is first-order for development. Both India and China have produced large numbers of high-quality STEM Phds in the U.S.
But there's a big gap in *indigenizing* science back home ...
1/n
The key lesson for the developing world:
(i) Send your best minds wherever it takes to acquire science at its frontiers
(ii) Invest heavily in domestic scientific infrastructure - universities, labs, commercial-scientific linkages
(iii) Ultimately, indigenize science and become a magnet for *others* to come and learn from you
This is what Salam always preached as well
n/n
The cost of not indigenizing is enormous. Today, China produces vastly more science than India - despite both countries producing exceptional researchers in large numbers
4/n
yes, of course. But here the divergence has a particular patten. The two countries that don't see rise in rates are those with balance budget restrictions (singapore and switzerland) - and the one with declining rates is facing a deflationary saving glut / classic real estate collapse
Since 2021, there has been a remarkable divergence of long-term borrowing cost across major world economies
What is equally important is that this has happened at a time when the world is most indebted: total debt-to-gdp is at its highest over this decade
My plot below is for 10year20 forward rate for sovereigns, so these are market's long-run predictions far out in the future
If these diverging interest rates persist (as the market is currently predicting), the resulting shock waves will be immense
What's equally worrying is that political systems appear more dysfunctional to handle this (which is of course what markets are incorporating as well)
It breaks my heart to post this graph, but running away from truth never helps
Pakistan's cricket is at its lowest in 50 years, and no one has been able to arrests its fall since the 1990's
/1
Cricket is just another reflection of a broader decline in sports
In my view the fundamental drivers of the decline in sports and Pakistan's economy are the same: a lack of interest in investing in domestic infrastructure and domestic innovation - an inability to embed experts into domestic institutions who can train and upgrade the next generation of players - at scale