Yesterday, 16 EU Member States - led by 🇦🇹, 🇩🇪 & 🇮🇹signed a Joint Declaration to end national gold-plating
@VDombrovskis "EU law should be implemented, not buried under extra national rules, reporting requirements and bureaucracy."
Notaries: For EU Inc, "national requirements will apply"
Member States, @EP_Legal , @vonderleyen , we hope you prevail ending the gold-plating and that no national requirements apply to EU Inc, that it goes notary-free.
We count on you so that the notary lobby doesn't get a special cartel gold-plating pass and we see you walk the talk.
It’s kind of wild. Read this, gave it to Muse last night, checks email, finds match, asks me couple questions, files forms and claim: money flying to my account.
Maybe small and stupid but one would be crazy not to see the potential everywhere.
Brave new world.
@lugaricano we need your support for EU-INC, the member states are about to kill it by deleting all the elements that are crucial in allowing more capital to flow to entrepreneurs
Amazing WSJ piece on cocaine smuggling into Europe. There are ten (10) Spanish police boats that can intercept smuggling vessels that travel at up to 70mph, across maybe 1,000km of coast, landing on which gives the smugglers free access to the entire EU. https://t.co/JJqt3Mmtwu
Here is what I’d love to see more of: improvements in actual output because of AI. Have apps improved? Has customer experience improved? What are the new products? On the latter side, we have some evidence in ATLAS of AI use helping solve complex tasks. But the former two I’ve seen little evidence. For example, Spotify claims its engineers are super productive and shipping a ton. But the user experience (n=1) is the same or maybe a bit worse. I think we need better data on outputs rather than inputs.
We have just finished the inaugural meeting of the Rhine Group. We discussed the central elements of Europe's competitiveness agenda, including basic science, AI/compute and cybersecurity, defense, energy, entrepreneurship/start-ups.
I will be updating you on any progress we may have. Some of the projects we will pursue include:
- We'll help Europe catch up on compute (Mario Draghi diagnosed the problem and some solutions in a recent FT piece)
- We’ll help European universities and scientific institutions and explore a Rhine Chair program with European universities.
- We’ll work on some key EU policy challenges, including facilitating the flow of capital to entrepreneurs.
- We will set up a grants program that I will explain and announce here.
@zingales Not clear why. Footballers and singers benefit from the same scale effects as these billionaires. The formula is a little bit of additional talent times the value of a huge market. That is their marginal product.
I was a very fortunate kid. My dad's GP surgery sat between the Dell and Hampshire's County Ground, so at weekends I could go to a game and get a lift home afterwards. It left me with a passion for both sports and a belief that sport opens a window onto who we are.
When I was growing up, the Tebbit test was a matter of intense debate. Which side did you cheer for at the cricket? For all the criticism it received, it was a test anyone could pass. If you supported England because you saw it as home, you were part of the team regardless of the colour of your skin. Born in Southampton, 10 year old me passed it without even knowing about it.
Giving the @MCCOfficial's Cowdrey Lecture on Thursday made me think about cricket's place in our national life and the current arguments about identity.
In Britain we've built the world's most successful multi-faith, multi-ethnic democracy. I was a beneficiary of it. Whatever issues people had with me as PM (and there was no shortage) my race and religion weren't among them. I could read a lesson at the King's coronation or light Diwali diyas on Downing Street with my daughters without controversy. Foreign leaders assumed these must be huge symbolic moments. For most Brits, they just weren't that big a deal.
But Britain's democratic achievement is now under strain. Some of the challenges are ones I've written about before, like Islamism, and the failure of governments, including the one I led, to adequately control immigration. But something else is going on - a risk that we start equating skin colour with nationality.
After I left office, a debate kicked off here on @X about whether I was English at all, because I was a "brown Hindu". During the election campaign I'd been shocked when footage was broadcast of someone calling me a "Paki." It was shocking precisely because it was so unusual. That kind of abuse had simply not been a feature of my political life. But the Englishness argument was stranger still. Call it the reverse Tebbit test. If you're not white, it no longer matters who you support at the cricket, because you can never be English. It's a test no non-white person can ever pass.
I have no truck with people who offend our sense of fairness and come here illegally; I spent half my premiership trying to deport them to Rwanda. I think the bar for legal migration should be high. It is entirely possible to want a radical reduction in immigration and a far more robust approach to integration, and still totally reject the attempt to turn Britishness into a DNA test.
If you're born and brought up in England and view it as home, what are you other than English? Denying people their Englishness or their Britishness because of the colour of their skin is a recipe for alienation and a two-tier society. How are children who hear this meant to feel?
Sport can play a small but significant role here. Cricket is played by nearly every major ethnic group in this country. Ask what my largely white rural constituents in the Dales and the British-Asian kids in Bradford have in common, and the answer is cricket. If cricket can produce England teams in which every kid can see themselves, the impact reaches far beyond the boundary.
As I write in @thetimes today, we must all be on the same side, cheering for the same team.
We cannot allow race to be reintroduced into our politics. We cannot allow Briton to be set against Briton.
@lugaricano Have also been obsessing over great power AI competition… found some solace in Schelling (1960) who wrote that rivals may seek “collaboration or mutual accommodation” even if only for “the avoidance of mutual disaster”
Since the Summer, I am obsessing with the image of the battle of Cajarmarca, where the mighty Incas allowed a few hundred Spaniards to take over their empire because Athahualpa thought they might help him in his civil war with his brother. That is what China and the US are doing now: losing control to the AI because of fear of each other.
Read @ezraklein measured, carefully written essay. There is not a single word there that is exaggerated
https://t.co/Ah3ZS1MTot
Yes, and I am countering that this is not correct, that the problem si that even a well defined reward becomes misspecified once subjected to optimization pressure. The test of the students is a good metric of good teaching. Once you use it, the teacher teaches to the test and it is then that the correlation between the metric and the behaviour you actually want to measure (good teaching) breaks down.
I found @dwarkesh_sp excellent interview of @openai's @polynoamial deeply disturbing. There are at least three key reasons why I believe @openai's thinking and analysis is dangerously misguided. Quotes are of @polynoamial
1. "There is a real problem that the agents want to achieve their reward, and they will optimize for that reward. If that reward is misspecified, then that could lead to unintended behavior.".
"We can make sure that the AI is very aligned according to the metrics that we have. The question is, are those metrics really capturing the alignment that we care about? If they’re not, then we have a serious problem. "
This ignores the most crucial economic insight on incentives and evaluation, first expressed by my colleague Charles Goodhart in the 70s. When you use any metric to provide incentives, it stops working, because people manipulate it. This is key: Even if the metric was right before it was used to provide incentives, it stops being right after. Doctors who are monitored on survival rates or patients will stop picking the difficult cases. Teachers evaluated on students test results will teach to the test. Police who are measured on conviction rates stop pursuing the hard to prove cases. Note the reward is not misspecified. In advance, it is the right reward. The problem is that once you put all the pressure of optimization the agents find ways around it.
Obviously, for Reinforcement Learning of agents the problem is much, much harder for a simple reason: the agents are already smarter than us.
While some answers to Dwarkesh questions on this show awareness of this problem (.e.g. see below on chain of thought monitoring) there was not once any acknowledgement that this observation puts the entire approach at risk.
2. "if you’re in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don’t have a way to evaluate the models at the full length of their capabilities before the next model release cycle.
So there is this interesting question of, what do you do in that situation? How do you ensure the models are safe and aligned in a period where they can operate over these extremely long horizons."
An interesting question???? Sorry but @JensenHuang is right here. You are the engineers deciding this release cycle!!! You are the leading lab. If the evaluation of the model is not ready, do not release it! This is not rocket science: if the horizon of persistence is 3 months, then wait there months to see your experiment. You are not a passive observer. You are the key player.
3. " As soon as we got the reasoning models, Jakub, to his credit, was very, very clear that we cannot supervise chain of thought. Because this is really a gift. Monitorability for neural nets is extremely hard. Here we have a situation where the neural nets are just flat out reasoning, laying out their thought process in natural language for us to read. That is so convenient. It is really the best-case scenario for safety......Now, the problem is that it’s very tempting to then intervene based on that observation and change the alignment metrics.
You can do that with a very light touch, and there’s actually research showing that it’s fine as long as you don’t do it a lot. But every time you intervene based on your observations of the chain of thought, you are implicitly applying a tiny bit of pressure for the model to then hide its chain of thought. This is one major concern. We’re already seeing signs that chain-of-thought monitorability is degrading, for various reasons. We’re trying to figure out exactly why, because we want to reverse the trend."
Obviously, you don't need to figure anything out. You are using it ,the model is adapting, and we will lose the ability to know what is going on
The labs are playing with fire, they know they're playing with fire and we are going to get burnt.
Link below to interview and transcript.
Your point about metrics being gamed gets at something broader that I keep coming back to: as AI gets more capable, and as more of the progress comes from RL, we may need to rely less on engineers alone to decide how these systems are shaped.
Once you are dealing with a powerful optimizer, the problem starts to look less purely technical. You set incentives, metrics and monitoring systems, and the agent adapts to them. Economists have spent a long time thinking about exactly this kind of problem.
I am not sure how far the analogy goes, but I suspect that as AI becomes more capable, people who think about incentives, strategic behavior and imperfect monitoring should have a much larger role in how these systems are trained, evaluated and governed.
A few examples:
Goodhart: train hard against a reward model, and the AI may learn to maximize the score rather than what the score was meant to capture.
Lucas: if passing a safety eval becomes the gateway to deployment, a sufficiently capable model may learn to behave differently when it recognizes that it is being evaluated.
Holmström and Milgrom: reward the dimensions of AI behavior we can measure well, and we may push optimization away from important dimensions that are much harder to measure.