One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
ESPAÑA ES VIGENTE CAMPEONA DEL MUNDO MASCULINA.
ESPAÑA ES VIGENTE CAMPEONA DEL MUNDO FEMENINA.
ESPAÑA ES VIGENTE CAMPEONA CONTINENTAL.
ESPAÑA ES VIGENTE CAMPEONA OLÍMPICA.
NO BUSQUEN PRECEDENTES DE ESTE HISTÓRICO POKER.
NO EXISTEN.
❤️🇪🇸
‼️Broche de oro a la bendición de la torre de Jesucristo de la Sagrada Familia: Sorpresa con el espectáculo de canto, luces y drones.
“Primero el amor, después la técnica” (Gaudí)
¡Qué preciosidad!
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.
It’s happening faster than we thought, and the implications deserve greater attention. https://t.co/OVVPJO7VQx
People who don't follow cancer research often ask me why we haven't cured cancer. That perception masks a wonderful reality: We make amazing, stepwise progress every year, and the result is that many people live much longer today than they would have previously.
Right now we're in the thick of the annual meeting of the American Society of Clinical Oncology, the biggest research meeting on new cancer medicines, and this morning a bunch of really important studies dropped. I'm going to review them here.
This first image is the result for daraxonrasib, a treatment for pancreatic cancer that is generating consdirable excitement. The green line is the probability of living for patients who got the new drug; the gray one is the chemo control group.
If you follow cancer drugs, a chart like this will make your breath hitch a little. I'm going to review these and some other data here.
AI has now solved a major open problem -- one of the best known Erdos problems called the unit distance problem, one of Erdos's favourite questions and one that many mathematicians had tried.
https://t.co/SD1vVPkrHR
Smartphones are not the explanation for the recent decline in fertility. Instead, they are an accelerator of deeper forces already at work.
Let’s start with the facts. Fertility is falling almost everywhere: in rich, middle-income, and poor countries; in secular and religious countries; and in countries with high and low levels of gender equality.
The decline accelerated around 2014. So, no country-specific explanation will work unless you are willing to believe that 200 distinct country-specific explanations arrived at roughly the same time.
Smartphones look like the obvious candidate: the first iPhone was released in 2007, and global adoption has been astonishingly fast.
Economists understand the first major decline in fertility in advanced economies, from 6 or 7 children per woman throughout most of human history to about 1.8, that occurred between the early 1800s and roughly 1970, well before smartphones. The main drivers were a sharp fall in child mortality (effective fertility was rarely above 3 and often close to 2) and the shift from a low-skill, rural agrarian economy to a high-skill, urban industrial one. We have quantitative models that fit these facts well.
Country-specific factors mattered too, of course. Proximity to low-fertility neighbors accelerated Hungary’s decline, while fragmented landowning structures accelerated France’s. But these were second-order mechanisms.
This is also why most economists long considered Paul Ehrlich’s doom scenarios implausible. We forecast that fertility in middle- and low-income economies would follow the same path as in the rich, probably faster, because reductions in child mortality reached India or Africa at lower income levels (medical technology is nearly universal, and most gains come from handwashing and cheap antibiotics, not Mayo Clinic-level care). Much of what we see in Africa or parts of Latin America today is still that old story.
But in the 1980s, a new pattern appeared. Japan and Italy fell below 1.8, the level we had thought was the new floor. By 1990, Japan was at 1.54 and Italy at 1.36.
This second fertility decline began in Japan and Italy earlier than elsewhere, driven by country-specific factors, but the underlying dynamics were widespread: secularization, an education arms race, expensive housing, the dissolution of old social networks, and the shift to a service economy in which women’s bargaining power within the household is higher. The U.S. lagged because secularization came later, suburban housing remained relatively cheap, and African American fertility was still high. U.S. demographic patterns are exceptional and skew how academics (most of whom are in the U.S.) and the New York Times see the world.
My best guess is that, without smartphones, Italy’s 2025 fertility rate would be about 1.24 rather than 1.14. I doubt anyone will document an effect larger than 0.1-0.2. Italy was at 1.19 in 1995, not far from today’s 1.14. The TFR is cyclical due to tempo effects, so I do not read too much into the rise between 1995 and 2007 or the decline from 1.27 in 2019 to 1.14 today. The direct effect of smartphones is not zero, but it is not, by itself, that large.
Where social media, in general, and smartphones, in particular, matter is in the diffusion of social norms. What would have taken 25 years now happens in 10. Social media are not the cause of fertility decline; modernity is. But they are a very fast accelerator.
That is why social media are a major part of the story behind Guatemala (yes, Guatemala) going from 3.8 children per woman in 2005 to 1.9 in 2025. Without them, Guatemala would also have reached 1.9, just 20 years later.
Modernity, in its current form, is incompatible with replacement-level fertility. By modernity, I do not mean capitalism: fertility fell earlier and faster in socialist economies than in market economies. Socialist Hungary fell below replacement in 1960, and socialist Czechoslovakia in 1966 (both experienced small, short-lived baby booms in the mid-1970s). By modernity, I mean a society organized around rational, large-scale systems and formalized knowledge.
Countries will not converge to the same fertility rate. East Asia is likely stuck near 1, possibly below, given its unbalanced gender norms and toxic education systems. Latin America faces the same gender problem plus weak growth prospects, so I expect something around 1.2. Northern Europe has more egalitarian family structures and might hold near 1.5. The very religious societies are probably the only ones that will sustain 1.8.
All of this could change with AI or changes in population composition. We will see. But on the current evidence, deep sub-replacement fertility is the “new new normal.” Unless we reorganize our societies, better learn to handle it as best we can.
I've recently got in on the act of getting AI to solve open problems in mathematics. More precisely, I gave some questions asked by Melvyn Nathanson to ChatGPT 5.5 Pro, to which I have been given access, and it answered them. 🧵
Judging by my tl there is a growing gap in understanding of AI capability.
The first issue I think is around recency and tier of use. I think a lot of people tried the free tier of ChatGPT somewhere last year and allowed it to inform their views on AI a little too much. This is a group of reactions laughing at various quirks of the models, hallucinations, etc. Yes I also saw the viral videos of OpenAI's Advanced Voice mode fumbling simple queries like "should I drive or walk to the carwash". The thing is that these free and old/deprecated models don't reflect the capability in the latest round of state of the art agentic models of this year, especially OpenAI Codex and Claude Code.
But that brings me to the second issue. Even if people paid $200/month to use the state of the art models, a lot of the capabilities are relatively "peaky" in highly technical areas. Typical queries around search, writing, advice, etc. are *not* the domain that has made the most noticeable and dramatic strides in capability. Partly, this is due to the technical details of reinforcement learning and its use of verifiable rewards. But partly, it's also because these use cases are not sufficiently prioritized by the companies in their hillclimbing because they don't lead to as much $$$ value. The goldmines are elsewhere, and the focus comes along.
So that brings me to the second group of people, who *both* 1) pay for and use the state of the art frontier agentic models (OpenAI Codex / Claude Code) and 2) do so professionally in technical domains like programming, math and research. This group of people is subject to the highest amount of "AI Psychosis" because the recent improvements in these domains as of this year have been nothing short of staggering. When you hand a computer terminal to one of these models, you can now watch them melt programming problems that you'd normally expect to take days/weeks of work. It's this second group of people that assigns a much greater gravity to the capabilities, their slope, and various cyber-related repercussions.
TLDR the people in these two groups are speaking past each other. It really is simultaneously the case that OpenAI's free and I think slightly orphaned (?) "Advanced Voice Mode" will fumble the dumbest questions in your Instagram's reels and *at the same time*, OpenAI's highest-tier and paid Codex model will go off for 1 hour to coherently restructure an entire code base, or find and exploit vulnerabilities in computer systems. This part really works and has made dramatic strides because 2 properties: 1) these domains offer explicit reward functions that are verifiable meaning they are easily amenable to reinforcement learning training (e.g. unit tests passed yes or no, in contrast to writing, which is much harder to explicitly judge), but also 2) they are a lot more valuable in b2b settings, meaning that the biggest fraction of the team is focused on improving them. So here we are.
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software.
It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans.
https://t.co/NQ7IfEtYk7
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: https://t.co/CDSQ8HpZoc