Me too. Their self-consistency is weak; it needs more weight to influence the confidence of their answers.
They do that to some extent; you will notice they use "likely" in their answers. Hedging.
It's all an illusion anyway; they're not thinking; it just looks like it
One thing that still annoys me about even the best AI models and agents:
They’ll give an overconfident opinion. Then if you push back even a little, they’ll usually reverse it right away.
It makes me feel like they weren’t thinking that hard in the first place.
A good human teammate might change their mind too, but they’ll at least explain why or have a more principled approach to their thinking.
Maybe there’s a way to fix this with a system prompt?
@paraschopra Local Health Metrics Extractor: to extract health metrics from different formats/files, so I can have a dashboard to track my health https://t.co/kCp3XOTVP2
And stock pulse: https://t.co/wvblVDuwCl
Track stocks with 200 & 365-day moving averages
Many people think any given ML project is 99% training.
In reality, it’s 50% evaluation, 40% data cleaning, 8% integration, and 2% training.
The first two set the noise floor for learning. No ML magic matters; the model cannot lower the noise floor, as that’s the optimal bound of Shannon encoding of your data.
Thus, not a single day goes by without me thinking about ontology. Even the old labels have to be constantly reviewed.
GLM-5.2 is Fully Open, Frontier Intelligence Belongs to Everyone
Today, the sudden restriction of certain frontier models is deeply regrettable. At a time when access to frontier models is abruptly cut off for non-technical reasons, we are even more convinced of one thing: science should be global.
The path to AGI (Artificial General Intelligence) must never be enclosed by high walls. We have always believed that AGI should be the cornerstone for all of humanity to collaboratively explore the boundaries of intelligence and solve complex challenges, rather than a privilege monopolized by a few rules and subject to revocation at any moment. In the face of external blockades and restrictions, our attitude is one of radical openness. Frontier intelligence must remain open-source, accessible, and buildable, serving every dedicated developer.
GLM-5.2 is Zhipu's most capable open-source model to date. It not only supports a truly usable 1M context window but also maintains a continuous lead in the independent completion of long-horizon tasks, providing solid foundational support for building complex agent applications. It also continues to be our main engine for creating the strongest domestic coding model.
Tonight at 5:21—at this special moment—GLM-5.2 will officially be available to all GLM Coding Plan users (including Lite / Pro / Max). The API will also go live next week.
A step closer to frontier intelligence for everyone.
The future of AI is open, and it is for the people.
ModelKey: GLM-5.2
I sat down with @lukaszkaiser to get into whether the architecture he helped invent is actually enough, and what's next in generalization, coding agents, RL and more. Lukasz co-authored "Attention Is All You Need," the paper that introduced the transformer and worked on reasoning models at OpenAI so he’s been a key part of major shifts in the field. We hit on:
▪️ The case for and against a new architecture coming after the transformer
▪️ What’s required for model generalization in the physical world
▪️ How much coding agents have improved his AI research productivity
▪️ The next domains for RL
▪️ Why Anthropic initially won coding
▪️ Future research directions he’s excited about
0:00 Intro
1:12 Transformers vs. Human Learning
8:37 How Do We Get Physical World Generalization?
10:52 What Comes After Transformers
13:59 How Much Have Agents Improved Lukasz's AI Research Productivity?
17:21 How Close Is an AI Research Intern?
26:06 RL Beyond Verifiable Tasks
35:38 App Companies: Build Models or Lean on Labs?
46:21 Multimodal Is Still Missing Something
49:46 OpenAI's Bet on Reasoning
55:26 The AI Coding Wars
59:26 Focus vs. Keeping Embers Burning
1:02:09 Open Source vs. Closed Source Gap
1:05:15 Quickfire
YouTube: https://t.co/wetErBD4B8
Spotify: https://t.co/FSYWR39ep2
Apple: https://t.co/j2omVI8xEa
A new and possibly controversial perspective:
In this video, I explain the sense in which generative AI trained by supervised learning is incapable of making novel discoveries.
https://t.co/zin5QbbT9N
The text of the speech:
AI Creativity and Discovery
Good day ladies and gentlemen. I regret that I am unable to be with you all today to engage in a back-and-forth discussion, but I am nevertheless pleased to be able to share with you, via this recording, some high-level thoughts about the current and future state of artificial intelligence, and in particular about AI’s relationship to science and mathematics, which is, as I understand it, the central focus of this meeting and of the SAIR Foundation.
I would like to start with an old joke; I am sure you have heard it before. It is the one about the researcher whose work is being evaluated, and the review comes back, and says “This work is both novel and good. Unfortunately, the parts that are good are not novel, and the parts that are novel are not good.”
My first point about AI is that this assessment applies exactly to large parts of AI as we know it today. Not all of today’s AI, but a large part of it. Pretty much all of what we mean by “Generative AI”---which includes large language models, and the images and video models, and even the new methods for learning world models. All of these AIs take large numbers of examples and produce a “model” which behaves similar to the examples, that is, which generates text like people, or images like artists or nature, and videos like we find on the internet. Don’t get me wrong, Generative AI can be extremely useful. No doubt about that. But the assessment of the joke still applies. These systems can produce output that is both novel and good, but not at the same time.
In many ways this is just absolutely not a problem. When we ask an AI for an answer from the internet, or to summarize a document, we don’t want it to be novel. We are happy if the quality of the answer, the goodness, comes from the source material—from the people who wrote the document or the articles on the internet. If the AI’s answer is novel it means it is going beyond the source material, adding something beyond it. This is what we call “hallucinations”. In most cases, we don’t like it when the AI makes something up, when it adds something novel.
One exception, of course, is when we are looking not for facts or reality, but for fiction and entertainment. We might ask for a bedtime story for a child, or an image based on existing images on the internet but which is nevertheless different and distinct from them. In these cases, it is never easy for us to know how creative the AI is actually being, as we do not know how close the AI’s story, poem, or image is to the source material. In a real practical sense we can not know this because the internet is too big, the possible sources that the AI may draw upon are too numerous.
When we ask for a fiction or novelty, the AI can give it to us because its processing is in part stochastic. Every decision can go multiple ways and will go different ways and produce a different trajectory every time. The trajectory can be random—and thus novel—or it can be based on the training data—and thus “good” because the training data is good, sourced from people or reality. Thus, the trajectory is either novel or good—based on randomness or based on data—but never both at the same time.
Really, I think it is okay if the output of Generative AI is never good and novel at the same time. For the researcher in the joke this is a devastating criticism, but for most things it is not, and for Generative AI it is not. Generative AI is meant to be a mimic. This is what supervised learning is for. Generative AI can be extremely useful, even when it just mimics, if it is faster, or cheaper, or smaller, or more customizable, or more copy-able, than the thing being mimicked. It is okay if Generative AI cannot be both novel and good at the same time. It is still a transformative technology.
But it is a limitation. And remember we are here to use AI for science and mathematics, and for these areas the assessment of the reviewer in the joke is devastating. For these areas we need true creativity and discovery. Generative AI—or Mimicking AI—will never get where us there. For these we need something more, and indeed we have something more in other parts of AI. We have many AI systems which can give us more. We have AlphaGo with its world-changing move 37, or AlphaZero with its brilliant original chess-playing style. We have GT-Sophy that drives simulated racecars better than any human. We have AlphaFold and AlphaProof and Claude-Code, which have brought true advances in science, mathematics, and programming. We have RL-Lyft which optimizes the assignment of cars to passengers in the ride-hailing business. All these systems have found things that are both novel and good. And, truth be told, some language models have been augmented in ways that make them more than Generative AI based on supervised learning.
All these systems have some additional features that make them capable of true creativity and true discovery. It is important for us to recognize what this is—and that it is not present in ordinary, garden-variety Generative AI. It is something that can not come from just supervised learning, from learning from examples. What is it? Well, it is a simple thing, a commonsense thing. It is not new. We have many names for it, but unfortunately none of them are very good names. I will call it Discovery. Basically, Discovery is just the idea of trying many things and seeing which of them work, then keeping those that worked the best. Evolution by natural selection works this way. The scientific method works this way. And just ordinary life and learning works this way. We try things and remember what works. What could be more obvious? In this behavioral case, psychology has two names for it— “instrumental learning” and “operant conditioning”—and in machine learning it is what we mean by “reinforcement learning”. We also see the idea of Discovery in planning and combinatorial search—anything that involves the idea of “generate and test”.
The essence of Discovery is to combine three steps:
1. Variation,
2. Evaluation, and
3. Selective retention.
Of course, I am not the first to say this. I am not the first to point out that this combination of steps is key to science, to evolution by natural selection, and to animal behavior. I think particularly of papers by Donald Campbell, by Daniel Dennett, and by Gary Cziko. What is new in my remarks is to directly relate the idea of Discovery to modern AI to help us see that it is not present in supervised learning or Generative AI—in particular, that Discovery is not present in backpropagation or gradient descent.
Let me say explicitly what is missing from Generative AI. As we have remarked, these systems do have a stochastic aspect, so they do generate a variety of trajectories and behavior. What is missing is the Evaluation step. The generator was pre-trained by supervised learning, leaving no way at runtime to Evaluate what it generates. And of course without Evaluation there can be no Selective retention, and thus no Discovery. The variation can bring novelty, but without evaluation there is no Discovery, and arguably, no creativity. That is, I would say that creativity requires that the new things generated be Evaluated. Without evaluation, and retention of the best, there is nothing created. The novelty flickers into existence but, if its value is unrecognized, it flickers away and is lost.
In many cases, Evaluation is done by people to make a discovery. As when we have Generative AI make many pictures for us, and then we pick the one that we like the best. The human+AI system completes the discovery.
In many other cases, the Evaluation comes from a clear objective. Some moves lead to checkmate, some steps lead to a proof, some actions result in high reward, some genotypes make more copies, some theories explain the data better.
Some prefer the Variation step to be called Blind variation, where “blind” here means that it is uninformed, a shot in the dark. It does not need to be completely uninformed; a good scientist does not select theories to test at random. But neither can it be completely informed and determined. There must be some uncertainty about where the answer lies in order for there to be a discovery. In practice, the variation is partly informed and partly blind, but it is the blind part that corresponds to the discovery.
Now let us briefly go all the way to modern deep learning, to the backpropagation algorithm. At first it might seem that backpropagation is incapable of discovery because it is deterministic and thus incapable of variation. But this is not correct. The weight updates of backprop are deterministic, but the weights are initialized to small random values. The random initialization is often downplayed, but in fact it is a necessary form of variation; it must be done properly to get good performance. In backprop this Variation is done once, at network initialization, so its effect is temporary, and later the network may lose its ability to learn. This is the weakness of deep learning that is alleviated with a new algorithm that my group presented in Nature a couple of years ago. Our “continual backpropagation” made one small change: every so often a less-used neuron would be re-initialized to small random weights. This allows the variation to continue and plasticity to be retained.
Although there is much more to be said about Creativity and Discovery, this is the key point: they are more than supervised learning, more than pattern recognition, more than prediction, and more than world modeling. Those things are important, but they alone will not bring us to discovery. Discovery requires Evaluation from a person or from an explicit goal, and only in the latter case will we attain full autonomy.
So that is my call to arms. If we want the full power of AI scientists, then we should share the goals with them so they can create, evaluate, discover, and in these ways fully participate in achieving the goals. Let’s be bold! Let’s fully automate Creativity and Discovery!
Weird thing about LLMs: "incorrect responses" are more expensive than correct ones.
If I go to a restaurant and they screw up my food, they usually refund me and remake the meal.
If an LLM gets stuck on a problem, it runs around in loops, burning tokens and costing money.
after working solo for almost two months, i now think the most valuable part of this experience is the freedom to explore who you really are
when you are working within a company, you are shaped by what's expected of you. you are a developer working on project X. you are the tech lead for team Y. you are the director of engineering for org Z
you play your role for years or decades, and that slowly becomes your identity. you start to tell others that's who you are, and you don't even realize that maybe you haven't actually explored who you could be
that was me, and i had no idea
in the past two months, i wrote blog posts, tweets, made a few youtube videos, published interesting data analyses, built and shared many open source projects, and made an app that my son started to use
through this experience i learned a lot more about what i care about, what i enjoy, what i can offer that's truly valuable, and what are some different ways of live a life
if you ever find yourself having some freedom on your hands, i think it's worth taking the opportunity to go try those things that you once thought could be interesting but never got the time to do, and just see where it leads you
somewhat similar to finding product-market fit, this is about finding the best fit between you and the world you're in. what could be more important?
Major cheat code for life: Be fully where your feet are. When you're at work, work. When you're with family, be with family. When you're resting, rest. Most people are physically present and mentally everywhere else.
Future devs will find it hard to believe; they will look back at us memorizing syntax the way we look at people writing assembly by hand - not wrong, just... unnecessarily hard.
Can’t believe I coded by hand for 15 years.
15 years of memorizing syntax, Vim, Stack Overflow, broken builds, cursed dependencies, merge conflicts, and “one last bug before sleep.”
All of that just to end up typing “fix this” into a chat box and watching an agent do crimes.
Wow, til: It [some microbe] ate a bacterium that had learned to use oxygen rather than die from it, and instead of digesting its meal, it kept it alive inside itself. That trapped bacterium became the mitochondria, the little engines that power your cells right now
Oxygen already killed most of the life on Earth once. The first time it filled the air, around 2.4 billion years ago, it was so poisonous that nearly everything alive died. Scientists call it the Oxygen Catastrophe.
Back then the oceans were full of tiny microbes, and none of them used oxygen. Then one kind, an ancestor of the green scum you still see on ponds, started giving off oxygen as a waste gas, the same way you breathe out air you don’t need. Oxygen is a wrecker. It rips apart the delicate machinery inside a living cell, including the DNA, and as it built up in the water and then the sky, it triggered the first mass extinction this planet had ever seen.
A few survivors hid in the mud and deep underground where the gas couldn’t reach, and some of their descendants are still down there. But one tiny cell did something nobody else did. It ate a bacterium that had learned to use oxygen rather than die from it, and instead of digesting its meal, it kept it alive inside itself. That trapped bacterium became the mitochondria, the little engines that power your cells right now. Almost every cell you are made of carries hundreds or thousands of them, all descended from that one strange truce with a poison.
The trade was worth it because burning food with oxygen releases about 18 times more energy than burning it without. It is the reason anything can swim fast or think hard. Every big, fast-moving animal on Earth, you included, runs on the gas that almost ended life.
Oxygen changed the sky too. Some of it floated up high and turned into ozone, a thin layer that blocks most of the sun’s harshest rays. Before that shield existed, raw sunlight was strong enough to fry the DNA of anything out in the open, so life had to stay underwater, where a few feet of sea soaked up the danger. For almost two billion years, nothing lived on land at all. Only once the ozone grew thick enough, a few hundred million years ago, did the first plants and animals crawl out of the water.
And the old poison never really left. Every second, the oxygen your cells burn throws off tiny broken bits called free radicals, and they keep nicking your DNA and the proteins around it. The damage adds up, slowly, your whole life. Back in 1956 a scientist named Denham Harman suggested this slow rusting from the inside is a big reason we get old. People still argue about how much it matters, and no antioxidant pill has ever been shown to make anyone live longer, but the basic idea has held up. The gas keeping you alive right now is also quietly wearing you down, year by year. The joke just got the timing wrong. Oxygen really does kill slowly, and billions of years before we showed up, it already proved it can kill fast.