My amazing conversation with a Vegas airport shoeshiner about his encounter with Googler Vint Cerf, the father of the internet, and the story of what happened next https://t.co/4u9PmZ2VPo
OpenAI HR told Leopold that a major reason for his firing was the memo he'd sent the board warning how broken OAI's security was.
What happened next makes me even prouder to call him a friend.
People may have forgotten, but up until 2024, OpenAI had this secret non-disparagement clause in their off-boarding agreement. If you didn’t sign it, they'd clawback all your equity.
As far as I'm aware, there were only two people ever who didn’t sign it - @DKokotajlo, and Leopold Aschenbrenner.
Of all the researchers who left OpenAI before 2024, who had far more to fall back on, Leopold, all of 22 years old, was one of the only ones willing to refuse the golden handcuffs, and say, "I want to preserve my right to talk about safety and security at OAI.”
We have a lot of disagreements, but he's among the most principled and resilient people I know.
Google just dropped a 2-hour course on Loops and Graphs: from agents to Loops to full automation
17:44 - Your first AI agent
39:30 - Run agents with loops
1:12:38 - Turn loops into graphs
1:34:26 - Self-throttling agents
1:55:05 - Graph engineering
This 1-hour watch replaces any $500 course you could pay for.
Watch it, then try your first loop or graph with the step-by-step guide below
Not sure what to make of this but I was able to completely AI generate this song including the lyrics, beats, voice, literally prompted a song on demand https://t.co/kPCvVrBUpS
A Taxonomy of LLM Halucinations
Large Language Models have transformed how we interact with artificial intelligence, yet their tendency to generate plausible but factually incorrect content—known as "hallucination"—remains one of the most critical challenges in AI deployment. Manuel Cossio's comprehensive taxonomy provides a timely analysis that fundamentally reframes our understanding of this phenomenon, moving beyond the optimistic assumption that hallucinations can be completely eliminated to a more nuanced approach focused on management and mitigation.
Theoretical Foundations and Inevitability
The paper's most significant contribution lies in its formal mathematical proof that hallucination is theoretically inevitable in any computable Large Language Model. Using diagonalization techniques from computability theory, Cossio demonstrates that "for any computably enumerable set of LLMs, there exists a computable ground truth function such that all states of all LLMs within that set will exhibit hallucination." This theoretical framework represents a shift from viewing hallucinations as engineering problems to be solved, to understanding them as fundamental limitations inherent to the computational nature of these systems.
This inevitability theorem carries profound practical implications. It suggests that the goal should not be perfect factual accuracy, but rather the development of robust detection mechanisms and effective management strategies. The paper argues that "without the integration of external aids such as guardrails, knowledge bases, or direct human control, LLMs cannot be autonomously used in safety-critical decision-making processes."
A Unified Taxonomical Framework
One of the paper's key strengths is its systematic approach to categorizing hallucinations. Cossio addresses the field's fragmentation by establishing clear distinctions between intrinsic hallucinations (contradicting input context) and extrinsic hallucinations (inconsistent with training data or reality), as well as between factuality hallucinations (absolute correctness) and faithfulness hallucinations (adherence to input). This unified framework provides researchers and practitioners with a common vocabulary for discussing and addressing different types of hallucination errors.
The taxonomy extends beyond these core distinctions to identify specific manifestations across domains. The paper catalogs everything from factual errors and temporal disorientation to ethical violations and task-specific hallucinations in areas like code generation and multimodal applications. This granular categorization is crucial because, as the author notes, "each type often arises from different underlying mechanisms," requiring tailored detection and mitigation strategies.
Complex Causation and Emergent Properties
Rather than treating hallucinations as simple bugs, Cossio presents them as emergent properties arising from complex interactions between data quality, model architecture, and user prompting. The paper identifies three major categories of causal factors: data-related issues (including quality, biases, and outdated information), model-related factors (such as auto-regressive nature and overconfidence), and prompt-related influences (including adversarial attacks and confirmatory bias).
This multi-factor analysis reveals why simple solutions have proven inadequate. The auto-regressive nature of LLMs fundamentally prioritizes generating plausible token sequences based on statistical patterns rather than ensuring factual accuracy. Combined with issues like exposure bias in training and the inherent randomness of sampling strategies, hallucinations become an inevitable consequence of current LLM design paradigms.
Human Factors and Cognitive Biases
A particularly innovative aspect of the paper is its extensive analysis of human factors in hallucination perception. Cossio identifies several cognitive biases that amplify hallucination risks, including automation bias (over-relying on AI outputs), confirmation bias (accepting information that confirms existing beliefs), and the illusion of explanatory depth (overestimating one's ability to evaluate AI-generated content).
The paper demonstrates that these biases persist even when users are explicitly warned about potential inaccuracies, making technical solutions insufficient. Instead, it advocates for user interface designs that incorporate calibrated uncertainty displays, source-grounding indicators, and justification prompts to promote more critical evaluation of LLM responses.
Evaluation Challenges and Limitations
The paper provides a comprehensive survey of existing evaluation benchmarks and metrics, from TruthfulQA and HalluLens to domain-specific tools like MedHallu for medical applications. However, it also reveals significant limitations in current assessment methods, including lack of standardization, task dependence, and insensitivity to subtle hallucinations.
Cossio argues that most automatic hallucination detection scores provide "little to no insight into why a particular output is deemed hallucinated," limiting their diagnostic value. This analysis points toward the need for more sophisticated evaluation frameworks that combine surface-level similarity measures with logic- and knowledge-aware assessments.
Mitigation Strategies and Future Directions
Given the theoretical inevitability of hallucinations, the paper advocates for hybrid mitigation systems that combine multiple complementary strategies. These include architectural approaches like Toolformer-style augmentation and Retrieval-Augmented Generation (RAG), as well as systemic approaches involving guardrails and symbolic integration.
The paper emphasizes that effective mitigation must be context-aware, adapting strategies based on application requirements. In high-stakes domains like medicine and law, systems should prioritize factual accuracy over fluency and enforce mandatory human oversight. In creative domains, more open-ended generation might be acceptable with appropriate uncertainty indicators.
Practical Monitoring and Real-World Deployment
Finally, the paper introduces practical resources for monitoring LLM performance in real-world deployments, including platforms like Artificial Analysis, the Vectara Hallucination Leaderboard, and LM Arena. These tools provide crucial infrastructure for tracking hallucination rates and model reliability as LLMs continue to evolve.
Conclusion
Cossio's comprehensive taxonomy represents a mature, nuanced approach to one of AI's most pressing challenges. By establishing the theoretical inevitability of hallucinations while providing practical frameworks for understanding, detecting, and mitigating them, the paper moves the field beyond simplistic solutions toward sophisticated, context-aware management strategies. This work will likely serve as a foundational reference for researchers and practitioners working to deploy LLMs safely and effectively in real-world applications.
The paper's contribution may be its reframing of the hallucination problem from a technical hurdle to be overcome to a fundamental characteristic to be managed. This perspective shift, supported by rigorous theoretical analysis and comprehensive practical guidance, provides a roadmap for responsible AI development in an era where large language models are becoming increasingly integrated into critical decision-making processes.
Someone just reminded me of this lecture I gave in 2009 that described the evolution of Google Search from 1999 to 2009. People who are interested in how our search systems work might find this interesting.
It touches on disk-based serving systems, in-memory indices, compression schemes for inverted indices, latency issues due to interference from background processes, queries of death, evolution of hardware, and more..
Video: https://t.co/OnFk1azDk9
Slides: https://t.co/q4WZXRWQg6
BREAKING: Text-to-FILM is now real.
SkyReels V2 is the world’s first open-source AI that creates full-length, movie-quality videos with unlimited duration from a single prompt.
Script, storyboard, voice, visuals — all in one flow.
Let me show you (full breakdown):
This lecture on LLMs is a must-watch for AI engineers.
This 1.5-hour lecture covers—tokenization, scaling laws, fine-tuning, evaluation, optimization, challenges, costs, and more.
A must-watch for AI engineers! Find it in the replies.
By Stanford with ~1M views.
First draft online version of The RLHF Book is DONE. Recently I've been creating the advanced discussion chapters on everything from Constitutional AI to evaluation and character training, but I also sneak in consistent improvements to the RL specific chapter.
https://t.co/2oveIOpdCD
RLHF has a long future ahead of it and this will do a lot to make it more accessible to the next generation.
What's next: Getting a physical copy in your hands (may not be exactly 1to1, we'll see) and minor fixes at a slower cadence (thanks to many github contributors, some of you will get a copy from me).
Here are all the chapters.
1.Introduction: Overview of RLHF and what this book provides.
2.Seminal (Recent) Works: Key models and papers in the history of RLHF techniques.
3.Definitions: Mathematical definitions for RL, language modeling, and other ML techniques leveraged in this book.
4.RLHF Training Overview: How the training objective for RLHF is designed and basics of understanding it.
5.What are preferences?: Why human preference data is needed to fuel and understand RLHF.
6.Preference Data: How preference data is collected for RLHF.
7.Reward Modeling: Training reward models from preference data that act as an optimization target for RL training (or for use in data filtering).
8.Regularization: Tools to constrain these optimization tools to effective regions of the parameter space.
9.Instruction Tuning: Adapting language models to the question-answer format.
10.Rejection Sampling: A basic technique for using a reward model with instruction tuning to align models.
11.Policy Gradients: The core RL techniques used to optimize reward models (and other signals) throughout RLHF.
https://t.co/IJ0TeZ6hps Alignment Algorithms: Algorithms that optimize the RLHF objective directly from pairwise preference data rather than learning a reward model first.
13.Constitutional AI and AI Feedback: How AI feedback data and specific models designed to simulate human preference ratings work.
14.Reasoning and Reinforcement Finetuning: The role of new RL training methods for inference-time scaling with respect to post-training and RLHF.
15.Synthetic Data: The shift away from human to synthetic data and how distilling from other models is used.
16.Evaluation: The ever-evolving role of evaluation (and prompting) in language models.
17.Over-optimization: Qualitative observations of why RLHF goes wrong and why over-optimization is inevitable with a soft optimization target in reward models.
https://t.co/BoIMqWyK9h and Information: How RLHF is often underestimated in its role in improving the user experience of models due to the crucial role that style plays in information sharing.
19.Product, UX, Character: How RLHF is shifting in its applicability as major AI laboratories use it to subtly match their models to their products.
🚨BREAKING: Stanford University just launched a FREE AI tool for researchers!
It writes Wikipedia-quality reports with 99% accuracy & citations.
Here’s how to access it for free:
#MachineLearning Systems — Principles and Practices of Engineering Artificially Intelligent Systems: https://t.co/LJXiuTNIDx by @profvjreddi
“open-source textbook focuses on how to design and implement AI systems effectively”
————
#ML#AI#MLOps#DataScience#DataScientist
We knew very little about how LLMs actually work...until now.
@AnthropicAI just dropped the most insane research paper, detailing some of the ways AI "thinks."
And it's completely different than we thought.
Here are their wild findings: 🧵
New short course: Vibe Coding 101 with Replit! Learn to build and host applications with an AI agent in this course, built in partnership with @Replit and taught by its President @pirroh and Head of Developer Relations @mattyp.
Coding agents are changing how we write code. "Vibe coding" refers to a growing practice where you might barely look at the generated code, and instead focus on the architecture and features of your application. However, contrary to popular belief, effectively coding this way isn't done by just prompting, accepting all recommendations, and hoping for the best. It requires structuring your work, refining your prompts, and having a systematic process that lead to a more efficient and effective workflow.
I code frequently using LLMs, and asking an LLM to do everything in one shot usually does not work. I'll typically take a problem, partition it into manageable modules, spend time creating prompts to specify each module, and use the model to produce the code one module at a time, and test/debug each module before moving on. A process like this is making me and many other developers faster and more efficient.
In this video-only course, you’ll learn how to use Replit’s cloud environment--with an integrated code editor, package manager, and deployment tools--to build and deploy web applications. Along the way, you’ll learn strategies for working effectively with agents and improve your development skills.
In detail, you’ll:
- Understand principles of agentic code development such as being precise, giving agents one task at a time, making prompts specific, keeping projects tidy, starting with fresh sessions for each new feature, and how to approach debugging.
- Learn how to get started with Replit, and key skills for vibe coding: Thinking, using frameworks, checkpoints, debugging, and providing context.
- Create a product requirement document (PRD) and wireframe for your agent to build a prototype of a website performance analyzer.
- See how to use an agent to make your prototype more visually appealing, and deploy it application others to access .
- Learn to build a head-to-head national park ranking app, from a sample dataset, with voting capabilities and persistent data storage, and refine further ask the assistant to recap and explain what it built to find room for improvement and reinforce your learning.
By the end of this course, you’ll have a solid foundation in building with coding agents, and a process you can use to keep vibe coding effectively.
Please sign up here: https://t.co/yDbX1QFTI7