I'm reflecting on how much research has changed since I've joined the PhD and wrote a short blog post about it (I joined in the tail end of the BERT era!). It seems pretty crazy how different processes are now, and I took the chance to do a retrospective before graduation:
https://t.co/OR7N8s4CgO
I was just reading this again this morning, and it is not just one of my favorite pieces of writing about science, it is one of my favorite pieces of writing, period. A masterclass in prose.
https://t.co/FkhNw7ogcU
Food for thought!
"When the Middle Class Undermines Progress: The Political Economy of AI Regulation" by Anna Denisenko and Konstantin Sonin.
"The classic literature argues that a large middle class is a prerequisite for economic development. Is this true when AI progress threatens to disrupt the labor market? On the one hand, uncertainty about the consequences can impede the passage of AI-friendly regulation: voters fear the prospect of losing when the uncertainty is resolved. On the other, granular information about reform consequences may reduce political support as well. In our model, a larger middle class hampers reform, compensatory redistribution is most politically effective when its welfare justification is weakest, and AI regulatory sandboxes encounter resistance because new information undermines reform coalitions."
https://t.co/QnRLl7dSfy
Research paper finds:
“Everyone gains with AI use in the short run. The long run depends on how you engage”
The 8 Modes of AI Engagement: How You Work With AI Determines Whether It Makes You Sharper or Duller
Most people treat AI as a smarter search engine or a tireless junior employee. They ask a question, receive a fluent answer, and move on. That pattern feels efficient. Research grounded in cognitive science, learning theory, and human-computer interaction suggests it is also quietly costly.
The AI Engagement Modes framework at Brigham Young University) and detailed in the research foundations at https://t.co/8OJzOrCzox, classifies human-AI interaction into eight observable modes. These modes sit on a clear gradient of cognitive agency—from near-total passivity to full human control over problem framing. The framework draws on Bloom’s revised taxonomy, Chi’s ICAP framework (Interactive, Constructive, Active, Passive), Zimmerman’s model of self-regulated learning, Vygotsky’s Zone of Proximal Development, distributed cognition, epistemic vigilance, and Schön’s reflective practice.
The central claim is straightforward: the way you engage AI shapes whether the technology amplifies your thinking or gradually erodes it.
The eight modes are grouped into three tiers of increasing human agency:
Passivity (Modes 1–2): The human largely cedes cognitive work to the AI.
Partnership (Modes 3–4): Human and AI form a joint cognitive system with shared effort.
Agency (Modes 5–8): The human retains or reasserts primary cognitive authority—verifying, challenging, expanding, or defining the problem itself.
Higher tiers map to deeper learning and better long-term skill retention. Lower tiers map to short-term performance gains that often reverse when AI is removed.
The 8 Ways of Engagement
Here are the eight modes, ordered from lowest to highest cognitive agency:
1. Oracle
You treat the AI as an authoritative source of answers. You ask broad or factual questions and largely accept the output. This corresponds to Bloom’s “remembering” level. It is efficient for quick facts but trains the habit of not checking or synthesizing.
2. Production (or Production Assistant)
You use the AI primarily as a generator of finished or near-finished output—text, code, summaries, plans—with minimal editing, verification, or reframing. The cognitive work of creation is mostly offloaded. This is still passive in the sense that the human does little constructive or critical processing.
3. Tutor
You use the AI as a scaffolded teacher. You ask for explanations, step-by-step guidance, or clarification within your Zone of Proximal Development. You remain engaged enough to learn rather than merely receive. This is the first genuine partnership mode.
4. Collaborative Problem-Solver
You and the AI work together as a distributed cognitive system. You break problems into steps, iterate, supply context, and build on the AI’s contributions while directing the overall direction. Cognition is shared rather than handed over.
5. Verification Agent
You actively evaluate the reliability of the AI’s claims before accepting them. You check assumptions, cross-reference, test edge cases, or demand evidence. This reactivates epistemic vigilance—the evolved capacity to scrutinize communicated information that fluent AI output can otherwise bypass.
6. Critical Challenger
You treat the AI as a sparring partner. You argue against its reasoning, force it to defend positions, surface weaknesses, or explore counter-arguments. This leverages the argumentative nature of human reasoning and produces stronger evaluation skills.
7. Creative Expander
You use the AI to push beyond the given frame—generating novel variations, analogies, alternative framings, or divergent possibilities. The human remains the director of creative direction while the AI multiplies options.
1 of 2
Real OGs know the strongest argument in favor of a public health insurance system is that public health insurers deny/ration care more often than free markets, and consumers are neurotic pussies who spend too much on healthcare that does nothing.
LLM interpretability is a rabbit hole. This repo gives you a map.
Awesome LLM Interpretability is a curated GitHub list of tools, papers, articles, groups, and a survey paper focused on understanding large language models.
It helps you study the field faster by grouping practical tools, research papers, explainers, and communities into scan-friendly sections instead of chasing random bookmarks.
Key features:
• Tool index – links to LLM interpretability and analysis tools like LIT, TransformerLens, Inseq, ecco, Pythia, and Automated Interpretability
• Paper list – collects academic and industry papers on topics like sparse probing, copy suppression, monosemanticity, causal tracing, and model editing
• Article section – includes explainers and interactive resources on grokking, the logit lens, activation patching, causal scrubbing, and evaluation pitfalls
• Community map – points to interpretability and alignment groups including PAIR, Alignment Lab AI, Nous Research, and EleutherAI
• Contribution path – includes fork, branch, PR, review, and merge guidelines so the list can keep improving
Free public GitHub repo.
Link in the reply 👇
Second for second, @tylercowen packs more substance into a talk than anyone I'm aware of. This is a clear, non-hysterical, and somewhat soothing discussion of our AI future.
World Labs CEO Dr. Fei-Fei Li: "The world is not made of words."
"Language models have given machines an extraordinary command of concepts, vocabulary, and reasoning, but the physical world, virtual or real, runs on a different substrate."
"Where language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, how a garden looks from an angle no camera has captured, how objects respond to force and follow the laws of physics."
"Language gave machines a way to talk about that world. World models are how machines will finally come to understand, imagine, reason and interact with it."
Full piece: https://t.co/C9qOJg5wuc
The early reporting on Canada's AI strategy. Carries the fingerprints of papers in the public consultation: -Cash for SME implementation -Subsidies for cloud/compute access especially Canadian suppliers -Keep-it-local data policies in some sectors but not others -Government as anchor customer for Canadian startups -Moonshot projects on A.I. in health, energy, ag, robotics, transport -Safety and Privacy Act upgrades. You can trace the lineage of these ideas to the expert submissions. The macro story?
The macro story is there was a sharp dichotomy in the public consultations. A contradiction. 1) More AI: The experts submitted papers with all sorts of ideas for scaling up AI use in Canada, which is slower than a number of peer countries. 2) Less AI: the broader public? A big thumbs down for AI. The public comments were mega-negative, based on the thousands I analyzed.
Based on this report, it sounds like the majority of Canada's strategy aims at addressing point #1, with some moves aimed at public concerns in point #2.
https://t.co/SjCnWs5U9S
🚨Carney - The Eleven Year Plan🚨
I spent the last week lining Mark's every move in 30 documented steps since 2015.
No more doubts! It was never about saving the planet, and I have proof.
Watch the house of cards getting built step by step.
Let me know if you see it too!
This week the only university in Newfoundland, @MemorialU, posted 5 tenured professor openings:
- AI-driven Navigation
- Computational Biochemistry
- Genomic Mapping
- Indigenous Knowledge
- Community Health and Substance Use
Each job stipulates that no white men may apply.
LLM Knowledge Bases
Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:
Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki includes summaries of all the data in raw/, backlinks, and then it categorizes data into concepts, writes articles for them, and links them all. To convert web articles into .md files I like to use the Obsidian Web Clipper extension, and then I also use a hotkey to download all the related images to local so that my LLM can easily reference them.
IDE:
I use Obsidian as the IDE "frontend" where I can view the raw data, the the compiled wiki, and the derived visualizations. Important to note that the LLM writes and maintains all of the data of the wiki, I rarely touch it directly. I've played with a few Obsidian plugins to render and view data in other ways (e.g. Marp for slides).
Q&A:
Where things get interesting is that once your wiki is big enough (e.g. mine on some recent research is ~100 articles and ~400K words), you can ask your LLM agent all kinds of complex questions against the wiki, and it will go off, research the answers, etc. I thought I had to reach for fancy RAG, but the LLM has been pretty good about auto-maintaining index files and brief summaries of all the documents and it reads all the important related data fairly easily at this ~small scale.
Output:
Instead of getting answers in text/terminal, I like to have it render markdown files for me, or slide shows (Marp format), or matplotlib images, all of which I then view again in Obsidian. You can imagine many other visual output formats depending on the query. Often, I end up "filing" the outputs back into the wiki to enhance it for further queries. So my own explorations and queries always "add up" in the knowledge base.
Linting:
I've run some LLM "health checks" over the wiki to e.g. find inconsistent data, impute missing data (with web searchers), find interesting connections for new article candidates, etc., to incrementally clean up the wiki and enhance its overall data integrity. The LLMs are quite good at suggesting further questions to ask and look into.
Extra tools:
I find myself developing additional tools to process the data, e.g. I vibe coded a small and naive search engine over the wiki, which I both use directly (in a web ui), but more often I want to hand it off to an LLM via CLI as a tool for larger queries.
Further explorations:
As the repo grows, the natural desire is to also think about synthetic data generation + finetuning to have your LLM "know" the data in its weights instead of just context windows.
TLDR: raw data from a given number of sources is collected, then compiled by an LLM into a .md wiki, then operated on by various CLIs by the LLM to do Q&A and to incrementally enhance the wiki, and all of it viewable in Obsidian. You rarely ever write or edit the wiki manually, it's the domain of the LLM. I think there is room here for an incredible new product instead of a hacky collection of scripts.
Agree; avoiding sycophancy and actively encouraging divergent analysis is essential for more trustworthy human–AI reasoning.
Just published: “An Antidote to Sycophancy: Toward Epistemic Divergence in Human–AI Reasoning”
It introduces MODAC identical prompts run across independent LLMs (different vendors, tabula-rasa conditions). Divergence is deliberately preserved as diagnostic signal, not noise. Human remains the final adjudicator. It is an LLM corollary to Medical Grand Rounds.
https://t.co/33HVff4QYP
https://t.co/ASUksMPgk2
Naming well is a great hack because the thing that’s more fun to say gets said more often, and the idea it’s associated with gets free infiltration into people’s thoughts