Finally, I met Yejin Choi at the conference!Thank you for accepting the photo request. I met her at the Ritz-Carlton with a view of the beautiful mosque last night. I was so honored that I couldn't sleep! As a Korean, I am proud of her! @YejinChoinka#EMNLP2022
interesting position paper throwing cold water on autoresearch/ai scientist: LLMs can't jump.
The thought experiment is this: Take an LLM with a 1905 knowledge cutoff. Feed it every paper, every dataset, every equation of that era. Could it invent general relativity?
No.
Discovery isn't one thing. It's three. You can induce — generalize from data, which lands you at Newton plus some epicycles to explain Mercury's weird orbit. You can deduce — derive rigorously from axioms you already have, which never gives you new axioms. Or you can jump — invent the frame itself, decide that spacetime curves. That third move is the one that matters, and it's exactly the one induction and deduction can't reach.
Penrose put it as three worlds: Physical, Mental, Platonic. Data flows from the world into a mind fine. But the new law has to be discovered into the Platonic world first — and that step is the jump. LLMs are induction machines running over what already exists. Structurally, they don't take it.
I think it’s a warning to AI scientists/autoresearch against collapsing two very different things into one word.
Hill-climbing: LLMs are already superhuman here, and autoresearch in this sense is real and moving fast.
Abduction/leap/jump: a new frame that reorganizes the field, that is a different act entirely, and nothing about scaling induction suggests you get there.
Most of what Autoresearch ships today will be spectacular hill-climbing. The jump is still ours for now.
🎉 Excited to share that our two papers have been accepted to COLM 2026!
1️⃣ New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
2️⃣ Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
Huge thanks to all collaborators! Looking forward to COLM 2026 🚀
// The Harness Effect //
(bookmark it)
Now more that ever pay very close attention to the orchestration harness and its effect on costs and performance.
This study ran 22 evaluation tasks on six foundation models (Claude Sonnet 4.6, Gemini 3.1, Qwen 3.6, GLM 5.1, and others), then change only the orchestration layer.
Holding models constant, the harness cuts blended cost per task 41%, tokens per task 38%, and median wall-clock 44%, with completion quality at parity.
Two results do the work. Efficiency is model-invariant, every model gets 33 to 61% cheaper. Quality gain correlates almost perfectly with baseline model strength (r=0.99 across six models), a effect they call harness leverage.
Why does it matter?
On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. The harness is the one component whose efficiency multiplies across every model an organization runs.
Paper: https://t.co/BkNoXp5ZLa
Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c
🏆 𝗕𝗲𝘀𝘁 𝗣𝗮𝗽𝗲𝗿 𝗔𝘄𝗮𝗿𝗱 @ 𝗔𝗖𝗟 𝟮𝟬𝟮𝟲
Thrilled to share that our linguistic paper "The Imperfective Paradox in LLMs" received a Best Paper Award at #ACL2026 ! 🎉
Hoping this encourages more linguistics-informed evaluation in NLP.
https://t.co/U2YlJwqpKX
I'm featured in this @TheAtlantic article about joining DeepMind to research "what it means to live in a world where cognitive agency is no longer uniquely human."
I do have a couple of minor critiques, though: first, the idea that "Someone Finally Wants to Hire Philosophers" is not completely true; philosophers have always been hot commodities 🙃 ; and second, the claim that Silicon Valley is leading the charge on hiring philosophers and ethicists; we have an incredible wealth of philosophers working at Google & Google DeepMind in London, a city which is hardly beatable 🙃
https://t.co/e39hHjeKaR
오늘 재밌는 논문이 나왔다.
Claude 안에 말로 출력되지 않는 개념들이 모이고, 붙잡히고, 조작되고, 여러 추론에 재사용되는 내부 작업공간이 있다는 걸 발견함.
이제 그 내부 상태를 읽고, 경우에 따라 조정할 수 있는 가능성도 생김.
https://t.co/vmlVBew3Jd
Yüksek lisans/doktora teziniz veya makaleniz için literatür taraması sürecini kolaylaştıracak yapay zeka tabanlı bir araçtan bahsetmek istiyorum. Bu aracın en güzel yanı, abonelik gerektirmemesi ve tamamen ücretsiz olması. Türkçe arama da yapabiliyorsunuz. Aracımızın adı: https://t.co/g8V7hlKnqO
Open Alex’in filtreleri sayesinde, bir çalışmanın yayın yılı, açık erişim olup olmadığını görebiliyorsunuz ve ayrıca konunuza uygun makale, kitap bölümü, tez vb. hangisini istiyorsanız onu seçebiliyorsunuz. Ayrıca atıf sayısı filtresi sayesinde, belirli bir sayı üzerinde atıf almış makaleleri de bulmanızı sağlıyor.
Open Alex’i güçlü kılan yanlarından biri de 'domain', ‘field’, ‘keyword’ gibi filtreleri sayesinde, çalışma alanınıza yönelik spesifik aramalar yapabilmeniz.
Bunların yanı sıra, sadece belirli yayıncıya (PubMed, Wiley vb.) ait dergileri de filtreleyebilirsiniz.
Ayrıca, arama sekmesinin hemen altında boolean ve semantic opsiyonları bulunuyor. Boolean opsiyonunu, tam olarak ne aradığınızı biliyor, hangi anahtar kelimeleri kullanarak araştırma yapmak istiyorsanız kullanabilirsiniz.
Semantic opsiyonu ise, konuyu genel çerçeve olarak biliyor, yapay zekanın bu konu ile anlamsal olarak ilintili çalışmalar önermesini istiyorsanız kullanabilirsiniz.
Türkçe arama yapabilsek de, malesef İngilizce kadar zengin filtre ayarları bulunmayabiliyor.
Sizler için hazırladığım, filtreleri görebileceğiniz kısa tanıtım videosunu aşağıda bulabilirsiniz.
📍Bu paylaşımım hesabımdaki diğer yapay zekâ araçları tanıtımında olduğu gibi, herhangi bir sponsorluk ya da ticari bir işbirliği değildir. Sadece bizzat deneyimlediğim akademik verimliliği artırdığına inandığım araçları araştırmacılara yol gösterme niyetiyle paylaşmaktayım.
“Linguistics is an unsung hero among sciences. It has initiated several groundbreaking trends in global history, yet it seldom receives the acknowledgment it deserves.”
– Gašper Beguš, https://t.co/dIhUGcAfiv
Must-read research by Anthropic.
Here is the simple explanation and why this is a big deal.
We suspect LLMs perform "internal reasoning". But little is known or do good methods exist to understand it.
Anthropic claims that J-Space (which differs from chain-of-thought or scratchpad), emerged on its own through training and provides a window into how Claude "reasons" internally.
In other words, this shows that Claude has a sort of internal workspace where information gets held, combined, and passed between different parts of the model. They can read from it, and they can steer the model by changing it.
As it is the case with these reports, the consciousness angle will get all the attention. However, the bigger story is that for the first time you can point to a specific place inside the model where reasoning is staged, rather than guessing at it from the text that comes out.
This, of course, changes what interpretability can be. We spent years inferring what a model was doing from what it said. Now there's a mechanism to observe directly, and a direct lever to move. This could enable even more advanced levels of "reasoning" in LLMs and bridges gaps in frontier intelligence and world models.
If you can see where a model holds an idea, you can also verify it, audit it, and catch it working toward a goal you never gave it. You can implement better guardrails and predict dangerous/unwanted scenarios better.
Great overview of always-on agents.
(bookmark it)
It's a new 130+ pages survey on always-on agents.
Simply put it, always-on agents are systems whose future behavior depends on durable state built up across earlier interactions. It treats that state as more than memory. Task ledgers, permissions, credentials, commitments, provenance, triggers, and externally committed effects all count.
The survey scores each state item on six axes, authority, scope, mutability, provenance, recoverability, and actionability, across a lifecycle that runs from write and retrieve through forget, audit, and rollback.
Paper: https://t.co/FRvZ7YGpaN
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
🚨Anthropic just showed a 24-minute workshop on how to actually do prompts for Claude.
Taught by the people who built it.
Free. No registration. No paywall.
I've seen $300 courses that don't cover what they teach in the first 8 minutes.
Watch it and bookmark it now.
Instead of watching an hour of Netflix, watch this 2-hour Stanford lecture, which will teach you more about how LLMs like ChatGPT and Claude are built than most people working at top AI companies learn in their entire careers.
It is apparently so long since @psresnik has tweeted, and so long since live-tweeting was a thing, that no one has posted on Philip’s @aclmeeting keynote? It was a great Philip-Resnik-style talk, balanced and realistic while imploring there to be more new and diverse NLP science.
AI Agents vs. Agentic AI
Interesting paper summarizing distinctions between AI Agents and Agentic AI.
It also talks about the key ideas, solutions, and the future.
Here are my notes:
Our paper "Polishing Every Facet of the GEM: Testing Linguistic Competence in LLMs and Humans" has been accepted at #ACL2025! I contributed to the paper as a co-first author.
We release a benchmark to evaluate linguistic competence of LLMs in Korean.
👉🏻https://t.co/mcf0C6l1i8
Thrilled, grateful, and humbled to have won 3 outstanding paper awards at #EMNLP2024!!! Not even in my wildest dreams. Immense thanks to my amazing students and collaborators!
All three works are on evaluating LLM’ abilities in creative narrative generation. 🧵👇
#PhD student: Makes $30k a year. Works on weekends as well. Zero life-work balance.
"Do you think I have a chance to become a professor?"
Prof: "Yes, of course! Finish this project and we will publish excellent papers. I am sure you will easily find a faculty position."
▫️
2 years later:
Student finishes the project. Professor writes a report. Papers are published.
Student: "Do you think my CV is strong enough?"
Prof: "Yes, you are the best!"
▫️
Next 4 months:
Student submits 50 well-tailored applications for faculty positions.
Zero interviews. A lot of broken dreams.
▫️
Key takeaways:
1. Make sure you distinguish encouragement from reality.
- By encouraging you, your advisor may unintentionally give you too much hope. Keep a cool head.
2. Always ask other faculties for external opinion on your case.
- Your advisor’s opinion is always biased. Look for more input outside your group.
3. Don’t expect fairness during candidate selection.
- Hiring process is subjective by definition. It is done by people with very different views on who is the best. You may put tons of efforts into a research statement only to find out later that no one really reads it.
4. The reality is brutal.
- Departments can receive 300-500 candidates per opening. Many have excellent CVs and cool ideas. At top- and mid-rank universities, selection criteria can become extremely questionable (like, who exactly was your PhD advisor? Is your recomm. letter 3 pages long? etc).
And there is no need to say “You don’t know anything about it. It’s not like this”.
I went through this myself. Many times. Along with many colleagues.
Do not expect fairness. See luck as a big factor.
Apply broadly but have a backdoor ready.
#AcademicTwitter