placed 1st, winning $5000 at the @GoogleDeepMind x @cerebral_valley hackathon bangalore! 😭
> we built an RPG game generator that uses @NanoBanana to create worlds and assets as you play and progress
> you can create your own worlds that are rooted in Indian culture
> every asset that you see is generated on the fly
> the gameplay assets, dialogues, and clues all change with the theme
> for instance, the first visual in the video is from the movie lagaan!
insane clutch at the last moment
shoutout to the team 💙: @im__apoorva@ahmedfahim21_@harsh_agw
People share deeply personal context with LLMs – it feels anonymous, judgment-free, and promises tailored advice and information.
In our #ACL2026 paper, we ask: how does an asker's role –– implicitly signaled in how they frame a question –– shape the response they get back? 🧵
Two rigorous studies. Opposite conclusions. And neither told us the one thing needed to interpret them.
Two weeks ago Nature Medicine reported frontier AI models beating specialized clinical tools. Now a stronger study, 149 blinded specialty-matched physicians, reports OpenEvidence beating the same frontier models on real clinical questions.
The design is excellent. The problem is simpler than the statistics.
AI models have settings. Reasoning effort. Thinking time. Response length. Change the settings and the quality of the answer changes, the same way changing the dose changes the drug. This study did not report the settings.
One clue they mattered: GPT-5.5, the worst performer by far, produced answers about a third shorter than every other system. Short answers are typical of default settings. And the study's own analysis showed longer answers scored better.
To be fair, OpenEvidence also won its head-to-head comparisons against all three frontier models. The finding is valid. But what it means depends on settings we were never shown. Did the specialized tool beat the frontier models, or their factory defaults? The paper cannot answer that. Neither can we.
I raised the same issue with the Nature Medicine study, which never specified which OpenEvidence mode it tested. Same standard in both directions.
In medicine we would never accept a trial reporting that patients received "a beta blocker." Drug, dose, route, frequency. AI evaluation needs the same discipline. Model, settings, effort level, web search on or off.
Until studies report configuration the way trials report dosing, the honest headline for both papers is the same. This setup beat that setup on these questions. Nothing more.
There are so many high-level, breakthrough studies coming out weekly about AI's role in medicine that it becomes really hard to keep up with the new findings. Here are three recent examples.
1) Towards autonomous medical artificial intelligence agents (Nature) https://t.co/lB1p0dut36
"EHR-integrated artificial intelligence agent can turn clinical intent into structured, actionable EHR operations, possibly making it a more effective decision-support partner for physicians."
2) Towards Conversational AI for Disease Management (Nature)
"AMIE was non-inferior to primary care physicians in management reasoning as assessed by specialists and scored better in both preciseness of treatments and investigations, and in its alignment with and grounding in clinical guidelines." The AI even outperformed physicians on higher difficulty questions.
3) General-purpose large language models outperform specialized clinical AI tools on medical benchmarks (Nature Medicine)
"Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ."
Amazing progress.
been asking others at Anthropic how they stay in the loop with Claude and fully understand the work being done
this is one of my favorites from Suzanne:
Most people “use” AI.
Almost nobody actually understands it.
Andrej Karpathy just dropped a 2-hour lecture that changes that.
No TensorFlow.
No PyTorch.
No shortcuts.
Just raw math + code → building a Neural Network from scratch.
This is the stuff 90% of AI engineers will NEVER learn.
Free on YouTube.
Watch it once…
and you’ll never see AI the same again.
- Drafted a blog post
- Used an LLM to meticulously improve the argument over 4 hours.
- Wow, feeling great, it’s so convincing!
- Fun idea let’s ask it to argue the opposite.
- LLM demolishes the entire argument and convinces me that the opposite is in fact true.
- lol
The LLMs may elicit an opinion when asked but are extremely competent in arguing almost any direction. This is actually super useful as a tool for forming your own opinions, just make sure to ask different directions and be careful with the sycophancy.
Hey @deepigoyal FWIW first you write absolute garbage with zero understanding of labor and economics
But wait these arent even your own words. It’s AI generated???
Still can’t believe it. Our work on uncovering plagiarism in AI generated research received the outstanding paper award at ACL!!
Effort led by, and envisioned by the amazing @tarungupta360!
AI is already at work in American newsrooms.
We examine 186k articles published this summer and find that ~9% are either fully or partially AI-generated, usually without readers having any idea.
Here's what we learned about how AI is influencing local and national journalism: