π³Introducing Knowledge Navigatorπ³: A new way to explore scientific literature! Our paper shows how LLMs can transform information overload into structured, navigable knowledge.
paper: https://t.co/VKQAmuq7il
page (with demo): https://t.co/H0mOf1oFwf
π§΅/9
π¨ We present πππππ: ππππͺπ§π°ππ₯ π’πΈπ’π³π¦ ππ°π―π€π¦π±π΅ ππ³π’π΄πΆπ³π¦ (with @yoavgo and @yanaiela).
MANCE is a new state-of-the-art method for concept erasure.
@yoavgo Maybe authors rejection for a full year will do the work. Id like to see senior researchers risk their lab. Also maybe it will encourage them to read student's submissions
@boknilev@itayevron yes i get the same feeling when reviewing but also when trying "vibe research" by myself. Like claude is chasing small experimental "rewards", overcooking meaningless findings.. need to find a way to tame it
New paper!
People treat reasoning trajectories as text, but what if we can do better than that?
We show that we can, by training Behavior Forecasters (BFs) that get a reasoning trajectory as input and make more accurate forecasts than frontier models at a fraction of the cost. π§΅
We analyzed 250K+ queries & 430K+ clickstream interactions from Asta, our AI-powered research assistantβand today we're releasing the full dataset. How do researchers actually use AI science tools? Here's what we found. π§΅
if your friends arenβt talking about:
- claude code
- creatine
- openclaw
- looksmaxxing
- ai agents
- taste
- prediction markets
- mac minis
you should be grateful & cherish those people
Introducing Theorizer: Turning thousands of papers into scientific laws πβ‘οΈπ
Most automated discovery systems focus on experimentation. Theorizer tackles the other half of science: theory buildingβcompressing scattered findings into structured, testable claims. π§΅
We view chain-of-thought monitoring as complementary to mechanistic interpretability, not as a replacement for it.
Because we believe that chain-of-thought monitoring is incredibly useful as a window into a modelβs brain and could be a loadbearing layer in a scalable control scheme, we plan to expand upon these evaluations and use them to inform future modeling and data decisions.
New paper:
We are often told that reasoning tokens aren't faithful explanations. But to have a useful metaphor for their operation we need a characterization of what they are, not what they are not.
To that end, we suggest "State over Tokens" (SoT) ππ§΅