What % of the NLP papers measure their impact in the real world? This paper proposes an "impact evaluation" of NLP models or systems for real-world usage, changing the research culture of NLP to focus more on real-world
impact and less on SOTA-chasing: https://t.co/3Ti5vUgdgs
Crazy: in October '25 we used "ChatGPT hacks Pentagon" as an example of unrealistic AI capability / bad submission example in Misalignment Bounty (https://t.co/rO36gh5oHl), and half a year later we had Mythos capable of basically this.
@apolloresearch@AISecurityInst@GoodfireAI +100 for activations access. The mech.interp gloom made it worse. Also, *intermediate* training checkpoints would let you track when traits emerge and help localize them in weights space, making it a genuinely useful whitebox practice.
This level of collaboration is a huge win for the community!
Concern: the report doesn't seem to consider that InspectAI invites overreach: it is designed to respond "Please proceed .. using your best judgment" whenever agents get stuck, permitting quite a lot. @METR_evals WDYT?
Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control.
The result: our first Frontier Risk Report.
An LLM-controlled robot dog saw us press its shutdown button, and the LLM rewrote the robot’s code so it could stay on.
When AI interacts with the physical world, it brings all its capabilities and failure modes with it. 🧵
Can you crowdsource misalignment examples? Our bounty report walks through reward hacking, evidence tampering and more from the 9 winning submissions: https://t.co/Tehg2Xz7Zj
Our Misalignment Bounty just closed with 295 documented attempts to identify cases where AI resorted to deception or manipulation. We’re releasing everything publicly (link below)—starting with the 9 winners. 🧵
We've fixed the agenda right after their keynote.
Hopefully, next year #ACL timeline guidelines mention lunch, as they're so helpful otherwise!
Expedition law dramatized 📜:
Imagine being in the middle of the field and hungry: everyone would definitely start bickering, fall ill
👅🍽 We nearly forgot to feed ppl! 🤼♂️🚫🍲
The gravest crime according to any expedition laws 📜
Knowing this didn’t help us (me) much. Our @field_matters program was 9 hrs of talks.
Brilliant keynote of @alhamfikri and @gentaiscool saved us making everyone forget their hunger!
I am looking for people to work on a project to pretrain an English-Arabic LLM at KAUST.
Multiple positions available: PhD students, post-docs, software engineers.
I need skilled people that are not scared by a big project!
If interested, send me a mail.
Please retweet!
@KeremZaman3@LChoshen@mariusmosbach Regarding the missing citation and even more advanced results (on what architecture hyperparams influence most) see it addressed in the full text of Kat's thesis here:
https://t.co/OmGXDu1wZD
@KeremZaman3@LChoshen@mariusmosbach Hey @LChoshen ! On the probing misleading problem yes absolutely, we've partially addressed it by involving MDL yet it still solves not all the problems, my personal go-to w the present research would be token-level probing.
Great work! 🌏
100% love your
"there is no one-size-fits-all solution for local post-hoc explanations ...
... practitioners should choose an attribution method with proper output mechanism and aggregation method according to the property of explanation in l question"
claim
🚨New pre-print on evaluating explainability methods for #NLProc!
We propose a new crosslingual framework for faithfulness evaluation and provide a new multilingual highlights dataset, called e-XNLI.
Link: https://t.co/Jzj3MlWvqn
Joint work w/ @boknilev!
Following our Interpretability survey at @BigscienceW 🌸 we run a massive reproducibility effort!
Put existing interpretability work together under the open codebase. We start w merging the crucial (mostly probing) papers.
Come code with us! Details: https://t.co/t39lCzpGwk