@alisawuffles I love this! subword tokens have some convenience advantages: for example if I have word labels that I want to map to token labels I can trust that it's a one-to-many mapping with subword tokenizers, but if there are multiple labels for a superword token I have to make a decision
Duolingo AI is hiring new PhDs and PhD interns! Come define the future of AI-driven education!
(New PhDs must graduate by Aug 2025, interns by Aug 2026)
https://t.co/ZmZqywIAvg
@BlancheMinerva oh interesting. I think I see the argument for this if "highly OOD contexts" means "(almost-)exclusively unseen tokens in the context" but I'm not clear if that would apply to observed tokens in very-low-probability sequences... would be interesting even just to test that part!
"you must make reasonable efforts to use the latest version" in the DBRX and Gemma licenses makes it feel so dishonest to call them "open models." like, yes, you can see the weights, but they can just take away all your usage rights by publishing a broken "update"
Wanna know gpt-3.5-turbo's embed size? We find a way to extract info from LLM APIs and estimate gpt-3.5-turbo’s embed size to be 4096. With the same trick we also develop 25x faster logprob extraction, audits for LLM APIs, and more!
📄 https://t.co/NdYU8ZhuVH
Here’s how 1/🧵
LMs are increasingly large🐘 and proprietary🔒 — what if we could “tune”🔧 them without accessing their internal weights?
Enter: proxy-tuning, which operates on only the *outputs* of LMs at decoding-time to achieve the effect of direct tuning!
📄: https://t.co/mx2SRlyTD2 1/
@mayhewsw yeah I'm pretty sure one is img2img of the other—everything has the same general shape and position but is different in details. black boxes / woman in the black sweater, light fixture, floor, table etc.
🚨 New Dataset Alert 🚨 I'm extremely excited to announce Universal NER v1, available now.
It is gold-standard human annotations of 18 datasets covering 12 languages, based on Universal Dependencies texts. This is the first data release of the UNER project.
1/3
at #ACL2023NLP ! I'll be presenting a poster at BEA on Thursday: work with @BenNaismithELT and @JillBurstein on discourse grading with GPT-4. come say hi :)
⚡️New paper!⚡️
It’s tempting to interpret chain-of-thought explanations as the LLM's process for solving a task. In this new work, we show that CoT explanations can systematically misrepresent the true reason for model predictions.
https://t.co/ecPRDTin8h
🧵
🌟Update🌟there's been a lot of debate of whether Theory of Mind (ToM) has emerged in new models (ChatGPT/GPT-3.5/GPT4), as people have reported qualitative/anecdotal evidence of good performance on these types of examples.
TLDR; still no neural ToM... 🧵
https://t.co/R6n4LgIwx5
wait,
what??
why do Bard AND ChatGPT *both* write an anodyne story about a young woman in idyllic "Willow Creek" at sundown???
(details: it's not deterministic, often you get a different town name, different phrasing, etc. broad strokes are similar though. gpt-4 does it too.)
@nostalgebraist Bard seems to prefer one of these two for me, but Willow Creek frequently shows up as draft #3, and Sarah and an old man named John recur (I got a young woman named Sarah and an old man named Mr. Johnson in ChatGPT).