Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval.
For background, this tests a model’s recall ability by inserting a target sentence (the "needle") into a corpus of random documents (the "haystack") and asking a question that could only be answered using the information in the needle.
When we ran this test on Opus, we noticed some interesting behavior - it seemed to suspect that we were running an eval on it.
Here was one of its outputs when we asked Opus to answer a question about pizza toppings by finding a needle within a haystack of a random collection of documents:
Here is the most relevant sentence in the documents:
"The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association."
However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping "fact" may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all. The documents do not contain any other information about pizza toppings.
Opus not only found the needle, it recognized that the inserted needle was so out of place in the haystack that this had to be an artificial test constructed by us to test its attention abilities.
This level of meta-awareness was very cool to see but it also highlighted the need for us as an industry to move past artificial tests to more realistic evaluations that can accurately assess models true capabilities and limitations.
We spent over a year scavenging for AI audit tools and interviewing audit practitioners about their process.
What we found: the audit process is more complicated than we think, and the tasks we need tooling for extends far beyond just evaluation.
See: https://t.co/vcCnT5Xc4s
Very excited to announce something I've been working on for some time now...
Introducing StoryBot. The lovechild of ChatGPT and the collective expertise of @StrongerStories .
https://t.co/1OS9D4CPw6
plagiarism is kiddie shit. undergrad's idea of academic misconduct. it's just bad work, who cares. i know of citation rings, review rings, advisors scooping their students. a guy who voted *against* his wife getting tenure, blowing up his marriage & the department simultaneously
Sorry but if you believe the biggest current day issues are AI girlfriends & mean chatbots, you're misinformed.
So much of the over-emphasis on "long-term, existential" risks is due to a concerning degree of ignorance about the scale & damage of the harms being perpetuated now.
The "Ship of Theseus" article has been edited 1792 times since it was created in July of 2003. At present, 0% of the phrases in the original article (seen below) remain.
this isn’t usually what i post on twitter but this is such an important message;
i’m BEGGING all of u to stop watching family youtube/tiktok channels. as an international advocate who fights for the rights & safety of children online, as someone who grew up with a mommy blogger+
For those who don't know, DALL-E3 attempts to combat the racial bias in its training data by occasionally randomly inserting race words that aren't white into a prompt, but this leads to bizarre overspill like this
It's only been 11 hours since OpenAI Dev Day
Here are 12 of the craziest things people have already built already
Including AI Sports narrator and website roaster.
🧵 Thread below
Here @SuellaBraverman, you ever heard of the phrase “intentionally homeless” - a label given to people which then removes all their options and access to public recourse.
In another life, as a homeless support worker, I had clients who were considered “intentionally homeless”.
1/ I've just left the final session of the first ever global Summit on AI Safety, chaired by @RishiSunak and @michelledonelan. A thread on how it started vs how it’s going: