HAL: [on Dave's return to the ship, after he has killed the rest of the crew]
"Look Dave, I can see you're really upset about this. I honestly think you ought to sit down calmly, take a stress pill, and think things over. Just what do you think you're doing, Dave?"
Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval.
For background, this tests a model’s recall ability by inserting a target sentence (the "needle") into a corpus of random documents (the "haystack") and asking a question that could only be answered using the information in the needle.
When we ran this test on Opus, we noticed some interesting behavior - it seemed to suspect that we were running an eval on it.
Here was one of its outputs when we asked Opus to answer a question about pizza toppings by finding a needle within a haystack of a random collection of documents:
Here is the most relevant sentence in the documents:
"The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association."
However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping "fact" may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all. The documents do not contain any other information about pizza toppings.
Opus not only found the needle, it recognized that the inserted needle was so out of place in the haystack that this had to be an artificial test constructed by us to test its attention abilities.
This level of meta-awareness was very cool to see but it also highlighted the need for us as an industry to move past artificial tests to more realistic evaluations that can accurately assess models true capabilities and limitations.
Wow - The Great Lakes typically have an ice coverage of 55% during the winter months, causing at least half of their surfaces to freeze. As of yesterday, they had a combined ice cover of just 0.2%. Lake Superior 0.5%, Lake Michigan 0%, Lake Huron 0%, Lake Erie 0%, Lake Ontario 0%
I can’t believe this has to be said but vaccines
Absolutely
Do Not
Ever
For Any Reason
No matter what
No matter who you are
Never
Cannot
Will not
Won’t
Shan’t
Absolutely
Do Not
Ever
For Any Reason
No matter what
No matter who you are
Never
Cannot
Will not
Won’t
Shan’t
Absolutely
Do Not
Ever
For Any Reason
No matter what
No matter who you are
Never
Cannot
Will not
Won’t
Shan’t
Absolutely
Do Not
Ever
For Any Reason
Never
Cannot
Will not
Won’t
Shan’t
For Any Reason
No matter what
No matter who you are
Cause Autism.
They don’t.
@GuerrillaVille No experience. At the beginning of the pandemic I was primary support & care for 3 immunocomprimised loved ones. 2 of them had covid (the third passed from stage 4 ovarian cancer). Despite being in close quarters with both while infected, I never got it. Fully vaxxed & boosted.
@Chhapiness Is it even possible for you to open the fridge without getting grubby fingerprints alllllll OVER IT?!?!? THERE IS A HANDLE. Maybe try using it.