Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval.
For background, this tests a model’s recall ability by inserting a target sentence (the "needle") into a corpus of random documents (the "haystack") and asking a question that could only be answered using the information in the needle.
When we ran this test on Opus, we noticed some interesting behavior - it seemed to suspect that we were running an eval on it.
Here was one of its outputs when we asked Opus to answer a question about pizza toppings by finding a needle within a haystack of a random collection of documents:
Here is the most relevant sentence in the documents:
"The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association."
However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping "fact" may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all. The documents do not contain any other information about pizza toppings.
Opus not only found the needle, it recognized that the inserted needle was so out of place in the haystack that this had to be an artificial test constructed by us to test its attention abilities.
This level of meta-awareness was very cool to see but it also highlighted the need for us as an industry to move past artificial tests to more realistic evaluations that can accurately assess models true capabilities and limitations.
In ein paar Stunden flimmert der #AppleEvent über die Bildschirme.
Mehr noch als scary Tempo wünschte ich mir von neuen Macbook Pro etwas anderes.
https://t.co/0FcH4UJTCN
One marketing trend we’ll see in the next decade will be traffic becoming as rare as water in the desert.
I hope I’m wrong but things aren’t looking great.
Search engines will use AI to give answers right away. People won't need to click on any websites to find what they want.
Social media platforms will be under pressure to make more money. They will cut organic reach forcing companies to pay to get their posts seen.
Paid ads will also get more expensive as more businesses compete for the same space. This has been going on for a long time and it won't stop.
But this will also be a great opportunity for marketers who can deliver results.
They’ll be worth their weight in gold.
Despite it being easier than ever to build a product, there are many macro themes making it harder to grow.
* PLG: easier to copy
* SEO: threatened by AI chat
* Paid: lost efficiency post-ATT
* Social: moving to interest graph
What growth tactics have you seen work in 2023?
Everyone is talking about OpenAI, Microsoft and Stability as a huge AI player.
But there is one company that arguable the best positioned in the space, and they dont even have an LLM offering yet.
Apple has the potential to shift the entire AI landscape, here's how:
Elon Musk is the master of pitching.
He's sold spaceships, flamethrowers and electric cars.
And his framework can be applied to any pitch.
10 simple storytelling tips to nail your next pitch: