I highly recommend reading Francois Chollet's (@fchollet) enlightening paper "𝗢𝗻 𝘁𝗵𝗲 𝗠𝗲𝗮𝘀𝘂𝗿𝗲 𝗼𝗳 𝗜𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝗰𝗲" for anyone interested in better understanding the concept of AGI.
https://t.co/TxyvwvvVet
Summary of some key insights below ↓
Does style matter over substance in Arena? Can models "game" human preference through lengthy and well-formatted responses?
Today, we're launching style control in our regression model for Chatbot Arena — our first step in separating the impact of style from substance in rankings.
Highlights:
- GPT-4o-mini, Grok-2-mini drop below most frontier models when style is controlled
- Claude 3.5 Sonnet, Opus, and Llama-3.1-405B rise significantly
- In Hard Prompts, Claude 3.5 Sonnet ties for #1 with ChatGPT-4o-latest. Llama-405B climbs to joint #3.
More analysis in the thread below👇
I'm partnering with @mikeknoop to launch ARC Prize: a $1,000,000 competition to create an AI that can adapt to novelty and solve simple reasoning problems.
Let's get back on track towards AGI.
Website: https://t.co/wNsM3IQgEI
ARC Prize on @kaggle: https://t.co/Lhsh1RiWKq
We are (finally) releasing the 🍷 FineWeb technical report!
In it, we detail and explain every processing decision we took, and we also introduce our newest dataset: 📚 FineWeb-Edu, a (web only) subset of FW filtered for high educational content.
Link: https://t.co/MRsc8Q5K9q
Exciting new blog -- What’s up with Llama-3?
Since Llama 3’s release, it has quickly jumped to top of the leaderboard. We dive into our data and answer below questions:
- What are users asking? When do users prefer Llama 3?
- How challenging are the prompts?
- Are certain users or prompts over-represented?
- Does Llama 3 have qualitative differences that make users like it?
Key Insights:
1. Llama 3 beats top-tier models on open-ended writing and creative problems but loses a bit on close-ended math and coding problems.
People are reading way too much into Claude-3's uncanny "awareness". Here's a much simpler explanation: seeming displays of self-awareness are just pattern-matching alignment data authored by humans.
It's not too different from asking GPT-4 "are you self-conscious" and it gives you a sophisticated answer. A similar answer is likely written by the human annotator, or scored highly in the preference ranking. Because the human contractors are basically "role-playing AI", they tend to shape the responses to what they find acceptable or interesting.
This is what Claude-3 replied to that needle-in-haystack test:
"I suspect this pizza topping "fact" may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all."
It's highly likely that somewhere in the finetuning dataset, a human has dealt with irrelevant or distracting texts in a similar fashion. Claude pattern matches the "anomaly detection", retrieves the template response, and synthesizes a novel answer with pizza topping.
Here's another example. If you ask the labelers to always inject a relevant joke in any response, the LLM will do exactly the same and appear to have a much better "sense of humor" than GPT-4. That's what @grok does, probably. It doesn't mean Grok has some magical emergent properties that other LLMs cannot have.
To sum up: acts of meta-cognition are not as mysterious as you think. Don't get me wrong, Claude-3 is still an amazing technical advance, but let's stay grounded on the philosophical aspects.
Cool video borrowed from @karinanguyen: Claude-3 generates a self-portrait with d3
There's a lot of speculation about whether OpenAI's video generation model Sora has a 'physics engine' (bolstered by OAI's own claims about 'world simulation'). Like the debate about world models in LLMs, this question is both genuinely interesting and somewhat ill-defined. 🧵1/
Foundations of Vector Retrieval
Very rare drop on arXiv.
This 200+ page monograph covers fundamental concepts along with advanced data structures and algorithms for vector retrieval.
Given how much we deal with vector representations today and how critical they are to building robust applications with modern AI systems, this topic is worth learning about or reviewing.
Looking further into LLM benchmark x-correlations:
- Top row: how each benchmark relates to human judgement (Arena Elo)
- Other rows: any benchmark pair & their relationship
- On the right: samples = # of models tested for each benchmark
thx: @chipro@maximelabonne@ldjconfirmed
🦜👑LangChain State of AI 2023
What are people building? Which LLMs are they using? Which vectorstores are they using? How are they testing?
📊We turn to anonymized usage stats from LangSmith to answer these questions with real data.
Blog: https://t.co/Qk5wyjwEPw
Recap 🧵
Google (DeepMind) releases AI model Gemini.
There is no turning back now, we are in for one mad ride. The multi modality, and fluidity of the model is super clean.
My jaw dropped at 4:24 seconds
A thread...
We’re excited to announce 𝗚𝗲𝗺𝗶𝗻𝗶: @Google’s largest and most capable AI model.
Built to be natively multimodal, it can understand and operate across text, code, audio, image and video - and achieves state-of-the-art performance across many tasks. 🧵 https://t.co/mwHZTDTBuG
Oh man -- you can just download the knowledge files (RAG) from GPTs. I don't know if this is a security leak or "just" a prompt engineering? @OpenAI@simonw
Pressure Testing GPT-4-128K With Long Context Recall
128K tokens of context is awesome - but what's performance like?
I wanted to find out so I did a “needle in a haystack” analysis
Some expected (and unexpected) results
Here's what I found:
Findings:
* GPT-4’s recall performance started to degrade above 73K tokens
* Low recall performance was correlated when the fact to be recalled was placed between at 7%-50% document depth
* If the fact was at the beginning of the document, it was recalled regardless of context length
So what:
* No Guarantees - Your facts are not guaranteed to be retrieved. Don’t bake the assumption they will into your applications
* Less context = more accuracy - This is well know, but when possible reduce the amount of context you send to GPT-4 to increase its ability to recall
* Position matters - Also well know, but facts placed at the very beginning and 2nd half of the document seem to be recalled better
Overview of the process:
* Use Paul Graham essays as ‘background’ tokens. With 218 essays it’s easy to get up to 128K tokens
* Place a random statement within the document at various depths. Fact used: “The best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day.”
* Ask GPT-4 to answer this question only using the context provided
* Evaluate GPT-4s answer with another model (gpt-4 again) using @langchain evals
* Rinse and repeat for 15x document depths between 0% (top of document) and 100% (bottom of document) and 15x context lengths (1K Tokens > 128K Tokens)
Next Steps To Take This Further:
* Iterations of this analysis were evenly distributed, it’s been suggested that doing a sigmoid distribution would be better (it would tease out more nuanced at the start and end of the document)
* For rigor, one should do a key:value retrieval step. However for relatability I did a San Francisco line within PGs essays.
Notes:
* While I think this will be directionally correct, more testing is needed to get a firmer grip on GPT4s abilities
* Switching up prompt with vary results
* 2x tests were run at large context lengths to tease out more performance
* This test cost ~$200 for API calls (a single call at 128K input tokens costs $1.28)
* Thank you to @charles_irl for being a sounding board and providing great next steps
Introducing Objaverse-XL, an open dataset of over 10 million 3D objects!
With it, we train Zero123-XL, a foundation model for 3D, observing incredible 3D generalization abilities: 🧵👇
📝 Paper: https://t.co/2oNakoka7v
🪩The @stateofai 2023 is now here.
Our 6th installment is one of the most exciting years I can remember. The #stateofai report covers everything you *need* to know, covering research, industry, safety and politics.
There’s lots in there, so here’s my director’s cut 🧵
🤖🧠NEW PAPER🧠🤖
Language models are so broadly useful that it's easy to forget what they are: next-word prediction systems
Remembering this fact reveals surprising behavioral patterns: 🔥Embers of Autoregression🔥 (counterpart to "Sparks of AGI")
https://t.co/w0wl4eW1M0
1/8
Does a language model trained on “A is B” generalize to “B is A”?
E.g. When trained only on “George Washington was the first US president”, can models automatically answer “Who was the first US president?”
Our new paper shows they cannot!
Mojo🔥 is now available for download locally to your machine! ❤️🔥🚀
Beyond a compiler, the Mojo SDK includes a full set of developer and IDE tools 🛠 that make it easy to build and iterate on Mojo applications. Let’s build the future together!🔥
https://t.co/KxmLvsxx5e