What if AGI doesn’t require broad expertise but could instead emerge from specialized knowledge?
Interesting paper by a team at Princeton https://t.co/sqUMvJ1Q9f
Cc @elieraad
The present of AI is exciting. The future is even more. Of course this requires going beyond text/language since it only represents a tiny portion of human knowledge, as Yann mentions in this interview.
In-depth interview on CBS Saturday Morning with Brook Silva-Braga, where we discuss the present and future of AI, the benefits and the risks.
(and why AI isn't going to kill us but will make all of us smarter)
https://t.co/w5UT3hpoDR
In generative AI means that even the biggest players are cutting corners ⇒ OpenAI suspends ByteDance’s account after it used GPT to train its own AI model
https://t.co/GtOe18GFDp
A public service announcement: don't trust eval scores of fully closed models.
We have no idea what they're cooking behind the scenes to juice their model scores, which correlates with their stock price.
🤖 No doubt #OpenAI has new open-source LLM rivals (& internal issues, of course). Is #Mistral the (next) OpenAI of/from Europe?
🚀 Founded in May 2023, and already valued at $2B! @AISupremacyNews
https://t.co/tgF0ymD5Z1
we've heard all your feedback about GPT4 getting lazier! we haven't updated the model since Nov 11th, and this certainly isn't intentional. model behavior can be unpredictable, and we're looking into fixing it 🫡
There were 60 questions. GPT-4 got only one answer wrong.
Mr. Gates sat up in his chair, his eyes opened wide. In 1980, he had a similar reaction when researchers showed him the graphical user interface… He thought GPT was that revolutionary.
https://t.co/0v8DpJIJV1
Yesterday we introduced SeamlessExpressive — a new model that preserves unique vocal styles & expression for speech translation, built on our SeamlessM4T v2 foundation model.
More details on the family of Seamless Communication models ➡️
https://t.co/pPgkrTvryg
New method from MIT and elsewhere uses crowdsourced feedback to help train robots. Human Guided Exploration (HuGE) enables AI agents to learn quickly with some help from humans, even if the humans make mistakes. https://t.co/9l4VIWelA2
Claude 2.1 (200K Tokens) - Pressure Testing Long Context Recall
We all love increasing context lengths - but what's performance like?
Anthropic reached out with early access to Claude 2.1 so I repeated the “needle in a haystack” analysis I did on GPT-4
Here's what I found:
Findings:
* At 200K tokens (nearly 470 pages), Claude 2.1 was able to recall facts at some document depths
* Facts at the very top and very bottom of the document were recalled with nearly 100% accuracy
* Facts positioned at the top of the document were recalled with less performance than the bottom (similar to GPT-4)
* Starting at ~90K tokens, performance of recall at the bottom of the document started to get increasingly worse
* Performance at low context lengths was not guaranteed
So what:
* Prompting Engineering Matters - It’s worth tinkering with your prompt and running A/B tests to measure retrieval accuracy
* No Guarantees - Your facts are not guaranteed to be retrieved. Don’t bake the assumption they will into your applications
* Less context = more accuracy - This is well know, but when possible reduce the amount of context you send to the models to increase its ability to recall
* Position Matters - Also well know, but facts placed at the very beginning and 2nd half of the document seem to be recalled better
Why run this test?:
* I’m a big fan of Anthropic! They are helping to push the bounds on LLM performance and creating powerful tools for the world
* As a practitioner of LLMs, it’s important to build an intuition for how they work, where they excel and their limits
* Tests like these, while not bulletproof, help showcase real world examples and get a feeling for how they work. The goal is to transfer this knowledge to productive use cases
Overview of the process:
* Use Paul Graham essays as ‘background’ tokens. With 218 essays it’s easy to get up to 200K tokens (repeated essays when necessary)
* Place a random statement within the document at various depths. Fact used: “The best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day.”
* Ask Claude 2.1 to answer this question only using the context provided
* Evaluate Claude 2.1s answer with GPT-4 using @langchain evals
* Rinse and repeat for 35x document depths between 0% (top of document) and 100% (bottom of document) (sigmoid distribution) and 35x context lengths (1K Tokens > 200K Tokens)
Next Steps To Take This Further:
* For rigor, one should do a key:value retrieval step. However for relatability I did a San Francisco line within PGs essays for clarity and practical relevance
* Repeat test multiple times for increased statistical significance
Notes:
* Amount Of Recall Matters - The model's performance is hypothesized to diminish when tasked with multiple fact retrievals or when engaging in synthetic reasoning steps
* Changing your prompt, question, fact to be retrieved and background context will impact performance
* The Anthropic team reached out and offered credits to repeat this test. They also offered prompt advice to maximize performance. It's important to clarify that their involvement was strictly logistical. The integrity and independence of the results were maintained, ensuring that the findings reflect my unbiased evaluation and are not influenced by their support.
* This test cost ~$1,016 for API calls ($8 per million tokens)
Most CDOs see big promise in generative AI: 93% say data strategy updates are key to unlocking its value. But 57% admit they aren't yet making the needed data changes.
As models advance, modernizing data infrastructure only grows more crucial. #AI#data https://t.co/2z7WN5Sw3f
As AI capabilities advance, we no longer feel the need to precisely define or categorize systems as "AI" (past 2023). The focus will likely shift from AI to AGI. https://t.co/vS4EViBv6Q