AMDAHL'S ARGUMENT FOR AI
The productivity speedup AI apps can provide is limited by how much human-in-the-loop work is required. Humans are ~1-3 tokens per second. They can't really be sped up – unless you're @neuralink.
So if your application requires a human completion for every LLM completion (i.e. ChatGPT or AI Copilots) then your maximum speedup is ~2x – even when LLMs become 10x faster.
@cognition_labs Devin is better, because it needs a human completion only every ~10 iterations. At current speeds, this feels about 2.9x better than raw ChatGPT, which is nice, but not mind-blowing. But because they're frugal with human tokens, they can go to ~10x productivity speedup just by waiting for models to get faster!
The fun stuff starts when AI agents get to the 100-1000x range, i.e. only require human input every 100-1000 iterations. It's going to be a long way there – but I'm excited every time I see something that will get us closer: Like code execution from @e2b_dev, browsing from @browserbase and a context engine from @sid_ai.
Many copilots & current ChatGPTs will seem silly in hindsight: Like doing a 1 on 1 with your intern every 15 minutes – when you could be managing a team that does a month's worth of progress between every meeting.
Today, developers are frugal with LLM tokens (I know: they're expensive) – alas we've built tools to use them wisely: @PareaAI, @humanloop, @langfuse, @langchain. But the most important thing to be frugal with are human tokens (both input and output) – they will define the overall productivity speedup your application can provide. Humans are insanely slow.
AI agents don't yet work well – but it won't be a competition once they do. If you can think of one that does or you're working on one, please post it below!
Naturally, there are many caveats here: Iterations are gameable, and reducing human tokens has been an important trend outside of agents, too: Google let you find information with fewer keystrokes and reading than anyone else – same holds for @perplexity_ai today. Button presses can be tokens (depending on the action they trigger) etc.
Some chart explanations:
0. I pin human completions at 1 token per second in all calculations. That is realistic for high quality human tokens, although some people are faster or slower.
1. @BCG put this number at 1.4x. I'm fine disagreeing. The 1.8x is at current GPT-4 speeds.
2. Devin doesn't fully realize it's potential yet.
3. Let's free that y axis! "Future agent" only needs a human completion every 100 iterations.
"Uncertainty routing" is the real news in the Gemini announcement.
Without it, GPT-4 still beats Gemini in CoT@32 on MMLU!
For people building apps: GPT-4 is still better in zero-shot.
(charts from the Deepmind Gemini Technical Report)
This is almost unreal. The president of Brown University made these edits, in real time, while delivering her speech.
She cut "Star of David" and "yarmulke" from her remarks. Instead, she said students and faculty should only be able to proudly wear a keffiyeh or hijab.
Been a while but excited to show off a small AgentGPT 1.0 update ✨
Features:
- A freshly redesigned theme 🚀
- Updated agent handling with a focus on search 🌍
- @sid_ai for personalized context and queries 🦥
- Increased usage limits 📈
Try it at https://t.co/F8Nz4LGC0e
At YC, we work with founders throughout the life of their startup and beyond, to build the world-changing companies of the future.
And we'd love to work with you.
The deadline to apply for YC W24 is this Friday, Oct 13 at 8pm PDT — apply today at https://t.co/gNl84El3BS.
🚀 These YC companies make building with AI 10x easier!
🧑💻 An underrated aspect of @ycombinator is the wealth of resources the community has created for building with LLMs & AI.
🔖 Here's a list of tools you can start using today. You don’t want to miss bookmarking this!
🧵From testing & fine-tuning to infrastructure, the range of AI dev tools crafted by YC founders is incredible.
❓What’s a hard thing you wish was easy when building AI apps?
0/n ⚡ @ycombinator LAUNCH ⚡
✨ Connect customers' data to any LLM in an afternoon with @sid_ai
Live demo link at the end of the thread.
⛏️ just one button to connect all customer data
🤖 specialized LLM-based extraction & summarization
⚡ world-class retrieval pipeline (<200ms) – benchmarks soon
⏱️ information is synced and deprecated automatically
🚀 made for AI apps that scale to millions of users
We want to make it exponentially easier to build AI apps that people want. This is the first step.
🇨🇭PS: We're Swiss, so we know how to keep information secret.