Web search is holding back AI 👇🏻
Web search costs two orders of magnitude too much and is holding back large-scale AI adoption in enterprises. Here's a story that illustrates this.
An agentic risk analyst agent does 100 searches a day to find lawsuits, cyber-attacks, recalls, strikes, and other material events for 25,000 companies.
A dataset of 10 million results with 500 tokens each would cost around $5,000 to process with grok-4-fast. It would also cost <$5 a day to maintain.
BUT ... how much would those web searches cost? Well at $5 CPM it would be $12,500/day ($5 ÷ 1000 x 100 x 25,000). Yes. $12,500/DAY!
And that's only 25,000 companies. Imagine you're a big US bank tracking ~3 million business customers. You'd be spending a million dollars a day.
It's obvious that web search is the optimal form factor for AI tooling. It's also obvious that $5CPM for search looks utterly ridiculous next to token costs.
This story is not made up. This is exactly what I tried to build years ago. I honestly didn't think I would need to build a search engine first ... but here we are.
On Oct 29 at @SHACK15sf I'll present a talk "The End of Blind Spots: Agentic Search as an Enterprise Risk Model" at the BrightData Web Discovery Summit.
So, if you're in SF, and you feel like geeking out over web search you should definitely try attend. @RichardSocher and @paraga are also speaking!
BTW this is just the start. NOSIBLE would not be possible without the heroic efforts of the open source community. Contributing back is something we have always wanted to do but didn't have bandwidth for. Now that is starting to change we have identified a bunch of internal tools we want to open source. We are starting small (pun intended) but will ramp up in a big way over the next year. Exciting times ✌🏻
Last night we ran our first large test on T2. It did 2 Mn searches in under 2 hours and yielded 5 Bn tokens of point-in-time text. What is T2? T2 is the reason why the smartest quant funds in the world use NOSIBLE.
Generally available Q4.
Forget fancy visuals. THIS is what building a world class search engine actually looks like. Evals all day, every day. Even at 20h30. We are zealots. We are going to win.
Cybernaut-1 is here 🥁
Cybernaut-1 combines our hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries!
If you're building with AI, our new blog is worth the read. In it:
1) We make the case for why AIs need their own search engine.
2) We walk through the 8-stage retrieval pipeline that we use.
3) We share an update regarding cybernaut-1 (codename, diablo).
Hope you enjoy ✌🏻
I have wanted to share this story for a very, very long time. Here is the story of how two quants - in South Africa of all places - found a solution to vector search that scales at 1/100th the cost of HNSW with pareto-optimal trade-offs and unique capabilities 🧵👇🏻
I am excited to share this. Today we added a new search algorithm to our search engine. It uses a company knowledge graph to massively improve hybrid search (lexical + semantic). This right here is why we need a web-scale search engine built by quants for quants! 🔥
#LelapaDemoDay was a game-changer! From insightful panels to groundbreaking tech, we showcased how #Vulavula & #InkubaLM are revolutionising African language communication. We can’t wait for what’s next!
Read more👇
🔗 https://t.co/77JBW344ch
🔗 https://t.co/O6cgbcm8Gp
We’re thrilled! 🎉Siya Xuza’s Lethabo Galactic Bot, powered by our #Vulavula technology, just won 1st prize at the SAB Foundation #SocailInnovation Awards!
Learn how our #AI platform is revolutionising service management across African languages
🔗https://t.co/MooRf0UPuw
Our CTO @alienelf has been named a global top AI Innovator Under 35 by MIT Tech Review. She’s on a mission to ensure African languages benefit from generative AI.
Read more: https://t.co/Qel3Ys6RaI
InkubaLM: A small language model for low-resource African languages
A small language model with 0.4 billion parameters, which achieves performance comparable to models with significantly larger parameter counts.
https://t.co/3hEhVvsJUD @LelapaAI
🪲 the @LelapaAI InkubaLM paper is up on archive! Indulge your eyes. 🕺🏽
Shout out to the team for the incredible work
https://t.co/XY3nb8CyPB
🤗 InkubaLM
https://t.co/BRnHOAWzCu
🤗 Inkuba-Instruct
https://t.co/ZmyXJhW5KJ
🤗 Inkuba-mono
https://t.co/ehqj3TDQU2
@LelapaAI is proud to present an open source language model that stands small and powerful against the rest for African languages. Read more about its capabilities here here —>
p.s it is 0.4B (smol and mighty)
https://t.co/HWd2TFYPfq
At @LelapaAI , we have a lil beta SDK out for our African language NLP Vulavula! We're releasing it here so y'all can try it out and give feedback! 💚
- 📃 docs here: https://t.co/Gx6nxXBILL
- 📒 notebook here: https://t.co/hpSvxxiKcr
- 👩🏾💻 pip install vulavula
Inexpensive token generation and agentic workflows for large language models (LLMs) open up intriguing new possibilities for training LLMs on synthetic data. Pretraining an LLM on its own directly generated responses to prompts doesn't help. But if an agentic workflow implemented with the LLM results in higher quality output than the LLM can generate directly, then training on that output becomes potentially useful.
Just as humans can learn from their own thinking, perhaps LLMs can, too. For example, imagine a math student who is learning to write mathematical proofs. By solving a few problems — even without external input — they can reflect on what does and doesn’t work and, through practice, learn how to more quickly generate good proofs.
Broadly, LLM training involves (i) pretraining (learning from unlabeled text data to predict the next word) followed by (ii) instruction fine-tuning (learning to follow instructions) and (iii) RLHF/DPO tuning to align the LLM’s output to human values. Step (i) requires many orders of magnitude more data than the other steps. For example, Llama 3 was pretrained on over 15 trillion tokens, and LLM developers are still hungry for more data. Where can we get more text to train on?
Many developers train smaller models directly on the output of larger models, so a smaller model learns to mimic a larger model’s behavior on a particular task. However, an LLM can’t learn much by training on data it generated directly, just like a supervised learning algorithm can’t learn from trying to predict labels it generated by itself. Indeed, training a model repeatedly on the output of an earlier version of itself can result in model collapse.
However, an LLM wrapped in an agentic workflow may produce higher-quality output than it can generate directly. In this case, the LLM’s higher-quality output might be useful as pretraining data for the LLM itself.
Efforts like these have precedents:
- When using reinforcement learning to play a game like chess, a model might learn a function that evaluates board positions. If we apply game tree search along with a low-accuracy evaluation function, the model can come up with more accurate evaluations. Then we can train that evaluation function to mimic these more accurate values.
- In the alignment step, Anthropic’s constitutional AI method uses RLAIF (RL from AI Feedback) to judge the quality of LLM outputs, substituting feedback generated by an AI model for human feedback.
A significant barrier to using LLMs prompted via agentic workflows to produce their own training data is the cost of generating tokens. Say we want to generate 1 trillion tokens to extend a pre-existing training dataset. Currently, at publicly announced prices, generating 1 trillion tokens using GPT-4-turbo ($30 per million output tokens), Claude 3 Opus ($75), Gemini 1.5 Pro ($21), and Llama-3-70B on Groq ($0.79) would cost, respectively, $30M, $75M, $21M and $790K. Of course, an agentic workflow that uses a design pattern like Reflection would require generating more than one token per token that we would use as training data. But budgets for training cutting-edge LLMs easily surpass $100M, so spending a few million dollars more for data to boost performance is quite feasible.
That’s why I believe agentic workflows will open up intriguing new opportunities for high-quality synthetic data generation.
[Original text: https://t.co/zOiuUFtmo3 ]