GLM-5.3-Flash is here.
We benchmarked it against GPT-5.6 Luna with a suite of different web search solutions to see just how different the cost and performance would be.
The results were stark:
- GLM-5.3-Flash w/ Parallel Fast performed better on cost per task (~7x) and accuracy vs. GPT-5.6 with built-in web_search (11 pts).
- GPT-5.6 Luna w/ Parallel Fast delivers equal accuracy (76%) but at twice the cost vs. GLM 5.3 ($0.02 per task).
GLM 5.3 Flash lives up to the hype.
similar performance to GPT-5.6 Luna, half the per-task cost
our team benchmarked GLM-5.3-Flash against GPT-5.6 Luna with a variety of different web search options to see where this new model stacks up.
no surprises: using GLM 5.3 with Parallel Search tops the leaderboard and is up to 14x cheaper on end-to-end tasks.
What happens when there are more agents than eyeballs browsing the internet? How will the economics of the internet be reinvented?
@paraga thinks Shapley values may hold the answer.
Parag built Twitter over a decade, eventually becoming CEO and selling the company to Elon. He's spent the last three years building @p0, a search engine built for agents instead of humans.
His core argument:
1) human click data is a bug. Agents aren't just a new technology, they're a distinct customer – and the feedback loop that made Google great is the wrong signal for the thing actually doing the work;
2) the ad-funded web assumed scarce human attention. If agents show up instead of eyeballs, the business model underneath the internet has to be rebuilt, or good content stops being published.
The conversation covers:
— why he shipped a search agent before a search engine, and how that let him grow the index incrementally instead of buying a full web crawl up front
— the billion-to-billion matching problem: pull the right 1,000 tokens out of a trillion web pages, and your agent uses under half the tokens
— going from a 3-second compute budget to 200 milliseconds
— why he doesn't think Parallel is a neo-lab: "our output is a complement to a model"
— the Google Cloud deal — on GCP, your grounding options are now Google Search or Parallel Search
— reinventing the economics of the internet, Shapley values as a payment rail for publishers, and why the company was originally incorporated as Shapley Inc
— why routing 2-10% of inference spend to web data would dwarf every content business outside the walled gardens
— the web going from pull to push: "call me if this happens"
00:00 Introduction
03:25 What Is Web Search
05:17 Why Start a New Index
07:52 Search Agents First
10:17 Not a Neolab
13:14 Agents vs Google Search
19:38 Inside the Search Stack
28:59 Search Multipliers With Agents
30:21 Meeting Prep Agent Workflows
31:46 Quality Cost Latency And Turbo
32:42 Are Agents Overtaking Humans
34:28 Ads Model Meets Agent Web
37:20 New Incentives For Content
40:48 Shapley Values Attribution
47:46 Parallel Web And Future Vision
Hosted with my very unwilling co-host @andrew__reed and @sequoia
What happens when there are more agents than eyeballs browsing the internet? How will the economics of the internet be reinvented?
@paraga thinks Shapley values may hold the answer.
Parag built Twitter over a decade, eventually becoming CEO and selling the company to Elon. He's spent the last three years building @p0, a search engine built for agents instead of humans.
His core argument:
1) human click data is a bug. Agents aren't just a new technology, they're a distinct customer – and the feedback loop that made Google great is the wrong signal for the thing actually doing the work;
2) the ad-funded web assumed scarce human attention. If agents show up instead of eyeballs, the business model underneath the internet has to be rebuilt, or good content stops being published.
The conversation covers:
— why he shipped a search agent before a search engine, and how that let him grow the index incrementally instead of buying a full web crawl up front
— the billion-to-billion matching problem: pull the right 1,000 tokens out of a trillion web pages, and your agent uses under half the tokens
— going from a 3-second compute budget to 200 milliseconds
— why he doesn't think Parallel is a neo-lab: "our output is a complement to a model"
— the Google Cloud deal — on GCP, your grounding options are now Google Search or Parallel Search
— reinventing the economics of the internet, Shapley values as a payment rail for publishers, and why the company was originally incorporated as Shapley Inc
— why routing 2-10% of inference spend to web data would dwarf every content business outside the walled gardens
— the web going from pull to push: "call me if this happens"
00:00 Introduction
03:25 What Is Web Search
05:17 Why Start a New Index
07:52 Search Agents First
10:17 Not a Neolab
13:14 Agents vs Google Search
19:38 Inside the Search Stack
28:59 Search Multipliers With Agents
30:21 Meeting Prep Agent Workflows
31:46 Quality Cost Latency And Turbo
32:42 Are Agents Overtaking Humans
34:28 Ads Model Meets Agent Web
37:20 New Incentives For Content
40:48 Shapley Values Attribution
47:46 Parallel Web And Future Vision
Hosted with my very unwilling co-host @andrew__reed and @sequoia
@zachmoskow@joshk@p0 the score here measures how agents do on benchmarks with access to our apis. if you are asking what all we do to achieve these results - let's chat
New independent benchmarking for search shows that @p0 search is: (1) Highest quality (2) Quality vs. cost Pareto frontier (3) Quality vs. latency Pareto frontier
Quality, Cost, Speed -- pick two (usually).
With Parallel you get the best possible combinations of all three. @p0 just took the top spots on @ArtificialAnlys's new search index. Fun to watch @paraga, @travers00, @utkarsh and the whole p0 team keep raising the bar.
Great to have rigorous independent benchmarking for search that measures end-to-end quality, cost, latency for agents.
Parallel search is
(1) Highest quality
(2) Quality vs. cost Pareto frontier
(3) Quality vs. latency Pareto frontier
Announcing the Artificial Analysis Search Index, benchmarking how search API providers perform on quality, cost, and speed when used by an agent. We are initiating coverage with Parallel, Exa, Firecrawl, You (dot) com, Tavily, Keenable, and Brave
Search is one of the most important tools for agents. Search providers make different choices about how they search, rank, and package results, and those choices change what the model reads and how it acts. We are expanding our benchmarking coverage to search APIs, so developers can pick a search provider on measured quality, cost, and speed.
Each provider result pairs a search API provider with the same model, GPT-5.6 Luna (medium). The model runs inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web - only the search provider behind the search tool changes.
At launch, the leaderboard covers 11 results across 7 search providers, and we’ll keep expanding coverage as we look to provide the most accurate and comprehensive benchmarking of search providers for AI agent usage.
Key elements of the Artificial Analysis Search Index:
➤ Three equally weighted benchmarks: the Search Index is the average of DeepSearchQA (900 broad research questions that need many searches, graded with an F1 score over answer items), BrowseComp (a 200-sample hard subset of facts that need multi-hop browsing), and AA-Omniscience (a 600 question private subset, balanced across 6 domains)
➤ Same agent, different search provider: the agent has 25 turns available to complete each task. Its web search tool returns the search provider's native response payload (with content modes standardized to snippets), with a maximum of 10 results and contamination sources filtered out
➤ Model-only baseline: we compare search agent results to the same model answering single-shot without tools, showing how much each provider lifts the model above its internal knowledge
➤ Cost and Time per Task: we aggregate the time and cost spent on both model inference and search. This is key - search APIs have different cost and latency structures, but these can be offset where they help an agent use fewer turns and save on costly language model inference
Key results:
➤ Parallel, Exa, and Firecrawl have the strongest overall performance, with Artificial Analysis Search Index scores of 75, 74, and 73 respectively at launch
➤ All search providers tested substantially improve knowledge-based benchmark performance: the model only baseline scores 33 on the Search Index, while search-included provider results score between 65 and 75
➤ Focused search results reduce spend on model inference: Parallel Search (advanced) search costs more per task than Parallel Search (basic) but less per task in total ($0.084 vs $0.11). Higher quality results cut the model's token use by over 40% in this case, more than offsetting increased search costs while reaching higher benchmark scores
➤ Fast search calls do not guarantee fast tasks: Parallel Search (turbo) has the fastest average search calls among Parallel's tiers (0.51s per query vs 1.03s for Parallel Search (basic)) but the basic tier scores higher on quality (73 vs 67) and the two land close on total time per task