Introducing our Browser Agent Evals.
We benchmark frontier and open models using various harnesses on computer-use tasks.
Our evals are fully open source and reproducible via our Evals CLI.
Jev wasn't meant to build Agents.
Instead, we added it to @Stagehanddev's AI-powered primitives Act, Extract, and Observe.
Rather than a fully autonomous agent, we use Jev as a decision layer: which element to click, field to fill, or text to extract.
it seems like agentic commerce is finally solved with stripe link. and we're excited to partner with stripe to enable all agents to transact online.
your agent can now natively make payments on the web using link cli + browserbase.
one shot prompt: https://t.co/KDHq9Ca6yT
we built blazing fast computer/browser use with Jev + @Stagehanddev.
this task cost $0.001 and executed at near instant speed (in a remote browser btw)
the loop: observe the page, send a11y tree as state + actions as questions, Jev decides the next action, then Stagehand executes it.
you can build a price-monitoring tool on @browserbase in <5 minutes using fetch w proxies
running this for 3 products costs roughly $1.50 a day if i price-check every 15 minutes.
Muse Spark 1.3 from @Meta tops the charts in our Browser Agent benchmarks using the @mastra agent harness.
In this harness, Muse outperforms Opus 5 in cost, speed, and accuracy.
Introducing our Browser Agent Evals.
We benchmark frontier and open models using various harnesses on computer-use tasks.
Our evals are fully open source and reproducible via our Evals CLI.
We just made Astra's computer use 2.5x faster with Stagehand.
Astra works by executing code against the a11y tree, we built a translator that turns Astra's playwright commands into @Stagehanddev.
Astra chose to batch the entire logo into one command and one-shotted it.
The Browserbase dashboard got a makeover.
In the last few months we've shipped a ton of new features and improved our dashboard's UI. Here are some of our favorites.
Google's response to Fable 5.1 is another flash model with Gemini 3.8 flash.
It landed an 83.33% accuracy on our browser agent evals, while being the cheapest model by wide margin at $0.07 per task.
They haven't launched a benchmark-topping model since February, and seems like they're more focused on real-life applications and price to performance.
WebMCP is a game changer for browser agents
I built a skill that adds WebMCP support to any codebase - in this demo, I turned Excalidraw into an agent-native experience in a few minutes
then I used @Stagehanddev to control the canvas and draw anything. works on any open-source codebase
We got early access to Fable 5.1 and evaluated it extensively across our internal benchmarks.
It outperformed all other frontier models (including Fable 5) on accuracy scoring 92.11%, while being 60% cheaper than Fable 5 per task.
Excited to announce we've been recognized as a 2026 IA40 winner for the second year in a row.
The IA40 recognizes the 40 most important private companies in applied AI. Thank you, Madrona, AWS Startups, Microsoft, Google for Startups, NYSE, Delta Air Lines, and McKinsey & Company for the recognition.
Agents need their own identity to do real work on the web. We partnered with @tryramp to automate our event logistics workflow.
Our agent logs into sites with a Browserbase Context, downloads receipts, and submits them via the Ramp CLI.
Create agent-ready web apps for the WebMCP Challenge → https://t.co/j3VdqlcWvD
We're excited to see ChatGPT supporting WebMCP, the experimental standard that lets users and AI agents navigate websites together. In the @OpenAIDevs WebMCP Challenge you can experiment with it and win prizes. Plus, Chrome's @sarah_edo is on the judging panel! What will you create?