Iโm Stan. I build @nodemaven , proxy infrastructure for people running scraping, browser automation and AI agents in prod.
I write about what breaks between the demo and the ops: IP quality, sessions, browser identity, reliability and cost.
A whole market sells one promise, you never touch a scraper
one vendor's homepage shows 10 billion records delivered, another counts 15,178 customers
here is the map for august 2026
who you call when you want the data and not the scraper
managed extraction, you send a spec and datasets come back
@scrapehero, @grepsr, @promptcloud, @DatamamScraping, ficstar, xbyte enterprise crawling, @actowizsolution, @crawl_now, autoscraping , @Prowebscraper
enterprise data on demand
@zytedata, BrightData, @Oxylabs_io, @nimble_search
ai-native extraction, the model figures out the page instead of a human writing selectors
@diffbot, @riveterhq, kadoa
the marketplace layer
@DataradeHQ, aws data exchange, @Snowflake marketplace
the product in every provider row is the same
somebody else runs the scraper, watches the quality and eats the anti-bot problem
even if you never touch the scraper, somebody still buys the addresses it runs through
that layer is what we provide at @nodemaven
for those of you running scrapers in production
what would you never outsource, and what are you happy to just buy?
Scrapling has 75,6k github stars
its stealth mode runs on patchright, which has 4k
5 open source scraping tools, tested and validated by our team
@Scrapling_dev, by @D4Vinci1
the adaptive parser re-finds elements after a site redesign, so markup changes do not always mean rewriting selectors
catch: the stealth engine switched from camoufox to patchright at 0.3.13, older write-ups describe a different browser
patchright
change the import and the rest of your playwright code stays as it is, with the automation tells patched out
catch: chromium only, and it disables the console api. anything you built on console output stops working
camoufox
fingerprints spoofed inside the firefox c++ layer, where javascript cannot reach them
catch: it is a patched firefox fork. firefox is less commonly used by regular web users, so anti-bots always score it with less trust
pydoll
chrome over cdp with no webdriver flag anywhere
catch: the captcha feature is not a solver. the docs say so plainly, it automates the same click a person would make
google-maps-scraper
self-hosted maps extraction in go with a rest api
catch: fast mode is still beta and returns up to 21 results per query. wider coverage means grid mode across a bounding box
all of that is fixable in code. a bad ip is not, which is why you still need decent residential proxies
we run those at @nodemaven
which of these are you running, and what broke first?
i spent 10 years on anti-detection
first @multilogin_en on the browser side, then @nodemaven on the ip side
we finally shipped the two as one product
scraping browser, a cloud browser you talk to over cdp
you generate a cdp url and point playwright or puppeteer at it
the scraping logic stays yours, the environment it runs in becomes our job
every session starts with a managed fingerprint, automatic captcha solving and a clean residential or mobile ip
in the dashboard
templates for e-commerce, b2b enrichment, real estate and price monitoring
ai prompts that return structured json
live browser and session recordings to debug runs
profiles that keep cookies and local storage
up to 50 parallel sessions
live on every rotating residential plan, start with a template
nextbrowser is alive
three months ago i wrote that it was dead
the cloud version still is
we closed it for very boring reasons
sessions ate 4 to 8 gigabytes of ram each, the codebase needed a full rewrite, and i did not see a market big enough to justify both
so we stopped
i think that was the right decision
the team did not stop
they rewrote it from scratch, 47 releases since the beginning of july, and they were right
so we opened it
agpl, every line public, free
now it is a desktop app that runs local ai agents in managed browser sessions
profiles, proxies, fingerprint rotation, schedules, live view in one place
in may i wrote that every closed product is a tuition payment
this is what the tuition bought
repo in the first reply
Five open source anti-bot repos, 32k github stars between them
tested by our team, here is what each one covers
a public page sees your ip and asn, then tls, then http/2. canvas and webgl matter only once a script runs
you install a second stealth library and nothing changes
both cover the same layer, and the block is somewhere else
curl_cffi
impersonates browser tls and http/2 fingerprints from python
runs no javascript, so js challenges are where it stops
fingerprint-suite
@apify toolkit for node
one call gives a playwright context where headers and js apis come from one generated fingerprint
obscura
independent rust and v8 engine without chromium, 21k stars at v0.2.0
rendering and web api behavior can differ from chromium
trawl
escalates from plain http to a cached session, then a fresh browser solve, then a residential proxy
if the early steps work it never reaches the proxy
stealth-browser-mcp by @ImVibhek
a stealth browser over mcp, pinned to nodriver 0.47.0
you get that driver's fixes only when someone bumps the pin
start with curl_cffi when you do not need page javascript
it is the cheapest thing to try
check the recent commits before you install any of them ๐
@OpenAI turned off Atlas on August 9, 9 months after launch, and keeps the agent side alive inside ChatGPT, including a cloud browser on its own servers.
Here is the map for August 2026: who runs the browser when your agent has to act on a page instead of reading it.
The Atlas category: @comet by Perplexity, @FellouAI, @strawberry, @genspark_ai
Cloud browser runtimes: @browserbase, @AnchorBrowser, @steeldotdev, @usekernel, @hyperbrowser, @browserless, @lightpanda_io
Open source agent frameworks: @browser_use, @Stagehanddev, @skyvernai, @nottecore
No-code task agents: @airtop, @browseract, @BrowseAI, @browserbots
The primitives all of them sit on: @playwrightweb, Puppeteer, CDP
On September 15 Cloudflare changes the defaults for new domains: the Agent category gets blocked on pages that display ads.
@Cloudflare defines that category as automated activity acting in real time on a person's behalf. That definition covers every tool above.
Each of those sessions ends at the same gate: an IP and a fingerprint the site judges as one unit, so one burned address kills the run no matter which of them started it.
Those of you that run agents or scraping in production: which of these is in your stack?
gpt-5.6 just went public ๐ฅ
the loudest number is on the pricing page
the intelligence you paid full price for in june now costs half
three models at once
sol for hard agentic work
terra for gpt-5.5 level output at half the price
luna for volume
what builders will feel first
ultra mode - sol spawns subagents and runs parts in parallel
speed - sol on cerebras at up to 750 tokens per second
gpt-live - full duplex voice that listens while it speaks and jumps in on its own
voice stops being a walkie-talkie
it becomes a real conversation
cheaper faster and now it talks
the web this lands on did not change
and that gap is where our whole industry lives
what are you pointing sol at first?
@ChatGPTapp
@Cloudflare blocked over 416 billion ai bot requests
and the industry that still gets the data out just produced a 275 million dollar exit
here is the map for july 2026
full stack platforms
@zytedata , @apify , @BrightData , @Oxylabs_io
scraping api specialists
@ScraperAPI , @ScrapingBee , scrapfly, https://t.co/ziMAIaSgAa , @ZenRowsHQ , @scrapingdog , crawlbase
ai native extraction
@firecrawl , @scrapegraphai , @JinaAI_
search apis for agents
@tavilyai , exa , @p0 , @serp_api , serper , @dataforseo
open source
scrapy , @crawl4ai , @crawleecloud , scrapling
the money went to the search layer
@tavilyai sold for 275m to @nebiusai
parallel raised at around 2b valuation
every request still exits through the same gate
an ip and a browser fingerprint the site judges as one unit
one burned ip kills the run no matter which tool sent it
that gate is where we work
@nodemaven filters the addresses before they reach your scraper
those of you running scraping in production
which of these tools is in your stack
@rbatista19 great article! seems this Verify wall is starting to massively affect scrapers. did you manage to get the JS payload? our team is looking into it too. if you have the payload - we can help you analyze it and share results what happens under the hood.
the ip was in berlin
the browser timezone said chicago
the run died on the first request
and of course the proxy got the blame
i see this every week
in most cases the ip was fine
the person just ran a mismatched browser profile
sites read everything at once
ip geolocation
accept-language
timezone from javascript
webrtc
together they describe a person that does not exist
anti-bot systems catch exactly that
years looking at fingerprints at @multilogin_en and ips at @nodemaven taught me one thing
browser environment
infrastructure
and ip
all have to agree
an expensive proxy will not fix a sloppy profile
a perfect profile will not fix burned ips
at @nodemaven we remove that last variable
every ip passes our quality filter before it reaches you
before you launch a profile check at least these
timezone vs exit ip country
locale vs exit ip country
webrtc for leaks
geolocation api for the same mismatch
takes about two minutes
if anything disagrees with the ip fix that first
scrapers what got you blocked more often
bad ips or bad profiles
detection builders which mismatch is the loudest tell right now
Four months ago I open-sourced webclaw. Today it just passed 2k GitHub stars
One founder. 57 releases. No funding, no ads. Developers spread it.
Turns any URL into clean Markdown or JSON for AI agents.
Rust, AGPL, self-hostable.
Four months ago I open-sourced webclaw. Today it just passed 2k GitHub stars
One founder. 57 releases. No funding, no ads. Developers spread it.
Turns any URL into clean Markdown or JSON for AI agents.
Rust, AGPL, self-hostable.
two days of q3 planning at nodemaven
i spoke maybe 17% of the time
thatโs what iโm proudest of
as ceo my goal is simple
speak as little as possible
negotiate achievable objectives with the team
leave the room only with shared understanding
every idea gets roasted including mine
product challenges marketing. marketing storms product. sales pushes back
in hybrid meetings the key rule
remote people must be able to interrupt
if the loudest voice in the room sets the pace you lose half the brains
we spent most time on the bigger picture
where the internet goes in the age of ai
clean verified data is becoming the rarest thing
nodemaven wants to be the layer that provides it
big things are coming
founders whatโs your one rule for great hybrid meetings