Before Drake became who he is today, he was spending $5k/month renting Rolls Royce Phantoms
He’d also borrow his uncle’s car and drive women through the wealthiest neighborhoods in Toronto
At face value, it may come off as if he was just larping, or perhaps acting out of insecurity
But the truth is much simpler
Drake understood the psychological advantage of removing the novelty from the things you desire
As long as a certain life remains completely foreign to you
Some part of your mind will continue rejecting the possibility of it ever becoming your “normal”
So the ulterior lesson is that manifestation will remain nothing but mental masturbation so long as it isn’t followed by aggressive, and oftentimes risk-laden action
if you're using claude to build web scrapers, feed it this context before letting it write another 800 lines of Playwright
some interesting rabbit holes worth going down:
HAR diffing
capture two network sessions, change ONE variable between them and have claude compare the request bodies, parameters, response structures and dependency order.
repeat across searches, filters, pagination etc.
you can start piecing together the underlying query model without having to automate every interaction through the browser.
AST analysis
rather than manually digging through 40,000 lines of compiled javascript, have claude parse the relevant bundles into an abstract syntax tree using Babel or Acorn.
trace request constructors, shared API clients, pagination handlers and response transformations. have claude correlate what it finds against your HAR captures.
source maps
check for exposed `.js.map` files containing `sourcesContent`.
depending on the build configuration, these can reconstruct portions of the original readable frontend source.
feed claude the recovered modules alongside your network captures and have it map the API clients, field transformations and application state.
minification ≠ encryption.
hydration payloads
inspect the initial HTML for `__NEXT_DATA__`, Nuxt payloads and React Server Component flight data.
some frameworks serialize application state into the initial document, meaning the records you're looking for might already be present before client-side hydration even begins.
have claude compare these against subsequent network responses.
schema inference
capture 20–30 DIFFERENT JSON responses. different searches, record types, filters, etc.
have claude generate a union schema, identify optional fields, nested attributes, inconsistent types and null behaviour.
last thing you want is an extractor that works beautifully on your first 100 records and silently drops useful attributes across the next million.
response transformations
compare raw server responses against the actual records displayed in the UI.
frontend code can merge objects, rename attributes, calculate additional values or discard fields entirely before rendering.
have claude reconstruct those transformations and map the original response fields.
occasionally the interface is showing you substantially less information than the server actually returned.
incremental acquisition
look into ETags, Last-Modified, update timestamps, stable identifiers and content hashing.
where supported, these can help identify which records actually changed instead of collecting and processing the same dataset every single run.
especially interesting once you're dealing with millions of records and recurring refreshes.
data reconciliation
take samples from multiple sources and have claude construct a canonical schema.
normalize company domains and identifiers, compare timestamps, preserve field-level provenance and assign confidence scores where information conflicts.
also please stop deduplicating companies by name alone. you are going to merge 14 unrelated businesses called Apex Solutions into one entity.
extraction observability
have claude build checks for schema drift, unusual null rates, unexpected pagination behaviour, duplicate inflation and sudden changes in record counts.
a scraper crashing is whatever.
a scraper running successfully for 3 weeks while quietly producing incomplete datasets is a much more expensive problem.
also worth having it generate a small validation report after every acquisition rather than trusting a successful exit code.
there's a LOT more to scraping than making a browser click buttons.
understand the data pipeline before you automate it.
anyone who's turbo autistic and riddled with ADHD has to actively resist starting a new company every time something interesting catches their attention
AI can now clone a competitor's frontend and core backend in a day
dangerous amount of freedom to give us LOL
@Dylanmadden I see grok bot as a solution for non-techy people who still think chatgpt in the browser is the best it gets
for your stack, nah. not worth adding