Introducing Prime Agent:
A self-improving RLM harness for coding and long-running autonomous tasks.
Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
@parmita I used to think everyone in a professional job was on the same level of thinking but most people have little agency and way too much negative steering
A talk on verifiable environments for agents in biology
0:00 Where does experimental data come from?
1:41 Modern bio research organized around measurement
2:44 Data analysis is an executable substrate for science
3:23 Five years of learnings building products for pharma
4:13 Trying to use coding models for biology workflows
5:40 Why frontier models can't be trusted (yet)
6:44 Sequencing based spatial analysis
7:38 SpatialBench design principles
8:26 Anatomy of a single evaluation
9:40 Human verification and early long horizon extensions
13:12 Grading functions and the case against rubrics
14:07 New work in multi-omics, therapeutics, biosecurity
@Dr_CSWright@AnthropicAI The ability to justify the direction of the experiment would be interesting
We might tell Claude that it’s a famous actress playing a corrupt politician haha
How Yoga Nidra puts you in "edit mode" for the unconscious mind:
"In Yoga Nidra, you're in a trance, and in that trance you're in the edit mode for your unconscious mind."
"A sankalpa is a resolve that you put in there, and that's where the rewriting comes in."
"A being statement is the pluripotent stem cell of where you want to go."
"Willpower is necessary when you are trying to not be narcissistic. It is not necessary when you are no longer narcissistic."
"Why not just change the tendency? That's literally what we do in psychotherapy every day."
@dr_alokkanojia on @hubermanlab
@chamath Delivering AI with the same consistency and amount and price would make AI commoditized. Model companies are simply the oil fields and the raw intelligence still needs to be refined to standardize across capability and behavior.
@yoyozhang06 My company has zero revenue but the police force and utility company in our area takes our advice as mandates. Does power count as traction?
100,000 stars on @graphify
First Indian developer to ship a 100k-star repo.
Only the 5th @ycombinator -backed open-source project to ever hit 100k.
Stars are vanity until they aren’t.
100k is where they stop being vanity.
@glangley@ycombinator Can you start telling the police departments you work with to follow protocol on those alerts?
Just because an officer is alerted about a stolen car doesn’t mean it’s stolen. Trust me, I know from experience, but either way I appreciate the pending settlement.
The demand for tokens is equal to the sum of all human intelligence as measured by salaries
Since most people aren’t working at full capacity, the range in cost actually varies wildly by task or role
Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems.
The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.
"GTM iS EaSY"
GTM:
> turn off open tracking
> no link in the first email, ever
> never send from your root domain
> verify emails at send, not at import
> 30 sends a day per mailbox, no more
> plain text, no images, no tracking pixels
> score positive replies, not total replies
> waterfall three data providers, never trust one
> enrich your tier one accounts, not the whole list
> use local data providers for emea and apac lists
> don't discard catch-alls, score them and send anyway
> track headcount deltas by department, not by company
> filter the list by mx record before you write any copy
> funding rounds are the most spammed signal in outbound
> 300 contacts per variant or you are just reading noise
> a first sales ops hire beats a series b as a trigger
> rebuild the list monthly, a third of titles rot every year
> a job post naming your competitor is a displacement signal
> read their job posts for the tech stack, it beats builtwith
> two customers asking for the same integration is a channel
> ai personalization on a bad list only amplifies the bad list
> write the first email to be forwarded down, not read at the top
> champion job changes are your warmest list and they cost nothing
> target whoever owns the pain, then ask them who owns the budget
> measure at meetings held, everything upstream of that is a proxy
> a new vp has 90 days to swap vendors, that window is your campaign
> soc2 on their trust center means they just started selling upmarket
> export closed-won, enrich it, and find the three attributes every account shares that your icp doc never mentions
> find where the champions from your last 20 lost deals work now, and open with the workflow you already know they run
> pull your competitor's sitemap, grab /customers and /case-studies, enrich every logo you find, that's the best list you'll ever build
> scrape every job post that names your competitor, filter to the last 30 days, and send the displacement email to the hiring manager, not the recruiter
> take your churned power users, find their new employers, and open with the exact workflow they used to run, nobody else hitting their inbox knows that