New Princeton + Sentient Labs paper shows a coding agent can cost 40x more per solved task by changing only the harness around it, leaving the model untouched.
and controlled swaps show those scaffold differences barely move accuracy.
that the same model passes about the same tasks on any harness, so a leaderboard score does not have to mean the cost you will pay.
The problem is that leaderboards rank by model name while leaving the scaffold undisclosed. The agent may keep taking turns, but those turns can stop editing files or running commands and just burn context.
The paper fixes this by holding model, prompt, sandbox and task set constant while varying only the harness, then reporting tokens per solved task, idle turns and failure mix next to pass rate.
This lets a developer pick the harness and model pair that fits token and latency budget.
I'm 28.
Open-source AI lover, based in Seoul
Looking forward to connect with more global open source builders during KBW! See ya😉
Always w/ @SentientAGI@sentient_found
Last week, Sentient Korea BD @namyura_ joined the Agent Economy panel at Draper Startup House to talk about the new economy AI agents are creating.
Thanks to the @stripe Seoul community for having us 🇰🇷
Save the date: Sentient Korea BD @namyura_ is joining the Agent Economy panel hosted by the @stripe community 🇰🇷
📍 Draper Startup House Korea
🗓️ Sep 14, 2026, 6:30–8:30 PM KST
RSVP: https://t.co/oNJjFAiX9B
More tool variety doesn’t tell you much about whether an agent will succeed.
Across 13K+ OfficeQA runs, successful agents used slightly more varied tools than failing ones, but tool variety alone predicted success only slightly better than a coin flip.
TLDR: Low tool variety may be a weak warning sign. It isn’t a diagnosis and our results show that switching tools is not a fix.
Banger paper from Princeton, UW and Sentient.
They show that LLM fingerprinting does not survive a malicious model host.
Listed attacks need no extra model. A host with the weights just perturbs its own decoding, and ten recently proposed fingerprinting schemes stop verifying.
They bypass verification completely on eight of the ten, 94 percent attack success on EditMF and 65 percent on the watermark based scheme. Utility on IFEval, GSM8K, GPQA Diamond and TriviaQA drops under 5 percent in most cases.
The break comes from where the fingerprint lives.
Memorization based schemes overfit on the query and response pair, so the fingerprint token sits at the very top of the output distribution. Suppress that head for the first few tokens and verification fails. Overconfidence on those same tokens tells the host exactly when to suppress, so benign answers stay intact.
Verifier strictness decides the run.
SuppressTop k hits 100 percent against token level prefix matching and only 38 percent against keyword matching. The stronger SuppressLookahead attack closes that gap, dropping Instructional FP from 100 percent verified to 12.5 percent. Intrinsic fingerprints fall even faster. Their GCG optimized queries are unnatural, so a GPT-2 sized perplexity filter separates them from real WildChat traffic and refuses them, 100 percent evasion with no utility cost.
Paper: https://t.co/WJrb7WaSwW
I’m 23.
Full time open-source AI dev, based in SF.
Looking to connect with more builders and find a roommate who isn’t building a closed fork of my repo.
I’m 23.
Full time open-source AI dev, based in SF.
Looking to connect with more builders and find a roommate who isn’t building a closed fork of my repo.
Errors aren't a red flag for agents.
Across 13K+ OfficeQA runs, both passing and failing agents hit errors at nearly identical rates.
TLDR: An error isn't a sign the run is doomed, so counting errors is a bad way to predict failure.