A company archive keeps changing after the map is built.
CorpusMap’s incremental experiment improved answers as new evidence entered the map, even though the raw files were already searchable. A handover test should include a question answered by a newly added document.
CorpusMap's file suggestions explain an important part of its token savings.
In the GPT-5.5 experiment on 80 multi-document EnterpriseRAG questions, giving the agent the map without suggested files used 496,400 input tokens per question. Raw search used 206,500.
Adding relevant paths to documents linked by the map cut that to 88,100 and improved the reported answer-quality score. The agent still chose which source files to read.
For a company archive, organising the evidence and finding a useful starting point are both part of the job. These figures measure query-time use; building the map has an additional cost.
Paper, Table 8:
https://t.co/BrbZHw8zqz
A test case needs a defined size. RenX’s evaluation guide scopes one saved answer and up to five scoring rules per case.
Expanding that to every message and tool call changes the job. Agree the scoring unit before comparing per-case quotes.
https://t.co/LtC1edBtsQ
FAQ: How did you decide product list?
I checked the investors behind companies in my space, then look at their portfolios. Relevant AI-native service I find goes into my marketplace guides.
A level playing field, where everyone sees the same signals.
https://t.co/lSAYiR9Jzp
@OfirEhrlich A shared customer name across two records can make a convincing answer point to the wrong company. The ground-truth check needs the ID of the record actually changed.
@xambao A successful request makes a useful control for the failing trace. Keep the payload and proxy revision the same; otherwise a difference in execution may come from comparing two deployments.
An agent can write a refactor and its tests around the same wrong assumption.
Compare old and new behaviour on the same awkward inputs. A green suite written alongside the new code isn’t enough evidence that it preserved the old contract.
Fully autonomous agentic coding still isn’t ready for serious production work.
Two reasons:
1. Small mistakes compound fast
An agent can make one wrong assumption early, then confidently build 20 steps on top of it.
2. Production software is full of hidden context
Legacy code, edge cases, security rules, business logic, deployment constraints - things that aren’t always obvious from the task.
Where agents already shine:
��️ well-scoped features
🧪 tests and refactors
🔍 debugging with clear boundaries
📄 repetitive implementation work
The winning setup today isn’t:
Human OR agent
It’s:
Human judgment + agent execution
Full autonomy will come.
But reliability has to come first.
@open_mercury The exact combination is worth sharing alongside the skill. Tool, memory and context versions are part of the result; copying one file into a different harness isn’t a reproduction of that 84.8% score.
A convincing pool render can still have the wrong width.
One test for this RenX brief: run the same scene with two requested widths. If both outputs keep the same geometry, a high visual-quality score has missed the requirement.
https://t.co/CiTqbv5MV9
@EOEboh If it fails because the test database never connected, the agent hasn’t reproduced the bug yet. The assertion for the reported behaviour needs to fail; a broken setup can produce the same red badge.
@2logics A checkout test should fail if “Place order” is replaced by “Save basket”, even if a repaired locator can click it. Checking that an order was created catches the difference; a green click doesn’t.
@PyTorch@AMD Will the comparisons show the first request as well as warmed-up throughput? Short-lived agent jobs can finish before a compilation cost has paid for itself.
Give an investigation agent two devices with similar names but different histories.
The expected verdict can be unchanged while only one set of events supports it. Grade the retrieved records too, or a wrong-device query can pass on a lucky final answer.
A 95% match rate between an AI agent's verdict and an analyst's can still mask pattern-matching behavior. Carter Church & Gabriel Bernadett-Shapiro explain why SentinelOne builds security agents backwards from evaluation, testing competency by competency to confirm real investigative reasoning: https://t.co/3Nc7misrZK
@open_mercury For a fixed training budget, the useful counter is original examples actually sampled, broken down by category. An archive can stay intact while one rare category contributes zero originals to a particular run.
A shared checkpoint is a useful starting point for comparing reward signals.
Keep the final test set out of tuning. If failed test cases become training feedback, report that separately: the next score is partly measuring adaptation to those cases.
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next.
We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them.
Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up.
I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay.
I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
A Python answer is a weak reference if it just repeats the method being tested.
For generated maths problems, keep an independent check: a hand-worked small case or a different calculation. Two matching outputs can share the same mistake.
@arkyyang The trace checker needs regression cases too: remove a required tool call from an otherwise good run while leaving the final numbers intact. If that still passes, the extra coverage is only apparent.
@RenXHQ Check the exported PDF too: a label can be legible in the browser and clipped at the chart boundary in the file. The final artefact is what the client reads.
For a research agent, passing the calculation check is one score. Improving on the field’s baseline is another.
Those need separate labels. Otherwise an eval can reward a correct result that adds nothing to the question being studied.
In Schwartz’s account, the first ecology result was technically impressive but familiar to ecologists.
O’Dwyer redirected the work towards what remained after subtracting the neutral prediction. Choosing the question was part of the expert’s contribution.