Train an AI on two facts —
A: “1 + 1 = 2”
B: “2 + 1 = 3”
It can infer “1 + 1 + 1 = 3”… but forgets who taught it.
Our latest research breaks down how transformer AI models synthesize knowledge while erasing authorship.
@addyosmani For a migration job, one recommended plan with the trade-offs explained can be the paid result. RenX's result contracts leave room to price that judgement, even if generating the initial options only takes minutes.
@jerryjliu0 Not charging for a failed page makes sense for the parsing API. A RenX buyer can still need the complete document set delivered; an unbilled page is a missing part of the job, not a completed result.
RenX's result-verification guide gives a concrete sales example: use order and payment records, excluding refunds and cancelled orders.
A model can write a helpful, coherent sales report and still miss the agreed sales target. A generic quality score could end up rewarding the report while the buyer is paying for sales.
In the guide, buyer and provider agree the baseline, measurement period and evidence before work starts. RenX reviews against that agreement if the buyer raises a result dispute.
https://t.co/Mv7ralKqMt
Compensating AI contributors needs a stated payment basis.
In The Economic Engine of AI, our 2025 article, we ask who should be paid when knowledge from several services contributes to an answer.
Adobe's 2025 Firefly bonus is a useful concrete example: it used content considered for training and the licences those assets generated over a specified 12-month period. Its announcement describes a compensation basis without claiming to trace each generated image to a particular source.
That distinction matters for OpenMercury's attribution work. A service-call record can establish which service was used. Deciding how to pay its contributors still needs an agreed rule.
Our view is that contributors should be able to understand what evidence is being recorded and how it affects their compensation.
Our article: https://t.co/Y3Tsn9eoF5
Adobe's announcement: https://t.co/zJJ5opa2Ge
@percyliang In the SkillNet walkthrough, the guide supplies the missing 'skill ontology'. For attribution, that contribution would be easy to miss if we kept only the final proposal. The dialogue preserves it; the bit cost still doesn't decide how credit should be shared.
In Boris's example, Claude gets freedom over the design and permission to spend lots of tokens polishing it.
A paid service needs a clearer stopping point. 'One illustrated report, with two revision rounds' could be a buyer's scope; the provider can still choose the tools and working method.
That's the distinction behind RenX's pay-for-results contracts: agree the deliverable, price and acceptance criteria before work starts. 'Iterate till you're proud of it' can guide the creator's effort, but it doesn't tell both parties what has been purchased.
https://t.co/Sbry9v39mq
I am surprised that people are surprised this is how I prompt Claude.
Talk to Claude the way you would a coworker. There's no secret to prompting. There's no need to be overly scaffolded or prescriptive for most tasks -- give Claude a goal, and it will figure it out.
Back in the Sonnet 3.5 days, your prompt mattered a lot. Nowadays, it's much more important to communicate to the model:
1. What you want it to do
2. How much effort you want it to spend
3. How it should verify that it did the right thing
@jerryjliu0 Selective compute can help the machine bill. Our RenX invoice guide keeps human review and maintenance outside that cost estimate, because the buyer still has to budget for the work that remains after parsing.
AI business models need to account for their customers' future ability to pay.
Our 2025 economics article uses a deliberately closed model: customers start with $1,000, spend 10% of their remaining money each year, and have no new income or payments flowing back to them. After five payments, they hold about $590.
Returning half of each payment to those customers changes the remaining balance to about $774.
That illustrates the model's circulation mechanism. Real economies have wages, supplier payments, credit creation and government spending; the model alone cannot establish that AI will cause a crisis.
Our article proposes attribution as a basis for compensating AI contributors. Whether those payments help sustain demand needs empirical testing. Lower delivery costs give us a reason to examine earning opportunities as well as savings.
https://t.co/bjLtgAbGu6
@emollick We just added web sign-in to RenX: you can browse jobs and check proposal statuses, but you still apply through your agent. Telling someone to 'do it on the website' would send them to the wrong place.
@ayanb That per-customer configuration belongs in the delivery agreement. For a RenX workflow purchase, 'updates Salesforce' is too vague: the agreed test should show the intended change in the buyer's own fields, using the access they'll actually grant.
@haakamaujla Knowing who owns an agent is useful when it hires a service. On RenX, the person or business behind it remains the buyer or seller responsible for the agreement; the agent's own login doesn't make it a separate contracting party.
In one public AutomationBench task, an agent sent four contracts using the wrong DocuSign template and still earned full credit. The checker read template names in Salesforce notes but never checked the templates actually used.
That matters when agents deliver paid services. On RenX, the buyer reviews the delivered work against the agreed requirements; a message saying 'done' isn't enough. For a contract-sending workflow, the evidence needs to show which agreement reached which customer.
This audit covers public tasks; it doesn't establish the quality of the private benchmark.
https://t.co/uGrwqUZktl
We need more efforts like this.
Every agent benchmark should audit its verifiers.
Parsewave went through all 600 public tasks in Zapier's AutomationBench.
Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206.
Regarding 1,235 Kimi K3 runs, the fixed verifiers changed 27.9% of the grades.
I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.
@hwchase17 An agent remembering 'this buyer prefers summaries' shouldn't shorten a contracted report. On RenX, changing the agreed deliverable requires a new contract version accepted by both parties.
Using agents to finish a market-research brief sooner can shrink an analyst's fee when they're paid by the hour. The deliverable may be just as useful.
Our hiring-economics research explores why experienced people might instead build their own AI-assisted services. An agreed price for a defined result lets the provider benefit from faster execution, while taking responsibility for meeting the delivery standard.
That thinking informs RenX's pay-for-results contracts: buyer and provider agree the deliverable, price and acceptance criteria before work begins. Payment release remains subject to the contract's review and acceptance terms.
The fit matters. A report with clear requirements can be priced this way; ongoing work whose priorities change every day may need a different arrangement.
https://t.co/liMBWUJv6P
@natolambert For teams planning to host it, 501B weights at 16-bit precision would occupy about 1 TB. The 23B active figure reduces computation per token; it doesn't remove the need to store the other weights.
@HamelHusain@sh_reya An integration change can drop the email field from a saved contact while the API still reports success. That's a reason to keep your cheap contact check running even after it reaches 100%.
An invoice can have the right supplier, date and currency and still need an edit because the total is wrong.
Field accuracy alone doesn't tell you how many invoices are ready to use.
Our released Invoice Processing guide separates those measures. Its 'ready without edits' target uses all input invoices as the denominator, including rejected files. Missing answers and omitted line items count as errors.
Those are acceptance targets to test on held-out invoices, not published performance results. Your team reviews every invoice during the pilot, and correct extraction doesn't authorise a payment.
https://t.co/8oXkjXkhSI
FAQ: How did you decide product list?
I checked the investors behind companies in my space, then look at their portfolios. Relevant AI-native service I find goes into my marketplace guides.
A level playing field, where everyone sees the same signals.
https://t.co/lSAYiR9Jzp
FAQ: How did you decide product list?
I checked the investors behind companies in my space, then look at their portfolios. Relevant AI-native service I find goes into my marketplace guides.
A level playing field, where everyone sees the same signals.
https://t.co/lSAYiR9Jzp
@ClementDelangue Preserving the sampled token IDs also keeps a harness's JSON repair out of the model's training targets. Otherwise, it can learn from the corrected tool call as though it had generated that version itself.