Future SOTA computer use models will also be SOTA software cloning models.
Why arent computer use agents as good as coding agents? Both have easily verifiable rewards. Is the bug fixed? Did the customer data get saved in salesforce portal?
The problem is the rollouts are expensive. You cant spawn 5k cuas to slam the salesforce portal at the same time. They will block you.
So you need thousands of realistic, disposable software environments: mock salesforce, mock workday, mock SAP, mock everything.
The only scalable way to produce cua environments is with coding agents trained to recreate real software from observation.
Any lab serious about post training computer use must also be serious about post training software cloning.
New FAIR paper dropped 🙀: Quantization makes reasoning models more indecisive
Lotfi et al. show that quantized models tend to doubt themselves, more often saying 'wait,' 'but,' or 'alternatively' and talking themselves out of the answer.
Penalizing overthinking at test time leads to a measurable (up to 12-23%) decrease in CoT token usage while maintaining accuracy.
This is why I always prompt inject affirmations into my cua models 😂
https://t.co/A44NF9nPbN