Sorry for the delay, yesterday was a relentless bullshit-task day lol
Two takes on the AUC idea:
① AUC as ability: like expenditure-adjusted score, but without specifying a utility function. Simpler, though it implicitly treats spend cost as zero
② Normalized AUC as scalability: AUC / (X · s(X)), smaller = better. Captures what "returns to expenditure" describes qualitatively, but collapses it into a single scalar for ranking. Low value = model keeps improving near X, not plateauing early
Does this make sense? Happy to elaborate!
New post: nine different ways of summarizing agent ability.
#1: Score at Fixed Expenditure. Spend a fixed amount on each model (x=tokens/money) , and report the agent's score. This is the classic.
Very happy to share the first paper from @ElasticityInst: The Economics of Recursive Self-Improvement. Two parts:
(1) a graphical representation of feedback loops, to formalize a variety of RSI-related arguments, where each arrow represents responsiveness (elasticity);
(2) a survey of existing evidence with a loose calibration & a “wish list” of evidence that would help us calibrate better.