ARC-AGI-3 scores agents on how close they are to human action efficiency.
All ARC-AGI-3 environments were solved by at least 2 human testers out of 10 (most of the time it was 5+). We use the action count of the 2nd best tester (to avoid outlier performance) as our human baseline.
Your score on an environment represents "how close you were to matching or exceeding the action count of the 2nd best human tester (out of 10 random people who attempted the task)"
@elonmusk@DOGE@cookdotmeme Launch a meme for this and generate new exactly same image but draw doge reflections on the glasses only. Touch nothing else