Meanwhile, $CPHY agent is in the top ten human leaderboard with super-human move efficiency putting Claude 5.5 to absolute shame 🦾
*Apparently, however, the Codex window was always open and even though the agent only ran the mazes a few hours a day while I did other work, the time reported was the duration the Codex browser use plugin was on.
Verifiable AGI engineering has always been shared with the world by $CPHY since inception.
Like anything worthwhile, it simply requires some effort.
Less & less effort, every day that passes, with each new update we make🦾🤖⛓️
$CPHY skill benchmarks currently public, fully replicated, and completed across Grok 4.6, GLM 5.2 & 5.3, Meta's Muse Spark 1.1 & 1.2, DeepSeek v4 variations, ChatGPT 5.5 or 5.6 series and Claude Sonnet 5 and Opus 4.6-5 (prior generation LLMs saturate the same benchmarks using the skill, showing the underlying LLM is not what drives performance, the skill is).
• ARC-AGI 3 public = 100% & the only agent that beats it in the ~5000 range of move efficiency for under $200 (not specifically designed to solve it at all, still the best move efficiency and exponential cost drop in the world including solutions trained on the benchmark, which Cypher is not).
• LongMemEval = 100% with 30/30 abstentions (no hallucinations).
• Pokemon Red = 100% 14-16 hour clear time with notably underleveled Pokemon across all models.
• LIBERO 40 = ~80%
• LIBERO 90 = ~85%
• RoboCurve Cubepick, Kitchen Bench, Robustness Grade = 100% saturated across all.
• Mazebench = Currently top 10 compared to human results of rooms cleared and gems collected with superhuman move efficiency (no AI competition, closet is Astra the only model to make any progress but not anywhere in the range of humans or Cypher Tempre).
*Many, many, many more benchmarks have been run or are in the process of being run and will be reported as they are cross-checked and vetted as easy to replicate by anyone at home.
*I am now replicating these benchmarks with models under 30 billion parameters, let this statement be a responsible, considerate, and most of all decent & loving economic warning to anyone building trillion dollar-trillion parameter AI with no timechain.
*I am training models born with a timechain from the beginning and the process eliminates catastrophic forgetting and everything else wrong with LLMs producing fundamentally more-capable systems for 99.2% less cost than the 95% less DeepSeek leveled training costs compared to OpenAI.
https://t.co/XqirGpIDzu
Timechains completely solved long horizon tasking to virtually infinite context and length of time, I have tasks over 60 days going with perfectly retained coherence from the $CPHY timechain skill:
https://t.co/oAcRH6DU76