Reinforcement learning rewards convergence, and convergence is the opposite of discovery.
We built the world’s first judgment model from the decisions papers erase: what to pursue, challenge, abandon, and revisit. Columbus-1 autonomously found a zero-click Bluetooth RCE and designed a 10-foot, self-landing rocket.
We've selected @Oracle Cloud Infrastructure to power the next phase of our AI co-scientist development.
Scaling with enterprise-grade AI infrastructure. Building AI systems that researchers can trust.
https://t.co/o7XwlhCBwp