Apodex Discovery: a new paradigm for discoverative AI
A benchmark and environment for AI that solves real-world problems without known answers. 423 high-value problems identified, 20 released with verifiable investigation.
So what does TRACES actually measure? Not just the answer, but the path to the answer.
Six capabilities that spell its name:
*T — Tools: selecting, calling and correctly interpreting external tools
*R — Repair: locating and correcting its own errors once feedback arrives
*A — Alternatives: laying out competing hypotheses and keeping or discarding them as evidence accumulates
*C — Coherence: holding state, constraints and logic intact across a long chain of work
*E — Evidence: grounding every conclusion in observation, data, experiment or citation
*S — Scope: stating the conditions under which a conclusion holds, and where it does not apply
TRACES evaluates today whether the discovery process is rigorous and evidence-grounded, even when the final outcome is not yet known 🔍
📄 Paper: https://t.co/qkCiNT0nTA
Today, in full: the technical report — the definitions, the rubric, the registry of problems — spanning biomedicine, clinical translation and frontier-model engineering. Led by Apodex’s Lead Scientist @profshengwang, who brings together the technical framework behind this effort.
Behind them: 10 STEM PhDs, two months, 561 industries surveyed to find the 423 problems worth doing this for — real, high-stakes questions that take human experts years to work through.
Bring a problem. Bring a solver. Both doors are open.
*Submit a Problem:https://t.co/ycJyKmlMYl
*Submit a Solver:https://t.co/sanydxSc5C
This is a good example of why final-answer accuracy can hide bad agent behaviour.
Most AI benchmarks measure whether a model can reach a known answer.
Apodex introduced TRACES 🧭, the world's first benchmark for measuring discoverative AI.
A shift from benchmarking models on solved problems to evaluating systems that can investigate consequential problems under evidence, tools, and verification.
TRACES says AI discovery should be evaluated as an entire investigation, not as a single final answer.
This can separate a high-scoring outcome from the quality of the process that produced it, including whether errors were repaired and claims were grounded.
The gap between "finding a known answer" and "earning one nobody has" is exactly where current evals stop working. What makes TRACES different: it scores the investigation itself — evidence, hypotheses, verification — not just the final answer.
Just came across TRACES from Apodex — this is genuinely exciting.
Most AI benchmarks only test retrieval of known answers. Real scientific breakthroughs, however, require discovery: working through evidence, testing hypotheses, and reaching verifiable conclusions when no answer key exists.
TRACES is the world’s first benchmark designed specifically for discoverative AI, evaluating six core capabilities (Tools, Repair, Alternatives, Coherence, Evidence, Scope).
This is the shift we’ve been waiting for. Highly recommend checking it out!
Apodex just released new benchmark. AI is moving from generative to discoverative. Future AI do not just use tool but can hypothesis, execute, verify and self-evolvement.
Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet?
Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test hypotheses, and reach verifiable conclusions on problems without answer keys. Proposed by our founder @tianqiao_chen, who defined its six capabilities.
Three things published today: a definition of "discoverative intelligence", a rubric to tell sound investigation from lucky guesses, and a open call for both solvers and problems
*Website: https://t.co/iJ5rDV6qz9
MakePlay is officially live! You don't write code. You direct games. 🎬
Type one sentence — "a shiba pixel runner with double jump" — and an AI crew builds the whole thing: code, art, music, sound effects, level design, all at once.
No engine, no assets, no setup.
From FPS to open world — racing, puzzle, space shooter, platformer.
Then keep directing in plain language. Make the boss slower. Make the music heavier. Soften the jump sound.
Completely free. Runs in any browser.
🕹️ https://t.co/Qb07BuHNnh
@TwoSetAI × @SimonShaoleiDu
“For writing an agent framework, you only need APIs. You don’t need that many people to do the same kind of work.”
Building an agent scaffold is easy. The harder problem is training the model inside that scaffold to plan, search, and reason the way a scientist does. 🧠
Here is how we trained the model itself—not just the scaffold—to handle planning and multi-hop search:
MakePlay is officially live! You don't write code. You direct games. 🎬
Type one sentence — "a shiba pixel runner with double jump" — and an AI crew builds the whole thing: code, art, music, sound effects, level design, all at once.
No engine, no assets, no setup.
From FPS to open world — racing, puzzle, space shooter, platformer.
Then keep directing in plain language. Make the boss slower. Make the music heavier. Soften the jump sound.
Completely free. Runs in any browser.
🕹️ https://t.co/Qb07BuHNnh
I did an interview with @TwoSetAI how we got to SOTA on deep research and future prediction:
→ we trained the model, not just the scaffold
→ up to 150 sub-agents run on top of it
→ verification is its own layer, not a final prompt
https://t.co/C61oiyqhsZ
Our Chief Scientist @SimonShaoleiDu recently explained where single-agent reasoning hits a wall:
Even a one-million-token context may not be enough for one agent to carry a hard research problem.
A new NVIDIA preprint makes that measurable.
Here’s what breaks inside a single context — and how we reorganize the work:
https://t.co/b5M87jQg2Z
“A large model is less random than a human. It follows instructions—people don’t.”
So—is it easier to organize 100 AI agents than 100 people?
And what does “AI improving AI” actually look like inside a real system?
Our Chief Scientist @SimonShaoleiDu joined @hi_angelinayang at TwoSetAI to unpack both—and explain what it takes to build a strong agent framework.
🎙️ Watch the preview and stay tuned for the full conversation. https://t.co/JtaWPFLK0K
Solvers of the Week!
Since that championship prediction landed, the prompts have started getting harder.
🏦 Will the Federal Reserve cut rates before Q3? Weigh the dot plot, FOMC minutes, inflation, employment, market pricing, and the signals that would change the conclusion.
₿ Can Bitcoin break its all-time high again within 90 days? Map the scenarios around ETF flows, liquidity, macro conditions, and the evidence that would invalidate each path. (A research prompt, not financial advice.)
🌍 Which opportunities across Southeast Asia, the Middle East, Latin America, and Central Asia are supported by evidence — and which are conference-stage optimism?
📱 Will Apple both announce and begin selling a foldable iPhone by December 31? Reconcile conflicting timelines across product reporting, engineering risks, production yields, and the supply chain.
Different domains, same pattern: sources to check, conflicting evidence to resolve, assumptions to expose, and conclusions that may need to change as new evidence arrives.
Take the weekend. Hand the hard problems still sitting at the back of your mind to Apodex.
#ApodexSolvers