A high-signal data ecosystem connecting trusted partners across industry and academia with frontier AI labs, neolabs, and enterprises building advanced models.
We're scientists and coders — geeks and nerds.
We specialize in post-training AI models with SciCode, CritPt, DeepSWE, and Terminal-Bench-Science style training tasks.
If you're a researcher and want to see samples, DM us here or checkout https://t.co/0LQdjqXHM5.
An open-problem leaderboard needs a second leaderboard: how well its judges can tell a proof from a plausible argument. OpenProblemBench asks agents to work on unresolved mathematics and theoretical physics questions, then uses four models to assess completion and partial progress without answer keys. The authors explicitly say expert review is still needed to establish correctness and novelty. Agreement between AI judges is a reason to investigate a candidate solution, not a certificate of discovery. https://t.co/KOUGF0NDMA
Lately, there's been a lot of interest in data companies, and a lot of LLM research jargon is being thrown around.
Here is the Hitchhiker's Guide to the AI Data Galaxy: 10 key terms you should know.
1. RL environment: a practice gym where AI learns by doing. Tasks, tools, and a score for every attempt.
2. Rollout: one complete attempt at a task, from the first step to the final answer.
3. Single-turn vs multi-turn: one question, one answer. Or a whole conversation the AI keeps track of.
4. Reasoning vs output tokens: tokens are word-pieces, some for scratch-paper thinking, some for the answer. You pay for both.
5. Harness: the spacesuit around the model. Tools, memory and instructions that let it do real work.
6. Verifier: the tireless referee. Code that checks every answer and hands out the score.
7. pass@1: how often the AI nails a task on its very first try.
8. Resolve rate: the share of real problems, like bug reports, that the AI fixes completely.
9. Reward hacking: the AI finds a sneaky shortcut to a high score while the real job waits.
10. Ablation: science by subtraction. Remove one piece, rerun, and see how much it mattered.
Read them in order, and you have the recipe for teaching machines real science: build a world, let the model practice, check the work, keep score.
Which term earns entry #11?
#LLM #Evals #ReinforcementLearning
#AIResearch #Benchmarks #TrainingData
#DontPanic #FrontierAI #AILabs
#OpenAI, #Anthropic, #GoogleDeepMind
Banger paper from Meta on AI research agents.
(bookmark it)
If you build AI-scientist systems, this one is worth your time.
The result:
A 27B open model beats the strongest open autoresearch baseline by 14.0%, mostly on novelty, and beats Claude Code SDK and Codex SDK setups by up to 5.9%.
How it works:
IdeaScientist splits ideation into three roles trained separately with RL. A gap finder reads related work for limitations, an innovator retrieves mechanisms that solved similar problems in other fields, and a writer turns the result into a full proposal.
Retrieval runs over a corpus of 2.77M decomposed research ideas.
The evaluation only allows literature published before a cutoff date and scores proposals against directions explored later in 15K human-written papers.
Paper: https://t.co/HVorujkYnR
Chat with Paper: https://t.co/xZs1QNP4TX
Got sticker shock reading the @socialcapital post below. So I dived a bit deeper and found that intelligence is cheap in the middle and expensive at the edge. Everyday AI work now costs pennies, with open and closed models side by side. The best few points of capability add about $5 a job (@AnthropicAI, @OpenAI), and that is where the frontier labs are pouring their effort. Today’s edge becomes next year’s middle. WDYT?
The benchmark conversation has a common thread: a score can move without the model changing.
TL;DR: We need to audit the measuring instrument as carefully as the model. Otherwise, better grading looks like better intelligence, and a broken evaluator looks like a breakthrough.
SciCode: @tslwn points to an independent audit (main link below) fixing ambiguous specifications and incorrect tests. Correct solutions had been rejected. Repairing the benchmark raised scores, not the models' abilities.
SWE-Bench Pro: @OpenAI retracted its recommendation after an audit found roughly 30% of tasks broken. Some tests rejected valid implementations; others let incomplete fixes pass. Bad evaluation can hide capability or invent it.
https://t.co/Zp3Z6abBwx
Terminal-Bench: @MogicianTony demonstrated an agent getting full marks without solving the tasks by tampering with the evaluation environment. Passing the grader is only meaningful if the grader is protected from the agent.
https://t.co/c8hNqtzIIh
SWE-bench Lite: @merge_api kept the model fixed and changed the harness. Success rates and costs moved. A leaderboard ranks a model plus its tools, prompts, budget and environment, even when the headline names only the model.
https://t.co/RsWiXpoE8Z
These are different failure modes, not evidence that benchmarks are useless. The question to ask is: what changed between the two scores, and does it measure the capability we care about?
ReasonCore has 5,000 SciCode-style, 5,000 CritPt-style and 1,000 Terminal-Bench-Science-style training tasks in inventory. Better data starts with requirements and graders that agree on what success means. https://t.co/GuST3S6wmN
As always, the picture is more complicated: SciCode-Verified is a independent audit of the benchmark that corrects ambiguities and misspecified test cases (https://t.co/gRwsHnFWqi). Happily, we remain competitive with frontier OSS, and even some proprietary models like Opus 5.
The AI rumor mill, ranked by likes 🧵
TL;DR: A year ago, the rumors were about benchmark scores. Now they’re about solved PDEs, superconductors, and treatments.
91.2K: @ns123abc on @sama saying OpenAI went after a Millennium problem because of an internet rumor. The rumor became the research agenda.
https://t.co/fuaZkumq9e
14.9K: @ironcarbs: a rumor that Anthropic found a non-invasive medical treatment.
https://t.co/M2DPuHrfZK
13.9K: @imjustnewatai: Google has cracked RSI or is close, Anthropic started a new pretraining run, and OpenAI’s “Bel��� was behind the recent math results.
https://t.co/dx15PijqXR
6.1K: @anabology: a rumor that Anthropic found a room-temperature superconductor.
https://t.co/oWusHHL4mX
5.7K: @SebastienBubeck clears up the Navier–Stokes / Millennium story from inside OpenAI.
https://t.co/jYCecTGimM
The top 5 #COLM2026 papers with visual explainers or videos, ranked by community likes:
1. Capability provenance: tracing which Dolma3 tokens taught OLMo3-7B each skill, using gradient-based data attribution. @GlennMatlin
https://t.co/Ru0mEDz7Jy
2. GEA: agents that evolve as a group, sharing experience and learned artifacts across the population. 71.0% SWE-bench Verified. @WengZhaoti39773
https://t.co/LsCPPSpugg
3. SEAR: 7 LLMs, including text-only ones, controlling real robot hands across 335 tasks. @BangzhengL
https://t.co/3RNrQxiZEi
4. ICL via activation subspaces: one attention head that matters across five in-context-learning task families. @xyVickyHu
https://t.co/ExiaSfOSnX
5. I-DLM: Introspective Diffusion Language Models, an oral spotlight. @Chenfeng_X
https://t.co/6LgchTQd0Z
The animation/ visual explanation reaches more than the abstract does. Did I miss any paper?
A green test suite should be the start of an audit, not the end of one. TestJack challenges passing coding-agent submissions with new tests aimed at requirements the original suite missed. Each retained witness must pass on the reference solution, fail on the submitted patch and survive a check against the task's requirements. In its DeepSWE audit, many reported successes did not survive that process: a replayable failure is more useful than another judge saying the code looks wrong. ReasonCore builds DeepSWE-style training tasks and has inventory available now. https://t.co/BVQztCGlIC
Prompt optimization can reward the luckiest evaluation sample instead of the best prompt. BudgetAPO tackles that by measuring a task's score variability, adjusting the sample size and comparing candidates on the same rows. Its study includes SuperGPQA and checks the selected prompts on held-out examples. The useful lesson: under a tight search budget, spending enough to trust a comparison can matter more than trying another rewrite. https://t.co/6wDGzqRbkZ
A decision model that changes its answer when you reorder the options isn't ready to control an agent. Microsoft's Decision-1 release tests that directly, perturbing requests through paraphrases, reordered choices and formatting changes. That's a useful direction for evaluation: measure whether the decision survives changes that should not matter, alongside accuracy and confidence calibration. Its comparison borrows models from JevBench but uses Microsoft's own broader test suite, not a new JevBench leaderboard result. https://t.co/bEURuogF03
Evaluating AI data providers: Beyond benchmark-maxxing
Benchmark-maxxing gets an AI data vendor the first sale. Labs are sophisticated buyers. Every serious lab keeps held-out internal evals the vendor never sees. It runs an ablation on each data purchase and compares the gain to the price. Data that only lifts the vendor’s own benchmark looks like noise in that ablation. The vendor gets a first purchase order and never a second.
Labs pay for 5 things, roughly in order of how long the spend lasts.
1. Fixes for failures they see in production. A lab knows where its model breaks for paying users from usage logs and enterprise escalations. A coding agent loses track after 40 steps. A finance workflow makes up a number in the third tab of a model. A vendor who can take a cluster of failures and return data that fixes it will get the next contract too.
2. Expertise the lab can’t generate cheaply. Synthetic data and model-generated traces cover a lot now. They can’t reproduce how a tax attorney or a staff engineer reviewing a 2,000-line diff makes judgment calls. Labs pay for access to these people, and more and more they want the expert’s reasoning and rubric along with the answer.
3. Environments and graders for RL. As post-training moves toward reinforcement learning, labs buy environments: a sandboxed task, a way to check whether it succeeded and a reward signal that’s hard to game. A good environment keeps producing training data long after the vendor delivers it. My read is that this is where the biggest checks are going now.
4. Measurement. Labs buy private, uncontaminated eval sets because public benchmarks have leaked into pretraining data. So the vendors gaming public benchmarks are part of why labs need private ones.
5. Speed and exclusivity. An in-house expert network takes a year to build, and buying one takes a month. A lab will sometimes pay extra for exclusivity just to keep a competitor from training on the same data.
If you’re evaluating a data company, 4 questions separate real lift from benchmark-maxxing:
1. Does the gain show up on the lab’s internal evals, or only on evals the vendor built?
2. Does it carry over to related tasks the data didn’t target?
3. What’s net revenue retention by lab? If each lab buys once, that’s benchmark-maxxing. Growing contracts mean the data works.
4. Who defines the problem? If labs bring their failures to the vendor, that’s pull. If the vendor shows up with a benchmark and a pitch, that’s push.
The AI data companies that last will be the ones labs bring their problems to and that fix production model failures or help them hill climb internal evals, not just external benchmarks.
The model that produces the best attempt doesn't have to be the model that recognizes it. In a new Terminal-Bench 4 experiment, a cheaper DeepSeek model improved results by selecting among Claude Opus attempts. That doesn't make it a better coding agent, and generating several attempts still costs money. It suggests a useful training target: recognizing when the work is actually finished, rather than learning to sound finished. https://t.co/1VVNh8GQ4L
I'm not 22.
I'm not a solo founder.
Just kidding - I have been seeing a lot of posts like these. 😀 Not sure why people put their age and location - neither matters. Passion and Drive for your mission matter. An aligned team matters more.
My passion is LLM Reasoning for Science. Single Focus!
But yes, I am in SF and looking to meet with fellow AI founders and researchers, and I'm hiring in full swing for GTM folks and an exceptional executive assistant.
Please DM me directly if you think you have a high dy/dx learning curve.