929 million edges in 10.8 minutes: structuring 40 million documents into an agentic knowledge graph ๐บ๏ธ
Building a knowledge graph usually means a model reads every document. At 40 million documents, that is where the usual approach breaks: GraphRAG indexing costs about $33,000 for a single corpus.
Some corpora already ship their edges. For those, graph construction becomes a parse rather than an extraction.
The numbers reported:
929,824,202 edges, built in 10.8 minutes
Zero model calls
46 Python modules, with no graph database and no extraction model
83.2% on 600 held-out questions
The build starts with six columnar edge tables parsed from 1,334 gzipped XML files. Citation retention is measured first, since a recent ten million slice keeps only 28.1% of its edges. The ontology layer follows, with 31,110 MeSH descriptors and 267,012 entry terms.
On top sits an agent written from scratch in Python. It walks the path through the graph, cites it, and refuses when there is none. One of the demonstrated questions is refused before the model is ever loaded.
A knowledge graph makes the route to an answer visible. The agent makes that route accurate.
By Fareed Khan
https://t.co/mMTqVpvICb
#GraphRAG #Ontology #AgenticAI #Python #DataEngineering #DataScience #EmergingTech
--
๐ฉ The Year of the Graph Autumn 2026 newsletter issue is out!
Meaning, Traversal, Reasoning, Action: The Graph Stack for Agentic AI. ๐
https://t.co/YfrPw2nFzS
All things #KnowledgeGraph, #GraphDB, Graph #Analytics / #DataScience / #AI and #SemTech.
Subscribe and follow to be in the know. Reach out if youโd like to be featured
Excited to share IdeaScientist, our research from my internship at Meta! ๐
Can AI agents learn to generate novel, grounded scientific ideas?
We introduce IdeaScientist, a framework that trains specialized agents to identify research gaps, discover useful connections across scientific domains, and turn them into concrete research proposals.
๐ Key findings:
โข +14% overall performance over open autoresearch baselines
โข 25% improvement in novelty, showing that explicit training can improve scientific ideation
โข 75โ92% human preference win rates against six open baselines in blind evaluations
โข Cross-domain retrieval substantially improves the transfer of ideas between research fields
We also introduce Svalbard Idea Vault, a collection of 2.77M decomposed research ideas for training and evaluating scientific ideation.
๐ Paper: https://t.co/OCFIyCKe4k
Introducing AG-UI 1.0 ๐ช
The open protocol that connects ANY agent to ANY app, now with a stable spec.
AG-UI gives you an entire ecosystem for free.
Adopted and supported by Google, Microsoft, Amazon, Oracle, Anthropic, LangChain, Mastra, and more.
New in 1.0:
โ Subagents: expose delegated agent work
โ Metadata: send data from agent to UI
โ Multimodality (images, audio, video, docs)
โ Interrupts
โ Token usage per run
Every SDK is now generated from one schema, so TypeScript, Python and .NET all behave the same.
AG-UI 1.0 is fully backwards compatible.
Full announcement: https://t.co/DeigmMD8KH
Sony Research introduces HakkenOSS, an open-source knowledge graph platform for scientific hypothesis generation featuring a complete pipeline from data ingestion to embedding-baโฆ
https://t.co/VW0McJYI6H
#MachineLearning#AI#LLM#DeepLearning#AgenticAI
We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation. https://t.co/hrqdNpm0g8
microsoft / data-formulator
๐ช Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data.
https://t.co/uiOtvGqaaS
Thrilled that #paper2agent is published in @nature today!
Scientific knowledge is traditionally stored in passive papers. Paper2Agent transforms papers into virtual authors that answer questions, apply its methods and collaborate w/ other paper agents to make new discoveries ๐งต
Real Deep Research (RDR) โ an AI-powered framework built to help actually keep up with modern science.
RDR bridges the gap between expert-written surveys and automated literature mining, featuring:
- A scalable pipeline: analyzes any research area
- Trend analysis: spots whatโs rising, whatโs fading
- Connecting dots across domains to reveal fresh opportunities
- Producing structured, high-quality summaries
Itโs the kind of research tool every researcher might need in their toolkit to see the bigger picture.
๐๐ผ๐ด๐ ๐๐ ๐ ๐ฒ๐๐ฟ๐ถ๐ฐ๐ ๐๐ ๐ง๐ฟ๐ฎ๐ฐ๐ฒ๐.
Logs, metrics, and traces can all point to the same problem, but they show you that problem from completely different perspectives.
๐๐ผ๐ด๐ = โ๐ช๐ต๐ฎ๐ ๐ต๐ฎ๐ฝ๐ฝ๐ฒ๐ป๐ฒ๐ฑ?โ
When something happens inside an application, logs capture a timestamped record of the event, from errors and warnings to requests and state changes. That detailed context helps you understand exactly what happened at a particular point in time.
๐ ๐ฒ๐๐ฟ๐ถ๐ฐ๐ = โ๐๐ผ๐ ๐ถ๐ ๐๐ต๐ฒ ๐๐๐๐๐ฒ๐บ ๐ฏ๐ฒ๐ต๐ฎ๐๐ถ๐ป๐ด?โ
Rather than recording individual events, metrics turn system behavior into numerical measurements over time, such as request rate, error rate, latency, CPU usage, and memory consumption. This makes patterns, trends, and abnormal behavior much easier to spot.
๐ง๐ฟ๐ฎ๐ฐ๐ฒ๐ = โ๐ช๐ต๐ฎ๐ ๐ฝ๐ฎ๐๐ต ๐ฑ๐ถ๐ฑ ๐๐ต๐ฒ ๐ฟ๐ฒ๐พ๐๐ฒ๐๐ ๐๐ฎ๐ธ๐ฒ?โ
Across a distributed system, a single operation can pass through many services. Traces connect spans from each step into an end-to-end journey, showing where time was spent, which services were involved, and where errors or latency appeared.
But during an incident, seeing the signals is only part of the problem.
The hard part is connecting them with deploys, commits, dependencies, and past incidents ๐๐ผ ๐ณ๐ถ๐ด๐๐ฟ๐ฒ ๐ผ๐๐ ๐๐ต๐ฎ๐ ๐ฎ๐ฐ๐๐๐ฎ๐น๐น๐ ๐ฏ๐ฟ๐ผ๐ธ๐ฒ ๐ฎ๐ป๐ฑ ๐๐ต๐.
Thatโs where agentic root cause analysis can help. Investigations by @incident_io starts working the moment an incident is declared, reasoning across that context to build a hypothesis backed by evidence and identify where to look next.
๐๐ป๐๐๐ฒ๐ฎ๐ฑ ๐ผ๐ณ ๐๐๐ฎ๐ฟ๐๐ถ๐ป๐ด ๐ณ๐ฟ๐ผ๐บ ๐๐ฐ๐ฟ๐ฎ๐๐ฐ๐ต and digging for answers, ๐๐ผ๐ ๐ฎ๐ฟ๐ฟ๐ถ๐๐ฒ ๐๐ถ๐๐ต ๐ฎ ๐๐ผ๐ฟ๐ธ๐ถ๐ป๐ด ๐๐ต๐ฒ๐ผ๐ฟ๐.
Check it out โ https://t.co/45CkvWsPva
What else would you add?
โโ
โป๏ธ Repost to help others learn and grow.
๐ Thanks to incident .io for sponsoring this post.
โ Follow me ( Nikki Siapno ) to improve at AI and system design.
To request an evaluation, submit a publicly available Hugging Face retrieval model through our evaluation request form: https://t.co/yGvVSxnrC3
The full technical report with methodology and findings is available here:
https://t.co/G0AKX2N8r5
PaperGym: Low-leakage training from research papers
Turns each arXiv paper into a training environment with rubric-based scoring. Cuts criterion leakage to 3.7% (vs 12โ34% in existing benchmarks), and boosts Qwen3-8B to 73.48 on ResearchQA.
The next step after Karpathy's autoresearch idea:
Tuning an agent is mostly manual work, done by editing prompts, tools, and control flow by hand and rerunning the evals to see what moved.
Researchers have been building several automated optimizers to do that outer loop instead.
The underlying process is the same:
- An LLM proposes a change
- An evaluator scores it
- And the proposer reads that result before proposing the next one.
They differ in what they edit and what feedback they get to read.
1) Berkeley built GEPA that optimizes the text of a system, like prompts, tool descriptions, or the agent's own code.
Instead of collapsing a run into one scalar reward the way RL does, it reads the full execution trace (errors, reasoning, tool output), diagnoses why the run failed, and proposes a targeted fix.
It also keeps every candidate that's best at some part of the task, not just the one with the highest average score. So a strong specialist survives even when a more balanced candidate beats it overall.
That trace-level feedback lets it converge in hundreds of rollouts instead of the thousands that GRPO needs.
2) AutoResearch, inspired by Karpathy, runs a narrower version of the loop, where a coding agent iterates on a program(.)md file, scores the outputs against fixed evals, and keeps whatever improves.
3) Meta-Harness points the loop at the harness itself, the scaffolding code that decides what to retrieve, how to format it, and what state to keep between calls.
The natural question is which of the three to use. And the answer is none of them alone.
On the Frontier-CS benchmark, holding the model, thinking effort, and budget fixed, no optimizer wins everywhere. Across 10 tasks, GEPA led on 3, AutoResearch on 3, and Meta-Harness on 4.
Each optimizer hill-climbs fast and then stalls, making most of its progress in a few iterations before flattening out.
But the researchers found that handing the stalled candidate to a different optimizer breaks the plateau, since each one attacks the problem differently.
To use this in practice, omni (open-source) already automates all of that.
It runs every optimizer on a fraction of the budget, takes the best candidate, and hands it to a fresh optimizer to keep going.
It scores 7.8 percentage points above the best standalone optimizer at the same budget, and finishes faster.
The whole meta-optimizer is around ten lines in the optimize_anything API, the same interface these optimizers already run through.
You can also point an agent at the gepa-ai/gepa repo and have it use the gepa-optimize-anything skill.
Here's the repo: https://t.co/cMuknES4Uz
For a first-principles guide to GEPA and the engine under most of this, we wrote a full breakdown of why reflective evolution beats GRPO by around 10 points with 35x fewer rollouts and no GPU training, purely from natural-language reflection on trajectories rather than backprop on weights.
Read it below.