Fireworks at $17.5B and Kimi K3 in the same news cycle is a good reminder that AI infra is becoming a digestive system for model churn. It ingests new models, specializes them for domain-specificity, serves them efficiently, and converts upstream intelligence into viable unit economics.
The best AI organizations may simply be the ones with the best model metabolism.
“Big computer” helps, but “large, flexible industrial load” may be the more useful framing. The real policy questions are whether the load pays for the grid capacity it requires, can curtail when the system is stressed, and induces new generation. A moratorium precludes the elasticity question and unfortunately replaces prices with prohibition.
This Venezuela earthquake story is the most concrete version of “AI for civic capacity” I’ve seen. Agents spinning up missing-person registries, hospital lists, donation coordination, WhatsApp interfaces, maps, in mere hours steered by diaspora developers.
The hard part is that these tools are filling a state-shaped hole. Faster software can route help but cannot confer legitimacy or durable accountability (yet)
https://t.co/ElFQFb8tP9
@karpathy Our recent work developed a bottom-up curriculum of practice problems for language models, utilizing medical knowledge graphs as a structured and clean proxy for medical textbooks 👇
https://t.co/KY9TQXbpoE
Training with tasks built from a medical knowledge graph turns a 32B model into a strong, reliable domain specialist.
Top down text pretraining misses deep structure.
This paper builds a bottom up curriculum from a medical knowledge graph, where each path composes simple relations into a reasoning chain.
Each path becomes a clinical multiple choice question with a step by step trace, then 2 graders verify both the answer and the trace.
This yields 24,000 high quality tasks.
They fine tune QwQ-32B with LoRA, producing QwQ-Med-3, and release ICD-Bench with 3,675 questions across 15 disease categories, mainly 2 to 5 hop chains.
The model wins across all categories and gains most on the hardest items.
At inference, they sample many parallel traces and vote, which beats iterative refinement once the curriculum gets deep.
The scaling curve shifts left, which means higher accuracy at lower token budgets.
Analyses show the model recalls more of the true hops and actually uses them to reason, not just to quote facts.
Training ran on 8 H100 GPUs for about 20 hours per run, which hints that compact specialists can be practical.
Net takeaway, a reliable knowledge graph plus a bottom up curriculum elicits domain specific superintelligence while cutting brute force inference.
----
Paper – arxiv. org/abs/2507.13966
Paper Title: "Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What They Need"
knowledge graphs can abstract a simulatable training env for some domains. from an RL POV, each edge of a KG encodes a local verifier, and multi-hop paths of a KG can be translated to synthesize natural language reasoning tasks with dense rewards. not enough bitterlessonpilled in the asymptote, but KGs can provide millions of structured domain-specific hill-climbing tasks to get started with
https://t.co/Sf8s9HdFeH
How do we use a KG? Our central insight is that paths in a KG can be translated into grounded natural language reasoning tasks, whose solution requires reasoning along the relational chain encoded in the paths. At the core of our approach, we map a KG path into a reasoning task, paired with thinking traces that are also grounded in the path. Training on such tasks can then enable an LM to explicitly acquire structured domain primitives and learn how to systematically compose them at inference time. (3/N)
How can we extend the recent success of LLMs at the IMO 🏅 to other domains 🩺🧬⚖️? If post-training on high-quality data is key, how do we curate data that imparts the right domain-specific primitives for reasoning?
Today, we're releasing a new paper on using a knowledge graph (KG) as a data foundry to generate dense reasoning curricula for post-training LLMs.
We use our approach to synthesize 24K reasoning tasks from a medical KG and obtain a reasoning model equipped with medical primitives that significantly improves reasoning across 15 medical sub-specialities.
🧵-->
I am very excited about this direction, with lots of cool possibilities arising from using a KG scaffold to generate reasoning curricula. The one direction that seems to be beneficial in an almost immediate sense is using a KG as a simulatable training environment for running RL fine-tuning on LLMs. Each KG primitive along a path can be operationalized as a local, verifiable reward, with the entire path providing a dense reward signal upon successful traversal, thereby encouraging structured reasoning in the process. Our work also lends support to the idea of a compositional model of AGI—one that emerges from the interaction of specialized, superintelligent agents, analogous to how human society develops deeper expertise hierarchically through the coordination of individuals across adjacent domains. (10/N)
I remain pretty convinced that Apollo program Hasselblad photography remains a civilizational high water mark re: vibes / accidental art. And the lesser known ones are in some ways greater.