Rust is a good prompt compilation target for the moment, but so is C++. And soon assembler. Then microcode. Myopic to think we're going to stop the agentic drill bit until it reaches computing bedrock.
It seems NVMe and DDR4-DDR5 are going to further skyrocket in price until someone starts producing more.
Learned ngrams help with model accuracy and can be efficiently hosted on cheaper memory.
Disaster, buying 128gb ecc ddr5 as we speak.
I made RDMA powered tools rdmasync and rdmapipe to connect Sparks and PCs with ConnectX networking. They let you move your docker images and models (and anything else) around super fast.
I turned them into a linuxbrew tap so its easy for anyone to install them.
https://t.co/kXDPThwQKI
Introducing OUI-1: the first open-weights model for Generative UI
71.7% on Generative UI Bench at 4B params. Beats Gemma 4 31B with 8× fewer active params, and scores 5.5× the base DiffusionGemma it was fine-tuned from.
Methodology, weights, and full benchmark results in the blog 👇
AI models are part of a living timeline.
Instead of retiring them completely, what if we preserved them in a public “Model Museum” — a lightweight, low-cost access layer for research, education, and comparison?
Possible structure: • Read-only or rate-limited access
• Lower-cost “archive tier” models
• Versioned snapshots for study & benchmarking
A living record of how intelligence has unfolded — not just progress, but memory.
I just finished testing Qwen 3.8 Flash.
GLM 5.3 Flash is the better model despite the lower tok/s, which don't matter because it's faster to get to result anyway.
GLM: best primary coding worker. Slightly better first-pass quality than DS4F and dramatically more efficient.
DS4F: ultimately matched GLM’s quality, but with much higher time/token usage.
Qwen: excellent for small parallel jobs, but not as reliable for substantial work.
Qwen is useful as a fast specialist for short parallel jobs, but its long-task reliability, stopping discipline and tool consumption remain substantially worse.
🔥 https://t.co/8P3upHAaGY just released a super cool project!
OpenVuln🛡️ a public vulnerability intelligence platform for open source.
https://t.co/Vy3xiGW6pQ
How it works:
- Submit your public GitHub repository
- VulnHunter AI engine scans for vulnerabilities
- Public aggregate security insights
- Detailed findings stay private for verified maintainers until disclosure
Wrote a piece on writing good evaluators, main take-aways:
- go for a code evaluator when you can
- don't rely on what the agent said it did, you need to actually verify
- calibrate your LLM as a judge against labeled data https://t.co/3wWoS5eDh6
The (software) factory is the product. Your product is only as good as the agents you set up to autonomously maintain it.
This is what @elonmusk figured out about Tesla, and it's now true for the software world.
We are pleased to highlight an excellent community model from developer : Qwen3.6-27B-MTP-pi-reasoning-GGUF.
Built on our Qwen3.6-27B base model, this release focuses on optimizing automated programming and debugging workflows for local coding agents.
If you are exploring local AI coding assistants, we encourage you to test this model in your environment.
a masterclass in coding agents from the head of anthropic.
there’s still a tonne of leverage in knowing how to use these systems optimally and this is the best i’ve seen.
make sure to bookmark so you can watch again and again chat
BREAKING: Microsoft just showed that the hardest part of AI research can't be automated yet.
An AI agent replicated 3 weeks of expert work in 1 day. But it plateaued at 70% quality. The jump to 100% required a human to look at failure patterns and make a structural decision the AI kept missing.
The last 30% is still a human job.
Microsoft Research built an AI system that evaluates whether computer-use agents actually completed their tasks.
Think of it as an automated judge that watches an AI browse the web and decides: did it succeed or fail?
Getting this right matters a lot.
If your judge is wrong, every benchmark score you've ever seen is wrong.
Every training signal your agent learned from is corrupted.
The existing judges WebVoyager and WebJudge had false positive rates above 45% and 22% respectively.
That means nearly half of all failed agent tasks were being marked as successes.
Microsoft's human expert spent 3 weeks iterating to fix this.
Across 32 experiments, he discovered four structural design principles that brought the false positive rate down to near zero.
Then Microsoft gave an AI agent the same starting point and the same goal.
> The AI finished in 1 day.
> It hit 70% of the human expert's quality.
> Then it stopped improving.
The gap between where the AI plateaued and where the human landed came down to one thing:
→ The AI made incremental edits — tightening thresholds, adjusting language for individual failure cases
→ The human made structural bets — looking at hundreds of failures and inventing new scoring categories
→ The AI's edits were conservative and safe — never increasing false positive rate
→ The human's biggest gains came from opinionated, high-level rules that required judgment, not data
→ One human insight alone — "separate nitpicks from critical failures" — drove a step-function jump the AI never discovered
The AI was given the same principles the human used.
It had the same experimental infrastructure.
It ran the same tests and committed changes to version control just like the human did.
But when the human saw an agent get penalized for rounding $5.95 to $6, he derived a general rule.
The AI saw the same failure and tightened the language for that specific case.
One approach scales. The other doesn't.
There is a twist though.
When the AI was given the human's best work as a starting point, it actually surpassed the human expert.
It found improvements the human couldn't find through fine-grained optimization of an already-strong foundation.
The lesson: human expertise and AI optimization play completely different roles.
Humans are essential for discovering the core structural principles.
AI is better at the fine-grained tuning that extracts the remaining performance once those principles exist.
The current framing of "AI replaces human researchers" misses this entirely.
The real workflow is: human does the hard structural thinking, AI does the exhaustive optimization on top.
The last 30% isn't a gap that closes with more compute or a stronger model.
It closes with judgment.
And judgment, for now, still belongs to the human.
Silicon Valley is quietly running on Chinese open source AI models.
Here are the receipts:
→ Cursor confirmed last month that Composer 2 is built on Moonshot's Kimi K2.5
→ Cognition's SWE-1.6 model is likely post-trained on Zhipu's GLM
→ Shopify saved $5M a year by switching to Alibaba’s Qwen model. Airbnb CEO Brian Chesky has also said: "We rely a lot on Qwen. It's very good, fast, and cheap."
And now Zhipu dropped GLM-5.1, an open source model that performs almost as well as Opus on coding benchmarks.
📌 More on the Anthropic + OpenClaw drama and what I'm learning about AI on the ground in China in my new post: https://t.co/cm9jYIZS8y
🎙️Introducing Max Agency
Max Agency is a new podcast where we go deep on how the best agents are actually being built: architecture decisions, tradeoffs, evals, and everything in between. Each episode, I sit down with engineering leaders who are doing this work in production.
Our first episode features Izzy Miller (@isidoremiller), AI Engineer at Hex (@_hex_tech). Hex has been shipping data agents since before most teams were even thinking about them, starting with single-cell text-to-SQL and graduating to a full Notebook agent that can work autonomously for 20 minutes on a complex analysis.
Izzy has a lot of perspective on what it actually takes to get agents working well in production, and what breaks along the way.
A few takeaways from our conversation:
- Keep your eval sets small enough to hold in your head: Izzy runs 30-50 handcrafted "traps" with multiple repetitions, rather than hundreds of variants. If you can't explain why your agent fails each one, your eval set is too big
- Day zero performance is almost irrelevant: The more interesting question is how the agent compounds. Izzy is building a 90-day simulation where the warehouse evolves and the agent has to accumulate understanding
- You can catch agent errors without seeing the raw outputs: By running an LLM-as-a-judge over production usage and clustering the results, you can surface places where something likely went wrong, without needing to read individual conversations
Watch the full episode on:
- Youtube: https://t.co/AdkQbV3Pq2
- Apple Podcasts: https://t.co/1MKF7mcYSr
- Spotify: https://t.co/DxACw24oob
A long time coming but new mlx-lm is here with better batching support in the server and Gemma 4.
pip install -U mlx-lm
Here is a video where a single M3 Ultra serves 5 opencode sessions with Gemma 4 26B that process ~130k tokens in ~1.5 minutes.
LLM Knowledge Bases
Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:
Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki includes summaries of all the data in raw/, backlinks, and then it categorizes data into concepts, writes articles for them, and links them all. To convert web articles into .md files I like to use the Obsidian Web Clipper extension, and then I also use a hotkey to download all the related images to local so that my LLM can easily reference them.
IDE:
I use Obsidian as the IDE "frontend" where I can view the raw data, the the compiled wiki, and the derived visualizations. Important to note that the LLM writes and maintains all of the data of the wiki, I rarely touch it directly. I've played with a few Obsidian plugins to render and view data in other ways (e.g. Marp for slides).
Q&A:
Where things get interesting is that once your wiki is big enough (e.g. mine on some recent research is ~100 articles and ~400K words), you can ask your LLM agent all kinds of complex questions against the wiki, and it will go off, research the answers, etc. I thought I had to reach for fancy RAG, but the LLM has been pretty good about auto-maintaining index files and brief summaries of all the documents and it reads all the important related data fairly easily at this ~small scale.
Output:
Instead of getting answers in text/terminal, I like to have it render markdown files for me, or slide shows (Marp format), or matplotlib images, all of which I then view again in Obsidian. You can imagine many other visual output formats depending on the query. Often, I end up "filing" the outputs back into the wiki to enhance it for further queries. So my own explorations and queries always "add up" in the knowledge base.
Linting:
I've run some LLM "health checks" over the wiki to e.g. find inconsistent data, impute missing data (with web searchers), find interesting connections for new article candidates, etc., to incrementally clean up the wiki and enhance its overall data integrity. The LLMs are quite good at suggesting further questions to ask and look into.
Extra tools:
I find myself developing additional tools to process the data, e.g. I vibe coded a small and naive search engine over the wiki, which I both use directly (in a web ui), but more often I want to hand it off to an LLM via CLI as a tool for larger queries.
Further explorations:
As the repo grows, the natural desire is to also think about synthetic data generation + finetuning to have your LLM "know" the data in its weights instead of just context windows.
TLDR: raw data from a given number of sources is collected, then compiled by an LLM into a .md wiki, then operated on by various CLIs by the LLM to do Q&A and to incrementally enhance the wiki, and all of it viewable in Obsidian. You rarely ever write or edit the wiki manually, it's the domain of the LLM. I think there is room here for an incredible new product instead of a hacky collection of scripts.