The race for LLM "cognitive core" - a few billion param model that maximally sacrifices encyclopedic knowledge for capability. It lives always-on and by default on every computer as the kernel of LLM personal computing.
Its features are slowly crystalizing:
- Natively multimodal text/vision/audio at both input and output.
- Matryoshka-style architecture allowing a dial of capability up and down at test time.
- Reasoning, also with a dial. (system 2)
- Aggressively tool-using.
- On-device finetuning LoRA slots for test-time training, personalization and customization.
- Delegates and double checks just the right parts with the oracles in the cloud if internet is available.
It doesn't know that William the Conqueror's reign ended in September 9 1087, but it vaguely recognizes the name and can look up the date. It can't recite the SHA-256 of empty string as e3b0c442..., but it can calculate it quickly should you really want it.
What LLM personal computing lacks in broad world knowledge and top tier problem-solving capability it will make up in super low interaction latency (especially as multimodal matures), direct / private access to data and state, offline continuity, sovereignty ("not your weights not your brain"). i.e. many of the same reasons we like, use and buy personal computers instead of having thin clients access a cloud via remote desktop or so.
Today's big biotech win is that we might be on the verge of a cure for type-1 diabetes🧵
Twelve diabetics were injected with stem cell-derived pancreatic islets.
They started producing insulin again.
One year in, 10/12 participants no longer needed to inject insulin.
Here's a recent talk I gave recapping the last 6-12 months of AI progress, why getting perfect models is hard, how labs are likely approaching the next phase of training (for agents), and other interesting tidbits across the reasoning landscape.
Topics:
00:00 Introduction & the state of reasoning
05:50 Hillclimbing imperfect evals
09:18 Technical bottlenecks
13:02 Sycophancy
18:08 The Goldilocks Zone
19:28 What comes next? (hint, planning)
26:40 Q&A
YouTube etc in replies.
Thanks @corbtt and @OpenPipeAI for hosting me.
10 years ago today, I launched the Math3ma blog. At the time, I wasn’t sure the site would resonate with anyone, but I’ve been amazed by all that’s happened over the past decade! To celebrate, here’s a new post on category theory and language models 🥳https://t.co/EO3ghmMjcT
SpatialLM is very interesting
it's a new model that encodes text prompts (e.g. detect windows) and point clouds projected to an LLM 🤯
the LLM outputs 3D bounding boxes, very simple yet effective approach
two models based on Llama and Qwen-0.5B and Llama-1B on @huggingface
5 AI agent frameworks to build multi-agent applications.
100% opensource.
They aren't LangChain, Crew AI, or OpenAI Agents SDK.
1. Motia is an AI agent framework is built for Software Engineers. Build agents in Python, TypeScript, JavaScript or Ruby.
https://t.co/lwnbOUQK6Q
Our MIT class “6.S184: Introduction to Flow Matching and Diffusion Models” is now available on YouTube!
We teach state-of-the-art generative AI algorithms for images, videos, proteins, etc. together with the mathematical tools to understand them.
https://t.co/wDJcM1YTxJ
(1/4)