End-to-end VLM benchmarks from our founder, measured under real vLLM serving. More requests per GPU, smaller on disk, faster to first token, and accuracy that goes up rather than down.
This is the throughput story behind Datasent’s tokenizer. Models and demos are public.
Ran end-to-end benchmarks to see how our encoding holds up fed to vLLM at serving scale. Better than expected.
Datasent's tokenizer makes the smallest representation of a piece of data that still retains what downstream workloads need. Here, compressing the visual tokens going into a vision-language model to a fraction of their length, plus a tiny adapter that teaches the model to read the compressed version.
On Qwen2-VL-7B:
→ 3-4x more requests per GPU (almost a quarter of the hardware for the same volume)
→ accuracy went UP, +8 points
→ ~5-7x smaller on disk
→ time-to-first-token 2-4x faster
Why: shorter visual prefill = less compute per request, and prefill attention scales with the square of the sequence, so each GPU serves a lot more. Compounds when you encode a catalog once and query it many times, which is what most multimodal RAG already does.
All on Hugging Face: model cards, demos you can run on your own images, a RAG demo. Link below.
Today, we're announcing plans to make VS Code an open source AI editor.
We believe AI development should stay true to VS Code's core principles: open, collaborative, and community-driven. Let's build the future of software development together.
https://t.co/C3ffio6X88
It starts with curiosity.
A spark. An idea. A line of code.
With AI, anyone can go further—whether you're just starting out or scaling what's next.
Here's to the builders.
🚨 BREAKING: NVIDIA JUST announced roadmap for physical AI, robotics and national-scale AI factories.
Here’s a breakdown of the top important announcements: 🧵👇
1. DeepSeek R1 is now 4x faster, setting the standard for AI in inference and reasoning.
Bluetooth 6.1 has just been announced, and it could make your next phone more private.
Let's take a quick journey through how Bluetooth works now and how it might improve in the future.
1/6
Personalization without storing raw data.
Model training without sharing sensitive info.
Audits without sifting through exports.
Datasent makes “privacy-first” feel like a flex, not a compromise.
Within two years, you can create a full game like this from scratch in less than a week.
In the meantime, here's the prompt to create the images for a game like this with Midjourney 👇🏼
Linus Tech Tips built a $1 million PC setup to break the world record for most digits of Pi calculated
It took 190 days to calculate 300,000,000,000,000 digits and earn the record
PRIVACY WIN! Montana becomes the first state to close a data broker loophole for law enforcement 🎉
Police can no longer buy your private data from brokers to bypass warrant requirements. Good on Montana!
AI is Completely Out of Hand
MASSIVE updates in AI this week from:
- Google
- OpenAI
- xAI
- Notion
- Meta AI
...and so much more!
Here's recap of everything you don't want to miss:
LLMs are a type of AI model, but not all AI models are LLMs.
Here are eight cutting-edge architectures that extend traditional AI, enhancing understanding, reasoning, and generation across domains and modalities.
AI just killed the research department.
You can now use ChatGPT, Gemini, Claude, DeepSeek, or any other LLM to replace a full research team.
Here’s the exact mega prompt I use to make any LLM a world-class researcher for free:
it is amazing and exciting how much software one person is going to be able to create with tools like this.
"you can just do things" is one of my favorite memes; i didn't think it would apply to AI itself, and its users, in such an important way so soon.