We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6.
We gave 4 models the same prompt:
Create a glass aquarium whose side panel develops a visible crack and then bursts.
1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.
GitHub repo: https://t.co/aZWYAtakBP
A 24-YEAR-OLD IN WARSAW BUILT A NETWORK OF EIGHT-DOLLAR CHIPS THAT DID EVERYTHING FORTY OFFICES WERE ABOUT TO PAY $3,000 FOR. HE MADE $45,000 IN THREE MONTHS. THE HARDWARE COST HIM LESS THAN THE FIRST INVOICE.
He is 24. Solo. No office.
He drives to a coworking space with a plastic bag of ESP32-S3 boards. $8 each. INMP441 microphones soldered on. 45 minutes on-site. He plugs one into every meeting room. Loads the 28-million parameter model into Flash. Leaves.
The office manager never asked what chip was in the room. She asked if the lights would still turn off if the internet dropped.
Everyone else buys $3,000 NVIDIA DGX Spark or stacks $800 Mac Minis. He deploys distributed swarms. One $8 chip listens. Another parses intent. A third executes over Bluetooth. Total node cost per office: under $200. Total feature set: same as the $3,000 box.
Investment $8 per ESP32-S3 chip (bulk from AliExpress)
Deployment setup fees $15,000 from 30 offices
Retainer revenue three months $7,500 from 20 active clients
Custom skill packages $22,500 from 15 clients
Total revenue three months $45,000
Hours worked per week 15
The enterprise budget was never about the chip. It was about the belief that intelligence had to live in one big box. He shipped the swarm before the invoice cleared.
When you don't know where to start, ship the thing nobody has shipped yet. Three thousand dollars for centralized. Eight dollars for distributed. He picked the swarm.
bookmark this and read the article below
CANCEL your weekend plans.
You NEED to:
• Learn vLLM + SGLang for high-throughput inference
• Build 2-3 optimized serving pipelines with paged attention
• Set up KV cache eviction strategies for long contexts
• Learn speculative decoding + draft model handoffs
• Master quantization tradeoffs (INT4, FP8, AWQ, GPTQ)
• Build your own model router by cost/latency/quality
• Create a token budgeting system per user request
• Experiment with edge deployment (ONNX, TensorRT, WebLLM)
• Try Ollama, LM Studio, LiteLLM for local testing
• Learn continuous batching + request queue management
• Build observability for latency, tokens, errors, costs
• Run load tests with 1000+ concurrent requests
• Learn Kubernetes for AI workloads (HPA, pod autoscaling)
• Use Grafana + Prometheus for inference dashboards
• Build one optimized inference service and benchmark it publicly
• Read inference research instead of model release news
• Start sharing your optimization benchmarks
• Learn how inference costs actually break unit economics
You have way too much to do.
Bookmark & Repost
happy building!
MCP is the embodiment of AI psychosis. Every iteration looked sane and went through peer review. But it ended up as an unmaintainable incoherent mess and now they finally full circled back to REST API
The most influential mathematician of his generation stands in a German lecture hall and explains a foundation he is rebuilding from scratch. Almost nobody watches it.
This is Peter Scholze at Bielefeld University, November 2025, opening lecture of a new public series called Ars Mathematica.
Scholze became a full professor at 24 and won the Fields Medal at 30. His work on perfectoid spaces reorganized entire areas of arithmetic geometry inside a decade.
Condensed mathematics is his attempt to fix something deeper: the way analysis and algebra refuse to sit together properly. He is proposing new foundations for how mathematical objects carry topology.
Watch how he pitches it to a general audience. No prerequisites assumed, no dilution, a working mathematician showing why abstraction opens rooms you could not otherwise enter.
A researcher I know rewatched the middle section twice and said it was the first time condensed mathematics felt like an idea rather than a rumour.
Free on YouTube from a university channel, subtitles on.
Some people prove theorems. He is replacing the floor.
La mejor compra que he hecho recientemente es este NUC, un miniPC que lo tengo como servidor en casa.
Puedo usar Claude Code desde el móvil con Termius + Tailscale como VPN, tiene 32gb de RAM y 1TB de almacenamiento. Va sobrado para todo lo que tengo instalado:
- Mis sideprojects para poder vibecodear desde cualquier lado en el móvil
- Loops con agentes (ej: mejoras de SEO automático cada 2 semanas para mis proyectos)
- Hermes agent
- Jellyfin como servidor multimedia (va increíblemente bien)
- Home Assistant para varias automatizaciones que quiero hacer en casa
Lo único que me falta es un teclado plegable bluetooth para completar el setup 🤣
En wallapop los hay a millones. Nuevos salen bastante más caros, los Beelink tienen bastante buena pinta.
Todo configurado en menos de 30-40 mins con claude, sin hacer prácticamente nada a mano.
This is just horrific… My God… they all need to pay dearly.
Dr. Fauci sowed the fingers and scalps of aborted babies into the backs of research mice, which the scalps grew fine baby hair.
https://t.co/ehYZnztgYQ
🚀 Qwythos-27B-v1 is here! The 27B you've been waiting for.
The bigger sibling to Qwythos-9B. Native MTP head intact, full vision tower, still uncensored, still 1M context. Apache-2.0.
https://t.co/sdX3UupN0H
Well now i'm sure Qwen3.6-27b is all you need no matter how much vram you have.
16GB: Tq3_4s
24GB: Q4_K_M
32GB: Q6_K
38GB: Q8_0
96GB: FP16
What a goated model @Alibaba_Qwen , it will be a shame to not replace it with something like 3.7 or 3.8 😏
A 2-bit quantization delivering ~100% of FP8 on a MoE model like Qwen 3.6-35B-A3B??? 👀
I need to try this on my 3090, as soon as @luceboxai will deliver it to me 😂