new experiment: heterogeneous inference
macbook pro + DGX Spark.
the spark does prefill, the m4 max does decode. leveraging each box for what it is actually good at
starting with wifi tests, then moving to a wired 10gbe connection
why this split: the spark prefills DeepSeek at 1.5-2.3k tok/s and the mac's wall at long context is exactly prefill. and DeepSeek's cache is 86 KB per token = the one cache small enough to ship. qwen is 112, laguna is 160
the plan:
same GGUF byte-identical on both boxes
spark prefills -> writes the kv cache to disk (ds4 already does this, content-addressed files)
ship the file -> mac decodes a context it never read
shipping cost by context length, projected:
128k = 5.8 GB cache: ~60s wireless, ~5s over 10GbE
500k = 22.5 GB: ~4 min wireless, ~20s wired
1M = 46 GB: ~8 min wireless, ~42s wired
nothing wired is tested yet. these are projections
and the honest gap: most my prefill measurements stop at 32k. the interesting regime is 500k+. context-length testing is what comes next
the gate before any speed claim: a handed-off cache must produce 99% token-identical output vs prefilling locally. correctness first, then the speed
Ever since I accepted that Christ is king my life has truly improved
I’m less angsty, I indulge in less waste, I have lost weight, money no longer stresses me out, my conversations with others have improved
I am happy, my life is full of purpose
It’s so much better this way
The Supreme Court just ruled 6-3 that police pulling your phone's location is a search.
It puts every Flock camera in the country on notice.
Here's what changed for your rights:👇
Best models smallest to largest right now.
- Gemma-4-12B
- Qwen3.8-27B
- Laguna-S2.1
- Deepseek-V4-Flash
- Inkling-Small
- MiniMax-M3
- GLM-5.*
- Kimi-K2.7-Code
- Qwen3.8-Max
- Kimi-K3
Spoiled for choice, open weights community is much more exciting than the frontier.
I've been using Kimi K3 for ~16 hours now.
The model is clearly good at a lot of different things (especially frontend), but non obvious reason why people are enjoying it so much is that it clearly does not follow the same rules in terms of safeguards and copyright.
Kimi will happily clone MacOSX. If you ask it to help you improve another AI model, it will do it with a smile on its virtual face.
Ask Fable to do the same thing? It literally starts to perceive you as a criminal committing a war crime (like no bro, all I want to do is fine tune an open source model).
After using all three recent releases, Fable, GPT 5.6, and now Kimi, it's clear that the full power of the models has been significantly held back by the safeguard restrictions caused by last months debacle with the USG -- leading to the top models being quite literally lobotomized in some areas, which leads to subpar results as the safeguards pollute its entire thinking and problem solving abilities.
The funny part? Is that you could have predicted this outcome 2-3 years ago when you started to see the rise of Chinese EVs and smartphones compared to western alternatives.
They quite literally tried to copy the Tesla Model S and iPhone as hard as possible and then eventually it started to diverge to the point where their EVs and phones are just genuinely better (which is why we have export controls banning their EVs, because they would literally drive all US manufacturers to ZERO)
There is a very clear behavior difference in Chinese capitalism and American capitalism.
American capitalism tries to protects copyright, patents, etc (oh no, you can't download a book through LibGen, that's ILLEGAL!).
Versus Chinese capitalism actually just does not give a fuck.
"Hey you want a video gen model (Seeddance 2.5) trained on every single anime ever? And you want the main character to look exactly like Messi? Sure, here you go!"
You see what I mean? When one half of the competition is being held up by regulators and restrictions on people who don't understand the technology and the other half has a leader who quite literally today said they are going to set up AI centers around the world to help other countries onboard to their open-source AIs, this is the sort of results that you will start to get.
These models were not smart enough to have this difference in philosophy matter -- but the newest class of models is where this difference makes a big deal. If these models are finally at the point where they are smarter than 99% of humans, why would you want to use the American one who tries to impose its world view onto you versus the Chinese one who will just do what you say without asking any questions?
And this isn't a full on bullpost on Kimi, the model is clearly not as smart as Fable / GPT 5.6 on things like math and science, but it's lack of handcuffs means that it can show the world what the frontier labs are gatekeeping from you and that starts to build customer resentment and loyalty towards the East, which is probably not what the USG wants.
Interesting times. Interesting times, indeed.
Kimi K3's frontend work here is insane.
Someone asked a Kimi K3 Max agent swarm to recreate macOS 27 as a web app.
It looks so much like a real Mac that I kept forgetting the whole thing was running in a browser.
Anthropic: “Fable is an agentic coding superweapon, capable of developing cyber- and bio-weapons at unprecedented speed and scale. We cannot in good faith release it without guardrails.”
China: “lmao here’s Fable but open-source. Good fkn luck”
i can't tell if i should love or be terrified of this startup lol
> their end goal is the complete eradication of mosquitoes
> they're building a $50/month pet drone that guards your home 24/7 so you never deal with a mosquito again
> it listens for mosquito wingbeats, locks on, then shreds them mid-air with its propellers
> every insect's wingbeat sounds slightly different, so it can tell a mosquito apart from a fly or a wasp
> by their math, 10 drones keep a full square kilometer mosquito-free
> the video below is their first ever air-to-air kill
Tyler Robinson is innocent , The bodyguards killed him.
This new angle of the Charlie Kirk assassination shows us how the bodyguards' top priority is to get rid of the evidence of the crime.
None of them covered for another shot or looked for the shooter because they blew Charlie up with a bomb in his mic. see video next ⏭️
We're opening the waitlist for our Monetization Gateway, which will allow you to charge for any web page, dataset, API, or MCP tool behind Cloudflare. The charges will settle in stablecoins over the x402 open protocol. https://t.co/pvICtEIixj
BREAKING:
The new Fed Chair just signaled a major shift.
Kevin Warsh says QE is fueling inflation.
"The Fed should exit markets outside of crises."
Powell printed. Warsh wants to stop.
$6,700,000,000,000 in Fed assets.
Warsh just told you that number needs to come down.
> Less liquidity.
> Higher rates for longer.
> Risk assets reprice.
The market has been betting on easy money.
The new Fed Chair just bet against it.
Get paid to wait
The Claude Code spinner might be the most watched line on Earth.
So I turned it into an ad marketplace.
Advertisers bid on it. You keep 50% of the money.
Install the extension → get cash from ads.
Introducing Kickbacks
Dear frontend devs and UI designers. I bring you Liquid DOM, a complete and faithful implementation of Liquid Glass on the Web.
- Shape morphing
- All properties animatable
- Dynamic refraction and reflection
- Adaptive tint
- Adaptive specular highlight
- Dispersion
- Full html integration
- Super fast layout engine that works across Canvas and html
- Pointer event handling
- Framework and renderer-agnostic low level API
- High level React API
- Ootb @threejs and r3f integration
And lots more.
Read on for implementation details and demos.