This is Gemma 4 12B QAT with MTP running local at ~76k tok/s with llama.cpp on my crusty old RTX3080. Slower than E4B, but hopefully better at coding due to billions more parameters.
I'm going to wire it into VScode through @opencode and the @cline extension next. #LocalLLM#AI
@opencode@cline This is how I got it working on Win10 w/Powershell:
llama-server -m "D:\_ai\llm\gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf" --model-draft "D:\_ai\llm\gemma-4-E4B-it-Q4_0-MTP.gguf" --port 8888 --temp 1.0 --top-p 0.95 --top-k 64 --spec-type draft-mtp --spec-draft-n-max 4 -ngl 999 -fa on
@hashbreaker The Post- prefix indicates obsolescence; especially after long-lasting stability of a dominant paradigm.
Post-truth is especially interesting because in took us our entire history (~200k years) to ruin our ability to trust our senses and destroy institutional authority.
Just got @opencode working with @UnslothAI's Studio. Had to patch a JSON file to add my own provider. Default only allowed LM Studio for #LocalLLM work.
Seems reliable and fast with Gemma 4 w/QAT and MTP. But, I don't trust it for coding. It pretends to understand pretty well.
#Veridium is going to turn AVX/NEON-capable hardware ownership into a yield bearing productive asset class with ~5% thermodynamic cap on the fast-prover advantage.
Maybe: Cheap and powerful SBCs for all.
Certainly: Disruption to incentives for labor, capital, and consumption.
@jamonholmgren 2.5 - Agentic engineering
- Human and agent collab on a strategy doc until cutting a multi-phase plan is viable
- Agent codes while driving plan adaptations with explicit human approval
- Human reviews code, commits, and decides on spec transitions (archived, future, etc)
- Ship
@charliermarsh I didn't choose Rust. I chose TypeScript as my go-to for client, server, and desktop long ago. Rust just kept popping up as the best tool.
Rust chose me.
I love what Rust does for me. Six months from now, I'll probably start buying the Ferris merch.
@luhelminger I'm currently building on the abandoned Binius v0 with FRI. Recursive STARKs over a tower of binary fields. Relevant!
But, I plan to write a GPU-based prover to verify my mining fairness scheme; so Binius64 might still be in the cards.
@therealdanvega We, the human race, are adapting to the GPT era by growing increasingly apathetic to all things that merely look incredible or sound amazing.
Once the attention-wall fully solidifies, disruptive technologies will be the only thing that can pierce it organically.
@Steph_Curdy I didn't know WolframResearch has a blockchain group.
I'd like to connect with researchers willing to discuss my Rule 30 VDF construction.
Irreducibility + addressable entry = provable sequencing
Assists to cap fast chip wins to ~5% under Monte Carlo. A true physics unlock.
Trying @GoogleDeepMind's Gemma 4 via @UnslothAI 's new Studio app on my Win10 host via a WSL2 Ubuntu guest in a terminal inside @antigravity IDE with Claude Code.
Gets ~69 tok/s (host side) on my RTX 3080. 200k context window!
Can it code? π€·ββοΈ Using it for docs. #LocalLLM#AI
@annkkitaaa perfect > statistical >computational π€
I might like this better: ideal > precise > effective
Are there applications to these distinctions?
One ZK concept I think deserves understanding is WHY the Fiat-Shamir transform matters; only the non-interactive proof is transmissible!
@realbarnakiss I'm using Claude Code via Antigravity IDE. It just realizes when it needs papers and grabs them. Opus 4.8 medium/high. The one bit of busy-work I used to enjoy; gone.
@EliBenSasson@lucapratadotcom@corygabrielsen Decentralized AI (open weights, localllm, bittensor, etc) is both a tool and a social movement towards self-empowerment and sovereignty over the inheritance of human knowledge.
Centralized AI is a system that controls your thoughts and actions; just not fully expressed yet.