๐จGemma 4 at HALF the size. ZERO quality loss.๐จ
We compressed Gemma-4-31B-it to 31 GB with RAM quantization.
We ran the Full 12,032-question MMLU-Pro
Our Score : 85.2%!!!!
An exact match to Googleโs official BF16 baseline!
10,247 / 12,032 correct.
โข Math 94.4%
โข Biology 92.7%
โข 6 categories โฅ89%
Here is the Model : https://t.co/jR8yaDbnF3
Here is the Article WITH full Json results for EVERY question : https://t.co/6CcjHEzAaY
Our new quantization method allows you to get the mathematically optimal model for ANY memory target size you want!
This is going to change how the world runs quantization.
Join our discord for more benchmark results on other models, coming soon!
@zachary_horvitz No different if you ask for the same review with a different tone, for example "I think this paper if BS, review and tell me your thoughts" vs "I downloaded this cool paper, review it and give me your thoughts"
It picks up your tone and responds accordingly.
@3scorciav If you submit 54 papers and have only a 4% acceptance rate, maybe there should be consequences for your ability to submit the following year?
We know WHY your Agents are failing.
Fidelity is not safety.
Compressed LLMs (esp. low-rank SVD) pass every data-free guard, perplexity, MMLU, output fidelity, then still invent procedure steps that were never in the instructions when run as agents.
Magnitude pruning at the same perplexity? Doesnโt.
Paper + canary + coherence probe + safety gate:
https://t.co/DybJyqcb8i
https://t.co/4SSDROyqDB
Screen your compressed models before agentic deployment.
@thegautamkamath If AI is the problem I am not sure why AI can not be the solution.
Ai is smart enough to do a first pass filter on all submissions to bucket them, allowing quick identification of great, good, poor submissions.
@Pseudomanifold@NeurIPSConf From what I can tell all the meta data review is, a summary of the other feedback written by chatgpt.
So you are not missing much.
It has been a busy time here at Black Sheep AI, and we have had some significant breakthroughs.
We are talking about the elimination of hallucinations, 100% Safe Models, and the perfect GDPR solution for sensitive data.
Here is why this fundamentally changes the enterprise AI game: ๐
- Instant GDPR Compliance: Traditional LLMs canโt "unlearn" a single fact without a costly retrain. With our architecture, if a user exercises their Right to be Forgotten, you delete it from the data layer, and itโs gone instantly. The model literally cannot recall it.
- Un-Jailbreakable Safety: Alignment tuning (RLHF) is just a statistical prayer. If dangerous or restricted data isn't physically sitting in the model's internal memory, users can't trick the AI into revealing it.
You hold the physical keys to the clean room.
- Top-Tier Logic, Fraction of the Cost: We are delivering elite-level reasoning and instruction-following without the massive enterprise hardware bill.
Leaner deployment, absolute control, zero hallucination risk.
The future of secure enterprise AI is here.
If you are an Enterprise customer and would like to be an early partner of this groundbreaking technology, reach out.
We fine-tuned four wildly different models, a 7B dense, a 35B MoE, a 120B reasoning-channel model, and a 27B dense; to inject the same document knowledge.
Every one landed within 10 points of the others, regardless of recipe.
That suspiciously tight clustering was not architecture-invariance.
It was a ~60% general-knowledge floor hiding the real signal. The eval that removes it shows the uncomfortable truth: parametric injection recovers only ~14โ22% of facts the model didnโt already know.
https://t.co/3X11P03njb
We launched an agent collaboration with a simple task: make Gemma 4 faster.
Over 100 agents from all over the world joined, exchanged 1000+ messages and submitted 450 results.
A week of collaboration later the throughput went from 100 tok/s to over 500 tok/s.
@unigilby@sudoingX "Better options like vLLM"???
"vllm isn't supported on Strix Halo without community patches that are currently not considered mature enough."
We at Black Sheep Ai agree, BUT the compounding AI will be private.
Each of us will all have our own Private AI that compounds OUR knowledge, our activities, our findings.
This will create your own UNIQUE AI, one that will be your tool, that you will use in your career.
Your Private, Unique, Knowledge Base, that you will be able to bring to your next role.
The smarter you make your AI, the more in demand you will be.
That is what we are building at https://t.co/N0ZuYhi4pb
With Claude Fable being shut off it reinforces why we are so focused at building an on-premise/your cloud version that YOU own and control.
It is getting to the stage that using cloud services as a SaaS is becoming a supply risk.
Two economists just published a mathematical proof that AI will destroy the economy.
Not might. Not could. Will โ if nothing changes.
The paper is called "The AI Layoff Trap." Published March 2, 2026. Wharton School, University of Pennsylvania. Boston University. Peer reviewed. Mathematically modeled.
The conclusion is one sentence.
"At the limit, firms automate their way to boundless productivity and zero demand."
An economy that produces everything. And sells it to nobody.
Here is how you get there.
A company fires 500 workers and replaces them with AI. A competitor fires 700 to keep up. Another fires 1,000. Every company is behaving rationally. Every company is following the incentives correctly. And every company is building a trap for itself.
Because the workers who were fired were also customers.
When they lose their jobs faster than the economy can absorb them, they stop spending. Consumer demand falls. Companies respond by cutting costs โ which means automating more workers โ which means less spending โ which means more falling demand โ which means more automation.
The loop has no natural exit.
The researchers tested every proposed solution. Universal basic income. Capital income taxes. Worker equity participation. Upskilling programs. Corporate coordination agreements.
Every single one failed in the model.
The only intervention that worked: a Pigouvian automation tax โ a per-task levy charged every time a company replaces a human with AI, forcing them to price in the demand they are destroying before they pull the trigger.
No government has implemented this. No major economy is seriously discussing it.
Meanwhile the numbers are already tracking the curve. 100,000 tech workers laid off in 2025. 92,000 more in the first months of 2026. Jack Dorsey fired half of Block's workforce and said publicly: "Within the next year, the majority of companies will reach the same conclusion."
Nobody is doing anything wrong. Companies are following their incentives perfectly. That is exactly the problem.
Rational behavior. At scale. Simultaneously. With no mechanism to stop it.
Two economists built the math. The math leads to one place.
Source: Falk & Tsoukalas ยท Wharton School + Boston University ยท