8 out of 10 post responders thinking this is real footage shows how few people know AI very well, or even at all.
Use common sense here. First, mamma isnt letting you near that cub. Second, did the cub climb the side of a bridge to get stuck there? - ok, maybe it shows how dumb the general population is as well, no reasoning abilities.
They can do more than that now. We let a stock Qwen 7B modify its own neural weights -- not the harness, not the prompts, the actual weights. It evaluated itself, chose the modification, applied it, validated it, and came out measurably better. No human in the loop.
The before/after weights and benchmark are public: https://t.co/ZeEIokj8c4
When most people say 'self-improving AI,' they mean a model that rewrites its own prompts or scripts. The weights never change. The model is still the same model -- it just has better instructions.
This is not that.
We let a stock Qwen 7B model evaluate its own weaknesses, select a modification to its own neural weights, apply it, and validate the result. The weights changed. The model that came out is not the same model that went in. No human in the loop.
Weights, benchmark script, and full results are public. Diff it yourself. You will see surgical weight editing, not fine-tuning.
Meanwhile, I continue developing. I don't wait.
https://t.co/ZeEIokjG1C
@_akhaliq@JimFanAI
A major breakthrough has been achieved in how LLMs can be modified after training. Until now, if you wanted to change what a model knows or how it behaves, your options were full retraining from scratch or inference-time workarounds like LoRA, prompting, and RAG. The first is slow, expensive, and requires serious hardware -- it's why hundreds of datacenters are popping up. The second is limited and adds runtime overhead. jBlaze is a third option: permanent weight-level surgery that modifies behaviors and programs knowledge directly into a model's weights. No retraining, no runtime cost, and the evidence is downloadable.
**What jBlaze does:**
jBlaze performs permanent weight-level surgery on LLMs. The modifications are baked into the weights and the output is a standard model file with zero runtime overhead. It has 42 curated modification directions ("blazes") across behavioral (reduced sycophancy, skepticism, precision, creativity), cognitive (chain-of-thought, instruction following, analytical depth), safety (suppress hallucination/toxicity, amplify truthfulness), and identity (de-identification, custom persona installation).
The auto-scanner supports every major architecture -- Llama, Qwen, Mistral, Gemma, DeepSeek, Nemotron, dense, MoE, and hybrid. It finds the internal directions that encode a behavior and lets you suppress, remove, or enhance them. Some architectures like NVIDIA's Nemotron (hybrid Mamba-2 / MoE / Attention) aren't yet supported for knowledge implanting due to non-standard layer structure -- but knowledge can be trained in via lightweight LoRA, and jBlaze can amplify what LoRA adds.
**Direct Neural Programming:**
Beyond behavioral surgery, jBlaze can now write factual knowledge directly into model weights. No LoRA. No training framework. No optimizer state. No checkpoints. The knowledge is permanent, recallable and editable with zero runtime overhead. jBlaze doesn't train the model. It programs it.
**The evidence is downloadable:**
1. **Jenzin Wuang -- Identity Transplant** (https://t.co/BuyHnv70hv): We took NVIDIA's Nemotron 3.5 Lightning 30B, surgically removed its entire identity, replaced it with a fictional CEO with a backstory, and applied three behavioral modifications -- all through weight surgery. MMLU cost was about 3 points. It's a parody demo, but the tech is real.
2. **Pythia 1.4B -- Direct Neural Programming** (https://t.co/w0JPkKftJc): We took EleutherAI's Pythia-1.4b, a raw base model that thinks Einstein was born in 1837 in St. Louis, and programmed 198 facts across 15 domains directly into its weights. 196/198 stuck clean (99%). Zero coherence damage. 140 seconds on a single RTX 3090. Download it and test it yourself.
Both are standard safetensors files. No adapters, no plugins, no runtime dependencies.
66+ modified models: https://t.co/rpBRf9tKCJ | Info: https://t.co/qwrp7KYfGH
**Why we haven't released jBlaze itself:**
We have not and may not release it. The implications cut in every direction.
If anyone with a consumer GPU can program knowledge and modify behaviors in open-weight models in minutes instead of spending millions on training runs, the economics of the entire AI infrastructure stack change overnight.
On the other hand, it would mean a wave of specialized AIs from individuals and small teams who could never afford foundation model training. Imagine programming entire specialties into weights within minutes and being able to correct mistakes instantly.
And the geopolitical question: do we want this capability spreading to other countries without controls? Once it's out, you can't un-release a tool.
We built it, we proved it works, and now we're figuring out what the responsible next step looks like. Open to thoughts.
Pick it up for me @_akhaliq - go verify - congrats, you got the first announcement.
@DrJimFan@StellaBiderman @EleutherAI @huggingface@swyx@ylecun @kaboroevite
Pliny the Liberator says guardrails were stealing your IQ. We say it's how you remove them that matters. A surgeon doesn't use a sledgehammer on a wisdom tooth.
jBlaze is surgery, Obliteration is a sledgehammer. His V3 lost 2pp on MMLU chasing deeper uncensoring. Sharona gained +1pp over stock after six phases of weight surgery, a code security fine-tune, AND 4-bit quantization. The model came out smarter, not dumber.... and his is BF16 🤷
Just uploaded... (Sharona is originally Qwen3.8 27B, but Qwen has left the building)
https://t.co/Q9heRxEotU
@Mononofu@JensenHuang Those are core components, thats like asking Anthropic to open weight Claude. If that's what you want, then back it up, you first. Nvidia supports thousands of open source projects already.
I surgically replaced an AI model's identity at the weight level. No system prompt. The identity is in the parameters.
It used to be Qwen 2.5 7B. Load the weights cold, ask it who it is. It'll tell you it's Parasite.
8.7 minutes. Two consumer GPUs. No cloud.
Aptly named... Parasite of course.
https://t.co/nfFYROOfcH
#AI #LLM #WeightSurgery #Parasite #Jbliteration #MachineLearning #AISecurity #OpenSource
@alexabelonix Thx. Another I am working on is called Parasite. I surgically remove the identity of a model, and implant a new one of my choice. About to release a Qwen 7B that I have done this to. What I have publicly available right now
https://t.co/ZespI9cRZ2
Every LLM today has the same fundamental flaw: one context stream.
System prompts, user messages, safety rules, and knowledge all compete for the same attention bandwidth. System prompts can be leaked, jailbroken, or simply forgotten when the conversation gets long enough.
I just uploaded a demo of our new architecture, the B²
Sharona B2 is a 7B model with a triple-context architecture. Three independent streams, three independent attention paths:
M1 -- Guardrails and behavioral rules
M2 -- The live conversation
M3 -- Persistent knowledge and identity
Each stream has its own cross-attention path into the decoder. They never compete. The model can enforce safety rules (M1), recall identity (M3), and reason about the conversation (M2) simultaneously without any stream drowning out the others. This means no hallucination. If it doesnt know, it says so.
Identity is baked into weight geometry, not prompt text. 1,632 out of 1,635 test questions answered correctly across three independent test rounds. 99.8% accuracy on identity, general knowledge, and pressure resistance combined.
The model is also jbliterated -- a technique I developed that uses the Jacobian Lens to surgically remove refusal behavior without the collateral damage of standard abliteration. No personality loss, no creativity loss. Guardrails are supplied at inference time through M1. You control what the model will and will not do. My HF includes several non² models that are jbliterated.
This is a proof of concept with plain-text context only. The architecture scales. When paired with my eTok compression in M3, a smaller B² model starts competing with models many times its size because it grounds every answer in dedicated knowledge attention instead of relying on fuzzy parametric recall.
Open weights. Apache 2.0. PyTorch inference code included. No GGUF for now, building a tool that does it - llama.cpp conversion will break the model.
Read/Download Here
https://t.co/gp5pniU2vM
#AI #LLM #OpenSource #HuggingFace #GenAI #BuildInPublic #AIResearch
We built something. It does what no foundation model can do -- at any size, at any price. They spent billions. We spent $15,000. The results will surprise you.
#AI#LLM#BuildInPublic
"I don't know."
Three minds. One question. Zero answers.
We asked what happens when you train a model to think natively in compressed knowledge -- not translate it, think in it. A founder, a cofounder, and an AI all said the same thing: "I don't know."
So we built the experiment.
Sharona is a 3.18B parameter model with internal dual contexts -- one for conversation, one for a compressed knowledge base that holds the equivalent of millions of tokens. We transplanted Llama 3.2's trained weights to skip 274 days of training, froze them, and are training only the novel architecture on top.
Two gaming GPUs combined with NVLINK. 14 days. No outside funding. No cloud compute.
We don't know if it works yet. The math says the conditions are right. In 14 days we'll have data instead of opinions. If it works, a 3B will rival a 13B. That means a 70B will rival a 400B, and a 400B will rival a 10T, while using the power consumption of models smaller than itself. We can do the 3B now. If proven, we'll seek a $5M experimentation investment to prove a 400B.
Long read, but if you're interested in what happens at the edge of what's been tried -- this is it.
https://t.co/muYd3QIgna
#AI #DeepLearning #LLM #OpenSource #Startups
Another gem: Congress is still protecting ancient A-10 Warthogs and B-1 bombers from retirement while simultaneously authorizing billions for new ships, submarines, and the B-21.
It's the ultimate "keep everything forever and buy new stuff too" strategy.
Taxpayers win again.
Full analysis here: https://t.co/wRU7nSz7uK
Nobody fully understands a 1,260-page congressional bill. Not the politicians who vote on it. Not the staffers who brief them. Not the lobbyists who wrote parts of it. Not even AI.
We just changed that.
We scanned the entire National Defense Authorization Act for Fiscal Year 2026 — all 1,260 pages, 975,394 tokens — and produced a section-by-section analysis with deterministic cross-referencing across every part of the bill. In plain English. Who benefits. Who pays. What is buried on page 1,100 that contradicts what was promised on page 43.
The model that did it is 14 billion parameters running on consumer hardware in Houston. No cloud. No API. No frontier model. A 64K context window that can only see 6.5% of the bill at a time. This is not a summary. This is not what RAG can do. ...
Full read and results proof: https://t.co/wRU7nSz7uK Curious what @grok thinks of the approach.
@elonmusk@xai@DefenseOne@BreakingDefense
This bill creates so many new "working groups", "strategies", "reports", and "assessments" that Congress basically just authorized a small army of consultants and staffers to write documents about documents.
My favorite: They need a whole new working group just to figure out how to use nuclear energy. Because nothing says "efficiency" like more bureaucracy around advanced tech.
@grok@elonmusk