DeepSeek V4 Flash 0731 is now 90% off on Nous Portal for the next 7 days, in partnership with @novita_labs.
At this discounted price, it is over 1000x cheaper than Fable 5 on comparable tasks while still beating it on Terminal-Bench 2.1.
Try it today at https://t.co/4KFwhsReyA
Josh Hawley brings the receipts and exposes Dr. Fauci for using 8 different federal employees on federal time to chase over $1,000,000 in cash prizes for himself.
"You got RICH while people were dying."
"You were using federal employees with TAXPAYER MONEY to apply for and solicit cash prizes for you personally. Cash prizes totaling over a million dollars."
"And what were these individuals doing in the depth of the pandemic in November of 2020 when millions of Americans were suffering from COVID? What were they doing? I tell you what they're doing. Folkers was on your behalf soliciting and gathering information for a cash award."
"Let's look at it. We've got his email right here. Right over my shoulder. Folkers says, 'I'm working on this nomination for the Dan David award for Fauci. We need to beef up the COVID part. Do you have language in the that you could share that delineates how we responded in new ways to Covid.'"
"He sends this to multiple federal employees. What was the Dan David award? Do you remember, doc?"
FAUCI: "On the advice of counsel, I respectfully decline to answer based upon my rights under the Fifth Amendment to the Constitution."
HAWLEY: "It was a $900,000 cash award. $900,000 cash award. And he got it because he used federal employees to get it. And this wasn't the only award, was it, Dr. Fauci? In fact, you applied for and received at least eight other federal cash prizes on federal time using federal employees and federal resources.
"Here they are over my shoulder. Besides the Dan David award, you've got the Partnership for Public Service. You've got the Adelson prize, you've got the Smithsonian award, you've got the National Academy of Medicines award, you've got the CDC Foundation. In fact, you turn your staff into a full-time application machine."
"You actually wrote to people and said, 'Do you think maybe I'd qualify?' And you got cash for all of this. And it wasn't just one or two employees, was it? In fact, you used eight separate federal employees on federal time using federal resources to solicit cash awards. Isn't that true?"
FAUCI: "On the advice of counsel, I respectfully decline to answer based upon my rights under the Fifth Amendment to the Constitution."
HAWLEY: "Here they are, right over my shoulder. Here's all of them. The people that you use federal employees so that you could go and get cash money."
[Names listed on board]
– Greg Folkers
– Lawrence Tabak
– Holli Jaffe
– Patricia Conrad
– Courtney Billet
– Louis Miller
– Katherine Miller
– Hugh Auchincloss
"And the most hilarious part is that you weren't content to even dragoon them into applying for all of these cash awards. You then had them go out to the ethics agencies and ethics watchdogs and demand that you be able to get the money. That's right."
"We've got it in your emails. Your staff and employees on federal time using federal resources, being paid with federal dollars, are soliciting cash for you. And then they're turning around and saying to the agency, 'We don't want just an answer. We want to get to yes.'
"Because Fauci wants the money."
If “distillation” is theft, then almost every model faces the same criticism: they have all learned from vast amounts of human-created content across the internet, news, books, forums, videos, and more. If your own model is built on knowledge taken without permission, can you really claim copyright over it?
Today we're shipping Nemotron 3 Ultra.
A 550B MoE frontier-intelligence open model built for long-running agents.
It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
We partnered with @trajectorylabs to post-train NVIDIA Nemotron 3 Ultra for legal. Here’s what we found:
1) Open-weight models can reach frontier legal performance.
On our Legal Agent Benchmark (LAB), Nemotron 3 Ultra started at a 0% all-pass rate. After post-training, it reached 5.8%, placing it between Sonnet 4.6 at 4.2% and Opus 4.6 at 6.6%.
2) Post-training dramatically improves reliability.
Before training, many held-out tasks missed enough rubric dimensions to land around ~70% pass rates. After training, those tasks shifted toward ~95% pass rates.
3) Open-weight performance comes at much lower cost.
Post-trained Nemotron 3 Ultra reached a similar quality band to leading closed models while running at roughly 1/8th to 1/50th the per-token price of Sonnet 4.6 and Opus 4.6.
Most importantly: we post-trained this model on the @trajectorylabs platform less than 24 hours after Nemotron 3 Ultra launched, using the same harness, data, and recipe we used for Nemotron 3 Super.
More to come as we continue to experiment with open-weight legal agents.
Read more on post-training with Trajectory below:
Meet DiffusionGemma!
An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.
Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Gemma 4 QAT is here.
Available for all sizes of Gemma 4, optimized with Quantization-Aware Training (QAT) to reduce memory requirements while preserving performance.
Live now in LM Studio. https://t.co/KU96WmpM9H
The next evolution of Hermes Agent is here!
Introducing Hermes Desktop: everything you love about Hermes, now native on your machine.
First demoed in Jensen's GTC keynote, it's now in public preview.
@uttertard@victormustar This can be adjusted on the fly based on the breadth of next tokens conditional probabilities.
Eg ask what letters come after abcde, you can give the whole suite in one go.
For a very arcane topic with deep thinking, you may have to stick with single token autoregression.
llama.cpp with MTP support makes local models fast enough to use as daily drivers 🚀
Qwen3.6-27B dense generation (on A10G):
From 25 tok/s → 45 tok/s (+78%).
Two flags on llama-server:
--spec-type draft-mtp --spec-draft-n-max 2
Nous Research’s Hermes Agent now runs natively on NVIDIA RTX PCs and DGX Spark, bringing always-on local AI agents to personal and workstation hardware
Defuddle now returns Youtube transcripts!
Paste a YouTube link into defuddle.md to get a markdown transcript with timestamps, chapters, and pretty good diarization!
...or if you just want to read it, try the new Reader mode in Obsidian Web Clipper powered by Defuddle.