Vals AI is providing benchmarks and associated data for evaluating AI systems to the National Institute of Standards and Technology (@NIST) for use by NIST’s Center for AI Standards and Innovation (CAISI) to conduct pre- and post-deployment evaluations of AI models.
NIST and CAISI’s use of any product, data, or service in
this Research Project does not imply a recommendation or endorsement of such product, data, or service, nor does it imply that any product, data, or service is necessarily the best available for the purpose used.
New from Vals: Center for a New American Security (@CNASdc), Virginia Tech’s National Security Institute (@vtnsi), and Vals AI are partnering to explore new methods to rigorously evaluate military AI.
These systems are being deployed across national security missions faster than anyone can measure them. Our partnership pairs deep operational expertise with independent, open evaluation, giving policymakers and the public an evidence-based read on where military AI can be trusted and where it falls short.
We look forward to sharing more about our work.
Excited to share that I've joined @ValsAI to lead Vals Public Sector: independent AI evaluation for government.
Government is adopting AI fast, but in high-stakes domains like national security, "vibe" checks won't cut it.
Vals builds rigorous, independent benchmarks trusted and cited by the leading AI labs. Now we're bringing that rigor to government, from public benefits to national security.
Learn more here: https://t.co/9Z9Hy8VKl2
🇺🇸
Today, we're launching Vals Public Sector: independent AI evaluation for government.
Recent events have made one thing clear: the government has a stake in understanding how frontier AI models perform on day-to-day work, and the risks they carry.
This work is central to who we are. We've spent years building industry-leading benchmarks alongside the top AI labs and enterprise domain experts. Now, we’re bringing that same rigor to support government where AI matters most, from public benefits to national security.
We are excited for the work ahead!
Today, we're launching Vals Public Sector: independent AI evaluation for government.
Recent events have made one thing clear: the government has a stake in understanding how frontier AI models perform on day-to-day work, and the risks they carry.
This work is central to who we are. We've spent years building industry-leading benchmarks alongside the top AI labs and enterprise domain experts. Now, we’re bringing that same rigor to support government where AI matters most, from public benefits to national security.
We are excited for the work ahead!
Today, we're launching Vals Public Sector: independent AI evaluation for government.
Recent events have made one thing clear: the government has a stake in understanding how frontier AI models perform on day-to-day work, and the risks they carry.
This work is central to who we are. We've spent years building industry-leading benchmarks alongside the top AI labs and enterprise domain experts. Now, we’re bringing that same rigor to support government where AI matters most, from public benefits to national security.
We are excited for the work ahead!
In conversations with our head of public sector:
- There may be other credible classified concerns that logistically could not be communicated to Ant in a timely manner. This means it’s possible Ant doesn’t yet have the full story intelligence wise. The blog post states USG didn’t go into detail on their NatSec concerns.
- Dario is a rational actor, I would expect Ant to resolve any known security risks that are newly possible with Fable 5. It would be odd to so severely raise security concerns that applies across frontier models but were only realized with the release of this one.
- USG is citing its export control authority. Meaning that USG is primarily concerned with who has access over the cyber capabilities itself. It’s unclear why foreign nationals in the US can’t have access - GPUs have had export controls but could still be used by foreign nationals in the US. Controls like this typically only apply to sensitive USG info, which is a high bar.
- The practical effect is that this is a ban. This makes it seem like the issue is either related to 1) insufficient access controls or 2) distillation (WH’s April memo specifically called out evidence of this).
- Tariffs:China :: This export control:Anthropic. An anti-regulation admin now sees the national security risk and wants to rapidly increase leverage. As with the DoW case, courts will set the precedents on what power USG has. Ant unlike China is non adversarial but even still this action forces all parties to the negotiating table. May give the USG authority to enforce stricter access controls.
- This is an incredibly high variance period. In the weak case, Fable will be restored Monday with this as a slap on the wrist. In the strong case, we have just seen the most capable model that will ever be generally accessible. At minimum I expect this starts an industry-wide change in governance that more explicitly covers AI software/models, not just hardware.
How does @OpenAI's new GPT-5 perform on US Military domains? We ran it against our JointStaffBench(v1).
More about JointStaffBench: https://t.co/1TsaDVyWE7
Really excited to announce my next thing after leading generative AI in the Pentagon...aligning AI to government! 😀
More to come soon 🇺🇸 Please go give our page a follow!
Today we’re thrilled to announce that GovBench is incorporating as a 501(c)(6) nonprofit association.
Our mission is clear: close the gap between frontier AI research and day‑to‑day government work so agencies can deploy AI that is safe, accurate, and transparent.
Billions of taxpayer dollars are earmarked for AI this year alone. GovBench is setting the standards to ensure those investments deliver real public value.
Get Involved:
- Donate or partner: [email protected]
- Follow our journey: @GovBench on LinkedIn & X
Together, we can make AI work for government—and for everyone it serves.
https://t.co/fq84qxbrYh