Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.
Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come.
But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility.
This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems.
Together, we are building the foundation of the AI economy.
Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://t.co/ugYWQ1MyRi
first came jev. then laya. then laya-mlx. and now kev.
all system one decision models. all within a span of 3 days. never seen a race this fierce before.
let us breathe, folks. let us breathe.
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. https://t.co/SHVzjpgnfx
My honest experiments after 10 days on DGX spark
- I have cancelled my frontier subscriptions temporarily. I was spending 300+ . 200 of claude code and 100 dollar of Chatgpt. Now I have moved to 20 dollar of chatgpt and 20$ cursor subscription.
- My main model in hermes agent is a local model. Currently testing Ornith 1.5 as my main model as I like the speed + decent intelligence.
- For complex tasks I still feel the need for frontier models. I just trust it more. Currently experimenting with Grok 4.6 and I quite like it. Alternatively I use luna max for tasks which I need to run from my mobile app.
- I use local models to update my G-brain. My second brain. Taking privacy a bit more seriously :)
- I love Qwen 3.8 27b but its too slow for spark at the moment. Will be adding RTX 5090 in the future or get another spark. ( I might be addicted haha)
- Need to experiment more with Deepseek flash 0731 but not sure if its too compromised for one spark.
I wonder if there are benchmarks for quantized models as somewhere I still feel that we get excited looking at the intelligence score on Artificial analysis but then we are using quantized models & I wonder if the intelligence is still the same.
I really like the local ai community, and the experiments they are running and at the same time, I use frontier models for heavy works. Trying to take the best of both the worlds :)
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Fun fact: OpenAI handles 800 million users on ChatGPT with just one PostgreSQL primary and 50 read replicas 🤯
Today, OpenAI published an engineering blog explaining how they scaled their Postgres setup to support a massive 800 million users using a single primary and 50 multi-region replicas.
They dive into details around their scaling approach, the PgBouncer proxy, cache locking, and cascading read replicas. It is genuinely neat and impressive.
Some time back, I published a video on my YouTube channel where I dissected the blog and broke down the nuances.
Give it a watch - it is short and fun.
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
BLR is buzzing with events. Buildathons, hackathons, meetups, confs - something new every day.
2015 to 2019 vibes are back. So good to see the rush, energy, and enthu again. Nothing like it.
Before you ping a peer, senior, or open a thread asking for help, put a reasonable amount of effort in first to solve the problem.
Not asking you to bang your head against something, but make a reasonable attempt - read the error message fully, check the docs, search the codebase, try one or two hypotheses, or better - ask Claude (at least).
This matters more than it sounds.
When you come to someone having already done that groundwork, the conversation is completely different. You are not asking them to do the thinking for you. You are asking them to help you get unstuck. That is a much more productive use of everyone's time, including yours.
It also builds something that will help you throughout your life (not just your career) - the muscle of debugging independently. Tbh, this is something I feel was, is, and will always be super crucial. Each time you try to solve it on your own before escalating, you get a little better at it. Over time, you stop needing to 'escalate' as often.
Again, there is no shame in asking for help. Asking is good. But the quality of the ask matters. "I tried X, then Y, I think the issue is Z, but I am not sure" is a completely different conversation than "it is broken, can you look?"
Hope this helps.
@LinghuaJ Interesting.. Chain of thought is a reduce (in addition to attention ofc), so I guess this can be seen as a bit more of a directed context compaction mechanism, inheriting structure from the preexisting idea of a wiki.
how do you take a deep agent to production? it's all outlined in this guide: you'll have to think about memory, execution environment, guardrails, durability, and more!
https://t.co/3t2x9O5vxU
- Drafted a blog post
- Used an LLM to meticulously improve the argument over 4 hours.
- Wow, feeling great, it’s so convincing!
- Fun idea let’s ask it to argue the opposite.
- LLM demolishes the entire argument and convinces me that the opposite is in fact true.
- lol
The LLMs may elicit an opinion when asked but are extremely competent in arguing almost any direction. This is actually super useful as a tool for forming your own opinions, just make sure to ask different directions and be careful with the sycophancy.
One common issue with personalization in all LLMs is how distracting memory seems to be for the models. A single question from 2 months ago about some topic can keep coming up as some kind of a deep interest of mine with undue mentions in perpetuity. Some kind of trying too hard.
If you have more than 5 years of experience, make sure you always keep a few small projects ready on the back burner. These are ideas you can pick up and run with at any time.
First of all, this is a great way to show that you care about the product and can spot gaps before others do. It also puts you in a good light with your leaders and increases your influence. It helps you stand out as someone who thinks ahead and stays prepared.
While these are already great wins for you, doing this also helps you
- be ready with good first projects for new joiners
- have meaningful work ready during slow cycles
- stay in touch with real customer pain points
- find small wins that build trust with stakeholders
More importantly, it lets you influence the product and your team's roadmap. This is a great way to earn extra points with leadership and climb the ladder faster.
Keep your back burner hot and your career hotter :) Hope this helps.