41 days ago I set up my first AI agent on OpenClaw.
The biggest problem I ran into wasn't my code. It was the web itself.
So I built the fix. It's called Agent Web Protocol.
We have partnered with @AnthropicAI to launch the Claude connector for Fusion.
With the Fusion MCP, you can now tap into Claude’s desktop capabilities—speech, vision, web search, and cross-service workflows—right inside your design flow.
Combined with #AutodeskAssistant, this opens up more flexible, customizable agentic workflows for how you build with Fusion + Claude.
Learn more here: https://t.co/X2YCUlAnnz
Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights
Muse Spark is a new model from @Meta evaluated on Artificial Analysis. We were given early access by Meta to independently benchmark the model. It is the first frontier-class model from Meta since Llama 4 Maverick was released in April 2025, and notably the first @AIatMeta model that is not being released as open weights. The release follows Meta's reorganization of its AI efforts under Meta Superintelligence Labs, and signals that Meta is re-entering the frontier race after roughly a year of relative quiet.
For context, Llama 4 Maverick and Scout scored 18 and 13 respectively on the Artificial Analysis Intelligence Index as non-reasoning models at the time of their release, while Muse Spark scores 52. Muse Spark essentially closes the gap between to the frontier in a single release.
The model is not open source and is not yet accessible via an API but Meta has shared they expect this to come soon. Meta is also integrating Muse Spark into their first party products including their Meta AI chat product, Facebook, Instagram and Threads.
Key takeaways from our benchmarks:
➤ Muse Spark scores 52 on the Artificial Analysis Intelligence Index, placing it within the top 5 models we have benchmarked. It sits ahead of Claude Sonnet 4.6, GLM-5.1, MiniMax-M2.7, Grok 4.20 and behind Gemini 3.1 Pro Preview, GPT-5.4 and Claude Opus 4.6
➤ Muse Spark is notably token efficient for its intelligence level. It used 58M output tokens to run the Intelligence Index, comparable to Gemini 3.1 Pro Preview (57M) and notably lower than Claude Opus 4.6 (Adaptive Reasoning, max effort, 157M), GPT-5.4 (xhigh, 120M) and GLM-5 (110M)
➤ Muse Spark is the second-most capable vision model we have benchmarked. It scores 80.5% on MMMU-Pro, behind only Gemini 3.1 Pro Preview (82.4%)
➤ Muse Spark performs strongly on reasoning and instruction-following evaluations. It scores 39.9% on HLE, trailing only Gemini 3.1 Pro Preview (44.7%) and GPT-5.4 (xhigh, 41.6%). The model also achieved 5th highest in CritPT with a score of 11%, an eval that is focused on difficult physics research questions. This is substantially above above Gemini 3 Flash (9%) and Claude 4.6 Sonnet (3%)
➤ Agentic performance does not stand out. On GDPval-AA, our evalaution focused on real world work tasks, Muse Spark scores 1427, behind both Claude Sonnet 4.6 at 1648 and GPT-5.4 at 1676, but ahead of Gemini 3.1 Pro Preview at 1320. On On TerminalBench Hard, Muse Spark trails Claude Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro. Muse Spark joins others in achieving a high τ²-Bench Telecom score of 92%
Key model details:
➤ Modalities: Multimodal including text and vision input, text output
➤ License: Proprietary, Meta's first frontier model not released as open weights
➤ Availability: No public API at the time of publishing. Meta expects to provide API access soon. Meta has started integration into their first party AI offering Meta AI and inside Facebook, Instagram, and Threads
Remember when we were all reworking websites to actually function on smartphones?
Now we’re doing it again - but this time, for agents 🤖
At our https://t.co/pSyO4rLwrj.SF Hackathon, the Injestor team - Benjamin Shyong, Vishal Verma , Alex Shirazi - built a tool to convert websites (originally made for humans) work better with agents - and took 1st place in a seriously competitive field (200+ attendees, 70+ projects).
Their stack:
- Models on Nebius Token Factory @nebiustf (they liked Nemotron3-Super for speed and good performance)
- @tavilyai for pulling web content
- And (my favorite) a Karpathy-style agent loop to iteratively optimize sites for agent use
Agents building sites for agents - Pretty meta 😄
As part of the win, we invited them join us at the @nebiusai booth at #nvidiagtc. It was great fun!
Checkout Benjamin's post for more details and a cool video : https://t.co/EuRFXKMDe5 ; And https://t.co/30jjKEJ8at
Great geeking out with you all - excited to see where this goes next.
@thesuprememarty@nvidia@nebiusai hard to quantify, hacked together a working prototype in 6 hours other parts of the code were buggy may not have been directly related to nemotron3super on nebius
Super impressed with @nvidia nemotron 3 super 120b. Used it in a hackathon event last weekend. Enabled a live demo thru entire product flow in 90 seconds. Vs previous baseline of 5min with llama3 70b. When you only have 3 minutes to present every second is crucial. Super excited to utilize and play around with nemotron3 Ultra when it comes out!!
NVIDIA has announced their first reasoning models, a new family of open weights Llama Nemotron models: Nano (8B), Super (49B) and Ultra (249B)
From our early testing, @nvidia's Nemotron Super 49B scores 64% on GPQA Diamond in reasoning mode and 54% in non-reasoning mode - we are still running our full set of evals and will share the complete results shortly!
Key model details:
➤ All three models are distilled and post-trained versions of open weights Llama models: Nano (8B parameter model distilled from Llama 3.1 8B), Super (49B parameter model distilled from Llama 3.3 70B), Ultra (253B parameter model distilled from Llama 3.1 405B)
➤ System prompt enables toggling between reasoning and non-reasoning modes (i.e. system prompt passed is either “detailed thinking on” or “detailed thinking off”)
➤ Super and Nano are released under the NVIDIA Open Model License, Ultra is coming soon
DGX Station. Holy cow. 748gb ram and 20 petaflops!! You can run a true 1b parameter model and do some serious fine tuning on these machines. Quotes look like around ~$97k to ~$100k+ dependent on oem manufacturer
NVIDIA DGX Station is now available to order from select OEMs🔥
Powered by the GB300 Grace Blackwell Ultra Desktop Superchip, DGX Station brings data-center-class AI performance to the desk — enabling developers to build and run autonomous AI agents locally.
⚡ 748GB of coherent memory
⚡ Up to 20 petaflops of AI compute performance
⚡ Run large open models up to one trillion parameters
Together with NVIDIA NemoClaw — an open source stack that simplifies running OpenClaw always-on assistants, more safely, with a single command, we are delivering a full-stack platform for secure, long-running agentic AI.
Learn more: https://t.co/PzpgN2amRi
Appreciate the acknowledgement on the risk we took haha. it was definitely a bold call. Initial benchmarks showed strong throughput and the latency held up better than expected for a 120B model in a live demo setting.
@nebiusai sponsored the hackathon and gave us credits to build on their cloud infrastructure — Blackwell GPUs with optimized NVLink bandwidth and memory coherence. Kept inference stable under sustained load with no latency spikes during the full product flow.
Cloud inference on purpose-built hardware is getting practical fast. The model and infrastructure layers are converging — what matters next is the protocol layer connecting agents to the web.
@nvidia Nemotron 3 Super 120B scoring 85.6% on PinchBench as an @openclaw coding agent is a big deal for the open model ecosystem.
Even bigger when you run it on @nebiusai Cloud's G300 instances — 72x Blackwell GPUs + 36x Grace CPUs. That's the kind of infrastructure that makes a 120B parameter model actually practical for long-running autonomous agents.
The model layer is getting solved. The infrastructure layer is getting solved. What's still missing is the web layer — how agents discover and interact with external sites. That's what we're building with Agent Web Protocol.
https://t.co/6a49T4fDtO
🦞These innovations come together to create a model that is well suited for long-running autonomous agents.
On PinchBench—a benchmark for evaluating LLMs as @OpenClaw coding agents—Nemotron 3 Super scores 85.6% across the full test suite, making it the best open model in its class.
Honest answer — agent.json declares auth requirements (method, token endpoint, scopes, refresh contract) so the agent knows what's needed upfront. The discovery step is what it solves.
The actual auth negotiation still uses existing standards like OAuth2. Where we need to keep building. more granular permission models, dynamic scope negotiation, and how agents handle multi-step auth flows like MFA.
It's v1. The discovery layer is solid. The execution layer still has work to do. Open to ideas! This is exactly the kind of feedback that shapes the spec!!
41 days ago I set up my first AI agent on OpenClaw.
The biggest problem I ran into wasn't my code. It was the web itself.
So I built the fix. It's called Agent Web Protocol.
No crawling. No scraping. No guessing. One file.
The web has robots.txt to tell crawlers where they CANNOT go.
agent.json tells agents what they CAN do.