OpenSand is an enterprise AI infrastructure platform that provides access to leading AI models through one API, making AI faster, easier, and more affordable.
Today, we are announcing a series of updates that give customers frontier-grade security at half the cost.
MAI-Cyber-1-Flash is our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases. When combined with MDASH, it delivers world-class performance at 50 percent of the cost of leading models.
We are bringing this capability to market through Project Perception, a complete agentic security offering grounded in real-world signals and security workflows. Teams of specialized agents work together to simulate attacks, detect and triage/investigate, and fix and remediate.
This is the benefit of building the harness, context/signals, and action space separate from one model family. By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome.
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community.
During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion.
That’s why we created the Open Secure AI Alliance.
Kimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intelligence Index
Moonshot has released the weights of their 2.6T parameter model under their 'Kimi K3 License' which we have labelled ‘Commercial Use Restricted’.
Restrictions compared with more permissive licenses such as MIT or Apache 2.0 include requiring model-as-a-service businesses with more than $20 million in revenue to enter into a separate agreement. Commercial products with more than 100 million monthly active users or $20 million in monthly revenue also need to display “Kimi K3” in the user interface.
@Kimi_Moonshot has also released their technical report with insights into model’s architecture and training approach.
Links below to the weights on Hugging Face, the Technical Report and further benchmarks 🔽
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Grok 4.5 just got even cheaper to run
As Elon promised, SpaceXAI has cut cached-input pricing from $0.50 to $0.30 per million tokens......a 40% reduction. For contexts above 200K tokens, cached input dropped from $1.00 to $0.60
This matters a lot for:
• Long-running coding agents
• Multi-turn conversations
• Repeated repository context
• Tool-heavy workflows
• Large system prompts
• Autonomous agent loops
Successful cache hits now cost nearly 7× less than regular input. At 100 million cached tokens, the cost falls from $50 to $30......pure savings that scale fast on real production workloads
Grok 4.5 was already fast and token-efficient. Now it’s even more cost-efficient to build with
chatgpt work is remarkable, and "work" undersells it.
from my phone i sent:
"use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready."
it...just worked.
Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available from @GoogleDeepMind@demishassabis . Onwards!
We have reset usage limits for all Codex and ChatGPT Work users.
Last night around 2am to 4am we suffered an almost global outage. All well and recovered, but you know what comes next.
We learn. We reset. Enjoy.
Opus 5 is a great model for coding, data analysis, design, biology, knowledge work.
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.
And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon.
https://t.co/Tc7z2FqJhQ
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56)
Key takeaways:
➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA
➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task
➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh)
➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra
➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50%
➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier
Other model details:
➤ Context window: 1 million tokens (equivalent to Opus 4.8)
➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens)
➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
ChatGPT Voice is now in the desktop app.
Control your computer and direct multiple agents running in ChatGPT Work or Codex, using just your voice.
It's powered by GPT-Live, so it can speak, listen, and coordinate work in the app at the same time.
Rolling out globally today on macOS and Windows to Plus, Pro, Business, Edu, and Enterprise plans.
Introducing the Qwen-Audio-3.0-TTS.
Our latest text-to-speech model, in two flavors:
• Flash: real-time interaction
• Plus: high-quality generation
What's new:
• Fine-grained inline tags-steer [whisper], [angry], [breaths] & [laughs]
• Free-style natural-language control-“read this slowly, like a bedtime story”
• 16 languages
• Clean output even from noisy reference audio
• One-pass long-form up to 3 min
now #1 on the Artificial Analysis TTS Leaderboard.
Blog: https://t.co/ACqTI2wGHI
API: https://t.co/Xh2Y19noVx
Voice mode now runs on Claude's more capable models and reaches the tools you've connected mid-conversation.
Talk through the hard problems out loud, in many more languages.
Grok Build just got major update, giving developers more control over AI tools, better account visibility, built-in diagnostics, stronger session reliability, and lower memory usage for Voice Mode on macOS.
Release Notes: v0.2.111
Features:
• Users can now disable image generation and video generation tools (and their slash commands) via config.toml or environment variables.
• /session-info now displays whether the session uses OAuth or an API key and where to manage the account.
• You can now run grok doctor fix commands directly from inside the TUI instead of only from the CLI.
Bug Fixes:
• !cmd commands now allow up to one hour before timing out.
• npm package now installs the native binary under $GROK_HOME/bin (honoring the same override as the Rust CLI).
• Startup warnings now point to /doctor for details and fixes.
• Dashboard hover and clicks no longer miss the gaps between items in wide mode.
• Shift/Alt+Enter now inserts a newline while editing a queued prompt.
• Queued prompt edits under combine mode no longer lose changes due to premature hold release.
• Forking a session that used compaction no longer causes later rewinds to fail with missing checkpoint errors.
• When a permission prompt appears while viewing scrollback, focus now correctly moves to the prompt so you can answer.
• Pressing Esc once now cancels the current agent turn (except in fullscreen vim scrollback mode).
• Grok now automatically stops a turn that keeps repeating the exact same tool call many times in a row.
• Configs using either spelling of the workspace teleport disable flag now load and save correctly.
• Background subagent completion messages no longer leak into unrelated sessions when multiple sessions are active.
• When the auto-permission classifier times out or fails, Grok now shows a normal permission prompt instead of silently denying.
• Managed MCP tools no longer time out prematurely on slow operations like Notion updates.
Performance:
• Voice dictation on macOS now uses less memory by running capture in a temporary helper process.
New for enterprises: OpenAI Presence helps companies deploy trusted voice and chat agents across customer and internal workflows.
AI agents can answer questions, use company systems, take approved actions, and escalate to people when needed—while improving over time.
OpenAI Presence is available to eligible enterprise customers through a limited general availability program.
https://t.co/aooz6Vzljd
The Claude Security plugin for Claude Code is now available in beta.
Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Claude inference you already run.
You can now ask Claude about the Anthropic Economic Index, our public dataset measuring how AI is used across the economy.
Ask which occupations use AI the most, or what kinds of tasks people are automating, and the answers draw directly from the Index data.