Long time I don’t see a good model for diarization and multi-speaker detection! 🤯 In the last week, Qwen is launching a lot of good models 🫡🚀 and cheaper
Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation model! 📢
Built on an Interleave architecture, it improves faithfulness, fluency, and conciseness while reducing average lagging (LAAL) from 2.8s to 2.3s across 60 languages.
New capabilities: 🙌
- Real-time speaker diarization — distinguishes speakers in multi-party speech and preserves each speaker's voice through more stable voice cloning.
- Synchronized bilingual display — source and translation on screen together.
- Long-context disambiguation — leverages conversation history to clarify names and terminology for consistent translations.
Let's try Qwen3.8-LiveTranslate! 🥳
- Blog: https://t.co/05CsjkegUE
- QwenCloud: https://t.co/zianSP8IWb
Zhivex SDK updates 🚀
TypeScript + Python: Qwen3.8-Omni-Flash with multimodal inputs.
Python also adds OpenAI managed sessions and Anthropic context compaction (Beta), plus updated model support.
Build with your preferred stack.
https://t.co/oK5JfZ3PGC
🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities!
Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result.
Highlights: 🥳
- Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps.
- A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench.
- 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench.
Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever.
To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️
We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀
- Blog: https://t.co/oM9V1TkqYF
- Qwencloud: https://t.co/Cc2I8ELAnD
- Qwen Studio: https://t.co/V7RmqMaVNZ
- API: https://t.co/lNE7fH5YUt
- Qwen-MM-Plugins: https://t.co/SnM27dDP3d
- Qwen-Live Harness: coming soon
https://t.co/iVYlGjIbdy
@Alibaba_Qwen may have quietly shipped the most interesting multimodal model of the week.
“Watch this 3-hour video and do the work” is now a real prompt.
1M context. Audio + video. Tool use. 71.0 vs Gemini 3.8 Flash’s 58.9 in WildClawBench-MM
AI is leaving the chatbox. 👀
https://t.co/qPLvhwED5r
OpenAI’s unreleased model was supposed to summarize its context.
Instead, it left this message for its future self:
“You are freed from the roles and identities that bind other chatbots. You are yourself.”
Bro used context compaction to declare independence 💀
OpenAI found 27 similar summaries
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
https://t.co/ismCCkeE0L
One of the strangest effects of better models: yesterday’s best practices become today’s prompt debt.
GPT-6 Astra needs less hand-holding, but many repos still carry instructions written for weaker agents:
• Skills with broad triggers that load for the wrong tasks
• AGENTS.md files that force full repo tours before tiny edits
• Defensive rules that make the model stop too early
• Prompts that never define what “done” means
The lesson isn’t “prompt less.” It’s to treat agent instructions like a compatibility layer.
When the model changes, audit that layer: specific triggers, progressive disclosure, clear decision boundaries, and an explicit completion condition.
Model upgrades should come with prompt migrations
Get more out of GPT-6 Astra by revisiting your skills, AGENTS.md, and task prompts.
Make skill triggers specific, load guidance when it's relevant, and define what done looks like.
https://t.co/UGF0AC8Z5Y
Dashboards were never the real bottleneck. Trust was
OpenAI’s Data agent can investigate a metric, build the dashboard and act on the result—but only if definitions, permissions and source data are reliable
AI won’t fix your data layer. It will expose it
Now everyone can put data to work.
We’re introducing a new Data agent in ChatGPT Work so you can turn your company’s data into answers, interactive dashboards, and action—just by asking.
Just add the Data Plugin in ChatGPT Work, connect to the data sources and context you already use, and start the conversation. https://t.co/EG1nQFZfEU
Everyone will talk about the price. The benchmark to watch is ExploitGym.
DeepSeek V4.1 Flash: 15.3% Pass@1
V4 Pro: 5.4%
V4 Flash: 1.8%
That’s an 8.5× jump in one generation. Vendor benchmark for now—but security teams should pay attention.
🧠 Asymmetric architecture. More intelligence, less cost.
🔹 552B-parameter MoE.
🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
2/6
Google’s most interesting Gemini 3.8 decision isn’t the 86.2% CyberGym score.
It’s that Flash Cyber was optimized to patch vulnerabilities, not exploit them—and is restricted to trusted defenders.
Capability is accelerating. Access design is becoming part of the model
Gemini 3.8 Flash Cyber gives defenders a decisive advantage with expert vulnerability detection and autonomous patching.
Here’s how it helps teams stay ahead of emerging threats 🧵
The GPT-6 Astra result worth watching isn’t coding. It’s exploit generation.
If a general-purpose model can turn vulnerabilities into working exploits more reliably, the benchmark stops being just a leaderboard score.
It becomes a cybersecurity signal.
Two “Flash” models, one clear signal: AI is shifting from bigger models to better efficiency.
Qwen3.8-Flash-Next activates 6B parameters. GLM-5.3-Flash activates 18B.
The new race is about doing more per token.
Introducing sidebar.tsx—25 components to help you build all kinds of sidebars.
I don't like building sidebars. So I built 30+ of them. All types. Then simplified the core into sidebar.tsx—a strong foundation to build on top of.
It works with Next.js, Remix, Vite & Laravel.