Microsoft's MAI-Transcribe-2-Streaming ranks #1 of 38 on Artificial Analysis.
• 2.5% WER at 0.13s latency
• 60 languages
• $0.54 per hour (intro price)
Caveat: public preview, no SLA, no open weights.
Read more: https://t.co/L3Ng4lGZpv
Reka's Rho-1 reads video and outputs both video and robot actions.
• 19B parameters
• Trained on 320 H100s for 3 months
• Video capped at 672×384
• Distilled: 99 to 8 denoising steps
• Research preview, no public weights
Read more: https://t.co/OCKF8t996X
Microsoft's MAI-Transcribe-2-Streaming ranks #1 of 38 on Artificial Analysis streaming WER.
• 2.5% WER
• 0.13s to final transcript
• 60 languages
• $0.54/hr intro price
Caveat: public preview, no SLA, no open weights.
Read more: https://t.co/L3Ng4lGZpv
Prime Intellect launched Prime Inference for frontier open models.
• Nearly a trillion tokens/day before launch
• GLM-5.3 on GB200 NVL72
• Serverless and reserved capacity
Caveat: its own GLM-5.3 pricing is not yet published.
Read more: https://t.co/ToIHrBwhND
Train a robot policy on 707 GB of Cosmos3-DROID without downloading it.
• HTTP byte-range reads
• Peak disk: a few hundred MB
• 48 episodes, 12 epochs
Caveat: AV1 decode may need FFmpeg fallback.
Read more: https://t.co/HlRAhs7xRj
IBM Bob now runs self-hosted and air-gapped.
• Isolated models: NVIDIA Nemotron, Poolside Laguna
• Includes IDE, BobShell, parallel tool calling
• Premium: Java Modernization, IBM i, IBM Z
Caveat: no pricing; multi-model routing is roadmap only.
Read more: https://t.co/zowwP8aMJI
Laya classifies intent with zero output tokens.
• 421M parameters, Apache 2.0
• CLINC150 banking: 0.878 zero-shot
• 16.1 ms median on CPU
• Blocks 89.3% of out-of-domain queries
Caveat: check calibration on your own data.
Read more: https://t.co/x7lfAGlmhV
Meta, OpenAI and Uber now ship agents that message you first.
• Meta Muse: launched Sept 8
• OpenAI Dots: Sept 29
• Uber driver assistant: Sept 24
Caveat: decision models judge only the context they are given.
Read more: https://t.co/Aqy4GWA11Q
NVIDIA's IsaacTeleop turns XR hand tracking into robot actions.
• 26 hand joints, OpenXR order
• Gripper closes below 3 cm, opens above 5 cm
• Outputs an 8-D action vector
Caveat: the demo runs on synthetic data, not real headsets.
Read more: https://t.co/gkmOSTk7Dr
Kyutai's Voice of Reason solves spoken math without transcription.
• Spoken GSM8K: 27.3% to 70.3%
• With STITCH: 77.1%
• Runs on a single H100
Caveat: spoken TriviaQA fell from 40.6% to 34.0%.
Read more: https://t.co/WlRCT1uqUb
Cantina's open apex-flash-1 solves 40 of 60 held-out bug tasks.
• ~$2.38 per run
• Opus 5 High: 71.7% at ~$74.68
Caveat: company-reported, internal eval set.
Read more: https://t.co/eCqqYbBqBJ
CLM-8B scores agent actions up to 9x faster than Jev.
• T-Rex: 16.5 ms vs 149.8 ms
• DeepSWE verifier: 81.6% vs 71.1%
• Apache-2.0, 20M-parameter head
Caveat: held-out subsets, not leaderboard submissions.
Read more: https://t.co/dgzDBbonIp
JEPA-Anything: one world-model recipe across 7 fields.
• Pong error down 34.83%
• Burgers error down ~44.7%
• Apache-2.0 code and checkpoints
Caveat: some gains are small (UK Biobank 0.718 vs 0.711).
Read more: https://t.co/Y4ysnYZuOo
Google's Gemini 3.8 Flash TTS designs new voices from text prompts.
• 2,000+ voices, 100+ languages
• Hume voice design score: 71.4, #1
• SynthID on every clip
Caveat: API-only; voice replication unavailable in UK, EEA, India.
Read more: https://t.co/D4uCAxuD3k
Together Link runs open models inside Claude Code and Codex.
• Free MIT CLI
• Kimi K3, GLM 5.3
• Claims 50 to 80% savings vs all-Opus
Caveat: beta; macOS and Linux only.
Read more: https://t.co/bhvqxEfLW1
NVIDIA's Nemotron 3 Diarization tracks 8 speakers with 100M parameters.
• #1 of 12 on Voice Arena: 14.72% DER
• Next best system: 19.3%
• OpenMDW 1.1 license
Caveat: struggles with heavy noise and reverberation.
Read more: https://t.co/AZ6uzWieTj
Mistral Large 4 is a 1.05T MoE with only 49B active per token.
• 1M-token context
• Cybench: 93%
• DeepSWE v1.1: 61.7%
• API: $1.36 in / $4.18 out per 1M
Caveat: API preview only; weights and license still pending.
Read more: https://t.co/9hU0bzCYLx