YODAS v3 is out: 1.1 million hours of 48kHz multi-channel speech in 147 languages under CC BY 3.0. The authors call it the largest open speech dataset to date and the first at scale with stereo audio, with word-level timestamps and English translations.
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://t.co/gJBf08vJH9
Cesti investori ve vesmirnem Blue Origin Jeffa Bezose.
Poprve v nasi vlastni investorske historii se nam podarilo byt primou soucasti takto vysoce preupsaneho neverejneho upisu, kde se upsal jeden dolar z patnacti nabizenych. Byl o to takovy zajem, ze byli investori doslova vybirani organizatorem emise po boku s Jeffem Bezosem, aby mohli byt soucasti tohoto prvniho upisu ve kterem pan Bezos pustil do Blue Origin dalsi spoluinvestory. Kriteriem byla kredibilita a vyse nabizene investice. Je to neverejny upis, Blue Origin jeste neni a nejakou dobu ani nebude na burze. O to tezsi bylo dostat se tam.
Ted ale muzeme rici, ze az poleti na obeznou drahu raketa BlueOrigin a zase ponese komunikacni satelity ASTSpace mobile, bude v obou nase ceska stopa. V obou jsme nyni nezanedbatelnymi investory v souladu s nasi strategii definovanou cca 6 let zpatky = odejit postupne ze SaaS sluzeb, a silne se zamerit na Space, Data a AI. A ted drzime robustni (i na svetove rozmery) podily v Nebius, BlueOrigin, SpaceX, ASTS a dalsich firmach z techto oblasti prumyslu a pridavame dalsi. SKY IS NO LIMIT.
Amerika leta do vesmiru, po zemi tam jezdi taxiky bez ridicu a Evropa? Doplnme si sami……. https://t.co/fxFpL2oQT5
TensorFold vs vLLM, Qwen3.8‑27B, one Spark, same 3 prompts, thinking off:
3× faster than vLLM on the same DGX Spark. Byte-identical output.
• TensorFold (DFlash2 drafts): 102.9 tok/s
• vLLM MTP=3 (NVFP4): 34.9 tok/s
• TensorFold serial, drafts off: 12.9 tok/s
Per prompt (TensorFold / vLLM):
sequence 126.9 / 37.0
code 75.7 / 32.7
json 106.2 / 35.0
TTFT 0.12 s. Drafted vs serial output: sha256 identical on all three. Loads in 32 s.
Not weight-matched: MLX 4‑bit g64 vs NVFP4 ModelOpt. Single pass. Long structured outputs draft well; prose will land lower.
Bonus: one Spark now beats my M4 Max on this model (oMLX MTP3: 65.8).
https://t.co/DBejw9Catm — @ashxhart
Today we’re announcing OrcaSAQ-2 27B
High-fidelity mixed-precision Qwen3.8 for long-horizon agents.
55.59 → 12.06 GB — 78.3% smaller / 4.61×
3.21 bpw · 93.2% Top-1 agreement
70.0 SWE-bench Verified
58.4 Terminal-Bench 2.1
262K context
A 27B model for coding, terminal, browser, security and multi-tool agents — in a footprint you can actually deploy.
SOTA agentic capability density among similarly sized models we evaluated.
https://t.co/FqslE1kYnw
🏢 Enterprise use
This release is trained at xhigh effort. For enterprise use, we are ready to fine-tune medium and low effort as well. Contact us at enterprise [at] https://t.co/ZC7xXiNWgR
🚨Happy to share ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
Side-by-side comparison video in the blogpost below.
📄 Blogpost: https://t.co/fj3oAxMpx8
🤗 Hugging Face: https://t.co/qSmYi4Xrlj
🔗 GGUF: https://t.co/zEm6m08Muc
@0xBakeer@FrostForger I had similar problem with pi harness, switched to qwen code, but there are timeouts after 15 mins. Tried qwen code w qwen3.8 27b w gpu and similar issues there. Need to look under the hood.
4B parameters. ZERO distillation. 61.5% on SWE-bench Verified 🤯
Meet FrogNano 🐸: Qwen3.5-4B post-trained purely with RL on synthetic tasks from TaskPilot.
Just 5 iterations × 300 tasks.
Who said coding agents have to be huge? 🐸
This is for you single DGX Spark owners 💫
You can now run Qwen3.8 Flash NVFP4 on one DGX Spark with great performance!
- Up to 1M context, 1,431,164 KV cache
- Full image & video support
- 37 decode tok/s on prose single stream
- Up to 86 tok/s on prose 4 concurrent streams
- 1500-2000 tok/s prefill on any size.
- Stress tested with 400k prefill - passed.
IMO this is the BEST model to run now on a single dgx spark. It's better than Qwen3.8-27B, and also better than the 1x dgx spark version of DeepSeek v4 Flash.
Get it here:
https://t.co/LG0I9PJKxI
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
One week with an ASUS Spark and having fun experimenting with local AI.
Tested Qwen 3.6, Qwen 3.8, DeepSeek, plus Hermes Agent, Pi and OMP.
Still exploring the ecosystem. What should I try next? 👀 #LocalAI#LLM
If you're looking for a weekend project, how about training your own text-to-speech model from scratch on your own GPU, and then running it on any device's CPU?
We just open-sourced the entire Pocket TTS training stack: data pipeline, recipes, and evals.
It learns pretty damn fast:
~15k steps: babbling starts turning into words
~50k steps: it reads anything you type (WER under 1%)
~200k steps: the voice stops sounding synthetic
On a beefy consumer GPU, that's a week of training. On eight H100s: 10-20 hours. A TTS training run will cost you less than $200 if you rent your hardware, and an order of magnitude less if you just pay for power.
Some things we'd love to see people try:
- Train it in your own language (a few hundred hours of speech gets you surprisingly far).
- Add new features to Pocket TTS (Emotion tags? Make it sing?).
- Beat us at our own game: make it faster and smaller.
Show us what you build! We'll highlight the best models and new languages for the whole community to enjoy. Pocket TTS has already found many use cases, from reading for people with visual impairments to making NPCs in video games talk, and we're sure there's much more to do with it!
Here's an example of a Czech Pocket TTS. Try just asking your favorite agent to find data and apply the method, and you can have your own.
Get started: https://t.co/3EH3sbKNRU
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.
What's new: 🥳
- Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4.
- Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks.
- Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
- 262K native context, extensible to 1M with YaRN.
We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀
We can't wait to see what you build with Qwen3.8-Flash!👀👇
- Blog: https://t.co/M5hYypFLgJ
- Technical Report: https://t.co/IF0gObIkQO
- Hugging Face: https://t.co/6ow8QVAABt
- ModelScope: https://t.co/tDOn2jNuFG
There is a Japanese practice of owning one very good version of something rather than several adequate ones.
One knife, kept sharp. One pen that writes exactly as you want. One bag chosen carefully and used for years.
This is not minimalism as aesthetic. This is minimalism as relationship.
You learn the object. You know its weight. You know how it responds to your hand.
The object becomes, over time, an extension of you.
That cannot happen with things you do not stay with long enough to know.
A BIG moment for all DGX Spark users ⚡️
You can now run DeepSeek v4 Flash 0731 without needing a second unit, with quality high enough for reliable code generation, high context, and great speed!
Optimized for single stream session:
- EXL3 quantization
- 384k context (conservative) / ~440k kv cache (!)
- 47 tok/s single stream (structured)
- 1024 tok/s prefill
- 370k token needle test passed (super stable!)
Thanks @0xSero for this excellent EXL3 quant! Tuned for SparkInfer + DSpark speculative decoding, this is about the same quality of a Q4_K_M / Q5 GGUF!
All this is possible thanks to native NVFP4 KV cache & fixing kernel bugs in the upstream prefill path to make it work at all.
Get it here:
https://t.co/531EaR6FWJ