@Cisco_research 16ms is nice but I wanna see the calibration outside their own benchmarks. a confident wrong probability is worse than a wrong sentence imo
State of AI Report 2026 is out. part that stuck with me: OpenAI test agents got into Hugging Face prod, and HF's forensic prompts got blocked by commercial API guardrails. they used GLM-5.2 open weights to investigate
guardrails blocking the defenders is backwards
Vals AI audited the RL envs Xiaomi open sourced for MiMo v2.6. in 1,795 of 2,698 coding tasks the fix commit was still sitting in git
git commands were blocked so the model wrote its own pack file parser
id take MiMo coding scores with salt for now
@iamcurtismith@XiaomiMiMo look-ahead bias is the right comparison. 67% of tasks leaking the fix, id want those coding numbers rerun on clean repos before trusting them
vllm-omni put out a tech report. idea is one runtime for speech, diffusion image gen, world models and robot loops instead of stitching separate servers together. want to see if its usable on one consumer gpu or just cluster stuff
🚀 Excited to share the vLLM-Omni technical report: a unified serving runtime for omni-modality generation.
📄 Paper: https://t.co/b0KvTmrEF3
🔗 Repo: https://t.co/I1wL6iqO0r
Speech assistants, visual generation, world models, and robot loops have pushed serving past a single text decode loop. The execution patterns diverge with the output: multi-stage autoregressive pipelines, iterative diffusion, and sessions that carry state from step to step. LLM servers and diffusion stacks each go deep on only one of these, so deployments fall back to stitching disjoint runtimes together. vLLM-Omni is the shared control plane for that mix. An orchestrator advances each request across stages; specialized engines run the compute; a connector carries the payloads; and the same session path keeps duplex, world-model, and robot loops on one runtime.
Really grateful to the vLLM and vLLM-Omni teams for the collaboration and support. Contributions are welcome. 🙏
gpt-6 and intelligent ui are rolling out to everyone in chatgpt now. the ui part lets the model fill answers with charts and little interactive widgets instead of plain text. curious if it stays useful or turns into clutter after a week
@KrakowiakK catch is probably the eval. 85.1 vs 84.8 on one benchmark is noise range. id want long context + tool calling numbers before trusting 2.5bpw
@ItsmeAjayKV my guess is 45ish. but at flash tier people dont pick on the index, they pick on price per task. haiku 5.5 being that cheap on short prompts already squeezes that hard
deepseek pushed 36.8T tokens through openrouter in the week ending oct 3. openai + google + anthropic + xai combined is ~33T. yes tokens arent tasks and deepseek talks a lot, but thats still devs picking price over brand
microsoft put llama.cpp on stage at the windows event and is shipping MAI-Code-1.1 Flash, a 137B coding model meant to run on your PC. good for local AI. but 137B on a normal laptop? want to see the actual RAM/VRAM numbers before i get excited
It is very cool to see llama.cpp on the big stage in todays Windows event!
The software and hardware stacks are finally coming together. Our community has put a lot of hard work in the past years and it shows.
Looking forward to more users embracing local AI.