Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog: https://t.co/ekQkO83jCv
@beffjezos Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks.
Initial training is complete and now we’re adding a massive amount of SpaceX company data in supplemental training. This will be something special.
Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark and cost substantially less than other leading models
AA-Briefcase tests models on long-horizon agentic knowledge work tasks. The test set is private to prevent contamination.
Grok 4.6 is neck and neck with Claude Fable 5, with overlapping confidence intervals. The model is also substantially cheaper than other leading models on the benchmark at a Cost per Task of $4.42 compared to Claude Fable 5's $22.30, Claude Opus 5's $17.79 and Kimi K3's $6.73.
Impressive release @SpaceXAI and @elonmusk.
@composio OpenCode salió más barato pero con menor success.
Yo lo he visto en mis propios flujos: el recovery es donde se va la plata y el tiempo.
Pi @pidotdev sigue siendo para mi la mejor opción en desarrollo.
El anuncio de DeepSeek Harness beta es interesante.
Modelo + harness = agente. (Eso ya lo sabíamos)
La pregunta real es si su "harness" va a manejar recovery y tool-calling mejor que OpenCode o Claude Code cuando el contexto se ensucia.
Yo todavía no lo tengo. Cuando salga lo voy a medir con las mismas tareas que uso en producción.
Actualizaron el modelo del chat y presumen un 68% menos de errores en sus "evals internas". Bonito número para la galería y el usuario casual. El motor que de verdad mueve Work y Codex sigue congelado en la versión de julio. Los que construimos agentes o pipelines en producción, este anuncio no aporta nada; es puro humo y maquillaje de interfaz :/ Menos PR de UI y más actualizaciones reales a la API.
We’re making better intelligence easier to access in ChatGPT for everyone:
- GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses.
- Free & Go users get unlimited text chats with GPT-5.6 Luna starting tomorrow.
Actualizaron el modelo del chat y presumen un 68% menos de errores en sus "evals internas". Bonito número para la galería y el usuario casual.
El motor que de verdad mueve Work y Codex sigue congelado en la versión de julio. Los que construimos agentes o pipelines en producción, este anuncio no aporta nada; es puro humo y maquillaje de interfaz :/ Menos PR de UI y más actualizaciones reales a la API.
@rootcausehq@barckcode@orca_build Tengo ese mismo setup en todos mis desarrollos usando orca + PI y es magnífico; solo agrego que uso grok como code review.
Tested Muse Spark 1.2, Kimi K3, and GLM 5.2 with our FlappyBench
3 models, same prompt with the /design command.
Reviewed gameplay features, UX/UI, and cost.
🔹 Kimi K3 → 9.5/10 · $0.0740
🔹 GLM 5.2 → 9/10 · $0.0480
🔹 Muse Spark 1.2 → 8/10 · $0.0187
Muse Spark 1.2 delivered a completely different game style despite being super cheap.
Kimi K3 and GLM 5.2 outputs are comparable and close in cost.
Qwen3.8-Max by @Alibaba_Qwen has reshaped the cost-performance Pareto frontier in Frontend Code Arena, with pricing of $2 per input MToken and $6 per output MToken.
Top models on the Pareto frontier:
- Claude-Opus-5
- Kimi-K3
- Qwen3.8-Max
- GLM-5.2
- DeepSeek-V4-Flash
Congrats to @Alibaba_Qwen on another major milestone!
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
🚀Hy3 is here.
295B MoE. Best in its size class. Rivals trillion-scale flagships.
Reliable and affordable for most agentic usecases.
Apache 2.0. Friendly for commercial use.
FREE API for 2 weeks → https://t.co/EyURKwTdgi
🤗 https://t.co/twqJpqb2SL
📖 https://t.co/4uEkIU1cW4
Learn to ship. Shipping is a skill distinct from coding. Shipping is designing, coding, QAing, story-telling, teaching, marketing, selling, pivoting, iterating…
It used to be that coding dominated in importance because of coding ability scarcity. AI will push you to go further.