DeepSeek releases DeepSeek-V4.1-Flash!! 🔥
This model is wild... the benchmarks are showcasing it's GPT-5.6 Sol level, yet it's only 552B params?! This feels like it shouldn't be possible, what's the catch? Perhaps benchmarkmaxxed?
Let's look into the architecture and training of the model:
The key focus seems to be on more aggressive KV cache compression:
"We adopt a Causal Encoder-Decoder (CED) architecture, in which decoder global KV is projected from the final encoder hidden states. This design enables the model to activate 8B parameters per token during prefill and 16B during decode, which is particularly cost-effective for input-heavy agentic scenarios."
"DeepSeek-V4 can be viewed as an SWA-based local-processing backbone augmented with compressed global context."
The idea behind CED is that the decoder's KV cache is constructed from the encoder output, bypassing full decoder computation.
They also introduce Compressed Sparse Attention 2 (CSA2) that has three operating modes that differ in how they obtain main KV, indexer K, and Top-K indices. (frankly I don't understand this part very well 😭)
DeepSeek-V4.1-Flash uses a variant of mHC called single-pass mHC, and they also incorporate a 196B Engram module to decouple memorization from computation.
The model is natively multimodal: a vision embeddings generated from "DeepSeek-ViT" (a pretty standard ViT arch) are passed jointly with the text tokens into the model. This is trained first with SigLIP loss then with autoregressive loss for the combined vision encoder+LLM.
"we train DeepSeek-V4.1-Flash on a large-scale multimodal corpus comprising 45T tokens."
They use FP4 KV cache with quantization-aware training to further save storage.
Regarding post-training:
"In this release, we refrain from introducing novel post-training algorithms."
"at the current stage, the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training."
They utilize the model itself to construct its own training environments, based on data they are getting from model use internally.
"As we transitioned from DeepSeek-V3 to V4, the rapidly growing number and diversity of agentic training environments motivated us to build DeepSeek Elastic Compute (DSec), a production-grade sandbox platform for large-scale agentic training and evaluation."
"We therefore introduce a scalar effort level 𝑏 as an explicit conditioning signal during reinforcement-learning."
max --> b=100, high --> b=75, low --> b=50.
"As the last stage of post-training, the final full-vocabulary OPD task is trained on datasets from all domains using over 40 teacher models."
Damn, this is a dense report, I've barely touched the surface tbh, very interesting!!
model: https://t.co/CXQzLhSk8h
paper: https://t.co/ULZRPXiarS
Sharing lecture videos for the **How to AI (Almost) Anything/Multimodal AI** course I taught at MIT in spring 2026.
This course became quite a hit the last time I shared it in spring 2025. Spring 2026's updated version contains updated topics on multimodal agents, reasoning, self-evolving AI, and new modalities like touch and smell. Also includes slides (not videos) of guest lectures on multimodal AI for health, design, manufacturing, cities, & transportation.
Youtube playlist: https://t.co/awhFUqHGBt
Course website and materials: https://t.co/m7LZaCQ4fw
Today's AI can be applied to almost anything - from language to vision, audio, sensors, medical data, music, art, smell, and taste. This course covers the principles of AI (focusing on deep learning and foundation models), how we can apply AI to novel real-world data modalities, and multimodal AI that can process many modalities at once, such as connecting language and multimedia, music and art, sensing and actuation, and more.
We started to post videos for CMU 11-768, AI Agents!
All of the videos will be posted to this playlist, so please bookmark/follow it if you want to be notified of new ones! I'll also try to post them on this thread too.
https://t.co/HAJgKXve0x
Retailers running shopping agents on Claude have seen carts up to 35% larger and shoppers 60% more likely to complete a purchase.
See our blog to learn more about the architecture, latency & cost techniques, and eval practices:
https://t.co/lAgsVD5JSL
We're open-sourcing Claude Commerce Agents.
This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.
I’m excited to finally announce the newest edition my Stanford course 𝗧𝗵𝗲 𝗠𝗼𝗱𝗲𝗿𝗻 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗗𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿. It has been 9 months in the making.
Last November, with the release of Claude Opus 4.5, coding agents experienced a step function improvement in capability. We all felt it. The LLMs were more powerful, could reason for longer, solve harder tasks.
This year’s iteration of my course reflects the 2026 metamorphosis of software engineering.
My core belief is simple: AI-native developers of the LLM era are going to become the most important members of any software organization. I have designed my course to train this next generation of engineers.
𝗪𝗵𝗮𝘁’𝘀 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝘁𝗵𝗶𝘀 𝘁𝗶𝗺𝗲 𝗮𝗿𝗼𝘂𝗻𝗱
First, 85% of my Fall 2025 class material is being thrown out. The Fall 2026 syllabus reflects the core capabilities AI-native engineers must have: agent skills, advanced context engineering, MCP portals, agent-ready codebase principles, agentic code review, security, parallelizing background agents, software factories, and more.
Second, I am going to teach my students how to have software taste. Every student will be required to ship pull requests to production-grade, real-world codebases. The course is collaborating with the top open-source AI repos who will offer support and mentorship to students on how to meaningfully contribute to their projects.
This has never been done before in any university course so I am incredibly grateful to our OSS Partners: @browserbase, @HeyGen, @CopilotKit, @semgrep, @OpenHandsDev, @milvusio, @marimo_io, Pi, @crewAIInc, @warpdotdev, @vercel, @cmux, @arizeai, @UnslothAI, and @anyscalecompute.
𝗪𝗵𝗮𝘁’𝘀 𝘀𝘁𝗮𝘆𝗶𝗻𝗴 𝘁𝗵𝗲 𝘀𝗮𝗺𝗲
I’m fortunate to again have AI software engineering leaders and founders as guest speakers to share their learnings from building top coding agent products. Thank you to @leerob from @cursor_ai, @bcherny of @claudeai code, @EnoReyes of @FactoryAI, @silasalberti of @cognition, @0xine of @semgrep, Rajesh Bhatia of @Cloudflare , @amasad of @Replit, and @eladgil.
All resources will be available online. All classes will be available to the public.
9/22 on Stanford campus. See you in class.
https://t.co/wTokHyUMsz
Nice benchmark to measure agentic e-commerce capabilities.
They ran an agent for one simulated year of e-commerce operations and it ends up with 27.3% of the money a human makes.
MerchantBench is a 365-day order-level simulation grounded in 98,843 real product records with 26 tools for agent interaction. Agents handle product sourcing, listing and pricing control, cash-flow management, and feedback that arrives at wildly different delays.
Scoring runs on cumulative net assets, so incoherence compounds rather than averaging out. Eight LLMs across two agent frameworks, 48 runs of 365 simulated days each, and the best configuration still lands far under the human baseline.
Bounded tasks with immediate success criteria have been flattering agents that cannot hold a plan for a month.
Paper: https://t.co/wEWNbzXUjb
Track more trending AI papers in our academy: https://t.co/LRnpZN7L4c
Anthropic Head of Product:
“Fable 5 - is our best model for self-improving agentic systems. It can run for days on a single /goal.
add /loops, dynamic workflows, dreaming and you become unstoppable.”
in 11 minutes, the Anthropic team shows how to build long-running systems with Fable 5 from scratch.
Worth more than a $500 agent-building course.
Live from Anthropic’s latest stage in Japan. Unpublished.
非常值得刷的视频课--麻省理工《How to AI (Almost) Anything》
这门MIT公开课教你如何用AI(几乎)做任何事情——共12节课,覆盖大量专业领域 https://t.co/iqUHxw9NLI
大部分内容容易忘,建议把这份中英双语版讲义+图文笔记(社区整理)保存到NotebookLM / Claude / ChatGPT构建你自己的知识库,方便随时对话检索👇
https://t.co/vQ5FmN9Jde