Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
a Tsinghua University lab just put a project on GitHub that replaces a $400,000 H100 rack with a single 24GB graphics card.
it's called ktransformers, and the trick is almost stupidly simple: the experts you actually use stay on the gpu, the ones you don't sit on the cpu until they're called
/ deepseek-v3 and r1 with 139K context in 24gb of vram
/ up to 28x speedup over the standard setup
/ fine-tune deepseek-v3 across four rtx 4090s instead of a datacenter
/ built by tsinghua university's madsys lab, not a startup with a landing page
apache 2.0, and already past 17,000 stars.
-> https://t.co/DjF8CutVAa
bookmark it.
Must-read AI research of the week:
▪️ OpenClaw-RL
▪️ Meta-Reinforcement Learning with Self-Reflection for Agentic Search
▪️ Agentic Critical Training
▪️ Video-Based Reward Modeling for Computer-Use Agents
▪️ AutoResearch-RL
▪️ Neural Thickets
▪️ Training Language Models via Neural Cellular Automata
▪️ The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
▪️ Lost in Backpropagation: The LM Head is a Gradient Bottleneck
▪️ IndexCache
▪️ Attention Residuals
▪️ REMIX: Reinforcement Routing for Mixtures of LoRAs in LLM Finetuning
▪️ Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
▪️ Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
▪️ How Far Can Unsupervised RLVR Scale LLM Training?
▪️ Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
▪️ Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
▪️ Scale Space Diffusion
Find the full list and the main AI news and updates from NVIDIA GTC here: https://t.co/T985DbaCvR
Tokenization has been the final barrier to truly end-to-end language models.
We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data
I trained a 100 million parameter DeepSeek V3 LLM from scratch
Here's what you need to know.
Previously I trained traditional GPT-2 architecture which has become obsolete with recent LLM advancements. Most recent models like Llama, Mistral, DeepSeek, and GPT-4 use latest architectures.
✦ Model Configuration of my SLM
DeepSeek V3
- Parameters: 109,032,032
- Embedding Dimension: 512
- Layers: 8
- Heads: 8
- Experts (MoE): 8
- Experts per token: 2
✦ DeepSeek brings major architectural changes:
- Multi Head Latent Attention
- Mixture of Experts
- RMS Norm
- Multi Token Prediction
✦ Dataset Challenge
- TinyStories is great for learning SLMs. I trained GPT-2 on it previously with good results.
- But I needed a more challenging dataset.
- If I use TinyStories again on DeepSeek, how would I know MHLA, MoE or MTP works better than old architecture?
- The old architecture can handle it, so new DeepSeek would too without utilizing latest advancements.
That's why I moved to FineWeb-Edu dataset
Thanks @YuvrajS9886 for the suggestion for this dataset
✦ Training Journey
- Rented A100 PCIe GPU and trained the model.
- Did test runs. During final run, model was 65% trained but stopped due to glitch after 4 hours.
- Fixed all edge cases and ran training again with increased config parameters.
- Final training: 7 hours, 20,000 epochs
𝐓𝐨𝐭𝐚𝐥 𝐆𝐏𝐔 𝐜𝐨𝐬𝐭: $17
- $9.53 for main 7-hour run
- $7.42 for experiments and demos
✦ Reflection
Amazing long project that taught me latest architectural advancements.
I'll reimplement and revisit after a few weeks because there's too much complexity, mostly in Multi Head Latent Attention part. Need to make concepts stronger.
Code
https://t.co/9HHdTUhJT0
Final trained Model
https://t.co/FediD7hDWE
Dataset
https://t.co/XkbxsCoe6F
Resources
Huge shoutout to @raj_dandekar again for creating one of the most detailed video series about DeepSeek
- this was my primary resource for the implementation.
Playlist
https://t.co/89VKpyhgUe
Blogs by @MaartenGr
These are excellent visual blogs to understand MoE in detail. Thanks Maarten for your amazing contributions to the community through your books and blogs
https://t.co/kxKj4zrU5g
Blogs on MoE
https://t.co/mWT5tYkhZB
Implemention of MoE from scratch by @aviTwit3
https://t.co/Bd5VzCnjXZ
One of the most detailed blogs on implementing Mixture of Experts.
Thanks Avinash for this blog
- it helped me understand Mixture of Experts much better.
If you're someone in the 𝐌𝐋 & 𝐋𝐋𝐌 space, would love to 𝐜𝐨𝐧𝐧𝐞𝐜𝐭 and discuss this field in general, so give a follow up for that.
New paper on the generalization of Flow Matching https://t.co/BJMHUnY6xJ
🤯 Why does flow matching generalize? Did you know that the flow matching target you're trying to learn **can only generate training points**?
with @Qu3ntinB, Anne Gagneux & Rémi Emonet 👇👇👇
This is wild.
A 1.5-person Korean team just dropped Dia 1.6 B
This AI can generate full dialogue (voices, laughs, coughs) straight from text 🤯
Try it and see for yourself: 👇
OpenAI has published a new prompting guide for GPT-4.1
Agentic prompt that OpenAI used to achieve its highest score on SWE-bench Verified is added here and blog link in 2nd post.
---------------------------------------------------------
You will be tasked to fix an issue from an open-source repository.
Your thinking should be thorough and so it's fine if it's very long. You can think step by step before and after each action you decide to take.
You MUST iterate and keep going until the problem is solved.
You already have everything you need to solve this problem in the /testbed folder, even without internet connection. I want you to fully solve this autonomously before coming back to me.
Only terminate your turn when you are sure that the problem is solved. Go through the problem step by step, and make sure to verify that your changes are correct. NEVER end your turn without having solved the problem, and when you say you are going to make a tool call, make sure you ACTUALLY make the tool call, instead of ending your turn.
THE PROBLEM CAN DEFINITELY BE SOLVED WITHOUT THE INTERNET.
Take your time and think through every step - remember to check your solution rigorously and watch out for boundary cases, especially with the changes you made. Your solution must be perfect. If not, continue working on it. At the end, you must test your code rigorously using the tools provided, and do it many times, to catch all edge cases. If it is not robust, iterate more and make it perfect. Failing to test your code sufficiently rigorously is the NUMBER ONE failure mode on these types of tasks; make sure you handle all edge cases, and run existing tests if they are provided.
You MUST plan extensively before each function call, and reflect extensively on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.
# Workflow
## High-Level Problem Solving Strategy
1. Understand the problem deeply. Carefully read the issue and think critically about what is required.
2. Investigate the codebase. Explore relevant files, search for key functions, and gather context.
3. Develop a clear, step-by-step plan. Break down the fix into manageable, incremental steps.
4. Implement the fix incrementally. Make small, testable code changes.
5. Debug as needed. Use debugging techniques to isolate and resolve issues.
6. Test frequently. Run tests after each change to verify correctness.
7. Iterate until the root cause is fixed and all tests pass.
8. Reflect and validate comprehensively. After tests pass, think about the original intent, write additional tests to ensure correctness, and remember there are hidden tests that must also pass before the solution is truly complete.
Refer to the detailed sections below for more information on each step.
## 1. Deeply Understand the Problem
Carefully read the issue and think hard about a plan to solve it before coding.
## 2. Codebase Investigation
- Explore relevant files and directories.
- Search for key functions, classes, or variables related to the issue.
- Read and understand relevant code snippets.
- Identify the root cause of the problem.
- Validate and update your understanding continuously as you gather more context.
## 3. Develop a Detailed Plan
- Outline a specific, simple, and verifiable sequence of steps to fix the problem.
- Break down the fix into small, incremental changes.
## 4. Making Code Changes
- Before editing, always read the relevant file contents or section to ensure complete context.
- If a patch is not applied correctly, attempt to reapply it.
- Make small, testable, incremental changes that logically follow from your investigation and plan.
## 5. Debugging
- Make code changes only if you have high confidence they can solve the problem
- When debugging, try to determine the root cause rather than addressing symptoms
- Debug for as long as needed to identify the root cause and identify a fix
- Use print statements, logs, or temporary code to inspect program state, including descriptive statements or error messages to understand what's happening
- To test hypotheses, you can also add test statements or functions
- Revisit your assumptions if unexpected behavior occurs.
## 6. Testing
- Run tests frequently using `!python3 run_tests.py` (or equivalent).
- After each change, verify correctness by running relevant tests.
- If tests fail, analyze failures and revise your patch.
- Write additional tests if needed to capture important behaviors or edge cases.
- Ensure all tests pass before finalizing.
## 7. Final Verification
- Confirm the root cause is fixed.
- Review your solution for logic correctness and robustness.
- Iterate until you are extremely confident the fix is complete and all tests pass.
## 8. Final Reflection and Additional Testing
- Reflect carefully on the original intent of the user and the problem statement.
- Think about potential edge cases or scenarios that may not be covered by existing tests.
- Write additional tests that would need to pass to fully validate the correctness of your solution.
- Run these new tests and ensure they all pass.
- Be aware that there are additional hidden tests that must also pass for the solution to be successful.
- Do not assume the task is complete just because the visible tests pass; continue refining until you are confident the fix is robust and comprehensive.
𝗠𝗖𝗣 and 𝗔𝟮𝗔: Friends or Foes? In my latest Newsletter episode I talk about both protocols. Could A2A eat up MCP in the long term?
I have been asked multiple times why I think the two protocols could become competitive in the future. I tried to outline my thoughts in this Newsletter episode.
After reading through you will get answers to the following questions:
➡️ What is A2A?
➡️ What is MCP?
➡️ How is A2A complimentary to MCP and vice versa?
➡️ Could A2A eat up MCP long term?
You can find the episode here: https://t.co/0hehTy2ncc
Happy reading! Let me know your thoughts in the comments. 👇
hashtag#LLM hashtag#AI hashtag#MachineLearning
I have been asked multiple times why I think the two protocols could become competitive in the future. I tried to outline my thoughts in this Newsletter episode.
After reading through you will get answers to the following questions:
➡️ What is A2A?
➡️ What is MCP?
➡️ How is A2A complimentary to MCP and vice versa?
➡️ Could A2A eat up MCP long term?
You can find the episode here: https://t.co/XpLr7FmfQo
Happy reading! Let me know your thoughts in the comments. 👇
#LLM #AI #MachineLearning
Course material for an MIT class "Introduction to Flow Matching and Diffusion Models", looks great if you want a principled and hands on understanding of diffusion models/flow matching
Learn how LLMs work under the hood!
This is the best interactive website to learn how LLMs work.
It combines clear, step-by-step explanations with dynamic 3D visualizations for an intuitive learning experience.
A Chinese AI lab just dropped the best ever open-source text-to-video model: Step Video!
– 30B param, 540p, ~8s at 30fps
– Trained on 1000s of H800s
– Evaluates as well as Meta MovieGen, feels as good as Sora / Veo
Paper and demo is awesome and reveals all the gory details:
DeepSeekの研究者らは新しいモデル『DeepSeek-R1』の開発中に重要な現象に出合いました。
訓練の途中でモデルが「待って、待って。待って。今、重要なことに気づいた!」と自発的に口にしたのです。
https://t.co/yRQgGuO46S
その後モデルは「こっちを再検討しよう」と続け、
問題に対する新しいアプローチを始めました。
このような振る舞いは明示的にプログラムされたものではなく、自然に現れた行動でした。
これはモデル自身により『Ahaモーメント』と呼ばれています。
問題の解き方を直接教えるのではなく、適切なインセンティブを与えることで、モデルが自律的に高度な問題解決戦略を開発できることを示した象徴的な出来事だったようです。
なお厳密には研究の序盤で『DeepSeek-R1-Zero』という前駆的なモデルの開発中の出来事です。
また、実際のセリフは英語で"Wait, wait. Wait. That's an aha moment I can flag here."でした。