Hi, I am PhD Student at @ELLISInst_Tue & @MPI_IS. Was MLSys Enginner at @Qualcomm & @DJIGlobal. Focusing on efficient and scalable MLSys for models and agents.
I do agree. In our recent work, we find that even with the raw trajactory, it is hard to understand the behaviors and decisions of coding agents. We are working on something for less comprehension debt and will share it very soon.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Reasoning should stay transparent, even as reasoning get more efficient.
FlashLoop explores the efficiency side: making looped architecture faster and more memory-efficient.
Can we have both?
FlashLoop: https://t.co/JH6iHOnsgT
it's interesting how much GDM talks about "preserving transparent architectures" in this blog post. i wouldn't be surprised if gemini 4 is a looped architecture similarly to gpt-6 astra.
https://t.co/rYytZilGFR
🔥 Frontier labs are redesigning the residual stream to beat the curse of depth: Kimi K3 uses AttnRes, DeepSeek V4 uses mHC, ByteDance proposed HC, and there are more: LNS, KEEL, MoDA…
🙋 But each method was tested with its own training budget, model design and codebase, so the results can't be compared directly. Which one actually works best? More importantly, which one turns architectural depth into effective computational depth?
👇 We built DepthBench to find out.
📄 https://t.co/DpVhu2b44F
In the age of agentic programming, the compiler's role will increasingly move from just supporting compilation to a flexible harness that agents interact with and rely upon. Compiler harness will be the new frontier in agentic GPU programming.
🚀 Kernel Design Agents to Optimize Kimi Delta Attention (upto 2.96x)
When the KDA (Kernel Design Agents) project was initiated, its name happened to coincide with another KDA (Kimi Delta Attention). More than once, people asked us , "Can we use KDA to write a version of KDA?" Today, right at a moment with a full moon good for shooting, we are thrilled to share the latest results of KDA(gent)-v0.6: using KDA(gent) to optimize KDA(ttn)!
KDA-v0.6 achieves this through
* Humanize2 Flow & LLM Collaboration Optimization: Introducing multi-round iterative evolution with Flame Chase, GPT-5.6-Sol, and Fable-5, paired with an automated Workspace cleanup mechanism to significantly broaden the operator search space.
* Multilingual Domain-Specific Skill Stack Support: Fully supporting CuTe-DSL, CUDA C++, as well as recent Agent-Native CAKE IR and TIRx. Combined with customized diagnostic tools (such as CuTe's dedicated IKET pipeline analyzer and TIRx's CPU-side numerical simulation and static checks), achieving deep performance bottleneck diagnostics.
* Self-Evolving Kernel Wiki: Building a dynamically updated knowledge base that automatically cleans up invalid/erroneous implementations, corrects classification tags, and streamlines search indexing, significantly improving the Agent's knowledge retrieval quality and code generation accuracy.
KDAgent ultimately optimized the KDAttn kernel to 2.96x of FlashKDA, and we face a lot of unexpected hacks during the optimization. More details shared in our blogs https://t.co/IQT9WcLktx
Release results of KDA for KDA: https://t.co/d5njJtzfL5
Community & Production Deployment:
All kernels are open-sourced for community verification! For deploying NVIDIA Agentic CUDA acceleration in production environments, we recommend using the production-level version rigorously verified by FlashInfer.
By the Team: Dongyun Zou, Yixin Dong, Hongyi Jin, Junxian Guo, Yahui Cui, Avery Huang, Zihao Ye, Junru Shao, Changye Li, Zijian Zhang, Sihao Liu, Song Bian, Ligeng Zhu.
@AlberFuen Some training-free looped Transformers research shows that wrapping mid-stack layers of frozen pretrained models with recurrence at test time improves results on benchmarks by treating blocks as ODE refinements with damped sub-steps.
Looped Transformers enable deeper computation without more parameters.
But what is the cost? 🙋
Every loop adds compute and KV cache.
The key to inference efficiency is exploiting cross-loop redundancy.
Get started: pip install flashloop
Plug and play with FlashLoopEngine. The attached code card shows an Ouro-1.4B chat example.
Paper: https://t.co/JH6iHOnsgT
Code: https://t.co/epk0JfRsXW
Project: https://t.co/IsnBarcuF0
How does FlashLoop use this redundancy without training?
Token-sparse updates reuse converged states; sparse attention selects important keys using the prior loop; 4-bit KV-residual quantization compresses the cache.
Accuracy stays close to baseline across most tested settings.
What makes FlashLoop's lazy updates possible?
As loops proceed, state changes concentrate on a few tokens; attention-output differences are dominated by sparse, stable key columns; and adjacent-loop KV residuals become easier to quantize.
One axis, three kinds of redundancy.
Ouro-2.6B R4 has fewer parameters than Qwen3-8B. Yet it uses more memory and 8K prefill FLOPs (214.4 vs 133.6 TFLOPs ❗️).
Repeated loops create inference redundancy.
Let's see how FlashLoop solve it with lazy updates
How should we allocate a limited bit budget across MoE experts and layers?
We uncover a key limitation of data-driven bit allocation: changing the calibration domain biases performance toward that domain.
Introducing AlphaQ, a calibration-free bit-allocation method. 🧵
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
Inference scaling part 1.
Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x)
00:00 Introduction and recap
00:31 Training-time and inference-time scaling
07:52 What we'll implement
11:47 Notebook setup and model loading
17:43 Building a flexible text generation function
24:40 Chain-of-thought prompting
28:26 Sampling and output diversity
33:43 Next-token logits and greedy decoding
38:20 Temperature scaling step by step
42:46 Softmax and token probabilities
47:42 Multinomial sampling
54:51 Adding temperature sampling to text generation
59:31 Top-p filtering step by step
1:10:23 Adding top-p filtering to text generation
1:13:43 Sampling and LLM watermarking
1:16:01 Self-consistency and majority voting
1:20:36 Implementing self-consistency
1:29:02 MATH-500 results
1:35:01 Accuracy and compute tradeoffs
1:36:50 Next steps and self-refinement
Introducing Bespoke Nimble: an open data, open model, open recipe for an open Jev.
Code and info: https://t.co/aC7kPejrcj
Model: https://t.co/snwKGdhn1I
Data:
* A new data curation recipe called contrastive data curation.
* Slightly change facts to generate negative data. This pushes the model to discriminate better and become a better decision maker. The calibration is implicit.
* Didn't do ablations but I think this is a critical piece!
* This also means training data doesn't need probabilities.
* Data covered 10 categories, and is fully synthetic.
* This data is split into train and eval.
Training
* LoRA finetune of Qwen3.5-9B.
* Distillation-free: we use Jev to only evaluate.
* No RL yet!
Serving
* Parallel constrained decoding as suggested by @NielsRogge and @harshagundal.
Results:
* The post-trained Qwen (Nimble) became substantially better on our curated eval: 66% for Qwen to 90% for Nimble. Jev is at 93%.
* 100ms on H100 and free to use on your macbook! Feel the AGI for free.
* 2 days of building in public. :)
Big caveat is that there is no standard benchmark to measure performance, and it's possible Nimble is much worse on other benchmarks compared to Jev. But it should be better than Qwen!
We thank @typesafeai for making Jev and the inspiring discussions in the community. Hope this release lifts all the boats and encourages more research and activity in this space.