7.31:
"we plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3's design."
8.3 : production ready video Gen M is open 🤗😍
https://t.co/kVs5xOH70d
@Francis_YAO_ Linear attention is amazing . Now with nope MLA linear attention has fundamental transformed the logics of KV cache D. Your kv cache will never scale with sequence length and work very well. So next question , where is K3 flash ?
quick things you should know about CDNA5 https://t.co/XY1PftvG9B for Mi455:
1. WarpSIze : 32 , no more 64 (4 cycles x 16 SIMD to 1x cycles x 32 SIMD)
2. WGP (2 CU)for multicast
3. New Tensor Data Mover analogue to TMA in Hopper
4. 1024 vgpr for simd
5. Better Noc cross XCU
@geekbb For more details of the insight see the following:
RedNote : https://t.co/XbZXFvf4d7
BlueBook : https://t.co/oM1v8u5XfS
X : https://t.co/m8u2Pr0pT7
"Apparently K3 linear attention has fundamentally transformed logics of SSD and KV cache, you cannot assume that SSD grows with KV cache", an investor visited me today in campus this morning.
"Apparently K3 linear attention has fundamentally transformed logics of SSD and KV cache, you cannot assume that SSD grows with KV cache", an investor visited me today in campus this morning.
Deepseek V4 flash (7-31) finally out ! And after few days thinking about K3 KDA and AttnRes I just found some new opportunity where we can apply optimization upon.
@suno For more technical details see :
For more info , see this article : https://t.co/NzWTtnqvuL
[1] Qwen-Music Report: https://t.co/A9tnD2y5h9
[2] BTC-LLM (our lab's quantization work): https://t.co/Nfdx7nYnRW
[3] StreamFlow (NeurIPS 2025): https://t.co/36hrZ6LafA
recently dived into Qwen-Music [1], a new open music generation system that achieves SOTA results in expert evaluations—outperforming Suno V5.5 and Mureka V8 on key metrics, and ranking 3rd on the Artificial Analysis leaderboard (as JazzCat).
A colleague from our research group contributed to this work, so to better guide our subsequent audio optimization efforts, I analyzed its core computational characteristics.
The work now is very close to the top AI music Suno V5.5 @suno
@a_karvonen Meta Monarch should be pushing for week 0 support. The recently released models like Think Machines and K3 are worth considering.
Monarch's support has been quite disappointing. I'm not sure if the management team even understands the critical priority.
@marius@joespeez
@a_karvonen@marius@joespeez This allows teams that are eager to deliver quickly to get immediate results on top of our work. It's a win-win.
On the other hand, Meta's solution was to develop TorchForge to address usability issues, but that ended up diverting their team's focus and attention.
@a_karvonen@marius@joespeez A few observations:
Most researchers prefer to modify code within their familiar frameworks. New lightweight frameworks mainly attract fresh graduates as early adopters.
Drawing from our experience with Miles last time, I believe the key to success is providing week 0 support.
@MiaAI_lab hi we have enabled deepseek-v4 q2/q4 in a single dgx, that means with 2 spark you can implement P1D1 disaggregation to get even better performance. Dgx has maximum 300 toks/sec prefill speed, and 16 tokens decoding speed . Did you try it ?
Your Muon Optimizer can be even more faster !
Accelerating SymGemm with Triangular Scheduler (friendly for NoC clusters and L2 cache affinity!)
## Intro
The author from MIT Laker New House once delivered a partial implementation of accelerated symmetric matrix multiplication