a Tsinghua University lab just put a project on GitHub that replaces a $400,000 H100 rack with a single 24GB graphics card.
it's called ktransformers, and the trick is almost stupidly simple: the experts you actually use stay on the gpu, the ones you don't sit on the cpu until they're called
/ deepseek-v3 and r1 with 139K context in 24gb of vram
/ up to 28x speedup over the standard setup
/ fine-tune deepseek-v3 across four rtx 4090s instead of a datacenter
/ built by tsinghua university's madsys lab, not a startup with a landing page
apache 2.0, and already past 17,000 stars.
-> https://t.co/DjF8CutVAa
bookmark it.
We’re entering a data-scarce regime; but is less data always worse? Our #ICML work shows that small-set training (e.g., with repetition) can be actively leveraged as a useful optimization bias for structured reasoning tasks. Come chat with Jingwen and Ezra in Seoul to know more!
When frontier labs suddenly cut costs, or say "we found a way to dramatically cut inference memory!!"
This is what they found. The secret sauce. Vision always wins!
Of course, my very capable ex-colleagues who now work on Gemini and Claude already found out years ago:
This is one of the papers I'm quite excited about in the past few weeks. It's a very simple but practical modification to the DINOv3 training framework.
Let me explain how it works.
based on our results under DiffusionBench, LDM converges much faster than pixel-space diffusion at 100k steps. curious why so many people are switching to pixel space, am i missing anything?
https://t.co/Rhmk7OUZv6
Scaling laws predict an LLM's pretraining loss, but not its capabilities. Abilities like in-context learning emerge abruptly and only past a certain scale. Our new paper traces this to one bottleneck: learning which tokens attention should focus on. 🧵https://t.co/ja0wc8aK2e
Latent-space models are a cage we’ve boxed ourselves into. The reason for using them in the first place was always efficiency, but we lost the plot and forgot that the speed costs us in terms of progress. It’s time to move on to pixel-space models for the next state of the art.
Speaking of recursive self improvement, @nayoung_nylee recently defended her thesis which among other things showed how transformers can learn progressively harder tasks by generating solutions to problems that sit *right at* the boundary of their capabilities.
https://t.co/UPYEYl7Do7
This paper helped me 1) overcome my obsession with transformers and arithmetic and 2) appreciate the value of environments.
She and @jackcai1206 did this before GRPO btw
"Can vision transformers learn without natural images?"
has just reached 50 citations! In this paper, we showed that Vision Transformers can be visually pre-trained without using any natural images.
Monica Lam (Stanford Professor):
"AI writes shallow reports for one reason, you ask it ONE question. Our method asks dozens, like a journalist, and the same chatbot starts writing articles 25% better organized than top AI."
paper presented at a leading AI research conference, her Stanford lab (OVAL) unveiled STORM - a method already used by 70,000+ people to generate Wikipedia-grade, fully-cited articles on topics it has never seen.
the secret formula: 6–8 expert perspectives + cited expert interviews + ruthless outline + grounded section-by-section writing + blind-spot red team = a report you can actually trust.
watch the full breakdown to copy all 5 prompts into Claude.
save this post so the formula is ready when you need it.
Earlier this year I was getting frustrated with Claude's charts, fed this book to claude and had it generate a Tufte skill. Instantly got simpler/more beautiful visualizations.
https://t.co/lfXwyQfmQG
그녀의 남자친구는 배우 이도현
어려운 형편 속에서도 발달장애를 가진 동생을
평생 책임지겠다고 다짐하며 살아온 사람이다.
원래 농구선수가 꿈이었으나 아버지의 반대로
좌절한 후, 연기에 매력을 느껴 아르바이트를 하며
연기 학원을 다녔다고 한다.
'슬기로운 감빵생활'로 데뷔한 이후, '호텔 델루나',
'더 글로리', '파묘' 등 수많은 히트작에 출연하며
연기력을 인정받고 집안의 빚을 모두 청산했다고.
배우 임지연과 공개열애 중인데,
너무나도 보기 좋은 선남선녀 🤎