SuperSplat 2.19.0 is here! ✨
Your free and open source Gaussian Splat Editor
🎞️ Create videos of 4D Gaussian splats
➡️ Import support for SPZ and KSPLAT added
🌐 HTML viewer export now based on SOG
⌨️ New hotkeys for animation authoring
[1 / 2]
I am seriously considering writing a blog to show attention sink circuits since I get some free time after my internship. Yeah, I also think massive activations (outliers) are used to rescale the non-outlier dimensions so that we can get a k bias, and then get attention sink. Good explorations from @Alibaba_Qwen
You should for sure check out Jeffrey's project, it goes into so much depth across the whole deep learning systems field, so bullish on the resources he's making 🙏
📜🚨
Introducing TensorLens! 🔎
Our new tool for Transformer & LLM interpretability.
The problem: attention matrices are (i) a shallow view that ignores embeddings, FFNs, and values, and (ii) there are too numerous (per head & layer), which quickly becomes overwhelming.
🧵 1/6
Uh guys. Remember COCONUT (Chain-of-Continuous-Thought)?
It doesn't work. It's LARP.
«latent tokens tend to act as placeholders rather than semantically meaningful representations.»
Perturbing latent COT ≈doesn't affect the output, which is dependent on shortcut reasoning.
Here is a Google NeurIPS paper on how to improve LLM results at virtually no cost:
Scaling Embedding Layers in Language Models
Normal LLMs have a fixed vocabulary, usually around 200k tokens, and each token has its own embedding. [1/N]
@carlesgelada Linear layouts solves some of the problems with CuTe layouts: https://t.co/3xJDORegzo
i.e swizzling isn't tacked on.
https://t.co/mDBgGoGqHA
Tomorrow Jan 17 at 10am PST we'll have @LoubnaBenAllal1 going over her book "The Smol Training Playbook: The Secrets to Building World-Class LLMs"
It's a wonderful and comprehensive reference for those of us that care about open models
https://t.co/GfQR8ucTmR
Wanting to properly understand LLMs, I built an educational, UI-based, from-scratch implementation of a decoder-only transformer, designed to close the gap between papers, lectures, code, and actual understanding.
You can:
• pre-train a small GPT-, LLaMA-, etc.-style language model on your own raw text
• fine-tune it on prompt/response data
• step through training and generation
• inspect attention, tokenization, and loss as they evolve
• read through code (with multiple implementations), equations, and architecture diagrams for your specific configuration
• run inference
The emphasis is clarity, not performance.
Big thank you to @stanfordnlp, @ChrisGPotts, @karpathy, @calsmcdougall, @amelia_f_hardy, @NeelNanda5, @rasbt, and everyone else who has taught me so much.
https://t.co/jt9AZvwYhf
We raised $250M to accelerate building AI's unified compute layer! 🔥 We’re now powering trillions of tokens, making AI workloads 4x faster 🚀 and 2.5x cheaper ⬇️ for our customers, and welcomed 10K’s of new developers 👩🏼💻. We're excited for the future!
[paper release!]
Did you know that you can
- speed up any LLM by 4x
- and reduce its memory footprint by 2x
- and improve its results
- without modifying the model at all
How???
Here is how we do it 🧵
Triton is nice if you want to get something onto a GPU but don't need full performance/TCO. However, if you want peak perf or other HW, then Mojo🔥 could be a better fit. I'm glad OpenAI folk are acknowledging this publicly, but I wrote about it here:
https://t.co/dzGlAUapOY
Interesting use of LLM models: Preparing for the Peter Carr Conference in Winterthur (ZHAU): ChatGPT explains to me MY OWN derivations and proofs of 10 y ago without straining my memory & why >> Bonferoni.
Also provides verifications, like a slightly dumb referee.
Batch inference at scale. Mojo in the Stack Overflow Developer Survey. Devs experimenting with everything from Gaussian splatting and probabilistic data structures to GPU puzzles. Catch up on everything that's happened in the community over the past month:
https://t.co/DnklK0hJYG