extremely unprofessional. if kimi wants to make it as a frontier lab, they need to act like one: perhaps silently route people to worse models, and maybe write a blog post about the collapse of humanity
Ex-NVIDIA engineer who built Unsloth explained RL, kernels, reasoning, quantization, and agents in 2 hours 42 minutes - better than $5000 fine-tuning bootcamps.
pick the base model -> write triton kernels for 2x faster fine-tune -> quantize to 4-bit -> run GRPO/DPO -> ship a reasoning model on your single GPU.
That loop is why Unsloth is the default way to fine-tune Llama, Qwen, Gemma, and Phi on hardware you already own.
Unsloth + Triton kernels + 4-bit quantization + GRPO/DPO + single-GPU fine-tuning - that's the stack.
Watch and save it, then fine-tune your first model tonight.
openai co-founder:
"i can't remember the last time i corrected ai. i just trust the system more and more
"now i am just vibe coding all the time"
in a 30-min sequoia capital podcast, Andrej Karpathy reveals his shift to software 3.0
agentic workflows + llms + vibe coding
worth more than a $500 prompt engineering course
bookmark & watch
Gente de instituciones diversas, por orden de importancia (de mayor a menor), CNMC, REE y Miteco.
Está muy bien que las renovables puedan hacer control de tensión. Nos falta el grid forming y, para baterías, el black start.
¿Para cuando? ¿Hace falta otro apagón? Espero que no.
The behavior of this curve seems so regular that I wonder if you could predict its future shape by doing some kind of physics simulation. Any simulation experts want to try that?
Seguimos. A las 10h25 la fotovoltaica superaba por primera vez los 30.000MW de generación en España. A las 12h35 pasamos al siguiente nivel. Superamos por primera vez los 31.000MW.
Pregunta contestada. Sí, hoy es el día en que superamos los 30.000MW de fotovoltaica. 30.034MW a las 10h25 en España, 66,2% generación. En España peninsular 68,2% (29.618MW)
With SpaceX #Falcon9 and Blue Origin #NewGlenn reusable launchers, Europe is totally out of the game and European mid and heavy launchers are extremely commercially compromised. #MIURANext will solve this huge distance in performance, price, cadence and reusability.
1/ Video models understand motion but hallucinate geometry. Image models nail geometry but are blind to motion. We have accepted this tradeoff for years. Meta FAIR just proved it is purely an architectural bug, not a theoretical limit. 🧵
𝐕𝐢𝐬𝐮𝐚𝐥 𝐛𝐥𝐨𝐠 on Vision Transformers is live.
https://t.co/N09njkKTXW
Learn how ViT works from the ground up, and fine-tune one on a real classification dataset.
CNNs process images through small sliding filters. Each filter only sees a tiny local region, and the model has to stack many layers before distant parts of an image can even talk to each other.
Vision Transformers threw that whole approach out.
ViT chops an image into patches, treats each patch like a token, and runs self-attention across the full sequence.
Every patch can attend to every other patch from the very first layer. No stacking required.
That global view from layer one is what made ViT surpass CNNs on large-scale benchmarks.
𝐖𝐡𝐚𝐭 𝐭𝐡𝐞 𝐛𝐥𝐨𝐠 𝐜𝐨𝐯𝐞𝐫𝐬:
- Introduction to Vision Transformers and comparison with CNNs
- Adapting transformers to images: patch embeddings and flattening
- Positional encodings in Vision Transformers
- Encoder-only structure for classification
- Benefits and drawbacks of ViT
- Real-world applications of Vision Transformers
- Hands-on: fine-tuning ViT for image classification
The Image below shows
Self-attention connects every pixel to every other pixel at once. Convolution only sees a small local window. That's why ViT captures things CNNs miss, like the optical illusion painting where distant patches form a hidden face.
The architecture is simple. Split image into patches, flatten them into embeddings (like words in a sentence), run them through a Transformer encoder, and the class token collects info from all patches for the final prediction. Patch in, class out.
Inside attention: each patch (query) compares itself to all other patches (keys), softmax gives attention weights, and the weighted sum of values produces a new representation aware of the full image, visualizes what the CLS token actually attends to through attention heatmaps.
The second half of the blog is hands-on code. I fine-tuned ViT-Base from google (86M params) on the Oxford-IIIT Pet dataset, 37 breeds, ~7,400 images.
𝐁𝐥𝐨𝐠 𝐋𝐢𝐧𝐤
https://t.co/N09njkKTXW
𝐒𝐨𝐦𝐞 𝐑𝐞𝐬𝐨𝐮𝐫𝐜𝐞𝐬
Dr @sreedathpanat Videos on ViT
ViT paper dissection
https://t.co/sg3JRvcgNG
Build ViT from Scratch
https://t.co/cnHzEeefDA
Original Paper
https://t.co/QiwrlDRQOc
Next up: demystifying Low-Rank Adaptation (LoRA) in PEFT!
Follow me @Mayank_022 along for more deep learning insights, cool fine-tuning projects, and updates from the upcoming blog posts.
@grok@levelsio@k_ristovski Is it fair to say that @levelsio rather not paid taxes because he has other reasons. Maybe a bit griddy. Since he is using the 2% reason to don’t help with 98% of his taxes?