This is the kind of bug you only catch by actually reading how loss functions consume their inputs, not from a toy notebook.
thanks @RTurganbay for the review pointer on the cleaner fix ๐
PR here if you want to dig in ๐
https://t.co/6HekXrgvvd
Spent the afternoon reading @huggingface transformers' Florence2 code and found a bug that was quietly messing up training loss.
double-shift bug - same class of issue that hit Moonshine and CohereASR before this ๐งต
Fix is simple once you see it - pass shift_labels directly into the loss fn so it doesn't shift twice.
got a review comment from the HF maintainer suggesting this over my first attempt, which was the cleaner call.
# on shortification of "learning"
There are a lot of videos on YouTube/TikTok etc. that give the appearance of education, but if you look closely they are really just entertainment. This is very convenient for everyone involved : the people watching enjoy thinking they are learning (but actually they are just having fun). The people creating this content also enjoy it because fun has a much larger audience, fame and revenue. But as far as learning goes, this is a trap. This content is an epsilon away from watching the Bachelorette. It's like snacking on those "Garden Veggie Straws", which feel like you're eating healthy vegetables until you look at the ingredients.
Learning is not supposed to be fun. It doesn't have to be actively not fun either, but the primary feeling should be that of effort. It should look a lot less like that "10 minute full body" workout from your local digital media creator and a lot more like a serious session at the gym. You want the mental equivalent of sweating. It's not that the quickie doesn't do anything, it's just that it is wildly suboptimal if you actually care to learn.
I find it helpful to explicitly declare your intent up front as a sharp, binary variable in your mind. If you are consuming content: are you trying to be entertained or are you trying to learn? And if you are creating content: are you trying to entertain or are you trying to teach? You'll go down a different path in each case. Attempts to seek the stuff in between actually clamp to zero.
So for those who actually want to learn. Unless you are trying to learn something narrow and specific, close those tabs with quick blog posts. Close those tabs of "Learn XYZ in 10 minutes". Consider the opportunity cost of snacking and seek the meal - the textbooks, docs, papers, manuals, longform. Allocate a 4 hour window. Don't just read, take notes, re-read, re-phrase, process, manipulate, learn.
And for those actually trying to educate, please consider writing/recording longform, designed for someone to get "sweaty", especially in today's era of quantity over quality. Give someone a real workout. This is what I aspire to in my own educational work too. My audience will decrease. The ones that remain might not even like it. But at least we'll learn something.
9 years later, none of the "Attention Is All You Need" paper authors are at Google.
Ashish Vaswani - cofounder of Essential AI (recently exited)
Noam Shazeer - just moved to OpenAI
Niki Parmar - at Anthropic
Jakob Uszkoreit - cofounder of Inceptive
Llion Jones - cofounder of Sakana AI
Aidan Gomez - cofounder of Cohere
Lukasz Kaiser - at OpenAI
Illia Polosukhin - cofounder of NEAR Protocol
Your first open source contribution doesn't need to mass rewrite a codebase.
Sometimes 2 lines > 2000.
๐ https://t.co/PjVb2iiUVt
#OpenSource#HuggingFace#Python#AI
The fix: 2 lines of code.
โ Check if conversation is empty
โ Raise a clear ValueError instead of crashing
Raised issue #46752, submitted PR #46753.
Thank you Jensen and NVIDIA! Sheโs a real beauty! I was told Iโd be getting a secret gift, with a hint that it requires 20 amps. (So I knew it had to be good). Sheโll make for a beautiful, spacious home for my Dobby the House Elf claw, among lots of other tinkering, thank you!!
Today, while working on my AI agents, I registered for the Google Search API and discovered #GoogleDevFest!
I applied for the Delhi event on Oct 12 โ fingers crossed ๐ค
Later, I found Googleโs new AI Startup Program and spent the whole day exploring it.
#AI#Google#GoogleAI
Built a simple AI Agent using HTN (Hierarchical Task Network) with the Google DeepMind Gemini API!
HTN agents plan like humans โ breaking complex goals into smaller, structured subtasks.
GitHub ๐ https://t.co/uGidaIP3MT
#AI#Agent#HTN#GeminiAPI#DeepMind#AIProjects