@JamesBaboono@NielsRogge@huggingface The 27B model is dense, whereas e.g. the larger 35B is MoE:
so the 27B a bit smarter, but the 35B is faster (less active params).
I guess this works (up to a certain size) for even larger models that are also MoE.
Here are a few images from a medieval manuscript to help you kick off your week 📖
This manuscript dates to 1460-70 and is FILLED with fascinating marginalia. Here are a few of my favs, including a medieval Yoda(?) and a quality unicorn 🦄
@laBnF Latin MS 1178
@proxy_vector@DanielLockyer @ChatGPTapp Hey Rohan, your idea about fixing search in ChatGPT got me thinking and I couldn’t resist putting together a quick prototype of Tagging.
I just sent you a quick DM (might be in Message Requests) and would love your thoughts 🙏
You spend more time on social media than you intend to, because time flows faster on these platforms, causing you to lose hours in what feels like minutes. This is no accident; it’s a result of a decades-long plot to steal your time.
My new essay.
https://t.co/7tnGcPxSK2
🚀 DeepSeek-R1 is here!
⚡ Performance on par with OpenAI-o1
📖 Fully open-source model & technical report
🏆 MIT licensed: Distill & commercialize freely!
🌐 Website & API are live now! Try DeepThink at https://t.co/v1TFy7LHNy today!
🐋 1/n
How artist Arthur Rackham, born on this day in 1867, revolutionized the business of visual art and the technology of books with his "Alice in Wonderland" illustrations, done when he turned 50 https://t.co/PIs3bp2Pih
How do we know what we know? Fascinating thought experiment about the limits of knowledge and the ongoing mystery of consciousness: https://t.co/XoevKbDR2Y
For folks who aren't able to attend in-person, we are excited to be able to stream all talks this year on our youtube channel:
https://t.co/UwIHP3wmzE
Papers can be found here:
https://t.co/nr5sGg8UEU
Please share!