Learn more about the latest advances in AI and systems, including LLM serving, efficient attentions, structured outputs, scaling up training, and more topics. Check out #MLSys2025. Accepted papers at https://t.co/sTsbrxWHlw and register today at https://t.co/2iRbuiDirc
Machine unlearning is taking off! There is a ton of interest in getting generative AI models to “unlearn” targeted undesirable information, often to meet policy aims in privacy, copyright, safety, etc.
Our new cross-disciplinary collaboration shows why this isn’t simple.
🦙QTIP 2, 3 and 4 bit Llama 3.3 models are now up on HF: https://t.co/mIkLRdbw76. Almost no zeroshot degradation at all sizes. 2 bit QTIP 3.3 70B fits on a single 4090 and gives pretty high quality generations (📹).
🧵
🏎️ Want faster, better quantized LLMs? Introducing QTIP, a new LLM quantization method that achieves a SOTA combination of quality and speed – outperforming methods like QuIP#!
🧑💻+🦙(w/ 2 bit 405B!): https://t.co/o0m95o0WWo
📜https://t.co/uwdhZB5OOB
🚀🌟🚀Excited to announce Samba-CoE v0.2, which outperforms DBRX by @DbrxMosaicAI and @databricks, Mixtral-8x7B from @MistralAI, and Grok-1 by @grok at a breakneck speed of 330 tokens/s.
These breakthrough speeds were achieved without sacrificing precision and only on 8 sockets, showcasing the true capabilities of dataflow! Why would you buy 576 sockets and go to 8 bits when you can run using 16 bits and just 8 sockets. Try out the model and check out the speed here - https://t.co/A3lB3RTEo2.
We are also providing a sneak peak of our next model, Samba-CoE v0.3, available soon with our partners at @LeptonAI. Read more about this announcement at https://t.co/DIiuJkPlKg
New this year, there will be a Young Professionals Symposium on Monday, which provides a forum for young professionals in industry and academia, to discuss important research & career questions/challenges: abstract submissions are now open (https://t.co/CEFUYSWRKu).
We are excited to announce the technical program for MLSys 2024! The provisional set of accepted papers is now available on the website at https://t.co/JPl5HbENr0.
Register for MLSys now at https://t.co/4yNtiwQkUQ
This year’s conference will be held at the Santa Clara Convention Center, from Mon May 13th through Thu the 16th, and features keynote speakers: Yejin Choi (University of Washington/Allen Institute for AI), Jeff Dean (Google), and Zico Kolter (CMU/Bosch Center for AI).
We have a fun attack that lets you extract the last-layer embedding weights of an LM via public APIs. It's really simple & uses SVDs!
https://t.co/GZnCrJUSNd
We discovered this for ChatGPT + PaLM-2. We privately disclosed, they fixed, now, we release :)
We are excited to announce this year’s keynote speakers for #MLSys2024: Jeff Dean @JeffDean, Zico Kolter @zicokolter, and Yejin Choi @YejinChoinka! MLSys this year will be held in Santa Clara on May 13–16. More details at https://t.co/HlUiTXLJTK.
🧵 (1/n)
👉 Introducing QuIP#, a new SOTA LLM quantization method that uses incoherence processing from QuIP & lattices to achieve 2 bit LLMs with near-fp16 performance! Now you can run LLaMA 2 70B on a 24G GPU w/out offloading!
💻 https://t.co/eM8Br2CqL3
Christopher De Sa and Yucheng Lu won an Outstanding Paper Award Honorable Mention at the 2021 International Conference on Machine Learning (ICML) for "Optimal Complexity in Decentralized Training.” @chrismdesa @yuchengthekid
https://t.co/w3X1aEjJT4
Cornell computer science professor Kavita Bala, a leading expert in computer graphics and computer vision, has been named dean of the Faculty of Computing and Information Science. @CornellCIS https://t.co/6LmZVRpbsl
Schools are shutting down today, but childcare is still a collective social responsibility. I've spent the last few days working with a crack team of coders to build this (incredibly cool) tool for scheduling MICRO childcare co-ops https://t.co/AuXwrzPZul