Apple presents Manzano: Simple & scalable unified multimodal LLM
• Hybrid vision tokenizer (continuous ↔ discrete) cuts task conflict
• SOTA on text-rich benchmarks, competitive in gen vs GPT-4o/Nano Banana
• One model for both understanding & generation
• Joint recipe: pretrain + refine + SFT
• Scales cleanly (300M → 30B) with consistent gains
At @Apple AIML, we always care about model cost and efficiency, including generation models. Some attempts we made recently to simplify MMDiT, moving towards faster and stronger generation models!
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
We address the long-standing problem of how to learn effective pixel-based image diffusion models at scale, introducing a remarkably simple greedy growing method for stable training of large-scale,
Leveraging #ChatGPT (GPT-4 model) to tackle complex #LaTeX package conflicts and resolve formatting issues in subfigure/table while drafting my thesis. It's been a game-changer! 📚🤖 Huge shoutout to #OpenAI for creating such a helpful tool!👏 #AcademicTwitter#ThesisWriting
Today (10/26) we're presenting our work on Fine-Grained Audiovisual Categorization with the SSW60 Dataset at #ECCV2022 poster #115 from 3:30-5:30pm.
What are the benefits of image, audio, and video modalities, and how should you train your model? Drop by and find out! 🐦🤖🐤
Today (10/26) we're presenting our work on Fine-Grained Audiovisual Categorization with the SSW60 Dataset at #ECCV2022 poster #115 from 3:30-5:30pm.
What are the benefits of image, audio, and video modalities, and how should you train your model? Drop by and find out! 🐦🤖🐤
Can we use audio and motion modality to improve open-vocabulary video classification?
We equip CLIP with cross-modal fusion to leverage multimodal information. Our method MOV archives SOTA results on UCF and HMDB zero-shot action recognition.
https://t.co/rEp4YRXE6n
Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models
abs: https://t.co/nw8EPuxWtp
achieves sota results on UCF and HMDB zero-shot video classification benchmarks, outperforming traditional zero-shot methods and recent methods based on VLMs