Can AI turn a picture of sheet music into editable, searchable, playable music? 🎼
Excited to share that our paper LEGATO was accepted to ICLR 2026!
LEGATO is an OMR model that converts sheet music images into machine-readable music notation.
Demo: https://t.co/OBApsH7Gg1
Our goal: unlock large sheet music collections for musicians, researchers, educators, libraries, and developers.
From scanned pages → symbolic music.
From static images → searchable, editable, playable scores! 🎶
Can AI turn a picture of sheet music into editable, searchable, playable music? 🎼
Excited to share that our paper LEGATO was accepted to ICLR 2026!
LEGATO is an OMR model that converts sheet music images into machine-readable music notation.
Demo: https://t.co/OBApsH7Gg1
Music notation is not ordinary visual text.
It has tiny symbols, dense layouts, multiple voices, rhythm, pitch, measures, ties, beams, and long-range structure.
That’s why OMR needs models built specifically for music notation.
OMR matters because so much music still lives as PDFs, scans, and images.
If we can convert those pages into symbolic music, scores become easier to search, edit, analyze, transpose, and play back.
General-purpose AI models like GPT/Gemini still struggle with OMR. LEGATO makes a big leap:
• 68% TEDn error reduction
• 47.6% OMR-NED error reduction
Paper: https://t.co/FNpCcWyuTE
Audio & music evals in multimodal generation are tough—noisy metrics, vague correctness. 🎧😵💫
Our new work MMMG improves this with clear tasks, reliable metrics, and human-aligned judgments. 💯 If you’re working on audio/music in multimodal models, this benchmark is a must-try!
We introduce MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
✅ Reliable: 94.3% agreement with human judgment
✅ Comprehensive: 4 modality combination × 49 tasks × 937 instructions
🔍Results and Takeaways:
> GPT-Image-1 from @OpenAI leads image generation at 78.3% accuracy—13.7% ahead of the next-best model. The top open-source model, BAGEL from #ByteDance , achieves 45.5% accuracy.
> Audio generation is still challenging: Top open-sourced models achieve only 48.7% accuracy in sound (Make-An-Audio 2 from #ByteDance) and 41.9% in music (MusicGen from @AIatMeta).
📜 Paper: https://t.co/pFsEkJZfw8
🛠��� Code and Evaluation Suite: https://t.co/QRU05NlGSO
🥇Leaderboard: https://t.co/oGEFw7YRpc
🧵1/N