SemiAnalysis is out with a deep dive on $META Superintelligence, and they’re clearly impressed with Meta’s pace of improvement.
They think Meta may be the only major AI player on track to be world-class across data, talent, and compute.
Meta has reportedly turned its internal workforce into an RL data engine, with around 3,000 engineers working on RL tasks/environments, while also building five 1GW+ “titan” compute clusters.
SemiAnalysis says Meta’s compute ramp could give it more AI compute than OpenAI and Anthropic by year-end, while its scale-across networking could connect campuses up to 2,000 km apart.
Muse Spark 1.1 still is not at OpenAI or Anthropic level, but the note says Meta could catch or pass Google within 6 months if the current ramp holds.
🎉 Thrilled to share that our paper "Lumos: Empowering Multimodal LLMs with Scene Text Recognition" has been accepted to #KDD2024 in Barcelona!
https://t.co/nvTjuXz1bm
We describe a low latency STR+Multimodal LLM system that significantly improves text understanding capability!
If this can one day reliably identify and answer q's about what you're looking at—plants, architecture, landmarks, storefronts, products—it has a good shot at being the first gen AI killer app. Extremely ironic since object recognition long predates gen AI
https://t.co/p6WsxgYbIJ
@simonw We experimented with this here : https://t.co/EZwQqTfEAC
We saw that augmenting a multimodal LLM with the text extracted using a separate OCR component works best for even tough text-in-the-wild scenarios.
Meta presents Lumos
Empowering Multimodal LLMs with Scene Text Recognition
paper page: https://t.co/pqJcHOHaHM
introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component that extracts text from first person point-of-view images, the output of which is used to augment input to a Multimodal Large Language Model (MM-LLM). While building Lumos, we encountered numerous challenges related to STR quality, overall latency, and model inference. In this paper, we delve into those challenges, and discuss the system architecture, design choices, and modeling techniques employed to overcome these obstacles. We also provide a comprehensive evaluation for each component, showcasing high quality and efficiency.
📢Hot Research Alert: LUMOS
Lumos is an end-to-end system for on-device multimodal text-understanding! No more need to send images to cloud for processing. It has a core Scene Text Recognition (STR) whose output is fed to a MLLM.
Promising results👇Lets wait for comm. demo!
I've been trying Meta smart glasses' new multimodal AI - while it's pretty basic right now, it's still sick to see it combine what it sees from the camera with the language model to describe what it's seeing! Already solid for accessibility
Full episode: https://t.co/f8zFYUbZ9T
Shout out to @mkbhd who shares an early look at the early access multimodal Meta AI assistant on the Ray-Ban Meta glasses here: https://t.co/ZlBCFpJOci. We are working on answer length and making the trigger more conversational lol.
In the meantime, I had to run my own very important benchmark for computer vision. Iykyk.
The AI assistant on Ray-Ban Meta Glasses are gaining the ability to answer questions about what the camera can see and also get realtime information.
https://t.co/ZWXomJqXjK
"Green Federated Learning" summarizes the effort on quantifying and minimizing carbon foot print for training language models at industry scale using FL. https://t.co/qyqa0SZvqj
@MetaAI#icml2023
I'll be at #icml2023 next week and would love to connect and chat about all things applied ML.
I'll be presenting our work on "Green Federated Learning" : https://t.co/qyqa0SZvqj at https://t.co/zlu9CefG7g on Friday (Jul 28). See you all in Honolulu! @MetaAI
What if Wes Anderson directed The Lord of the Rings? We asked the community which video they want to see next and Lord of the Rings took the cake… or should we say Elven bread. We hope you enjoy this Midjourney to Middle-Earth.
#LordOfTheRings#WesAnderson#MovieTrailer#LOTR
Management technique we learned early on: short-term machine learning deadlines can be set based on inputs (e.g. high-quality execution on a set of experiments) but not outputs (e.g. reaching some level of performance). Science does not bend easily to the wishes of managers.
Flew air vistara multiple times over the last couple of weeks and the customer service at trivandrum airport was excellent. Also probably the best among the domestic airlines right now @airvistara