Google AI Introduces STATIC: A Sparse Matrix Framework Delivering 948x Faster Constrained Decoding for LLM Based Generative Retrieval
STATIC (Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding) addresses the hardware inefficiency of standard prefix trees in LLM-based generative retrieval by replacing pointer-chasing traversals with vectorized sparse matrix operations. By flattening trie structures into Compressed Sparse Row (CSR) matrices, the framework achieves O(1) I/O complexity, enabling hardware accelerators like TPUs and GPUs to enforce business logic without the typical latency bottlenecks associated with irregular memory access. Deployed at scale on YouTube, STATIC delivers a 948x speedup over CPU-offloaded tries with a negligible per-step overhead of 0.033 ms, directly increasing fresh video consumption by 5.1% and significantly improving cold-start recommendation performance.....
Full analysis: https://t.co/r12tCUfLJN
Paper: https://t.co/piB5r9tMFb
Code: https://t.co/P7QNBBvsfI
@YouTube@GoogleDeepMind@GoogleAI@GoogleResearch
It's been a lot of fun working on this research, enabling extremely efficient constrained decoding for YouTube's LLM-based recommender system (https://t.co/dAqqLfuLxx).
It's also a great add-on to our PLUM generative retrieval framework (https://t.co/bgy193ayWe).
Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
YouTube presents a constrained decoding method that flattens prefix trees into sparse CSR matrices.
📝 https://t.co/a99bstsiXK
👨🏽💻 https://t.co/P2d1fAfxmH
@swyx@eugeneyan@aiDotEngineer YouTube and Google Deepmind also released a paper a week ago describing this work:
https://t.co/mpvMwhRkWX
PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations
LLM-Powered Nuanced Video Attribute Annotation for Enhanced Recommendations
@BoyuanLong et al. at Google use LLMs as annotators to achieve nuanced content understanding at scale for video recommendations.
📝https://t.co/kzjPCyXtHr
@shuchaobi It will also help us solve many puzzles in the universe - AGI/ASI agents on silicon hardware can travel/explore across the universe much more easily than human as they don’t need oxygen, water, etc for survival, and are not limited by a single lifespan.
We explored incorporating negative user feedback into the training objective of sequential retrieval models, and demonstrated its benefits using both live experiments and a counterfactual simulation framework that measures the recommender's responsiveness to user feedback.
Our recent work on "Learning from Negative User Feedback and Measuring Responsiveness for Sequential Recommenders" has been accepted at the #RecSys2023 Industry Track!
We'll be presenting this work at the @ACMRecSys conference in Singapore.
https://t.co/23sNjPTsAu
Our paper is out on @NatureComms! We generated human telencephalic organoids from stem cell-derived single neural rosettes, and studied autism-associated SHANK3 deficiency using transcriptomic and electrophysiological analysis.
https://t.co/KtC8GVrYMI
and https://t.co/h7UInLcv82
Dr. @AlexShcheglovit's lab (@shcheglovitov) led by Dr. Yueqi Wang (@yueqiw) investigated human telencephalic #organoids from stem cell-derived single neural 🧠 rosettes and the hemizygous deletion of an autism associated gene, SHANK3.
Out in @NatureComms: https://t.co/0mjKMmvAlK
New preprint out! Modeling autism-associated SHANK3 deficiency using human cortico-striatal organoids generated from single neural rosettes | bioRxiv https://t.co/VusOdAVvs6
Check out our new paper in Mol Psychiatry (https://t.co/dpaoBOaiDH)! SHANK3 hemizygosity disrupts AMPA synapses and neuronal morphology in human neurons transplanted in the mouse brain. Huge kudos to @SimoneChiola, @KandyNapan, @yueqiw! #SHANK3, #22q13, #autism, #synapses, #iPSC