Announcing a cool new feature in @TensorFlow: differentiable map ops! ๐โจ
If you've ever wanted to store and train embeddings in an embedding model, this should make the process *much* simpler. ๐
https://t.co/0ybR7BB6QN
Amazing work from our TF intern, @kattian_ (@Harvard)!
Want to learn about SAE feature sensitivity, and our new method for measuring it efficiently? Catch Claireโs spotlight talk and poster at the NeurIPS Mech Interp workshop tomorrow, 11-12:30 in Upper Level Room 30A-E. Super fun project by @cla_tian, mentored by @NathanHu12 and me!
Activating examples of SAE features are often interpretable, but does a feature reliably activate on all inputs of a given concept? We built an automated evaluation to study feature sensitivity: whether features activate on text similar to their top examples.
one of the most important things I know about deep learning I learned from this paper: "Pretraining Without Attention"
this what I found so surprising:
these people developed an architecture very different from Transformers called BiGS, spent months and months optimizing it and training different configurations, only to discover that at the same parameter count, a wildly different architecture produces identical performance to transformers
this may imply that as long as there are enough parameters, and things are reasonably well-conditioned (i.e. a decent number of nonlinearities and and connections between the pieces) then it really doesn't matter how you arrange them, i.e. any sufficiently good architecture works just fine
i feel there's something really deep here, and we may be already very close to the upper bound of how well we can approximate a given function given a certain amount of compute. so we should spend more time thinking about other questions, such as what that function should actually look like (what data? which objective function?) and how to make it more efficient
Iโm really excited to be starting a new adventure with multiple amazing friends & colleagues.
Our company is called Physical Intelligence (Pi or ฯ, like the policy).
A short thread ๐งต
@nibnalin@amirbolous thatโs a reasonable hypothesis but it doesnโt seem like thatโs the full story - evan chen reviewed a subset of proofs and verified theyโre not all just coordinate bashing
We wrote a position piece on how LMs have expanded the toolkit for social equity researchers, especially in health - check it out, and feel free to share thoughts! https://t.co/1k4FG4BszE.
(and feeling very lucky to have worked with so many cool co-authors who I look up to!)