Check out our new pre-print: “Schrödinger’s Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition”, work with @ryandcotterell https://t.co/6XzFTwDY0R
Code is available at https://t.co/XZReGggxO2
How much does an LM depend on information provided in-context vs its prior knowledge?
Check out how @vesteinns, @niklas_stoehr, @JenniferCWhite, @AaronSchein, @ryandcotterell + I answer this by measuring a *context's persuasiveness* and an *entity's susceptibility*🧵
@youngvulgarian I loved this and it chimed so much with my own experiences and nostalgia for an internet that doesn't exist anymore. I read it on Kindle and was just wondering last night whether the Paperback was available yet as I'd like to have it on my bookshelf!
@Dr_CMingarelli is there potential to use this data (now or when higher sensitivity data is available) to learn more about the geometry of the universe through looking at e.g. wave interference?
Update: we've started replicating their experiments directly with GPT4 calls, and somehow it only gets worse.
We've finished running zero-shot GPT 4 on the dataset, and after hand grading the first 30% of the dataset, the results don't seem to match the paper.
🧵
A recent work from @iddo claimed GPT4 can score 100% on MIT's EECS curriculum with the right prompting.
My friends and I were excited to read the analysis behind such a feat, but after digging deeper, what we found left us surprised and disappointed.
https://t.co/mpDqlenk04
🧵
As an AI researcher, I find the glorification of closed-source, proprietary models problematic. We should be emphasizing sharing and open-sourcing AI models, datasets and code, be it in conferences or in the press (4/4).
We need to stop conflating open/gated access and opensource.
ChatGPT is *not* open source -- we don't know what model is under the hood, how it works, or any other tweaks/filters that are applied. (1/n)
This suggests that the encoding of a polysemous word consists of a sum of a contribution from each sense vector, plus a component in the nullspace of the subspace they span. Combined with our first finding, this offers one possible explanation for homonym duplication.
Check out our new pre-print: “Schrödinger’s Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition”, work with @ryandcotterell https://t.co/6XzFTwDY0R
Code is available at https://t.co/XZReGggxO2
Next, we investigate how polysemous words are encoded by CLIP. We identify approximate vectors corresponding to senses of words like “bat”, and find that by editing the prompt representation in the subspace spanned by these vectors we can influence which sense is generated.
To what extent do neural networks learn compositional behaviour? Together with @nsaphra, Jon Rawski, @ryandcotterell and @adinamwilliams we take a lesson from formal language theory to answer this question.
https://t.co/xZpi9bKbGy
As my internship at @MetaAI comes to an end, I want to say a big thank you to my host @adinamwilliams, as well as @_dieuwke_ and Shubham Toshniwal. It's been great having the opportunity to work with you and hopefully there will be chances for more collaboration in the future 😊
Check out "Equivariant Transduction Through Invariant Alignment", my new paper with @ryandcotterell, which will be presented at COLING 2022! https://t.co/Xqb3RzEN1Q
We hope that this will motivate more analysis of the properties of equivariant architectures, and that the decoupling enabled by hard alignment may make it easier to adapt these architectures for more realistic tasks.