@krishnanrohit either the 7B-chat is extra sensitive to RLHF or the system prompt is keeping it very nerfed. i would suggest trying out the 70B-chat on huggingchat
https://t.co/1xxKmebeOO
@proetrie to be fair it's a dev conference. it would look clownish to constantly talk about "AI" so they used the term "Transformer model" a couple of times
@Sim89776996@ValueRaider@brickroad7 iβm surprised how little attention the Falcon models have received but it looks like this news is a direct response
an open source 175B GPT-3 will obliterate any competition
Another work that chips away at architecture complexity.
Get rid of positional encodings in generative transformers, and just let it learn to do its thing β much better generalization to longer context sizes.
The bitter lesson of βlet neural nets be neural netsβ strikes again!