@carad0 I think it stems from the fact that people don't believe current systems possess sufficiently dangerous capabilities. I'm sure most of them would take the matter seriously if they encountered even hints of a real threat from the models.
@sir_deenicus@xlr8harder It's generally hard to judge performance of language models. Standard benchmarks can mislead on how it will perform in real tasks. I usually go for giving a 1k token story and trying to spot incoherences in flow and logic, but it's not a scientific approach by any means.
@sir_deenicus@xlr8harder I've compared with GPT-J with repetition penalty on for both models. On prompts that I've tested 7B LLaMA looked more coherent, but I'm not sure that my testing was adequate.