@RyanGreenblatt I ran a cross-entropy comparison of all raw text responses from numerous models using data I already had from a benchmark I run. I leaned on Fable for the stats-know-how; certainly seems very suspicious. The results/code are here for others to inspect:
https://t.co/pSbxYS6bWb