I've completed the World War II language model evaluation set and it's corresponding leaderboard. Anyone can view. I will add more open-weights models as time permits.
This leaderboard tests a broad range of capabilities all focusing on World War II.
Peacebell is in the lead among the models I've scored thus far.
This evaluation set and leaderboard is, so far as I can tell, the only one out there focusing on WWII.
https://t.co/Ge6T0TfUx8