You found a real bug on our side.
Our public API was replacing descriptive true/false criteria with literal "true" and "false" labels. That pushed noul from 80.5% to 65.2% (483/600 -> 391/600).
We fixed the adapter, added regression tests, and pinned the repro. The fix is committed, but not pushed to the Hub yet.
Commit:
https://t.co/hoDaparVI2
Also, your CPU results matched ours exactly. The remaining CPU/FP32 vs CUDA/BF16 gap is still unresolved, so we are keeping those results separate.
@ziyacivan Thanks for your honest feedback, ziya
We will continue to work on improving our models and shipping better results! :D
Stay tuned for what's next!
Resolvi fazer alguns testes com o Jev vs Julia 1 da @supersonicai
- Quadrante político (politicalcompass)
- Personalidade (16Personalities)
- Personalidade (BigFive)
Os mesmos prompts em português pros dois, usando as mesmas perguntas originais:
Small note for everyone following us:
With all the hype and how fast this account grew, we could’ve easily sold people on something that doesn’t even exist yet.
Instead, we’ve tried to be transparent about what we’ve built, who we are, what still sucks, and what we want to improve.
That felt like the right way to do it
We’ve aimed to be as transparent as possible with everything around this launch. We’re not perfect, so feedback and questions are always welcome. We’ll keep updating this page with more examples, details, and relevant updates.