@OfficialLoganK PLEASE just make a Model better than Gemini 1.5 Pro GA on vertex AI (no 2.0 Flash is not better for long context tasks).
We are stuck on a 6 month old model …
@airesearch12 Or was he tricked by @csahil28? Matt didn‘t know what LORA was and in one of his early posts he mentioned, that benchmarks were done by sahil?
@ArtificialAnlys@mattshumer_ How do you run the MMLU benchmarks? The default implementation is logprob and not „generation“. I guess you run it as a generation benchmark and prompt the models to genrate the letter of the correct answer?
@cognitivecompai@FernandoNetoAi@maximelabonne I wonder how that would work for eg MMLU with lm-eval-harness?! MMLU is a logprob Benchmark which does not actually generate the response / answer? In my opinion there is no way to benchmark the Reflection model on MMLU and compare the results?!