We're launching High Performance AI Lab today. We're a boutique AI lab focused on the AI stack for local and private inference, verifiable evaluations, and reproducible environments.
https://t.co/2whgTSJrOA
@tunguz But is it that good? At least what is available, in the scorecards they might use deep thinking features. In my exp using all the major closed sourced models gemini is the worst performing, also the harness plays a big role. It is a shame because I want gemini to be good
@k_dense_ai Wondering how you manage the plugin release lifecycle, not the SWE part but tje evals part, since these are skills that are used by agents do you have tests to measure if quality improve/decreased?
@fernandezpablo Sisi nadie sabe que va a suceder igual la diferencia de o1 mini y QwQ es chica. Y QwQ tiene 32b. Serรญa interese ver una version reasoning del nuevo DeepSeek de 605b