@OpenRouter@GoogleColab@posthog Using AI now is like having an intern/fresher at all time. Just tell it to "hey, use colab cli to evaluate jeeves on our dataset", and it do everything then report back reuslts.
OK so @OpenRouter jev is up so I cannot wait to run a test on our niche problem: classify 203 classes, small test data set of 488.
- jev: 67.0%
- Gemini 3.5 Flash Lite: 59.4%
fine-tuned:
- ModernBERT-base: 66.6%
- ModernBERT-large: 69.5%
- Qwen 3.5 2B LoRA: 72.5%
@Gabriel1860256@PawelHuryn When look at any benchmark, look at: effort, turns, price. Yes sometimes smaller models outperform bigger models, but at what cost?
@GergelyOrosz Accessible is one problem but at least cursor, codex, agy has some free tier; and some give students pro plans (they don't even know). I ask them to read the news, open twitter, reddit whatever, catchup with the cutting edge but seems they don't really care.
@GergelyOrosz Funny, I have talked as alumni to students in college last week: Most students only use chatgpt and claude, gemini free; never use Claucde Code or Codex.