if you're going to say that your finetune performs better than jev on some baseline I hope your benchmark is trustworthy because I'm willing to be most of these "benchmarks" are
- generated in the spirit of "let's make this look good"
- overfit
- extremely narrow
@ziwenxu_ jev built for years in silence, and...
- "I built for 2 hours and.."
- "a day later we have.."
- "less than a week later we've..."
- "we built for two days and.."
@athletic_coder I have been pointing my harnesses to my own locally-hosted metasearch engine for months now with 0 degradation
all metasearch engines are wrappers around paid search results
no thanks
@eptwts > it technically can't "hallucinate" since it can't make things up
it can still be confidently wrong so "can't hallucinate" is very much marketing
@madiator@AlexGDimakis on the benchmark in question:
2,000 questions on probably unseen/held-out tasks:
jev: 83.7%
my local 4b finetune: 78.4%
nimble 9b: 78.1%
honestly at a point base weights stop mattering as much as you might think. your (fable's/astra's) synthetic data starts mattering more
- pick a domain from a list of 1,000 domains
- finetune a LoRA for that domain on predefined weights
your API becomes:
- infer domain
- route to domain classifier
infinite growth just means training cheap LoRAs?
hold my beer
@madiator@AlexGDimakis according to whose benchmark :)
my local 4b finetune, benchmarked against a large exam dataset against nimble on Modal for $1 says my 4b finetune outperforms it