A 14B open model on a single RTX 4090 matched our hosted frontier model on text-to-SQL.
That wasn’t supposed to happen.
We ran 28 configurations, graded 25,000+ answers, and benchmarked against Snowflake + Databricks.
What we learned:
model size, model reputation, and vendor benchmark scores are bad shortcuts for predicting performance on your database.
So today we’re open-sourcing mnemiq.
Test it on your own data ↓