Beautiful from @gretel_ai!
Synthetic_text_to_sql stands as the largest and most diverse synthetic Text-to-SQL dataset available to-date.
The dataset includes:
- 105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
- Coverage across 100 distinct domains/verticals
- Comprehensive array of SQL tasks: data definition, retrieval, manipulation, analytics & reporting
- Wide range of SQL complexity levels, including subqueries, single joins, multiple joins, aggregations, window functions, set operations
- Database context, including table and view create statements
- Natural language explanations of what the SQL query is doing
- Contextual tags to optimize model training
Blogpost: https://t.co/SVZ9iN6cVs
Dataset: https://t.co/6Nqhjwvzfz
With @gretel_ai, limited data supplies and data access issues are no longer blockers to creating state-of-the-art #AI apps and unlocking valuable insights in your data.
If you’re a #VertexAI user, you can tap into Gretel’s API-driven #generativeAI modeling capabilities today!