Databricks spent $1-2 BILLION dollars to acquire a ~30 person company from the creators of Apache Iceberg.
A revolution is going on in the Big Data space and its centering around Iceberg. 🧊
Why would Databricks spend this outrageous amount of money on such a small company?
Figure out here (2-minute read) 👇
The story of Iceberg is a classic disruption story. You set out to solve a relatively niche problem and the solution ends up “accidentally” solving a larger problem. 🏆
The relatively niche problem in this case was Apache Hive. 🐝
Apache Hive was a popular query engine for big data sets, and it implicitly used a simple table format.
✋ Pause. Quick primer on the terminology used here:
• 📁 file format - a format which adds additional metadata to a file to help you organize the file’s data (e.g so you can read only what you need, so you can modify and evolve the file’s contents, etc) - stuff like CSV, Parquet, ORC, Avro
• 🗃️ table format - a format which adds additional metadata to a COLLECTION of files, so that again you enhance what you can do with them - e.g only read the file you need, modify the structure and organization of the files, add ACID capabilities, etc.
This includes open table formats like Iceberg, Delta, Hudi and many implicit ones (e.g MySQL’s InnoDB, PostgreSQL, Snowflake)
Hive had done a few things quite well:
• its format was simple and easy to understand 👍
• this made it ubiquitous - Hive tables are WIDELY supported in most query engines - Hive, Spark, Presto, Flink, Pig
• the gain was that the whole ecosystem could use the same at-rest data 👌
But Hive also had a few problems, succinctly:
• non-atomic writes when writing to multiple partitions of the data, resulting in mishaps (deleted data, half-done jobs) 😨
• inefficient with relation to cloud object storage ☁️
• scale challenges
Ryan Blue and Daniel Weeks, the creators of Iceberg, figured out that the main bottleneck was the table format itself.
So, while at Netflix, they set out to create a new format that solved for these issues.
The new format - Iceberg - improved on the following:
1. all changes are atomic with serializable isolation
2. support for many concurrent writers ⚡️
3. native cloud object store support 🌤️
4. no gotchas & surprises (e.g renaming a Parquet column in Hive breaks a ton of stuff)
5. … a lot more
In classic disruption fashion, the first three improvements ended up solving a much bigger problem. 🏆
Which problem was that?
The problem of Shared Database Storage.
With an open table format like Iceberg, you can store your data in one single source of truth (e.g S3) and have many different engines access and modify the data at the same time. 🤯
This is the rise of the so-called headless data architecture, where the storage layer (data) is decoupled from the query layers (engines) that use it. 💡
It is the key enabler of the growing trend called zero copy.
Zero Copy means that you do NOT have to spend millions in expensive cloud networking costs to copy petabytes of data to have it be used by the right processing engine – you can use the same set of data in the standardized Iceberg table format. 🧊
The two layers have always been tightly coupled because the query layer relies deeply on optimizations in the storage layer which allow for data to be fetched efficiently for faster querying.
And because the table format essentially defines the storage layer, you have big dogs like Snowflake and Databricks outbidding each other for Tabular (a company founded by the Iceberg creators) and aggressively competing with each other on the table formats. 💸
Just in the last few weeks we had some major announcements:
• June 3: Snowflake’s Open Source Polaris Iceberg Catalog announced
• June 4: Databricks acquires Tabular
• June 13: Databricks’ Unity Delta Catalog open sourced
And it seems like this is just the beginning…
Interested in more concise, simple content around the table format wars and the lakehouse revolution?
1. Follow me here - ✅ @kozlovski
2. Retweet this story so your network learns too. It takes 5 seconds to do, and it takes me 5 hours to write 🙏
L'AWS Summit de Paris 2023 fut une belle expérience tout au long de la journée... des milliers de visiteurs, plein de sessions, de témoignages clients et de stands intéressants . J'ai eu la chance d'échanger avec des clients et partenaires passionnés et h…https://t.co/DZKAmFUi16
L'AWS Summit de Paris, le 4 avril prochain, approche à grands pas.
C'est l'occasion pour vous de suivre de nombreux témoignages clients et de rencontrer également nos partenaires et nos experts.
J'aurai la chance d'accompagner Simon Parisot qui vous exp…https://t.co/qCtoOih3BM
For those of you in Belgium interested to find out more about the new stuff AWS announced at the last re:Invent, wait no more ... there are just a few spots left to attend the recap organized by the AWS meetup group on 1st Feb in Gent. https://t.co/VAzvyeGAE4
Fresh out of the press... Jeremy Ware and I just published a major refresh of my original blogpost on authenticating external applications with AWS. Hope it helps you folks more easily deal with such machine-to-machine scenarios.…https://t.co/pYXlG4TEVF https://t.co/Rh7BVrcSJD
Feeling uncomfortable doing in-place upgrades of your production database on Amazon RDS or Aurora? Then this new blue-green deployment feature might make your day...
#aws#amazonrds#amazonaurora https://t.co/ZcEKQay1D3
It was great seeing some good old faces again in person and follow very interesting sessions at the latest Data Science Leuven group meetup. Kudos to the speakers @bernardsacre @mLavaert Wim Vancuyck and organizers @peeterskris Wannes Rosiers #datascience https://t.co/8JCwWVoNF7
Amazon Lex now has a visual interface to help you more easily build great conversations. Check it out! #amazon#aws#lex#visual#bot https://t.co/NOUPK4FpV5
Interested to find out how a financial services institution like NatWest Group streamlined their end-to-end ML lifecycle using Amazon Sagemaker?
Please connect on 20th July. #ml#financialservices#aws https://t.co/yTrVC4bvY7
Inspiring keynote and sessions with great stories from customers and partners. Looking forward to the 2023 edition! #aws#AWSSummit https://t.co/jzEkqdkSgn
Hopefully a useful initiative for our M&E customers to more easily find their way among the various services and solutions AWS provides to support you in your challenges in this fast-paced industry #aws#mediaandentertainment https://t.co/PdNaoQTkhn
@chrisdlangton @chrisdlangton Thanks for your feedback.
Are you perhaps referring to the fact that both protocols rely on an external Identity Provider for authentication?
This is so cool, so many use cases 👉 Introducing Amazon S3 Object Lambda – Use Your Code to Process Data as It Is Being Retrieved from S3 ✍️ https://t.co/UmEna297kz #AWS#Storage#Serverless
ever wanted to test the resilience of your solutions before rolling them out ?
then this is a good motivation to check out AWS Fault Injection Simulator #aws#resilience https://t.co/k2BvFmnEJ8