Fundamental concepts every data engineer should know because they don’t really change
- ANSI SQL
- distributed compute
- OLTP vs OLAP
- CAP theorem
- slowly-changing dimensional modeling
- fact data modeling
- logging best practices
- AVRO / Thrift schemas
- idempotent pipelines
- job orchestration
- flexible schema vs defined schema
- data quality testing
#dataengineering
Here's how to systematically think about data lake query optimizations:
First think about business needs:
- Ask "what is the impact of sampling this dataset on downstream consumers?"
If the answer to this is minimal, then add sampling to your pipeline and that will probably make all other optimizations unnecessary.
- Ask "what are all the query patterns you're looking for?"
Adding in long-tail query patterns can sometimes make your pipeline and data model extremely bloated without adding much additional value.
Then think about logic bugs/flaws:
- Ask "what are my join types?"
You might have a bad join condition or could minimize the number of comparisons each JOIN operation is making. Remember string comparisons are much slower than integer comparisons
- Ask "am I bringing in too make additional columns?"
Select * belongs in adhoc queries only. Production pipelines should always list out every column
- Ask "am I filtering as soon as I can?"
Leveraging predicate pushdown with Spark and Trino will allow you to avoid shuffling so much data. The earlier you can slap a WHERE clause in, the better
Then think about window sizing errors:
- Ask "how many partitions am I scanning?"
If the answer is a lot (>30), then your query could most definitely benefit from cumulative table design. Tutorial: https://t.co/uQjykwRVqo
- Ask "Can I run this pipeline more often? Maybe hourly? to make it more reliable" Tutorial on how to run an hourly dedupe batch pipeline as efficiently as possible: https://t.co/x6SNyEoJGL
Then think about production problems:
- Ask "does my Spark job balance partitioning and memory correctly?"
Bumping up spark.sql.shuffle.partitions will increase the parallelism of your Spark job. But it also increases the network overhead so it's a balancing act with spark.executor.memory
- Ask "does my Spark job need to be skew aware?"
If your job is heavily skewed, adding adaptive execution with spark. Setting this setting to true will fix your problems immediately: spark.sql.adaptive.enabled
- Ask "do my upstream data sets produce enough files or splittable files to keep my initial parallelism high"
Slow initial reads can be a painful bottle neck for some Spark jobs. Working with your upstream producers so they output many files or splittable files will keep the job zooming along.
🚨🚨🚨 Claudia Sheinbaum Disturbing Truth EXPOSED ⚠️ ALL HELL IS ABOUT TO BREAK LOOSE ⚠️ Share this video quickly before they take it down: https://t.co/z1k4y81X07
I am 31 years old. 🧓
69 hours ago, I sold my AI startup for $100M 💸
I started working on it *just* 4 months ago.
Here's the step-by-step guide to building a great AI startup 🧵
@OliLondonTV It's ironic that rich people were aboard a vessel that ignored safety meaaures to look at a ship that ignored safety measures that was filled with rich people.
1/🧵🔍 Making sense of Principal Component Analysis (PCA), Eigenvectors & Eigenvalues: A simple guide to understanding PCA and its implementation in R! Follow this thread to learn more! #RStats#DataScience#PCA
Hyperparameter tracking is only 1% of experiment tracking.
When you use experiment tracking tools, you're sitting on a full suite of ML project management tools without even realizing it.
Here are 9 tools you get (with a one-line integration):