Stop asking stakeholders what they want.
One notebook page: why the wish list question fails, who has to be in the room, the four questions that surface the real problem, how to capture it, and how to end with a ranked list and an owner.
The recap that makes it real: Seven lines, sent within 24 hours. Fill the brackets while the room still remembers.
Trap: the requirement is a solution: Three things you will hear in every workshop, and the question that opens them up.
Your query scans 2 TB to read one day.
One notebook page: what a partition is, how pruning skips the days you do not need, the filters that silently turn it off, clustering as the second lever, and the BigQuery guardrails.
Coarse partition, fine cluster: Partition on the date everyone filters by. Cluster on the columns they filter next.
Trap: the filter is there, the scan is full: Three BigQuery filters that look selective and read every partition anyway.
DISTINCT does not dedupe your table.
One notebook page: what a duplicate really is, how to find them, the ROW_NUMBER pattern that keeps the latest row, the tie-breaker rule, and how to stop them at the source.
Keep the latest row per key: Number the rows inside each key, newest first, and keep number one.
Trap: dedupes that do not dedupe: Each one passes a quick look. None of them gives one row per key, every run.
LIMIT 1 OFFSET 1 fails the interview.
One notebook page: the second highest salary question, why the naive answer breaks on ties and one-row tables, the three robust answers, and the nth and per-department variants.
The answer that survives a tie: Rank the salary values, not the rows. Ties share a rank, and an empty result becomes NULL.
Trap: same question, three tables: The naive query is right on clean data. Interview data is never clean.
Full sheet as a free print-ready PDF, no signup:
https://t.co/x7uJgftMbp
which of these need scaling: KNN, random forest, logistic regression, XGBoost? π
You scaled your data for a random forest.
One notebook page: the two formulas, which models need scaling and which ignore it, which scaler to pick, and the leakage mistake that makes your CV score a lie.
The scaler lives inside the Pipeline: Fit on train, transform test, never the other way round. The Pipeline makes that impossible to get wrong.
Trap: three scaling mistakes: Two are wasted work, one is a leak that inflates your score.
Upstream renamed a column. You found out at 3am.
One notebook page: what a data contract is, what it must cover, where it lives so it is enforced, which changes break it, and how to roll a breaking change out without a pager.
A contract is a file, not a meeting: Twelve lines next to the producer code. CI reads it, consumers read it, nobody is surprised.
Trap: the same change, with and without: Three real incidents and how a contract turns each one into a blocked pull request.
Your customer moved. Your warehouse disagrees.
One notebook page: why dimensions need versions, the three SCD types with what each one loses, the surrogate-key rule, and the join that silently doubles revenue.
Type 2, in two statements: Close the current version, open a new one. Facts written from now on join the new surrogate key.
Trap: the report that changed overnight: Three SCD mistakes and how each one shows up in a dashboard.