𝗪𝗵𝗶𝗰𝗵 𝗰𝗹𝗼𝘂𝗱 𝗽𝗿𝗼𝘃𝗶𝗱𝗲𝗿 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝘂𝘀𝗲𝗱 𝘄𝗵𝗲𝗻 𝗯𝘂𝗶𝗹𝗱𝗶𝗻𝗴 𝗮 𝗯𝗶𝗴 𝗱𝗮𝘁𝗮 𝘀𝗼𝗹𝘂𝘁𝗶𝗼𝗻?
The diagram below illustrates the detailed comparison of AWS, Google Cloud, and Microsoft Azure, created by Satish Chandra Gupta.
They all use similar parts of solutions:
1. 𝗗𝗮𝘁𝗮 𝗶𝗻𝗴𝗲𝘀𝘁𝗶𝗼𝗻 of structured or unstructured data.
2. 𝗥𝗮𝘄 𝗱𝗮𝘁𝗮 𝘀𝘁𝗼𝗿𝗮𝗴𝗲.
3. 𝗗𝗮𝘁𝗮 𝗽𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴, including filtering, transformation, normalization, etc.
4. 𝗗𝗮𝘁𝗮 𝘄𝗮𝗿𝗲𝗵𝗼𝘂𝘀𝗲, including key-value storage, relational database, OLAP database, etc.
5. 𝗣𝗿𝗲𝘀𝗲𝗻𝘁𝗮𝘁𝗶𝗼𝗻 𝗹𝗮𝘆𝗲𝗿, with dashboards and real-time notifications.
Hyperscalers use different names for the same service type, e.g., AWS "lambda" is Azure "function."
Which services did you use, and for what kind of solutions? What are your experiences with these three providers?
The full text is in the comments.
--
If you like my posts, please follow me @milan_milanovic and hit the 🔔 on my profile to get a notification for all my new posts.
Learn something new every day 🚀!
#aws #technology #cloudomputing #data #softwareengineering
SQL, NoSQL, or something else — how do you decide which database?
The performance of your application can suffer if you choose the incorrect database type, and going back on a bad choice can be time-consuming and expensive.
Before we dive into the factors to take into account when choosing an appropriate database, let’s examine the characteristics of the most widely used databases.
In a relational database, data is organized in rows and columns where a row represents a record and its data fields are stored in columns. They are ideal for when ACID compliance is required, and a predefined schema can be created.
With columnar databases, records are stored as columns rather than rows. This makes them very performant for analytical purposes where complex queries are run across large datasets; especially those that contain aggregate functions.
In a document database, data is stored in a semi-structured format such as JSON. They offer a flexible and schema-less approach which makes them a great choice for data with complex or continually changing structures.
Graph databases are optimized for storing and querying highly connected data. Records are represented as nodes and relationships as edges. Under the hood, they use graph theory to traverse relationships between nodes to power performant queries.
Key-value stores are a simple form of storage where values are inserted, updated, and retrieved using a unique key. They are more commonly used for small datasets and often temporary purposes such as caching or session management.
Time-series databases are ideal for time-stamped data that are queried and analyzed in relation to time. They provide built-in time-based functions that assist in analyzing large datasets over time.
Each database type has been optimized for specific use cases.
It's important to thoroughly consider the correct database for your use case as it can impact your application’s performance.
Below are the considerations that should be made:
🔸 How structured is your data?
🔸 How often will the schema change?
🔸 What type of queries do you need to run?
🔸 How large is your dataset and do you expect it to grow?
🔸 How large is each record?
🔸 What is the nature of the operations you need to run? Is it read-heavy or write-heavy?
Use these questions as a starting point for your analysis. Take the time to investigate your use case and ask questions to your stakeholders and end-users when necessary.
It's important to invest as much time in this decision as needed. Choosing the wrong database type can be detrimental to your application’s performance, and difficult to reverse.
——
Want more engineering insights like this?
Subscribe to our free newsletter for a weekly deep-dive and roundup of all our best content → https://t.co/oSSc0CUweH
Thinking, Fast and Slow is a book everyone should read.
But, only 7% of people who’ve started it ever made it to the end.
21 insights from the best book you never finished:
What do you need to know about 𝗖𝗗𝗖 (𝗖𝗵𝗮𝗻𝗴𝗲 𝗗𝗮𝘁𝗮 𝗖𝗮𝗽𝘁𝘂𝗿𝗲)?
𝗖𝗵𝗮𝗻𝗴𝗲 𝗗𝗮𝘁𝗮 𝗖𝗮𝗽𝘁𝘂𝗿𝗲 is a software process used to replicate actions performed against 𝗢𝗽𝗲𝗿𝗮𝘁𝗶𝗼𝗻𝗮𝗹 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲𝘀 for use in downstream applications.
𝗧𝗵𝗲𝗿𝗲 𝗮𝗿𝗲 𝘀𝗲𝘃𝗲𝗿𝗮𝗹 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀 𝗳𝗼𝗿 𝗖𝗗𝗖. 𝗧𝘄𝗼 𝗼𝗳 𝘁𝗵𝗲 𝗺𝗮𝗶𝗻 𝗼𝗻𝗲𝘀:
➡️ 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲 𝗥𝗲𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻 (refer to 3️⃣ in the Diagram).
👉 𝗖𝗗𝗖 can be used for moving transactions performed against 𝗦𝗼𝘂𝗿𝗰𝗲 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲 to a 𝗧𝗮𝗿𝗴𝗲𝘁 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲. If each transaction is replicated - it is possible to retain all ACID guarantees when performing replication.
👉 𝗥𝗲𝗮𝗹 𝘁𝗶𝗺𝗲 𝗖𝗗𝗖 is extremely valuable here as it enables 𝗭𝗲𝗿𝗼 𝗗𝗼𝘄𝗻𝘁𝗶𝗺𝗲 𝗦𝗼𝘂𝗿𝗰𝗲 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲 𝗥𝗲𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗮𝗻𝗱 𝗠𝗶𝗴𝗿𝗮𝘁𝗶𝗼𝗻. E.g It is extensively used when migrating 𝗼𝗻-𝗽𝗿𝗲𝗺 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲𝘀 serving 𝗖𝗿𝗶𝘁𝗶𝗰𝗮𝗹 𝗔𝗽𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻𝘀 that can not be shut down for a moment to the cloud.
➡️ Facilitation of 𝗗𝗮𝘁𝗮 𝗠𝗼𝘃𝗲𝗺𝗲𝗻𝘁 𝗳𝗿𝗼𝗺 𝗢𝗽𝗲𝗿𝗮𝘁𝗶𝗼𝗻𝗮𝗹 𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲𝘀 𝘁𝗼 𝗗𝗮𝘁𝗮 𝗟𝗮𝗸𝗲𝘀 (refer to 1️⃣ in the Diagram) 𝗼𝗿 𝗗𝗮𝘁𝗮 𝗪𝗮𝗿𝗲𝗵𝗼𝘂𝘀𝗲𝘀 (refer to 2️⃣ in the Diagram) 𝗳𝗼𝗿 𝗔𝗻𝗮𝗹𝘆𝘁𝗶𝗰𝘀 𝗽𝘂𝗿𝗽𝗼𝘀𝗲𝘀.
👉 There are currently two Data movement patterns widely applied in the industry: 𝗘𝗧𝗟 𝗮𝗻𝗱 𝗘𝗟𝗧.
👉 𝗜𝗻 𝘁𝗵𝗲 𝗰𝗮𝘀𝗲 𝗼𝗳 𝗘𝗧𝗟 - data extracted by CDC can be transformed on the fly and eventually pushed to the Data Lake or Data Warehouse.
👉 𝗜𝗻 𝘁𝗵𝗲 𝗰𝗮𝘀𝗲 𝗼𝗳 𝗘𝗟𝗧 - Data is replicated to the Data Lake or Data Warehouse as is and Transformations performed inside of the System.
𝗧𝗵𝗲𝗿𝗲 𝗶𝘀 𝗺𝗼𝗿𝗲 𝘁𝗵𝗮𝗻 𝗼𝗻𝗲 𝘄𝗮𝘆 𝗼𝗳 𝗵𝗼𝘄 𝗖𝗗𝗖 𝗰𝗮𝗻 𝗯𝗲 𝗶𝗺𝗽𝗹𝗲𝗺𝗲𝗻𝘁𝗲𝗱, 𝘁𝗵𝗲 𝗺𝗲𝘁𝗵𝗼𝗱𝘀 𝗮𝗿𝗲 𝗺𝗮𝗶𝗻𝗹𝘆 𝘀𝗽𝗹𝗶𝘁 𝗶𝗻𝘁𝗼 𝘁𝗵𝗿𝗲𝗲 𝗴𝗿𝗼𝘂𝗽𝘀:
➡️ 𝗣𝘂𝗹𝗹 𝗕𝗮𝘀𝗲𝗱 𝗖𝗗𝗖
👉 A client queries the Source Database and pushes data into the Target Database.
❗️Downside 1: There is a need to augment all of the source tables to include indicators that a record has changed.
❗️Downside 2: Usually - not a real time CDC, it might be performed hourly, daily etc.
❗️Downside 3: Source Database suffers high load when CDC is being performed.
❗️Downside 4: It is extremely challenging to replicate Delete events.
➡️ 𝗣𝘂𝘀𝗵 𝗕𝗮𝘀𝗲𝗱 𝗖𝗗𝗖
👉 Triggers are set up in the Source Database. Whenever a change event happens in the Database - it is pushed to a target system.
❗️ Downside 1: This approach usually causes highest database load overhead.
✅ Upside 1: Real Time CDC.
➡️ 𝗟𝗼𝗴 𝗕𝗮𝘀𝗲𝗱 𝗖𝗗𝗖
👉 Transactional Databases have all of the events performed against the Database logged in the transaction log for recovery purposes.
👉 A Transaction Miner is mounted on top of the logs and pushes selected events into a Downstream System. Popular implementation - Debezium.
❗️ Downside 1: More complicated to set up.
❗️ Downside 2: Not all Databases will have open source connectors.
✅ Upside 1: Least load on the Database.
✅ Upside 2: Real Time CDC.
--------
Follow me to upskill in #MLOps, #MachineLearning, #DataEngineering, #DataScience and overall #Data space.
Also hit 🔔to stay notified about new content.
𝗗𝗼𝗻’𝘁 𝗳𝗼𝗿𝗴𝗲𝘁 𝘁𝗼 𝗹𝗶𝗸𝗲 💙, 𝘀𝗵𝗮𝗿𝗲 𝗮𝗻𝗱 𝗰𝗼𝗺𝗺𝗲𝗻𝘁!
Join a growing community of Data Professionals by subscribing to my 𝗡𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿.
What are the most common 𝗨𝘀𝗲 𝗖𝗮𝘀𝗲𝘀 𝗳𝗼𝗿 𝗞𝗮𝗳𝗸𝗮?
We have covered lots of concepts around Kafka already. But what are the most common use cases for The System that you are very likely to run into as a Data Engineer?
𝗟𝗲𝘁’𝘀 𝘁𝗮𝗸𝗲 𝗮 𝗰𝗹𝗼𝘀𝗲𝗿 𝗹𝗼𝗼𝗸:
𝗪𝗲𝗯𝘀𝗶𝘁𝗲 𝗔𝗰𝘁𝗶𝘃𝗶𝘁𝘆 𝗧𝗿𝗮𝗰𝗸𝗶𝗻𝗴.
➡️ The Original use case for Kafka by LinkedIn.
➡️ Events happening in the website like page views, conversions etc. are sent via a Gateway and piped to Kafka Topics.
➡️ These events are forwarded to the downstream Analytical systems or processed in Real Time.
➡️ Kafka is used as an initial buffer as the Data amounts are usually big and Kafka guarantees no message loss due to its replication mechanisms.
𝗗𝗮𝘁𝗮𝗯𝗮𝘀𝗲 𝗥𝗲𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻.
➡️ Database Commit log is piped to a Kafka topic.
➡️ The committed messages are executed against a new Database in the same order.
➡️ Database replica is created.
𝗟𝗼𝗴/𝗠𝗲𝘁𝗿𝗶𝗰𝘀 𝗔𝗴𝗴𝗿𝗲𝗴𝗮𝘁𝗶𝗼𝗻.
➡️ Kafka is used for centralized Log and Metrics collection.
➡️ Daemons like FluentD are deployed in servers or containers together with the Applications to be monitored.
➡️ Applications send their Logs/Metrics to the Daemons.
➡️ The Daemons pipe Logs/Metrics to a Kafka Topic.
➡️ Logs/Metrics are delivered downstream to storages like ElasticSearch or InfluxDB for Log/Metrics discovery respectively.
➡️ This is also how you would track your IoT Fleets.
𝗦𝘁𝗿𝗲𝗮𝗺 𝗣𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴.
➡️ This is usually coupled with ingestion mechanisms already covered.
➡️ Instead of piping Data to a certain storage downstream we mount a Stream Processing Framework on top of Kafka Topics.
➡️ The Data is filtered, enriched and then piped to the downstream systems to be further used according to the use case.
➡️ This is also where one would be running Machine Learning Models embedded into a Stream Processing Application.
𝗠𝗲𝘀𝘀𝗮𝗴𝗶𝗻𝗴.
➡️ Kafka can be used as a replacement for more traditional messaging brokers like RabbitMQ.
➡️ Kafka has better durability guarantees and is easier to configure for several separate Consumer Groups to consume from the same Topic.
❗️Having said this - always consider the complexity you are bringing with introduction of a Distributed System. Sometimes it is better to just use traditional frameworks.
--------
Follow me to upskill in #MLOps, #MachineLearning, #DataEngineering, #DataScience and overall #Data space.
Also hit 🔔to stay notified about new content.
𝗗𝗼𝗻’𝘁 𝗳𝗼𝗿𝗴𝗲𝘁 𝘁𝗼 𝗹𝗶𝗸𝗲 👍, 𝘀𝗵𝗮𝗿𝗲 𝗮𝗻𝗱 𝗰𝗼𝗺𝗺𝗲𝗻𝘁!
Join a growing community of Data Professionals by subscribing to my 𝗡𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿.
20 Terraform Best Practices to Improve your TF workflow
Explore best practices for managing #Infrastructurea Code (#IaC) with #Terraform. Terraform enables us to safely and predictably apply changes to our infrastructure.
👀https://t.co/0KgYYKszDO #DevOps
What is the TPM role (Technical Program Manager)?
Sometimes a picture is worth a thousand words. So here's a picture:
And here's more than a thousand words on the same: https://t.co/pCODtpa5vS
❗️Wrote something new today:
Event-driven architectures vs. event-based compute in serverless applications.
I see these terms used interchangeably a lot, but they are different, and the implications are important!
https://t.co/MRBtdamY76