Enjoying your summer vacations? Here's a #DataEngineeringRoadmap that helps you learn the basics to be a data engineer. The schedule is three weeks; it's very ambitious, but you can schedule at your own pace.
Yesterday, CloudFlare dropped a bomb that I believe may change the future of Lakehouse storage.
R2 + Iceberg should become the de-facto choice for hybrid and multi-cloud data lakehouse architectures.
Here's why it may break the cloud monopoly 🧵
1/ There are an overwhelming number of options for data engineering today. Here's today's choices at @OSObserver :
For cloud infrastructure: @googlecloud
For data orchestration: @dagster
For data transformations: @SQLMesh
For data ingest: @dltHub
For data lake tables: @ApacheIceberg
For OLAP database: @ClickHouseDB
For OLTP database: @supabase
For GraphQL APIs: @HasuraHQ
For low-code frontend builder: @plasmicapp
For frontend framework: @nextjs
For frontend hosting: @vercel
Si eres programador necesitas conocer esta web.
Decenas de hojas de referencia rápida y trucos para lenguajes de programación y herramientas de desarrollo.
quickref․me
Today, I'm sharing «The Data Engineering Toolkit», a set of tools and fundamentals for any data engineer who is getting started or wants to understand the role of a modern data engineer.
Maybe you remember the book: The Data Warehouse Toolkit. First released in 1996, it holds strong to this day for all related to data modeling. In the article (https://t.co/TLAyaWaIn5), I take a step back and focus on the essential toolset and knowledge working as a data engineer.
We'll discuss OS choices, Linux commands, and how to leverage command-line tools for simple and complex data tasks on local machines or servers. We'll explore developer productivity boosters like modern IDEs, Codespaces, and Notebooks, along with essential programming languages for data engineering.
Como es tradición, he escrito mis recortes del 2024, una colección de enlaces, con sus respectivos comentarios de lo que ma ha gustado este año --->
RBAC (Role-Based Access Control) is key for managing permissions in Snowflake, but scaling it with SQL scripts can get complex. Titan Core simplifies this process.
Querying the whole social network from @duckdb? Yes, it's possible. Check out the exciting discussion linked.
@TobiM set up an endpoint we can query with a few lines. Currently there's data since last Friday: 12.6 million posts and 165 million from Jetstream.
Really cool stuff is happening. Follow along on: https://t.co/pAtHhG3WI1 (more to come soon).
Tired of BI dashboards that are slow to build and prone to errors?
There's a better way.
Here's how BI-as-Code tools like @Evidence_dev, @RillData, and @streamlit—powered by @duckdb—are changing the game.
Using @duckdb to analyze open data from Chicago has been a game-changer: extract with @dltHub , store in R2 as Parquet, query in DuckDB, transform with @dbt_labs , and publish the database in R2 for remote access—all from my laptop!
Hoy me he acordado de este tuit @xoelipedes.
No sé si llegaste a usar @dltHub, pero con ingestr puedes hacer justo lo que buscabas en tu post original. Es un CLI que por debajo usa dlt y SQLAlchemy.
cc @franloza93 😃
https://t.co/o4yBi8UPgT
Cual es la mejor forma de sincronizar los datos de Stripe con Postgres? No me quiero picar nada ni pagar nada si es posible
Idealmente algo que cree las tablas que haga falta automáticamente y que con un comando pueda hacer sync, si es open source mejor
Today I tested how to handle changes in data schema with @dltHub. It manages tables, columns, and data types based on simple settings. It’s easy, and it automatically saves schema versions. This will save me some extra work in the future. Here’s an example 👇