Hi 👋, I am Ayo a Data Engineer building optimized data systems that enables organization makea data driven decisions.
Currently expanding my expertise in modern data engineering practices, including orchestration, data warehousing, cloud platforms, and data governance.
@LeoOliemans91 The post-fix window is where things get risky.
My check would be:
1. Re-read after the write.
2. Compare row counts, hashes, or key aggregates.
3. Mark recovered only if they match
Zero-row writes and race conditions are too common to trust the first attempt.
AI agents + Data engineering is a serious combo....imagine an agent detecting a failed pipeline, investigating the logs, identifying the bad dataset, and opening a fix automatically.
But your revenue dashboard is now wrong.
A simple data-quality check could catch it:
order_amount IS NOT NULL
order_amount >= 0
This is why a successful pipeline ≠ trustworthy data.
Have tax questions?
We have answers!
Join us at The Bridge for 90 minutes of practical answers and jargon free explanation.
CEO, C- Suite exec, Accountant, Tax & Finance managers
Everyone gets a seat.
With idempotency, you use a unique order_id and an UPSERT/merge strategy:
Order #1025 → ₦50,000 ✅
Run it once or run it five times, the final state remains the same.
That’s idempotency:
Safe retries without corrupting your data.
Imagine a pipeline processes yesterday’s orders and inserts:
"Order #1025 → ₦50,000"
The job fails after inserting the record, then retries.
Without idempotency:
"Order #1025 → ₦50,000"
"Order #1025 → ₦50,000" ❌
Now your revenue is overstated.
Imagine a pipeline processes yesterday’s orders and inserts:
"Order #1025 → ₦50,000"
The job fails after inserting the record, then retries.
Without idempotency:
"Order #1025 → ₦50,000"
"Order #1025 → ₦50,000" ❌
Now your revenue is overstated.
This matters when you have:
• Retries
• Failed jobs
• Duplicate events
• Backfills
• Scheduled runs
A pipeline shouldn’t just move data.
It should be designed to recover safely when things go wrong.
This matters when you have:
• Retries
• Failed jobs
• Duplicate events
• Backfills
• Scheduled runs
A pipeline shouldn’t just move data.
It should be designed to recover safely when things go wrong.
If the same job runs twice and creates duplicate records, you have a problem.
That’s where idempotency comes in.
An idempotent pipeline can safely process the same input multiple times without changing the final result.
If the same job runs twice and creates duplicate records, you have a problem.
That’s where idempotency comes in.
An idempotent pipeline can safely process the same input multiple times without changing the final result.
It’s becoming less about simply building pipelines and more about building data systems that AI agents can actually reason over, interact with, and use effectively.
It’s becoming less about simply building pipelines and more about building data systems that AI agents can actually reason over, interact with, and use effectively.
What stood out to me is how the rise of agentic systems is pushing us to rethink the relationship between AI, data, databases, and the architecture that connects them.
What stood out to me is how the rise of agentic systems is pushing us to rethink the relationship between AI, data, databases, and the architecture that connects them.
AI agents are changing more than how we use data. They’re changing how we architect it.
Just listened to “From RAG to Relational: How Agentic AI is Reshaping Data Architecture.”
AI agents are changing more than how we use data. They’re changing how we architect it.
Just listened to “From RAG to Relational: How Agentic AI is Reshaping Data Architecture.”
The rollout of Rev 360 has begun, and it’s happening in phases.
First, medium and emerging taxpayers are being onboarded.
Then, large taxpayers and government institutions will follow.