Bacalhau is an open-source platform for fast and distributed computation that enables users to run compute jobs where the data is generated and stored.
If you:
- Are knee-deep in pipelines & general data chaos.
- Use existing cloud/data platforms (think Redshift, Databricks, Splunk, etc).
... please click below!
👉🏻 https://t.co/eHzssGlmqD
Thanks for helping us shape the future of data - no duct tape required.
Hiya all 👋🏻 -
We’re cooking up some new features and we want your spicy data opinions. If you’ve got pipelines to wrangle and bills that make you weep, you’re exactly who we need (and yes, we'll pay (or donate to a charity of your choice!)).
Code runs. Results appear.
- No keys to manage
- No storage config to maintain
- No hunting for outputs
Bacalhau 1.8 makes storage feel like it should: invisible.
https://t.co/VWEw6xMmYF
We’ve supercharged daemon jobs in the latest Bacalhau release—rebuilt the orchestration logic to be smarter, faster, and more responsive.
What it means for you:
✅ Instant deployment to new nodes as they join
✅ Stronger reliability—cluster-wide coverage without fail
✅ Effortless scaling—daemons grow with your cluster automatically
And it’s all fully backward compatible.
Read the deep dive → https://t.co/2YWkP70eK0
Remember j-f47ac10b-58cc-4372-a567-0e02b2c3d479?
Of course not.
Bacalhau 1.8 lets you name your jobs, rerun them, version and even diff your jobs before you run them.
Bacalhau v1.8.0 is live.
We’ve made distributed computing radically more usable and radically more cost-efficient:
• 📊 Native Splunk integration: Slash logging costs by up to 80%
• 🏷️ Name-based jobs: Human-readable jobs, not UUIDs
• 🛠️ Enhanced daemon reliability for services at scale
Built for the platform engineers actually running modern infra.
As Data + AI Summit rolls on—with more features, more data, and more spend—we’re focused on something else:
You should be paying less.
Expanso cuts data infrastructure costs by up to 80%. Without disrupting your stack.
You need proof, so we let you run it. Anywhere, at any scale.
Try Expanso on 10 servers.
Or 10,000. See how much you save in 60 days.
We’re moving toward a world where data sovereignty is table stakes.
If you’re building data infra, you need to think globally - and legally.
This new Bacalhau tutorial walks through setting up multi-region compute and anonymization using @Microsoft Presidio, while keeping compliance front and center.
Open source. Real code. Built for real-world constraints.
Cross-border data flows are tricky.
We wanted to see if we could build a real compliant pipeline from scratch - with open-source tools.
✅ Generate sensitive data in the EU
✅ Anonymize it w/ @Microsoft Presidio
✅ Send to US for processing
@BacalhauProject handles the orchestration. Presidio handles privacy.
Here’s how it all works (Code included!) → https://t.co/pZyfHHiOQ4
Streamlining Bacalhau Development with the Power of Docker-in-Docker.
Our latest blog explores a practical solution using Docker Compose and Docker-in-Docker to create a self-contained, local Bacalhau environment.
Learn how to get your local Bacalhau instance running quickly and efficiently.👇
.@ExpansoIO has created @BacalhauProject, an open source architecture that allows users to run compute jobs where the data is generated and stored.
https://t.co/qbdiUlsn92
Want to set up an open-source, distributed ML pipeline that respects geographical and regulatory restrictions and runs compute in the same location as your data?
This post gets you started with Bacalhau to set up nodes in three different regions, analyze data, all while respecting data sovereignty.