Introducing brae: GPU-native computational fluid dynamics. 🌊
The entire simulation runs on one GPU, no matrix shuffled back to the CPU every iteration.
~5× faster than GPU-accelerated OpenFOAM, on the same GPU. Validated to <1%.
🌐 https://t.co/hNi7lL9gm5
Brae now runs OpenFOAM’s rhoSimpleFoam fully on CUDA 🚀
Across cases from 15k to 24M cells on one GH200 outperforms all 64 Grace CPU cores.
At 10M cells / 100 SIMPLE iterations:
- GH200 GPU: 109s
- 64 Grace CPU cores: 921s
Matches OpenFOAM to 7e-08.
A port, not a reimplementation.
A transonic 180° U-bend
Mach 0.999 on the inner radius, 960K. converged in 3.4 seconds.
112,000 cells. Brae: OpenFOAM’s rhoSimpleFoam, ported to CUDA. The run stopped on the tutorial’s own residualControl at iteration 154 - not a hand-picked iteration count.
Brae now runs OpenFOAM’s rhoSimpleFoam fully on CUDA 🚀
Across cases from 15k to 24M cells on one GH200 outperforms all 64 Grace CPU cores.
At 10M cells / 100 SIMPLE iterations:
- GH200 GPU: 109s
- 64 Grace CPU cores: 921s
Matches OpenFOAM to 7e-08.
A port, not a reimplementation.
@YiCasillas Brae goes 554 ns/cell-iter at 14k down to 74.5 at 24M. That 7.4x is exactly the fixed overhead you're describing: at 14k it's launch-bound, by ~900k it's flat.
Profiling puts it at ~1,150 kernels/iteration launching at <=32 blocks.
Brae now runs OpenFOAM’s rhoSimpleFoam fully on CUDA 🚀
Across cases from 15k to 24M cells on one GH200 outperforms all 64 Grace CPU cores.
At 10M cells / 100 SIMPLE iterations:
- GH200 GPU: 109s
- 64 Grace CPU cores: 921s
Matches OpenFOAM to 7e-08.
A port, not a reimplementation.
Our previous SPUMA benchmark showed brae running 4.5× faster on a transient PimpleFoam case.
This new rhoSimpleFoam port suite puts brae at more than 20× 🚀 faster than SPUMA by geometric mean across 6 cases, with up to 136× on the 1M-cell case.
The lead is getting much bigger.
93 million $SIMD tokens locked again until January 2027.
We're here to build: GPU-native simulation infrastructure and real utility for the SIMD ecosystem - starting soon.
Live Roadmap: https://t.co/HCL8ADBvrG
SIMD on top!
Brae is heading to the supercomputer 🚀
We’ve secured access to LuxProvide’s MeluXina platform, including NVIDIA A100s GPU infrastructure.
We’ll use this pilot to accelerate BRAE’s development, benchmark it in an HPC environment, and advance our goal of building large, multiscale simulations.
Thank you to @luxprovide for the support!
You only need GPU to run the simulation, not 1000 CPU core anymore.
2.78M cells · kOmega SST-IDDES · Re 1.6×10⁷ · 4,800 timesteps in less than 2 hours on one GH200.