In 2006, Edward Feigenbaum spoke of an “Einstein in a box”. How does that ambition look alongside today’s AI discoveries? Our educational series @UCompsci connects the pioneers’ ideas with the technology, questions and debates of today.
https://t.co/uOmVJP6ZGp
Back to the roots: Scale-conditioned structure-based closure for
homogeneous turbulence: Ray–Column Interacting Particle
Representation Mode.
Physics of Fluids 38, 075186 (2026)
https://t.co/qYaTfmc5Rc
The George Gilder Technology Report : “The End of the Microchip” https://t.co/4BInTzKMKh
My take: “The Beginning of Locality”.
The center of gravity is shifting toward integration and locality. We are seeing wafer‑scale engines, box‑scale UMA, and rack‑scale fabrics that cut data motion instead of just adding FLOPS. Data motion dominates energy draw, so every meter of wire saved is fewer joules per result.
Across today’s AI/HPC systems, the pattern is consistent: move compute and memory closer; keep communication local. Some examples where we see the trend:
➡️ Wafer scale: Cerebras WSE‑3. One intact 300 mm wafer: 900 k AI cores, 44 GB on‑wafer SRAM, ~125 PFLOPS (BF16/FP16), ~21 PB/s on‑chip memory bandwidth, and ~214 Pbit/s (~26.8 PB/s) core‑to‑core fabric (TSMC 5 nm). CS‑3 racks scale out while keeping intra‑wafer traffic local.
➡️ Rack scale: NVIDIA GB200 NVL72. 36 Grace CPUs + 72 Blackwell GPUs in one 72‑GPU NVLink domain, delivering ~130 TB/s of in‑rack GPU communication—treating the rack as a tightly‑coupled accelerator. Blackwell itself merges two reticle‑limited dies via a 10 TB/s die‑to‑die link.
➡️ Desk scale: single‑node developer/workstation boxes.
• Apple Mac Studio (M3 Ultra): up to 512 GB unified memory at ~819 GB/s. Unified memory + MLX lets arrays live in shared memory and run on CPU or GPU without copies, and independent testing has shown large‑model inference well under 200 W.
• NVIDIA DGX Spark (GB10): 128 GB LPDDR5X (≈273 GB/s), ConnectX‑7 NIC up to 200 Gb/s, 240 W power supply (GB10 TDP ~140 W), and ≈1 PFLOP FP4 (theoretical, sparsity‑on) for local inference/fine‑tuning.
➡️ Lab scale: workstation clusters. Our Apple‑silicon (M4 Max and M3 Ultra) cluster uses unified memory (e.g. 512 GB/node, ~819 GB/s) so arrays stay in shared memory and ops hop CPU↔GPU without copies, translating to direct energy and latency gains. This isn’t academic: the 1‑trillion‑parameter Kimi K2 Thinking model (QAT INT4) recently ran natively on two M3 Ultras using MLX pipeline parallelism, sustaining ≈ 15 tokens/s with no reported quality loss (@awnihannun Awni Hannun, Nov 2025).
➡️ Mesh / card scale: Tenstorrent Blackhole. Add‑in boards with 4× QSFP‑DD 800 Gb/s ports for card‑to‑card links and memory pooling, pushing neighbor‑to‑neighbor communication on‑chip and between boards.
These are different engineering choices, but they all point to the same direction: cut data motion. This matters for both inference and training, across LLMs and physics‑informed neural networks (PINNs).
That’s why our teaching in the “HPC‑AI Synergies Lab” emphasizes locality‑first in code design, precision literacy (use reduced-FP operations when appropriate), and fabric‑aware parallelism. Students learn to reason in joules per result, not just FLOPS per chip. Our lab uses single and clustered Mac Studio M‑series nodes to bring these ideas into a setting accessible even to undergraduates.
From wafer‑scale engines to Apple M‑series unified memory, the future of performance is “locality per joule”, i.e., fewer joules to reach the target result and Kimi K2 on two M3 Ultra shows it’s already happening 🚀🚀 thanks to MLX and @awnihannun
Walking towards the future… 🚀
& the HPC-AI Synergies Lab awaits 😉 — join us to learn how you can use MLX on Apple Silicon for your AI/ML applications in Engineering & beyond. Nov. 5th is the official inauguration of the new Engineering Complex at @UCYOfficial 🎉 @awnihannun@tim_cook
Huge thanks @Korata_hiu for kicking the tires on Kourkoutas‑β! 🦎
Love hearing it works on SDXL/T2I and becomes your default vs fixed β₂.
Paper → https://t.co/FZBOLB5iQo
Code → https://t.co/AbgdWtbC3a
PyPI → https://t.co/xN2G7mvyqh
P.S. Availability diagnostics & efficiency plots coming soon, will keep you posted 👀
Happy to have contributed to the development/validation of this product. Fun memories from the collaboration. Solid engineering team on the @aptar side.
🚀 🎉Just dropped our first public release of Kourkoutas-β! An Adam-style optimiser with dynamic β₂ memory for bursty gradients.
🔗 Paper: https://t.co/E987G119nQ
🔗 Code: https://t.co/Y9Ba9dZEAS
🔗 PyPI: https://t.co/1DEv0dbMZB
🦎🌞 “we code in MLX” @awnihannun@angeloskath
CDA Mangis: Pleased to join the Stanford Alumni Association Cyprus (SAAC) dinner. Thanks to cofounders Suzi Abdel-Malak & Effie Kokkinou. SAAC promotes networking among Stanford alumni in Cyprus, strengthening the 🇺🇸🇨🇾 connection to the Stanford University family.
🥳🚀 1st installment of Mac Studios M4 Max for the “HPC+AI Synergies” teaching lab has arrived. 🚀🚀
Benefits:
👉🏻Ultra low energy footprint
👉🏻 Quiet student-friendly environment for hands on experimentation w/ parallel connectivity topologies
👉🏻Students get trained on machines they can afford on their own as professionals @awnihannun@angeloskath
I’m sooo 🦎🦎happy 🦎🦎w/ Kourkoutas-β. Completed 1st full test on training a heavy Transformer model as PDE surrogate in data-driven mode. Kourkoutas-β blows vanilla Adam out of the desert, especially for smaller training datasets & quantized models where the grads tend to be more spiky. This far exceeded my expectations. Kourkoutas-β is a version specifically designed for training Transformers. @awnihannun@angeloskath
The second Podcast 🎧 in our series dives into our efforts @UCompsci with collaborators to create a computationally efficient model of the deep lung 🫁 by taking some clever shortcuts: The cited article is https://t.co/SqVHeygCWY #AerosolScience#InhalationTherapies#LungHealth
Starting today we initiate a series of science podcasts 🎧 on X showcasing some of the most exciting research output of @UCompsci@respihub & collaborators. Thanks to Google for making this possible. These podcasts are themselves a testimony of the real value of AI 🚀🚀
🫁🌊 Collaborating with Josué Sznitman on this Annual Review was fun & rewarding. It is finally live in early release @AnnualReviews : Multiscale Modeling of Respiratory Transport Phenomena and Intersubject Variability | Annual Reviews - https://t.co/EA2CDdLfbS
Thermo Therapist with CoolProp is being released in the GPT store for my Engineering Thermo students. It explains concepts, uses a custom API to access CoolProp, solves problems & will convince you that you should ❤️ Thermo. Not responsible for any Transference issues 😂 @OpenAI
🚀🚀🚀🐢 We verified that if you mix an M1 Max w/ several M2 Ultra you can almost recover full performance of using all M2U by proportionally reducing batch size on M1, increasing it on M2, keeping total # of epochs the same. Might depend on case/size. @awnihannun@angeloskath
Some like CalDigit Thunderbolt 4 Element Hub do offer 40 Gb/s on all TB ports. But as correctly pointed out by @elegyals you can connect up to 7 MacStudio M2 Ultra using either an iPad as monitor over WiFi or other non TB monitor (over usb or hdmi)