Hello guys 👋 ,
My name is MIcheal, but u can call me mikkyCode, and I am a full stack software engineer, a founder and a 3rd year studying CS at Lagos State University, and I love building cool stuffs and shipping useful products that people use..let's connect 💯
Incredible work by @lucasmaes_ and the team! Stable, end-to-end JEPA training from raw pixels is a massive leap forward.
I took the 15M parameter LeWorldModel, trained it from scratch, and pushed the local inference speed to the absolute limit.
The Engineering Pipeline:
• Cloud Training: 100k steps on the 43GB PushT dataset using an H100.
• The Bottleneck: To run local MPC planning on a consumer gaming GPU, the native PyTorch Epps-Pulley math for the SIGReg loss was memory-bound.
•I wrote a custom Warp-Reduction CUDA kernel, fusing the mean, variance, and hinge-loss calculations directly into the SM registers.
• The Result: Crushed execution latency from 6.6ms down to 92µs (a 72.4x speedup) and achieved an exact 86.0% zero-shot success rate.
What This Means for Robotics:
Giving a robot a "World Model"—a brain that can actually imagine and predict the physics of its environment before making a move—used to require massive, expensive compute. By optimizing the underlying CUDA math, this just proved these forward-thinking robotic brains can run smoothly in real-time on a standard consumer laptop. Fast, cheap, and accessible robotic AI is here.
Custom CUDA kernel & local-to-cloud pipeline available in my fork here:
https://t.co/wZwHawIfFS
DeepSeek just dropped a banger paper to wrap up 2025
"mHC: Manifold-Constrained Hyper-Connections"
Hyper-Connections turn the single residual “highway” in transformers into n parallel lanes, and each layer learns how to shuffle and share signal between lanes.
But if each layer can arbitrarily amplify or shrink lanes, the product of those shuffles across depth makes signals/gradients blow up or fade out.
So they force each shuffle to be mass-conserving: a doubly stochastic matrix (nonnegative, every row/column sums to 1). Each layer can only redistribute signal across lanes, not create or destroy it, so the deep skip-path stays stable while features still mix!
with n=4 it adds ~6.7% training time, but cuts final loss by ~0.02, and keeps worst-case backward gain ~1.6 (vs ~3000 without the constraint), with consistent benchmark wins across the board
I got the chance to personally pitch our startup , WakaMate Ai to the CEO of MUST Company, Cha JuHun, alongside Amani Kanu and their team — all the way from South Korea 🇰🇷. This opportunity wouldn’t have happened without @RemoStart Ai-Labas. So Glad to be among the first startups
🚨👀 | BREAKING:
The Glazers are considering selling Man Utd before 2027 to maximize returns!
Potential buyers include Qatar's Sheikh Jassim, UAE consortia, Saudi interests despite denials, and private equity firms.
[@TheAthleticFC]