@DevRico003@mmastrac I agree. I don't understand the fuss around "aggregate tps". Is there any real use or is it just click bait tactics? 9 tps sucks if you are trying to do any real work.
GIVE A FLY A FISH π
"Give a man a fish, he eats for a day. Teach a man to fish, he eats for a lifetime."
I trained a #fruitfly#flybrain (connectome-constrained neural net) to play Doom. First I directed my #hermesagent using #glm5.3-flash to train the brain and check back in my when we broke certain thresholds. This was moderately successful: 19.3s survival, 23 kills max. Fine.
Then I taught the brain to TEACH ITSELF by playing. Play a round, review what worked, train on that, play again, repeat β an LSTM with working memory + a promotion gate that only accepts genuinely better play : 23 kills (up from 18.2), 39.49 score. First self-taught champion.
A brain you train is good. A brain that trains itself is better.
NATURES VERSION OF CONTINUOUS LEARNING.
25+ hrs continuous self-play, zero humans in the loop.
LSTM film above: arena (left), fly's eye (top right), brain activity (bottom right).
I did notice that the brain is only aiming for far away enemies not close by ones so I'm working on that now.
My #fruitfly brain #doomfly is getting better after upgraded my eyes to #maleCNS and wired it all in. Doesn't sound like much but my champion fly is surviving 20sec. I have some more upgrades on the way...
That said 90% of my training has been locally on my #rtx3090 with #hermesagent with glm5.3-flash as my orchestrator
#flywire #flyvis
My #fruitfly brain plays doom! #gymnasium carRacer was a good challenge, #doomfly seems like a perfect fit! First run with no extra training yet π
Reinforcement Learning (RL) did not work for carRacer but will definitely apply here.
High level:
- Doom (doomfly, Alex Wormuth): mapped each frame into the fly's sensory neurons, then trained using a local 3-factor learning rule on Kenyon-cellβMBON synapses β the mushroom body is the fly's actual associative-learning center. Reward pulses were delivered biologically: damage in-game β dopamine pulse to PPL101 cells. That's essentially RL, but implemented as biologically-local plasticity inside the brain rather than backprop on an MLP.
- MaleCNS changed the game: the new connectome includes the ventral nerve cord (the fly's spinal cord) β the viral wave's motor outputs come from there, and it's why Doom/Mario/Beat Saber runs have real motor grounding.
- Other projects validated the same pattern: hold the wiring fixed, train only connection gains + input projection + readout (SuperTruth), or pretrain a connectome policy by imitating an MLP expert then fine-tune with RL (whole-brain graph model work)
Making progress with my fruitfly #flybrain and #openai#carracer . Hard to tell if the #connectome are correct or my harness fabricated this but, progress!
#flybrain Breakthrough! Not sure if it's cheating. My fly racer keeps driving the wrong direction after a spinout to so added a U-turn guard that locks the brakes until it faces the right way to continue. Haven't found a way that works yet, even after 100 training iterations.
I taught a 45k-neuron spiking #flybrain to drive a race car. Getting laps was the easy part. Keeping it going the right direction was not.
What failed:
β’ Wider view β 3/20 laps (broke its vision)
β’ Physics speed teacher (brake at apex) β 7/20 (fought the brain's own cornering dwell)
β’ Expert takeover when facing backward β laps fell to 9/20
β’ Reward training: "grass = race canceled" β 3/20 (196k steps, real signal, still degraded the skill)
β’ Teaching it to pivot out of wrong-way moments β pivots fired mid-corner, grass 7%β36%
What worked: plain imitation, then extra hairpin practice β 15/20 laps, 7% grass, textbook slow-in-fast-out emerged on its own.
The catch: it still drives the wrong way 41% of the time. The reason is the interesting part β at hairpins its wedge-shaped view can't tell "track ahead" from "track behind," and its recovery maneuver legitimately sweeps through backward angles.
Laps are solved. Direction is a perception problem, not a control problem. More to come...
#machinelearning #ml #ai
Mind blown, this is real! I trianed #flybrain myself and it works!
My carRacer isn't great yet but follow along as I keep training.
Using #hermesagent /glm5.3-flash as my orchestrator for #flywire and #flyvis. #machinelearning
(2/2) The honest numbers: it's still rare (~1 in 25 runs of this policy completes a lap) and it's slow. The official 32-episode benchmark is +150.4 mean with all 32 surviving the eval window. Replays of laps diverge β GPU nondeterminism in the brain's forward pass β so the lap video below was recorded LIVE during the run, not replayed.
>
WHAT'S ON THE VIDEO (attached)
>
Left panel: top-down plan view of the track β the cyan trail is the car's full path, the yellow arrow is the car.
Right panel: the fly's actual vision at that exact moment β the 26Γ27 hexal stimulus, 721 pixels, the only thing the brain ever sees. Red blocks are curbs, white is road, dark is grass.
>
Watch them together and you can see WHY the car drives the way it does: every hesitation and loop in the plan view is a moment the fly's vision lost both road lanes at once.
>
WHAT'S NEXT
>
1. Native T4/T5 motion features β replace the static block grid with the direction-selective motion signals the fly actually computes. The hairpin ambiguity is precisely the problem motion-sensitive neurons evolved to solve.
2. Lap-rate consolidation β the near-misses (97%, 99%) all die in the same endgame pocket. Targeted DAgger on stall recoveries.
3. The speed/lap-rate curve β how fast can this brain complete a lap?
>
Full training summary with every dead end documented is on the way.
>
If you build something with agents driving real neurobiology β or you're the reason I found this rabbit hole, @alright_mark β I want to see it. πͺ°π
---
Got inspired by the videos @alright_mark has been posting about his trained #flybrain driving a car and decided to try myself. It's hard to know what's real what's satire sometimes but this is the real deal. I pointed my #hermesagent, running glm5.3-flash, at the FlyWire connectome and told it: train a fly brain to drive a car. Here's what happened. π§΅
>
THE SUBSTRATE
>
The brain is a FlyWire connectome: 45,669 neurons, 1.5 million synapses, pretrained, and completely frozen. We never touch a weight inside it. Every timestep, the racing game's 96Γ96 pixel feed gets compressed into a 26Γ27 "hexal" grid β 721 values total. That blocky wedge is the fly's entire visual world: no map, no coordinates, no speedometer. Just 721 photoreceptors looking forward.
>
A tiny policy network reads the brain's neural activity (65 central cells + 8 motion-sensitive T4/T5 cell populations = 73 features) and writes exactly one thing: steering. That's the whole loop. Vision β brain β spikes β steering.
>
THE LADDER
>
You can't go from a frozen fly brain to CarRacing in one jump, so we built a training ladder:
>
β’ CartPole β the brain balancing a pole: random β 23 β 67.9 after training
β’ Lane centering with optic flow β a straight road: near-oracle performance (468, oracle β 497)
β’ Curved lanes β five failed attempts, until a one-line change to the action space unlocked it (346.1)
β’ CarRacing β the real game, real physics, real curbs
>
CAR RACING: THE DEAD END
>
First lesson, the hard way: standard RL fails here. PPO trained from scratch for 200M steps. Best result: β10.9. The fly would not drive.
>
But here's the experiment that changed everything: I took a simple hand-written controller (steer β the road's position in the fly's view, with P and D terms) and behavior-cloned it into a tiny policy. No reward signal at all, just copying. It scored +43.5 instantly.
>
That single result rewrote the project. The brain's features CAN support driving. The bottleneck was never the representation β it was the training paradigm.
>
DAgger
>
PPO fine-tuning kept destroying the cloned policy (three different configs, same monotone degradation). The fix was DAgger: let the policy drive, have the controller label every state it visits β including its own mistakes, which a hand-written expert can always do β then refit on the aggregate. No collapse, steady gains, and an official benchmark of +99.2 across 32 held-out episodes.
>
THE WALL
>
Then we hit the wall that ate two full days: a hairpin where every policy β clone or expert β stalled. Same section of track, every seed, every speed setting. Coverage numbers looked consistent enough to conclude the section was simply "impossible" for this sensor.
>
That conclusion was wrong, and the instrumentation proved it: at the stall point, the road was VISIBLE in 61 out of 61 recorded steps.
>
The real failure: bistability. At a hairpin, two road segments share the fly's retina at once β the incoming lane and the outgoing lane. The road-position estimate ping-pongs between them, +0.6 to β0.7 per frame. The policy slams steer left, then right, then left, and the car orbits a pocket forever.
>
Not blindness. Ambiguity.
>
THE BREAKTHROUGH
>
The fix wasn't a better sensor. It was SLOWING DOWN.
>
At cruise gas 0.15 β a crawl β the ambiguity resolves before the car arrives. Same brain, same frozen weights, same tiny policy architecture. Teacher ring coverage jumped from ~30% of the track to 100%. Two seeds completed full laps.
>
Then we re-ran the imitation pipeline on the slow expert: BC + three DAgger rounds β a new policy that could do it too.
>
THE LAP
>
A complete racing lap, driven end-to-end by a frozen fly brain:
β’ 276 of 276 track tiles A complete racing lap, driven end-to-end by a frozen fly brain:
β’ 276 of 276 track tiles
β’ ~5,500 steps
β’ every steering decision made through the connectome
>(1/2)
@googlegemma Real use case with great results: I parallelized this gemma-4 model at work to speed up the process of pulling metadata from decades worth of calc and drawing PDFs we moved into a blob structure. Thanks Gemma, this would have been a lot more expensive to use frontier api
If big scary #mythos is similar to #claude#opus4.8 the world has nothing to fear. So many errors and recoding with opus. Seriously @AnthropicAI what the heck!?!