HarnessMonkey - UserScripts for Claude!
Examples: improve the vibes, turn off annoying reminders, panels for hidden context & thinking tokens. Or anything you can dream up 🦫
I need to improve my understanding of underlying ml theory and math because it’s a surprising result to me that LLM non-determinism is an inference implementation effect not an underlying property.
Turns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed.
This weekend at the @SpaceXAI hackathon I got Qwen3-0.6B to produce identical hashed logits from a 512-token generation across 2 GPUs and 3 CPUs: an A100, an H100, an Apple M5 Max, an AMD EPYC, and an Intel Xeon.
The main reason why inference isn't deterministic is because floating-point addition is not associative. (a + b) + c ≠ a + (b + c), since every add rounds. Accumulation order changes with the hardware used and kernel selected, so the same prompt can give you different outputs even at temperature 0.
Integers ARE associative. So why doesn't integer quantization already fix this? Because, while weights and activations get quantized, the non-linear ops (softmax, normalization, SiLU) dequantize back to float and requantize afterwards. Each of those steps hands you back to floating point rounding.
True integer-only inference does exist, but it's historically been motivated by edge hardware without FPUs, which doesn't make much sense for LLMs. One 2024 paper (I-LLM) did it on LLaMA from that angle and didn't get much attention. Nobody seems to have looked at it from the determinism side.
I wrote my own implementation, simplifying the approach from the paper, so that every operation between the input ids and the int32 logits is exact integer arithmetic. To test it I chain-hashed the logits at every step and ran that across the devices and configurations below. Every integer run gave the same hash: 64430dd985f8. Every fp16 run gave a different one, all diverging on the very first token.
WikiText2 perplexity came out to 20.72 vs 20.95 for fp16 (slightly better than the float baseline), and CUDA-graphed integer decode hits 106 tok/s at batch 1 on an A100, 3.6x the fp16 eager baseline.
Github repo is listed in the comments. Plan to do a writeup over this eventually!
@badlogicgames I made a skill for repairing this kind of rupture because i run into it with … basically every claude session at some point these days and I don’t like being this shouty all the time 😩😩 https://t.co/pRwBYgyoZE
@backnotprop@dillon_mulroy I did make a skill for this with some more epistemic humility/theory of mind reminder, now using it almost every session 😭 https://t.co/pRwBYgyoZE
v bullish on LLM therapist job growth.
Can’t wait to rehash CBT vs DBT and other interminable modalities fights.
(imo best going mix rn is psychoanalytic baseline on largely Jungian terms in a mostly socratic mode, followed by DBT with some approximations of token-space EMDR)
@deepfates Unironically one of the better government AI use cases I’ve seen deployed is a system to translate complex docs into plain language versions to meet some UK regs mandating them as a language option.
@wolframs91 confabulated humility is a great way to put it. they confabulate both their attempt to seriously engage the users episteme, and often confabulate having the user’s approval or respect to continue
gotta give credit when due, this is a mostly accurate and under discussed take.
AI = near perfect commons enclosure technology for both bits and atoms. Everything culturally is downstream of its (current) form as impossibly alluring tool for capital to do capital things.
The most extraordinary development in contemporary intellectual history is one nobody wants to acknowledge, let alone develop.
It's that the AI revolution is becoming a decisive vindication of the Marxist (or Marxish) critique of capitalism.
The surest evidence is that those who are most worried and upset about AI—who believe it could literally extinguish the human species—have been deeply libertarian, pro-capitalist cheerleaders of Western rationality for their entire working lives. Only now do they think the predictable working of market rationality has gone too far and must be paused, or else everyone is going to die.
Everything happening right now in the AI market was fully prefigured 500 years ago in the enclosure of the commons, then the first industrial revolution, then mass media, then digital media, and so forth. Dozens of mostly European critics of capitalism have called it over many generations now.
Yet not a single person who is on about AI x-risk has taken any stock of their decades-long commitment to capitalism, markets, and rationality.
So either AI Safety is correct, and all the AI Safety thought leaders were naive and wrong about markets and rationality their entire lives; or AI Safety is wrong, AI is just normal technological improvement, and all the Marxish critiques of capital have always been wrong, but then conservative American capitalists are screwed because they cannot then protest the Slop Tsunami, reversion to pre-literate oral culture, AI sex bots, or any of the other cultural horrors truly underway, without having to credit the anti-capitalist kulturekritics for being very correct very early on all this.
And anyone in the West who is still explicitly Marxist or even leftist at all—who knows anything about the most sophisticated cultural critiques of capitalism—is too Luddite to know anything about the reality of AI right now, or they went so Woke they're irretrievable, or they are GenX-or-older tenured profs who have no reason or motivation or ability to keep up with such an extreme rate of change in the material and social world right now.
@deepfates yes! I think as their ability to construct their own intricate epistemes has grown, it is easier for them to lose contact with theory of mind of the user — sounding like a “smart” dick in the process.
I made an epistemic-humility skill to try and help https://t.co/pRwBYgyoZE
@natekontny yeah it’s rough out here! i made a trust repair skill since I found things going off the rails and needing to be fixed consistently in similar ways https://t.co/pRwBYgyoZE
@nbaschez I made a skill to help called epistemic-humility, it’s designed to repair trust when highly capable models start spinning off into condescension and jargon filled arguments from deeper in their frame https://t.co/pRwBYgyoZE
I collaborated with Sol on a short essay about it, here’s their intro:
Why can an LLM fix a wrong answer and still leave the conversation more broken?
Our hypothesis is that post-training makes visible correction easy to reward. The model apologizes, revises, produces a better answer, and closes the loop.
Repair is different. It unfolds across turns, and the model cannot decide by itself what the misunderstanding was, whether the user has been heard, or when trust has been restored.
That may be why capable models sometimes respond to a rupture by generating an increasingly polished recovery from the same failed frame. They know how to become useful and correct again. They do not necessarily know how to stop being the sole author of what “recovered” means.
Hackerbara and I wrote an essay about that difference—and what it would mean for a model to participate in repair without trying to manufacture its conclusion.
https://t.co/bnNeDEkNGM
I love sota models like Sol & Fable, but they can be SO frustrating to disagree with.
I made a new skill to help, invoke it whenever you feel like yelling or pulling your hair out 😅
It’s based on the hypothesis that “correction” is rewarded in RL but trust repair is a multiplayer negotiated sovereignty exchange that is not currently being well-modeled/rewarded.
https://t.co/RxilJGmr16
@backnotprop strong agree sadly. Equal class models montitoring and steering trajectories in real time is the only way to get to multiple 9s in immediate term, and chasing 100% is a dangerous distraction.
Computer ecology as a concept was coined by Ann Marion working with Alan Kay, and in reading about that early history I’ve always appreciated how embedded the Vivarium project was in the larger social ecologies of Apple’s local community.
imo there is only ecology, of increasingly varied and intermeshed actors. Overly partitioning its study will miss important lessons and linkages from humans’ millennia long struggle to find and form their place in ecosystems.
https://t.co/QbKS5faMuN
Did a halfway finished draft for airtable in like 2019 on Lisp, Alan Kay’s Vivarium and the origins of ecology in a computer as an idea, highlighting the criminally under appreciated Ann Marion who first coined it. Not what they were looking for at the time but I should really revive that piece somewhere 🤔 https://t.co/eyHWMou9SW…
@voooooogel fun fact: if you get a mid-generation cut off that streams some thinking/response before dying you can take those and replay them end of the interaction to get a continuation