Stop renting intelligence. I build Private Enterprise LLMs & agentic workflows for businesses that care about data privacy.
↓ Get my AI Architecture Blueprint
Most "AI features" are just an LLM call wearing a UI.
I spent time building the opposite: an AI layer for a personal notes app that knows exactly when not to call an LLM.
A thread on what disciplined AI engineering actually looks like.🧵
You’re not paying for intelligence. You’re paying for LLM calls nobody scoped.
Local writes. Reuse embeddings. Call a model only when the task is narrow enough to constrain.
Reply BLUEPRINT and I’ll send the PDF.
Or free Enterprise AI Security Blueprint - link in bio.
Most "AI features" are just an LLM call wearing a UI.
I spent time building the opposite: an AI layer for a personal notes app that knows exactly when not to call an LLM.
A thread on what disciplined AI engineering actually looks like.🧵
None of this is AI magic. It's picking the cheapest tool that actually solves the problem, sometimes a trained classifier, sometimes a tightly scoped LLM call, sometimes no model at all. That judgment is the actual skill, not the API key.
I’ve always had trouble with note-taking. I write everything down - ideas, article drafts, voice memos- but almost none of it ever made it into a finished piece. My note-taking app had turned into a graveyard. Every time I opened it, I was faced with a pile of hundreds of unprocessed notes. Nothing but stress, zero clarity.
I realized that the problem was never “where to write it down”- it was that nothing ever forced a note to turn into a solution. So I created an AI system for myself that reads a raw thought and independently transforms it into what’s needed: a task, a solution, a plan or permanently archives it.
Tomorrow a thread about how it actually works.🧵
@jun_song Saying you paused because 'it’s too dangerous' sounds way better to investors than admitting RL scaling hit diminishing returns. Safety is the ultimate PR shield.
@WiFiMoneyGuy 100% agree on orchestration over raw model size, but the verifier is still the bottleneck. If the evaluating model misses subtle edge cases, running 20 fast iterations just leads to confident compounding errors.
@gregisenberg The lazy prompt wrappers definitely died. The ones printing money today are the ones fixing messy internal workflows and plugging into actual company data. At that point, it's just real software.
@TheAhmadOsman Open the codebase of literally any top-tier paper. It’s hardcoded CUDA paths, notebook spaghetti, and zero tests. Great minds? Absolutely. Good engineers? Almost never.
@birdabo They'll release Fable 5.1, everyone will be thrilled, and they'll be blown away by its intelligence, but in two weeks they'll nerf it again and impose even worse limits to boost financial metrics ahead of the IPO.