A small thing that trips people up when they start working with LLMs: the difference between the model and the product wrapped around it.
Ask a raw model, through the API, with no system context, what today's date is. It will answer with something close to its training cutoff. Not because it's broken, but because it genuinely doesn't know. The model has no clock. It has no calendar. What it has is a frozen snapshot of the internet up to a certain point, and when you ask it "what's today's date," it's doing next token prediction against that snapshot. Its best guess is whatever date pattern appeared most often near the end of its training data.
Now ask the same underlying model through a consumer product, ChatGPT, Claude's app, whatever, and it gives you the correct date. That's not the model suddenly knowing something new. That's the product layer doing its job. Every one of these apps injects the current date into the system prompt before your message ever reaches the model. The model isn't recalling today's date, it's reading it, because the application handed it over as context.
The same logic extends further than just the date. Ask the raw model who won an election last week, or what a company's stock did this morning, and it has no path to that answer. It's frozen. But ask the same question through the product, and it can actually get it right, not because the model got smarter, but because the product gave it a tool. Web search, browsing, retrieval, whatever you want to call it, is the product layer letting the model reach outside its own frozen weights, run a query, read the result, and reason over something it was never trained on. The model isn't remembering, it's looking it up, the same way you or I would.
This distinction matters more than it looks like on the surface, especially if you're building on top of these models rather than just chatting with one.
If you're calling a model raw through an API, you own that context. No injected date, no injected location, no browsing, no idea what's "current" unless you build that pipeline yourself.
Training cutoff is not the same as capability cutoff. A model can reason perfectly well about events after its cutoff if you feed it the information, whether that's injected context or a live search result.
This is also why hallucinated "current events" answers happen. Without a grounding date or retrieved context, the model defaults to pattern completion from stale data, and it will do so confidently, because confidence isn't the same as correctness in these systems. Give it a tool to check instead of guess, and that failure mode mostly disappears.
The takeaway for anyone building products on top of LLMs: never assume the model knows what day it is, who's currently in a role, or what shipped last week. If your product needs that, you either inject it or give the model a way to go find it. The intelligence is powerful, but it is not automatically current, and conflating the two is where a lot of subtle bugs in AI products come from.
#LLMEngineering #AIProducts #MachineLearning
Spent the last few months building and rebuilding LLM routing layers. Some thoughts on abstraction that I think get lost in the hype.
Everyone wants a router. Send simple queries to a cheap model, hard ones to a frontier model, save money, look smart in the cost dashboard. Reasonable in theory. In practice, most routing logic I've seen is a classifier bolted onto a system prompt, guessing at task difficulty before the task has even been attempted. You're making a routing decision with less information than the model itself would have after one forward pass.
The abstraction layer problem is subtler and matters more long term. We wrap every provider behind a common interface so we can swap models without rewriting the app. Good instinct. But LLMs are not interchangeable behind a clean interface the way databases are. A prompt tuned for one model's quirks doesn't transfer cleanly to another. Tool calling formats differ. Context handling differs. The abstraction leaks constantly, and when it leaks, it leaks into production behavior nobody notices until a support ticket shows up.
What I've landed on, at least for now:
Treat the router as a cost optimization, not a quality guarantee. If correctness matters, don't let a routing layer make silent tradeoffs on your behalf.
Keep the abstraction thin enough that you can see through it. If your interface is hiding which model actually ran, you've traded flexibility for blindness.
Version prompts per model, not per feature. The model is part of the behavior, not an implementation detail underneath it.
None of this is a solved problem yet. The tooling is still catching up to how differently these systems actually behave under the same interface. Curious how others are handling this, especially anyone running multi-model setups at real production scale.
#LLMEngineering #AIInfrastructure #MachineLearning
Before Transformers, sequence modeling went through several architectural generations. Here's a look at how we arrived at the design underpinning every LLM in production today.
For years, two families of architectures dominated:
RNNs / LSTMs Processed text one token at a time, carrying forward a hidden state as memory. This worked reasonably well, but training was inherently sequential — you couldn't parallelize across a sentence. Long sequences also suffered from vanishing gradients, limiting the model's ability to retain context from earlier in the input.
CNNs for text Borrowed from computer vision, convolutional models captured local patterns in language using sliding filters. They trained faster than RNNs, but struggled with long-range dependencies, since a word's meaning often depends on context well outside a fixed-size window.
Attention as an add-on Around 2014-2015, researchers began augmenting RNN-based encoder-decoder models with attention mechanisms (the seq2seq era). This let models weigh which earlier tokens mattered most when predicting the next one — a meaningful improvement, but attention was still a component bolted onto a fundamentally sequential architecture.
Then, in 2017, a team at Google published a paper that reset the field:
"Attention Is All You Need" (Vaswani et al.)
The core proposal: remove recurrence and convolutions entirely, and build the model using attention alone.
The result was the Transformer:
1. Self-attention allows every token to attend to every other token simultaneously, capturing long-range dependencies without sequential bottlenecks
2. Full parallelization during training, since there's no step-by-step recurrence — a key enabler for training on internet-scale datasets
3. Positional encodings reintroduce a notion of word order, since attention itself is order-agnostic
4. Multi-head attention lets the model attend to different types of relationships in parallel ��� syntactic, semantic, coreferential, and so on
This wasn't an incremental refinement. It was a fundamental rethink of how to structure a neural network for sequence data.
It's also the direct ancestor of everything that followed. GPT, BERT, T5, Claude, LLaMA, Gemini — every modern LLM is, at its core, a stack of Transformer blocks scaled up with more data, more parameters, and improved training techniques.
One paper. One architectural bet that recurrence wasn't necessary. It reshaped the field.
What stands out in hindsight is how much simpler the winning idea was — a reminder that the highest-leverage breakthroughs often come from removing complexity, not adding it.
#MachineLearning #AI #Transformers #DeepLearning #LLM #NLP
Today I learned that the OpenAI Python client isn't just for OpenAI models
Turns out the openai Python library lets you swap in a base_url parameter and suddenly you're calling completely different LLM providers through the exact same interface.
Why does this work? The OpenAI API format became so widely adopted that other providers (Gemini, and plenty more) started exposing "OpenAI-compatible" endpoints. So instead of learning a new SDK for every model, you can often just:
1. Keep your existing openai client code
2. Point base_url at the provider's compatible endpoint
3. Swap the API key and model name
That's it. No new SDK, no rewritten integration logic.
It's a nice reminder of how powerful de facto standards are in software. OpenAI didn't just ship a product, it accidentally shipped an interface that the rest of the industry decided was easier to adopt than to compete against.
Makes me wonder what other "just build on top of the popular thing" patterns are quietly saving developers time out there.
#AI #LLM #SoftwareEngineering #OpenAI #DeveloperTools
Your backend shouldn't know your frontend exists.
Working with full stack teams for a while now, I often hear this from backend developers: "Just handle it on the frontend, mate." If only it were that simple.
A good system takes shape when the backend stops relying on the frontend to cover its gaps:
- Every list that can grow infinitely should be paginated from the backend by default.
- Any piece of information not visible on the UI screen should not be returned from the API.
- Frontend shouldn't have to "connect dots" to pull together the data it needs. It should just call the API, store the result, and render it.
Frontend already has enough on its plate — UI, state, rendering. Don't make data modeling one of them too.
#SoftwareEngineering #BackendDevelopment #FrontendDevelopment #APIDesign #SystemDesign
Today’s focus: rate limiting.
I added a Token Bucket–based rate limiter to the service.
How it works:
- Each bucket has a fixed capacity
- Tokens are refilled at a steady rate per second
- Every request consumes one token
- If no token is available:
--the request can be rejected immediately
--or wait until a token becomes available
Why Token Bucket?
- Allows short bursts while still enforcing a steady rate
- More realistic than fixed windows
- Commonly used in API gateways and distributed systems
What I liked about this addition:
- The limiter is framework-agnostic
- Works well with the existing:
--circuit breaker
--caching
--sequential/parallel execution
- Makes the health checker behave closer to a real client, not a stress tool
At this point the service has:
- caching (LRU + TTL)
- circuit breakers per dependency
- rate limiting
- memory observability
- configurable execution modes
Small project, but it’s becoming a nice playground for resilience, performance, and traffic shaping patterns.
#softwareengineering #backend #nodejs #ratelimiting #systemdesign #distributedSystems #learninginpublic
After working on caching, execution modes, and circuit breakers, today I shifted focus to observability.
I added a memory-tracking middleware that wraps every request and logs memory usage:
What it does:
- Captures memory before and after each request
- Computes delta to show how much memory was actually consumed
- Tracks peak heap usage across requests
- Highlights significant memory changes (>1MB)
- Can be toggled on/off without touching the core logic
Metrics tracked include:
- RSS
- Heap total / heap used
- External memory
- ArrayBuffers
This made it much easier to answer questions like:
“Is this request leaking memory?”
“Which code paths cause heap growth?”
“Are retries, caching, or failures changing memory behavior?”
What I liked about this addition is that it’s middleware-based:
- no coupling with business logic
- easy to remove in prod or enable during debugging
- reusable beyond this service
Small system, but it’s slowly turning into a mini playground for:
performance, resilience, and observability.
#softwareengineering #backend #nodejs #observability #performance #systemsdesign #learninginpublic
After adding caching and sequential/parallel execution, today I focused on resilience.
I added a circuit breaker to the health check service.
How it behaves:
- Each API endpoint now has its own circuit breaker
- After 3 consecutive failures, the circuit opens
While open:
- No HTTP calls are made
- Requests are short-circuited immediately
After 10 seconds, the circuit automatically closes and traffic is tried again
No change needed in how the service is called — the circuit breaker is applied transparently per API.
What this adds:
- Prevents hammering already-failing services
- Reduces unnecessary latency and noise
- Makes health checks reflect real system behavior, not just retries
This was a fun reminder that reliability isn’t about adding more requests — it’s often about knowing when not to make one.
Small system, big lessons:
- isolation per dependency
- failure thresholds
- recovery windows
- and why “fast failure” is a feature, not a bug
#softwareengineering #backend #nodejs #resilience #systemdesign #learninginpublic #engineering
Yesterday I shared a small health check service I was working on.
Today I pushed it a step further.
I added:
Sequential vs Parallel execution modes
→ sequential for predictability and rate-limited APIs
→ parallel for faster diagnostics using Promise.allSettled
A from-scratch LRU cache with:
- configurable size
- TTL-based freshness
- cache hits clearly marked in logs
Smarter logging & summaries so it’s easier to see:
- what failed
- what was slow
- what came straight from cache
What I enjoyed most about this iteration wasn’t the code itself, but the trade-offs:
- when caching actually helps (and when it shouldn’t)
- why failed requests shouldn’t poison a cache
- why execution mode matters in real systems, not just benchmarks
Small systems like this are great reminders that good engineering is mostly about decisions, not frameworks.
#softwareengineering #backend #nodejs #systemsdesign #performance #learninginpublic
How healthy are your APIs?
Building a complex system is great, but knowing when it’s failing is better. For today's project, I built a Health Check Service in TypeScript to monitor API uptime and performance.
It’s a lightweight CLI tool that pings a set of endpoints, measures latency, and provides a color-coded status report.
What it does:
- Parallel Pinging: Hits multiple APIs simultaneously.
- Latency Tracking: Measures response times in milliseconds (critical for catching "slow-failures").
- Visual Feedback: Uses color-coded indicators (Green/Yellow/Red) based on performance thresholds.
- Detailed Logging: Captures status codes and ISO timestamps for every request.
The Tech Stack:
- TypeScript (for that sweet type safety)
- Node.js
- Axios for handling HTTP requests
One interesting thing I noticed while testing: even "public" APIs vary wildly in latency depending on the time of day. It’s a good reminder that your frontend is only as fast as the slowest service it depends on!
GitHub Repo: https://t.co/aHTgb5fALj
#100DaysOfCode #TypeScript #NodeJS #WebDevelopment #Backend #BuildInPublic #APIMonitoring
Stop writing your https://t.co/2RnKOMAAQL by hand. 🛑
We’ve all been there: You’re ready to push a release, but then you realize you have to manually scrub through 50+ commits to figure out what actually changed for the user.
I decided to fix that workflow for myself. I built a Python-based Changelog Generator that turns Git history into a clean, organized https://t.co/2RnKOMAAQL in seconds.
Why it’s useful:
🚀 Auto-Versioning: Uses Git tags to create version sections.
🧠 Conventional Commits: Automatically groups feat, fix, and docs into readable categories.
🛠 Customizable: Works on any local repo with simple CLI flags.
Check out the GitHub: https://t.co/fEwY2reePQ
#Python #Git #OpenSource #DevTools #Programming
Stop guessing what’s in your node_modules. 📦
We’ve all been there: you install one small package, and suddenly your dependency tree looks like an uncontrolled forest. I wanted a way to see exactly what’s happening without leaving the terminal.
So for today’s build, I created deps-graph-visualizer.
It’s a lightweight CLI tool that gives you a color-coded, recursive view of your project’s architecture.
Key Features:
🌳 Tree and List visualization modes.
🔄 Handles circular dependencies (no infinite loops!).
📊 Depth control to keep things readable.
⚡ Zero configuration—just run it and see the "why" behind your bundle size.
Next time you're wondering where that mystery package came from, try: npx deps-graph-visualizer
Check out the README below. Feedback on the tree logic is welcome! 👇
#BuildInPublic #SoftwareEngineering #JavaScript #NPM #CodingChallenge
I’m tired of Googling the same Regex and SQL snippets every week. So, I built a "Second Brain" for my Terminal.
As a Senior Engineer, the most expensive thing I own is my context.
Every time I leave my IDE to search for a "TIL" (Today I Learned) snippet in Notion, Slack, or a random .txt file, I lose momentum. Browsers are distraction traps.
For Day 7 of my 30-day build challenge, I implemented a CLI TIL (Today I Learned) Manager using Python.
The Requirements:
✅ Local-First & Fast: It needed to be faster than opening a browser.
✅ Fuzzy Search: I shouldn't have to remember the exact title. "dkr prn" should find my Docker Prune notes.
✅ Clipboard Integration: Find the snippet, hit enter, and have it immediately ready to paste into my terminal or editor.
The Tech Stack:
- Python for the logic.
- Typer for the CLI interface.
- Rich for the formatted Markdown rendering in the terminal.
- Thefuzz for the fuzzy string matching logic.
Why build this when Notion exists? Because high-performance engineering is about reducing friction. Having my technical "cheat sheets" accessible in < 200ms without leaving the terminal isn't just a "nice to have"—it’s a flow-state preserver.
#SoftwareEngineering #Python #BuildInPublic #Productivity #Coding #Terminal #DeveloperExperience
Stop manually writing TypeScript interfaces for huge JSON objects. 🛑
I got tired of the constant back-and-forth between API responses and my IDE, so I built a JSON-to-TypeScript converter to automate the grunt work.
It’s built with Next.js and handles the tricky stuff:
Deep Nesting: Recursively generates interfaces for complex objects.
Enum Support: Identifies repeating patterns and suggests Enums instead of just strings.
Instant Deployment: Live on Vercel for zero-latency conversions.
Efficiency is about building tools that let you focus on the actual logic, not the boilerplate.
#NextJS #TypeScript #WebDev #ProductivityTools #Vercel
Stop writing Dockerfiles from scratch. 🐳
Writing a Dockerfile shouldn’t feel like a chore. Every time I start a new project—whether it's Next.js, FastAPI, or Rust—I find myself hunting through old repos to copy-paste the same "optimized" boilerplate.
So, I built dockerme.
It’s an npx CLI tool that analyzes your imports and project structure to auto-generate a production-ready Dockerfile for you.
🔍 How it works:
Run npx dockerme in your root folder.
It detects your language (Node, Python, Go, etc.) and framework (React, Django, etc.).
It spits out an optimized, multi-stage Dockerfile.
No more "Which base image should I use?" or "How do I configure Nginx for Vite?" Just code, containerize, and ship.
#Docker #WebDev #DevOps #OpenSource #JavaScript #CodingLife