@kartikb753 Yea the orchestrator should control all the retry and counting logic
Prompt doesn't provide guarantee
We can rely on probabilistic model to
Enforce the security logic
Your agent just made 47 tool calls, burned $30 and gave the wrong answer.
Confidently. With formatting.
No step limit. No cost ceiling. No success test. Nobody told it to stop, so it didn't.
Most agent failures are not model failures. The model is the smallest part.
A chain is control flow you wrote. An agent is control flow the model decides. The line is who holds the steering wheel.
What keeps an agent alive in production:
β Orchestrator in code that owns retries and budgets. "Please stop after 5 steps" in the prompt is a wish, not a limit.
β Typed tools with idempotency keys, so a retry doesn't email the customer 4 times
β Durable state, so a deploy doesn't kill the run at step 38
β Hard limits on steps, tokens, time and cost
β Human approval before anything irreversible
If you already know the workflow, you don't need an agent. You need a pipeline. Cheaper, faster, debuggable.
Start deterministic. Add autonomy only where measured failures justify it.
A demo is a model and a loop. Production is everything else.
mathematics path is where this gets interesting
The computation problems with normal attention and
the tricks and techniques that are being proposed are
interesting to learn about .
Have you directly started learning it or any prerequisite ?
AI Engineering Day - 01: Attention
Spent today understanding how attention actually works in LLMs. The idea is simpler than the formula makes it look.
Take "The animal didn't cross the road because it was too tired."
To figure out what "it" means, the token "it" sends out a Query: what am I looking for?
Every other token holds a Key: here's what I contain.
The best matches get a high score, and the model blends their Values into a new representation of "it".
That's the whole trick. Softmax just turns the scores into weights that add up to 1.
I wrote my notes on this with diagrams, including self vs cross vs causal attention and multi-head. If you're learning this too, it might save you some time:
https://t.co/v3wU67y81Z