@devayush__ O_O save to maine bhi kiya tha like kr diya
Ek help chahiye thi - how to find startup and all be it early start or series a ,b . Mil hi nhi rhe
you change the prompt. eval pass rate goes 82% to 87%. looks good right, ship it?
nah slice it by intent first
lets say ur dataset is 70% refund queries and 30% cancellations. refunds went from 78 to 94. cancellations went from 90 to 71. average still goes up but a whole workflow just broke. big slice drowns the small one and u never see it
(numbers are made up but this exact thing happens, i have seen it)
what actually helps
pin ur dataset. if cases change between runs ur comparing nothing with nothing
diff per case not per run. dont ask whats the score, ask which cases flipped from pass to fail. 10 wins and 4 new failures is very different from "+6 net"
report per task slice. one number per intent, side by side for every version
also watch sample size. a slice with 30 cases has like +-15 points of noise. add temperature and llm judge variance on top and a 5 point jump can be literally nothing. rerun it or add more cases before u trust small delta
one aggregate pass rate is just a vibe check with decimals
Yup so to your question
1. Leetcode Krna but only leetcode mat kr - do neetcode pattern wise wali sheet
2. Build project , I can see you have worked with javascript, typescript to usme kis domain mai kiya hai backend ya sirf frontend
3. Java is not so good to build projects switch to ts or go .
4. Kitna time hai tumpe on campus ke liye jaoge ya offcampus
@kartikb753 Yea the orchestrator should control all the retry and counting logic
Prompt doesn't provide guarantee
We can rely on probabilistic model to
Enforce the security logic
Your agent just made 47 tool calls, burned $30 and gave the wrong answer.
Confidently. With formatting.
No step limit. No cost ceiling. No success test. Nobody told it to stop, so it didn't.
Most agent failures are not model failures. The model is the smallest part.
A chain is control flow you wrote. An agent is control flow the model decides. The line is who holds the steering wheel.
What keeps an agent alive in production:
β Orchestrator in code that owns retries and budgets. "Please stop after 5 steps" in the prompt is a wish, not a limit.
β Typed tools with idempotency keys, so a retry doesn't email the customer 4 times
β Durable state, so a deploy doesn't kill the run at step 38
β Hard limits on steps, tokens, time and cost
β Human approval before anything irreversible
If you already know the workflow, you don't need an agent. You need a pipeline. Cheaper, faster, debuggable.
Start deterministic. Add autonomy only where measured failures justify it.
A demo is a model and a loop. Production is everything else.