@Vtlabs1 Agents have no instinct for backing off. Every retry is "free" to them, so one slow response turns into a synchronized wave. Same fixes as always: exponential backoff with full jitter, a retry budget per client, and request coalescing so 1,000 identical misses become one rebuild.
Your cache TTL is a scheduled outage. One hot key expired at noon, 4,000 requests hit Postgres in the same second, and the connection pool died before Redis refilled anything.
A longer TTL is not the safe choice. It only moves the outage to when nobody is watching. Refresh hot keys before they expire.
More Java and backend deep dives: follow @code_w1th_me
Fix 2: lock the rebuild. SET lock:key 1 NX EX 10. One request goes to Postgres, the rest wait or serve the stale value. Trade-off: if the lock holder crashes, everyone waits out the 10 seconds.
Do you rely on a direct publish call in your event-driven service, or did you already learn this the hard way?
More Java and backend deep dives: follow @code_w1th_me
Order service wrote to Postgres, then published to Kafka in a separate call. One day the publish failed silently. Twelve thousand order-confirmation emails never went out.
Honest trade-off: at-least-once delivery, so every consumer must be idempotent. Plus one more table and one more process to run. Worth it once a lost event costs you real customers.
Honest trade-off: "up to 200x faster, 400x cheaper" is the vendor's best case. Calibration is a claim until you test it on labeled data. Experimental package. I have not run it in prod.
How many of your agent's LLM calls are really classification?
#AIAgents#LLM