I’m Apoorav Malik, a full-stack + AI builder.
Working on AI workflow systems at ContraVault.
Building WePlay for grassroots esports organizers.
Documenting what I build, learn, break, and improve.
@itssubhamoy I love these kinds of interviews, in which they actually see what work you did and how are you going to do it for them. I dont think implementing a LRU cache is gonna help both of the parties 😅
Been a while since I posted here.
Life lately:
1/ Saw the Jev boom - started digging deeper into it.
2/ Helped my brother with DSA - somehow started learning it myself too. Still wondering why tf are we doing this in 2026 😭
3/ Going deeper into Harness Engineering and how much the setup around an agent actually matters.
4/ Building a system design workflow simulation project - trying to simulate real architectural decisions, tradeoffs, failures, and scaling.
Basically: went quiet, kept learning, kept building.
If you're interested in learning how to scale APIs, this is a MUST WATCH.
@Joseph_Cododev makes it incredibly fun to follow, and you get to see the engineering thought process behind breaking down and solving scaling problems.
Going to try recreating some of it this weekend, minus spinning up that many EC2 instances! 🚀
3/ Also measure TTFT per stage.
query processing: 40ms
retrieval: 120ms
reranking: 350ms
prompt assembly: 80ms
model queue/prefill: 700ms
Without stage-level timings, “our LLM is slow” is mostly guessing.
TTFT is probably the most underrated metric in AI search.
Your app can finish in 4s and still feel faster than one that finishes in 2s.
The trick: optimize the path to the first useful token, not the entire pipeline.
For AI search, that usually means:
→ fire retrieval immediately
→ parallelize independent searches
→ retrieve small, expand later
→ don’t block on rerankers
→ stream as soon as evidence is sufficient
→ push citations + secondary retrieval off the critical path
Fast AI search isn’t about doing less.
It’s about doing less before token #1.
2/ Reranking is a common offender.
For the initial answer, top-k from hybrid retrieval may already be good enough.
Start generation.
Run the heavier reranker concurrently and use it for follow-up retrieval, citations, or correction if needed.
1/ One useful mental model:
Split your AI search pipeline into:
must happen before token #1
vs.
can happen after token #1
You’ll usually find way more work in the first bucket than actually belongs there.
Starting to learn more about inference engineering.
Feels like this is the path from “vibe coding” to actually understanding the engineering behind the models.
The more you understand inference, the more it feels like engineering