I have 12 years of experience and working as a Principal Engineer @Atlassian and I have seen concurrency scaring the hell out of a lot of junior engineers.
It’s one of the most feared topics in system design & backend interviews — race conditions, deadlocks, thread pools… you name it.
But once you internalize these 20 must-know concepts, everything clicks.
Save this thread. Read till the end.
Your future interviews and production systems will thank you.
BREAKING: AI can now analyze stocks like Wall Street analysts (for free).
Here are 10 insane Claude prompts that replace $2,000/month Bloomberg terminals (Save for later)
35 practical thoughts on system design
Core Principles
1. Every system is a trade-off -> you never get speed, cost, and simplicity all at once.
2. Latency compounds -> every millisecond added across layers adds up to user pain.
3. Scalability ≠ performance -> one handles growth, the other handles speed.
4. Read vs. write paths -> scaling each requires completely different strategies.
5. Design for change, not perfection -> requirements will shift.
Databases & Storage
6. Indexes are levers -> high-selectivity columns are worth indexing, low ones often aren’t.
7. Replication helps reads, partitioning helps writes -> don’t confuse the two.
8. Dual writes are a lie -> without coordination, you will see drift.
9. Event stores > queues (sometimes) -> better traceability, worse simplicity.
10. Cache invalidation is still the hardest problem -> freshness vs. performance is the eternal fight.
Reliability & Consistency
11. Idempotency saves you -> retries without it will hurt.
12. Fail fast, fail loud -> silent failures sink systems.
13. Eventual consistency is a feature -> not a bug, but only if the business allows it.
14. Conflict resolution in active-active is business logic, not infra.
15. Durability isn’t free -> syncing across regions has cost and latency.
Architecture Patterns
16. Microservices are an org structure, not a tech goal.
17. Monolith first, modular second, microservices last -> don’t jump too early.
18. Choreography scales, orchestration simplifies -> pick based on team maturity.
19. Serverless buys you focus, costs you control.
20. Queues don’t remove work, they smooth it.
Observability & Operations
21. Tracing > logging -> logs tell you “what”, traces tell you “why”.
22. Metrics rot without ownership -> measure what someone actually uses.
23. Retries without backoff = DDoS on yourself.
24. Dead-letter queues aren’t optional -> every system has poison messages.
25. Levers > knobs -> design quick kill switches and controls to contain blast radius.
Performance & Cost
26. Optimize the hot path, not the cold path.
27. Most bottlenecks live in the database, not the code.
28. Horizontal scaling beats vertical scaling ->until coordination kills it.
29. Warm caches mask bad queries -> measure against cold starts too.
30. The cheapest resource is disk, the most expensive is time.
People & Process
31. The best architecture dies without documentation.
32. System design reviews aren’t about diagrams, they’re about trade-offs.
33. Small PRs are for speed, big PRs are for context -> balance both.
34. A senior engineer’s role in design is asking the uncomfortable “what if”.
35. No design survives first contact with production -> but good ones bend instead of break.
What else?
BREAKING: AI can now analyze any stock like a Wall Street analyst (for free).
Here are 10 insane Claude prompts that replace $2,000/month Bloomberg terminals:(Save for later)
What's the role of an API gateway?
It's like the front door to your backend services and APIs.
Key benefits of API gateways:
- Centralized access
- Service abstraction
- Performance enhancement
- Authentication and authorization
I want to focus on Authentication for a moment.
Did you know you can do request-response with messaging?
One service, the requester, sends a request message and waits for a corresponding response message. This is a synchronous communication approach from the requester's side.
Here's a diagram of what the flow looks like. 👇
As a Backend dev , how many concepts can you explain from below :
1. Event-Driven Architecture
2. Saga Pattern
3. CQRS (Command Query Responsibility Segregation)
4. Event Sourcing
5. Circuit Breaker Pattern
6. Distributed Tracing
7. CAP Theorem
8. Idempotency
9. Data Sharding
10. API Gateway
SOLID Principles Explained with Clear Examples:
𝐒 - 𝐒𝐢𝐧𝐠𝐥𝐞 𝐑𝐞𝐬𝐩𝐨𝐧𝐬𝐢𝐛𝐢𝐥𝐢𝐭𝐲 𝐏𝐫𝐢𝐧𝐜𝐢𝐩𝐥𝐞
A class should have only one reason to change.
- Example: Instead of one giant User class that handles authentication, profile updates, and sending emails, split it into UserAuth, UserProfile, and EmailService.
𝐎 - 𝐎𝐩𝐞𝐧/𝐂𝐥𝐨𝐬𝐞𝐝 𝐏𝐫𝐢𝐧𝐜𝐢𝐩𝐥𝐞
Classes should be open for extension but closed for modification.
- Example: Define a Shape interface with an area() method. When you need a new shape, just add a Circle or Triangle class that implements it.
𝐋 - 𝐋𝐢𝐬𝐤𝐨𝐯 𝐒𝐮𝐛𝐬𝐭𝐢𝐭𝐮𝐭𝐢𝐨𝐧 𝐏𝐫𝐢𝐧𝐜𝐢𝐩𝐥𝐞
Subtypes must be substitutable for their base types without breaking behavior.
- Example: If Bird has a fly() method, then Eagle and Sparrow should both work anywhere a Bird is expected.
𝐈 - 𝐈𝐧𝐭𝐞𝐫𝐟𝐚𝐜𝐞 𝐒𝐞𝐠𝐫𝐞𝐠𝐚𝐭𝐢𝐨𝐧 𝐏𝐫𝐢𝐧𝐜𝐢𝐩𝐥𝐞
Don't force classes to implement interfaces they don't use.
- Example: Instead of one fat Machine interface with print(), scan(), and fax(), break it into Printable, Scannable, and Faxable. A SimplePrinter only implements Printable.
𝐃 - 𝐃𝐞𝐩𝐞𝐧𝐝𝐞𝐧𝐜𝐲 𝐈𝐧𝐯𝐞𝐫𝐬𝐢𝐨𝐧 𝐏𝐫𝐢𝐧𝐜𝐢𝐩𝐥𝐞
High-level modules should not depend on low-level modules. Both should depend on abstractions.
- Example: Your OrderService should depend on a PaymentGateway interface, not directly on Stripe or PayPal.
The real power of SOLID is not in following each principle in isolation. It's in how they work together to make your code easier to change, test, and extend.
♻️ Repost to help others in your network
p50, p95, and p99.
How many p's are there, and what do they mean?
Short answer: p = percentile.
It shows how performance is distributed, not the average.
How many “p’s” exist?
There are percentiles from p0–p100; in practice we usually look at :
• p50 → median
• p90 → early slowdowns
• p95 → SLA baseline
• p99 → tail latency
• p99.9 → ultra reliability systems (Google and Amazon watch p99.9)
Imagine 100 users hitting your API
p50 = 120ms → 50 users were at or faster than 120ms, 50 were slower.
p95 = 800ms → 5 users were at or faster than 800ms, 5 were slower.
p99 = 2.5s → 99 users were at or faster than 2.5s, 1 was slower (this is where frustration lives).
Why track percentiles instead of averages?
Average latency hides pain.
Example:
100ms
110ms
120ms
130ms
5000ms
Average ≈ 1,092ms ❌
p50 = 120ms
p95 = 5000ms ✅
Knowing p50, p95, and p99 helps you:
• spot latency spikes before outages
• understand real user experience
• detect scaling and contention issues
• design safer retries, timeouts, and backpressure
• set meaningful SLAs
• prevent “fast but unreliable” systems
Users don’t experience your average.
They experience your worst moments.
Ignore the tail, and production will remind you.
When did p99 last surprise you?
The computer science major is going through an identity crisis.
ChatGPT can finish any programming assignment with a single prompt so what's the value of teaching students how to write a function in Python?
Here’s the point: we still need architects, not button-pushers.
The next decade will belong to people who understand theory and how to break complex systems into smaller components even without Wifi.
Imagine a degree that drills algorithmic thinking. Weekly closed-book exams. CLRS becomes the most important book in the major.
And coding? You’ll exercise that muscle only at internships and real jobs.
The computer science major will look more and more like a math degree.
CQS and CQRS are often confused.
CQS = Command Query Separation
- Concerned with class design, methods either change state (commands) or return a value (query)
CQRS = Command Query Responsibility Segregation
Most engineers overbuild their system design portfolios.
You only need three projects that prove how you think.
Not tools. Decisions.
1. CRUD System that Scales
Goal: show fundamentals.
Build: multi-tenant SaaS, auth, rate limits, caching.
Show: schema decisions, read vs write patterns, bottlenecks.
👉 proves you understand the boring foundations that scale.
2. Event-Driven System
Goal: show async thinking.
Build: order pipeline, notifications, audit log.
Include: broker, retries, idempotency, DLQs, delivery guarantees.
👉 proves you understand real distributed workflows.
3. Read-Heavy System
Goal: show performance instincts.
Build: feed, analytics, trending service.
Demonstrate: caching, precompute vs query, partitioning, latency control.
👉 proves you understand how systems survive traffic.
These three projects map to real production patterns:
CRUD → foundation of 80% of business systems
Events → how systems communicate
Read-heavy → where scale & latency matter
That’s the whole game.
What would you add, or remove?
"Competitive Programmer’s Handbook"
This book of ~300 pages covers 30 important topics. A MUST for tech interviews.
Download for FREE → https://t.co/BdDvKWWdq4