Data Consistency & Idempotency — Microservices Interview
Target: Senior Engineer · Engineering Lead · Pre-Architect Focus: Saga pattern, event sourcing, idempotency, eventual consistency, distributed transactions
Q: How do you ensure data consistency across multiple services?
Why interviewers ask this: This is the fundamental challenge of microservices. Tests understanding of eventual consistency, saga patterns, and trade-offs vs traditional ACID.
Answer
In a monolith with a single database, you use ACID transactions. In microservices, each service has its own database — ACID transactions across services are impossible.
Three approaches:
| Approach | Consistency Model | Complexity | Use Case |
|---|---|---|---|
| Saga Pattern | Eventual consistency | Medium | Distributed business transactions (orders, payments) |
| Event Sourcing | Eventual consistency | High | Audit trail, temporal queries, event-driven systems |
| 2-Phase Commit | Strong consistency | High complexity, poor performance | Rare — only tightly coupled legacy systems |
Saga Pattern — Recommended:
graph LR
subgraph Saga["Order Saga · Compensating Transactions"]
Step1["Order Service\nCreate Order"]
Step2["Payment Service\nProcess Payment"]
Step3["Inventory Service\nReserve Stock"]
Comp1["Compensate:\nRefund Payment"]
Comp2["Compensate:\nRelease Reservation"]
end
Step1 -->|Success| Step2
Step2 -->|Success| Step3
Step3 -->|Fail| Comp1
Comp1 -->|Undo| Step1
Step2 -->|Fail| Comp1
Comp1 -->|Execute| Comp2
Saga orchestrates a sequence of local (per-service) transactions. On failure, compensating transactions execute in reverse order: 1. Order Service creates order 2. Payment Service processes payment 3. If payment fails → trigger refund (compensating tx) 4. If refund fails → escalate to manual intervention
Architect Insight
Sagas don't guarantee atomicity like a database transaction — they guarantee eventual consistency and durability (no money lost, no order lost). Design each saga step to be idempotent so retries are safe.
Q: Orchestration vs Choreography Sagas — Which should you use?
Answer
Orchestration — Central coordinator explicitly manages the flow:
Choreography — Services react to events, flow emerges from interactions:
Client → OrderService (publishes: OrderCreated)
↓
PaymentService (listens, publishes: PaymentProcessed)
↓
InventoryService (listens, publishes: InventoryReserved)
↓
ShippingService (listens)
Trade-off comparison:
| Orchestration | Choreography | |
|---|---|---|
| Coupling | Medium — coordinator knows all steps | Low — each service independent |
| Debugging | Easy — one place to trace flow | Hard — distributed event logic |
| Complexity | Single orchestrator service | Many event listeners |
| Testing | Mock dependencies in orchestrator | Integration test entire flow |
| Scalability | Orchestrator can become bottleneck | Scales better |
| Failure recovery | Timeouts + compensations in one place | Scattered across services |
Recommendation:
- Orchestration for critical, complex business processes (order → payment → inventory → shipping)
- Choreography for loosely coupled events (user signup → send welcome email, create analytics record)
- Hybrid for best of both: orchestrate order creation, choreograph notifications
Q: Why is 2-Phase Commit problematic in microservices?
Why interviewers ask this: Tests whether you understand why old database patterns don't scale in distributed systems.
Answer
2PC (Two-Phase Commit) is a database algorithm for atomic updates across multiple databases:
Phase 1 (Prepare): Coordinator asks all participants "Can you commit?" — they lock resources and respond. Phase 2 (Commit): If all yes, commit. If any no, rollback.
Why it fails in microservices:
| Problem | Impact |
|---|---|
| Resource locks | Prepare phase holds locks on Payment, Inventory for the entire duration — other requests blocked |
| Blocking | If one service is slow, all others wait (cascading slowness) |
| Network partitions | If coordinator crashes, participants hold locks forever (distributed deadlock) |
| Scalability | Doesn't work across service boundaries reliably — services should be independent |
| Latency | Synchronous, multi-round-trip protocol — slow |
Example failure:
Coordinator: "All commit?"
PaymentService: "Yes" ✓
InventoryService: [network partition, no response]
Coordinator: [waiting indefinitely]
Other orders: [blocked on payment/inventory locks]
When 2PC is safe:
- Single database with distributed transactions (no network partitions)
- Tightly coupled legacy systems where you control all components
- Very low scale (not at production microservices scale)
Use Saga instead — allows partial success and recovers gracefully.
Q: How do you implement idempotent API endpoints?
Why interviewers ask this: Idempotency is critical for safe retries. Tests system thinking about failure scenarios.
Answer
Idempotency means calling the same operation multiple times with the same input produces the same result as calling it once.
Implementation strategy:
-
Client provides idempotency key (UUID, unique per logical operation):
-
Server checks if key was already processed:
@PostMapping("/payments") public ResponseEntity<PaymentResponse> pay( @RequestBody PaymentRequest request, @RequestHeader("Idempotency-Key") String key) { // Check if already processed Optional<PaymentResponse> cached = idempotencyStore.get(key); if (cached.isPresent()) { return ResponseEntity.ok(cached.get()); // Return cached result } // Process payment Payment payment = paymentService.process(request); // Store result with idempotency key idempotencyStore.put(key, response, ttl = 24_hours); return ResponseEntity.ok(response); } -
Store the result for a time window (24 hours typical):
For payment services specifically:
@Service
public class PaymentService {
@Transactional
public PaymentResponse processPayment(PaymentRequest req, String idempotencyKey) {
// All in one transaction — atomic
PaymentResponse response = callPaymentGateway(req);
idempotencyRepo.save(new IdempotencyRecord(idempotencyKey, response));
return response;
}
}
Common Mistake
Don't check idempotency AFTER processing — if the check fails and you process twice, you've already duplicated the charge. Check BEFORE, and only proceed if not found.
Q: How do you implement exactly-once event processing?
Why interviewers ask this: Event-driven systems at scale face duplicate message challenges. Tests understanding of messaging guarantees and application-level deduplication.
Answer
Message delivery guarantees:
| Guarantee | How it works | When duplicates occur |
|---|---|---|
| At-most-once | Sent once, may be lost | Network failure before server acks |
| At-least-once | Resent until acked | Server crashes after processing but before acking |
| Exactly-once | At-least-once + deduplication | Application responsibility |
Most message brokers offer at-least-once (safer than at-most-once). To achieve exactly-once:
graph LR
Producer["Producer"]
Kafka["Kafka · At-least-once"]
Consumer["Consumer · Event Handler"]
DB["Database"]
Producer -->|Send event| Kafka
Kafka -->|Deliver · may retry| Consumer
Consumer -->|1. Check if processed| DB
Consumer -->|2. If not: Process + Store ID| DB
DB -->|Atomic transaction| DB
Consumer -->|3. Acknowledge| Kafka
Implementation:
@KafkaListener(topics = "orders")
public void handleOrderEvent(OrderEvent event, Acknowledgment ack) throws Exception {
String eventId = event.getEventId(); // UUID
try {
// Step 1: Check if already processed
if (deduplicationStore.hasProcessed(eventId)) {
log.info("Event {} already processed, skipping", eventId);
ack.acknowledge();
return;
}
// Step 2: Process event + store ID in SAME transaction
@Transactional
void processInTransaction() {
// Process the event
Order order = orderService.createOrder(event);
// Store the event ID to prevent reprocessing
deduplicationStore.markProcessed(eventId, Instant.now());
// ^^ Both writes happen atomically in the DB transaction
}
// Step 3: Acknowledge after successful processing
ack.acknowledge();
} catch (Exception e) {
log.error("Failed to process event {}, will retry", eventId, e);
// Don't acknowledge — Kafka will redeliver
throw e;
}
}
Deduplication store schema:
CREATE TABLE processed_events (
event_id VARCHAR(36) PRIMARY KEY,
processed_at TIMESTAMP NOT NULL,
expires_at TIMESTAMP NOT NULL
);
CREATE INDEX idx_event_id ON processed_events(event_id);
Key principle: Process event + store event ID in the same database transaction. If either fails, both rollback. On retry, the event ID check prevents reprocessing.
Q: A message queue is backing up. How do you diagnose and fix it?
Answer
Diagnose:
| Metric | Healthy | Warning | Critical |
|---|---|---|---|
| Queue depth | < 1K | 1K-10K | > 10K |
| Consumer lag | < 1 sec | 1-30 sec | > 1 min |
| Throughput | > 1K msg/sec | 100-1K | < 100 |
Common causes and fixes:
Queue backing up?
├─ Consumer slow?
│ ├─ Optimize handler code (profile, add caching)
│ ├─ Parallelize: increase consumer instances
│ └─ Batch: process multiple messages together
├─ Downstream service down?
│ ├─ Circuit break: stop consuming, prevent cascading
│ └─ Retry: exponential backoff + dead letter queue
├─ Message poison (repeatedly fails)?
│ ├─ Move to dead letter queue
│ ├─ Alert ops
│ └─ Manual review and fix
└─ Sustained high volume?
├─ Add more partitions (Kafka)
├─ Auto-scale consumers
└─ Archive old messages
Production example — Spring Cloud Stream:
@Configuration
public class ConsumerConfig {
@Bean
public Consumer<Message<OrderEvent>> handleOrder() {
return message -> {
OrderEvent event = message.getPayload();
try {
orderService.processOrder(event);
} catch (TemporaryException e) {
// Retry with backoff
throw new AmqpRejectAndDontRequeueException("Temporary failure", e);
} catch (PermanentException e) {
// Send to dead letter
deadLetterQueue.send(event);
log.error("Unrecoverable error, moved to DLQ", e);
}
};
}
}
Diagram — Complete Idempotent Order Processing
graph LR
Client["Client\n· Idempotency Key"]
OrderSvc["Order Service"]
IdempStore["Idempotency Store\nRedis + DB"]
PaymentSvc["Payment Service"]
InventorySvc["Inventory Service"]
EventBus["Event Bus"]
Client -->|Request + Key| OrderSvc
OrderSvc -->|Check: already processed?| IdempStore
IdempStore -->|Cache hit: return cached| OrderSvc
IdempStore -->|Cache miss: process| OrderSvc
OrderSvc -->|Reserve inventory| InventorySvc
OrderSvc -->|Process payment| PaymentSvc
PaymentSvc -->|Idempotent call| PaymentSvc
OrderSvc -->|Store result + key| IdempStore
OrderSvc -->|Publish: OrderCreated| EventBus
style IdempStore fill:#51cf66
style OrderSvc fill:#4ecdc4
style EventBus fill:#ffe066
Distributed Systems Fundamentals
Q: What are SQL transaction isolation levels, and how do they affect microservices?
Why interviewers ask this: Isolation level bugs cause phantom reads, lost updates, and data races that are extremely hard to reproduce. Tests whether a candidate truly understands database concurrency, not just "wrap it in @Transactional".
Answer
Four isolation levels (weakest to strongest):
| Level | Dirty Read | Non-Repeatable Read | Phantom Read | Performance |
|---|---|---|---|---|
| Read Uncommitted | ✅ possible | ✅ possible | ✅ possible | Highest |
| Read Committed | ❌ prevented | ✅ possible | ✅ possible | High (default for most DBs) |
| Repeatable Read | ❌ prevented | ❌ prevented | ✅ possible | Medium |
| Serializable | ❌ prevented | ❌ prevented | ❌ prevented | Lowest |
Phenomena explained:
Dirty Read: Transaction A reads data written by uncommitted Transaction B.
B rolls back → A has read data that never officially existed.
Non-Repeatable Read: Transaction A reads a row. Transaction B updates it.
A reads same row again → different value. Same query, different results.
Phantom Read: Transaction A queries "orders WHERE status=PENDING".
Transaction B inserts a new PENDING order.
A re-runs query → sees more rows. Same predicate, different row count.
Spring Boot — setting isolation level:
// Default: inherits DB default (usually READ_COMMITTED)
@Transactional
public void processOrder(String orderId) { ... }
// For financial ledger — prevent phantom reads
@Transactional(isolation = Isolation.SERIALIZABLE)
public void transferFunds(String fromId, String toId, BigDecimal amount) {
Account from = accountRepo.findById(fromId).orElseThrow();
Account to = accountRepo.findById(toId).orElseThrow();
from.debit(amount);
to.credit(amount);
}
// Optimistic locking — alternative to SERIALIZABLE for read-heavy scenarios
@Entity
public class Product {
@Id private String id;
private int stockCount;
@Version
private Long version; // JPA checks version on UPDATE — throws OptimisticLockException if changed
}
Practical guide for microservices:
| Use Case | Recommended Isolation | Why |
|---|---|---|
| General CRUD (order reads, profile updates) | Read Committed | Best performance; dirty reads prevented |
| Inventory reservation (check stock then decrement) | Repeatable Read + SELECT FOR UPDATE |
Prevent lost updates under concurrent writes |
| Financial transfers, double-entry ledger | Serializable | No phantom can create phantom balance |
| Read-only reporting queries | Read Uncommitted | Acceptable for approximate analytics |
Common Mistake
Using @Transactional without thinking about the isolation level is the default of most developers. READ_COMMITTED is fine for most CRUD, but concurrent stock checks (if (stock > 0) { reserve(); }) need SELECT FOR UPDATE or optimistic locking — otherwise two transactions both see stock=1 and both proceed, resulting in stock=-1.
Q: How do you implement distributed locking in a microservices environment?
Why interviewers ask this: Without distributed locks, concurrent service instances can create duplicate records, double-charge customers, or corrupt shared state. Tests understanding of synchronisation beyond single-process concurrency.
Answer
The problem: Multiple instances of a service run concurrently. A local synchronized block protects against thread races within one JVM, but does nothing when the same code runs on 10 pods.
Three approaches:
| Approach | Tool | Strengths | Weaknesses |
|---|---|---|---|
| Redis SETNX | Redisson / Lettuce | Fast, simple, TTL-based auto-release | Redis restart loses lock state |
| Database lock | PostgreSQL advisory lock | ACID-backed, no extra infra | Slower, DB connection held |
| Consensus-based | etcd / Zookeeper | Strongest guarantees | Operational complexity |
Redis distributed lock with Redisson:
@Service
public class InventoryService {
@Autowired
private RedissonClient redisson;
public boolean reserveStock(String productId, int quantity) {
RLock lock = redisson.getLock("inventory-lock:" + productId);
try {
// Try to acquire lock: wait 3s max, auto-release after 10s
boolean acquired = lock.tryLock(3, 10, TimeUnit.SECONDS);
if (!acquired) {
throw new LockAcquisitionException("Could not acquire lock for product: " + productId);
}
// Critical section — only one instance executes at a time
int stock = inventoryRepo.getStock(productId);
if (stock < quantity) return false;
inventoryRepo.decrementStock(productId, quantity);
return true;
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
return false;
} finally {
if (lock.isHeldByCurrentThread()) {
lock.unlock();
}
}
}
}
PostgreSQL advisory lock (no extra infrastructure):
// Uses database's built-in advisory lock mechanism
@Transactional
public void withAdvisoryLock(long lockId, Runnable action) {
// Blocks until lock acquired (for this DB session / transaction)
jdbcTemplate.execute("SELECT pg_advisory_xact_lock(" + lockId + ")");
action.run();
// Lock auto-released when transaction ends
}
// Usage
withAdvisoryLock(productId.hashCode(), () -> {
int stock = inventoryRepo.getStock(productId);
if (stock >= quantity) inventoryRepo.decrementStock(productId, quantity);
});
Architect Insight
Distributed locks are often a symptom of missing domain design. Ask first: can idempotency keys prevent the duplicate? Can optimistic locking (@Version) resolve the race? Can you make the operation naturally idempotent (e.g., SET stock = MAX(0, stock - qty))? Locks introduce latency, lock-holder crashes cause timeouts, and they serialise throughput. Use them as a last resort, not a first instinct.
Q: Explain the CAP Theorem. How does it affect your database choices in microservices?
Why interviewers ask this: A foundational distributed systems concept. Tests whether a candidate can reason about trade-offs rather than just picking a database by habit.
Answer
The CAP Theorem states that a distributed system can only guarantee two of three properties simultaneously:
| Property | Meaning |
|---|---|
| Consistency (C) | Every read receives the most recent write or an error |
| Availability (A) | Every request receives a response (not necessarily the most recent data) |
| Partition Tolerance (P) | System continues operating despite network partitions |
Network partitions in real distributed systems are unavoidable — so the real trade-off is always CP vs AP:
CP (Consistency + Partition Tolerance):
Sacrifice availability — reject reads/writes during a partition
Examples: HBase, Zookeeper, etcd, traditional RDBMS with quorum writes
AP (Availability + Partition Tolerance):
Sacrifice strong consistency — serve potentially stale data during a partition
Examples: Cassandra, DynamoDB, Couchbase, CouchDB
Practical decision table:
| Use Case | Trade-off | Choice |
|---|---|---|
| Financial transactions (payments, ledgers) | CP — never show stale balance | PostgreSQL, CockroachDB |
| Product catalog, user sessions | AP — stale data is acceptable | DynamoDB, Cassandra |
| Distributed config / leader election | CP — correctness critical | etcd, Zookeeper |
| Shopping cart, activity feeds | AP — availability > consistency | Redis, DynamoDB |
graph LR
CAP["CAP Theorem"]
C["Consistency\nAll nodes see same data"]
A["Availability\nAlways responds"]
P["Partition Tolerance\nSurvives network split"]
CP["CP Systems\nPostgreSQL · HBase · etcd"]
AP["AP Systems\nCassandra · DynamoDB · Redis"]
CAP --> C
CAP --> A
CAP --> P
C & P --> CP
A & P --> AP
Architect Insight
CAP is a simplification — in practice, use PACELC which extends it: even when no partition exists, there is a trade-off between Latency and Consistency. Most NoSQL databases trade consistency for lower latency, not just for partition tolerance.
Q: What is the difference between ACID and BASE? When do you apply each model?
Why interviewers ask this: Tests understanding of why microservices move away from traditional database guarantees and the consequences of that choice.
Answer
ACID — traditional relational database guarantees:
| Property | Meaning |
|---|---|
| Atomicity | All operations in a transaction succeed or all are rolled back |
| Consistency | Database moves from one valid state to another |
| Isolation | Concurrent transactions don't interfere |
| Durability | Committed data survives crashes |
BASE — the distributed systems alternative:
| Property | Meaning |
|---|---|
| Basically Available | System guarantees availability per CAP |
| Soft state | State may change over time without input (due to eventual consistency) |
| Eventually Consistent | System will become consistent over time, given no new input |
When to use each:
| Scenario | Model | Example |
|---|---|---|
| Payment processing, financial records | ACID | PostgreSQL, MySQL |
| Order placement spanning multiple services | BASE + Saga | Eventual consistency with compensating txns |
| User profile reads, product catalog | BASE | DynamoDB, Cassandra |
| Inventory reservation | Depends — prefer ACID for stock counts | PostgreSQL with optimistic locking |
// ACID: single-service local transaction
@Transactional
public void transferFunds(String fromId, String toId, BigDecimal amount) {
Account from = accountRepo.findById(fromId).orElseThrow();
Account to = accountRepo.findById(toId).orElseThrow();
from.debit(amount);
to.credit(amount);
// All-or-nothing — atomicity guaranteed
}
// BASE: cross-service eventual consistency via Saga
public void placeOrder(OrderRequest req) {
orderService.createOrder(req); // Local ACID txn per service
eventBus.publish(new OrderPlaced(req)); // Trigger downstream saga steps
// Eventual consistency — inventory and payment settle asynchronously
}
Common Mistake
Don't assume BASE means "no consistency." It means eventual consistency — your system must be designed so that all services converge to a consistent state given no new failures. Design every saga step to be idempotent and define explicit compensating transactions for failure paths.