ddia/ch03.md - DDIA

I worked through this chapter while preparing a presentation for a book club, so you can check out the presentation instead: PDF.

The notes below are incomplete, as kept most of it in Russian, while preparing.

Event sourcing

  • When:
    • You need a complete audit history
    • You need to reconstruct past state.
    • You need multiple independent projections from the same data.
    • Higher write throughput.
  • Problems:
    • GDPR:
      • crypto-shredding
      • Pseudonymization
      • Tombstone events
      • puncturable encryption
  • Schema evolution
    • Upcasting
    • Version strategies

Implementation:

  • If we implement (or end-up) ES the following way database (postgres) as an event store + kafka as a bus => this leads to dual-write problem
  • Instead we need projections to read from the event store directly.

Approaches to do ES:

  • KurrentDB (EventStoreDB)
  • Postgres as event store Simple to operate if you already run Postgres. Rehydration, optimistic concurrency via version, and stream reads are all standard SQL. Consumer notification requires extra work (polling, LISTEN/NOTIFY, or CDC).
  • Postgres + Kafka Problem: dual-write problem (above)
  • Kafka alone Problems: topics with infinite retention act as the event log. Aggregates are rehydrated by consuming from offset 0. Works, but Kafka is optimized for throughput not for per-stream reads — replaying a single aggregate's history means scanning a whole partition.

Event sourcing => CQRS

  • ES is purely about how you store state — instead of overwriting current state, you append immutable events and derive state by replaying them, so a balance is computed by summing Deposited and subtracting Withdrawn events with no separate read model.
  • Works for reading a single aggregate, but don't when you need cross-aggregate queries like "all orders over $500 this week," because load and replay every order stream in the database (snapshots help but not really)
  • Solution => CQRS = pre-compute projections built from the event stream

CQRS

  • Command side: validates invariants, emits events, cares about consistency
  • Query side: pre-materialized, denormalized, optimized for access patterns
  • These can live in different databases, different services, different physical machines

When to apply CQRS

  • Heavily asymmetric read/write loads. You want to scale the read side independently.
    • Can it be mitigated on DB side? Choosing a suitable DB?
    • Maybe scale is the problem?
  • Having a single write path but many projections.