Dual-Write Problem
In the distributed world, it's often the case that two external systems need to be synchronized and updated simultaneously to maintain consistency. It’s called the dual-write problem. A classic example of this is when data needs to be stored both in a database and in Kafka.
So what's the best solution for this problem? What are common pitfalls and how to avoid them? These questions are addressed in the article 'Solving the Dual-Write Problem` by W.Waldron.
Identified antipatterns:
📍Relying on operation order
📍Wrapping dual writes into a database transaction
📍Retrying failed operations
Possible solutions:
📍Transactional outbox pattern. In this approach, an outbox table is set up in the database. Changes are made to the target tables and the outbox table within the same transaction. A separate process then reads the outbox table rows and sends them to Kafka, retrying if there's a failure.
📍Change Data Capture (CDC) if supported by the database. Modification of p.1
📍Event sourcing. Event sourcing: Every change is recorded in the database as an event. Since each event is written to a single row in a single table, transactions are unnecessary. A separate process can then read these events and send them to Kafka.
📍The listen-to-yourself pattern. Any change is directly sent to Kafka. A separate process can listen to these events and use them to update the database, retrying as needed. Database will be eventually consistent.
The core concept behind these solutions is to divide writes into two separate processes and establish a dependency between them. While this isn't an exhaustive list of all potential solutions, it provides a solid set of practices to begin with.
#architecture #systemdesign #patterns
Post #22
434