In the past weeks, there have been two big articles discussing caching strategies:
- Mastering Caching in Distributed Applications
- 9 Caching Strategies for System Design Interviews
So it's a good time to review the patterns we use.
Let’s start with the definition of what caching is. I really like the following one:
Caching is the action of storing data in a temporary medium where the data is either cheaper, faster, or more optimal to retrieve rather than retrieving it from its original storage
Caches can be local or distributed.
5 implementation strategies:
✔️Cache-Aside (Lazy-Loading): The app controls reading and writing to the cache. If data's in the cache, it's used; if not, it's fetched from storage and added to the cache.
✔️Write-Through. Both the cache and the database get updated together within the same transaction when you write data. Reading happens only from the cache.
✔️Write-Around. Data is written directly to storage. Usually, it's combined with Cache-Aside for reading.
✔️Write-behind (write-back). Data is first written to the cache and then sent to the datastore asynchronously. The cache product usually handles syncing with the datastore.
✔️Read-Through. The cache is the main place to get data from. If it's not there, it's fetched from storage. Unlike Cache-Aside, the cache product, not the app, decides when and how to fetch data.
Cache invalidation can be performed time-based or event-based.
Cache-eviction strategies:
- Least Recently Used (LRU)
- First In First Out (FIFO)
- Least Frequently Used (LFU)
- Time To Live (TTL)
- Random Replacement
Promising paths for further development in that field:
- AI and Machine Learning-Driven Caching. It can enhance caching mechanisms by predicting data usage patterns and preemptively caching data based on the needs.
- In-Memory Data Grids. They not only cache data but also provide a range of data processing capabilities, real-time analytics, and decision-making directly within the cache layer.
#architecture #patterns