Snowflake ID
One more popular algorithm for generating IDs in distributed systems is Snowflake.
It was initially created by Twitter (X) to generate IDs for tweets.
In 2010, Twitter migrated their infrastructure from MySQL to Cassandra. Since Cassandra doesn’t have an in-built id generator, the team needed an approach for ID generation that met the following requirements:
- Generate tens of thousands of ids per second in a highly available manner
- IDs need to be roughly sortable by time, the accuracy of the sort should be about 1 second (tweets within the same second are not sorted)
- ID have to be compact and fit into 64 bits
As a solution Snowflake service was introduced that generates IDs with the following structure:
✔️ Sign bit: The first bit is always 0 to keep the ID a positive integer.
✔️ Timestamp (41 bits): Time when ID was generated
✔️ Node ID (10 bits): Unique identifier of the worker node generating the ID
✔️Step (12 bits): A counter that is incremented for each ID generated within the same timestamp
Snowflake IDs are sortable by time, because they are based on the time they were created. Additionally, the creation time can be extracted from the ID itself. This can be used to get objects that were created before or after a particular date.
There is no official standard for Snowflake approach, but there are several implementations available on Github. The approach has also been adopted by major companies like Discord, Instagram and Mastodon.
#architecture #systemdesign
Post #45
203
- 👍 2
- 🔥 2