1. No branching
2. Don't process byte by byte, use SIMD
3. Avoid memory allocations
4. Measure the performance (CI performance tests)
How simdjson is made that parses gigabytes of JSON per second. There are also a few performance tricks related to parsing.
And there is a very interesting comment under the video:
43:37 "cause you're assuming that the person running your program is not switching the CPU under you". The audience might be laughing, but this actually is sometimes a case, even in consumer hardware. Non-US Samsung Galaxy S9 has a heterogenous CPU, with some cores supporting the atomic increment instruction LDADDAL and others not, and with Linux kernel modified by Samsung to report that all cores support that instruction. Your program would crash after being rescheduled to another core.
Which is even more true nowadays with a wider adoption of e-cores and p-cores in modern CPUs.
https://www.youtube.com/watch?v=wlvKAT7SZIQ
#performance
