Latte: a single-header latency measurement for quick insights
Hey /r/cpp, I've been working on a single-header latency measurement lib. Not meant for bottleneck detection or replacing Tracy/OpenTelemetry/perf/callgrind, but just for situations where I needed simple and trustable latency numbers during development.
\## The pitch:
* **2.5x faster per call than chrono** (RDTSC: \~60 cycles vs \~154 cycles)
* **Built-in statistics** (mean, median, stddev, skew, min, max, range, outliers)
* **Thread-safe** (per-thread ring buffers, zero contention)
* **Header-only** - nothing to do except placing monitoring beacons
\## Basic usage:
```cpp
Latte::Fast::Start(__func__);
DoWork(); // block of logic to measure
Latte::Fast::Stop(__func__);
// For loops/toroidal events:
for (;;) {
// ... work ...
LATTE_PULSE("MyLoop"); // records delta between successive calls
}
Technical implementation:
Three capture modes with different serialization guarantees:
\- Fast: __rdtsc
\- Mid: __rdtscp
\- Hard: _LFENCE+__rdtscp
Storage model:
\- Per-thread std::map<const char*, RingBuffer> (keys compared by pointer address, not string content)
\- Each ring buffer: alignas(64) for cache-line isolation, fixed 65k samples (configurable via BUFFER_PWR)
\- Zero allocations in hot path, ring buffers overwrite on wrap
\- Supports 64-deep nesting via per-thread SoA stack (stores ID, timestamp, capture mode)
Statistical cleaning:
Before computing stats, samples are bucketed (groups of 1000), bucket-max is recorded, then IQR filtering is applied to the maxima. This is more robust against long-tail outliers than raw IQR on samples.
Calibration:
When using Parameter::Calibrated, the framework measures and subtracts instrumentation overhead for each Start/Stop mode combination (e.g., Fast→Mid, Hard→Hard). Mixed-mode nesting is handled by storing the capture mode on the stack.
Example output:
| Component Samples Avg Median StdDev Min Max Range Outliers |
|--------------------------------------------------------------------------------------------|
| DoWork 10000 0.82 ms 0.81 ms 0.05 ms 0.79 ms 1.20 ms 0.41 ms 12 |
| InnerLoop 50000 0.15 ms 0.14 ms 0.02 ms 0.12 ms 0.35 ms 0.23 ms 3 |
With overhead correction enabled, it also dumps a calibration table showing measured overhead for each mode permutation.
Design decisions I’m curious about:
[0\] Pointer-as-key: Using const char* directly (string literals) as map keys to skip hashing. Feels hacky but saves cycles. Better alternatives?
[1\] Thread-local maps: Each thread owns its own std::map for ID→RingBuffer lookups. O(log N) per Stop(), but N is typically small. Considered flat_map but insertion cost for new IDs felt worse. Any thoughts ?
[2\] No RAII wrapper: Deliberately avoided RAII (Latte::Scope s("id")) to allow mixing Start/Stop modes and finer control. Trade-off worth it ?
Caveats:
\- x86_64 only (needs RDTSC/RDTSCP)
\- C++17
\- DumpToStream() not thread-safe - call only after workers stop
\- IDs must be string literals or stable static const char* (pointer comparison)
Nothing extravagant. Use case might be niche, I personally find it very useful for HFT/gamedev/tooling work where I just want to have quick insights without the measurement distorting results.
What do you think? Any obvious optimizations or design flaws I’m missing?
https://github.com/MoonFlowww/Latte\##Bench
\### **ASM**
| Function | Avg (cycles) | Median (cycles) | StdDev (cycles) | Min (cycles) | Max (cycles) | Δ Min-Max (cycles) |
|:-----------------|-------------:|----------------:|----------------:|-------------:|-------------:|-------------------:|
| __rdtsc | 30.1 | 29.9 | 0.4 | 29.7 | 31.2 | 1.5 |
| __rdtscp | 57.7 | 57.5 | 0.9 | 57.3 | 62.6 | 5.2 |
| _LFENCE |