TGViewer
C++ - Reddit C++ - Reddit @r_cpp · 230 subscribers
Post #25258 13
Latte: a single-header latency measurement for quick insights

Hey /r/cpp, I've been working on a single-header latency measurement lib. Not meant for bottleneck detection or replacing Tracy/OpenTelemetry/perf/callgrind, but just for situations where I needed simple and trustable latency numbers during development.

\## The pitch:

* **2.5x faster per call than chrono** (RDTSC: \~60 cycles vs \~154 cycles)
* **Built-in statistics** (mean, median, stddev, skew, min, max, range, outliers)
* **Thread-safe** (per-thread ring buffers, zero contention)
* **Header-only** - nothing to do except placing monitoring beacons

\## Basic usage:

```cpp
Latte::Fast::Start(__func__);
DoWork(); // block of logic to measure
Latte::Fast::Stop(__func__);

// For loops/toroidal events:
for (;;) {
// ... work ...
LATTE_PULSE("MyLoop"); // records delta between successive calls
}

Technical implementation:

Three capture modes with different serialization guarantees:
\- Fast: __rdtsc
\- Mid: __rdtscp
\- Hard: _LFENCE+__rdtscp

Storage model:

\- Per-thread std::map<const char*, RingBuffer> (keys compared by pointer address, not string content)
\- Each ring buffer: alignas(64) for cache-line isolation, fixed 65k samples (configurable via BUFFER_PWR)
\- Zero allocations in hot path, ring buffers overwrite on wrap
\- Supports 64-deep nesting via per-thread SoA stack (stores ID, timestamp, capture mode)

Statistical cleaning:
Before computing stats, samples are bucketed (groups of 1000), bucket-max is recorded, then IQR filtering is applied to the maxima. This is more robust against long-tail outliers than raw IQR on samples.

Calibration:
When using Parameter::Calibrated, the framework measures and subtracts instrumentation overhead for each Start/Stop mode combination (e.g., Fast→Mid, Hard→Hard). Mixed-mode nesting is handled by storing the capture mode on the stack.

Example output:

| Component Samples Avg Median StdDev Min Max Range Outliers |
|--------------------------------------------------------------------------------------------|
| DoWork 10000 0.82 ms 0.81 ms 0.05 ms 0.79 ms 1.20 ms 0.41 ms 12 |
| InnerLoop 50000 0.15 ms 0.14 ms 0.02 ms 0.12 ms 0.35 ms 0.23 ms 3 |

With overhead correction enabled, it also dumps a calibration table showing measured overhead for each mode permutation.

Design decisions I’m curious about:

[0\] Pointer-as-key: Using const char* directly (string literals) as map keys to skip hashing. Feels hacky but saves cycles. Better alternatives?
[1\] Thread-local maps: Each thread owns its own std::map for ID→RingBuffer lookups. O(log N) per Stop(), but N is typically small. Considered flat_map but insertion cost for new IDs felt worse. Any thoughts ?
[2\] No RAII wrapper: Deliberately avoided RAII (Latte::Scope s("id")) to allow mixing Start/Stop modes and finer control. Trade-off worth it ?

Caveats:
\- x86_64 only (needs RDTSC/RDTSCP)
\- C++17
\- DumpToStream() not thread-safe - call only after workers stop
\- IDs must be string literals or stable static const char* (pointer comparison)

Nothing extravagant. Use case might be niche, I personally find it very useful for HFT/gamedev/tooling work where I just want to have quick insights without the measurement distorting results.

What do you think? Any obvious optimizations or design flaws I’m missing?

https://github.com/MoonFlowww/Latte

\##Bench
\### **ASM**
| Function | Avg (cycles) | Median (cycles) | StdDev (cycles) | Min (cycles) | Max (cycles) | Δ Min-Max (cycles) |
|:-----------------|-------------:|----------------:|----------------:|-------------:|-------------:|-------------------:|
| __rdtsc | 30.1 | 29.9 | 0.4 | 29.7 | 31.2 | 1.5 |
| __rdtscp | 57.7 | 57.5 | 0.9 | 57.3 | 62.6 | 5.2 |
| _LFENCE |
GitHub GitHub - fior512/Latte: Latency Telemetry with ultra low overhead Latency Telemetry with ultra low overhead. Contribute to fior512/Latte development by creating an account on GitHub.
More from @r_cpp
  1. Sep 29, 2026Boost.Graph 1.95 will be C++17 Dear Boost.Graph community, In two release cycles (Boost re…
  2. Sep 26, 2026Token Sequence Injection & Modern Macros: The Most Game-Changing Compile-Time Feature in C…
  3. Sep 25, 2026myStringStream.str("") Considered Harmful Under C++20 I was recently looking at some (rath…
  4. Sep 20, 2026A clever branch free optimization I'm the developer of memlz which is an extremely fast co…
  5. Sep 15, 2026Inside Boost.PolyCollection https://bannalia.blogspot.com/2026/09/inside-boostpolycollecti…
  6. Sep 11, 2026C++26: Standard Library Hardening Experiments https://www.cppstories.com/2026/hardening-ex…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →