TGViewer
C++ - Reddit C++ - Reddit @r_cpp · 230 subscribers
Post #25683 14
detour<false>(dst + done, tile - done, w, f, mask);
prev = rem * probe / (tile - done);
}
} else {
prev = detour<false>(dst, tile, w, f, mask) * probe / tile;
}
}
} else if (cnt < xlo) {
prev = detour_compact<true>(dst, tile, w, f, mask) * probe / tile;
} else if (cnt < lo) {
prev = detour_compact<false>(dst, tile, w, f, mask) * probe / tile;
} else {
prev = detour<false>(dst, tile, w, f, mask) * probe / tile;
}
dst += tile;
}
}

And v4 easily passes that test. pow10:

| thd | 0 | 1 |
|--------------|--------:|-------:|
| tiled detour | 38.18 | 6.94 |
| pilot v3 | 68.11 | 1.04 |
| pilot v3.5 | 68.67 | 3.84 |
| pilot v4 | 69.51\* | 7.05\* |
| BSL | 1.04 | 1.04 |

btw here thd no longer matches the density. For thd = 0, density = 0, and for thd = 1, density = 9.5%

It beats v3.5 by over 80%. It's also faster than plain tiled detour, because v4 picks detour_compact. This is the final
version.

## Results

v4 vs BSL:

| thd | 0 | 0.05 | 0.1 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 | 1 |
|:-----------------|--------:|--------:|--------:|--------:|--------:|--------:|--------:|--------:|--------:|--------:|--------:|--------:|
| pilot_v4 sqrt | 69.81\* | 31.99\* | 31.84 | 31.73 | 31.83 | 32.07\* | 31.98 | 31.96\* | 31.96 | 31.89\* | 31.91 | 31.77 |
| BSL sqrt | 31.88 | 31.93 | 31.95\* | 31.80\* | 31.89\* | 31.99 | 32.06\* | 31.92 | 31.98\* | 31.75 | 31.93\* | 31.91\* |
| pilot_v4 frfrexp | 69.41\* | 17.21\* | 16.57 | 15.79 | 15.99 | 16.06 | 15.93 | 15.99 | 16.39 | 15.84 | 15.87 | 15.80 |
| BSL frfrexp | 16.79 | 16.87 | 16.82\* | 16.71\* | 16.78\* | 16.88\* | 16.89\* | 16.83\* | 16.89\* | 16.79\* | 16.75\* | 16.75\* |
| pilot_v4 sin10 | 68.19\* | 18.82\* | 14.84\* | 10.84\* | 8.93\* | 7.55\* | 6.56\* | 5.81\* | 5.19\* | 4.69 | 4.68 | 4.68 |
| BSL sin10 | 4.78 | 4.75 | 4.78 | 4.72 | 4.75 | 4.74 | 4.75 | 4.77 | 4.75 | 4.76\* | 4.77\* | 4.77\* |
| pilot_v4 pow10 | 64.17\* | 10.76\* | 6.85\* | 4.04\* | 2.89\* | 2.25\* | 1.85\* | 1.57\* | 1.36\* | 1.20\* | 1.08\* | 1.00 |
| BSL pow10 | 1.02 | 1.02 | 1.02 | 1.01 | 1.01 | 1.01 | 1.02 | 1.02 | 1.02 | 1.02 | 1.02 | 1.01\* |

Against BSL, it loses at worst 6%, but wins big much more often.
To reduce the loss, dispatch can be sped up: use every 4th register in density and bsl_verified. But that helps only if
the density is uniform.

The worst v4 miss I found: 512 dense, 512 empty, then everything is dense until the end of the cycle (17 * 4096).

const size_t cycle = 17 * 4096;
for (size_t i = 0; i < n; ++i) {
dst[i] = i % cycle >= 512 && i % cycle < 1024;
}

This hurts most with the cheapest `f` (with `d_max > 0`, of course):

const auto a = vdupq_n_f32(0.5f);
#pragma unroll
for (int i = 0; i < 13; ++i) x = vfmaq_f32(a, x, a);
return x;

| thd | 0 | 1 |
|--------------|--------:|--------:|
| tiled detour | 38.29 | 9.44 |
| pilot v3 | 74.72 | 15.81\* |
| pilot v3.5 | 74.84 | 15.04 |
| pilot v4 | 74.89\* | 9.84 |
| BSL | 15.61 | 15.60 |

v3.5's biggest loss is limited by BSL, while v4 is limited by detour. And a wrong detour is the cheaper mistake.
So v4 isn't always better, but its misses are less severe.

This problem has no perfect solution. Any dispatch algo can be countertested.

Full code: [godbolt](https://godbolt.org/z/nKhPjEn6v).


https://redd.it/1v94b7w
@r_cpp
godbolt.org Compiler Explorer - C++ (armv8-a clang (trunk)) consteval auto compress_table() { std::array<std::array<uint8_t, 16>, 16> index; std::array<uint8_t, 16> row; row.fill(255); index.fill(row); for (size_t idx = 0; idx < 16; ++idx) { size_t j = 0; for (size_t i = 0; i <…
More from @r_cpp
  1. Sep 26, 2026Token Sequence Injection & Modern Macros: The Most Game-Changing Compile-Time Feature in C…
  2. Sep 25, 2026myStringStream.str("") Considered Harmful Under C++20 I was recently looking at some (rath…
  3. Sep 20, 2026A clever branch free optimization I'm the developer of memlz which is an extremely fast co…
  4. Sep 15, 2026Inside Boost.PolyCollection https://bannalia.blogspot.com/2026/09/inside-boostpolycollecti…
  5. Sep 11, 2026C++26: Standard Library Hardening Experiments https://www.cppstories.com/2026/hardening-ex…
  6. Sep 11, 2026MSVC C++23: constexpr cmath with LLVM Libc https://devblogs.microsoft.com/cppblog/msvc-c23…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →