Title: UltrafastSecp256k1 — Zero-dependency C++20 secp256k1 with ASM backends, CUDA/ROCm, and 9 platform targets
After 3+ years of development, I'm sharing **UltrafastSecp256k1** — a complete secp256k1 implementation in modern C++20 with zero external dependencies.
**What makes it different:**
* **Zero dependencies** — SHA-256, SHA-512, Keccak-256, RIPEMD-160, Base58, Bech32 all implemented from scratch
* **Multi-backend assembly** — x64 MASM/GAS (BMI2/ADX), ARM64 (MUL/UMULH), RISC-V (RV64GC)
* **SIMD** — AVX2/AVX-512 batch operations, Montgomery batch inverse
* **GPU** — CUDA (4.63M kG/s), OpenCL (3.39M kG/s), ROCm/HIP
* **Constant-time** — Separate `secp256k1::ct` namespace, Montgomery ladder, no flag switching
* **Hot path contract** — Zero heap allocations, explicit buffers, fixed-size POD, in-place mutation
* **9 platforms** — x64, ARM64, RISC-V, ESP32, STM32, WASM, iOS, Android, CUDA/ROCm
**Protocol coverage:**
ECDSA (RFC 6979, low-S, recovery), Schnorr (BIP-340), MuSig2, FROST, Adaptor sigs, Pedersen commitments, Taproot, BIP-32/44, Silent Payments (BIP-352), ECDH, multi-scalar multiplication, batch verification, 27+ coin address generation.
**Design philosophy:**
**Algorithm > Architecture > Optimization > Hardware**
**Memory traffic > Arithmetic**
**Correctness is absolute; performance is earned**
Everything is `constexpr`\-friendly where possible. No exceptions, no RTTI, no virtual calls in compute paths.
200+ tests, fuzz harnesses, known vector verification.
Happy to answer questions about the architecture, assembly backends, or GPU implementation.
https://redd.it/1r5c9pe
@r_cpp
Post #24789
23