I’ve been working on a Vulkan optimisation layer and static library called GoCL. The C++ side ended up more interesting than I expected, so I wanted to share a few design choices.
It’s a Vulkan proxy (
vulkan_proxy.so / vulkan‑1.dll) and a static library (GoCL_core.a) that adapts rendering to the GPU’s actual capabilities at runtime. The proxy intercepts Vulkan calls and modifies them before they reach the driver, while the static library provides the same logic directly to engines that link against it. All C++20.C++ details that might interest this sub:
Capability oracle – on device creation, the engine queries every relevant physical device feature (texture compression, float16 support, descriptor indexing tiers, memory budget, etc.) and stores the result in a plain `DeviceCapabilities` struct. Every subsystem reads from that struct; there’s no virtual dispatch or runtime capability checks on the hot path. Decisions are data‑driven and branch‑predictable.
Proxy dispatch chain – the proxy hooks Vulkan entry points by patching a dispatch table (the loader’s
vkGetDeviceProcAddr pattern). The table is populated at vkCreateDevice time. On Linux the library is linked with -Wl,-z,now so all dynamic symbols are resolved eagerly at load time. Callgrind confirms zero lazy‑binding overhead in the render loop – dlopen/dlsym calls appear only in initialisation functions.SPIR‑V rewriting – unsupported shader features (FP16, oversized descriptor sets) are detected and the binary is rewritten before the driver sees it. The rewriter is a lightweight in‑memory pass that widens FP16 ops to FP32 and splits large descriptor bindings. No external compiler, no LLVM dependency; it’s just a handful of C++ functions operating on the SPIR‑V binary blob.
Lazy‑loaded transcoder – the heavy Basis Universal encoder (used for ASTC→ETC2 texture conversion) is separated into
gocl_transcoder.so, loaded via dlopen / LoadLibrary only when an ASTC texture is encountered. The static initialisers of the transcoder are therefore never invoked unless actually needed, keeping the proxy’s baseline footprint minimal.Zero per‑frame overhead – Callgrind instruction counts on a GTX 960M running 75+ Sascha Willems Vulkan examples: the proxy has \~3.8% fewer total instructions than native (3.806B vs 3.955B). FPS and frame times are identical within noise. All the work happens once at pipeline creation or texture load; the render loop is untouched.
Dual‑licensed under Apache‑2.0 OR MIT. Tested in CI with Mesa Lavapipe.
Architecture write‑up (covers the capability model, shader emulation, proxy internals, and the lazy‑load design):
GoCL – Architecture
GitHub: GoCL
Happy to discuss the C++ side – the capability oracle and the SPIR‑V rewriter in particular were fun to keep simple and predictable.
https://redd.it/1uu3392
@r_cpp