[Show and Tell] u8ility: A C++20/23 header-only, zero-allocation UTF-8 view library
Hi r/cpp,
This is my first post in this subreddit, so I'll try my best to follow the guidelines and provide some value!
Back in 2021, I started working on a simple UTF-8 library to avoid the dependency overhead required for basic codepoint interaction in my project. \
I recently dusted it off and decided to completely rewrite it to create the thinnest possible wrapper around std::string_view, to meet modern C++ standards that provide correct iteration of code points, focusing solely on performance and ergonomics.
### Key Design Principles:
- **Header-only**: Ease of use by providing complete details on what's under the hood
- **Zero-Allocation**: The core character type (`u8::mchar`) is a small, stack-based value type (max 5 bytes). It avoids heap allocation entirely during iteration
- **Cache-Friendly**: By avoiding pointers and virtual calls, it ensures high cache locality when iterating
- **Constexpr**: Allows encoding, decoding, and basic character validation at compile-time
- **Ergonomic**: Provides an u8::u8_view that works flawlessly with range-based for loops
I believe this offers an efficient alternative to full-featured libraries *when you just need quick, safe access to UTF-8 characters within existing `std::string` data*.
I'd love your roast/feedback on the current implementation. I'm especially interested in whether the char8_t vs char interoperability feels correct and how I could further improve validation logic without breaking the zero-allocation rule.
Here is [the Github link](https://github.com/lmela0/u8ility) 🙏
`https://github.com/lmela0/u8ility`
https://redd.it/1pilv04
@r_cpp
Post #24458
20