Announcing TooManyCooks: the C++20 coroutine framework with no compromises
[TooManyCooks](https://github.com/tzcnt/TooManyCooks) aims to be the fastest general-purpose C++20 coroutine framework, while offering unparalleled developer ergonomics and flexibility. It's suitable for a variety of applications, such as game engines, interactive desktop apps, backend services, data pipelines, and (consumer-grade) trading bots.
It competes directly with the following libraries:
* tasking libraries: libfork, oneTBB, Taskflow
* coroutine libraries: cppcoro, libcoro, concurrencpp
* asio wrappers: boost::cobalt (via [tmc-asio](https://github.com/tzcnt/tmc-asio))
# TooManyCooks is Fast (Really)
I maintain a comprehensive suite of benchmarks for competing libraries. You can view them here: ([benchmarks repo](https://github.com/tzcnt/runtime-benchmarks)) ([interactive results chart](https://fleetcode.com/runtime-benchmarks/))
TooManyCooks beats every other library (except libfork) across a wide variety of hardware. I achieved this with cache-aware work-stealing, lock-free concurrency, and many hours of obsessive optimization.
TooManyCooks also doesn't make use of any ugly performance hacks like busy spinning ([unless](https://www.fleetcode.com/oss/tmc/docs/v1.4/executors/ex_cpu.html#_CPPv4N3tmc6ex_cpu9set_spinsE6size_t) you [ask it to](https://www.fleetcode.com/oss/tmc/docs/v1.4/data_structures/channel.html#_CPPv4N3tmc8chan_tok18set_consumer_spinsE6size_t)), so it respects your laptop battery life.
# What about libfork?
I want to briefly address libfork, since it is typically the fastest library when it comes to fork/join performance. However, it is arguably not "general-purpose":
* ([link](https://github.com/ConorWilliams/libfork?tab=readme-ov-file#-welcome-to-libfork--)) it requires arcane syntax (as a necessity due to its implementation)
* it requires every coroutine to be a template, slowing compile time and creating bloat
* limited flexibility w.r.t. task lifetimes
* no I/O, and no other features
Most of its performance advantage comes from its custom allocator. The recursive nature of the benchmarks prevents HALO from happening, but in typical applications (if you use Clang) HALO will kick in and prevent these allocations entirely, negating this advantage.
TooManyCooks offers the best performance possible without making any usability sacrifices.
# Killer Feature #1 - CPU Topology Detection
As every major CPU manufacturer is now exploring disaggregated / hybrid architectures, legacy work-stealing designs are showing their age. TooManyCooks is designed for this new era of hardware.
It uses the CPU topology information exposed by the [libhwloc](https://www.open-mpi.org/projects/hwloc/) library to implement the following automatic behaviors:
* ([docs](https://www.fleetcode.com/oss/tmc/docs/v1.4/executors/ex_cpu.html#hardware-optimized-defaults)) locality-aware work stealing for disaggregated caches (e.g. Zen chiplet architecture).
* ([docs](https://www.fleetcode.com/oss/tmc/docs/v1.4/executors/topology.html#_CPPv4N3tmc8topology12cpu_topology19container_cpu_quotaE)) Linux cgroups detection sets the number of threads according to the CPU quota when running in a container
* If the CPU quota is set instead by selecting specific cores (`--cpuset-cpus`) or with [Kubernetes Guaranteed QoS](https://docs.starlingx.io/r/stx.6.0/usertasks/kubernetes/using-kubernetes-cpu-manager-static-policy.html), the hwloc integration will detect the allowed cores (and their cache hierarchy!) and create locality-aware work stealing groups as if running on bare metal.
Additionally, the topology can be queried by the user ([docs](https://www.fleetcode.com/oss/tmc/docs/v1.4/executors/topology.html)) ([example](https://github.com/tzcnt/tmc-examples/blob/44ce66de8e74107c803645c82ddd88a7be35fd29/examples/hwloc/topo.cpp)) and APIs are provided that let you do powerful things:
*
Post #24739
13