Writing gRPC Clients and Servers with C++20 Coroutines (Part 1)
# The Callback Hell of C++ Async Programming
If you have ever written network programs in C++, you are probably very familiar with this scenario —
You need to send request B after request A completes, then send request C after B completes. So you end up writing code like this:
client.send(requestA, [](Response a) {
client.send(requestB(a), [](Response b) {
client.send(requestC(b), [](Response c) {
// Finally made it here...
});
});
});
This is the infamous "callback hell". The deeper the nesting, the harder the code is to read, and error handling becomes a nightmare — each layer of callback needs to handle errors independently. One missed callback call in a branch, and the entire request silently disappears without a trace.
The callback pattern also introduces another thorny problem: lifetime management. When a callback executes, every captured object must still be alive, yet their lifetimes are often hard to predict. You end up littering the code with `shared_ptr` and `shared_from_this()` just to keep things alive — ugly and unnecessarily expensive.
The most maddening part is cancellation. Imagine a user closes a page and you want to cancel an in-progress multi-step operation, but every step in the callback chain may have already started. How do you propagate the cancel signal? How do you ensure all resources are properly cleaned up? This is nearly an unsolvable problem.
If you have worked with async programming in `Go` or `Python`, you have probably envied their concise coroutine syntax — expressing asynchronous logic with synchronous-looking code. The good news is that C++20 introduced coroutines, and C++ programmers finally have a proper tool for the job.
# The Async World of gRPC
Before diving in, let's take a look at what gRPC's async model looks like and what new challenges it brings.
The C++ implementation of gRPC provides two async APIs: the classic `CompletionQueue`\-based model and the Reactor pattern callback model. Clients typically use the Reactor pattern — inheriting various `ClientXxxReactor` classes and implementing callbacks; servers typically use the `CompletionQueue` model — registering operations with the queue and polling for results.
For a client-side server-streaming RPC, you need to inherit `grpc::ClientReadReactor<T>` and implement several callbacks:
class MyReader : public grpc::ClientReadReactor<Response> {
public:
void OnReadDone(bool ok) override {
if (ok) {
// Process the received data, then call StartRead again
StartRead(&response_);
}
}
void OnDone(const grpc::Status &status) override {
// RPC finished
}
private:
Response response_;
};
The server side is even more complex. With the `CompletionQueue`\-based model, you need to manually maintain a state machine:
enum class State { WAIT, READ, WRITE, FINISH };
class Handler {
void Proceed() {
switch (state_) {
case State::WAIT:
// Register the next request, transition to READ state
break;
case State::READ:
// Read data, decide whether to continue reading or writing
break;
// ...
}
}
};
This state machine must be maintained by hand. A single wrong state transition can cause data loss or a crash. Every time you need to implement this logic for a new RPC method, you have to go through the same ordeal.
Worse still, both approaches make it very hard to support cancellation cleanly. In the Reactor model, cancellation means calling `context->TryCancel()` and waiting for the `OnDone` callback; in the `CompletionQueue` model, state transitions after cancellation require extra care.
The root of the problem is: **gRPC's async API is designed around callbacks, while we want to organize code around
Post #25031
15