Lessons learned from 2 years of operating a C++ MMO game server in production
I've been working as a server team lead and technical director on a large-scale MMO for the past 2+ years, from launch through live service. Our server runs on Windows, written in C++20 (MSVC) with an in-house Actor model, protobuf serialization, and our own async framework built on C++20 coroutines. I wanted to share some hard-won lessons that might be useful to others writing performance-critical C++ in production.
# 1. A data race in std::sort comparators that led to memory corruption
This was our most painful category of bugs. **The root cause was a data race**, but it manifested in a surprising way: a violation of [strict weak ordering](https://en.cppreference.com/w/cpp/named_req/Compare.html) inside `std::sort`, which invokes **undefined behavior** — and in practice, this means **memory corruption**, not a nice crash.
**The trap:** We had NPCs sorting nearby targets by distance:
std::sort(candidates.begin(), candidates.end(),
[&self](const auto& lhs, const auto& rhs) {
auto lhs_len = Vec2(self->pos() - lhs->pos()).Length();
auto rhs_len = Vec2(self->pos() - rhs->pos()).Length();
return lhs_len < rhs_len;
});
This looks correct at first glance — `<` on floats is a valid strict weak ordering, right? The real problem was that **the objects' positions were being updated by other threads while** `std::sort` **was running.** Each comparator call recomputed distances from live positions, so `compare(a, b)` could return a different result depending on *when* it was called. This breaks transitivity: if `a < b` and `b < c` were evaluated at different moments with different position snapshots, `a < c` is no longer guaranteed. `std::sort` assumes a consistent total order — violate that and it reads out of bounds, corrupting memory.
We fixed it by pre-computing distances into a `vector<pair<Fighter*, float>>` *before* sorting, so the comparator operates on stable, snapshot values:
for (auto& [fighter, dist] : candidates) {
dist = Vec2(self->pos() - fighter->pos()).Length();
}
std::sort(candidates.begin(), candidates.end(),
[](const auto& lhs, const auto& rhs) {
return lhs.second < rhs.second; // comparing cached values — safe
});
We found **multiple instances** of this pattern in our codebase, all causing intermittent memory corruption that took weeks to track down.
**Lesson:** A comparator that calls any non-const external state is a ticking time bomb. Pre-compute, then sort.
# 2. Re-entrancy is a silent killer
In an MMO, game logic chains are deeply interconnected. Example:
1. Player uses a skill that heals self on cast
2. HP recovery triggers a debuff check
3. Debuff applies stun, which cancels the skill
4. Skill cancel calls `Stop()` which deletes the current state
5. Control returns to the skill's `BeginExecution()` which accesses the now-deleted state
We developed three strategies:
|Strategy|When to use|
|:-|:-|
|**Prevent** (guard with assert on re-entry)|FSMs, state machines — crash immediately if re-entered|
|**Defer** (queue for later execution)|When completeness matters more than immediacy|
|**Allow** (remove from container before processing)|Hot paths where performance is critical|
The "prevent" approach uses a simple RAII guard:
class EnterGuard final {
explicit EnterGuard(bool* entering) : entering_(entering) {
DEBUG_ASSERT(*entering_ == false); // crash on re-entry
*entering_ = true;
}
~EnterGuard() { *entering_ = false; }
bool* entering_;
};
This caught a real live-service crash where an effect was being duplicated in a heap due to re-entrant processing during abnormal-status stacking policy evaluation.
# 3. Actor model on a 72-core machine: beware the Monolithic Actor
Our server uses a **single-process Actor model on 72-core machines** — not a distributed actor system across multiple nodes. The goal is to maximize utilization of all cores within one
Post #24772
15