A small change in xdisp.c reduced allocations by ~42% in a boxed-text workload
**I think there is a performance issue worth investigating in how** xdisp.c **checks box boundaries.** In the path I tested, determining whether a visual neighbor still has a box repeatedly invokes full face resolution, including font-related merging and temporary allocation.
A small, experimental change that bypasses some of these queries reduced vector-cell allocation by approximately **42%** and mean synchronous edit/redraw time by **23–26%** in a dense org-modern workload.
The interesting part is not just that org-modern became faster. **The allocation behavior also reproduces in a plain fundamental-mode buffer, without Org, org-modern, SVGs, overlays, or display replacements.**
This is an experimental core change, not a finished patch or a general “Emacs is 25% faster” claim.
# The question: how much work should a box-boundary check need?
The relevant path in the tested revision is roughly:
get_next_display_element
→ check whether the current element ends a box run
→ face_before_or_after_it_pos
→ face_at_buffer_position
→ merge face/font attributes
→ copy font specifications
→ inspect the resulting face's box state
For the ordinary branch under investigation, the result is used to decide whether the neighboring face has no box and the current box run should end.
That appears to create an opportunity for disproportionate work: **resolving and merging a complete face repeatedly to answer a much narrower boundary question.**
Emacs already has a realized-face cache. This is not a claim that face caching is missing. The issue is that, along this path, attribute merging and temporary allocation can happen **before** the final face-cache lookup.
# This reproduces without org-modern
I started with org-modern because it provides a useful real-world source of densely decorated text. I then reduced the case to repeated text with face properties in fundamental mode.
No Org, font-lock, SVG, overlays, or `display` substitutions were involved.
For 60 forced redraws:
|Configuration|Allocated vector cells|
|:-|:-|
|Colors only, no box|13,320|
|Colors only, with box|13,320|
|Font-related attributes, no box|2,016,360|
|Same font-related attributes, with box|32,188,320|
|Inherited font-related attributes, with box|32,188,320|
The relevant font-attribute cases had the same selected font, visible character count, and line height.
Adding a box increased vector-cell allocation by approximately **16×** in the font-attribute case, while adding a box to the colors-only case did not.
So the finding is specifically about the interaction between **box-boundary checking and font-related face processing**, not that every box is expensive or that face inheritance alone explains the allocation.
“Vector cells” here means cumulative allocation, not vector objects or retained memory.
# Direct counters identify the allocation source
To check whether this was actually the suspected path, I instrumented the native code.
For 30 redraws of the minimal fixture:
|Configuration|Buffer face queries|`copy_font_spec` calls|Allocated vector cells|
|:-|:-|:-|:-|
|Font attributes, no box|51,960|77,490|1,008,630|
|Same font attributes, with box|464,460|1,237,950|16,094,610|
In the boxed case, vector cells allocated directly by `copy_font_spec` accounted for approximately **99.992%** of the measured vector-cell allocation.
More importantly, across three repetitions, the increase in total vector-cell allocation between the unboxed and boxed cases exactly matched the increase attributed to font-spec copying inside the instrumented box-check scope.
That gives a more specific explanation than “decorated buffers allocate a lot”:
Box-boundary checks
→ repeated neighboring-face resolution
→ font-spec copying
→ temporary allocation
→ more GC pressure
These allocation percentages apply to the minimal fixture, not to all Emacs memory use or all org-modern costs.
# A small fast path
Post #28689
7