makes a measurable difference
The experimental change is deliberately narrow.
It remembers a previously resolved buffer interval and whether the resolved neighboring face has a box. At the dedicated box-end check, a valid cache hit confirming that the neighbor still has a box can bypass the full face lookup.
Otherwise, it uses the original path.
It does not remove borders, simplify the requested font attributes, replace text with images, or return a cached face ID to arbitrary callers. Strings and display-vector paths are not bypassed.
The cache checks window/buffer identity, modification counters, narrowing bounds, and face-change flags, and is invalidated on iterator initialization.
**The actual bypass requires only a small change to the existing path.** That is why I think the underlying behavior deserves attention: a narrowly scoped change can eliminate a substantial amount of work without intentionally reducing the decorations being displayed.
I do not mean that a production-quality fix is necessarily simple. Most of the remaining difficulty is proving that the shortcut is correct across redisplay’s edge cases.
# Results in the org-modern workload
The final comparison used 1,000 dense Org content groups, with baseline and fast-path modes in the same instrumented executable.
Each mode ran in a fresh GUI process, serially, with mode order rotated across three repetitions. GC remained enabled.
An edit sample includes a modification and restoration, each followed by redisplay.
|Metric|Baseline|Experimental fast path|
|:-|:-|:-|
|Mean body-edit pair|15.60 ms|11.86 ms|
|Mean timestamp-edit pair|15.11 ms|11.59 ms|
|Mean forced redraw|6.99 ms|5.20 ms|
Across these phases:
* Buffer face queries fell by approximately **26.5%**.
* Font-spec copying fell by approximately **42.3%**.
* Vector-cell allocation fell by **41.6–42.2%**.
* Paired mean-time reductions ranged from **23% to 26%**.
These are synchronous operation timings, not physical input-to-screen latency or FPS.
**The long tail remains.** Editing p95 stayed around 49 ms. The patch reduced work and GC frequency; it did not eliminate all pauses.
# How much correctness checking has been done?
Before actually skipping queries, I ran a verification mode that made the cache prediction and then executed the original lookup for comparison.
Crucially, this verification mode used the fast path’s cache-maintenance behavior: it did not refresh the cache on hits that would have skipped the original query.
Across 18 scenarios, it compared approximately **25.54 million cache hits** without a box-state mismatch. The scenarios included bidi, composition, overlays, display replacements, invisibility, edits, face changes, remapping, narrowing, two windows, and scrolling.
Window ranges, visible-text pixel dimensions, and sampled character geometry also agreed between modes. Deliberately incorrect predictions were detected by the verification harness.
However:
* Repeated cache hits are not millions of independent test cases.
* Geometry checks do not establish pixel-perfect border rendering.
* Multi-frame, different-buffer, dynamic-face, and long-running interactive cases need more coverage.
* A few small query-count residuals remain unexplained and are retained in the results.
* An uninstrumented performance comparison is still needed.
# One correction from the investigation
An earlier experiment appeared to improve allocation by flattening face inheritance. That conclusion was wrong: the flattening happened before org-modern updated the parent face’s box attribute, so the supposedly equivalent configuration was missing a border.
Preserving the box removed that apparent benefit.
That mistake helped isolate the actual trigger, but the inheritance-flattening explanation has been withdrawn. The fast-path results above come from a separate experiment that retains the requested box and font attributes.
# Why I am posting this
My main question is about **the internal redisplay path**, rather than how to optimize org-modern’s Lisp code:
>Does box-boundary detection
Post #28690
9