Make the remaining shared node lookups faster

Encode the shared nodes as quadkeys, as the comment on struct node
already said they were, with a branch-free bit interleave, so that
vertices that are near each other are near each other in the sorted list
too. Search the list with an inlined std::lower_bound instead of bsearch.

Size the Bloom filter by the number of nodes, at about 16 bits each and
at most 32MB, instead of always using 34MB, so that it can usually stay
in the cache, and set three bits for each node, chosen by a mixing hash,
within a single 64-bit word, so that each check still touches only one
cache line but has far fewer false positives than a single bit.

In the pass that marks the vertices with whether they are shared nodes,
this is about 1.7x faster with 70 thousand nodes and 1.9x faster with
3 million.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P2nBqZisxNQfmEmon3vE9v
This commit is contained in:
Claude
2026-09-23 19:46:09 +00:00
parent 175036930d
commit e1795ebe32
6 changed files with 88 additions and 12 deletions
+5
View File
@@ -14,6 +14,11 @@
exactly on one of these new points, and means that only
vertices that come back from a prefilter, or that are created by polygon
cleaning, still need to be looked up in the list of shared nodes during tiling.
* Make the remaining lookups of shared nodes faster: sort the list of nodes
by quadkey, as its comment always said, so that nearby vertices are near
each other in the list; search it with `std::lower_bound` instead of `bsearch`;
and size the Bloom filter in front of it by the number of nodes, with three
bits for each node within one 64-bit word, so that it usually fits in the cache.
# 2.82.0