Spread shared node marking across all CPUs regardless of readers

Marking used one thread per reader, so input that was read serially,
which all goes to the first reader, was marked by a single thread.
Instead, use the readers' indexes to divide all the geometry into
ranges of about the same size at feature boundaries, and have each
thread take the next unmarked range until there are none left.

With 4 CPUs and input read by one reader, this makes marking about
3x faster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P2nBqZisxNQfmEmon3vE9v
This commit is contained in:
Claude
2026-09-23 20:02:49 +00:00
parent 58ad6e3008
commit 9ff3f92e6e
2 changed files with 93 additions and 16 deletions
+5
View File
@@ -19,6 +19,11 @@
each other in the list; search it with `std::lower_bound` instead of `bsearch`;
and size the Bloom filter in front of it by the number of nodes, with three
bits for each node within one 64-bit word, so that it usually fits in the cache.
* Spread the marking of shared nodes across all the CPUs even when all
the features were read by one reader, by dividing the geometry into
ranges at feature boundaries instead of giving each reader's geometry
to its own thread.
* Speed up `encode_quadkey` with a branch-free bit interleave.
# 2.82.0