Make the remaining shared node lookups faster

Encode the shared nodes as quadkeys, as the comment on struct node
already said they were, with a branch-free bit interleave, so that
vertices that are near each other are near each other in the sorted list
too. Search the list with an inlined std::lower_bound instead of bsearch.

Size the Bloom filter by the number of nodes, at about 16 bits each and
at most 32MB, instead of always using 34MB, so that it can usually stay
in the cache, and set three bits for each node, chosen by a mixing hash,
within a single 64-bit word, so that each check still touches only one
cache line but has far fewer false positives than a single bit.

In the pass that marks the vertices with whether they are shared nodes,
this is about 1.7x faster with 70 thousand nodes and 1.9x faster with
3 million.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P2nBqZisxNQfmEmon3vE9v
This commit is contained in:
Claude
2026-09-23 19:46:09 +00:00
parent 175036930d
commit e1795ebe32
6 changed files with 88 additions and 12 deletions
+13 -6
View File
@@ -2136,8 +2136,8 @@ std::pair<int, metadata> read_input(std::vector<source> &sources, char *fname, i
fprintf(stderr, "Merging nodes \r");
}
// This is sized once the number of nodes is known, below
std::string shared_nodes_bloom;
shared_nodes_bloom.resize(34567891); // circa 34MB, size nowhere near a power of 2
// Sort nodes that can't be simplified away; scan the list to remove duplicates
@@ -2204,11 +2204,6 @@ std::pair<int, metadata> read_input(std::vector<source> &sources, char *fname, i
fwrite_check((void *) &here, sizeof(here), 1, shared_nodes, &nodepos, "shared nodes");
written = here;
size_t bloom_ix = here.index % (shared_nodes_bloom.size() * 8);
unsigned char bloom_mask = 1 << (bloom_ix & 7);
bloom_ix >>= 3;
shared_nodes_bloom[bloom_ix] |= bloom_mask;
#if 0
unsigned wx, wy;
decode_quadkey(here.index, &wx, &wy);
@@ -2227,6 +2222,18 @@ std::pair<int, metadata> read_input(std::vector<source> &sources, char *fname, i
perror("mmap nodes");
exit(EXIT_MEMORY);
}
// Size the Bloom filter at about 16 bits per node, so that for a moderate
// number of nodes it can stay in the cache while it is being checked,
// but no more than 32MB.
size_t nnodes = nodepos / sizeof(node);
size_t bloom_size = std::min(nnodes * 2, (size_t) 32 * 1024 * 1024);
bloom_size = std::max((bloom_size + 7) / 8 * 8, (size_t) 64);
shared_nodes_bloom.resize(bloom_size);
for (size_t i = 0; i < nnodes; i++) {
add_shared_node_to_bloom(shared_nodes_bloom, shared_nodes_map[i].index);
}
}
fclose(node_out);