Commit Graph
100 Commits
Author SHA1 Message Date
4f2621186a Convert jsonpull to C++ with shared_ptr and std::vector/std::string (#388)
* Rename to jsonpull.cpp

* Clear for merge

* Convert jsonpull to C++ with shared_ptr and std::vector/std::string

Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.

The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Subclass json_object so primitives shrink from 168 to 24 bytes

The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:

  json_object (base, TRUE / FALSE / NULL)   24 bytes
  json_number                                48 bytes
  json_string                                48 bytes
  json_array  (empty)                        48 bytes
  json_hash   (empty)                        72 bytes

Other size wins along the way:

* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
  was 16 bytes per node). json_pull now keeps an explicit
  container_stack and the parser no longer needs to resurrect a
  shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
  original std::make_shared<json_xxx> call, so destroying a
  shared_ptr<json_object> still runs the right subclass dtor.

The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Store hash key/value pairs in one ordered vector

Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:

    for (auto &[k, v] : o->entries()) { ... }

Side effects:

* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
  header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
  single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
  (`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
  original pattern relied on `nprop = 0` to short-circuit iteration on
  a null or non-hash `properties`, the rewrite now guards the loop
  explicitly with `if (o->type == JSON_HASH)` so that calling
  entries() doesn't trip the asserting downcast.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Move parser-only `expect` state out of json_object

`expect` was only meaningful while the parser was building a container,
and only ever read or written from jsonpull.cpp itself; once parsing
finished it was dead weight on every JSON_ARRAY and JSON_HASH (and
present-but-unused on every primitive too). Move it into the parser's
container stack, alongside the shared_ptr to the container it pertains
to:

    struct json_pull::parse_frame {
        json_object_ptr container;
        json_type       expect;
    };
    std::vector<parse_frame> container_stack;

The base class now only carries data-model state (parent, parser, type).
No external caller depended on `expect`, so no sweep was needed outside
jsonpull.cpp.

This change does not, in itself, shrink any json_object: the 4-byte
`expect` field used to live at offset 20 inside the base, where it was
already being eaten by alignment padding for the 8-byte-aligned first
member of every subclass (std::string, std::vector, double). The win is
in the data model, not the byte count -- the 4-byte hole is still
there, but it is now available for a future subclass whose first member
is small enough to slot into it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Discriminate json_number's three numeric slots into one union

json_number used to carry three parallel 8-byte fields (a double plus
both a 64-bit unsigned and a 64-bit signed slot for the large-integer
cases) even though at most one of the integer slots is ever the
canonical value for any given number. Collapse them into a
discriminated union:

    enum repr_t { REPR_DOUBLE, REPR_LARGE_UNSIGNED, REPR_LARGE_SIGNED };
    repr_t repr;
    union { double d; unsigned long long u; long long s; } value;

Callers keep the same read API: number() returns the appropriate
double, large_unsigned() returns the ull (or 0 if not currently stored
that way), large_signed() likewise. Writes go through new set_number /
set_large_unsigned / set_large_signed methods that keep the
discriminator and the union value in sync.

This was prompted by an observation that moving json_type to the end
of the object should shrink things via tail-padding reuse. Empirically
the type-at-end rearrangement saves nothing on its own (every
subclass payload is 8-byte aligned so it can't slot into the 4-byte
tail), but the discriminated-number redesign hits the same idea from
a different direction: adding the 4-byte `repr` to json_number makes
the class non-standard-layout, which lets the Itanium ABI pack `repr`
into the base's 4-byte tail padding at offset 20. The union value
then starts at the natural offset 24, and json_number ends at offset
32 -- a 33% reduction.

Per-node sizes:
  json_object (TRUE/FALSE/NULL)  24 bytes
  json_number                    32 bytes  (was 48)
  json_string                    48 bytes
  json_array                     48 bytes
  json_hash                      48 bytes

Numbers dominate real GeoJSON (every coordinate is one), so the net
memory win on a typical parse is substantial.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix bugs flagged in code review of jsonpull C++ port

- jsontool.cpp `out()`: route JSON_NUMBER (and anything else non-string)
  through `json_stringify` instead of `o->string()`, which now asserts
  on a non-string type and would crash `--extract` on numeric attributes.
- geojson.{hpp,cpp} `json_end_map`: take `json_pull_ptr` by reference so
  the caller's shared_ptr is released, null-guard before touching
  `jp->source`, and clear `jp->source` after delete to avoid a dangling
  pointer.
- jsonpull/jsonpull.cpp: low-surrogate range check was comparing the
  outer-loop byte `c` instead of the parsed code unit `ch`, breaking
  surrogate-pair decoding for some \\uXXXX escapes. Pre-existing bug
  preserved across the port.
- tile-join.cpp `handle_vector_layers`: require the field value to have
  type JSON_STRING (and the key to be non-null) before calling
  `string()`; the previous truthy `type` check would assert on a
  non-string value.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add jsonpull regression test for surrogate-pair decoding

Covers the `c` vs `ch` bug fixed in the previous commit: parsing
"\uD83D\uE000" (a valid high surrogate followed by a non-surrogate
BMP code point) used to mis-classify U+E000 as a low surrogate and
combine the two units into U+1F400 (F0 9F 90 80). The fixed code
flushes the stale high surrogate as standalone CESU-8 (ED A0 BD)
and then encodes U+E000 normally as EE 80 80. Verified the test
fails under the pre-fix logic.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Cheap perf wins in jsonpull C++ port

Profiling tl_2022_us_county.json (sample(1) on Apple Silicon) showed
~38% of parse time in allocator work and ~14% in std::string::push_back
during string-token construction. These changes target the low-hanging
fruit from that profile:

- Pre-reserve 2 slots in json_array and 4 slots in json_hash so
  coordinate `[x, y]` pairs and typical GeoJSON property maps avoid
  the 0 -> 1 -> 2 -> 4 vector-growth chain (and the shared_ptr copies
  it incurs).
- Reuse a parser-wide std::string buffer for JSON_STRING tokens
  instead of constructing a fresh local std::string per token. The
  buffer is cleared (capacity preserved) at the start of each token
  and copied into the final json_string, so once it has grown to the
  longest string seen it stops reallocating entirely.
- std::move the freshly-created container shared_ptr into the parser
  container stack in the `[` and `{` handlers, and move it out of the
  frame on the matching `]` / `}`. Each move skips one atomic
  inc/dec round-trip per container open and close.

On a tl_2022_us_county.json benchmark (4-iter user-time mean, Apple
Silicon, /usr/bin/time):
- main baseline:                              ~8.17s
- jsonpull-cpp before these changes:          ~10.90s  (+33%)
- jsonpull-cpp with these changes:            ~9.33s   (+14%)

So this commit recovers roughly half of the post-port regression.
The remaining gap is dominated by shared_ptr atomic refcount traffic
on the parse tree and per-node heap allocations, which would require
the larger unique_ptr/arena reworks to address.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Make json_free actually free the subtree

In the C++ port, json_free was just `o.reset()`, which dropped the
caller's reference but left the subtree alive: the parent's vector
slot kept it allocated, and for line-delimited streams the parser's
jp->root co-owned it until the next top-level value started parsing.
That defeated the geojson-loop pattern of calling json_free on each
feature after serializing it, which is supposed to release the
feature so it doesn't sit in memory while subsequent ones are parsed.

Restore the historical "remove this from the tree" semantics by
splicing the node out of its parent (sharing splice_from_parent with
json_disconnect) and clearing jp->root when the node is the parser's
current top-level value, then dropping the caller's reference.

Two unit tests pin this down: a pruning test parses
"[[1, 2], [3, 4], [5, 6]]" element-wise and confirms that calling
json_free on [3, 4] leaves the outer array with just [1, 2] and
[5, 6]; a top-level test uses a weak_ptr observer to confirm that
json_free on the parser's root really destroys the tree.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Migrate jsonpull to unique_ptr ownership

Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.

API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
  json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
  takes ownership; back-pointers are cleared so the subtree can
  outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
  unique_ptr ownership of the most recent top-level value.

Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.

In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.

Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix preprocessor mistakes identified by Copilot

* Make indent

* Skip non-string metadata.json entries instead of reading them as strings

dirmeta2tmp() warned about a metadata entry that was not a string/string
pair and then read it as a string anyway. Under the new type-tagged
accessors that trips the assert in json_object::string(); before them it
reinterpreted the node's storage as a char pointer, which segfaulted for
most values. Either way, tippecanoe-decode and tile-join could not read a
directory tileset whose metadata.json had a numeric minzoom or a nested
object, which is common in metadata.json files written by other tools.

Add the missing continue, and cover it in raw-tiles-test.

pmtilesmeta2tmp() handles the same case correctly but read the key with
string() before its own JSON_STRING check, so the assert would have fired
ahead of the check meant to catch a bad key. Hoist the check above the
read. The parser rejects non-string hash keys, so this is unreachable in
practice; the ordering is what makes the check meaningful.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Don't redefine _GNU_SOURCE in the C++ jsonpull port

The `#define _GNU_SOURCE` carried over from jsonpull.c, where it was
needed to get asprintf() declared. g++ already defines _GNU_SOURCE on the
command line for C++ translation units, so redefining it warns:

    jsonpull/jsonpull.cpp:1: warning: "_GNU_SOURCE" redefined

Guard the define rather than drop it, so platforms whose C++ driver does
not predefine it still get asprintf() declared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add a unit test for json_disconnect

json_disconnect() is documented in jsonpull.h as the supported way to
splice a subtree out of the parser's tree and take ownership of it, but
nothing calls it: read_filter() and parse_filter() used to, and now get
the same guarantee from json_read_tree() clearing back-pointers on the way
out. Cover the behavior rather than leave the primitive dead and untested.

The test pins that the subtree is removed from its parent, that the parser
keeps the rest of the tree, and that the detached subtree stays readable
after the json_pull is destroyed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Correct two stale comments in the jsonpull port

jsonpull.h said a json_number is 40 bytes; it is 32 (json_object is 24,
and the repr discriminator fits in the base class's tail padding, so the
8-byte union lands at offset 24).

plugin.cpp's parse_feature() said `j` is freed only just before returning
or as jp->root at end of stream, but there is a third json_free(j) at the
bottom of the loop, for a complete Feature whose geometry came out empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add a changelog entry and bump the version for the jsonpull rewrite

The rewrite is meant to be behavior-preserving, but it carries four
user-visible bug fixes that warrant release notes: tippecanoe-json-tool
--extract on a numeric attribute, surrogate-pair decoding, tile-join
reading a non-string tilejson field type, and non-string values in a
directory tileset's metadata.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Encode U+FFFF as three bytes instead of an overlong four

The \uXXXX decoder tested `ch < 0xFFFF` before taking the three-byte UTF-8
path, so U+FFFF itself fell through to the four-byte branch and came out as
F0 8F BF BF -- an overlong, and therefore invalid, encoding of a code point
that fits in three bytes.

check_utf8() only checks that continuation bytes look like continuation
bytes, not that a sequence is the shortest form, so nothing downstream
noticed: a GeoJSON attribute containing U+FFFF put invalid UTF-8 into the
output tile, where a strict consumer would reject it.

Since `ch` is parsed from exactly four hex digits it cannot exceed 0xFFFF
on its own, so after this change the four-byte branch is reached only for a
code point assembled from a surrogate pair, which is the only way to name
one above the BMP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Keep the parent links inside a detached jsonpull subtree

json_read_tree() and json_disconnect() cleared both back-pointers on every
node of the subtree they handed out. Clearing `parser` throughout is
necessary -- the json_pull can be destroyed while the subtree lives on, so
a surviving `parser` would dangle -- but clearing `parent` throughout cost
more than it bought.

`parent` is a non-owning raw pointer, so keeping it cannot form a reference
cycle or keep anything alive; there is nothing to leak. And within a
detached subtree it refers to nodes the caller now owns as a single unit,
so it stays valid for exactly as long as the subtree itself. Clearing it
only made the tree unwalkable upwards, and made json_free() and
json_disconnect() silently no-ops on interior nodes of a detached tree,
since both find a node's owner through o->parent.

So clear `parser` everywhere and clear `parent` on the detached root alone,
which is the one that pointed out of the subtree at a node the parser still
owns. Split the old clear_back_pointers() into clear_parser_pointers() plus
a detach_subtree() wrapper that adds the root's `parent`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Cover the U+FFFF encoding and detached-tree parent links

Each of the new assertions fails against the previous behavior, so they
pin the two fixes rather than merely passing alongside them:

  - the U+FFFF test, plus the U+FFFE boundary below it and a surrogate pair
    above it, so the three-byte and four-byte paths are both held in place
  - json_disconnect() leaving the parent links inside the subtree intact
    while clearing the root's
  - json_free() pruning an interior node of a tree whose parser is already
    gone, which only works because those links survive
  - json_free() of a hash value leaving the key paired with a JSON_NULL
    placeholder, which is the documented behavior and not a removal

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Address review notes in jsonpull itself

- json_stringify walked c_str(), so it truncated at an embedded NUL even
  though values are std::string now and carry one through faithfully.
  Range over the string instead; the existing control-character branch
  already escapes a NUL like any other, so the output stays valid JSON.
- Assert that the hash has an entry waiting before add_object() assigns to
  entries().back(). It always does -- JSON_VALUE is only set by a colon,
  which requires a pushed key -- but the derivation is not local.
- Drop fabricate_object(), a pass-through to make_object() with the
  arguments reordered, kept only to preserve the old C name.
- Inline the string_append / string_append_c wrappers over push_back and
  append, and note that json_print_one's JSON_HASH and JSON_ARRAY branches
  are unreachable, since json_print handles both itself.
- json_hash_get's comment said nullptr meant "the matching value is null",
  which reads as JSON null. A JSON null comes back as a JSON_NULL node;
  nullptr means the value slot is not filled in yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Tidy jsonpull call sites flagged in review

- geojson.cpp and attribute.cpp passed key_pool::pool() and
  set_attribute_accum() a c_str() from a std::string, forcing a needless
  reconstruction (and truncating at an embedded NUL). Both overloads take
  std::string, so pass it directly. The geojson.cpp one is the hottest
  loop in the program.
- Replace the hand-maintained counters beside range-for loops in
  attribute.cpp, main.cpp and tile-join.cpp with indexed loops, since the
  index is only wanted for error messages.
- parse_json_args took json_pull_ptr by value and then copied it, costing
  two refcount bumps per construction. Move it.
- Assert that the parser is still attached where geojson.cpp reads
  geometry->parser->line. Only json_read results reach it today, but
  json_read_tree and json_disconnect now clear every parser pointer, so a
  detached tree would null-deref there instead of tripping an assert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Pin the array-splicing fix, and close four test gaps

The existing pruning test does not discriminate: json_read hands back each
container as it completes, so the node it frees is always the most recently
added element of its parent -- the one case the old element-count-vs-byte-
count memmove got right, because it then moved zero bytes. Widening that
test to more elements does not change this; the shape is what matters, not
the size. Verified: the eight-element streaming variant still passes
against the pre-fix code.

Add a test that builds the array first and then prunes element 0 of eight,
asserting the identity of every survivor rather than just the resulting
count. That fails against the pre-fix code deterministically, with
arr[0] == arr[1] and the last element dropped. Note the limitation on the
streaming test so the next reader does not try to strengthen it in place.

Also cover, all previously untested:

- json_free of a hash key, and of both halves of a pair, where the entry
  survives with a JSON_NULL stand-in until both are gone
- repeated json_read_tree over a line-delimited stream, which is what the
  filter loaders and -L / -E do, asserting each detached tree survives the
  next read and the parser's destruction
- json_stringify of a partially-parsed tree, the json_context error path
- json_stringify across an embedded NUL, which fails against the c_str()
  walk this branch replaces

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add changelog entries for three unadvertised fixes

The array-splicing fix goes first: it is a memory-corruption fix, and it is
the strongest illustration of why the ownership model is worth having,
since it is exactly the failure the model makes unrepresentable.

Also the uninitialized read when a filter emitted "properties": null, and
the evaluator.hpp include guard that defined EVALUATOR HPP and so never
guarded anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Consolidate the jsonpull comments

Comments were 25% of the added lines, and the ownership model was spelled
out in five places. Collect it into one block at the top of jsonpull.h and
point at it from the rest, cutting the ratio to 14% and the total by about
200 lines.

Removed the duplicate explanations of the deleter dispatch, of what detach
does to the back-pointers, and of "json_read returns intermediate
containers, do not free them". Trimmed the comments that argued for a
choice rather than described the code -- the reserve(2) / reserve(4)
rationales, the string-buffer copy, the pmtiles check ordering -- to a line
each, and shortened the test preambles, keeping the parts that say why a
test is shaped the way it is.

No code changes; the test suite is unchanged in both configurations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-13 21:27:07 -07:00
Erica FischerandClaude Opus 5 63fcac725a Keep variable-depth tile pyramids consistent when a zoom level has to drop features (#407)
* Only skip polygon cleaning if we are still at very high resolution

* Remove collinear points and clean polygons even at high resolution

* If we truncated but still have the data and need to drop, revive

* Deduplicate by ID even when the duplicate is clipped away

* Test that deduplication works across tile boundaries

* Write out children of a tile revived after its parent truncated

A tile writes the geometry for its children on pass 0 of its zoom level,
and the later passes, which are only retries with new thresholds, must
not write it again. But a tile whose parent truncated its pyramid is
skipped on pass 0, and is only revived on a later pass, once the zoom
has had to start dropping features. Gating on pass 0 meant its children
were never written at all, so a revived tile was always a dead end: it
appeared in the output at ordinary detail with nothing below it, even
though its truncated ancestor still held the full-detail geometry.

Write the children on whichever pass first tiles the tile instead. The
dropping thresholds only ever increase within a zoom, so for a revived
tile that is exactly the pass on which the zoom started dropping.

Also collect the three thresholds into dropping_features(), since
write_tile() and run_thread() have to agree about when truncation is
disabled and when a skipped tile comes back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Treat dropping by attribute like the other ways of dropping features

--drop-by-attribute-as-needed was added after variable-depth pyramids,
and minattribute never made it into the test for whether a zoom level is
discarding features. A zoom that was dropping by attribute could still
truncate pyramids, so some of its tiles became full-detail leaves while
the rest of the zoom had features dropped out of them, and tiles skipped
because an ancestor had truncated stayed missing.

Unlike the other thresholds, minattribute starts at the infinity on
whichever side is being kept rather than at zero, so dropping_features()
now takes the direction too.

On tests/tl_2022_11_tract at -Z10 -M15000, zoom 11 was dropping by
attribute and truncating two pyramids at the same time; now it truncates
none of them and the tile that had been skipped under zoom 10's
truncation is written out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Delete the merged tile that the deduplication test leaves behind

overzoom-test removes merged-dedup.pbf.json.check but not the
merged-dedup.pbf it was decoded from, so the file was left in the working
tree after every test run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Document what variable-depth pyramids now do to geometry and to dropping

Truncated tiles are no longer left uncleaned: they keep every vertex that
isn't collinear with its neighbors, but their polygons are cleaned so
that overlapping areas are merged instead of stacked. Say so, and say
that dropping features at a zoom level now suppresses truncation for the
whole zoom rather than for individual tiles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Start minattribute out at the infinity that excludes nothing

dropping_features() reads write_tile_args::minattribute, and the in-class
default of 0 decodes as a threshold that has already been chosen. Every
path assigns it from zoom_minattribute before anything reads it, so this
changes no behavior, but a future one that didn't would silently suppress
pyramid truncation rather than fail visibly.

-HUGE_VAL is the value that excludes nothing for the ascending order that
drop_by_attribute_descending also defaults to, so the two members agree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Regenerate the drop-by-attribute fixture through the test harness

The Makefile can't be asked for a target whose name contains an =, since
make reads that as a variable assignment, so this fixture was generated by
hand into a scratch directory. The output path ends up in the tileset's
name, description, and generator_options, and tippecanoe-decode is only
passed -x generator, so all three were compared against the harness's
.check.mbtiles path and could never match. make test failed on it.

Regenerated with the same output path the rule uses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Restore the -z14 -M25000 variable-depth fixture

This configuration was dropped rather than regenerated when the -z17
-M10000 fixture was added. It still runs, so it was losing a passing
regression test for no stated reason. Regenerated against current
behavior.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Encode the = in the drop-by-attribute fixture name as %3d

A test output name containing an = can't be asked for on the make command
line, because make reads that argument as a variable assignment, so the
fixture couldn't be regenerated through its own rule. Add %3d to the
punctuation escapes that testargs decodes and use it here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

* Regenerate the man page for the README change

The variable-depth pyramid option's description changed, and man/tippecanoe.1
is generated from README.md, so the committed page no longer matched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KewFM6XCt5W9QBZkTZDjCp

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 12:26:33 -07:00
Erica FischerandClaude Opus 5 734bba7c78 Fix three latent defects exposed by compiler warnings, and clear the rest (#406)
* Fix variable-length-array and uninitialized-union compiler warnings

Clang warns about every variable-length array in C++ (-Wvla-cxx-extension,
on by default), since VLAs are a compiler extension rather than standard
C++. Replace all 57 of them with std::vector, or with std::string for the
mkstemp() template buffers built from tmpdir. Add -Wvla to WARNING_FLAGS so
new ones don't creep back in.

Separately, mvt_value's numeric_value union is 16 bytes wide (the size of
string_value), but both constructors only wrote the 8 bytes of the member
they were setting, leaving the rest indeterminate. The implicit copy
constructor copies the union as a whole, so copying any non-string value
read uninitialized bytes, which GCC reports as

  mvt.hpp:83:8: warning: 'v.mvt_value::numeric_value. ... .len' may be
  used uninitialized [-Wmaybe-uninitialized]

Give string_value, the widest member, a default member initializer so the
union's full width is initialized however it is later used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

* Fix remaining float-conversion and format-truncation warnings

Clang's -Wimplicit-const-int-float-conversion flagged two comparisons
against LLONG_MAX, which is not representable as a double and rounds up
to 2^63.

In serial.cpp this was a real latent overflow, not just noise: the guard
`extent <= LLONG_MAX` was really `extent <= 2^63`, so an extent of exactly
2^63 passed it and then hit `(long long) extent`, which is undefined for
that value and yields LLONG_MIN in practice -- the opposite of the clamp
the else branch intends. Make the bound exclusive so the conversion is
always in range. Requires a polygon area at the very top of the double
range to reach, but the clamp now behaves as written.

In mbtiles.cpp the value is only a stand-in for infinity on its way into
JSON, so cast explicitly; the emitted number is unchanged.

Separately, g++ at -O0 warned that `char abbrev[20]` can be truncated by
"%lld", which is correct: the most negative long long needs 21 bytes with
the NUL. That branch is only reached when point_count < 1000, so it cannot
happen today, but size the buffer to fit rather than rely on that, and
replace the garbled comment about how the size was derived.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

* Clamp the low end of extent before converting to long long too

The upper bound was fixed in the previous commit; the same overflow exists
on the negative side. get_area() returns a signed shoelace area, so inner
rings contribute negatively, and a polygon whose holes outweigh its rings
drives extent below zero. Far enough below and `(long long) extent` is
undefined again.

The bounds are asymmetric, so this is not simply the mirror of the upper
one: LLONG_MIN is exactly -2^63 and converts exactly, so unlike LLONG_MAX
it can be an inclusive bound.

Verified with -fsanitize=float-cast-overflow that the previous form traps
on 2^63 and on doubles just below -2^63, and that this one is clean across
both boundaries, the infinities, and NaN (which falls to LLONG_MAX, as it
did before).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

* Add CHANGELOG entries for 2.81.0 and bump the version

CHANGELOG.md was last updated for 2.80.0 (#361), and version.hpp has not
moved since. Twelve PRs have landed in the meantime with no entry: #365,
#368, #375, #382, #384, #385, #391, #395, #397, #399, #400, and #401.

Document all of them, plus this PR, under a single 2.81.0 heading. They are
not given separate version numbers because none of them was ever released
under one -- version.hpp read v2.80.0 throughout -- so assigning a version
per PR would invent release history. 2.81.0 is the version that will
actually carry them.

Where an unreleased PR was corrected by a later one (#384 by #385, #397 by
#399), the pair is described as the single behavior that ships, since the
intermediate behavior was never in a release.

Minor rather than patch bump: the batch adds command-line options.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

* Review feedback: enforce the union-width assumption, describe both clamp ends

The comment on mvt_value's union claimed string_value is the widest member.
That is true on LP64 (16 bytes against 8) but not on ILP32, where size_t is
4 and it ties with double and long long. The default member initializer still
covers the full union either way, so the fix held, but the justification did
not travel. Replace the claim with a static_assert that checks it on whatever
target is being built, so a platform where it stops holding is a compile
error rather than silently indeterminate bytes. Verified the assert is not
vacuous by widening the union in a scratch copy and watching it fail.

The changelog described only the upper end of the extent clamp. Describe both:
the old guard admitted everything below LLONG_MIN too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

* Add 2.81.0 changelog entries for the four PRs merged from main

#404, #408, #409, and #410 landed while this branch was open. None of them
bumped version.hpp, so they belong under the same 2.81.0 heading as the rest
of the unreleased work rather than getting versions of their own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 17:06:07 -07:00
Erica FischerandClaude Opus 5 1820630392 Fix the radix sort, and check that it agrees with the in-memory sort (#404)
* Don't write an extra byte when the radix sort writes a bucket directly

radix1() writes out a sorted bucket in two places. merge() writes all but
the last byte of each serialized feature and then appends the byte for the
feature minzoom, since the minzoom is the last byte of the feature. The path
taken when a bucket holds only one feature, or when the recursion has
consumed every bit of the index, instead writes the feature's whole
serialized length and then appends another minzoom byte, which is one byte
more than the feature's length prefix says it is. Everything read from the
geometry afterward is then misaligned by a byte.

--prefer-radix-sort lowers the memory limit to 8K so that this code gets
exercised, and any bucket that has to be written directly is enough to
desynchronize the stream, so it fails on several of the existing test
inputs:

    $ ./tippecanoe -q -f -o out.mbtiles -z4 -aR tests/ne_110m_ocean/in.json
    wrong length decoding feature: used 10, len is 33

Write one byte less here too, as merge() does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011wLk2itWETPBAS9a9yE8zu

* Keep the radix sort from recursing forever when it runs out of files

radix1() subdivides a bucket by the next splitbits bits of the index, and
stops recursing once prefix + splitbits reaches the width of the index.
The number of buckets comes from the number of files still available, which
shrinks at every level, so deep enough recursion reaches availfiles / 4 == 1
and therefore splitbits == 0. At that point the recursion consumes no bits
of the index and availfiles stops shrinking, so prefix never advances and
the recursion has no way to terminate.

A splitbits of 0 also makes the shift that chooses a feature's bucket a
shift by the full width of the index, which is undefined. In practice it
leaves the shift count masked to zero, so the bucket number is the whole
index rather than 0, and writing to that bucket runs off the end of the
arrays of open files.

Require at least two buckets so that each subdivision always consumes at
least one bit of the index and the shift is always in range, and don't
recurse at all when the next level would not have enough files to split
with: sort that bucket in memory instead, even though it is larger than
the memory limit asked for, since that is the only way left to get it
sorted.

--prefer-radix-sort, which lowers the memory limit to 8K so that this code
gets exercised, segfaults on tests/feature-filter/in.json without this:

    $ ./tippecanoe -q -f -o out.mbtiles -z0 -aR tests/feature-filter/in.json
    Segmentation fault

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011wLk2itWETPBAS9a9yE8zu

* Check that the radix sort and the in-memory sort agree

The result of a sort shouldn't depend on how the sort was performed, so
rather than checking the sorted output against a committed copy of it,
check that --prefer-radix-sort, which lowers the memory limit to 8K to
force the radix subdivision to recurse, produces the same tiles as sorting
in memory. Nothing new has to be kept up to date, and the comparison holds
regardless of how deeply the subdivision recurses on a given machine, which
depends on how many files it will let us open at once.

What sends the sort down the paths that are otherwise almost never taken is
the shape of the input rather than the size of it, so two small inputs are
generated for the purpose: several well-separated features that are each
too big to sort in memory, which are each written out as a bucket of their
own, and many features at one location, which have to be subdivided until
there are no index bits left. Between them and tests/feature-filter, all
three of radix1()'s branches are covered, including sorting in memory
because there are no files left to subdivide with.

Both of these inputs fail without the two preceding commits, and every
input here failed before them.

The whole target runs in about ten seconds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011wLk2itWETPBAS9a9yE8zu

* Record what the shift by the full index width actually did

Say in the comment that masking the shift count to zero makes the bucket
number come out as the whole shifted index, so the writes go somewhere
past the end of the arrays of buckets, rather than only that the shift is
undefined.

Also correct the note on the test: --prefer-radix-sort sets the memory
limit to 8K, but radix() halves it again, so the subdivision is working
against 4K.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011wLk2itWETPBAS9a9yE8zu

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 16:47:45 -07:00
Erica FischerandClaude Opus 5 905fe84459 Generate the man page with go-md2man instead of md2man-roff (#408)
* Generate the man page with go-md2man instead of md2man-roff

md2man-roff is distributed only as a Ruby gem -- it is in neither Homebrew
nor apt -- so in practice nobody has it installed and man/tippecanoe.1
drifts away from README.md. It was stale again as of #400: the man page
still had the dead All Streets link that commit fixed.

Switch to go-md2man, the maintained Go port of the same converter (it is
what Docker, podman and runc use). It is packaged as a single static
binary for Homebrew, apt, Fedora and Alpine, and it renders inline code as
bold the same way md2man-roff did, so the man page still reads the way it
used to.

It also emits valid roff, which md2man-roff did not. `mandoc -T lint` goes
from 621 errors and warnings to 1 (an empty .TH date, left empty on
purpose so that generation stays reproducible). 590 of those were
`invalid escape sequence: \fC`, from md2man-roff wrapping every inline
code span in `\fB\fC` -- `\fC` is not a font escape.

md2man-roff was losing content, too:

  README:      1/(2^32) of the size of Earth
  md2man-roff: 1/(2 of the size of Earth
  go-md2man:   1/(2^32) of the size of Earth

  README:      '{"attr": "operation", "attr2": "operation2"}'
  md2man-roff: '{"attr": "operation", "attr2", "operation2"}'
  go-md2man:   '{"attr": "operation", "attr2": "operation2"}'

Prepend a title block and a NAME section during generation rather than
adding them to README.md, where they would render as noise on GitHub.
The man page had neither, so its header rendered as "tippecanoe()" with no
section, and `man -k tippecanoe` and `whatis tippecanoe` found nothing.
It now renders as TIPPECANOE(1) and is indexed.

Finally, add a CI job that regenerates the man page and fails if the
committed copy differs, so a README edit that needs `make docs` gets
caught rather than sitting stale until someone notices. This is what
makes the missing-tool problem stop mattering: contributors no longer
need go-md2man installed to keep the man page current, since CI will
tell them when it needs regenerating. The go-md2man version is pinned
there because different versions produce different roff for the same
input.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015B6PcMbNY779iaczo6S6Pu

* Strip the redundant blank lines go-md2man puts between paragraphs

go-md2man separates paragraphs with a blank line as well as a .PP macro.
A blank line is itself a break in roff, so the two together double-space
the page: every paragraph was followed by two blank lines rather than one.
md2man-roff did not do this, so it showed up as a regression -- the source
went from 13 blank lines to 200.

Filter them out after generation. Blank lines inside .EX and .TS blocks
are kept, since there they are part of the example or the table rather
than spacing around it; that is all 13 of the ones md2man-roff emitted.

The rendered page loses 186 blank lines and the source loses 187, with
byte-identical non-blank output under both groff -t -man and mandoc.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015B6PcMbNY779iaczo6S6Pu

* Decouple the man page from version.hpp, and name the first section

Review feedback on #408.

Making man/tippecanoe.1 depend on version.hpp turned the docs job into a
hard CI failure on any release commit that bumps the version without
regenerating -- #406, which is open and moves version.hpp to v2.81.0
without touching the man page, would have tripped it as soon as either
merged. The only thing the dependency bought was the version in the page
footer, so every release would have had to regenerate the whole file to
rewrite that one line, gated by CI. Drop it: the source field is now just
"tippecanoe", and the page depends on README.md alone.

Separately, README.md's own title heading became the second .SH, directly
below the NAME section this branch adds, so the page opened with a stray
"tippecanoe" section. Rename it to DESCRIPTION, which is where that text
belongs and what a reader expects after NAME.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015B6PcMbNY779iaczo6S6Pu

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 16:47:13 -07:00
Erica FischerandClaude Opus 5 ec727172b1 docs: correct README statements that don't match the code (#410)
* docs: correct README statements that don't match the code

Cross-checked README.md against the option tables in main.cpp,
tile-join.cpp, decode.cpp, jsontool.cpp and overzoom.cpp, plus
options.hpp for the -pX/-aX letter assignments.

Incorrect:

* -aD and -aS were swapped. options.hpp assigns 'D' to
  A_COALESCE_FRACTION_AS_NEEDED and 'S' to
  A_COALESCE_DENSEST_AS_NEEDED, the opposite of what was documented.
* --limit-base-zoom-to-maximum-zoom was given as -Pb. It is a
  prevent flag (P_BASEZOOM_ABOVE_MAXZOOM = 'b'), so it is -pb; -P
  is --read-parallel and takes no letters.
* --retain-points-multiplier referred to --tile-size-limit, which
  is not an option. The limit it extends is --maximum-tile-bytes.
* The dot-dropping description said tippecanoe "drops 1/2.5 of the
  dots for each zoom level above the point base zoom". It keeps
  1/2.5 of them, at zooms below the base zoom (prep_drop_states
  sets interval only where i < basezoom).
* The default tileset name was given as "file.json". make_metadata
  sets both name and description from the output file or directory
  name.
* tile-join -r/--read-from was described as a "list of input
  mbtiles"; it names a file to read that list from, one per line.
* tippecanoe-decode's -I and -F were given as --integer and
  --fraction. Those work only as getopt abbreviations; the real
  names are --integer-coordinates and --fractional-coordinates.
* Development notes said C++11 and suggested g++-5. The Makefile
  builds with -std=c++17.
* Malformed references: "-quiet" and "no-simplification-of-shared-nodes".

Undocumented options now covered:

* tippecanoe: -aa/--keep-point-cluster-position,
  --preserve-multiplier-density-threshold, -H/--help, the count
  operation for --accumulate-attribute, and the
  point_count_abbreviated cluster attribute.
* tile-join: -O as the short form of --overzoom, -q/--quiet,
  --exclude-all-tile-attributes, --exclude-all-tile-geometries.
* tippecanoe-decode: -y/--include, -x/--exclude-metadata-row.
* tippecanoe-overzoom: -x/--exclude, --exclude-prefix, -J,
  -S/--line-simplification, --tiny-polygon-size,
  --deduplicate-by-id, --no-tile-compression, -t/--source-tile,
  -o/--output, and the long names for -b, -d, -y, -j, -m and -E.

Also noted that CSV latitude/longitude columns are matched
case-insensitively as substrings, added file.csv to the usage
synopsis, and explained the -a/-p letter-bundle syntax that the
short forms throughout the document rely on.

Every newly documented flag was run against a built binary. The
man page is regenerated from README.md per the Makefile rule; that
also picks up the All Streets link fix from #400, which had not
been regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MvSCpD1yQZhRT5iMU9yufQ

* Hide --unidecode-data from the generated usage messages

The option has done nothing since 533e000 removed the only caller of
unidecode_smash(), so listing it advertises behavior the tools don't
have. Move it after the empty-name entry that ends the usage listing,
the same place --no-polygon-splitting and the debug options sit, so it
is still accepted but no longer offered.

This is the situation #409 already fixed for tile-join's
--use-attribute-for-id, but it applies to all three tools that take
--unidecode-data, not just tile-join: main.cpp listed it under
"Filtering features by attributes" and overzoom.cpp under "Modifying
feature attributes", both ahead of the terminator.

tile-join's "Modifying feature attributes" heading covered only this
option, so it goes too rather than being left empty. overzoom.cpp had
no hidden group at all, so one is added. In main.cpp and overzoom.cpp
the heading keeps its other options and stays.

strip_usage_headings() copies every entry with a non-zero val, so the
moved option still reaches getopt_long(); confirmed by running each
tool with --unidecode-data and checking it is absent from --help.
make test passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MvSCpD1yQZhRT5iMU9yufQ

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 16:36:13 -07:00
Erica FischerandClaude Opus 5 e6e1ec3263 Generate the usage message of each tool from its long_options (#409)
* Generate the usage message of each tool from its long_options

The usage messages of tile-join, tippecanoe-overzoom,
tippecanoe-json-tool, tippecanoe-decode, and tippecanoe-enumerate were
hand-written lists of options that had drifted years out of date, since
nothing tied them to the options that are really accepted. Move the
option-list printing that tippecanoe already does into a shared
print_usage(), and use it in all the tools, so that the message is
derived from the same long_options table that getopt_long() gets and
can't fall behind it again.

The tables now carry section headings, as tippecanoe's does, and the
options that were only reachable by their short names (tile-join's -O,
-b, -R, and -r among them) are listed for the first time.

Also state the non-option arguments the way each tool really treats
them: tile-join takes source tilesets unless --read-from names a file to
read them from, tippecanoe-decode takes a tileset either alone or with a
zoom/x/y, tippecanoe-json-tool reads standard input when no files are
named, and tippecanoe-overzoom's two forms are the ones its argument
parsing recognizes. tippecanoe-overzoom now reports the missing -o
instead of passing NULL to fopen(), and tippecanoe-enumerate goes
through getopt_long() so that it will pick up any options added later.

The shared getopt_string() replaces the identical loop that four of the
tools each had for building the short option string, and strip_usage_headings()
the one for dropping the headings before getopt_long() sees them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

* Print the usage message when tippecanoe is run with no arguments

Running `tippecanoe` with nothing at all reported the missing output
file, which is true but is not what someone who typed the bare command
needs to know. Check for the empty command line before parsing and print
the general usage message instead, and leave the specific complaint for
the case where an input file was named but an output file wasn't.

To make the message reachable from there, the options table and the
usage printing move out of main() into a usage() function, as in the
other tools.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

* Address review: alternation, the dead tile-join option, and --version

Four fixes from review of the generated usage messages:

* `--output` and `--output-to-directory` are one-of, not one required and
  one optional, in both tippecanoe and tile-join. A `usage_required_option`
  can now name an alternation that it belongs to, and the options in one
  are listed together as `(--output=... | --output-to-directory=...)`,
  which is what the runtime check enforces.

* tile-join's `--use-attribute-for-id` has had no implementation since
  533e000 removed it; only the table entry was left behind, so the option
  parsed and then exited with "Unrecognized option". Generating the usage
  message from the table turned that into a documented option that doesn't
  work, so remove the leftover entry too.

* `--version` was grouped under "Progress indicator", in the options table
  and in the README both. Give it a heading of its own now that the
  headings are something users see.

* print_usage() left `width` holding the length of the last synopsis line,
  and only got away with it because every table so far begins with a
  heading, which resets it. Start the option list on a line of its own
  instead of depending on that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 16:19:03 -07:00
Erica Fischer 9a7ac5733f Remove unused Dockerfiles and lambda to avoid security warnings (#365) 2025-09-03 13:12:59 -07:00
Erica Fischer 533e000faa Remove undocumented command-line options (#361)
* Remove --accumulate-numeric-attributes

* Remove join-sqlite, etc.

* Remove --accumulate-numeric-attributes from overzoom

* Remove --assign-to-bins and --bin-by-id-list

* Remove --clip-polygon and --clip-bounding-box

* Remove FSL expressions

* Update version and changelog
2025-07-31 17:03:01 -07:00
Erica Fischer 68ab8dcc22 Deduplicate in tippecanoe-overzoom even when the duplicate is clipped away (#353)
* Deduplicate by ID even when the duplicate is clipped away

* Test that deduplication works across tile boundaries

* Update version and changelog
2025-07-24 13:21:10 -07:00
Erica Fischer 2d548bed06 Infinite loop fixes, minimizing changes to behavior (#345)
* Divide-and-conquer polygon cleaning

* Catch the case where the gap can't be increased further

* Catch the case where we try to keep impossibly many features

* Make label points earlier in the tiling process

* Another case where it could try to drop even after already limiting.

* And do not coalesce on impossibly small geometries

* Add missing return

* Update version and changelog
2025-05-09 09:08:28 -07:00
Erica Fischer 94929b048c Add --deduplicate-by-id option to tippecanoe-overzoom (#331)
* Add option to deduplicate by feature ID in overzoom

* Add test, fix default

* Update version and changelog
2025-04-03 12:48:29 -07:00
Erica Fischer bfb62ee2db Add missing case for accumulating the mean of attributes that are inconsistently present (#329)
* Add missing case for accumulating the mean of attributes that are inconsistently present

* Add more specific test
2025-03-20 14:53:04 -07:00
Erica Fischer 423fea1544 Reduce memory consumption during tiling and tile encoding (#319)
* Improve the memory spike during tile construction

* Remove `need_tilestats`. Add a bunch of debug logging

* Remove debug logging

* Update version and changelog
2025-01-31 12:57:41 -08:00
Erica Fischer 583fc3744a Reduce attribute accumulation memory consumption (#318)
* Add a flag to use an H3 index for the feature index

* Give mvt_value and serial_val a double-with-count concept

* Switch mean over to internal accumulation state

* Get rid of the attribute accumulation map

* Change vectors of features to vectors of pointers to features

* Fix --coalesce

* Revert "Add a flag to use an H3 index for the feature index"

This reverts commit b9b48f42c9.

* Update version and changelog
2025-01-30 21:58:35 -08:00
Erica Fischer 10f7f0a3c2 Support joins from sqlite tables in tile-join (#308)
* Sketching out sqlite options for tile-join

* Enable sqlite3 serialized multithreading

* Fix some unnecessary round trips from std::string to char * and back

* Gathering join keys for sql query

* Make a query

* Actually open the gpkg. Fix the query quoting.

* Actually do the query and get results back

* Join the attributes onto the feature

* Set matched if the sql join matches

* Add a flag to get the feature ID from the query

* Observe attribute exclusion when joining from sql queries

* Add a flag to exclude all attributes from the tile side of the join

* Make tile-join bounding boxes reflect feature bounds, not tile bounds

* More tests

* An empty tileset has empty bounds at null island

* Gather the results from each thread *after* the thread finishes

* Missed a test

* Fix accidental inclusion of the top left of the tile in the bbox

* Case smashing and prefix trimming in the select

* Add test of sql join

* Make the join column option a join expression option

* Fix antimeridian adjustment. Z0 can't wait until the end of the tile

* Checkpoint on accepting multiple joined rows per tiled feature

* Adding the joined attribute should be per-feature, not per-attribute

* Forgot to update the test. Order of joined attributes has changed.

* Get the attributes back in the right order

* Add test of sql join with limit

* Allow multiple tile features to have the same join key

* Update tests for a country name with two distinct geometries

* Update version and changelog

* Forgot to mention the bounding box improvements
2025-01-13 12:25:00 -08:00
Erica Fischer 7165ae6999 Fix clipping bug when the clip region doesn't intersect the tile (#312) 2024-12-11 10:24:12 -08:00
Erica Fischer bdfb06cd6a Make tippecanoe-overzoom accept filters from a file (#307)
* Make tippecanoe-overzoom accept filters from a file

* Accept clip polygons from a file too

* Add test of clipping by polygon from file

* Add a test of reading an overzoom filter from a file

* Update version and changelog
2024-12-05 12:05:50 -08:00
Erica Fischer dcc616d3d4 Adding optional clipping to tippecanoe-overzoom (#298)
* Plumb a clip bounding box around through overzoom

* Actually do some clipping

* Add a test

* Fix post-binning clipping

* Factoring out geometry parsing from feature parsing

* Accept a clip polygon argument to tippecanoe-overzoom

* Progress in the direction of polygon clipping

* Fix the wagyu flags. We need intersection, not union

* Remove debug spew

* Clip points to polygon bounds too

* Copy the geometric binning code to serve as intersection-finding code

* Add clipper2 for linestring clipping

* Compiles, but does not actually seem to clip. Hmm.

* Oh, it helps if I actually call the function

* Add clipping tests

* Add missing fixture, and don't crash if it is missing

* Remember to do polygon clipping after binning too

* Fix scaling before post-binning clipping. Add test.

* Remove unused parts of clipper

* Rename for consistency

* Revert accidentally added line

* Clip the clip regions to the tile bounds to reduce their complexity

* Add a test of clipping the clip region down to the tile boundary

* Update version and changelog
2024-11-27 13:43:40 -08:00
Erica Fischer 11e3196c9a Raise tippecanoe-decode tile size limit to 250 MB (#299)
* Raise tippecanoe-decode tile size limit to 250 MB

* Update version and changelog
2024-11-21 09:17:22 -08:00
Erica Fischer 905b58cdf4 Reducing attribute tagging within tippecanoe-overzoom (#296)
* Strip out unwanted attributes earlier in the process

* Skip aggregations whose attributes have been excluded

* Forgot one

* Add some more tests

* Update version and changelog
2024-11-15 13:14:30 -08:00
Erica Fischer 6d46578ad8 Avoid crash if the first bin gets clipped away (#294)
* Avoid crash if the first bin gets clipped away

* Add test

* Update version and changelog
2024-11-13 10:48:33 -08:00
Erica Fischer 794ae2b1eb Tippecanoe-decode and tippecanoe-overzoom changes to improve binning performance (#292)
* Add a tippecanoe-decode option to restrict which attributes to decode

* Plumb buffer and feature limit around

* Check the feature limit

* Clarifying cases where output detail can be unspecified

* Clip bins to the tile buffer instead of just passing them through

* Add missing include

* Missed some tests

* Add --no-tile-compression option to tippecanoe-overzoom

* Update version and changelog
2024-11-08 15:46:04 -08:00
Erica Fischer 23667bb8eb Reduce memory consumption from attribute accumulation (#290)
* Progress on plumbing a string pool for full_keys through

* More plumbing for key_pool

* Don't keep features with identical locations as multiplier features

* Revert "Don't keep features with identical locations as multiplier features"

This reverts commit 413f0c8024.

* Adjust calculated maxzoom to account for duplicate feature locations

* Update changelog and version

* Add a test affected by the maxzoom change with duplicate locations

* Round the drop rate a little for cross-platform test consistency
2024-11-05 14:39:41 -08:00
Erica Fischer b3b89e1e07 Performance optimizations for binning (#283)
* Add an all-mvt_value attribute accumulation path

* Only bin by ID, not geometrically

* A little cleanup; changelog and version; test

* Remove accidental double-conversion

* Replace duplicated code with template

* Update changelog
2024-11-01 15:27:13 -07:00
Erica Fischer 28efc40e6e Choose the megatile features from those that will be in the next N zooms (#280)
* Choose the megatile features from those that will be in the next N zooms

* Take fractional zooms into account in multiplier feature choices

* Fix more tests

* Add a flag to retain multiplier features by minimum distance

* Limit feature expansion from multiplier density to 2x

* The multiplier cap was a bad idea

* Revert "The multiplier cap was a bad idea"

This reverts commit 6f8273a4c8.

* Revert "Limit feature expansion from multiplier density to 2x"

This reverts commit a26e41309d.

* Revert "Add a flag to retain multiplier features by minimum distance"

This reverts commit 01f14a4255.

* Remove the multiplier sequence, which should no longer matter

* Revert "Revert "Add a flag to retain multiplier features by minimum distance""

This reverts commit 776da4a1b8.

* Revert "Revert "Limit feature expansion from multiplier density to 2x""

This reverts commit 44a683d808.

* Revert "Revert "The multiplier cap was a bad idea""

This reverts commit 80f7cb1c0e.

* Track two kinds of previous index for next_feature

* Fix multiplier density threshold, I think

* Oh, I didn't git add the code changes

* Update version and changelog

* Try to install sqlite3 to fix the automated build

* Deleted too much

* Only let --preserve-point-density-threshold shift density around

* Remove the density debt concept, since it doesn't help

* Make the drop states a vector instead of an array

* Revert "Make the drop states a vector instead of an array"

This reverts commit 66c7abb6fa.

* Revert "Remove the density debt concept, since it doesn't help"

This reverts commit 707bb0c562.

* Revert "Only let --preserve-point-density-threshold shift density around"

This reverts commit ecf01f2231.
2024-10-17 15:59:44 -07:00
Erica Fischer 78661f1cb2 Binning by ID (#276)
* Binning by ID

* Add test of binning by ID
2024-10-03 10:56:58 -07:00
Erica Fischer 8409bf9ce4 Top-level null filter now evaluates to true (#277) 2024-10-02 15:19:30 -07:00
Erica Fischer 23c6559b03 Remove buggy optimization to avoid reclipping in overzoom (#275)
* Remove buggy optimization to avoid reclipping in overzoom

* Add clarifying comment
2024-10-01 10:22:41 -07:00
Erica Fischer 66e5d66300 Another round of binning fixes (#274)
* Bin more aggressively if a point doesn't meet the pnpoly test

* Don't clip the points if we are binning

* Fix output of features added to the bin after its closure

* Update tests

* Update version and changelog
2024-09-30 08:57:13 -07:00
Erica Fischer 4822d02f24 Fix count accumulation in overzoom (#272)
* Fix count accumulation in overzoom

* Add test

* Increment version and changelog
2024-09-24 23:55:43 -07:00
Erica Fischer c5f2f0da34 More work on plumbing attribute accumulation through (#263)
* Plumb bounding boxes through potential intersections

* Quick bbox reject for bins that can't possibly intersect

* Inching toward attribute accumulation in megatile handling

* Some sort of test for how all these things interact with each other.

Automatic numeric attribute accumulation does *not* apply to attributes
that have an explicit attribute accumulator set, because the order of
operations is too messy and weird

* More sketching

* More sketching

* Actually do some accumulation

* Put all that behind an --accumulate-numeric flag

* Use the same attribute accumulation logic in binning as in megatiles

* Fix backwards conditional

* Add means, but somehow I have some counts of 0

* Handle aggregated attributes with no base attribute in the feature

* Checkpoint before I break everything

* Found a flaw, now to debug

* Fix a typo that broke accumulation

* Add binning tests

* Make sure IDs make it through on the bins

* Fix count/mean accumulation

* Make the numeric accumulation prefix configurable

* Make sure the accumulate test still works with a different prefix

* Forgot to update this test

* More testing to make sure cluster sizes make it all the way through

* Fix neglected --accumulate-attribute when binning

* Mark unexercised attribute accumulation cases as "can't happen"

* Factor out numeric preservation

* Attrs with the accumulation prefix are just preserved, not accumulated

* Test behavior of prefixed attributes

* Plumbing for exclude and exclude-prefix

* Implement and test attribute prefix stripping in overzoom

* Update version and changelog

* For debugging, make an attribute list of source feature IDs

* Revert "For debugging, make an attribute list of source feature IDs"

This reverts commit 65fc99c9d1.
2024-09-20 15:30:29 -07:00
Erica Fischer 534242887d Pass bin IDs through to the output (#266)
* Pass bin IDs through to the output

* Update version and changelog
2024-09-16 17:08:56 -07:00
Erica Fischer 84f6e887a7 Another try at fixing longitude wraparound for bins (#261)
* Another try at fixing longitude wraparound for bins

* Gonna get it right this time

* Forgot to update the comment

* The filter case was not supposed to reinterpret geometry

* Copy antimeridian-crossing geometries to the other side too

* Getting closer to getting antimeridian-crossing polygons right

* Update changelog and version
2024-09-10 14:25:57 -07:00
Erica Fischer 51fcf142df Fix bad interaction between dynamic dropping and limiting by truncation (#260)
* Fix bad interaction between dynamic dropping and limiting by truncation

* Don't enforce the tile size limit here if they said not to
2024-09-05 13:39:53 -07:00
Erica Fischer 5b18eea673 Work in progress on binning features in overzoom (#258)
* Factoring out tilestats management from GeoJSON file reading

* Move code around so overzoom can link against parse_layers

* Read the file of bins

* Plumb the bins through to overzoom()

* Some zip code bins to test with

* (Currently non-functional) test of binning

* Starting to spell out the bin matching loop

* Can't flatten points, so don't flatten bins either

* More fleshing out bin traversal

* Bounding box of tile-relative mvt geometry

* Smallest enclosing tile from bbox

* Most of the bin scan

* Add point in polygon check. It crashes.

* Find the matching bins

* GDAL-style bounding boxes have eaten my brain

* Make some features to bin into

* Increment a count as features are found to be within the bins

* Fix longitude wraparound in overzoom bins

* Fix the tests

* Push off attribute copying until after bin assignment

* Carry sum of numeric attributes into the bins

* Also add mean, min, and max

* Add --calculate-feature-index since I keep needing it for testing

* Add an option to accumulate sum/mean/max/min/count of all numeric attrs

* Don't bake in tippecanoe:mean, since we redo it from sum and count

* Forgot to update this test fixture after removing tiled mean

* Update version and changelog
2024-09-05 12:06:51 -07:00
Erica Fischer 40bb4ff732 Be more careful to retry when the feature count is exceeded (#257)
* Clip before dealing with multiplier or filters in overzoom

* Be more careful to retry when the feature count is exceeded

* Adjust the estimated total feature count for the multiplier too

* Fix the feature count estimates, I think

* Pass build info into the version string

* Report the actual max zoom of any tiles as the metadata maxzoom

* Revert unneeded renaming to make the diff more readable

* Clean up the adjustments to tile sizes and feature counts

* Update version and changelog

* Dropping a feature into a multiplier cluster still effectively drops it

* Update changelog

* Rethink the changelog description

* Don't try to truncate zooms if we are still tiling at z18
2024-08-20 10:54:16 -07:00
Erica Fischer d891ca7289 Fix latitude bboxes for features that extend beyond the mercator plane (#254)
* Fix latitude bboxes for features that extend beyond the mercator plane

* Update some more test expectations

* Update version and changelog
2024-08-08 15:30:52 -07:00
Erica Fischer bc3ef87c3f Add --generate-variable-depth-tile-pyramid option (#251)
* Track output position at the file level instead of within each tile

* Track file position where the child tile data begins

* Add option and document its intended behavior

* Changing the detail loop to account for stopping early

* I forgot I already added an option for this

* Stop early if we can make a complete tile

* Add a test of zoom truncation with limited feature count

* Forgot to commit the actual code change

* Make room for a vertex count in the header of each serialized tile

* Estimate tile complexity; don't try truncating when unlikely to work

* Be more conservative, because ever retrying a tile is a big speed hit

* If stopping early, don't simplify or clean; leave that to overzoom

* Add tiny polygon reduction / dust to overzoom

* Don't try to stop early in the children if we dropped anything by rate

* Fflush here too before pwriting

* Don't stop early if we ended up dropping any features.

Rework the can-the-next-zoom-stop-early logic to avoid going
one zoom further than needed.

* Fix warning

* Fix warnings

* Oops, checking for the wrong expected return value

* Cleanup from adding line simplification in overzoom

* Current (wrong) behavior when combining coalescing and truncating

* Keep a list of parent tiles to skip rather than truncating

* Now the coalesced tiles in z12 get children in z13

* Don't double-count feature dropping when the zoom level is retried

* Correct README description

* Remove todo about special case below basezoom, which is accounted for

* Be a little more aggressive in drop-densest determination

* Scale tile feature limit for megatiles in the same way as byte limit

* Fully deprecate -detect-shared-borders into an alias

* Track the distances found in the douglas-peucker recursion

* Serialize and deserialize the distance with the vertices

* Revert "Serialize and deserialize the distance with the vertices"

This reverts commit 753f1b7909.

* Revert "Track the distances found in the douglas-peucker recursion"

This reverts commit e5361f8c22.

* Revert "Fully deprecate -detect-shared-borders into an alias"

This reverts commit 0698aeb766.

* Better tracking of whether we failed to make a full-detail tile

* Put a bloom filter in front of the binary search for shared nodes

* Forgot to take out this printf

* Improve dispatch of tiling tasks

* Still dispatch the biggest tasks first

* Track zoom truncation in the strategies list in the tileset metadata

* Prescan for small deltas before doing proper simplification

* Revert "Prescan for small deltas before doing proper simplification"

This reverts commit d1d8238b83.

* Update version and changelog

* Rename to --generate-variable-depth-tile-pyramid
2024-08-06 16:05:52 -07:00
Erica Fischer 50deb9ce63 Add multi-tile input to tippecanoe-overzoom (#249)
* Reviving multi-source-tile overzoom: the clip.cpp side

* Reviving multi-source-tile overzoom: the overzoom.cpp side

* Update readme

* Update version and changelog
2024-07-23 13:21:10 -07:00
Erica Fischer 47e774adc4 Improve the appearance of coalesce-densest-as-needed tiles (#247)
* Start to distinguish fixed cluster density setting from as-needed density

* Make consistent {drop,coalesce}-densest decisions between zooms

* Actually track the previous index instead of just intending to

* Clean up collinearities in coalesced features

* To determine densest, look at actual physical distance, not just index

* Don't actually need the previous index in serial_feature now

* Center of mass of one feature to most distant point of the next

* Add apologetic comment

* Wait, how did the tests pass before?

* Revert "Wait, how did the tests pass before?"

This reverts commit f73c8ee543.

* Add --maximum-string-attribute-length option

* Update version and changelog

* A little more testing to make sure
2024-07-16 13:23:23 -07:00
Erica Fischer e11583df1c Fix hash collision in the string pool (#239)
* Current broken behavior

* More blatant test

* Fix hash collision in string pool

* Update version and changelog
2024-06-07 15:01:14 -07:00
Erica Fischer bb4f220678 Reduce tiling memory (#227)
* Trying to reduce memory in tiling

* Let the simplification workers go out of scope earlier

* Bail out quickly once the maximum feature count is reached

* Update changelog and version
2024-04-03 07:12:26 -07:00
Erica Fischer bd48ba8ea1 Fix accidental loss (at all zooms) of features with an explicit minzoom (#221)
* Fix accidental loss (at all zooms) of features with an explicit minzoom

* Try to stabilize centers and bounding boxes between architectures
2024-03-22 10:20:10 -07:00
Erica Fischer e796400377 Fix null behavior in in/ni expressions (#214)
* Fix null behavior in in/ni expressions

* Add test for null attributes in "ni" filters

* Remove unidecode from "in" expression evaluation again
2024-03-08 13:01:12 -08:00
Erica Fischer 655eccfcc3 Restore the use of unidecode for in/ni operators (#212)
* Restore the use of unidecode for in/ni operators

* Update changelog and version
2024-03-08 09:22:55 -08:00
Erica Fischer 021bb96aeb Allow non-string types to be used in "in" expressions (#211)
* Allow non-string types to be used in "in" expressions

* No unidecode smashing in "in" expressions, though!

* Update version and changelog
2024-03-04 10:16:48 -08:00
Erica Fischer 6a8f1b83d8 Fix some undefined behavior (#209)
* Fix some undefined behavior

* Avoid overflow in line simplification calculations

* Oops, missed a test

* Didn't mean to add that to the Makefile

* Revert "Revert "[ci] test in debug mode (#202)""

This reverts commit c95c328e47.

* Fix reference to out-of-scope pointer

* Fix invalid shift and out-of-bounds vector element reference

* Update changelog and version
2024-03-01 10:11:41 -08:00
Erica Fischer 312e1560a3 Stabilize feature order in overzoom (#210)
* Stabilize feature order in overzoom

* I want my sorts to be stable, please

* Revert "[ci] test in debug mode (#202)"

This reverts commit 853ada87b5.

* No need to reinitialize here
2024-02-29 10:39:10 -08:00
Erica Fischer a987197eed Don't swap attributes when reducing tiny polygon dust (#207)
* Don't swap attributes when reducing tiny polygon dust

Because the dust placeholder may be misleadingly far from the feature
that contributed the most area to it

* Remove unused arguments; update changelog
2024-02-26 15:41:26 -08:00
Erica Fischer 2b6630c42c Allow features that would be dropped dynamically to become multiplier features (#199)
* Postpone tagging features as being the first of a multiplier cluster

* Upgrade some dynamically dropped features to multiplier features

* Still don't let it put more features in a cluster than is allowed

* Update tests

* Make the current multiplier cluster size per-layer

* Make sure the first non-empty-geometry in the layer is marked as primary

* Improve comments

* Update changelog and version

* Remove commented out debugging printf

* Factor out duplicated code
2024-02-15 11:34:09 -08:00
Erica Fischer 96f126dd59 FSL-style expressions can use unidecode data to smash case and diacritics (#197)
* Read unidecode data, do some plumbing of it

* More unidecode plumbing

* Do the unidecode smashing, but it doesn't seem to be working

* Ah, that's better!

* Add missing header

* And reorder the includes too

* Shortcut when there is no unidecode data to work with

* Update version and changelog

* Avoid repeated unidecode smashing of the same constant string
2024-02-13 14:18:30 -08:00
Erica Fischer 4e52cbd957 Drop or retain whole multiplier clusters when dropping as needed (#198)
* Prep to track conditions other than just "dropped" or "kept"

* Count up instead of down

* Drop or retain whole multiplier clusters based on their first feature

* Calculate a global feature dropping sequence

* Switch over to using the drop sequence for drop-fraction

* Remove unused arguments for the old drop-fraction implementation

* Fix copy-and-paste bugs, update tests

* Properly incorporate feature_minzoom into the drop sequence, I hope

* Rename drop_by to drop_sequence

* See if sorting within clusters fixes filter stability between zooms

* Remove very chatty debug print

* Update changelog and version

* Use named constants instead of numbers for feature dropping/keeping

* Add comment to explain purpose and method of bit reversal
2024-02-12 10:58:49 -08:00
Erica Fischer e2a7a409c7 Improve tiling speed (#195)
* Add a way to run tippecanoe single-threaded for profiling

* Do less work when the tilestats sample values list is already full

* Save a copy when retrieving the attribute key

* Fewer atomic operations

* Move string hashing from mbtiles to text

* Only do approximate attribute deduplication when writing tiles

* Feature dropping tests are sensitive to exact tile size

* All tile creators now create a string pool for the tile

* Features clipped away to nothing should not participate in that tile

* Revert "Only do approximate attribute deduplication when writing tiles"

This reverts commit c42b34b498.

* Also revert the related test changes

* Revert "Revert "Only do approximate attribute deduplication when writing tiles""

This reverts commit 18509876c3.

* Be more specific about the string hash function

* Use fnv1a instead of std::hash for everything

* Reduce the chance of hash collisions

* Stick a hash search on the front of the tree search in addpool

* Eliminate repeated hashing of the same string

* Switch instead of ifs in json parsing

* A few more cases to populate the hash in addpool

* Store the hash in the tree instead of recalculating

* Add explanatory comment for mysterious argument

* Fewer copies in attribute stringification

* Clean up ancient weirdness in JSON attribute stringification

* More serial_val cleanup

* Pass a serial_feature to rewrite instead of many broken-down arguments

* Get rid of the multiple geometries within `partial`

* Revert "Pass a serial_feature to rewrite instead of many broken-down arguments"

This reverts commit 6f4ab9b725.

* Goodbye, struct coalesce

* Revert "Features clipped away to nothing should not participate in that tile"

This reverts commit 124462fbdc.

* Migrating fields from partial to serial_feature

* Name reconciliation between serial_feature and partial

* Replace struct partial with an augmented serial_feature

* Fix some overzealous search-and-replace renaming

* Don't say struct so often

* Remove more of the former partial construction

* Commenting and cleaning up

* Trying again to avoid all these arguments to rewrite

* I swear I did this same thing before and it didn't work.

* More rewrite cleanup

* Exile --detect-shared-borders to its own file

* Add missing headers

* More commenting and cleanup

* More comments

* Sprinkle consts around

* Emplacing and std::moving

* More cleanup

* That shouldn't have worked after a std::move

* Don't need to allocate memory to compare keys

* Reduce use of the global string pool in tiling

* Another avoidable mvt_value construction

* Further reduction to explicit string pool passing

* These reverses are no longer optimizations

* These layernames can all be references

* Don't drag an unused layername string around with every feature

* Heed a compiler warning about potential buffer overflow

* Fix my confusion about which feature's string pool is relevant

* Avoid some unnecessary allocations in attribute accumulation

* Maybe faster serialization?

* Eliminate a comparison

* Do the same here

* Save a couple of allocations when parsing numbers in JSON

* Immediately assign features to layers instead of subdividing later

* Maintain tilestats for tippecanoe:retain_points_multiplier_sequence

* Crunch out more duplicate attribute values when writing out the tile

* Do tilestats for tippecanoe:retain_points_multiplier_first too

* Shell filters need to be real threads, even if nothing else does

* Simplify tippecanoe_minzoom/maxzoom representation

* Update version and changelog
2024-02-07 17:50:35 -08:00
Erica Fischer 6a2bce8164 Scale the tile size limit up with the multiplier at low zooms (#192)
* Scale the tile size limit up with the multiplier at low zooms

* Add a test to demonstrate that high zoom tiles can't be extra large

* Update changelog and version

* Fail more cleanly when a tile can't be made small enough

* Guard against a cluster where the start marker has been dropped

* Look harder for a working feature interval instead of giving up

* That change to the drop-smallest logic changed a test output

* Update changelog

* Add explanatory comment
2024-01-31 10:23:05 -08:00
Erica Fischer 7d4d264d50 Speeding up tippecanoe-overzoom (#191)
* Speed up mvt_value comparison

* Converting repetitive ifs to cases

* More conversions from ifs to cases

* Optimize the always-true filter case

* Don't convert types of attributes without accumulators

* Unordered map seems to be faster than map

* Add missing header

* Fix some warnings

* Fix the warnings better

* Avoid an int->string->int conversion

* Lazily initialize layer key and values maps when actually needed

* More switches from maps to unordered_maps

* Sure, I'll take the microoptimization

* More emplacement

* Save some copies

* Emplaces and moves

* Lazy linear scan of attributes instead of building a map

* Extra printfs, missing header

* Avoid clipping if the input and output tiles are the same

* But do clip if the tile extent is being reduced

* Make sure I'm not constructing std::strings here at runtime

* More worrying about runtime string construction

* A couple more std::moves

* Const references!

* More const references

* Another std::move

* Make the string_value of mvt_value std::optional

* Reserve storage when decoding

* Provision for different mvt_values to share a string pool

* Use the string pool when decoding

* Avoid another string construction

* Try limiting the depth of the search for duplicate attributes

* Revert "Try limiting the depth of the search for duplicate attributes"

This reverts commit 9ec94a15ff.

* Update changelog

* Fix typo noticed during code review
2024-01-29 11:33:56 -08:00
Erica Fischer cbf222754b Add --accumulate-attribute to tippecanoe-overzoom (#189)
* Starting to factor out attribute accumulation into its own file

* Continuing to factor out attribute accumulation

* Reduce duplicate code

* Plumbing the accumulate-attribute option around

* Call the attribute accumulator

* Test that accumulation works

* Add missing #includes

* Don't sort within individual multiplier clusters

Doing so throws off the spatial distribution of the low zooms

* Docs and changelog

* Add comments
2024-01-23 15:41:36 -08:00
Erica Fischer f957f30f90 Clean up internal naming related to tilestats (#190)
* Get rid of the type_and_string near-synonym for serial_val

* Rename file_keys to the more familiar tilestats

* "tas" (type_and_string) => "sv" (serial_val)

* "fk" (tile_keys) => "ts" (tilestats)

* Revert ""tas" (type_and_string) => "sv" (serial_val)"

This reverts commit 4854c57e22.

* More carefully this time: "tas" (type_and_string) => "sv" (serial_val)
2024-01-23 14:17:19 -08:00
Erica Fischer 679a0d62f2 Make feature ordering cooperate with --retain-points-multiplier (#188)
* Make feature ordering cooperate with --retain-points-multiplier

* Forgot to check in the actual code changes???

* Sort within each multiplier cluster as well as between clusters

* Correct description of behavior in changelog

* Drag original feature sequence along in megatiles for post-filter sort

* Plumb the preserve-input-order flag through overzoom

* Sort in overzoom if requested

* Use within-tile input sequence numbers, not global sequence numbers

* Documentation

* Reverse direction of search to prevent accidental skipping

* Add some comments about converting between attribute representations
2024-01-21 12:09:05 -08:00
Erica Fischer 5d92a17193 Add point retention multiplier (#179)
* Add an option to retain N times as many points as usual at each zoom

* Tests for point multipler with specified and guessed maxzooms

* Work in progress on inverse spatial ordering

* Fix inverse spatial feature order

* --reorder was depending on a feature index that wasn't being preserved

* Separate ordering by feature_minzoom from ordering inverse-spatially

* Add a test for the inverse spatial ordering

* Store the basezoom/droprate/multiplier decisions in tileset metadata

* Progress on adding filters to tippecanoe-overzoom

* Type promotion for comparison

* Look up the attribute value for ordering

* Add test of thinning and ordering features

* Plumb tippecanoe_decisions metadata through pmtiles

* Be careful not to put infinities in JSON

* Fix accidental dropping in what is meant to preserve sparse points

* Start distinguishing true, false, and null in expressions

* Most of the type conversions

* Add boolean conversions

* Literals and conjunctions

* Add filtering to tippecanoe-overzoom

* Add a test of filtering in overzoom

* Fix boolean conjunctions

* Handle the combination of cluster size and filtering

* Rework dot dropping to reconcile density threshold and multiplier

* Revert "Rework dot dropping to reconcile density threshold and multiplier"

This reverts commit f253a66382.

* Retain points by multiplier within each tile, not in global probability

* Test that intends to verify that the multiplier is reversible

* Get the test to detect the discrepancy

* Mark the start of multiplier clusters with a magic attribute

* Add string-contains

* Add in and ni operators

* Revert "Look up the attribute value for ordering"

This reverts commit 56bc73e49a.

* Revert "Type promotion for comparison"

This reverts commit 6f3256f5af.

* Make number formatting in tippecanoe_decisions consistent

* Revert "Add a test for the inverse spatial ordering"

This reverts commit c8047de9ab.

* Revert "Separate ordering by feature_minzoom from ordering inverse-spatially"

This reverts commit 35b19a223c.

* Revert "Fix inverse spatial feature order"

This reverts commit 5978ecdb44.

* Revert "Work in progress on inverse spatial ordering"

This reverts commit fdf230f632.

* Somehow missed the tests associated with that last revert

* Round-robin assign attributes to partials from across the multiplier

* Count the multiplier separately in each layer

* Fix distribution of accumulated attribute across multiplier features

* Update changelog, version, and docs

* Add "is null" and "isnt null" expressions

* Update interpretation of FSL expressions to pass the tests

* Test to assert that polygons are unaffected by the multiplier

* Clean up and comment

* Remove accidental unused case
2024-01-18 15:56:09 -08:00
Erica Fischer e8ca6c6de3 Slightly less compression makes as-needed dropping twice as fast (#182)
* Slightly less compression makes as-needed dropping twice as fast

* Update changelog
2024-01-05 11:07:14 -08:00
Erica Fischer 02e3bac2c0 Reduce memory consumption during tiling (#177)
* Drop duplicate geometries sooner in coalescing-as-needed

* Concatenate geometries before partial cleaning

* Spend less time with two feature representations in memory

* Further reduce in-memory duplication

* Remove another copy

* Don't keep duplicates in memory while coalescing

* Also clear out the mvt layer once it is no longer needed

* Update changelog
2023-12-22 14:20:21 -08:00
Erica Fischer 15a3e313ff Tolerate polygon rings with insufficiently many points in input (#175)
* Tolerate polygon rings with insufficiently many points in input

* Fix changelog typo
2023-12-11 11:14:26 -08:00
Erica Fischer d7bdbe363b Reduce maximum memory used for vertex sorting (#170)
* Add logging for out-of-memory debugging

* Use less memory for sub-sorting

* Reduce maximum memory used for vertex sorting
2023-11-30 13:48:50 -08:00
Erica Fischer d359461e61 Reduce tile-join overzooming memory usage (#162)
* 16 bits is enough for tile numbers

* Revert "16 bits is enough for tile numbers"

This reverts commit 71a0c4e1cf.

* Check what child tiles each overzoomed tile will have

and don't queue further overzooming of empty tiles

* Clean up naming; use std::move to avoid copying large arrays
2023-11-08 12:51:28 -08:00
Erica Fischer 2c56187530 Make tile-join distrust source tilesets' metadata maxzoom and minzoom (#161)
* Special-case a longitude wraparound of exactly 360°

* Update version and changelog

* Make tile-join distrust source tilesets' metadata maxzoom and minzoom
2023-11-07 14:41:05 -08:00
Erica Fischer e1c665b35d Only detect longitude wraparound within each ring. (#150)
* Only detect longitude wraparound within each ring.

Between rings it's OK to jump around from one side of the world
to the other.

* Add test

* Update changelog
2023-10-17 12:04:47 -07:00
Erica Fischer 0ca14086a9 Pull from the back of the list of tiles to overzoom, not the front (#149) 2023-10-04 14:59:19 -07:00
Erica Fischer cc5c1c79df Speed up overzooming in tile-join (#147)
* Clip away entire features by bbox. Avoid unnecessary recompression.

* Move parent tile decoding in tile-join out of overzoom proper

* An ever-growing cache of parent tiles

* Limit the size of the cache

* Remove the current reader *before* checking if we can run the queue

* Clean up

* Add missing #include

* Add comment

* When the tile-join cache fills up, evict the least recently used

* Fix microsecond math

* Factoring out tile-join's cache for testing

* Add unit tests for tile-join cache

* Update changelog and version
2023-10-04 12:14:22 -07:00
Erica Fischer 26bf08deb2 Externalizing polygon shard detection (#146)
* Start of externalizing polygon shard detection

* Completely untested external quicksort

* Add unit test for external quicksort

* Remember to clean up temporary files

* Sort and scan the vertices

* Bring over more vertex logic

* Make nodes from vertices

* Checkpoint on switching over to global shared nodes

* Do the thing

* Revert unintended change to coalesced linestring behavior

* Take shared nodes into account in early simplification

* Let it do more sorting in memory

* Fix overnoding of collinear linestrings

* Remove duplicate nodes, since only duplicate vertices now matter

* Remember to delete temporary files

* Fix out of bounds memory access below, apparently

* Fix the actual undefined behavior

* Still running out of memory in one case. Find out where.

* Forgot the conditional

* Try again to make it not run out of memory

* Update version and changelog
2023-10-03 10:53:12 -07:00
Erica Fischer f7dc7faf31 Reduce memory use of polygon shard detection (#139)
* Crunch down memory required per polygon joint

* *Actually* reduce the size of the structure

* Update version and changelog
2023-09-21 10:59:12 -07:00
Erica Fischer 390771f855 Polygon shards (#105)
* Put all of this back in geometry.cpp for conflict resolution

* Stabilize line simplification to behave the same regardless of winding

* Round instead of truncating when clipping lines

* Restore non-Wagyu polygon clipping from prior to 2fdec7d2

* Make it round, not truncate, which reverts the last commit's test diffs

* Clip in floating point, not integers, which makes no difference

* Track nodes added at tile edges during clipping

* Scale geometry up before wagyu to prevent changes from precision loss

* Actually do the shared edge detection

* Fix cases where nodes were not being added at the tile boundary

* One more place I should have rounded

* Narrow down where the discrepancy comes in

* Revert "Narrow down where the discrepancy comes in"

This reverts commit 221c4c5fc0ac9a6567e091c6a94b3d87dc8ade83.

* Another attempt to narrow it down

* The discrepancy seems to be introduced in reordering. No obvious bug

* Was still truncating instead of rounding in projection

* Also makes no difference...

* Just forget that line reversal exists for a minute

* Just reversal no coalescing

* Try clipping in integers instead of floating point

* Don't simplify after coalescing if they said no simplification

* Check whether behavior is consistent with intentional simplification

* Replace more floating point with integer

* Are these three features enough to demonstrate the problem?

* Add a few more nearby borders

* All the features that touch tile 6/16/23

* Stay in integers in line simplification

* More attempts to solve failures to simplify consistently

* Fix most of the overflow errors

* Fix known cases of integer overflow

* All the tests change again

* Pull clipping and scaling code back out into clip.cpp

* Resolve the test conflicts

* Stabilize choice of which three points to keep with different windings

* Almost right, I think!

* Fix collapse of islands to shards

* Self-intersections in the same feature don't count

* Revert "Self-intersections in the same feature don't count"

This reverts commit e04b19916e.

* Don't scale down geometry if we are going to look for shared nodes

* Fix the missing multiply that was keeping simplification from happening

* Fix one more opportunity for overflow

* Somehow I deleted this test?

* Lost this test too

* Clean up debugging printfs

* Restore code sequence from main to make it reviewable

* Remove unneeded rounding

* Update documentation

* This test is no longer useful

* Back to floating point Douglas-Peucker to fix undersimplification

* Try an older ubuntu

* Revert "Try an older ubuntu"

This reverts commit 13fefacfd7.

* Log OS info

* Remove tests that are no longer needed

* Oops, did need that one after all

* Fix the arm vs x86 discrepancy?

* Try another quantization

* Cleanup from review

* Add a test for the actual purpose of this PR
2023-09-14 16:51:38 -07:00
Erica Fischer e6d05bc317 Fix tile-join crash when trying to merge empty tilesets with --overzoom (#138)
* Fix tile-join crash when trying to merge empty tilesets with --overzoom

* Add an option not to reduce tiny polygons to dust at maxzoom

* Add test for prevention of tiny polygon reduction at maxzoom

* Change version number
2023-08-31 13:02:03 -07:00
Erica Fischer 2ec6180003 Fix "strategies" accounting for 0-length linestrings and degenerate polygons (#137)
* Dropping a 0-length feature doesn't count as dropping-as-needed

* Add an option not to reduce tiny polygons to dust at maxzoom

* Add test for prevention of tiny polygon reduction at maxzoom

* Fix accounting for tiny polygons not to include degenerate geometries

* Revert "Add test for prevention of tiny polygon reduction at maxzoom"

This reverts commit f931bbd73e.

* Revert "Add an option not to reduce tiny polygons to dust at maxzoom"

This reverts commit 03f0882bb6.

* Fix tests

* Another test that no longer has any really tiny polygons

* Oops, that broke LineString simplification

* This time for sure!

* Update changelog and version
2023-08-29 10:54:39 -07:00
Erica Fischer 6778aeac52 Add an option to extend zooms if still dropping, but with a limit (#131)
* Add an option to extend zooms if still dropping, but with a limit

* At least when to overzoom, even if not actually doing it yet

* Refactor to give tile-join access to overzoom()

* Didn't work, but *might* have worked

* OK, it did something now

* Ah, there's the bug!

* Hook up pmtiles and dirtiles as overzooming sources

* Add command line option to enable or disable overzooming

* Add (currently broken) test of overzooming in tile-join

* Slightly more abstraction for the tile-join readers

* Factor out duplicated code

* Move construction into a constructor

* More changing accessors to methods

* Reduce magic

* Start tracking a list of the tiles at maxzoom

* I think it worked?

* Add missing #include

* Fix sequence of overzoomed tiles (Y sorts backwards for TMS)

* Don't spend memory on overzooming when we aren't going to use it

* Diff rather than cmp, in the hope of figuring out this broken test

* Keep full coordinate precision if we might extend zooms

* Try a slightly different byte limit

* Make drop-densest more consistent across tile boundaries

* Also affects this test

* Does it behave any differently if it can extend forever?

* I think the discrepancy is a thread-safety problem here

* Revert "Does it behave any differently if it can extend forever?"

This reverts commit 0dff0a0acc.

* Lost this change to the test

* This time for sure!

* Revert "Also affects this test"

This reverts commit cd1f7c2e78.

* Revert "Make drop-densest more consistent across tile boundaries"

This reverts commit 563f7d2bc2.

* Revert "Try a slightly different byte limit"

This reverts commit 2e271213d6.

* Add some more explanatory comments

* Amend the join-test to detect my current bug

* Allow overzooming to complete the zoom if it ever starts

* Forgot to correct the test

* Update changelog and version

* Cleanups from code review

* Remove version number from fixture to fix test
2023-08-25 12:42:38 -07:00
Erica Fischer 430d8edd17 Tool for overzooming individual tiles (#121)
* Starting work on overzooming

* Factor clipping out of geometry.cpp to simplify linkage

* Pull out more geometry functions into now-badly-named clip.cpp

* Not surprisingly, there is a bug

* Found the bug

* Make indent

* Pass attributes through

* Make formatting more consistent

* Fix typos in comments

* Fix the typos better

* Add tippecanoe-overzoom to the install list

* Give overzoom a predictable exit status

* Forgot to translate to and from polygon ring closepaths

* Fix geometry collapse at z21

* Add attribute stripping; don't generate layers if they have no features

* Add docs for tippecanoe-overzoom (as it will be, not as it is)

* Change overzoom to accept input and output files as arguments

* Working on overzooming tests

* Oops

* Hook up and test the detail and buffer options, and the empty-tile case

* Also test attribute stripping

* Fix error message

* Update changelog and version
2023-08-09 11:48:16 -07:00
Erica Fischer 55b9aa501c Allow --set-attribute to override existing attribute values (#123)
* Allow --set-attribute to override existing attribute values

* Bump version number
2023-07-26 11:05:35 -07:00
Erica Fischer 4318e964e3 Add maximum pseudocluster size option for low-zoom points (#119)
* Trying to improve cluster positions

* Fix centroid calculation

* Just move points to cluster centroids; add point gap threshold

* Add maximum-point-gap option

* Fix formatting

* Rename options for clarity; add tests; update changelog

* Add missing break to switch

* Revert unneeded constructor cleanup that broke a test somehow

* Clean up comments

* Remove unused --move-points-to-cluster-centroids

* Mention tile-join change in changelog
2023-07-17 15:33:07 -07:00
Erica Fischer 1c3bd5352c Add a new option to set an initial value for an attribute in each feature (#115)
* Starting on --set-attributes

* Add a JSON form to accumulate-attribute

* Fix formatting

* Set attributes, JSON form

* Also handle setting attributes to unquoted values

* Factoring out character-mangling in the names of tests

* Rename tests for new quoting convention

* Test for JSON form of attribute setting

* Test for JSON form of attribute accumulation

* Forgot to increment the version
2023-07-14 10:47:55 -07:00
Erica Fischer 48e769c94c More codeowners (#118)
* Add more code owners

* Fix typo
2023-07-12 10:34:53 -07:00
Erica Fischer b5ec3c46f0 Reduce excessive progress logging during pmtiles conversion (#111)
* Don't log progress so often during pmtiles conversion

And turn off the `catch` tests, which have bit-rotted

* Fix accidental double-multiplication-by-100

* Reenable catch

* Update changelog
2023-07-10 15:59:52 -07:00
Erica Fischer c1051232ad Do more of line simplification with integer coordinates (#113)
* Turn off unit, which has bit-rotted

* Instrument line simplification

* Do more work in integer space

* Update tests with simplification done with more integers

* Clean up

* Update changelog

* Eradicate calls to deprecated sprintf()

* Revert "Turn off unit, which has bit-rotted"

This reverts commit 998779313f.

* Fix catch by switching to an unsigned type
2023-07-10 15:54:21 -07:00
Erica Fischer b09a880822 Add CODEOWNERS (#114)
* Add CODEOWNERS

* Narrower list?
2023-07-10 15:53:13 -07:00
Erica Fischer 20eb9c8dae Avoid crashing if there is a polygon ring containing only one point (#102) 2023-05-11 12:20:00 -07:00
Erica Fischer 523ca55243 Fix bugs in --no-simplification-of-shared-nodes (#99)
* Fix bugs in --no-simplification-of-shared-nodes

* Update changelog
2023-05-08 15:31:46 -07:00
Erica Fischer 605640e138 Add --include/-y to tile-join (#86)
* Add --include/-y to tile-join

* Update documentation and changelog
2023-05-04 11:55:36 -07:00
Erica Fischer afc7ca01a2 Calculate an antimeridian-adjusted bounding box in tileset metadata (#82)
* Calculate a new antimeridian-adjusted bounding box

* Add antimeridian bounding box to pmtiles, dirtiles, and tile-join

* Don't take out-of-bounds latitudes into account in the adjusted bbox

* Update changelog

* Forgot to adjust tests after the last change
2023-03-09 20:26:48 -08:00
Erica FischerandAurele Nitoref 329dd86bfd Add --cluster-maxzoom option (#75)
* Add `--cluster-maxzoom` option

* Add tests, update changelog

* Update manpage

* Update tests

---------

Co-authored-by: Aurele Nitoref <aurele.nitoref@icloud.com>
2023-02-24 16:05:45 -08:00
Erica FischerandAurele Nitoref 3d714e32b0 Add point_count_abbreviated property to clustered features (#74)
* Add `point_count_abbreviated` property to clustered features

* Update tests and changelog

---------

Co-authored-by: Aurele Nitoref <aurele.nitoref@icloud.com>
2023-02-24 15:59:00 -08:00
Erica Fischer f2127fec97 Reduce thrashing during feature ingestion and tiling (#56)
* Remove the concept of "separate metadata"

This was an extra level of attribute indirection (features point
to metadata records which point to key and value strings) which was
intended to reduce the size of temporary storage for features with
large numbers of attributes that were also spread across large numbers
of tiles at maxzoom.

For other kinds of features, the extra indirection slowed things down
instead, and, especially when maxzoom guessing was being used, many more
features were having their metadata externalized than could actually
benefit from it.

* Shave a few bytes off temporary files by using more unsigned integers

* Flush stderr after logging progress

* Revert "Shave a few bytes off temporary files by using more unsigned integers"

This reverts commit eef29084ec.

* Limit the size of the string pools and trees to fit in memory

* Add missing #include

* Move the string pool and search tree from mmap to allocated memory

* Sort in allocated rather than mapped memory too

* Also use pread instead of mapping to read in the data to sort

* When the pool gets too big, switch to just the file, not memory

* Switch string pool from memory to disk when memory is 10% full

* Add to-memory versions of the serialization functions

* Crashy work in progress toward compression

* Fix the pointer bug that was causing the crash

* Serialize features into memory rather than straight to disk

* Compress individual features in the temporary files

* Don't need to store the length of the geometry

* Remove per-feature compression; move minzoom back into the object

* Start adding a stream compressor object

* Track file position within fwrite_check()

* Add compressed stream writer functions

* Pull the writing of the serialized feature out to the callers

* Starting toward compression again from a different point

* Hook up more compression functions

* Remove unused code from the other day

* Make enough deflate calls to flush out all the buffered data

* Start on decompression

* Tile number is uncompressed, tile content is compressed

* Work on alternating compressed and uncompressed in decompression

* Closer, but still doesn't work

* Sort of works

* Works until we get to concatenated tiles

* More attempts that don't work

* One bug down

* It made a tileset!

* Handle nonzero initial zooms

* Fix seeking within compressed feature streams

* Tests pass!

* Remove debug spew

* Oops: remember to delete the temporary files so they don't hang around

* Test that fails with the current compression code

* Properly account for bytes read while closing the compressed stream

* Limit the number of warnings about bad label points

* A little more armor when closing decompression

* This time for sure

* A different, less fragile, test that failed previously with compression

* Move feature stream compression to its own file

* Remove now-unused code to deserialize from a file

* Forgot to add the new files

* Remove a little debugging logging

* Add a couple of comments on what it means to be within decompression

* Fix indentation

* Update changelog. Remove stray debugging comment.
2023-02-14 12:47:40 -08:00
Erica Fischer b155b4671b It only needs to look for other small features when coalescing, not dropping (#64)
* Back out the slow search for other small features when dropping features

* Actually, *do* look for small features still when coalescing

* Update changelog
2023-01-27 12:04:36 -08:00
Erica Fischer e615668475 Avoid placing polygon labels in holes (#62)
* Avoid placing polygon labels in holes

* Traverse vertices in hilbert order to find potential label points

* Update test to reflect new label placements

* Sorting by Y coordinate is better and easier than Hilbert order

* Give a bonus for being near the center of mass

* Make label points a little more border-shy

* Tall places, not just wide places; border lines, not just border points

* Clean up

* Try diagonals through the features too

* Limit the search for label points to prevent slowing down too much

* At this point the center of mass bonus is doing more harm than good

* Forgot to update test
2023-01-27 11:31:57 -08:00
Erica Fischer c58a8e3b97 Round coordinates instead of truncating them (#60)
* Round coordinates instead of truncating them

* Update all the tests for coordinate rounding changes

* Curses, integer division still truncates

* Fix tests

* Also round instead of shifting when looking for no-op linetos

* Should I worry that the same change for moveto doesn't change any tests?

* Also round instead of shifting when scaling down to maxzoom resolution

* Replace another explicit shift, for origin point

* Round instead of shift when writing clipped geometries to the next zoom

* Fix low-zoom gridding and smaller-than-a-pixel checks

* Don't guess an excessively large maxzoom when there is only one feature

* Add a test for guessing the maxzoom of a single point

* Explicitly sort by index if no other order distinguishes features

* Another affected test
2023-01-27 10:57:22 -08:00
Erica Fischer 3cb6e4dd5e Set base zoom if requested as part of maxzoom guessing (#63)
* Set base zoom if requested as part of maxzoom guessing

* Add a test for guessing the maxzoom of a single point

* Don't guess an excessively large maxzoom when there is only one feature

* Update changelog
2023-01-26 15:27:06 -08:00
Erica Fischer 54b47a6757 Fix crash when using tile-join to copy an empty pmtiles tileset (#61)
* Fix crash when using tile-join to copy an empty pmtiles tileset

* Update changelog
2023-01-17 15:16:22 -08:00
Erica FischerandBrandon Liu 3b2599f587 Add support for pmtiles output format (#55)
* add pmtiles.hpp from github.com/protomaps/PMTiles [#10]

* tippecanoe main writes pmtiles output. [#10]

* detect output format using suffix
* after mbtiles is done writing, replace with pmtiles based on map/image tables.
* add method to write_json for writing json sub-object.

* tippecanoe-decode reads pmtiles input. [#10]

* tile-join reads and writes pmtiles. [#10]

* pmtiles test suite for decode and tile-join [#10]

* add base GitHub CI action for compiling and test suite.

* update pmtiles.hpp with z>15 fix

* Fix some ordering problems with pmtiles decode

* Pmtiles should also pass the raw tiles tests

* Eradicate spaces from tileset metadata JSON fields

* Eradicate spaces from more test fixtures

* Update more tests

* Pmtiles tests pass now too

* Remove unnecessary sort (and make indent)

* Update changelog

* The allow-existing test for pmtiles needs -o, not -e

* Declare --allow-existing to be unsupported for pmtiles.

It was always a bad idea even for mbtiles.

Co-authored-by: Brandon Liu <bdon@bdon.org>
2022-12-29 12:00:13 -08:00
Erica Fischer 079929717e Avoid spending gigabytes of memory on statistics for -as-needed dropping (#50)
* Downsample indices and areas during tiling if they get too big

* Cap the indices rather than downsampling them
2022-12-20 12:54:21 -08:00
Erica Fischer 2b7bbbdbca Report JSON object context on the same line as the error message (#49)
* Report JSON object context on the same line as the error message

* Update changelog
2022-12-14 12:34:21 -08:00
Erica Fischer 9c77029b63 Add a data structure for tileset metadata (#44)
* Progress toward making a tileset metadata structure

* Write metadata from structure to mbtiles

* Write dirtiles metadata.json from metadata structure

* Update changelog
2022-12-12 13:32:02 -08:00
Erica Fischer f1df09f147 Generate fewer duplicate label points at high zoom levels (#42) 2022-12-02 11:02:33 -08:00