mirror of
https://github.com/felt/tippecanoe.git
synced 2026-10-02 16:35:40 +02:00
9ff3f92e6ec318e51e8cd875e862a7a2db88b8a8
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4f2621186a |
Convert jsonpull to C++ with shared_ptr and std::vector/std::string (#388)
* Rename to jsonpull.cpp * Clear for merge * Convert jsonpull to C++ with shared_ptr and std::vector/std::string Replace the manual malloc/realloc/free memory management in jsonpull with std::shared_ptr ownership. Each json_object now owns its children through std::vector<json_object_ptr>; raw back-pointers to parent and parser remain valid by structural invariant and are cleared on json_disconnect so detached subtrees can outlive their parser. Strings become std::string, child arrays become std::vector, and the old union becomes a struct so non-trivial members can coexist while preserving the existing o->value.xxx access paths. The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now returns std::string, and all callers across tippecanoe, tile-join, tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the unit tests are updated to use json_object_ptr / json_pull_ptr. Co-authored-by: Cursor <cursoragent@cursor.com> * Subclass json_object so primitives shrink from 168 to 24 bytes The previous "every member in a struct" layout cost 168 bytes per json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that have no payload. Splitting json_object into a small base class plus json_number / json_string / json_array / json_hash subclasses brings each instance down to just the size of its actual contents: json_object (base, TRUE / FALSE / NULL) 24 bytes json_number 48 bytes json_string 48 bytes json_array (empty) 48 bytes json_hash (empty) 72 bytes Other size wins along the way: * Drop enable_shared_from_this<json_object> (its embedded weak_ptr was 16 bytes per node). json_pull now keeps an explicit container_stack and the parser no longer needs to resurrect a shared_ptr from a raw `parent` walk. * Remove the unused `refcon` slot from the string variant. * No virtual destructor: shared_ptr keeps the deleter from the original std::make_shared<json_xxx> call, so destroying a shared_ptr<json_object> still runs the right subclass dtor. The base class exposes type-tagged accessors (o->string(), o->number(), o->array(), o->keys(), o->values(), o->large_signed(), o->large_unsigned()) that assert the type matches and downcast to the appropriate subclass storage. All call sites were swept from the old `o->value.X.Y` field paths to these accessors. A raw-pointer overload of json_hash_get() replaces the few external uses of shared_from_this() that survived in geojson-loop.cpp. Co-authored-by: Cursor <cursoragent@cursor.com> * Store hash key/value pairs in one ordered vector Replace the parallel std::vector<json_object_ptr> keys / values on json_hash with a single std::vector<json_entry>, where json_entry is a small {key, value} aggregate. This still preserves insertion order (the property the parallel vectors were providing) but removes the "keep two vectors in lockstep" pattern, and call sites can now use range-for with structured bindings: for (auto &[k, v] : o->entries()) { ... } Side effects: * sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector header), matching json_array. * The keys() and values() accessors on json_object are replaced by a single entries() accessor returning std::vector<json_entry>&. * All call sites were swept from the old paired-index pattern (`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the original pattern relied on `nprop = 0` to short-circuit iteration on a null or non-hash `properties`, the rewrite now guards the loop explicitly with `if (o->type == JSON_HASH)` so that calling entries() doesn't trip the asserting downcast. Co-authored-by: Cursor <cursoragent@cursor.com> * Move parser-only `expect` state out of json_object `expect` was only meaningful while the parser was building a container, and only ever read or written from jsonpull.cpp itself; once parsing finished it was dead weight on every JSON_ARRAY and JSON_HASH (and present-but-unused on every primitive too). Move it into the parser's container stack, alongside the shared_ptr to the container it pertains to: struct json_pull::parse_frame { json_object_ptr container; json_type expect; }; std::vector<parse_frame> container_stack; The base class now only carries data-model state (parent, parser, type). No external caller depended on `expect`, so no sweep was needed outside jsonpull.cpp. This change does not, in itself, shrink any json_object: the 4-byte `expect` field used to live at offset 20 inside the base, where it was already being eaten by alignment padding for the 8-byte-aligned first member of every subclass (std::string, std::vector, double). The win is in the data model, not the byte count -- the 4-byte hole is still there, but it is now available for a future subclass whose first member is small enough to slot into it. Co-authored-by: Cursor <cursoragent@cursor.com> * Discriminate json_number's three numeric slots into one union json_number used to carry three parallel 8-byte fields (a double plus both a 64-bit unsigned and a 64-bit signed slot for the large-integer cases) even though at most one of the integer slots is ever the canonical value for any given number. Collapse them into a discriminated union: enum repr_t { REPR_DOUBLE, REPR_LARGE_UNSIGNED, REPR_LARGE_SIGNED }; repr_t repr; union { double d; unsigned long long u; long long s; } value; Callers keep the same read API: number() returns the appropriate double, large_unsigned() returns the ull (or 0 if not currently stored that way), large_signed() likewise. Writes go through new set_number / set_large_unsigned / set_large_signed methods that keep the discriminator and the union value in sync. This was prompted by an observation that moving json_type to the end of the object should shrink things via tail-padding reuse. Empirically the type-at-end rearrangement saves nothing on its own (every subclass payload is 8-byte aligned so it can't slot into the 4-byte tail), but the discriminated-number redesign hits the same idea from a different direction: adding the 4-byte `repr` to json_number makes the class non-standard-layout, which lets the Itanium ABI pack `repr` into the base's 4-byte tail padding at offset 20. The union value then starts at the natural offset 24, and json_number ends at offset 32 -- a 33% reduction. Per-node sizes: json_object (TRUE/FALSE/NULL) 24 bytes json_number 32 bytes (was 48) json_string 48 bytes json_array 48 bytes json_hash 48 bytes Numbers dominate real GeoJSON (every coordinate is one), so the net memory win on a typical parse is substantial. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix bugs flagged in code review of jsonpull C++ port - jsontool.cpp `out()`: route JSON_NUMBER (and anything else non-string) through `json_stringify` instead of `o->string()`, which now asserts on a non-string type and would crash `--extract` on numeric attributes. - geojson.{hpp,cpp} `json_end_map`: take `json_pull_ptr` by reference so the caller's shared_ptr is released, null-guard before touching `jp->source`, and clear `jp->source` after delete to avoid a dangling pointer. - jsonpull/jsonpull.cpp: low-surrogate range check was comparing the outer-loop byte `c` instead of the parsed code unit `ch`, breaking surrogate-pair decoding for some \\uXXXX escapes. Pre-existing bug preserved across the port. - tile-join.cpp `handle_vector_layers`: require the field value to have type JSON_STRING (and the key to be non-null) before calling `string()`; the previous truthy `type` check would assert on a non-string value. Co-authored-by: Cursor <cursoragent@cursor.com> * Add jsonpull regression test for surrogate-pair decoding Covers the `c` vs `ch` bug fixed in the previous commit: parsing "\uD83D\uE000" (a valid high surrogate followed by a non-surrogate BMP code point) used to mis-classify U+E000 as a low surrogate and combine the two units into U+1F400 (F0 9F 90 80). The fixed code flushes the stale high surrogate as standalone CESU-8 (ED A0 BD) and then encodes U+E000 normally as EE 80 80. Verified the test fails under the pre-fix logic. Co-authored-by: Cursor <cursoragent@cursor.com> * Cheap perf wins in jsonpull C++ port Profiling tl_2022_us_county.json (sample(1) on Apple Silicon) showed ~38% of parse time in allocator work and ~14% in std::string::push_back during string-token construction. These changes target the low-hanging fruit from that profile: - Pre-reserve 2 slots in json_array and 4 slots in json_hash so coordinate `[x, y]` pairs and typical GeoJSON property maps avoid the 0 -> 1 -> 2 -> 4 vector-growth chain (and the shared_ptr copies it incurs). - Reuse a parser-wide std::string buffer for JSON_STRING tokens instead of constructing a fresh local std::string per token. The buffer is cleared (capacity preserved) at the start of each token and copied into the final json_string, so once it has grown to the longest string seen it stops reallocating entirely. - std::move the freshly-created container shared_ptr into the parser container stack in the `[` and `{` handlers, and move it out of the frame on the matching `]` / `}`. Each move skips one atomic inc/dec round-trip per container open and close. On a tl_2022_us_county.json benchmark (4-iter user-time mean, Apple Silicon, /usr/bin/time): - main baseline: ~8.17s - jsonpull-cpp before these changes: ~10.90s (+33%) - jsonpull-cpp with these changes: ~9.33s (+14%) So this commit recovers roughly half of the post-port regression. The remaining gap is dominated by shared_ptr atomic refcount traffic on the parse tree and per-node heap allocations, which would require the larger unique_ptr/arena reworks to address. Co-authored-by: Cursor <cursoragent@cursor.com> * Make json_free actually free the subtree In the C++ port, json_free was just `o.reset()`, which dropped the caller's reference but left the subtree alive: the parent's vector slot kept it allocated, and for line-delimited streams the parser's jp->root co-owned it until the next top-level value started parsing. That defeated the geojson-loop pattern of calling json_free on each feature after serializing it, which is supposed to release the feature so it doesn't sit in memory while subsequent ones are parsed. Restore the historical "remove this from the tree" semantics by splicing the node out of its parent (sharing splice_from_parent with json_disconnect) and clearing jp->root when the node is the parser's current top-level value, then dropping the caller's reference. Two unit tests pin this down: a pruning test parses "[[1, 2], [3, 4], [5, 6]]" element-wise and confirms that calling json_free on [3, 4] leaves the outer array with just [1, 2] and [5, 6]; a top-level test uses a weak_ptr observer to confirm that json_free on the parser's root really destroys the tree. Co-authored-by: Cursor <cursoragent@cursor.com> * Migrate jsonpull to unique_ptr ownership Replaces the shared_ptr-based json_object_ptr with a unique_ptr that has a stateless custom deleter dispatching on json_object::type before calling the right subclass destructor. Eliminates per-node atomic reference-counting and the control-block allocation that shared_ptr required for every node in the tree. API now distinguishes owning and borrowing pointers explicitly: - json_read / json_read_separators / json_hash_get return raw json_object * (borrowed from the parser-owned tree). - json_read_tree / json_disconnect return json_object_ptr (caller takes ownership; back-pointers are cleared so the subtree can outlive the parser). - json_free / json_context / json_stringify take raw pointers. - The parser's container_stack holds raw pointers; jp->root keeps unique_ptr ownership of the most recent top-level value. Internally, take_from_owner moves the unique_ptr out of whichever parent vector / hash entry / parser root owned it, which both json_free and json_disconnect rely on. In the streaming parsers (parse_feature, parse_layers, the geojson-loop callback), we are careful to free `j` only after we have processed a complete Feature: json_read returns each token as the tree is being built up, and freeing an intermediate node would splice it out of the surrounding hash and corrupt the in-progress feature. Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping, median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr and 8.7s on the pre-refactor C baseline. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix preprocessor mistakes identified by Copilot * Make indent * Skip non-string metadata.json entries instead of reading them as strings dirmeta2tmp() warned about a metadata entry that was not a string/string pair and then read it as a string anyway. Under the new type-tagged accessors that trips the assert in json_object::string(); before them it reinterpreted the node's storage as a char pointer, which segfaulted for most values. Either way, tippecanoe-decode and tile-join could not read a directory tileset whose metadata.json had a numeric minzoom or a nested object, which is common in metadata.json files written by other tools. Add the missing continue, and cover it in raw-tiles-test. pmtilesmeta2tmp() handles the same case correctly but read the key with string() before its own JSON_STRING check, so the assert would have fired ahead of the check meant to catch a bad key. Hoist the check above the read. The parser rejects non-string hash keys, so this is unreachable in practice; the ordering is what makes the check meaningful. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Don't redefine _GNU_SOURCE in the C++ jsonpull port The `#define _GNU_SOURCE` carried over from jsonpull.c, where it was needed to get asprintf() declared. g++ already defines _GNU_SOURCE on the command line for C++ translation units, so redefining it warns: jsonpull/jsonpull.cpp:1: warning: "_GNU_SOURCE" redefined Guard the define rather than drop it, so platforms whose C++ driver does not predefine it still get asprintf() declared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Add a unit test for json_disconnect json_disconnect() is documented in jsonpull.h as the supported way to splice a subtree out of the parser's tree and take ownership of it, but nothing calls it: read_filter() and parse_filter() used to, and now get the same guarantee from json_read_tree() clearing back-pointers on the way out. Cover the behavior rather than leave the primitive dead and untested. The test pins that the subtree is removed from its parent, that the parser keeps the rest of the tree, and that the detached subtree stays readable after the json_pull is destroyed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Correct two stale comments in the jsonpull port jsonpull.h said a json_number is 40 bytes; it is 32 (json_object is 24, and the repr discriminator fits in the base class's tail padding, so the 8-byte union lands at offset 24). plugin.cpp's parse_feature() said `j` is freed only just before returning or as jp->root at end of stream, but there is a third json_free(j) at the bottom of the loop, for a complete Feature whose geometry came out empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Add a changelog entry and bump the version for the jsonpull rewrite The rewrite is meant to be behavior-preserving, but it carries four user-visible bug fixes that warrant release notes: tippecanoe-json-tool --extract on a numeric attribute, surrogate-pair decoding, tile-join reading a non-string tilejson field type, and non-string values in a directory tileset's metadata.json. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Encode U+FFFF as three bytes instead of an overlong four The \uXXXX decoder tested `ch < 0xFFFF` before taking the three-byte UTF-8 path, so U+FFFF itself fell through to the four-byte branch and came out as F0 8F BF BF -- an overlong, and therefore invalid, encoding of a code point that fits in three bytes. check_utf8() only checks that continuation bytes look like continuation bytes, not that a sequence is the shortest form, so nothing downstream noticed: a GeoJSON attribute containing U+FFFF put invalid UTF-8 into the output tile, where a strict consumer would reject it. Since `ch` is parsed from exactly four hex digits it cannot exceed 0xFFFF on its own, so after this change the four-byte branch is reached only for a code point assembled from a surrogate pair, which is the only way to name one above the BMP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Keep the parent links inside a detached jsonpull subtree json_read_tree() and json_disconnect() cleared both back-pointers on every node of the subtree they handed out. Clearing `parser` throughout is necessary -- the json_pull can be destroyed while the subtree lives on, so a surviving `parser` would dangle -- but clearing `parent` throughout cost more than it bought. `parent` is a non-owning raw pointer, so keeping it cannot form a reference cycle or keep anything alive; there is nothing to leak. And within a detached subtree it refers to nodes the caller now owns as a single unit, so it stays valid for exactly as long as the subtree itself. Clearing it only made the tree unwalkable upwards, and made json_free() and json_disconnect() silently no-ops on interior nodes of a detached tree, since both find a node's owner through o->parent. So clear `parser` everywhere and clear `parent` on the detached root alone, which is the one that pointed out of the subtree at a node the parser still owns. Split the old clear_back_pointers() into clear_parser_pointers() plus a detach_subtree() wrapper that adds the root's `parent`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Cover the U+FFFF encoding and detached-tree parent links Each of the new assertions fails against the previous behavior, so they pin the two fixes rather than merely passing alongside them: - the U+FFFF test, plus the U+FFFE boundary below it and a surrogate pair above it, so the three-byte and four-byte paths are both held in place - json_disconnect() leaving the parent links inside the subtree intact while clearing the root's - json_free() pruning an interior node of a tree whose parser is already gone, which only works because those links survive - json_free() of a hash value leaving the key paired with a JSON_NULL placeholder, which is the documented behavior and not a removal Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Address review notes in jsonpull itself - json_stringify walked c_str(), so it truncated at an embedded NUL even though values are std::string now and carry one through faithfully. Range over the string instead; the existing control-character branch already escapes a NUL like any other, so the output stays valid JSON. - Assert that the hash has an entry waiting before add_object() assigns to entries().back(). It always does -- JSON_VALUE is only set by a colon, which requires a pushed key -- but the derivation is not local. - Drop fabricate_object(), a pass-through to make_object() with the arguments reordered, kept only to preserve the old C name. - Inline the string_append / string_append_c wrappers over push_back and append, and note that json_print_one's JSON_HASH and JSON_ARRAY branches are unreachable, since json_print handles both itself. - json_hash_get's comment said nullptr meant "the matching value is null", which reads as JSON null. A JSON null comes back as a JSON_NULL node; nullptr means the value slot is not filled in yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Tidy jsonpull call sites flagged in review - geojson.cpp and attribute.cpp passed key_pool::pool() and set_attribute_accum() a c_str() from a std::string, forcing a needless reconstruction (and truncating at an embedded NUL). Both overloads take std::string, so pass it directly. The geojson.cpp one is the hottest loop in the program. - Replace the hand-maintained counters beside range-for loops in attribute.cpp, main.cpp and tile-join.cpp with indexed loops, since the index is only wanted for error messages. - parse_json_args took json_pull_ptr by value and then copied it, costing two refcount bumps per construction. Move it. - Assert that the parser is still attached where geojson.cpp reads geometry->parser->line. Only json_read results reach it today, but json_read_tree and json_disconnect now clear every parser pointer, so a detached tree would null-deref there instead of tripping an assert. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Pin the array-splicing fix, and close four test gaps The existing pruning test does not discriminate: json_read hands back each container as it completes, so the node it frees is always the most recently added element of its parent -- the one case the old element-count-vs-byte- count memmove got right, because it then moved zero bytes. Widening that test to more elements does not change this; the shape is what matters, not the size. Verified: the eight-element streaming variant still passes against the pre-fix code. Add a test that builds the array first and then prunes element 0 of eight, asserting the identity of every survivor rather than just the resulting count. That fails against the pre-fix code deterministically, with arr[0] == arr[1] and the last element dropped. Note the limitation on the streaming test so the next reader does not try to strengthen it in place. Also cover, all previously untested: - json_free of a hash key, and of both halves of a pair, where the entry survives with a JSON_NULL stand-in until both are gone - repeated json_read_tree over a line-delimited stream, which is what the filter loaders and -L / -E do, asserting each detached tree survives the next read and the parser's destruction - json_stringify of a partially-parsed tree, the json_context error path - json_stringify across an embedded NUL, which fails against the c_str() walk this branch replaces Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Add changelog entries for three unadvertised fixes The array-splicing fix goes first: it is a memory-corruption fix, and it is the strongest illustration of why the ownership model is worth having, since it is exactly the failure the model makes unrepresentable. Also the uninitialized read when a filter emitted "properties": null, and the evaluator.hpp include guard that defined EVALUATOR HPP and so never guarded anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r * Consolidate the jsonpull comments Comments were 25% of the added lines, and the ownership model was spelled out in five places. Collect it into one block at the top of jsonpull.h and point at it from the rest, cutting the ratio to 14% and the total by about 200 lines. Removed the duplicate explanations of the deleter dispatch, of what detach does to the back-pointers, and of "json_read returns intermediate containers, do not free them". Trimmed the comments that argued for a choice rather than described the code -- the reserve(2) / reserve(4) rationales, the string-buffer copy, the pmtiles check ordering -- to a line each, and shortened the test preambles, keeping the parts that say why a test is shaped the way it is. No code changes; the test suite is unchanged in both configurations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
ec727172b1 |
docs: correct README statements that don't match the code (#410)
* docs: correct README statements that don't match the code
Cross-checked README.md against the option tables in main.cpp,
tile-join.cpp, decode.cpp, jsontool.cpp and overzoom.cpp, plus
options.hpp for the -pX/-aX letter assignments.
Incorrect:
* -aD and -aS were swapped. options.hpp assigns 'D' to
A_COALESCE_FRACTION_AS_NEEDED and 'S' to
A_COALESCE_DENSEST_AS_NEEDED, the opposite of what was documented.
* --limit-base-zoom-to-maximum-zoom was given as -Pb. It is a
prevent flag (P_BASEZOOM_ABOVE_MAXZOOM = 'b'), so it is -pb; -P
is --read-parallel and takes no letters.
* --retain-points-multiplier referred to --tile-size-limit, which
is not an option. The limit it extends is --maximum-tile-bytes.
* The dot-dropping description said tippecanoe "drops 1/2.5 of the
dots for each zoom level above the point base zoom". It keeps
1/2.5 of them, at zooms below the base zoom (prep_drop_states
sets interval only where i < basezoom).
* The default tileset name was given as "file.json". make_metadata
sets both name and description from the output file or directory
name.
* tile-join -r/--read-from was described as a "list of input
mbtiles"; it names a file to read that list from, one per line.
* tippecanoe-decode's -I and -F were given as --integer and
--fraction. Those work only as getopt abbreviations; the real
names are --integer-coordinates and --fractional-coordinates.
* Development notes said C++11 and suggested g++-5. The Makefile
builds with -std=c++17.
* Malformed references: "-quiet" and "no-simplification-of-shared-nodes".
Undocumented options now covered:
* tippecanoe: -aa/--keep-point-cluster-position,
--preserve-multiplier-density-threshold, -H/--help, the count
operation for --accumulate-attribute, and the
point_count_abbreviated cluster attribute.
* tile-join: -O as the short form of --overzoom, -q/--quiet,
--exclude-all-tile-attributes, --exclude-all-tile-geometries.
* tippecanoe-decode: -y/--include, -x/--exclude-metadata-row.
* tippecanoe-overzoom: -x/--exclude, --exclude-prefix, -J,
-S/--line-simplification, --tiny-polygon-size,
--deduplicate-by-id, --no-tile-compression, -t/--source-tile,
-o/--output, and the long names for -b, -d, -y, -j, -m and -E.
Also noted that CSV latitude/longitude columns are matched
case-insensitively as substrings, added file.csv to the usage
synopsis, and explained the -a/-p letter-bundle syntax that the
short forms throughout the document rely on.
Every newly documented flag was run against a built binary. The
man page is regenerated from README.md per the Makefile rule; that
also picks up the All Streets link fix from #400, which had not
been regenerated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MvSCpD1yQZhRT5iMU9yufQ
* Hide --unidecode-data from the generated usage messages
The option has done nothing since
|
||
|
|
e6e1ec3263 |
Generate the usage message of each tool from its long_options (#409)
* Generate the usage message of each tool from its long_options
The usage messages of tile-join, tippecanoe-overzoom,
tippecanoe-json-tool, tippecanoe-decode, and tippecanoe-enumerate were
hand-written lists of options that had drifted years out of date, since
nothing tied them to the options that are really accepted. Move the
option-list printing that tippecanoe already does into a shared
print_usage(), and use it in all the tools, so that the message is
derived from the same long_options table that getopt_long() gets and
can't fall behind it again.
The tables now carry section headings, as tippecanoe's does, and the
options that were only reachable by their short names (tile-join's -O,
-b, -R, and -r among them) are listed for the first time.
Also state the non-option arguments the way each tool really treats
them: tile-join takes source tilesets unless --read-from names a file to
read them from, tippecanoe-decode takes a tileset either alone or with a
zoom/x/y, tippecanoe-json-tool reads standard input when no files are
named, and tippecanoe-overzoom's two forms are the ones its argument
parsing recognizes. tippecanoe-overzoom now reports the missing -o
instead of passing NULL to fopen(), and tippecanoe-enumerate goes
through getopt_long() so that it will pick up any options added later.
The shared getopt_string() replaces the identical loop that four of the
tools each had for building the short option string, and strip_usage_headings()
the one for dropping the headings before getopt_long() sees them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY
* Print the usage message when tippecanoe is run with no arguments
Running `tippecanoe` with nothing at all reported the missing output
file, which is true but is not what someone who typed the bare command
needs to know. Check for the empty command line before parsing and print
the general usage message instead, and leave the specific complaint for
the case where an input file was named but an output file wasn't.
To make the message reachable from there, the options table and the
usage printing move out of main() into a usage() function, as in the
other tools.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY
* Address review: alternation, the dead tile-join option, and --version
Four fixes from review of the generated usage messages:
* `--output` and `--output-to-directory` are one-of, not one required and
one optional, in both tippecanoe and tile-join. A `usage_required_option`
can now name an alternation that it belongs to, and the options in one
are listed together as `(--output=... | --output-to-directory=...)`,
which is what the runtime check enforces.
* tile-join's `--use-attribute-for-id` has had no implementation since
|
||
|
|
533e000faa |
Remove undocumented command-line options (#361)
* Remove --accumulate-numeric-attributes * Remove join-sqlite, etc. * Remove --accumulate-numeric-attributes from overzoom * Remove --assign-to-bins and --bin-by-id-list * Remove --clip-polygon and --clip-bounding-box * Remove FSL expressions * Update version and changelog |
||
|
|
94929b048c |
Add --deduplicate-by-id option to tippecanoe-overzoom (#331)
* Add option to deduplicate by feature ID in overzoom * Add test, fix default * Update version and changelog |
||
|
|
7165ae6999 | Fix clipping bug when the clip region doesn't intersect the tile (#312) | ||
|
|
bdfb06cd6a |
Make tippecanoe-overzoom accept filters from a file (#307)
* Make tippecanoe-overzoom accept filters from a file * Accept clip polygons from a file too * Add test of clipping by polygon from file * Add a test of reading an overzoom filter from a file * Update version and changelog |
||
|
|
dcc616d3d4 |
Adding optional clipping to tippecanoe-overzoom (#298)
* Plumb a clip bounding box around through overzoom * Actually do some clipping * Add a test * Fix post-binning clipping * Factoring out geometry parsing from feature parsing * Accept a clip polygon argument to tippecanoe-overzoom * Progress in the direction of polygon clipping * Fix the wagyu flags. We need intersection, not union * Remove debug spew * Clip points to polygon bounds too * Copy the geometric binning code to serve as intersection-finding code * Add clipper2 for linestring clipping * Compiles, but does not actually seem to clip. Hmm. * Oh, it helps if I actually call the function * Add clipping tests * Add missing fixture, and don't crash if it is missing * Remember to do polygon clipping after binning too * Fix scaling before post-binning clipping. Add test. * Remove unused parts of clipper * Rename for consistency * Revert accidentally added line * Clip the clip regions to the tile bounds to reduce their complexity * Add a test of clipping the clip region down to the tile boundary * Update version and changelog |
||
|
|
794ae2b1eb |
Tippecanoe-decode and tippecanoe-overzoom changes to improve binning performance (#292)
* Add a tippecanoe-decode option to restrict which attributes to decode * Plumb buffer and feature limit around * Check the feature limit * Clarifying cases where output detail can be unspecified * Clip bins to the tile buffer instead of just passing them through * Add missing include * Missed some tests * Add --no-tile-compression option to tippecanoe-overzoom * Update version and changelog |
||
|
|
78661f1cb2 |
Binning by ID (#276)
* Binning by ID * Add test of binning by ID |
||
|
|
c5f2f0da34 |
More work on plumbing attribute accumulation through (#263)
* Plumb bounding boxes through potential intersections
* Quick bbox reject for bins that can't possibly intersect
* Inching toward attribute accumulation in megatile handling
* Some sort of test for how all these things interact with each other.
Automatic numeric attribute accumulation does *not* apply to attributes
that have an explicit attribute accumulator set, because the order of
operations is too messy and weird
* More sketching
* More sketching
* Actually do some accumulation
* Put all that behind an --accumulate-numeric flag
* Use the same attribute accumulation logic in binning as in megatiles
* Fix backwards conditional
* Add means, but somehow I have some counts of 0
* Handle aggregated attributes with no base attribute in the feature
* Checkpoint before I break everything
* Found a flaw, now to debug
* Fix a typo that broke accumulation
* Add binning tests
* Make sure IDs make it through on the bins
* Fix count/mean accumulation
* Make the numeric accumulation prefix configurable
* Make sure the accumulate test still works with a different prefix
* Forgot to update this test
* More testing to make sure cluster sizes make it all the way through
* Fix neglected --accumulate-attribute when binning
* Mark unexercised attribute accumulation cases as "can't happen"
* Factor out numeric preservation
* Attrs with the accumulation prefix are just preserved, not accumulated
* Test behavior of prefixed attributes
* Plumbing for exclude and exclude-prefix
* Implement and test attribute prefix stripping in overzoom
* Update version and changelog
* For debugging, make an attribute list of source feature IDs
* Revert "For debugging, make an attribute list of source feature IDs"
This reverts commit
|
||
|
|
84f6e887a7 |
Another try at fixing longitude wraparound for bins (#261)
* Another try at fixing longitude wraparound for bins * Gonna get it right this time * Forgot to update the comment * The filter case was not supposed to reinterpret geometry * Copy antimeridian-crossing geometries to the other side too * Getting closer to getting antimeridian-crossing polygons right * Update changelog and version |
||
|
|
5b18eea673 |
Work in progress on binning features in overzoom (#258)
* Factoring out tilestats management from GeoJSON file reading * Move code around so overzoom can link against parse_layers * Read the file of bins * Plumb the bins through to overzoom() * Some zip code bins to test with * (Currently non-functional) test of binning * Starting to spell out the bin matching loop * Can't flatten points, so don't flatten bins either * More fleshing out bin traversal * Bounding box of tile-relative mvt geometry * Smallest enclosing tile from bbox * Most of the bin scan * Add point in polygon check. It crashes. * Find the matching bins * GDAL-style bounding boxes have eaten my brain * Make some features to bin into * Increment a count as features are found to be within the bins * Fix longitude wraparound in overzoom bins * Fix the tests * Push off attribute copying until after bin assignment * Carry sum of numeric attributes into the bins * Also add mean, min, and max * Add --calculate-feature-index since I keep needing it for testing * Add an option to accumulate sum/mean/max/min/count of all numeric attrs * Don't bake in tippecanoe:mean, since we redo it from sum and count * Forgot to update this test fixture after removing tiled mean * Update version and changelog |
||
|
|
bc3ef87c3f |
Add --generate-variable-depth-tile-pyramid option (#251)
* Track output position at the file level instead of within each tile * Track file position where the child tile data begins * Add option and document its intended behavior * Changing the detail loop to account for stopping early * I forgot I already added an option for this * Stop early if we can make a complete tile * Add a test of zoom truncation with limited feature count * Forgot to commit the actual code change * Make room for a vertex count in the header of each serialized tile * Estimate tile complexity; don't try truncating when unlikely to work * Be more conservative, because ever retrying a tile is a big speed hit * If stopping early, don't simplify or clean; leave that to overzoom * Add tiny polygon reduction / dust to overzoom * Don't try to stop early in the children if we dropped anything by rate * Fflush here too before pwriting * Don't stop early if we ended up dropping any features. Rework the can-the-next-zoom-stop-early logic to avoid going one zoom further than needed. * Fix warning * Fix warnings * Oops, checking for the wrong expected return value * Cleanup from adding line simplification in overzoom * Current (wrong) behavior when combining coalescing and truncating * Keep a list of parent tiles to skip rather than truncating * Now the coalesced tiles in z12 get children in z13 * Don't double-count feature dropping when the zoom level is retried * Correct README description * Remove todo about special case below basezoom, which is accounted for * Be a little more aggressive in drop-densest determination * Scale tile feature limit for megatiles in the same way as byte limit * Fully deprecate -detect-shared-borders into an alias * Track the distances found in the douglas-peucker recursion * Serialize and deserialize the distance with the vertices * Revert "Serialize and deserialize the distance with the vertices" This reverts commit |
||
|
|
50deb9ce63 |
Add multi-tile input to tippecanoe-overzoom (#249)
* Reviving multi-source-tile overzoom: the clip.cpp side * Reviving multi-source-tile overzoom: the overzoom.cpp side * Update readme * Update version and changelog |
||
|
|
96f126dd59 |
FSL-style expressions can use unidecode data to smash case and diacritics (#197)
* Read unidecode data, do some plumbing of it * More unidecode plumbing * Do the unidecode smashing, but it doesn't seem to be working * Ah, that's better! * Add missing header * And reorder the includes too * Shortcut when there is no unidecode data to work with * Update version and changelog * Avoid repeated unidecode smashing of the same constant string |
||
|
|
7d4d264d50 |
Speeding up tippecanoe-overzoom (#191)
* Speed up mvt_value comparison
* Converting repetitive ifs to cases
* More conversions from ifs to cases
* Optimize the always-true filter case
* Don't convert types of attributes without accumulators
* Unordered map seems to be faster than map
* Add missing header
* Fix some warnings
* Fix the warnings better
* Avoid an int->string->int conversion
* Lazily initialize layer key and values maps when actually needed
* More switches from maps to unordered_maps
* Sure, I'll take the microoptimization
* More emplacement
* Save some copies
* Emplaces and moves
* Lazy linear scan of attributes instead of building a map
* Extra printfs, missing header
* Avoid clipping if the input and output tiles are the same
* But do clip if the tile extent is being reduced
* Make sure I'm not constructing std::strings here at runtime
* More worrying about runtime string construction
* A couple more std::moves
* Const references!
* More const references
* Another std::move
* Make the string_value of mvt_value std::optional
* Reserve storage when decoding
* Provision for different mvt_values to share a string pool
* Use the string pool when decoding
* Avoid another string construction
* Try limiting the depth of the search for duplicate attributes
* Revert "Try limiting the depth of the search for duplicate attributes"
This reverts commit
|
||
|
|
cbf222754b |
Add --accumulate-attribute to tippecanoe-overzoom (#189)
* Starting to factor out attribute accumulation into its own file * Continuing to factor out attribute accumulation * Reduce duplicate code * Plumbing the accumulate-attribute option around * Call the attribute accumulator * Test that accumulation works * Add missing #includes * Don't sort within individual multiplier clusters Doing so throws off the spatial distribution of the low zooms * Docs and changelog * Add comments |
||
|
|
679a0d62f2 |
Make feature ordering cooperate with --retain-points-multiplier (#188)
* Make feature ordering cooperate with --retain-points-multiplier * Forgot to check in the actual code changes??? * Sort within each multiplier cluster as well as between clusters * Correct description of behavior in changelog * Drag original feature sequence along in megatiles for post-filter sort * Plumb the preserve-input-order flag through overzoom * Sort in overzoom if requested * Use within-tile input sequence numbers, not global sequence numbers * Documentation * Reverse direction of search to prevent accidental skipping * Add some comments about converting between attribute representations |
||
|
|
5d92a17193 |
Add point retention multiplier (#179)
* Add an option to retain N times as many points as usual at each zoom * Tests for point multipler with specified and guessed maxzooms * Work in progress on inverse spatial ordering * Fix inverse spatial feature order * --reorder was depending on a feature index that wasn't being preserved * Separate ordering by feature_minzoom from ordering inverse-spatially * Add a test for the inverse spatial ordering * Store the basezoom/droprate/multiplier decisions in tileset metadata * Progress on adding filters to tippecanoe-overzoom * Type promotion for comparison * Look up the attribute value for ordering * Add test of thinning and ordering features * Plumb tippecanoe_decisions metadata through pmtiles * Be careful not to put infinities in JSON * Fix accidental dropping in what is meant to preserve sparse points * Start distinguishing true, false, and null in expressions * Most of the type conversions * Add boolean conversions * Literals and conjunctions * Add filtering to tippecanoe-overzoom * Add a test of filtering in overzoom * Fix boolean conjunctions * Handle the combination of cluster size and filtering * Rework dot dropping to reconcile density threshold and multiplier * Revert "Rework dot dropping to reconcile density threshold and multiplier" This reverts commit |
||
|
|
d359461e61 |
Reduce tile-join overzooming memory usage (#162)
* 16 bits is enough for tile numbers
* Revert "16 bits is enough for tile numbers"
This reverts commit
|
||
|
|
cc5c1c79df |
Speed up overzooming in tile-join (#147)
* Clip away entire features by bbox. Avoid unnecessary recompression. * Move parent tile decoding in tile-join out of overzoom proper * An ever-growing cache of parent tiles * Limit the size of the cache * Remove the current reader *before* checking if we can run the queue * Clean up * Add missing #include * Add comment * When the tile-join cache fills up, evict the least recently used * Fix microsecond math * Factoring out tile-join's cache for testing * Add unit tests for tile-join cache * Update changelog and version |
||
|
|
6778aeac52 |
Add an option to extend zooms if still dropping, but with a limit (#131)
* Add an option to extend zooms if still dropping, but with a limit * At least when to overzoom, even if not actually doing it yet * Refactor to give tile-join access to overzoom() * Didn't work, but *might* have worked * OK, it did something now * Ah, there's the bug! * Hook up pmtiles and dirtiles as overzooming sources * Add command line option to enable or disable overzooming * Add (currently broken) test of overzooming in tile-join * Slightly more abstraction for the tile-join readers * Factor out duplicated code * Move construction into a constructor * More changing accessors to methods * Reduce magic * Start tracking a list of the tiles at maxzoom * I think it worked? * Add missing #include * Fix sequence of overzoomed tiles (Y sorts backwards for TMS) * Don't spend memory on overzooming when we aren't going to use it * Diff rather than cmp, in the hope of figuring out this broken test * Keep full coordinate precision if we might extend zooms * Try a slightly different byte limit * Make drop-densest more consistent across tile boundaries * Also affects this test * Does it behave any differently if it can extend forever? * I think the discrepancy is a thread-safety problem here * Revert "Does it behave any differently if it can extend forever?" This reverts commit |
||
|
|
430d8edd17 |
Tool for overzooming individual tiles (#121)
* Starting work on overzooming * Factor clipping out of geometry.cpp to simplify linkage * Pull out more geometry functions into now-badly-named clip.cpp * Not surprisingly, there is a bug * Found the bug * Make indent * Pass attributes through * Make formatting more consistent * Fix typos in comments * Fix the typos better * Add tippecanoe-overzoom to the install list * Give overzoom a predictable exit status * Forgot to translate to and from polygon ring closepaths * Fix geometry collapse at z21 * Add attribute stripping; don't generate layers if they have no features * Add docs for tippecanoe-overzoom (as it will be, not as it is) * Change overzoom to accept input and output files as arguments * Working on overzooming tests * Oops * Hook up and test the detail and buffer options, and the empty-tile case * Also test attribute stripping * Fix error message * Update changelog and version |