Commit Graph
13 Commits
Author SHA1 Message Date
4f2621186a Convert jsonpull to C++ with shared_ptr and std::vector/std::string (#388)
* Rename to jsonpull.cpp

* Clear for merge

* Convert jsonpull to C++ with shared_ptr and std::vector/std::string

Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.

The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Subclass json_object so primitives shrink from 168 to 24 bytes

The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:

  json_object (base, TRUE / FALSE / NULL)   24 bytes
  json_number                                48 bytes
  json_string                                48 bytes
  json_array  (empty)                        48 bytes
  json_hash   (empty)                        72 bytes

Other size wins along the way:

* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
  was 16 bytes per node). json_pull now keeps an explicit
  container_stack and the parser no longer needs to resurrect a
  shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
  original std::make_shared<json_xxx> call, so destroying a
  shared_ptr<json_object> still runs the right subclass dtor.

The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Store hash key/value pairs in one ordered vector

Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:

    for (auto &[k, v] : o->entries()) { ... }

Side effects:

* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
  header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
  single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
  (`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
  original pattern relied on `nprop = 0` to short-circuit iteration on
  a null or non-hash `properties`, the rewrite now guards the loop
  explicitly with `if (o->type == JSON_HASH)` so that calling
  entries() doesn't trip the asserting downcast.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Move parser-only `expect` state out of json_object

`expect` was only meaningful while the parser was building a container,
and only ever read or written from jsonpull.cpp itself; once parsing
finished it was dead weight on every JSON_ARRAY and JSON_HASH (and
present-but-unused on every primitive too). Move it into the parser's
container stack, alongside the shared_ptr to the container it pertains
to:

    struct json_pull::parse_frame {
        json_object_ptr container;
        json_type       expect;
    };
    std::vector<parse_frame> container_stack;

The base class now only carries data-model state (parent, parser, type).
No external caller depended on `expect`, so no sweep was needed outside
jsonpull.cpp.

This change does not, in itself, shrink any json_object: the 4-byte
`expect` field used to live at offset 20 inside the base, where it was
already being eaten by alignment padding for the 8-byte-aligned first
member of every subclass (std::string, std::vector, double). The win is
in the data model, not the byte count -- the 4-byte hole is still
there, but it is now available for a future subclass whose first member
is small enough to slot into it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Discriminate json_number's three numeric slots into one union

json_number used to carry three parallel 8-byte fields (a double plus
both a 64-bit unsigned and a 64-bit signed slot for the large-integer
cases) even though at most one of the integer slots is ever the
canonical value for any given number. Collapse them into a
discriminated union:

    enum repr_t { REPR_DOUBLE, REPR_LARGE_UNSIGNED, REPR_LARGE_SIGNED };
    repr_t repr;
    union { double d; unsigned long long u; long long s; } value;

Callers keep the same read API: number() returns the appropriate
double, large_unsigned() returns the ull (or 0 if not currently stored
that way), large_signed() likewise. Writes go through new set_number /
set_large_unsigned / set_large_signed methods that keep the
discriminator and the union value in sync.

This was prompted by an observation that moving json_type to the end
of the object should shrink things via tail-padding reuse. Empirically
the type-at-end rearrangement saves nothing on its own (every
subclass payload is 8-byte aligned so it can't slot into the 4-byte
tail), but the discriminated-number redesign hits the same idea from
a different direction: adding the 4-byte `repr` to json_number makes
the class non-standard-layout, which lets the Itanium ABI pack `repr`
into the base's 4-byte tail padding at offset 20. The union value
then starts at the natural offset 24, and json_number ends at offset
32 -- a 33% reduction.

Per-node sizes:
  json_object (TRUE/FALSE/NULL)  24 bytes
  json_number                    32 bytes  (was 48)
  json_string                    48 bytes
  json_array                     48 bytes
  json_hash                      48 bytes

Numbers dominate real GeoJSON (every coordinate is one), so the net
memory win on a typical parse is substantial.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix bugs flagged in code review of jsonpull C++ port

- jsontool.cpp `out()`: route JSON_NUMBER (and anything else non-string)
  through `json_stringify` instead of `o->string()`, which now asserts
  on a non-string type and would crash `--extract` on numeric attributes.
- geojson.{hpp,cpp} `json_end_map`: take `json_pull_ptr` by reference so
  the caller's shared_ptr is released, null-guard before touching
  `jp->source`, and clear `jp->source` after delete to avoid a dangling
  pointer.
- jsonpull/jsonpull.cpp: low-surrogate range check was comparing the
  outer-loop byte `c` instead of the parsed code unit `ch`, breaking
  surrogate-pair decoding for some \\uXXXX escapes. Pre-existing bug
  preserved across the port.
- tile-join.cpp `handle_vector_layers`: require the field value to have
  type JSON_STRING (and the key to be non-null) before calling
  `string()`; the previous truthy `type` check would assert on a
  non-string value.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add jsonpull regression test for surrogate-pair decoding

Covers the `c` vs `ch` bug fixed in the previous commit: parsing
"\uD83D\uE000" (a valid high surrogate followed by a non-surrogate
BMP code point) used to mis-classify U+E000 as a low surrogate and
combine the two units into U+1F400 (F0 9F 90 80). The fixed code
flushes the stale high surrogate as standalone CESU-8 (ED A0 BD)
and then encodes U+E000 normally as EE 80 80. Verified the test
fails under the pre-fix logic.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Cheap perf wins in jsonpull C++ port

Profiling tl_2022_us_county.json (sample(1) on Apple Silicon) showed
~38% of parse time in allocator work and ~14% in std::string::push_back
during string-token construction. These changes target the low-hanging
fruit from that profile:

- Pre-reserve 2 slots in json_array and 4 slots in json_hash so
  coordinate `[x, y]` pairs and typical GeoJSON property maps avoid
  the 0 -> 1 -> 2 -> 4 vector-growth chain (and the shared_ptr copies
  it incurs).
- Reuse a parser-wide std::string buffer for JSON_STRING tokens
  instead of constructing a fresh local std::string per token. The
  buffer is cleared (capacity preserved) at the start of each token
  and copied into the final json_string, so once it has grown to the
  longest string seen it stops reallocating entirely.
- std::move the freshly-created container shared_ptr into the parser
  container stack in the `[` and `{` handlers, and move it out of the
  frame on the matching `]` / `}`. Each move skips one atomic
  inc/dec round-trip per container open and close.

On a tl_2022_us_county.json benchmark (4-iter user-time mean, Apple
Silicon, /usr/bin/time):
- main baseline:                              ~8.17s
- jsonpull-cpp before these changes:          ~10.90s  (+33%)
- jsonpull-cpp with these changes:            ~9.33s   (+14%)

So this commit recovers roughly half of the post-port regression.
The remaining gap is dominated by shared_ptr atomic refcount traffic
on the parse tree and per-node heap allocations, which would require
the larger unique_ptr/arena reworks to address.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Make json_free actually free the subtree

In the C++ port, json_free was just `o.reset()`, which dropped the
caller's reference but left the subtree alive: the parent's vector
slot kept it allocated, and for line-delimited streams the parser's
jp->root co-owned it until the next top-level value started parsing.
That defeated the geojson-loop pattern of calling json_free on each
feature after serializing it, which is supposed to release the
feature so it doesn't sit in memory while subsequent ones are parsed.

Restore the historical "remove this from the tree" semantics by
splicing the node out of its parent (sharing splice_from_parent with
json_disconnect) and clearing jp->root when the node is the parser's
current top-level value, then dropping the caller's reference.

Two unit tests pin this down: a pruning test parses
"[[1, 2], [3, 4], [5, 6]]" element-wise and confirms that calling
json_free on [3, 4] leaves the outer array with just [1, 2] and
[5, 6]; a top-level test uses a weak_ptr observer to confirm that
json_free on the parser's root really destroys the tree.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Migrate jsonpull to unique_ptr ownership

Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.

API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
  json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
  takes ownership; back-pointers are cleared so the subtree can
  outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
  unique_ptr ownership of the most recent top-level value.

Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.

In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.

Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix preprocessor mistakes identified by Copilot

* Make indent

* Skip non-string metadata.json entries instead of reading them as strings

dirmeta2tmp() warned about a metadata entry that was not a string/string
pair and then read it as a string anyway. Under the new type-tagged
accessors that trips the assert in json_object::string(); before them it
reinterpreted the node's storage as a char pointer, which segfaulted for
most values. Either way, tippecanoe-decode and tile-join could not read a
directory tileset whose metadata.json had a numeric minzoom or a nested
object, which is common in metadata.json files written by other tools.

Add the missing continue, and cover it in raw-tiles-test.

pmtilesmeta2tmp() handles the same case correctly but read the key with
string() before its own JSON_STRING check, so the assert would have fired
ahead of the check meant to catch a bad key. Hoist the check above the
read. The parser rejects non-string hash keys, so this is unreachable in
practice; the ordering is what makes the check meaningful.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Don't redefine _GNU_SOURCE in the C++ jsonpull port

The `#define _GNU_SOURCE` carried over from jsonpull.c, where it was
needed to get asprintf() declared. g++ already defines _GNU_SOURCE on the
command line for C++ translation units, so redefining it warns:

    jsonpull/jsonpull.cpp:1: warning: "_GNU_SOURCE" redefined

Guard the define rather than drop it, so platforms whose C++ driver does
not predefine it still get asprintf() declared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add a unit test for json_disconnect

json_disconnect() is documented in jsonpull.h as the supported way to
splice a subtree out of the parser's tree and take ownership of it, but
nothing calls it: read_filter() and parse_filter() used to, and now get
the same guarantee from json_read_tree() clearing back-pointers on the way
out. Cover the behavior rather than leave the primitive dead and untested.

The test pins that the subtree is removed from its parent, that the parser
keeps the rest of the tree, and that the detached subtree stays readable
after the json_pull is destroyed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Correct two stale comments in the jsonpull port

jsonpull.h said a json_number is 40 bytes; it is 32 (json_object is 24,
and the repr discriminator fits in the base class's tail padding, so the
8-byte union lands at offset 24).

plugin.cpp's parse_feature() said `j` is freed only just before returning
or as jp->root at end of stream, but there is a third json_free(j) at the
bottom of the loop, for a complete Feature whose geometry came out empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add a changelog entry and bump the version for the jsonpull rewrite

The rewrite is meant to be behavior-preserving, but it carries four
user-visible bug fixes that warrant release notes: tippecanoe-json-tool
--extract on a numeric attribute, surrogate-pair decoding, tile-join
reading a non-string tilejson field type, and non-string values in a
directory tileset's metadata.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Encode U+FFFF as three bytes instead of an overlong four

The \uXXXX decoder tested `ch < 0xFFFF` before taking the three-byte UTF-8
path, so U+FFFF itself fell through to the four-byte branch and came out as
F0 8F BF BF -- an overlong, and therefore invalid, encoding of a code point
that fits in three bytes.

check_utf8() only checks that continuation bytes look like continuation
bytes, not that a sequence is the shortest form, so nothing downstream
noticed: a GeoJSON attribute containing U+FFFF put invalid UTF-8 into the
output tile, where a strict consumer would reject it.

Since `ch` is parsed from exactly four hex digits it cannot exceed 0xFFFF
on its own, so after this change the four-byte branch is reached only for a
code point assembled from a surrogate pair, which is the only way to name
one above the BMP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Keep the parent links inside a detached jsonpull subtree

json_read_tree() and json_disconnect() cleared both back-pointers on every
node of the subtree they handed out. Clearing `parser` throughout is
necessary -- the json_pull can be destroyed while the subtree lives on, so
a surviving `parser` would dangle -- but clearing `parent` throughout cost
more than it bought.

`parent` is a non-owning raw pointer, so keeping it cannot form a reference
cycle or keep anything alive; there is nothing to leak. And within a
detached subtree it refers to nodes the caller now owns as a single unit,
so it stays valid for exactly as long as the subtree itself. Clearing it
only made the tree unwalkable upwards, and made json_free() and
json_disconnect() silently no-ops on interior nodes of a detached tree,
since both find a node's owner through o->parent.

So clear `parser` everywhere and clear `parent` on the detached root alone,
which is the one that pointed out of the subtree at a node the parser still
owns. Split the old clear_back_pointers() into clear_parser_pointers() plus
a detach_subtree() wrapper that adds the root's `parent`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Cover the U+FFFF encoding and detached-tree parent links

Each of the new assertions fails against the previous behavior, so they
pin the two fixes rather than merely passing alongside them:

  - the U+FFFF test, plus the U+FFFE boundary below it and a surrogate pair
    above it, so the three-byte and four-byte paths are both held in place
  - json_disconnect() leaving the parent links inside the subtree intact
    while clearing the root's
  - json_free() pruning an interior node of a tree whose parser is already
    gone, which only works because those links survive
  - json_free() of a hash value leaving the key paired with a JSON_NULL
    placeholder, which is the documented behavior and not a removal

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Address review notes in jsonpull itself

- json_stringify walked c_str(), so it truncated at an embedded NUL even
  though values are std::string now and carry one through faithfully.
  Range over the string instead; the existing control-character branch
  already escapes a NUL like any other, so the output stays valid JSON.
- Assert that the hash has an entry waiting before add_object() assigns to
  entries().back(). It always does -- JSON_VALUE is only set by a colon,
  which requires a pushed key -- but the derivation is not local.
- Drop fabricate_object(), a pass-through to make_object() with the
  arguments reordered, kept only to preserve the old C name.
- Inline the string_append / string_append_c wrappers over push_back and
  append, and note that json_print_one's JSON_HASH and JSON_ARRAY branches
  are unreachable, since json_print handles both itself.
- json_hash_get's comment said nullptr meant "the matching value is null",
  which reads as JSON null. A JSON null comes back as a JSON_NULL node;
  nullptr means the value slot is not filled in yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Tidy jsonpull call sites flagged in review

- geojson.cpp and attribute.cpp passed key_pool::pool() and
  set_attribute_accum() a c_str() from a std::string, forcing a needless
  reconstruction (and truncating at an embedded NUL). Both overloads take
  std::string, so pass it directly. The geojson.cpp one is the hottest
  loop in the program.
- Replace the hand-maintained counters beside range-for loops in
  attribute.cpp, main.cpp and tile-join.cpp with indexed loops, since the
  index is only wanted for error messages.
- parse_json_args took json_pull_ptr by value and then copied it, costing
  two refcount bumps per construction. Move it.
- Assert that the parser is still attached where geojson.cpp reads
  geometry->parser->line. Only json_read results reach it today, but
  json_read_tree and json_disconnect now clear every parser pointer, so a
  detached tree would null-deref there instead of tripping an assert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Pin the array-splicing fix, and close four test gaps

The existing pruning test does not discriminate: json_read hands back each
container as it completes, so the node it frees is always the most recently
added element of its parent -- the one case the old element-count-vs-byte-
count memmove got right, because it then moved zero bytes. Widening that
test to more elements does not change this; the shape is what matters, not
the size. Verified: the eight-element streaming variant still passes
against the pre-fix code.

Add a test that builds the array first and then prunes element 0 of eight,
asserting the identity of every survivor rather than just the resulting
count. That fails against the pre-fix code deterministically, with
arr[0] == arr[1] and the last element dropped. Note the limitation on the
streaming test so the next reader does not try to strengthen it in place.

Also cover, all previously untested:

- json_free of a hash key, and of both halves of a pair, where the entry
  survives with a JSON_NULL stand-in until both are gone
- repeated json_read_tree over a line-delimited stream, which is what the
  filter loaders and -L / -E do, asserting each detached tree survives the
  next read and the parser's destruction
- json_stringify of a partially-parsed tree, the json_context error path
- json_stringify across an embedded NUL, which fails against the c_str()
  walk this branch replaces

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Add changelog entries for three unadvertised fixes

The array-splicing fix goes first: it is a memory-corruption fix, and it is
the strongest illustration of why the ownership model is worth having,
since it is exactly the failure the model makes unrepresentable.

Also the uninitialized read when a filter emitted "properties": null, and
the evaluator.hpp include guard that defined EVALUATOR HPP and so never
guarded anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

* Consolidate the jsonpull comments

Comments were 25% of the added lines, and the ownership model was spelled
out in five places. Collect it into one block at the top of jsonpull.h and
point at it from the rest, cutting the ratio to 14% and the total by about
200 lines.

Removed the duplicate explanations of the deleter dispatch, of what detach
does to the back-pointers, and of "json_read returns intermediate
containers, do not free them". Trimmed the comments that argued for a
choice rather than described the code -- the reserve(2) / reserve(4)
rationales, the string-buffer copy, the pmtiles check ordering -- to a line
each, and shortened the test preambles, keeping the parts that say why a
test is shaped the way it is.

No code changes; the test suite is unchanged in both configurations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-13 21:27:07 -07:00
Brandon Keepers 1b060b7faf Drop a hole that no ring can parent instead of failing the run (#401)
* wagyu: drop a hole no remaining ring can parent instead of throwing

correct_tree() throws "Could not properly place hole to a parent" when
topology correction leaves a hole whose parent ring was removed (degenerate
input such as stacked duplicate rings from coalesced tiny-polygon
placeholders). That aborts the entire tiling run over one unrepresentable
sliver. Remove the ring and its points instead, matching how other
unresolvable degeneracies are handled.

* Add a regression test for dropping an unplaceable hole

A fuzzer-minimized pair of mutually reversed self-intersecting rings that
makes wagyu's correct_tree fail to find a parent for a hole — the same
failure reported in mapbox/tippecanoe#761. Before the topology_correction
change, running this test exits with EXIT_IMPOSSIBLE via the polygon
cleaning error handler; with it, the clean returns.
2026-08-05 08:45:39 -07:00
Mike Jones c82e4beee3 Fix: Respect -t temporary directory option in sorting operations (#368)
Enhance fqsort function to accept a temporary directory parameter for file handling. Update calls to fqsort in main.cpp, sort.cpp, sort.hpp, and unit.cpp to utilize the new parameter, ensuring temporary files are created in the specified directory.
2025-09-24 09:09:40 -07:00
Erica Fischer 533e000faa Remove undocumented command-line options (#361)
* Remove --accumulate-numeric-attributes

* Remove join-sqlite, etc.

* Remove --accumulate-numeric-attributes from overzoom

* Remove --assign-to-bins and --bin-by-id-list

* Remove --clip-polygon and --clip-bounding-box

* Remove FSL expressions

* Update version and changelog
2025-07-31 17:03:01 -07:00
Erica Fischer 2d548bed06 Infinite loop fixes, minimizing changes to behavior (#345)
* Divide-and-conquer polygon cleaning

* Catch the case where the gap can't be increased further

* Catch the case where we try to keep impossibly many features

* Make label points earlier in the tiling process

* Another case where it could try to drop even after already limiting.

* And do not coalesce on impossibly small geometries

* Add missing return

* Update version and changelog
2025-05-09 09:08:28 -07:00
Erica Fischer 5b18eea673 Work in progress on binning features in overzoom (#258)
* Factoring out tilestats management from GeoJSON file reading

* Move code around so overzoom can link against parse_layers

* Read the file of bins

* Plumb the bins through to overzoom()

* Some zip code bins to test with

* (Currently non-functional) test of binning

* Starting to spell out the bin matching loop

* Can't flatten points, so don't flatten bins either

* More fleshing out bin traversal

* Bounding box of tile-relative mvt geometry

* Smallest enclosing tile from bbox

* Most of the bin scan

* Add point in polygon check. It crashes.

* Find the matching bins

* GDAL-style bounding boxes have eaten my brain

* Make some features to bin into

* Increment a count as features are found to be within the bins

* Fix longitude wraparound in overzoom bins

* Fix the tests

* Push off attribute copying until after bin assignment

* Carry sum of numeric attributes into the bins

* Also add mean, min, and max

* Add --calculate-feature-index since I keep needing it for testing

* Add an option to accumulate sum/mean/max/min/count of all numeric attrs

* Don't bake in tippecanoe:mean, since we redo it from sum and count

* Forgot to update this test fixture after removing tiled mean

* Update version and changelog
2024-09-05 12:06:51 -07:00
Erica Fischer 47e774adc4 Improve the appearance of coalesce-densest-as-needed tiles (#247)
* Start to distinguish fixed cluster density setting from as-needed density

* Make consistent {drop,coalesce}-densest decisions between zooms

* Actually track the previous index instead of just intending to

* Clean up collinearities in coalesced features

* To determine densest, look at actual physical distance, not just index

* Don't actually need the previous index in serial_feature now

* Center of mass of one feature to most distant point of the next

* Add apologetic comment

* Wait, how did the tests pass before?

* Revert "Wait, how did the tests pass before?"

This reverts commit f73c8ee543.

* Add --maximum-string-attribute-length option

* Update version and changelog

* A little more testing to make sure
2024-07-16 13:23:23 -07:00
Erica Fischer 4e52cbd957 Drop or retain whole multiplier clusters when dropping as needed (#198)
* Prep to track conditions other than just "dropped" or "kept"

* Count up instead of down

* Drop or retain whole multiplier clusters based on their first feature

* Calculate a global feature dropping sequence

* Switch over to using the drop sequence for drop-fraction

* Remove unused arguments for the old drop-fraction implementation

* Fix copy-and-paste bugs, update tests

* Properly incorporate feature_minzoom into the drop sequence, I hope

* Rename drop_by to drop_sequence

* See if sorting within clusters fixes filter stability between zooms

* Remove very chatty debug print

* Update changelog and version

* Use named constants instead of numbers for feature dropping/keeping

* Add comment to explain purpose and method of bit reversal
2024-02-12 10:58:49 -08:00
Erica Fischer cc5c1c79df Speed up overzooming in tile-join (#147)
* Clip away entire features by bbox. Avoid unnecessary recompression.

* Move parent tile decoding in tile-join out of overzoom proper

* An ever-growing cache of parent tiles

* Limit the size of the cache

* Remove the current reader *before* checking if we can run the queue

* Clean up

* Add missing #include

* Add comment

* When the tile-join cache fills up, evict the least recently used

* Fix microsecond math

* Factoring out tile-join's cache for testing

* Add unit tests for tile-join cache

* Update changelog and version
2023-10-04 12:14:22 -07:00
Erica Fischer 26bf08deb2 Externalizing polygon shard detection (#146)
* Start of externalizing polygon shard detection

* Completely untested external quicksort

* Add unit test for external quicksort

* Remember to clean up temporary files

* Sort and scan the vertices

* Bring over more vertex logic

* Make nodes from vertices

* Checkpoint on switching over to global shared nodes

* Do the thing

* Revert unintended change to coalesced linestring behavior

* Take shared nodes into account in early simplification

* Let it do more sorting in memory

* Fix overnoding of collinear linestrings

* Remove duplicate nodes, since only duplicate vertices now matter

* Remember to delete temporary files

* Fix out of bounds memory access below, apparently

* Fix the actual undefined behavior

* Still running out of memory in one case. Find out where.

* Forgot the conditional

* Try again to make it not run out of memory

* Update version and changelog
2023-10-03 10:53:12 -07:00
Eric Fischer 5a09fcc35e Some basic unit tests for string truncation 2017-07-21 14:27:30 -07:00
Eric Fischer d4d966893c Forgot to test the emoji case 2016-10-05 15:01:47 -07:00
Eric Fischer 9806db3c0a Make UTF-8 checking into a unit test with Catch 2016-10-05 14:55:32 -07:00