Commit Graph
13 Commits
Author SHA1 Message Date
Claude 4ab8c66b86 Consolidate the jsonpull comments
Comments were 25% of the added lines, and the ownership model was spelled
out in five places. Collect it into one block at the top of jsonpull.h and
point at it from the rest, cutting the ratio to 14% and the total by about
200 lines.

Removed the duplicate explanations of the deleter dispatch, of what detach
does to the back-pointers, and of "json_read returns intermediate
containers, do not free them". Trimmed the comments that argued for a
choice rather than described the code -- the reserve(2) / reserve(4)
rationales, the string-buffer copy, the pmtiles check ordering -- to a line
each, and shortened the test preambles, keeping the parts that say why a
test is shaped the way it is.

No code changes; the test suite is unchanged in both configurations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r
2026-08-14 04:17:06 +00:00
Claude b5c3cd6642 Skip non-string metadata.json entries instead of reading them as strings
dirmeta2tmp() warned about a metadata entry that was not a string/string
pair and then read it as a string anyway. Under the new type-tagged
accessors that trips the assert in json_object::string(); before them it
reinterpreted the node's storage as a char pointer, which segfaulted for
most values. Either way, tippecanoe-decode and tile-join could not read a
directory tileset whose metadata.json had a numeric minzoom or a nested
object, which is common in metadata.json files written by other tools.

Add the missing continue, and cover it in raw-tiles-test.

pmtilesmeta2tmp() handles the same case correctly but read the key with
string() before its own JSON_STRING check, so the assert would have fired
ahead of the check meant to catch a bad key. Hoist the check above the
read. The parser rejects non-string hash keys, so this is unreachable in
practice; the ordering is what makes the check meaningful.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r
2026-08-12 18:16:21 +00:00
Erica FischerandCursor 65beb3f0d7 Migrate jsonpull to unique_ptr ownership
Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.

API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
  json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
  takes ownership; back-pointers are cleared so the subtree can
  outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
  unique_ptr ownership of the most recent top-level value.

Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.

In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.

Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 20:44:25 -07:00
Erica FischerandCursor 4d9a48c3d4 Store hash key/value pairs in one ordered vector
Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:

    for (auto &[k, v] : o->entries()) { ... }

Side effects:

* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
  header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
  single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
  (`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
  original pattern relied on `nprop = 0` to short-circuit iteration on
  a null or non-hash `properties`, the rewrite now guards the loop
  explicitly with `if (o->type == JSON_HASH)` so that calling
  entries() doesn't trip the asserting downcast.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:13 -07:00
Erica FischerandCursor f366b2c4aa Subclass json_object so primitives shrink from 168 to 24 bytes
The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:

  json_object (base, TRUE / FALSE / NULL)   24 bytes
  json_number                                48 bytes
  json_string                                48 bytes
  json_array  (empty)                        48 bytes
  json_hash   (empty)                        72 bytes

Other size wins along the way:

* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
  was 16 bytes per node). json_pull now keeps an explicit
  container_stack and the parser no longer needs to resurrect a
  shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
  original std::make_shared<json_xxx> call, so destroying a
  shared_ptr<json_object> still runs the right subclass dtor.

The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:08 -07:00
Erica FischerandCursor 3da03c6075 Convert jsonpull to C++ with shared_ptr and std::vector/std::string
Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.

The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:48:32 -07:00
Erica Fischer 312e1560a3 Stabilize feature order in overzoom (#210)
* Stabilize feature order in overzoom

* I want my sorts to be stable, please

* Revert "[ci] test in debug mode (#202)"

This reverts commit 853ada87b5.

* No need to reinitialize here
2024-02-29 10:39:10 -08:00
Erica Fischer 7d4d264d50 Speeding up tippecanoe-overzoom (#191)
* Speed up mvt_value comparison

* Converting repetitive ifs to cases

* More conversions from ifs to cases

* Optimize the always-true filter case

* Don't convert types of attributes without accumulators

* Unordered map seems to be faster than map

* Add missing header

* Fix some warnings

* Fix the warnings better

* Avoid an int->string->int conversion

* Lazily initialize layer key and values maps when actually needed

* More switches from maps to unordered_maps

* Sure, I'll take the microoptimization

* More emplacement

* Save some copies

* Emplaces and moves

* Lazy linear scan of attributes instead of building a map

* Extra printfs, missing header

* Avoid clipping if the input and output tiles are the same

* But do clip if the tile extent is being reduced

* Make sure I'm not constructing std::strings here at runtime

* More worrying about runtime string construction

* A couple more std::moves

* Const references!

* More const references

* Another std::move

* Make the string_value of mvt_value std::optional

* Reserve storage when decoding

* Provision for different mvt_values to share a string pool

* Use the string pool when decoding

* Avoid another string construction

* Try limiting the depth of the search for duplicate attributes

* Revert "Try limiting the depth of the search for duplicate attributes"

This reverts commit 9ec94a15ff.

* Update changelog

* Fix typo noticed during code review
2024-01-29 11:33:56 -08:00
Erica Fischer 5d92a17193 Add point retention multiplier (#179)
* Add an option to retain N times as many points as usual at each zoom

* Tests for point multipler with specified and guessed maxzooms

* Work in progress on inverse spatial ordering

* Fix inverse spatial feature order

* --reorder was depending on a feature index that wasn't being preserved

* Separate ordering by feature_minzoom from ordering inverse-spatially

* Add a test for the inverse spatial ordering

* Store the basezoom/droprate/multiplier decisions in tileset metadata

* Progress on adding filters to tippecanoe-overzoom

* Type promotion for comparison

* Look up the attribute value for ordering

* Add test of thinning and ordering features

* Plumb tippecanoe_decisions metadata through pmtiles

* Be careful not to put infinities in JSON

* Fix accidental dropping in what is meant to preserve sparse points

* Start distinguishing true, false, and null in expressions

* Most of the type conversions

* Add boolean conversions

* Literals and conjunctions

* Add filtering to tippecanoe-overzoom

* Add a test of filtering in overzoom

* Fix boolean conjunctions

* Handle the combination of cluster size and filtering

* Rework dot dropping to reconcile density threshold and multiplier

* Revert "Rework dot dropping to reconcile density threshold and multiplier"

This reverts commit f253a66382.

* Retain points by multiplier within each tile, not in global probability

* Test that intends to verify that the multiplier is reversible

* Get the test to detect the discrepancy

* Mark the start of multiplier clusters with a magic attribute

* Add string-contains

* Add in and ni operators

* Revert "Look up the attribute value for ordering"

This reverts commit 56bc73e49a.

* Revert "Type promotion for comparison"

This reverts commit 6f3256f5af.

* Make number formatting in tippecanoe_decisions consistent

* Revert "Add a test for the inverse spatial ordering"

This reverts commit c8047de9ab.

* Revert "Separate ordering by feature_minzoom from ordering inverse-spatially"

This reverts commit 35b19a223c.

* Revert "Fix inverse spatial feature order"

This reverts commit 5978ecdb44.

* Revert "Work in progress on inverse spatial ordering"

This reverts commit fdf230f632.

* Somehow missed the tests associated with that last revert

* Round-robin assign attributes to partials from across the multiplier

* Count the multiplier separately in each layer

* Fix distribution of accumulated attribute across multiplier features

* Update changelog, version, and docs

* Add "is null" and "isnt null" expressions

* Update interpretation of FSL expressions to pass the tests

* Test to assert that polygons are unaffected by the multiplier

* Clean up and comment

* Remove accidental unused case
2024-01-18 15:56:09 -08:00
Erica Fischer b5ec3c46f0 Reduce excessive progress logging during pmtiles conversion (#111)
* Don't log progress so often during pmtiles conversion

And turn off the `catch` tests, which have bit-rotted

* Fix accidental double-multiplication-by-100

* Reenable catch

* Update changelog
2023-07-10 15:59:52 -07:00
Erica Fischer afc7ca01a2 Calculate an antimeridian-adjusted bounding box in tileset metadata (#82)
* Calculate a new antimeridian-adjusted bounding box

* Add antimeridian bounding box to pmtiles, dirtiles, and tile-join

* Don't take out-of-bounds latitudes into account in the adjusted bbox

* Update changelog

* Forgot to adjust tests after the last change
2023-03-09 20:26:48 -08:00
Erica Fischer f2127fec97 Reduce thrashing during feature ingestion and tiling (#56)
* Remove the concept of "separate metadata"

This was an extra level of attribute indirection (features point
to metadata records which point to key and value strings) which was
intended to reduce the size of temporary storage for features with
large numbers of attributes that were also spread across large numbers
of tiles at maxzoom.

For other kinds of features, the extra indirection slowed things down
instead, and, especially when maxzoom guessing was being used, many more
features were having their metadata externalized than could actually
benefit from it.

* Shave a few bytes off temporary files by using more unsigned integers

* Flush stderr after logging progress

* Revert "Shave a few bytes off temporary files by using more unsigned integers"

This reverts commit eef29084ec.

* Limit the size of the string pools and trees to fit in memory

* Add missing #include

* Move the string pool and search tree from mmap to allocated memory

* Sort in allocated rather than mapped memory too

* Also use pread instead of mapping to read in the data to sort

* When the pool gets too big, switch to just the file, not memory

* Switch string pool from memory to disk when memory is 10% full

* Add to-memory versions of the serialization functions

* Crashy work in progress toward compression

* Fix the pointer bug that was causing the crash

* Serialize features into memory rather than straight to disk

* Compress individual features in the temporary files

* Don't need to store the length of the geometry

* Remove per-feature compression; move minzoom back into the object

* Start adding a stream compressor object

* Track file position within fwrite_check()

* Add compressed stream writer functions

* Pull the writing of the serialized feature out to the callers

* Starting toward compression again from a different point

* Hook up more compression functions

* Remove unused code from the other day

* Make enough deflate calls to flush out all the buffered data

* Start on decompression

* Tile number is uncompressed, tile content is compressed

* Work on alternating compressed and uncompressed in decompression

* Closer, but still doesn't work

* Sort of works

* Works until we get to concatenated tiles

* More attempts that don't work

* One bug down

* It made a tileset!

* Handle nonzero initial zooms

* Fix seeking within compressed feature streams

* Tests pass!

* Remove debug spew

* Oops: remember to delete the temporary files so they don't hang around

* Test that fails with the current compression code

* Properly account for bytes read while closing the compressed stream

* Limit the number of warnings about bad label points

* A little more armor when closing decompression

* This time for sure

* A different, less fragile, test that failed previously with compression

* Move feature stream compression to its own file

* Remove now-unused code to deserialize from a file

* Forgot to add the new files

* Remove a little debugging logging

* Add a couple of comments on what it means to be within decompression

* Fix indentation

* Update changelog. Remove stray debugging comment.
2023-02-14 12:47:40 -08:00
Erica FischerandBrandon Liu 3b2599f587 Add support for pmtiles output format (#55)
* add pmtiles.hpp from github.com/protomaps/PMTiles [#10]

* tippecanoe main writes pmtiles output. [#10]

* detect output format using suffix
* after mbtiles is done writing, replace with pmtiles based on map/image tables.
* add method to write_json for writing json sub-object.

* tippecanoe-decode reads pmtiles input. [#10]

* tile-join reads and writes pmtiles. [#10]

* pmtiles test suite for decode and tile-join [#10]

* add base GitHub CI action for compiling and test suite.

* update pmtiles.hpp with z>15 fix

* Fix some ordering problems with pmtiles decode

* Pmtiles should also pass the raw tiles tests

* Eradicate spaces from tileset metadata JSON fields

* Eradicate spaces from more test fixtures

* Update more tests

* Pmtiles tests pass now too

* Remove unnecessary sort (and make indent)

* Update changelog

* The allow-existing test for pmtiles needs -o, not -e

* Declare --allow-existing to be unsupported for pmtiles.

It was always a bad idea even for mbtiles.

Co-authored-by: Brandon Liu <bdon@bdon.org>
2022-12-29 12:00:13 -08:00