Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.
API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
takes ownership; back-pointers are cleared so the subtree can
outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
unique_ptr ownership of the most recent top-level value.
Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.
In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.
Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.
Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:
for (auto &[k, v] : o->entries()) { ... }
Side effects:
* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
(`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
original pattern relied on `nprop = 0` to short-circuit iteration on
a null or non-hash `properties`, the rewrite now guards the loop
explicitly with `if (o->type == JSON_HASH)` so that calling
entries() doesn't trip the asserting downcast.
Co-authored-by: Cursor <cursoragent@cursor.com>
The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:
json_object (base, TRUE / FALSE / NULL) 24 bytes
json_number 48 bytes
json_string 48 bytes
json_array (empty) 48 bytes
json_hash (empty) 72 bytes
Other size wins along the way:
* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
was 16 bytes per node). json_pull now keeps an explicit
container_stack and the parser no longer needs to resurrect a
shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
original std::make_shared<json_xxx> call, so destroying a
shared_ptr<json_object> still runs the right subclass dtor.
The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.
Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.
The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Stabilize feature order in overzoom
* I want my sorts to be stable, please
* Revert "[ci] test in debug mode (#202)"
This reverts commit 853ada87b5.
* No need to reinitialize here
* Speed up mvt_value comparison
* Converting repetitive ifs to cases
* More conversions from ifs to cases
* Optimize the always-true filter case
* Don't convert types of attributes without accumulators
* Unordered map seems to be faster than map
* Add missing header
* Fix some warnings
* Fix the warnings better
* Avoid an int->string->int conversion
* Lazily initialize layer key and values maps when actually needed
* More switches from maps to unordered_maps
* Sure, I'll take the microoptimization
* More emplacement
* Save some copies
* Emplaces and moves
* Lazy linear scan of attributes instead of building a map
* Extra printfs, missing header
* Avoid clipping if the input and output tiles are the same
* But do clip if the tile extent is being reduced
* Make sure I'm not constructing std::strings here at runtime
* More worrying about runtime string construction
* A couple more std::moves
* Const references!
* More const references
* Another std::move
* Make the string_value of mvt_value std::optional
* Reserve storage when decoding
* Provision for different mvt_values to share a string pool
* Use the string pool when decoding
* Avoid another string construction
* Try limiting the depth of the search for duplicate attributes
* Revert "Try limiting the depth of the search for duplicate attributes"
This reverts commit 9ec94a15ff.
* Update changelog
* Fix typo noticed during code review
* Add an option to retain N times as many points as usual at each zoom
* Tests for point multipler with specified and guessed maxzooms
* Work in progress on inverse spatial ordering
* Fix inverse spatial feature order
* --reorder was depending on a feature index that wasn't being preserved
* Separate ordering by feature_minzoom from ordering inverse-spatially
* Add a test for the inverse spatial ordering
* Store the basezoom/droprate/multiplier decisions in tileset metadata
* Progress on adding filters to tippecanoe-overzoom
* Type promotion for comparison
* Look up the attribute value for ordering
* Add test of thinning and ordering features
* Plumb tippecanoe_decisions metadata through pmtiles
* Be careful not to put infinities in JSON
* Fix accidental dropping in what is meant to preserve sparse points
* Start distinguishing true, false, and null in expressions
* Most of the type conversions
* Add boolean conversions
* Literals and conjunctions
* Add filtering to tippecanoe-overzoom
* Add a test of filtering in overzoom
* Fix boolean conjunctions
* Handle the combination of cluster size and filtering
* Rework dot dropping to reconcile density threshold and multiplier
* Revert "Rework dot dropping to reconcile density threshold and multiplier"
This reverts commit f253a66382.
* Retain points by multiplier within each tile, not in global probability
* Test that intends to verify that the multiplier is reversible
* Get the test to detect the discrepancy
* Mark the start of multiplier clusters with a magic attribute
* Add string-contains
* Add in and ni operators
* Revert "Look up the attribute value for ordering"
This reverts commit 56bc73e49a.
* Revert "Type promotion for comparison"
This reverts commit 6f3256f5af.
* Make number formatting in tippecanoe_decisions consistent
* Revert "Add a test for the inverse spatial ordering"
This reverts commit c8047de9ab.
* Revert "Separate ordering by feature_minzoom from ordering inverse-spatially"
This reverts commit 35b19a223c.
* Revert "Fix inverse spatial feature order"
This reverts commit 5978ecdb44.
* Revert "Work in progress on inverse spatial ordering"
This reverts commit fdf230f632.
* Somehow missed the tests associated with that last revert
* Round-robin assign attributes to partials from across the multiplier
* Count the multiplier separately in each layer
* Fix distribution of accumulated attribute across multiplier features
* Update changelog, version, and docs
* Add "is null" and "isnt null" expressions
* Update interpretation of FSL expressions to pass the tests
* Test to assert that polygons are unaffected by the multiplier
* Clean up and comment
* Remove accidental unused case
* Don't log progress so often during pmtiles conversion
And turn off the `catch` tests, which have bit-rotted
* Fix accidental double-multiplication-by-100
* Reenable catch
* Update changelog
* Calculate a new antimeridian-adjusted bounding box
* Add antimeridian bounding box to pmtiles, dirtiles, and tile-join
* Don't take out-of-bounds latitudes into account in the adjusted bbox
* Update changelog
* Forgot to adjust tests after the last change
* Remove the concept of "separate metadata"
This was an extra level of attribute indirection (features point
to metadata records which point to key and value strings) which was
intended to reduce the size of temporary storage for features with
large numbers of attributes that were also spread across large numbers
of tiles at maxzoom.
For other kinds of features, the extra indirection slowed things down
instead, and, especially when maxzoom guessing was being used, many more
features were having their metadata externalized than could actually
benefit from it.
* Shave a few bytes off temporary files by using more unsigned integers
* Flush stderr after logging progress
* Revert "Shave a few bytes off temporary files by using more unsigned integers"
This reverts commit eef29084ec.
* Limit the size of the string pools and trees to fit in memory
* Add missing #include
* Move the string pool and search tree from mmap to allocated memory
* Sort in allocated rather than mapped memory too
* Also use pread instead of mapping to read in the data to sort
* When the pool gets too big, switch to just the file, not memory
* Switch string pool from memory to disk when memory is 10% full
* Add to-memory versions of the serialization functions
* Crashy work in progress toward compression
* Fix the pointer bug that was causing the crash
* Serialize features into memory rather than straight to disk
* Compress individual features in the temporary files
* Don't need to store the length of the geometry
* Remove per-feature compression; move minzoom back into the object
* Start adding a stream compressor object
* Track file position within fwrite_check()
* Add compressed stream writer functions
* Pull the writing of the serialized feature out to the callers
* Starting toward compression again from a different point
* Hook up more compression functions
* Remove unused code from the other day
* Make enough deflate calls to flush out all the buffered data
* Start on decompression
* Tile number is uncompressed, tile content is compressed
* Work on alternating compressed and uncompressed in decompression
* Closer, but still doesn't work
* Sort of works
* Works until we get to concatenated tiles
* More attempts that don't work
* One bug down
* It made a tileset!
* Handle nonzero initial zooms
* Fix seeking within compressed feature streams
* Tests pass!
* Remove debug spew
* Oops: remember to delete the temporary files so they don't hang around
* Test that fails with the current compression code
* Properly account for bytes read while closing the compressed stream
* Limit the number of warnings about bad label points
* A little more armor when closing decompression
* This time for sure
* A different, less fragile, test that failed previously with compression
* Move feature stream compression to its own file
* Remove now-unused code to deserialize from a file
* Forgot to add the new files
* Remove a little debugging logging
* Add a couple of comments on what it means to be within decompression
* Fix indentation
* Update changelog. Remove stray debugging comment.
* add pmtiles.hpp from github.com/protomaps/PMTiles [#10]
* tippecanoe main writes pmtiles output. [#10]
* detect output format using suffix
* after mbtiles is done writing, replace with pmtiles based on map/image tables.
* add method to write_json for writing json sub-object.
* tippecanoe-decode reads pmtiles input. [#10]
* tile-join reads and writes pmtiles. [#10]
* pmtiles test suite for decode and tile-join [#10]
* add base GitHub CI action for compiling and test suite.
* update pmtiles.hpp with z>15 fix
* Fix some ordering problems with pmtiles decode
* Pmtiles should also pass the raw tiles tests
* Eradicate spaces from tileset metadata JSON fields
* Eradicate spaces from more test fixtures
* Update more tests
* Pmtiles tests pass now too
* Remove unnecessary sort (and make indent)
* Update changelog
* The allow-existing test for pmtiles needs -o, not -e
* Declare --allow-existing to be unsupported for pmtiles.
It was always a bad idea even for mbtiles.
Co-authored-by: Brandon Liu <bdon@bdon.org>