- json_stringify walked c_str(), so it truncated at an embedded NUL even
though values are std::string now and carry one through faithfully.
Range over the string instead; the existing control-character branch
already escapes a NUL like any other, so the output stays valid JSON.
- Assert that the hash has an entry waiting before add_object() assigns to
entries().back(). It always does -- JSON_VALUE is only set by a colon,
which requires a pushed key -- but the derivation is not local.
- Drop fabricate_object(), a pass-through to make_object() with the
arguments reordered, kept only to preserve the old C name.
- Inline the string_append / string_append_c wrappers over push_back and
append, and note that json_print_one's JSON_HASH and JSON_ARRAY branches
are unreachable, since json_print handles both itself.
- json_hash_get's comment said nullptr meant "the matching value is null",
which reads as JSON null. A JSON null comes back as a JSON_NULL node;
nullptr means the value slot is not filled in yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r
json_read_tree() and json_disconnect() cleared both back-pointers on every
node of the subtree they handed out. Clearing `parser` throughout is
necessary -- the json_pull can be destroyed while the subtree lives on, so
a surviving `parser` would dangle -- but clearing `parent` throughout cost
more than it bought.
`parent` is a non-owning raw pointer, so keeping it cannot form a reference
cycle or keep anything alive; there is nothing to leak. And within a
detached subtree it refers to nodes the caller now owns as a single unit,
so it stays valid for exactly as long as the subtree itself. Clearing it
only made the tree unwalkable upwards, and made json_free() and
json_disconnect() silently no-ops on interior nodes of a detached tree,
since both find a node's owner through o->parent.
So clear `parser` everywhere and clear `parent` on the detached root alone,
which is the one that pointed out of the subtree at a node the parser still
owns. Split the old clear_back_pointers() into clear_parser_pointers() plus
a detach_subtree() wrapper that adds the root's `parent`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r
jsonpull.h said a json_number is 40 bytes; it is 32 (json_object is 24,
and the repr discriminator fits in the base class's tail padding, so the
8-byte union lands at offset 24).
plugin.cpp's parse_feature() said `j` is freed only just before returning
or as jp->root at end of stream, but there is a third json_free(j) at the
bottom of the loop, for a complete Feature whose geometry came out empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r
Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.
API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
takes ownership; back-pointers are cleared so the subtree can
outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
unique_ptr ownership of the most recent top-level value.
Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.
In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.
Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.
Co-authored-by: Cursor <cursoragent@cursor.com>
Profiling tl_2022_us_county.json (sample(1) on Apple Silicon) showed
~38% of parse time in allocator work and ~14% in std::string::push_back
during string-token construction. These changes target the low-hanging
fruit from that profile:
- Pre-reserve 2 slots in json_array and 4 slots in json_hash so
coordinate `[x, y]` pairs and typical GeoJSON property maps avoid
the 0 -> 1 -> 2 -> 4 vector-growth chain (and the shared_ptr copies
it incurs).
- Reuse a parser-wide std::string buffer for JSON_STRING tokens
instead of constructing a fresh local std::string per token. The
buffer is cleared (capacity preserved) at the start of each token
and copied into the final json_string, so once it has grown to the
longest string seen it stops reallocating entirely.
- std::move the freshly-created container shared_ptr into the parser
container stack in the `[` and `{` handlers, and move it out of the
frame on the matching `]` / `}`. Each move skips one atomic
inc/dec round-trip per container open and close.
On a tl_2022_us_county.json benchmark (4-iter user-time mean, Apple
Silicon, /usr/bin/time):
- main baseline: ~8.17s
- jsonpull-cpp before these changes: ~10.90s (+33%)
- jsonpull-cpp with these changes: ~9.33s (+14%)
So this commit recovers roughly half of the post-port regression.
The remaining gap is dominated by shared_ptr atomic refcount traffic
on the parse tree and per-node heap allocations, which would require
the larger unique_ptr/arena reworks to address.
Co-authored-by: Cursor <cursoragent@cursor.com>
json_number used to carry three parallel 8-byte fields (a double plus
both a 64-bit unsigned and a 64-bit signed slot for the large-integer
cases) even though at most one of the integer slots is ever the
canonical value for any given number. Collapse them into a
discriminated union:
enum repr_t { REPR_DOUBLE, REPR_LARGE_UNSIGNED, REPR_LARGE_SIGNED };
repr_t repr;
union { double d; unsigned long long u; long long s; } value;
Callers keep the same read API: number() returns the appropriate
double, large_unsigned() returns the ull (or 0 if not currently stored
that way), large_signed() likewise. Writes go through new set_number /
set_large_unsigned / set_large_signed methods that keep the
discriminator and the union value in sync.
This was prompted by an observation that moving json_type to the end
of the object should shrink things via tail-padding reuse. Empirically
the type-at-end rearrangement saves nothing on its own (every
subclass payload is 8-byte aligned so it can't slot into the 4-byte
tail), but the discriminated-number redesign hits the same idea from
a different direction: adding the 4-byte `repr` to json_number makes
the class non-standard-layout, which lets the Itanium ABI pack `repr`
into the base's 4-byte tail padding at offset 20. The union value
then starts at the natural offset 24, and json_number ends at offset
32 -- a 33% reduction.
Per-node sizes:
json_object (TRUE/FALSE/NULL) 24 bytes
json_number 32 bytes (was 48)
json_string 48 bytes
json_array 48 bytes
json_hash 48 bytes
Numbers dominate real GeoJSON (every coordinate is one), so the net
memory win on a typical parse is substantial.
Co-authored-by: Cursor <cursoragent@cursor.com>
`expect` was only meaningful while the parser was building a container,
and only ever read or written from jsonpull.cpp itself; once parsing
finished it was dead weight on every JSON_ARRAY and JSON_HASH (and
present-but-unused on every primitive too). Move it into the parser's
container stack, alongside the shared_ptr to the container it pertains
to:
struct json_pull::parse_frame {
json_object_ptr container;
json_type expect;
};
std::vector<parse_frame> container_stack;
The base class now only carries data-model state (parent, parser, type).
No external caller depended on `expect`, so no sweep was needed outside
jsonpull.cpp.
This change does not, in itself, shrink any json_object: the 4-byte
`expect` field used to live at offset 20 inside the base, where it was
already being eaten by alignment padding for the 8-byte-aligned first
member of every subclass (std::string, std::vector, double). The win is
in the data model, not the byte count -- the 4-byte hole is still
there, but it is now available for a future subclass whose first member
is small enough to slot into it.
Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:
for (auto &[k, v] : o->entries()) { ... }
Side effects:
* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
(`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
original pattern relied on `nprop = 0` to short-circuit iteration on
a null or non-hash `properties`, the rewrite now guards the loop
explicitly with `if (o->type == JSON_HASH)` so that calling
entries() doesn't trip the asserting downcast.
Co-authored-by: Cursor <cursoragent@cursor.com>
The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:
json_object (base, TRUE / FALSE / NULL) 24 bytes
json_number 48 bytes
json_string 48 bytes
json_array (empty) 48 bytes
json_hash (empty) 72 bytes
Other size wins along the way:
* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
was 16 bytes per node). json_pull now keeps an explicit
container_stack and the parser no longer needs to resurrect a
shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
original std::make_shared<json_xxx> call, so destroying a
shared_ptr<json_object> still runs the right subclass dtor.
The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.
Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.
The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Read unidecode data, do some plumbing of it
* More unidecode plumbing
* Do the unidecode smashing, but it doesn't seem to be working
* Ah, that's better!
* Add missing header
* And reorder the includes too
* Shortcut when there is no unidecode data to work with
* Update version and changelog
* Avoid repeated unidecode smashing of the same constant string
* Add a way to run tippecanoe single-threaded for profiling
* Do less work when the tilestats sample values list is already full
* Save a copy when retrieving the attribute key
* Fewer atomic operations
* Move string hashing from mbtiles to text
* Only do approximate attribute deduplication when writing tiles
* Feature dropping tests are sensitive to exact tile size
* All tile creators now create a string pool for the tile
* Features clipped away to nothing should not participate in that tile
* Revert "Only do approximate attribute deduplication when writing tiles"
This reverts commit c42b34b498.
* Also revert the related test changes
* Revert "Revert "Only do approximate attribute deduplication when writing tiles""
This reverts commit 18509876c3.
* Be more specific about the string hash function
* Use fnv1a instead of std::hash for everything
* Reduce the chance of hash collisions
* Stick a hash search on the front of the tree search in addpool
* Eliminate repeated hashing of the same string
* Switch instead of ifs in json parsing
* A few more cases to populate the hash in addpool
* Store the hash in the tree instead of recalculating
* Add explanatory comment for mysterious argument
* Fewer copies in attribute stringification
* Clean up ancient weirdness in JSON attribute stringification
* More serial_val cleanup
* Pass a serial_feature to rewrite instead of many broken-down arguments
* Get rid of the multiple geometries within `partial`
* Revert "Pass a serial_feature to rewrite instead of many broken-down arguments"
This reverts commit 6f4ab9b725.
* Goodbye, struct coalesce
* Revert "Features clipped away to nothing should not participate in that tile"
This reverts commit 124462fbdc.
* Migrating fields from partial to serial_feature
* Name reconciliation between serial_feature and partial
* Replace struct partial with an augmented serial_feature
* Fix some overzealous search-and-replace renaming
* Don't say struct so often
* Remove more of the former partial construction
* Commenting and cleaning up
* Trying again to avoid all these arguments to rewrite
* I swear I did this same thing before and it didn't work.
* More rewrite cleanup
* Exile --detect-shared-borders to its own file
* Add missing headers
* More commenting and cleanup
* More comments
* Sprinkle consts around
* Emplacing and std::moving
* More cleanup
* That shouldn't have worked after a std::move
* Don't need to allocate memory to compare keys
* Reduce use of the global string pool in tiling
* Another avoidable mvt_value construction
* Further reduction to explicit string pool passing
* These reverses are no longer optimizations
* These layernames can all be references
* Don't drag an unused layername string around with every feature
* Heed a compiler warning about potential buffer overflow
* Fix my confusion about which feature's string pool is relevant
* Avoid some unnecessary allocations in attribute accumulation
* Maybe faster serialization?
* Eliminate a comparison
* Do the same here
* Save a couple of allocations when parsing numbers in JSON
* Immediately assign features to layers instead of subdividing later
* Maintain tilestats for tippecanoe:retain_points_multiplier_sequence
* Crunch out more duplicate attribute values when writing out the tile
* Do tilestats for tippecanoe:retain_points_multiplier_first too
* Shell filters need to be real threads, even if nothing else does
* Simplify tippecanoe_minzoom/maxzoom representation
* Update version and changelog
* Change JSON objects to a union type to use less memory
* Stop storing the string representation of JSON numbers
* Restore the ability to create features with large integer attributes
* Make sure large-integer feature IDs still behave as before
* Add missing #include
* Don't preallocate as much space for arrays and objects
* Treat inability to check free disk space as a warning, not an error
* Update changelog and version