Commit Graph
26 Commits
Author SHA1 Message Date
Erica Fischer 03a63976b4 Merge remote-tracking branch 'origin/main' into jsonpull-cpp 2026-08-12 09:58:14 -07:00
Erica FischerandClaude Opus 5 e6e1ec3263 Generate the usage message of each tool from its long_options (#409)
* Generate the usage message of each tool from its long_options

The usage messages of tile-join, tippecanoe-overzoom,
tippecanoe-json-tool, tippecanoe-decode, and tippecanoe-enumerate were
hand-written lists of options that had drifted years out of date, since
nothing tied them to the options that are really accepted. Move the
option-list printing that tippecanoe already does into a shared
print_usage(), and use it in all the tools, so that the message is
derived from the same long_options table that getopt_long() gets and
can't fall behind it again.

The tables now carry section headings, as tippecanoe's does, and the
options that were only reachable by their short names (tile-join's -O,
-b, -R, and -r among them) are listed for the first time.

Also state the non-option arguments the way each tool really treats
them: tile-join takes source tilesets unless --read-from names a file to
read them from, tippecanoe-decode takes a tileset either alone or with a
zoom/x/y, tippecanoe-json-tool reads standard input when no files are
named, and tippecanoe-overzoom's two forms are the ones its argument
parsing recognizes. tippecanoe-overzoom now reports the missing -o
instead of passing NULL to fopen(), and tippecanoe-enumerate goes
through getopt_long() so that it will pick up any options added later.

The shared getopt_string() replaces the identical loop that four of the
tools each had for building the short option string, and strip_usage_headings()
the one for dropping the headings before getopt_long() sees them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

* Print the usage message when tippecanoe is run with no arguments

Running `tippecanoe` with nothing at all reported the missing output
file, which is true but is not what someone who typed the bare command
needs to know. Check for the empty command line before parsing and print
the general usage message instead, and leave the specific complaint for
the case where an input file was named but an output file wasn't.

To make the message reachable from there, the options table and the
usage printing move out of main() into a usage() function, as in the
other tools.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

* Address review: alternation, the dead tile-join option, and --version

Four fixes from review of the generated usage messages:

* `--output` and `--output-to-directory` are one-of, not one required and
  one optional, in both tippecanoe and tile-join. A `usage_required_option`
  can now name an alternation that it belongs to, and the options in one
  are listed together as `(--output=... | --output-to-directory=...)`,
  which is what the runtime check enforces.

* tile-join's `--use-attribute-for-id` has had no implementation since
  533e000 removed it; only the table entry was left behind, so the option
  parsed and then exited with "Unrecognized option". Generating the usage
  message from the table turned that into a documented option that doesn't
  work, so remove the leftover entry too.

* `--version` was grouped under "Progress indicator", in the options table
  and in the README both. Give it a heading of its own now that the
  headings are something users see.

* print_usage() left `width` holding the length of the last synopsis line,
  and only got away with it because every table so far begins with a
  heading, which resets it. Start the option list on a line of its own
  instead of depending on that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016frkRY1xXtiWjxYuCJ8vZY

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 16:19:03 -07:00
Erica FischerandCursor 65beb3f0d7 Migrate jsonpull to unique_ptr ownership
Replaces the shared_ptr-based json_object_ptr with a unique_ptr that
has a stateless custom deleter dispatching on json_object::type before
calling the right subclass destructor. Eliminates per-node atomic
reference-counting and the control-block allocation that shared_ptr
required for every node in the tree.

API now distinguishes owning and borrowing pointers explicitly:
- json_read / json_read_separators / json_hash_get return raw
  json_object * (borrowed from the parser-owned tree).
- json_read_tree / json_disconnect return json_object_ptr (caller
  takes ownership; back-pointers are cleared so the subtree can
  outlive the parser).
- json_free / json_context / json_stringify take raw pointers.
- The parser's container_stack holds raw pointers; jp->root keeps
  unique_ptr ownership of the most recent top-level value.

Internally, take_from_owner moves the unique_ptr out of whichever
parent vector / hash entry / parser root owned it, which both
json_free and json_disconnect rely on.

In the streaming parsers (parse_feature, parse_layers, the
geojson-loop callback), we are careful to free `j` only after we
have processed a complete Feature: json_read returns each token
as the tree is being built up, and freeing an intermediate node
would splice it out of the surrounding hash and corrupt the
in-progress feature.

Benchmark (tl_2022_us_county.json, -z0 --extend-zooms-if-still-dropping,
median of 5 runs on macOS arm64): 8.5s, vs 10.6s with shared_ptr
and 8.7s on the pre-refactor C baseline.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 20:44:25 -07:00
Erica FischerandCursor 05b1762077 Fix bugs flagged in code review of jsonpull C++ port
- jsontool.cpp `out()`: route JSON_NUMBER (and anything else non-string)
  through `json_stringify` instead of `o->string()`, which now asserts
  on a non-string type and would crash `--extract` on numeric attributes.
- geojson.{hpp,cpp} `json_end_map`: take `json_pull_ptr` by reference so
  the caller's shared_ptr is released, null-guard before touching
  `jp->source`, and clear `jp->source` after delete to avoid a dangling
  pointer.
- jsonpull/jsonpull.cpp: low-surrogate range check was comparing the
  outer-loop byte `c` instead of the parsed code unit `ch`, breaking
  surrogate-pair decoding for some \\uXXXX escapes. Pre-existing bug
  preserved across the port.
- tile-join.cpp `handle_vector_layers`: require the field value to have
  type JSON_STRING (and the key to be non-null) before calling
  `string()`; the previous truthy `type` check would assert on a
  non-string value.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:28 -07:00
Erica FischerandCursor 1bf18d39ce Discriminate json_number's three numeric slots into one union
json_number used to carry three parallel 8-byte fields (a double plus
both a 64-bit unsigned and a 64-bit signed slot for the large-integer
cases) even though at most one of the integer slots is ever the
canonical value for any given number. Collapse them into a
discriminated union:

    enum repr_t { REPR_DOUBLE, REPR_LARGE_UNSIGNED, REPR_LARGE_SIGNED };
    repr_t repr;
    union { double d; unsigned long long u; long long s; } value;

Callers keep the same read API: number() returns the appropriate
double, large_unsigned() returns the ull (or 0 if not currently stored
that way), large_signed() likewise. Writes go through new set_number /
set_large_unsigned / set_large_signed methods that keep the
discriminator and the union value in sync.

This was prompted by an observation that moving json_type to the end
of the object should shrink things via tail-padding reuse. Empirically
the type-at-end rearrangement saves nothing on its own (every
subclass payload is 8-byte aligned so it can't slot into the 4-byte
tail), but the discriminated-number redesign hits the same idea from
a different direction: adding the 4-byte `repr` to json_number makes
the class non-standard-layout, which lets the Itanium ABI pack `repr`
into the base's 4-byte tail padding at offset 20. The union value
then starts at the natural offset 24, and json_number ends at offset
32 -- a 33% reduction.

Per-node sizes:
  json_object (TRUE/FALSE/NULL)  24 bytes
  json_number                    32 bytes  (was 48)
  json_string                    48 bytes
  json_array                     48 bytes
  json_hash                      48 bytes

Numbers dominate real GeoJSON (every coordinate is one), so the net
memory win on a typical parse is substantial.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:24 -07:00
Erica FischerandCursor 4d9a48c3d4 Store hash key/value pairs in one ordered vector
Replace the parallel std::vector<json_object_ptr> keys / values on
json_hash with a single std::vector<json_entry>, where json_entry is
a small {key, value} aggregate. This still preserves insertion order
(the property the parallel vectors were providing) but removes the
"keep two vectors in lockstep" pattern, and call sites can now use
range-for with structured bindings:

    for (auto &[k, v] : o->entries()) { ... }

Side effects:

* sizeof(json_hash) drops from 72 to 48 bytes (one fewer vector
  header), matching json_array.
* The keys() and values() accessors on json_object are replaced by a
  single entries() accessor returning std::vector<json_entry>&.
* All call sites were swept from the old paired-index pattern
  (`o->keys()[i]` / `o->values()[i]`) to entry-based access. Where the
  original pattern relied on `nprop = 0` to short-circuit iteration on
  a null or non-hash `properties`, the rewrite now guards the loop
  explicitly with `if (o->type == JSON_HASH)` so that calling
  entries() doesn't trip the asserting downcast.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:13 -07:00
Erica FischerandCursor f366b2c4aa Subclass json_object so primitives shrink from 168 to 24 bytes
The previous "every member in a struct" layout cost 168 bytes per
json_object, even for JSON_NULL / JSON_TRUE / JSON_FALSE nodes that
have no payload. Splitting json_object into a small base class plus
json_number / json_string / json_array / json_hash subclasses brings
each instance down to just the size of its actual contents:

  json_object (base, TRUE / FALSE / NULL)   24 bytes
  json_number                                48 bytes
  json_string                                48 bytes
  json_array  (empty)                        48 bytes
  json_hash   (empty)                        72 bytes

Other size wins along the way:

* Drop enable_shared_from_this<json_object> (its embedded weak_ptr
  was 16 bytes per node). json_pull now keeps an explicit
  container_stack and the parser no longer needs to resurrect a
  shared_ptr from a raw `parent` walk.
* Remove the unused `refcon` slot from the string variant.
* No virtual destructor: shared_ptr keeps the deleter from the
  original std::make_shared<json_xxx> call, so destroying a
  shared_ptr<json_object> still runs the right subclass dtor.

The base class exposes type-tagged accessors (o->string(),
o->number(), o->array(), o->keys(), o->values(), o->large_signed(),
o->large_unsigned()) that assert the type matches and downcast to
the appropriate subclass storage. All call sites were swept from
the old `o->value.X.Y` field paths to these accessors. A raw-pointer
overload of json_hash_get() replaces the few external uses of
shared_from_this() that survived in geojson-loop.cpp.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:49:08 -07:00
Erica FischerandCursor 3da03c6075 Convert jsonpull to C++ with shared_ptr and std::vector/std::string
Replace the manual malloc/realloc/free memory management in jsonpull
with std::shared_ptr ownership. Each json_object now owns its children
through std::vector<json_object_ptr>; raw back-pointers to parent and
parser remain valid by structural invariant and are cleared on
json_disconnect so detached subtrees can outlive their parser.
Strings become std::string, child arrays become std::vector, and the
old union becomes a struct so non-trivial members can coexist while
preserving the existing o->value.xxx access paths.

The old jsonpull.c is replaced by jsonpull.cpp, json_stringify now
returns std::string, and all callers across tippecanoe, tile-join,
tippecanoe-decode, tippecanoe-json-tool, tippecanoe-overzoom and the
unit tests are updated to use json_object_ptr / json_pull_ptr.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-30 17:48:32 -07:00
Erica Fischer c1051232ad Do more of line simplification with integer coordinates (#113)
* Turn off unit, which has bit-rotted

* Instrument line simplification

* Do more work in integer space

* Update tests with simplification done with more integers

* Clean up

* Update changelog

* Eradicate calls to deprecated sprintf()

* Revert "Turn off unit, which has bit-rotted"

This reverts commit 998779313f.

* Fix catch by switching to an unsigned type
2023-07-10 15:54:21 -07:00
Erica Fischer af1a7ed7ae Once features can't possibly fit in a tile, stop trying (#9)
* Stop adding features to a tile if it can't possibly work

* Add --integer and --fraction options to tippecanoe-decode

* Carry the strategies field from tileset metadata through tile-join

* Update changelog

* Assign different codes to different kinds of error exits
2022-09-23 20:01:13 -04:00
Erica Fischer 67cd9d8d85 Reduce tippecanoe memory usage (#5)
* Change JSON objects to a union type to use less memory

* Stop storing the string representation of JSON numbers

* Restore the ability to create features with large integer attributes

* Make sure large-integer feature IDs still behave as before

* Add missing #include

* Don't preallocate as much space for arrays and objects

* Treat inability to check free disk space as a warning, not an error

* Update changelog and version
2022-08-09 15:18:51 -07:00
Monty Taylor efa40d20ab Clean a few warnings emited by GCC v8
An instance of catching an exception by value:

  https://blog.knatten.org/2010/04/02/always-catch-exceptions-by-reference/

In jsontool, there's a warning about writing 8 bytes into a 7
byte buffer, potentially truncating or losing the nul-terminator.

There are a couple of instances of an unused bool variable, not
really a big deal.
2020-05-17 09:43:24 -05:00
Eric Fischer 19c132e79c Be more cautious about use of a null feature for bare geometries 2018-08-08 14:17:08 -07:00
Eric Fischer 98cf4d94aa Don't accept features or geometries inside another object's properties 2018-08-08 13:42:49 -07:00
Eric Fischer ba8966a4ae Use the same GeoJSON parsing loop in tippecanoe-json-tool 2018-08-08 13:11:25 -07:00
Eric Fischer 87a1bb7851 Add an option to treat empty CSV columns as nulls, not empty strings 2018-07-19 14:33:33 -07:00
Eric Fischer 54532795f6 Trailing commas in CSVs are now treated as empty fields.
Empty fields are now treated as empty strings rather than nulls
in tippecanoe-json-tool, for consistency with tile-join.
2018-05-24 13:54:00 -07:00
Eric Fischer afb5cece96 Verify that CSV input is encoded as UTF-8 2017-12-06 13:32:44 -08:00
Eric Fischer 8ac7c46788 Make the same null pointer fix in jsontool.cpp as in geojson.cpp 2017-11-22 13:06:39 -08:00
Eric Fischer 2b1cba0b53 Warn during json-tool extraction if the extracted field isn't found 2017-11-17 13:52:45 -08:00
Eric Fischer 9ebeb47d24 Don't duplicate the join key in JSON tool output 2017-10-10 17:51:16 -07:00
Eric Fischer 68a55b8749 Follow JSON rules for what looks like a number in a CSV 2017-10-10 16:22:47 -07:00
Eric Fischer 86a4ce67a6 Joining basically works 2017-10-10 16:12:40 -07:00
Eric Fischer 19117d8060 Move CSV code into its own file 2017-10-10 14:57:38 -07:00
Eric Fischer ebb26ee14c Add property extraction for sorting 2017-10-10 14:03:24 -07:00
Eric Fischer d9c22135e5 Rename geojson2nd to tippecanoe-json-tool 2017-10-10 11:37:30 -07:00