Tippecanoe formatted every double it wrote through milo::dtoa_milo, a
vendored Grisu2. Grisu2 is fast, but it guarantees neither the shortest
digit string nor the correctly rounded one: it only guarantees that what
it prints parses back to the value it came from. In practice it prints a
digit more than necessary about 0.16% of the time, and picks a neighbor
of the correctly rounded digits about 32% of the time.
This ports Russ Cox's fpfmt (https://github.com/rsc/fpfmt) to C++ in
fpfmt/ and formats through it instead. fpfmt is both shortest and
correctly rounded, and it is faster:
full std::string formatting Grisu2 fpfmt speedup
random bit patterns 156.62 ns 66.83 ns 2.34x
geo coordinates 124.07 ns 58.62 ns 2.12x
short decimals 69.37 ns 49.16 ns 1.41x
small integers 44.18 ns 38.06 ns 1.16x
digit generation only Grisu2 fpfmt speedup
random bit patterns 90.07 ns 20.81 ns 4.33x
geo coordinates 80.64 ns 20.18 ns 4.00x
short decimals 55.61 ns 21.90 ns 2.54x
small integers 40.23 ns 22.50 ns 1.79x
(Intel Xeon @ 2.80GHz, g++ 13.3 -O3. `make fpfmt-bench` reproduces this,
and `./fpfmt-bench -check` reruns the correctness sweep, which is why
milo/dtoa_milo.h is kept even though nothing links it any more.)
The port is deliberately literal, so it can be diffed against fpfmt.go.
Its Short() agrees bit for bit with the Go original's on 445,640 values
covering powers of ten, small integers and reciprocals, subnormals, and
random bit patterns. Over 38.5 million values, fpfmt::dtoa always round
trips, is never longer than Grisu2's output, and is shorter 61,329 times.
Output is otherwise formatted exactly as before, including the choice
between plain and exponential notation, so 26 expected test outputs
change: some numbers lose digits (-26.170044999999999 becomes
-26.170045), and some have a corrected final digit (9.823748927348929e+55
becomes 9.823748927348928e+55). Every changed token was checked to parse
back to the identical double; none of the values themselves moved.
milo/milo.h, whose only job was to declare the C shim jsonpull calls, is
replaced by fpfmt/fpfmt.h, and the shim is renamed dtoa_shortest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014wJRAuhMninQE4wK2TUfuZ
* Fix variable-length-array and uninitialized-union compiler warnings
Clang warns about every variable-length array in C++ (-Wvla-cxx-extension,
on by default), since VLAs are a compiler extension rather than standard
C++. Replace all 57 of them with std::vector, or with std::string for the
mkstemp() template buffers built from tmpdir. Add -Wvla to WARNING_FLAGS so
new ones don't creep back in.
Separately, mvt_value's numeric_value union is 16 bytes wide (the size of
string_value), but both constructors only wrote the 8 bytes of the member
they were setting, leaving the rest indeterminate. The implicit copy
constructor copies the union as a whole, so copying any non-string value
read uninitialized bytes, which GCC reports as
mvt.hpp:83:8: warning: 'v.mvt_value::numeric_value. ... .len' may be
used uninitialized [-Wmaybe-uninitialized]
Give string_value, the widest member, a default member initializer so the
union's full width is initialized however it is later used.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
* Fix remaining float-conversion and format-truncation warnings
Clang's -Wimplicit-const-int-float-conversion flagged two comparisons
against LLONG_MAX, which is not representable as a double and rounds up
to 2^63.
In serial.cpp this was a real latent overflow, not just noise: the guard
`extent <= LLONG_MAX` was really `extent <= 2^63`, so an extent of exactly
2^63 passed it and then hit `(long long) extent`, which is undefined for
that value and yields LLONG_MIN in practice -- the opposite of the clamp
the else branch intends. Make the bound exclusive so the conversion is
always in range. Requires a polygon area at the very top of the double
range to reach, but the clamp now behaves as written.
In mbtiles.cpp the value is only a stand-in for infinity on its way into
JSON, so cast explicitly; the emitted number is unchanged.
Separately, g++ at -O0 warned that `char abbrev[20]` can be truncated by
"%lld", which is correct: the most negative long long needs 21 bytes with
the NUL. That branch is only reached when point_count < 1000, so it cannot
happen today, but size the buffer to fit rather than rely on that, and
replace the garbled comment about how the size was derived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
* Clamp the low end of extent before converting to long long too
The upper bound was fixed in the previous commit; the same overflow exists
on the negative side. get_area() returns a signed shoelace area, so inner
rings contribute negatively, and a polygon whose holes outweigh its rings
drives extent below zero. Far enough below and `(long long) extent` is
undefined again.
The bounds are asymmetric, so this is not simply the mirror of the upper
one: LLONG_MIN is exactly -2^63 and converts exactly, so unlike LLONG_MAX
it can be an inclusive bound.
Verified with -fsanitize=float-cast-overflow that the previous form traps
on 2^63 and on doubles just below -2^63, and that this one is clean across
both boundaries, the infinities, and NaN (which falls to LLONG_MAX, as it
did before).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
* Add CHANGELOG entries for 2.81.0 and bump the version
CHANGELOG.md was last updated for 2.80.0 (#361), and version.hpp has not
moved since. Twelve PRs have landed in the meantime with no entry: #365,
#368, #375, #382, #384, #385, #391, #395, #397, #399, #400, and #401.
Document all of them, plus this PR, under a single 2.81.0 heading. They are
not given separate version numbers because none of them was ever released
under one -- version.hpp read v2.80.0 throughout -- so assigning a version
per PR would invent release history. 2.81.0 is the version that will
actually carry them.
Where an unreleased PR was corrected by a later one (#384 by #385, #397 by
#399), the pair is described as the single behavior that ships, since the
intermediate behavior was never in a release.
Minor rather than patch bump: the batch adds command-line options.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
* Review feedback: enforce the union-width assumption, describe both clamp ends
The comment on mvt_value's union claimed string_value is the widest member.
That is true on LP64 (16 bytes against 8) but not on ILP32, where size_t is
4 and it ties with double and long long. The default member initializer still
covers the full union either way, so the fix held, but the justification did
not travel. Replace the claim with a static_assert that checks it on whatever
target is being built, so a platform where it stops holding is a compile
error rather than silently indeterminate bytes. Verified the assert is not
vacuous by widening the union in a scratch copy and watching it fail.
The changelog described only the upper end of the extent clamp. Describe both:
the old guard admitted everything below LLONG_MIN too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
* Add 2.81.0 changelog entries for the four PRs merged from main
#404, #408, #409, and #410 landed while this branch was open. None of them
bumped version.hpp, so they belong under the same 2.81.0 heading as the rest
of the unreleased work rather than getting versions of their own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8gsGMjK78TQiCGKTZ2PyR
---------
Co-authored-by: Claude <noreply@anthropic.com>
* Clip before dealing with multiplier or filters in overzoom
* Be more careful to retry when the feature count is exceeded
* Adjust the estimated total feature count for the multiplier too
* Fix the feature count estimates, I think
* Pass build info into the version string
* Report the actual max zoom of any tiles as the metadata maxzoom
* Revert unneeded renaming to make the diff more readable
* Clean up the adjustments to tile sizes and feature counts
* Update version and changelog
* Dropping a feature into a multiplier cluster still effectively drops it
* Update changelog
* Rethink the changelog description
* Don't try to truncate zooms if we are still tiling at z18
* Track output position at the file level instead of within each tile
* Track file position where the child tile data begins
* Add option and document its intended behavior
* Changing the detail loop to account for stopping early
* I forgot I already added an option for this
* Stop early if we can make a complete tile
* Add a test of zoom truncation with limited feature count
* Forgot to commit the actual code change
* Make room for a vertex count in the header of each serialized tile
* Estimate tile complexity; don't try truncating when unlikely to work
* Be more conservative, because ever retrying a tile is a big speed hit
* If stopping early, don't simplify or clean; leave that to overzoom
* Add tiny polygon reduction / dust to overzoom
* Don't try to stop early in the children if we dropped anything by rate
* Fflush here too before pwriting
* Don't stop early if we ended up dropping any features.
Rework the can-the-next-zoom-stop-early logic to avoid going
one zoom further than needed.
* Fix warning
* Fix warnings
* Oops, checking for the wrong expected return value
* Cleanup from adding line simplification in overzoom
* Current (wrong) behavior when combining coalescing and truncating
* Keep a list of parent tiles to skip rather than truncating
* Now the coalesced tiles in z12 get children in z13
* Don't double-count feature dropping when the zoom level is retried
* Correct README description
* Remove todo about special case below basezoom, which is accounted for
* Be a little more aggressive in drop-densest determination
* Scale tile feature limit for megatiles in the same way as byte limit
* Fully deprecate -detect-shared-borders into an alias
* Track the distances found in the douglas-peucker recursion
* Serialize and deserialize the distance with the vertices
* Revert "Serialize and deserialize the distance with the vertices"
This reverts commit 753f1b7909.
* Revert "Track the distances found in the douglas-peucker recursion"
This reverts commit e5361f8c22.
* Revert "Fully deprecate -detect-shared-borders into an alias"
This reverts commit 0698aeb766.
* Better tracking of whether we failed to make a full-detail tile
* Put a bloom filter in front of the binary search for shared nodes
* Forgot to take out this printf
* Improve dispatch of tiling tasks
* Still dispatch the biggest tasks first
* Track zoom truncation in the strategies list in the tileset metadata
* Prescan for small deltas before doing proper simplification
* Revert "Prescan for small deltas before doing proper simplification"
This reverts commit d1d8238b83.
* Update version and changelog
* Rename to --generate-variable-depth-tile-pyramid
* Add a way to run tippecanoe single-threaded for profiling
* Do less work when the tilestats sample values list is already full
* Save a copy when retrieving the attribute key
* Fewer atomic operations
* Move string hashing from mbtiles to text
* Only do approximate attribute deduplication when writing tiles
* Feature dropping tests are sensitive to exact tile size
* All tile creators now create a string pool for the tile
* Features clipped away to nothing should not participate in that tile
* Revert "Only do approximate attribute deduplication when writing tiles"
This reverts commit c42b34b498.
* Also revert the related test changes
* Revert "Revert "Only do approximate attribute deduplication when writing tiles""
This reverts commit 18509876c3.
* Be more specific about the string hash function
* Use fnv1a instead of std::hash for everything
* Reduce the chance of hash collisions
* Stick a hash search on the front of the tree search in addpool
* Eliminate repeated hashing of the same string
* Switch instead of ifs in json parsing
* A few more cases to populate the hash in addpool
* Store the hash in the tree instead of recalculating
* Add explanatory comment for mysterious argument
* Fewer copies in attribute stringification
* Clean up ancient weirdness in JSON attribute stringification
* More serial_val cleanup
* Pass a serial_feature to rewrite instead of many broken-down arguments
* Get rid of the multiple geometries within `partial`
* Revert "Pass a serial_feature to rewrite instead of many broken-down arguments"
This reverts commit 6f4ab9b725.
* Goodbye, struct coalesce
* Revert "Features clipped away to nothing should not participate in that tile"
This reverts commit 124462fbdc.
* Migrating fields from partial to serial_feature
* Name reconciliation between serial_feature and partial
* Replace struct partial with an augmented serial_feature
* Fix some overzealous search-and-replace renaming
* Don't say struct so often
* Remove more of the former partial construction
* Commenting and cleaning up
* Trying again to avoid all these arguments to rewrite
* I swear I did this same thing before and it didn't work.
* More rewrite cleanup
* Exile --detect-shared-borders to its own file
* Add missing headers
* More commenting and cleanup
* More comments
* Sprinkle consts around
* Emplacing and std::moving
* More cleanup
* That shouldn't have worked after a std::move
* Don't need to allocate memory to compare keys
* Reduce use of the global string pool in tiling
* Another avoidable mvt_value construction
* Further reduction to explicit string pool passing
* These reverses are no longer optimizations
* These layernames can all be references
* Don't drag an unused layername string around with every feature
* Heed a compiler warning about potential buffer overflow
* Fix my confusion about which feature's string pool is relevant
* Avoid some unnecessary allocations in attribute accumulation
* Maybe faster serialization?
* Eliminate a comparison
* Do the same here
* Save a couple of allocations when parsing numbers in JSON
* Immediately assign features to layers instead of subdividing later
* Maintain tilestats for tippecanoe:retain_points_multiplier_sequence
* Crunch out more duplicate attribute values when writing out the tile
* Do tilestats for tippecanoe:retain_points_multiplier_first too
* Shell filters need to be real threads, even if nothing else does
* Simplify tippecanoe_minzoom/maxzoom representation
* Update version and changelog
* Get rid of the type_and_string near-synonym for serial_val
* Rename file_keys to the more familiar tilestats
* "tas" (type_and_string) => "sv" (serial_val)
* "fk" (tile_keys) => "ts" (tilestats)
* Revert ""tas" (type_and_string) => "sv" (serial_val)"
This reverts commit 4854c57e22.
* More carefully this time: "tas" (type_and_string) => "sv" (serial_val)
* Add an option to retain N times as many points as usual at each zoom
* Tests for point multipler with specified and guessed maxzooms
* Work in progress on inverse spatial ordering
* Fix inverse spatial feature order
* --reorder was depending on a feature index that wasn't being preserved
* Separate ordering by feature_minzoom from ordering inverse-spatially
* Add a test for the inverse spatial ordering
* Store the basezoom/droprate/multiplier decisions in tileset metadata
* Progress on adding filters to tippecanoe-overzoom
* Type promotion for comparison
* Look up the attribute value for ordering
* Add test of thinning and ordering features
* Plumb tippecanoe_decisions metadata through pmtiles
* Be careful not to put infinities in JSON
* Fix accidental dropping in what is meant to preserve sparse points
* Start distinguishing true, false, and null in expressions
* Most of the type conversions
* Add boolean conversions
* Literals and conjunctions
* Add filtering to tippecanoe-overzoom
* Add a test of filtering in overzoom
* Fix boolean conjunctions
* Handle the combination of cluster size and filtering
* Rework dot dropping to reconcile density threshold and multiplier
* Revert "Rework dot dropping to reconcile density threshold and multiplier"
This reverts commit f253a66382.
* Retain points by multiplier within each tile, not in global probability
* Test that intends to verify that the multiplier is reversible
* Get the test to detect the discrepancy
* Mark the start of multiplier clusters with a magic attribute
* Add string-contains
* Add in and ni operators
* Revert "Look up the attribute value for ordering"
This reverts commit 56bc73e49a.
* Revert "Type promotion for comparison"
This reverts commit 6f3256f5af.
* Make number formatting in tippecanoe_decisions consistent
* Revert "Add a test for the inverse spatial ordering"
This reverts commit c8047de9ab.
* Revert "Separate ordering by feature_minzoom from ordering inverse-spatially"
This reverts commit 35b19a223c.
* Revert "Fix inverse spatial feature order"
This reverts commit 5978ecdb44.
* Revert "Work in progress on inverse spatial ordering"
This reverts commit fdf230f632.
* Somehow missed the tests associated with that last revert
* Round-robin assign attributes to partials from across the multiplier
* Count the multiplier separately in each layer
* Fix distribution of accumulated attribute across multiplier features
* Update changelog, version, and docs
* Add "is null" and "isnt null" expressions
* Update interpretation of FSL expressions to pass the tests
* Test to assert that polygons are unaffected by the multiplier
* Clean up and comment
* Remove accidental unused case
* Calculate a new antimeridian-adjusted bounding box
* Add antimeridian bounding box to pmtiles, dirtiles, and tile-join
* Don't take out-of-bounds latitudes into account in the adjusted bbox
* Update changelog
* Forgot to adjust tests after the last change
* add pmtiles.hpp from github.com/protomaps/PMTiles [#10]
* tippecanoe main writes pmtiles output. [#10]
* detect output format using suffix
* after mbtiles is done writing, replace with pmtiles based on map/image tables.
* add method to write_json for writing json sub-object.
* tippecanoe-decode reads pmtiles input. [#10]
* tile-join reads and writes pmtiles. [#10]
* pmtiles test suite for decode and tile-join [#10]
* add base GitHub CI action for compiling and test suite.
* update pmtiles.hpp with z>15 fix
* Fix some ordering problems with pmtiles decode
* Pmtiles should also pass the raw tiles tests
* Eradicate spaces from tileset metadata JSON fields
* Eradicate spaces from more test fixtures
* Update more tests
* Pmtiles tests pass now too
* Remove unnecessary sort (and make indent)
* Update changelog
* The allow-existing test for pmtiles needs -o, not -e
* Declare --allow-existing to be unsupported for pmtiles.
It was always a bad idea even for mbtiles.
Co-authored-by: Brandon Liu <bdon@bdon.org>
* Progress toward making a tileset metadata structure
* Write metadata from structure to mbtiles
* Write dirtiles metadata.json from metadata structure
* Update changelog
Change mbtiles tile_id inserts from blob to text to enable querying by text.
* previous blob type worked for internal join, but made querying by text impossible without casting.
* Working on eliminating preflighting
* Adjust for the map/images schema change
* Avoid generating duplicate tiles with the detail reduction strategy
* Do error checking if tiles in a directory can't be written
* Update changelog and version
* Revert unintentional code reordering
* Neglected to add the exit on error here
* Add an option to limit geometry vertex count
* Update docs and changelog
* Track desired feature count and geometry size in strategies
* Remove dead code for a long-forgotten inaccessible option
* Fix the option name in changelog
* Refine coalesce-smallest to only coalesce onto other small features
* Clean polygons before coalescing-as-needed
* Don't accumulate tiny polygon holes as negative dust
* Add the option not to limit the feature count at maxzoom
* Add options to limit feature count more abruptly in each tile
* Revert "Add the option not to limit the feature count at maxzoom"
This reverts commit ace173ab7a.
* Revert "Fix the option name in changelog"
This reverts commit 38dbce8405.
* Revert "Add an option to limit geometry vertex count"
This reverts commit d379fdf06b.
* Remove test for reverted option
* Remove more leftovers from geometry size limiting
* Update changelog
* Change sqlite3 schema to deduplicate identical tiles
* Limit guessed maxzoom to avoid spending too many tiles on polygon fill
* Fix test.
These dust polygons now have their area calculated because their
maxzoom is being guessed, so the attributes from the largest one
rather than the last one are preserved.
* Increase polygon limit to a million tiles
* Two million tiles ought to be enough for anyone, right?
* Add explanatory comments for mysterious numbers
* Stop adding features to a tile if it can't possibly work
* Add --integer and --fraction options to tippecanoe-decode
* Carry the strategies field from tileset metadata through tile-join
* Update changelog
* Assign different codes to different kinds of error exits
* Add an option to retain extra coordinate precision at maxzoom
* Make sure not to shift away the extra detail from coordinates
* Add an option to convert double-precision attributes to single
* Sort attribute values in tiles to make them compress a little better
* Slightly improve polygon simplification
By choosing a point that would be retained after simplification
to be the start/end point that always gets retained
* I regret making all of these tests involve polygons
* Add an option to specify the size of tiny polygons
* Fix accidental requiring of argument for --single-precision
* Guard against duplicate points when generating "sizes" for them
* Restore the intended behavior that tiny polygons don't get simplified
* Make the extra detail settable rather than always maximizing it
* Revert "Improve maxzoom guessing for tightly-clustered point data sources (#4)"
This reverts commit fec5e8354c.
* Add an option to prevent choosing a base zoom higher than the maxzoom
* Keep the drop rate high enough when the basezoom gets constrained
* Revert "Revert "Improve maxzoom guessing for tightly-clustered point data sources (#4)""
This reverts commit db6bc27d9e.
* Add --order-by and --order-descending options
* Accept multiple --order-by and --order-descending-by sort keys
* Handle points too when dropping or coalescing the "smallest" features.
* Add statistics of tile size reduction strategies to tileset metadata
* Update changelog and version
* Update documentation
An instance of catching an exception by value:
https://blog.knatten.org/2010/04/02/always-catch-exceptions-by-reference/
In jsontool, there's a warning about writing 8 bytes into a 7
byte buffer, potentially truncating or losing the nul-terminator.
There are a couple of instances of an unused bool variable, not
really a big deal.
This moves filtering from the serialization stage to the
tiling stage so that the zoom level can be known to the filter.
The side effect is to carry null attributes much further through
the pipeline than previously.
(The feature count when filtering will be the sum of features
across tiles instead of filters from the original input, since
the filter reader doesn't know what the original input feature
set was.)