Tippecanoe formatted every double it wrote through milo::dtoa_milo, a
vendored Grisu2. Grisu2 is fast, but it guarantees neither the shortest
digit string nor the correctly rounded one: it only guarantees that what
it prints parses back to the value it came from. In practice it prints a
digit more than necessary about 0.16% of the time, and picks a neighbor
of the correctly rounded digits about 32% of the time.
This ports Russ Cox's fpfmt (https://github.com/rsc/fpfmt) to C++ in
fpfmt/ and formats through it instead. fpfmt is both shortest and
correctly rounded, and it is faster:
full std::string formatting Grisu2 fpfmt speedup
random bit patterns 156.62 ns 66.83 ns 2.34x
geo coordinates 124.07 ns 58.62 ns 2.12x
short decimals 69.37 ns 49.16 ns 1.41x
small integers 44.18 ns 38.06 ns 1.16x
digit generation only Grisu2 fpfmt speedup
random bit patterns 90.07 ns 20.81 ns 4.33x
geo coordinates 80.64 ns 20.18 ns 4.00x
short decimals 55.61 ns 21.90 ns 2.54x
small integers 40.23 ns 22.50 ns 1.79x
(Intel Xeon @ 2.80GHz, g++ 13.3 -O3. `make fpfmt-bench` reproduces this,
and `./fpfmt-bench -check` reruns the correctness sweep, which is why
milo/dtoa_milo.h is kept even though nothing links it any more.)
The port is deliberately literal, so it can be diffed against fpfmt.go.
Its Short() agrees bit for bit with the Go original's on 445,640 values
covering powers of ten, small integers and reciprocals, subnormals, and
random bit patterns. Over 38.5 million values, fpfmt::dtoa always round
trips, is never longer than Grisu2's output, and is shorter 61,329 times.
Output is otherwise formatted exactly as before, including the choice
between plain and exponential notation, so 26 expected test outputs
change: some numbers lose digits (-26.170044999999999 becomes
-26.170045), and some have a corrected final digit (9.823748927348929e+55
becomes 9.823748927348928e+55). Every changed token was checked to parse
back to the identical double; none of the values themselves moved.
milo/milo.h, whose only job was to declare the C shim jsonpull calls, is
replaced by fpfmt/fpfmt.h, and the shim is renamed dtoa_shortest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014wJRAuhMninQE4wK2TUfuZ
* Plumb bounding boxes through potential intersections
* Quick bbox reject for bins that can't possibly intersect
* Inching toward attribute accumulation in megatile handling
* Some sort of test for how all these things interact with each other.
Automatic numeric attribute accumulation does *not* apply to attributes
that have an explicit attribute accumulator set, because the order of
operations is too messy and weird
* More sketching
* More sketching
* Actually do some accumulation
* Put all that behind an --accumulate-numeric flag
* Use the same attribute accumulation logic in binning as in megatiles
* Fix backwards conditional
* Add means, but somehow I have some counts of 0
* Handle aggregated attributes with no base attribute in the feature
* Checkpoint before I break everything
* Found a flaw, now to debug
* Fix a typo that broke accumulation
* Add binning tests
* Make sure IDs make it through on the bins
* Fix count/mean accumulation
* Make the numeric accumulation prefix configurable
* Make sure the accumulate test still works with a different prefix
* Forgot to update this test
* More testing to make sure cluster sizes make it all the way through
* Fix neglected --accumulate-attribute when binning
* Mark unexercised attribute accumulation cases as "can't happen"
* Factor out numeric preservation
* Attrs with the accumulation prefix are just preserved, not accumulated
* Test behavior of prefixed attributes
* Plumbing for exclude and exclude-prefix
* Implement and test attribute prefix stripping in overzoom
* Update version and changelog
* For debugging, make an attribute list of source feature IDs
* Revert "For debugging, make an attribute list of source feature IDs"
This reverts commit 65fc99c9d1.