Commit Graph
49 Commits
Author SHA1 Message Date
Claude d3a98170bf Look up shared nodes once per vertex instead of once per tile
With --no-simplification-of-shared-nodes, simplify_lines() used to offset
every vertex of every feature to world coordinates and check it against
the Bloom filter and the global sorted list of shared nodes, in every
tile at every zoom level.

Whether a vertex is a shared node depends only on its world coordinates,
so now it is found once, after the list of shared nodes has been made
and before the geometry is sorted, by a parallel pass over each reader's
geometry that marks each vertex in place, in the upper bits of its
serialized operation byte. Decoding puts that state into a new field of
draw (which still fits in 16 bytes), and it is carried through clipping,
the copies across the antimeridian at z0, and the geometry written for
the next zoom level. Polygon cleaning of coalesced features restores the
state of any output vertex with the same coordinates as an input vertex.

Vertices whose state is still unknown, because they were created by
clipping or polygon cleaning or came back from a prefilter, are still
looked up in the global list when they are simplified, so the output is
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P2nBqZisxNQfmEmon3vE9v
2026-09-23 17:39:48 +00:00
Erica Fischer 423fea1544 Reduce memory consumption during tiling and tile encoding (#319)
* Improve the memory spike during tile construction

* Remove `need_tilestats`. Add a bunch of debug logging

* Remove debug logging

* Update version and changelog
2025-01-31 12:57:41 -08:00
Erica Fischer 583fc3744a Reduce attribute accumulation memory consumption (#318)
* Add a flag to use an H3 index for the feature index

* Give mvt_value and serial_val a double-with-count concept

* Switch mean over to internal accumulation state

* Get rid of the attribute accumulation map

* Change vectors of features to vectors of pointers to features

* Fix --coalesce

* Revert "Add a flag to use an H3 index for the feature index"

This reverts commit b9b48f42c9.

* Update version and changelog
2025-01-30 21:58:35 -08:00
Erica Fischer 23667bb8eb Reduce memory consumption from attribute accumulation (#290)
* Progress on plumbing a string pool for full_keys through

* More plumbing for key_pool

* Don't keep features with identical locations as multiplier features

* Revert "Don't keep features with identical locations as multiplier features"

This reverts commit 413f0c8024.

* Adjust calculated maxzoom to account for duplicate feature locations

* Update changelog and version

* Add a test affected by the maxzoom change with duplicate locations

* Round the drop rate a little for cross-platform test consistency
2024-11-05 14:39:41 -08:00
Erica Fischer b3b89e1e07 Performance optimizations for binning (#283)
* Add an all-mvt_value attribute accumulation path

* Only bin by ID, not geometrically

* A little cleanup; changelog and version; test

* Remove accidental double-conversion

* Replace duplicated code with template

* Update changelog
2024-11-01 15:27:13 -07:00
Erica Fischer 28efc40e6e Choose the megatile features from those that will be in the next N zooms (#280)
* Choose the megatile features from those that will be in the next N zooms

* Take fractional zooms into account in multiplier feature choices

* Fix more tests

* Add a flag to retain multiplier features by minimum distance

* Limit feature expansion from multiplier density to 2x

* The multiplier cap was a bad idea

* Revert "The multiplier cap was a bad idea"

This reverts commit 6f8273a4c8.

* Revert "Limit feature expansion from multiplier density to 2x"

This reverts commit a26e41309d.

* Revert "Add a flag to retain multiplier features by minimum distance"

This reverts commit 01f14a4255.

* Remove the multiplier sequence, which should no longer matter

* Revert "Revert "Add a flag to retain multiplier features by minimum distance""

This reverts commit 776da4a1b8.

* Revert "Revert "Limit feature expansion from multiplier density to 2x""

This reverts commit 44a683d808.

* Revert "Revert "The multiplier cap was a bad idea""

This reverts commit 80f7cb1c0e.

* Track two kinds of previous index for next_feature

* Fix multiplier density threshold, I think

* Oh, I didn't git add the code changes

* Update version and changelog

* Try to install sqlite3 to fix the automated build

* Deleted too much

* Only let --preserve-point-density-threshold shift density around

* Remove the density debt concept, since it doesn't help

* Make the drop states a vector instead of an array

* Revert "Make the drop states a vector instead of an array"

This reverts commit 66c7abb6fa.

* Revert "Remove the density debt concept, since it doesn't help"

This reverts commit 707bb0c562.

* Revert "Only let --preserve-point-density-threshold shift density around"

This reverts commit ecf01f2231.
2024-10-17 15:59:44 -07:00
Erica Fischer 47e774adc4 Improve the appearance of coalesce-densest-as-needed tiles (#247)
* Start to distinguish fixed cluster density setting from as-needed density

* Make consistent {drop,coalesce}-densest decisions between zooms

* Actually track the previous index instead of just intending to

* Clean up collinearities in coalesced features

* To determine densest, look at actual physical distance, not just index

* Don't actually need the previous index in serial_feature now

* Center of mass of one feature to most distant point of the next

* Add apologetic comment

* Wait, how did the tests pass before?

* Revert "Wait, how did the tests pass before?"

This reverts commit f73c8ee543.

* Add --maximum-string-attribute-length option

* Update version and changelog

* A little more testing to make sure
2024-07-16 13:23:23 -07:00
Erica Fischer 4e52cbd957 Drop or retain whole multiplier clusters when dropping as needed (#198)
* Prep to track conditions other than just "dropped" or "kept"

* Count up instead of down

* Drop or retain whole multiplier clusters based on their first feature

* Calculate a global feature dropping sequence

* Switch over to using the drop sequence for drop-fraction

* Remove unused arguments for the old drop-fraction implementation

* Fix copy-and-paste bugs, update tests

* Properly incorporate feature_minzoom into the drop sequence, I hope

* Rename drop_by to drop_sequence

* See if sorting within clusters fixes filter stability between zooms

* Remove very chatty debug print

* Update changelog and version

* Use named constants instead of numbers for feature dropping/keeping

* Add comment to explain purpose and method of bit reversal
2024-02-12 10:58:49 -08:00
Erica Fischer e2a7a409c7 Improve tiling speed (#195)
* Add a way to run tippecanoe single-threaded for profiling

* Do less work when the tilestats sample values list is already full

* Save a copy when retrieving the attribute key

* Fewer atomic operations

* Move string hashing from mbtiles to text

* Only do approximate attribute deduplication when writing tiles

* Feature dropping tests are sensitive to exact tile size

* All tile creators now create a string pool for the tile

* Features clipped away to nothing should not participate in that tile

* Revert "Only do approximate attribute deduplication when writing tiles"

This reverts commit c42b34b498.

* Also revert the related test changes

* Revert "Revert "Only do approximate attribute deduplication when writing tiles""

This reverts commit 18509876c3.

* Be more specific about the string hash function

* Use fnv1a instead of std::hash for everything

* Reduce the chance of hash collisions

* Stick a hash search on the front of the tree search in addpool

* Eliminate repeated hashing of the same string

* Switch instead of ifs in json parsing

* A few more cases to populate the hash in addpool

* Store the hash in the tree instead of recalculating

* Add explanatory comment for mysterious argument

* Fewer copies in attribute stringification

* Clean up ancient weirdness in JSON attribute stringification

* More serial_val cleanup

* Pass a serial_feature to rewrite instead of many broken-down arguments

* Get rid of the multiple geometries within `partial`

* Revert "Pass a serial_feature to rewrite instead of many broken-down arguments"

This reverts commit 6f4ab9b725.

* Goodbye, struct coalesce

* Revert "Features clipped away to nothing should not participate in that tile"

This reverts commit 124462fbdc.

* Migrating fields from partial to serial_feature

* Name reconciliation between serial_feature and partial

* Replace struct partial with an augmented serial_feature

* Fix some overzealous search-and-replace renaming

* Don't say struct so often

* Remove more of the former partial construction

* Commenting and cleaning up

* Trying again to avoid all these arguments to rewrite

* I swear I did this same thing before and it didn't work.

* More rewrite cleanup

* Exile --detect-shared-borders to its own file

* Add missing headers

* More commenting and cleanup

* More comments

* Sprinkle consts around

* Emplacing and std::moving

* More cleanup

* That shouldn't have worked after a std::move

* Don't need to allocate memory to compare keys

* Reduce use of the global string pool in tiling

* Another avoidable mvt_value construction

* Further reduction to explicit string pool passing

* These reverses are no longer optimizations

* These layernames can all be references

* Don't drag an unused layername string around with every feature

* Heed a compiler warning about potential buffer overflow

* Fix my confusion about which feature's string pool is relevant

* Avoid some unnecessary allocations in attribute accumulation

* Maybe faster serialization?

* Eliminate a comparison

* Do the same here

* Save a couple of allocations when parsing numbers in JSON

* Immediately assign features to layers instead of subdividing later

* Maintain tilestats for tippecanoe:retain_points_multiplier_sequence

* Crunch out more duplicate attribute values when writing out the tile

* Do tilestats for tippecanoe:retain_points_multiplier_first too

* Shell filters need to be real threads, even if nothing else does

* Simplify tippecanoe_minzoom/maxzoom representation

* Update version and changelog
2024-02-07 17:50:35 -08:00
Erica Fischer 7d4d264d50 Speeding up tippecanoe-overzoom (#191)
* Speed up mvt_value comparison

* Converting repetitive ifs to cases

* More conversions from ifs to cases

* Optimize the always-true filter case

* Don't convert types of attributes without accumulators

* Unordered map seems to be faster than map

* Add missing header

* Fix some warnings

* Fix the warnings better

* Avoid an int->string->int conversion

* Lazily initialize layer key and values maps when actually needed

* More switches from maps to unordered_maps

* Sure, I'll take the microoptimization

* More emplacement

* Save some copies

* Emplaces and moves

* Lazy linear scan of attributes instead of building a map

* Extra printfs, missing header

* Avoid clipping if the input and output tiles are the same

* But do clip if the tile extent is being reduced

* Make sure I'm not constructing std::strings here at runtime

* More worrying about runtime string construction

* A couple more std::moves

* Const references!

* More const references

* Another std::move

* Make the string_value of mvt_value std::optional

* Reserve storage when decoding

* Provision for different mvt_values to share a string pool

* Use the string pool when decoding

* Avoid another string construction

* Try limiting the depth of the search for duplicate attributes

* Revert "Try limiting the depth of the search for duplicate attributes"

This reverts commit 9ec94a15ff.

* Update changelog

* Fix typo noticed during code review
2024-01-29 11:33:56 -08:00
Erica Fischer f957f30f90 Clean up internal naming related to tilestats (#190)
* Get rid of the type_and_string near-synonym for serial_val

* Rename file_keys to the more familiar tilestats

* "tas" (type_and_string) => "sv" (serial_val)

* "fk" (tile_keys) => "ts" (tilestats)

* Revert ""tas" (type_and_string) => "sv" (serial_val)"

This reverts commit 4854c57e22.

* More carefully this time: "tas" (type_and_string) => "sv" (serial_val)
2024-01-23 14:17:19 -08:00
Erica Fischer 679a0d62f2 Make feature ordering cooperate with --retain-points-multiplier (#188)
* Make feature ordering cooperate with --retain-points-multiplier

* Forgot to check in the actual code changes???

* Sort within each multiplier cluster as well as between clusters

* Correct description of behavior in changelog

* Drag original feature sequence along in megatiles for post-filter sort

* Plumb the preserve-input-order flag through overzoom

* Sort in overzoom if requested

* Use within-tile input sequence numbers, not global sequence numbers

* Documentation

* Reverse direction of search to prevent accidental skipping

* Add some comments about converting between attribute representations
2024-01-21 12:09:05 -08:00
Erica Fischer 26bf08deb2 Externalizing polygon shard detection (#146)
* Start of externalizing polygon shard detection

* Completely untested external quicksort

* Add unit test for external quicksort

* Remember to clean up temporary files

* Sort and scan the vertices

* Bring over more vertex logic

* Make nodes from vertices

* Checkpoint on switching over to global shared nodes

* Do the thing

* Revert unintended change to coalesced linestring behavior

* Take shared nodes into account in early simplification

* Let it do more sorting in memory

* Fix overnoding of collinear linestrings

* Remove duplicate nodes, since only duplicate vertices now matter

* Remember to delete temporary files

* Fix out of bounds memory access below, apparently

* Fix the actual undefined behavior

* Still running out of memory in one case. Find out where.

* Forgot the conditional

* Try again to make it not run out of memory

* Update version and changelog
2023-10-03 10:53:12 -07:00
Erica Fischer 390771f855 Polygon shards (#105)
* Put all of this back in geometry.cpp for conflict resolution

* Stabilize line simplification to behave the same regardless of winding

* Round instead of truncating when clipping lines

* Restore non-Wagyu polygon clipping from prior to 2fdec7d2

* Make it round, not truncate, which reverts the last commit's test diffs

* Clip in floating point, not integers, which makes no difference

* Track nodes added at tile edges during clipping

* Scale geometry up before wagyu to prevent changes from precision loss

* Actually do the shared edge detection

* Fix cases where nodes were not being added at the tile boundary

* One more place I should have rounded

* Narrow down where the discrepancy comes in

* Revert "Narrow down where the discrepancy comes in"

This reverts commit 221c4c5fc0ac9a6567e091c6a94b3d87dc8ade83.

* Another attempt to narrow it down

* The discrepancy seems to be introduced in reordering. No obvious bug

* Was still truncating instead of rounding in projection

* Also makes no difference...

* Just forget that line reversal exists for a minute

* Just reversal no coalescing

* Try clipping in integers instead of floating point

* Don't simplify after coalescing if they said no simplification

* Check whether behavior is consistent with intentional simplification

* Replace more floating point with integer

* Are these three features enough to demonstrate the problem?

* Add a few more nearby borders

* All the features that touch tile 6/16/23

* Stay in integers in line simplification

* More attempts to solve failures to simplify consistently

* Fix most of the overflow errors

* Fix known cases of integer overflow

* All the tests change again

* Pull clipping and scaling code back out into clip.cpp

* Resolve the test conflicts

* Stabilize choice of which three points to keep with different windings

* Almost right, I think!

* Fix collapse of islands to shards

* Self-intersections in the same feature don't count

* Revert "Self-intersections in the same feature don't count"

This reverts commit e04b19916e.

* Don't scale down geometry if we are going to look for shared nodes

* Fix the missing multiply that was keeping simplification from happening

* Fix one more opportunity for overflow

* Somehow I deleted this test?

* Lost this test too

* Clean up debugging printfs

* Restore code sequence from main to make it reviewable

* Remove unneeded rounding

* Update documentation

* This test is no longer useful

* Back to floating point Douglas-Peucker to fix undersimplification

* Try an older ubuntu

* Revert "Try an older ubuntu"

This reverts commit 13fefacfd7.

* Log OS info

* Remove tests that are no longer needed

* Oops, did need that one after all

* Fix the arm vs x86 discrepancy?

* Try another quantization

* Cleanup from review

* Add a test for the actual purpose of this PR
2023-09-14 16:51:38 -07:00
Erica Fischer 523ca55243 Fix bugs in --no-simplification-of-shared-nodes (#99)
* Fix bugs in --no-simplification-of-shared-nodes

* Update changelog
2023-05-08 15:31:46 -07:00
Erica Fischer afc7ca01a2 Calculate an antimeridian-adjusted bounding box in tileset metadata (#82)
* Calculate a new antimeridian-adjusted bounding box

* Add antimeridian bounding box to pmtiles, dirtiles, and tile-join

* Don't take out-of-bounds latitudes into account in the adjusted bbox

* Update changelog

* Forgot to adjust tests after the last change
2023-03-09 20:26:48 -08:00
Erica Fischer f2127fec97 Reduce thrashing during feature ingestion and tiling (#56)
* Remove the concept of "separate metadata"

This was an extra level of attribute indirection (features point
to metadata records which point to key and value strings) which was
intended to reduce the size of temporary storage for features with
large numbers of attributes that were also spread across large numbers
of tiles at maxzoom.

For other kinds of features, the extra indirection slowed things down
instead, and, especially when maxzoom guessing was being used, many more
features were having their metadata externalized than could actually
benefit from it.

* Shave a few bytes off temporary files by using more unsigned integers

* Flush stderr after logging progress

* Revert "Shave a few bytes off temporary files by using more unsigned integers"

This reverts commit eef29084ec.

* Limit the size of the string pools and trees to fit in memory

* Add missing #include

* Move the string pool and search tree from mmap to allocated memory

* Sort in allocated rather than mapped memory too

* Also use pread instead of mapping to read in the data to sort

* When the pool gets too big, switch to just the file, not memory

* Switch string pool from memory to disk when memory is 10% full

* Add to-memory versions of the serialization functions

* Crashy work in progress toward compression

* Fix the pointer bug that was causing the crash

* Serialize features into memory rather than straight to disk

* Compress individual features in the temporary files

* Don't need to store the length of the geometry

* Remove per-feature compression; move minzoom back into the object

* Start adding a stream compressor object

* Track file position within fwrite_check()

* Add compressed stream writer functions

* Pull the writing of the serialized feature out to the callers

* Starting toward compression again from a different point

* Hook up more compression functions

* Remove unused code from the other day

* Make enough deflate calls to flush out all the buffered data

* Start on decompression

* Tile number is uncompressed, tile content is compressed

* Work on alternating compressed and uncompressed in decompression

* Closer, but still doesn't work

* Sort of works

* Works until we get to concatenated tiles

* More attempts that don't work

* One bug down

* It made a tileset!

* Handle nonzero initial zooms

* Fix seeking within compressed feature streams

* Tests pass!

* Remove debug spew

* Oops: remember to delete the temporary files so they don't hang around

* Test that fails with the current compression code

* Properly account for bytes read while closing the compressed stream

* Limit the number of warnings about bad label points

* A little more armor when closing decompression

* This time for sure

* A different, less fragile, test that failed previously with compression

* Move feature stream compression to its own file

* Remove now-unused code to deserialize from a file

* Forgot to add the new files

* Remove a little debugging logging

* Add a couple of comments on what it means to be within decompression

* Fix indentation

* Update changelog. Remove stray debugging comment.
2023-02-14 12:47:40 -08:00
Erica Fischer 622084a507 Limit guessed maxzoom to avoid spending too many tiles on polygon fill (#23)
* Change sqlite3 schema to deduplicate identical tiles

* Limit guessed maxzoom to avoid spending too many tiles on polygon fill

* Fix test.

These dust polygons now have their area calculated because their
maxzoom is being guessed, so the attributes from the largest one
rather than the last one are preserved.

* Increase polygon limit to a million tiles

* Two million tiles ought to be enough for anyone, right?

* Add explanatory comments for mysterious numbers
2022-11-07 10:14:17 -08:00
Erica Fischer b763625862 Add an option to generate label points in place of polygons (#20)
* Add an option to generate label points in place of polygons

* Change all these places where I said "extent" but really meant "area"

* Revert "Change all these places where I said "extent" but really meant "area""

This reverts commit 403828d2f7.

* Add --order-smallest-first and --order-largest-first options

* Use Turf's center-of-mass algorithm for polygon label points

* If the label point isn't within the polygon, find one that is

* Don't choose a label point that is too close to a border

* Try a little harder to find an optimal label point

* Checkerboard which tiles labels are generated in, to reduce adjacency

* Use a label point for the general representative point for polygons

(Skipping the iteration to find one that is as far as possible from
the borders)

This makes the labels look better in many cases (like France at z1)
but unfortunately ripples into changing the sequence of polygons in
many tests, so the diff is big.

* Revert "Use a label point for the general representative point for polygons"

This reverts commit 2261adf05e.

* Checkpoint work on spiral labels

* Clip label spirals to the feature bounds

* Fix label test

* Be careful not to place spiral labels too close to borders either

* For spiral anchors, only check tile scale, not feature size

* Update test

* Only try to find a central label point for the largest ring

* In tiny polygon dust, keep the attributes of the largest feature
2022-10-13 14:21:36 -07:00
Eric Fischer e805d113d0 Remove now-unused filter argument during parsing 2019-01-22 14:26:06 -08:00
Eric Fischer f070c7464b Add missing #include 2018-05-07 15:10:18 -07:00
Eric Fischer 59dd095607 Make file positions and lengths thread-safe 2018-05-07 14:42:49 -07:00
Eric Fischer 1b26becad9 Clear up some confusion about attribute count and external references
Now the count is always adjacent to whereever the key/value pair is
stored, and is not kept in the serial feature object other than as
the length of the vectors of keys and values.
2018-04-05 15:40:14 -07:00
Eric Fischer fac0ebbf52 All the other places where I used volatile but really wanted atomic 2018-03-13 15:21:21 -07:00
Eric Fischer a8a342f701 Send dot-dropping through the same pipeline.
The first feature in a tile can never be dropped, since there is
no previous feature to attach its properties to.

Remove the previous special case that reset the dropping counter
at the first feature within each tile proper (as opposed to the
first feature in each tile, including its buffer, which is now
the one that is guaranteed to be preserved).
2018-02-23 17:19:54 -08:00
Eric Fischer 0152db4a20 More initializers 2017-11-07 15:57:56 -08:00
Eric Fischer 4f974b3dc6 Less verbose initializer syntax 2017-11-07 15:25:54 -08:00
Eric Fischer ba62ab8596 More structure initializers 2017-11-07 15:20:17 -08:00
Eric Fischer 0fd4454129 Allow filter expressions during tippecanoe as well as during tile-join 2017-09-01 11:51:12 -07:00
Eric Fischer 5943c82457 Move file-format-neutral code out of JSON-specific source file 2017-08-28 11:10:57 -07:00
Eric Fischer b98bf6e8c7 Get attribute value decoding working 2017-08-25 15:46:32 -07:00
Eric Fischer e7ee83f27b Move attribute type coercion out of parsing and into serialization 2017-08-24 17:27:30 -07:00
Eric Fischer f4818ffb07 Move attribute include/exclude logic into serialization 2017-08-24 17:10:15 -07:00
Eric Fischer 34b1b215f4 Move tilestats management out of parsing and into serialization 2017-08-24 16:30:01 -07:00
Eric Fischer ed8fbd0236 Split more serialization details out from being parsing parameters 2017-08-24 15:57:33 -07:00
Eric Fischer 6caf20b9c8 Put the pieces back together 2017-08-23 11:43:48 -07:00
Eric Fischer 6cea2d5db6 Progress on factoring out serialization state into a single object 2017-08-22 18:10:52 -07:00
Eric Fischer c79f19e3ca Merge branch 'master' into plugins 2017-08-08 11:08:10 -07:00
Eric Fischer 65c095cc2b Clean up #includes and add fields for counting attributes 2017-07-14 16:56:23 -07:00
Eric Fischer 6a5461763c Fix reordering of attributes and failure to update layer name table 2016-12-20 16:41:23 -08:00
Eric Fischer 0e5b513637 Start getting features (just geometry so far) back from the prefilter 2016-12-09 15:35:57 -08:00
Eric Fischer 569825324a Factor out feature deserialization 2016-12-08 17:11:37 -08:00
Eric Fischer 013e6512b4 Add an option to drop the smallest features to make tiles small enough 2016-11-09 17:09:05 -08:00
Eric Fischer 2e3ba8f374 Retain original feature index rather than recalculating
For better density calculation of clipped features
2016-11-02 15:11:22 -07:00
Eric Fischer c8a1b082e0 Don't serialize the per-feature minzoom until geometry merging time 2016-10-10 15:31:09 -07:00
Eric Fischer 475ce9dd23 Fix g++ compiler warnings 2016-08-08 17:14:48 -07:00
Eric Fischer bf571571a9 Factor out (initial) feature serialization 2016-08-08 15:36:49 -07:00
Eric Fischer 488dff0efb Encode the "id" field of GeoJSON objects as vector tile feature ID 2016-07-15 15:00:56 -07:00
Eric Fischer f3b9e15267 Move serialization code to its own file 2016-04-27 14:19:10 -07:00