Commit Graph
247 Commits
Author SHA1 Message Date
Erica Fischer 6a6d47eb07 Add a tippecanoe-decode option to restrict which attributes to decode 2024-11-06 16:59:53 -08:00
Erica Fischer 23667bb8eb Reduce memory consumption from attribute accumulation (#290)
* Progress on plumbing a string pool for full_keys through

* More plumbing for key_pool

* Don't keep features with identical locations as multiplier features

* Revert "Don't keep features with identical locations as multiplier features"

This reverts commit 413f0c8024.

* Adjust calculated maxzoom to account for duplicate feature locations

* Update changelog and version

* Add a test affected by the maxzoom change with duplicate locations

* Round the drop rate a little for cross-platform test consistency
2024-11-05 14:39:41 -08:00
Erica Fischer 28efc40e6e Choose the megatile features from those that will be in the next N zooms (#280)
* Choose the megatile features from those that will be in the next N zooms

* Take fractional zooms into account in multiplier feature choices

* Fix more tests

* Add a flag to retain multiplier features by minimum distance

* Limit feature expansion from multiplier density to 2x

* The multiplier cap was a bad idea

* Revert "The multiplier cap was a bad idea"

This reverts commit 6f8273a4c8.

* Revert "Limit feature expansion from multiplier density to 2x"

This reverts commit a26e41309d.

* Revert "Add a flag to retain multiplier features by minimum distance"

This reverts commit 01f14a4255.

* Remove the multiplier sequence, which should no longer matter

* Revert "Revert "Add a flag to retain multiplier features by minimum distance""

This reverts commit 776da4a1b8.

* Revert "Revert "Limit feature expansion from multiplier density to 2x""

This reverts commit 44a683d808.

* Revert "Revert "The multiplier cap was a bad idea""

This reverts commit 80f7cb1c0e.

* Track two kinds of previous index for next_feature

* Fix multiplier density threshold, I think

* Oh, I didn't git add the code changes

* Update version and changelog

* Try to install sqlite3 to fix the automated build

* Deleted too much

* Only let --preserve-point-density-threshold shift density around

* Remove the density debt concept, since it doesn't help

* Make the drop states a vector instead of an array

* Revert "Make the drop states a vector instead of an array"

This reverts commit 66c7abb6fa.

* Revert "Remove the density debt concept, since it doesn't help"

This reverts commit 707bb0c562.

* Revert "Only let --preserve-point-density-threshold shift density around"

This reverts commit ecf01f2231.
2024-10-17 15:59:44 -07:00
Erica Fischer c5f2f0da34 More work on plumbing attribute accumulation through (#263)
* Plumb bounding boxes through potential intersections

* Quick bbox reject for bins that can't possibly intersect

* Inching toward attribute accumulation in megatile handling

* Some sort of test for how all these things interact with each other.

Automatic numeric attribute accumulation does *not* apply to attributes
that have an explicit attribute accumulator set, because the order of
operations is too messy and weird

* More sketching

* More sketching

* Actually do some accumulation

* Put all that behind an --accumulate-numeric flag

* Use the same attribute accumulation logic in binning as in megatiles

* Fix backwards conditional

* Add means, but somehow I have some counts of 0

* Handle aggregated attributes with no base attribute in the feature

* Checkpoint before I break everything

* Found a flaw, now to debug

* Fix a typo that broke accumulation

* Add binning tests

* Make sure IDs make it through on the bins

* Fix count/mean accumulation

* Make the numeric accumulation prefix configurable

* Make sure the accumulate test still works with a different prefix

* Forgot to update this test

* More testing to make sure cluster sizes make it all the way through

* Fix neglected --accumulate-attribute when binning

* Mark unexercised attribute accumulation cases as "can't happen"

* Factor out numeric preservation

* Attrs with the accumulation prefix are just preserved, not accumulated

* Test behavior of prefixed attributes

* Plumbing for exclude and exclude-prefix

* Implement and test attribute prefix stripping in overzoom

* Update version and changelog

* For debugging, make an attribute list of source feature IDs

* Revert "For debugging, make an attribute list of source feature IDs"

This reverts commit 65fc99c9d1.
2024-09-20 15:30:29 -07:00
Erica Fischer 51fcf142df Fix bad interaction between dynamic dropping and limiting by truncation (#260)
* Fix bad interaction between dynamic dropping and limiting by truncation

* Don't enforce the tile size limit here if they said not to
2024-09-05 13:39:53 -07:00
Erica Fischer 5b18eea673 Work in progress on binning features in overzoom (#258)
* Factoring out tilestats management from GeoJSON file reading

* Move code around so overzoom can link against parse_layers

* Read the file of bins

* Plumb the bins through to overzoom()

* Some zip code bins to test with

* (Currently non-functional) test of binning

* Starting to spell out the bin matching loop

* Can't flatten points, so don't flatten bins either

* More fleshing out bin traversal

* Bounding box of tile-relative mvt geometry

* Smallest enclosing tile from bbox

* Most of the bin scan

* Add point in polygon check. It crashes.

* Find the matching bins

* GDAL-style bounding boxes have eaten my brain

* Make some features to bin into

* Increment a count as features are found to be within the bins

* Fix longitude wraparound in overzoom bins

* Fix the tests

* Push off attribute copying until after bin assignment

* Carry sum of numeric attributes into the bins

* Also add mean, min, and max

* Add --calculate-feature-index since I keep needing it for testing

* Add an option to accumulate sum/mean/max/min/count of all numeric attrs

* Don't bake in tippecanoe:mean, since we redo it from sum and count

* Forgot to update this test fixture after removing tiled mean

* Update version and changelog
2024-09-05 12:06:51 -07:00
Erica Fischer 40bb4ff732 Be more careful to retry when the feature count is exceeded (#257)
* Clip before dealing with multiplier or filters in overzoom

* Be more careful to retry when the feature count is exceeded

* Adjust the estimated total feature count for the multiplier too

* Fix the feature count estimates, I think

* Pass build info into the version string

* Report the actual max zoom of any tiles as the metadata maxzoom

* Revert unneeded renaming to make the diff more readable

* Clean up the adjustments to tile sizes and feature counts

* Update version and changelog

* Dropping a feature into a multiplier cluster still effectively drops it

* Update changelog

* Rethink the changelog description

* Don't try to truncate zooms if we are still tiling at z18
2024-08-20 10:54:16 -07:00
Erica Fischer bc3ef87c3f Add --generate-variable-depth-tile-pyramid option (#251)
* Track output position at the file level instead of within each tile

* Track file position where the child tile data begins

* Add option and document its intended behavior

* Changing the detail loop to account for stopping early

* I forgot I already added an option for this

* Stop early if we can make a complete tile

* Add a test of zoom truncation with limited feature count

* Forgot to commit the actual code change

* Make room for a vertex count in the header of each serialized tile

* Estimate tile complexity; don't try truncating when unlikely to work

* Be more conservative, because ever retrying a tile is a big speed hit

* If stopping early, don't simplify or clean; leave that to overzoom

* Add tiny polygon reduction / dust to overzoom

* Don't try to stop early in the children if we dropped anything by rate

* Fflush here too before pwriting

* Don't stop early if we ended up dropping any features.

Rework the can-the-next-zoom-stop-early logic to avoid going
one zoom further than needed.

* Fix warning

* Fix warnings

* Oops, checking for the wrong expected return value

* Cleanup from adding line simplification in overzoom

* Current (wrong) behavior when combining coalescing and truncating

* Keep a list of parent tiles to skip rather than truncating

* Now the coalesced tiles in z12 get children in z13

* Don't double-count feature dropping when the zoom level is retried

* Correct README description

* Remove todo about special case below basezoom, which is accounted for

* Be a little more aggressive in drop-densest determination

* Scale tile feature limit for megatiles in the same way as byte limit

* Fully deprecate -detect-shared-borders into an alias

* Track the distances found in the douglas-peucker recursion

* Serialize and deserialize the distance with the vertices

* Revert "Serialize and deserialize the distance with the vertices"

This reverts commit 753f1b7909.

* Revert "Track the distances found in the douglas-peucker recursion"

This reverts commit e5361f8c22.

* Revert "Fully deprecate -detect-shared-borders into an alias"

This reverts commit 0698aeb766.

* Better tracking of whether we failed to make a full-detail tile

* Put a bloom filter in front of the binary search for shared nodes

* Forgot to take out this printf

* Improve dispatch of tiling tasks

* Still dispatch the biggest tasks first

* Track zoom truncation in the strategies list in the tileset metadata

* Prescan for small deltas before doing proper simplification

* Revert "Prescan for small deltas before doing proper simplification"

This reverts commit d1d8238b83.

* Update version and changelog

* Rename to --generate-variable-depth-tile-pyramid
2024-08-06 16:05:52 -07:00
Erica Fischer 47e774adc4 Improve the appearance of coalesce-densest-as-needed tiles (#247)
* Start to distinguish fixed cluster density setting from as-needed density

* Make consistent {drop,coalesce}-densest decisions between zooms

* Actually track the previous index instead of just intending to

* Clean up collinearities in coalesced features

* To determine densest, look at actual physical distance, not just index

* Don't actually need the previous index in serial_feature now

* Center of mass of one feature to most distant point of the next

* Add apologetic comment

* Wait, how did the tests pass before?

* Revert "Wait, how did the tests pass before?"

This reverts commit f73c8ee543.

* Add --maximum-string-attribute-length option

* Update version and changelog

* A little more testing to make sure
2024-07-16 13:23:23 -07:00
Erica Fischer bb4f220678 Reduce tiling memory (#227)
* Trying to reduce memory in tiling

* Let the simplification workers go out of scope earlier

* Bail out quickly once the maximum feature count is reached

* Update changelog and version
2024-04-03 07:12:26 -07:00
Erica Fischer bd48ba8ea1 Fix accidental loss (at all zooms) of features with an explicit minzoom (#221)
* Fix accidental loss (at all zooms) of features with an explicit minzoom

* Try to stabilize centers and bounding boxes between architectures
2024-03-22 10:20:10 -07:00
Erica Fischer 6a8f1b83d8 Fix some undefined behavior (#209)
* Fix some undefined behavior

* Avoid overflow in line simplification calculations

* Oops, missed a test

* Didn't mean to add that to the Makefile

* Revert "Revert "[ci] test in debug mode (#202)""

This reverts commit c95c328e47.

* Fix reference to out-of-scope pointer

* Fix invalid shift and out-of-bounds vector element reference

* Update changelog and version
2024-03-01 10:11:41 -08:00
Erica Fischer 312e1560a3 Stabilize feature order in overzoom (#210)
* Stabilize feature order in overzoom

* I want my sorts to be stable, please

* Revert "[ci] test in debug mode (#202)"

This reverts commit 853ada87b5.

* No need to reinitialize here
2024-02-29 10:39:10 -08:00
Erica Fischer a987197eed Don't swap attributes when reducing tiny polygon dust (#207)
* Don't swap attributes when reducing tiny polygon dust

Because the dust placeholder may be misleadingly far from the feature
that contributed the most area to it

* Remove unused arguments; update changelog
2024-02-26 15:41:26 -08:00
Erica Fischer 2b6630c42c Allow features that would be dropped dynamically to become multiplier features (#199)
* Postpone tagging features as being the first of a multiplier cluster

* Upgrade some dynamically dropped features to multiplier features

* Still don't let it put more features in a cluster than is allowed

* Update tests

* Make the current multiplier cluster size per-layer

* Make sure the first non-empty-geometry in the layer is marked as primary

* Improve comments

* Update changelog and version

* Remove commented out debugging printf

* Factor out duplicated code
2024-02-15 11:34:09 -08:00
Erica Fischer 96f126dd59 FSL-style expressions can use unidecode data to smash case and diacritics (#197)
* Read unidecode data, do some plumbing of it

* More unidecode plumbing

* Do the unidecode smashing, but it doesn't seem to be working

* Ah, that's better!

* Add missing header

* And reorder the includes too

* Shortcut when there is no unidecode data to work with

* Update version and changelog

* Avoid repeated unidecode smashing of the same constant string
2024-02-13 14:18:30 -08:00
Erica Fischer 4e52cbd957 Drop or retain whole multiplier clusters when dropping as needed (#198)
* Prep to track conditions other than just "dropped" or "kept"

* Count up instead of down

* Drop or retain whole multiplier clusters based on their first feature

* Calculate a global feature dropping sequence

* Switch over to using the drop sequence for drop-fraction

* Remove unused arguments for the old drop-fraction implementation

* Fix copy-and-paste bugs, update tests

* Properly incorporate feature_minzoom into the drop sequence, I hope

* Rename drop_by to drop_sequence

* See if sorting within clusters fixes filter stability between zooms

* Remove very chatty debug print

* Update changelog and version

* Use named constants instead of numbers for feature dropping/keeping

* Add comment to explain purpose and method of bit reversal
2024-02-12 10:58:49 -08:00
Erica Fischer e2a7a409c7 Improve tiling speed (#195)
* Add a way to run tippecanoe single-threaded for profiling

* Do less work when the tilestats sample values list is already full

* Save a copy when retrieving the attribute key

* Fewer atomic operations

* Move string hashing from mbtiles to text

* Only do approximate attribute deduplication when writing tiles

* Feature dropping tests are sensitive to exact tile size

* All tile creators now create a string pool for the tile

* Features clipped away to nothing should not participate in that tile

* Revert "Only do approximate attribute deduplication when writing tiles"

This reverts commit c42b34b498.

* Also revert the related test changes

* Revert "Revert "Only do approximate attribute deduplication when writing tiles""

This reverts commit 18509876c3.

* Be more specific about the string hash function

* Use fnv1a instead of std::hash for everything

* Reduce the chance of hash collisions

* Stick a hash search on the front of the tree search in addpool

* Eliminate repeated hashing of the same string

* Switch instead of ifs in json parsing

* A few more cases to populate the hash in addpool

* Store the hash in the tree instead of recalculating

* Add explanatory comment for mysterious argument

* Fewer copies in attribute stringification

* Clean up ancient weirdness in JSON attribute stringification

* More serial_val cleanup

* Pass a serial_feature to rewrite instead of many broken-down arguments

* Get rid of the multiple geometries within `partial`

* Revert "Pass a serial_feature to rewrite instead of many broken-down arguments"

This reverts commit 6f4ab9b725.

* Goodbye, struct coalesce

* Revert "Features clipped away to nothing should not participate in that tile"

This reverts commit 124462fbdc.

* Migrating fields from partial to serial_feature

* Name reconciliation between serial_feature and partial

* Replace struct partial with an augmented serial_feature

* Fix some overzealous search-and-replace renaming

* Don't say struct so often

* Remove more of the former partial construction

* Commenting and cleaning up

* Trying again to avoid all these arguments to rewrite

* I swear I did this same thing before and it didn't work.

* More rewrite cleanup

* Exile --detect-shared-borders to its own file

* Add missing headers

* More commenting and cleanup

* More comments

* Sprinkle consts around

* Emplacing and std::moving

* More cleanup

* That shouldn't have worked after a std::move

* Don't need to allocate memory to compare keys

* Reduce use of the global string pool in tiling

* Another avoidable mvt_value construction

* Further reduction to explicit string pool passing

* These reverses are no longer optimizations

* These layernames can all be references

* Don't drag an unused layername string around with every feature

* Heed a compiler warning about potential buffer overflow

* Fix my confusion about which feature's string pool is relevant

* Avoid some unnecessary allocations in attribute accumulation

* Maybe faster serialization?

* Eliminate a comparison

* Do the same here

* Save a couple of allocations when parsing numbers in JSON

* Immediately assign features to layers instead of subdividing later

* Maintain tilestats for tippecanoe:retain_points_multiplier_sequence

* Crunch out more duplicate attribute values when writing out the tile

* Do tilestats for tippecanoe:retain_points_multiplier_first too

* Shell filters need to be real threads, even if nothing else does

* Simplify tippecanoe_minzoom/maxzoom representation

* Update version and changelog
2024-02-07 17:50:35 -08:00
Erica Fischer 6a2bce8164 Scale the tile size limit up with the multiplier at low zooms (#192)
* Scale the tile size limit up with the multiplier at low zooms

* Add a test to demonstrate that high zoom tiles can't be extra large

* Update changelog and version

* Fail more cleanly when a tile can't be made small enough

* Guard against a cluster where the start marker has been dropped

* Look harder for a working feature interval instead of giving up

* That change to the drop-smallest logic changed a test output

* Update changelog

* Add explanatory comment
2024-01-31 10:23:05 -08:00
Erica Fischer 7d4d264d50 Speeding up tippecanoe-overzoom (#191)
* Speed up mvt_value comparison

* Converting repetitive ifs to cases

* More conversions from ifs to cases

* Optimize the always-true filter case

* Don't convert types of attributes without accumulators

* Unordered map seems to be faster than map

* Add missing header

* Fix some warnings

* Fix the warnings better

* Avoid an int->string->int conversion

* Lazily initialize layer key and values maps when actually needed

* More switches from maps to unordered_maps

* Sure, I'll take the microoptimization

* More emplacement

* Save some copies

* Emplaces and moves

* Lazy linear scan of attributes instead of building a map

* Extra printfs, missing header

* Avoid clipping if the input and output tiles are the same

* But do clip if the tile extent is being reduced

* Make sure I'm not constructing std::strings here at runtime

* More worrying about runtime string construction

* A couple more std::moves

* Const references!

* More const references

* Another std::move

* Make the string_value of mvt_value std::optional

* Reserve storage when decoding

* Provision for different mvt_values to share a string pool

* Use the string pool when decoding

* Avoid another string construction

* Try limiting the depth of the search for duplicate attributes

* Revert "Try limiting the depth of the search for duplicate attributes"

This reverts commit 9ec94a15ff.

* Update changelog

* Fix typo noticed during code review
2024-01-29 11:33:56 -08:00
Erica Fischer cbf222754b Add --accumulate-attribute to tippecanoe-overzoom (#189)
* Starting to factor out attribute accumulation into its own file

* Continuing to factor out attribute accumulation

* Reduce duplicate code

* Plumbing the accumulate-attribute option around

* Call the attribute accumulator

* Test that accumulation works

* Add missing #includes

* Don't sort within individual multiplier clusters

Doing so throws off the spatial distribution of the low zooms

* Docs and changelog

* Add comments
2024-01-23 15:41:36 -08:00
Erica Fischer f957f30f90 Clean up internal naming related to tilestats (#190)
* Get rid of the type_and_string near-synonym for serial_val

* Rename file_keys to the more familiar tilestats

* "tas" (type_and_string) => "sv" (serial_val)

* "fk" (tile_keys) => "ts" (tilestats)

* Revert ""tas" (type_and_string) => "sv" (serial_val)"

This reverts commit 4854c57e22.

* More carefully this time: "tas" (type_and_string) => "sv" (serial_val)
2024-01-23 14:17:19 -08:00
Erica Fischer 679a0d62f2 Make feature ordering cooperate with --retain-points-multiplier (#188)
* Make feature ordering cooperate with --retain-points-multiplier

* Forgot to check in the actual code changes???

* Sort within each multiplier cluster as well as between clusters

* Correct description of behavior in changelog

* Drag original feature sequence along in megatiles for post-filter sort

* Plumb the preserve-input-order flag through overzoom

* Sort in overzoom if requested

* Use within-tile input sequence numbers, not global sequence numbers

* Documentation

* Reverse direction of search to prevent accidental skipping

* Add some comments about converting between attribute representations
2024-01-21 12:09:05 -08:00
Erica Fischer 5d92a17193 Add point retention multiplier (#179)
* Add an option to retain N times as many points as usual at each zoom

* Tests for point multipler with specified and guessed maxzooms

* Work in progress on inverse spatial ordering

* Fix inverse spatial feature order

* --reorder was depending on a feature index that wasn't being preserved

* Separate ordering by feature_minzoom from ordering inverse-spatially

* Add a test for the inverse spatial ordering

* Store the basezoom/droprate/multiplier decisions in tileset metadata

* Progress on adding filters to tippecanoe-overzoom

* Type promotion for comparison

* Look up the attribute value for ordering

* Add test of thinning and ordering features

* Plumb tippecanoe_decisions metadata through pmtiles

* Be careful not to put infinities in JSON

* Fix accidental dropping in what is meant to preserve sparse points

* Start distinguishing true, false, and null in expressions

* Most of the type conversions

* Add boolean conversions

* Literals and conjunctions

* Add filtering to tippecanoe-overzoom

* Add a test of filtering in overzoom

* Fix boolean conjunctions

* Handle the combination of cluster size and filtering

* Rework dot dropping to reconcile density threshold and multiplier

* Revert "Rework dot dropping to reconcile density threshold and multiplier"

This reverts commit f253a66382.

* Retain points by multiplier within each tile, not in global probability

* Test that intends to verify that the multiplier is reversible

* Get the test to detect the discrepancy

* Mark the start of multiplier clusters with a magic attribute

* Add string-contains

* Add in and ni operators

* Revert "Look up the attribute value for ordering"

This reverts commit 56bc73e49a.

* Revert "Type promotion for comparison"

This reverts commit 6f3256f5af.

* Make number formatting in tippecanoe_decisions consistent

* Revert "Add a test for the inverse spatial ordering"

This reverts commit c8047de9ab.

* Revert "Separate ordering by feature_minzoom from ordering inverse-spatially"

This reverts commit 35b19a223c.

* Revert "Fix inverse spatial feature order"

This reverts commit 5978ecdb44.

* Revert "Work in progress on inverse spatial ordering"

This reverts commit fdf230f632.

* Somehow missed the tests associated with that last revert

* Round-robin assign attributes to partials from across the multiplier

* Count the multiplier separately in each layer

* Fix distribution of accumulated attribute across multiplier features

* Update changelog, version, and docs

* Add "is null" and "isnt null" expressions

* Update interpretation of FSL expressions to pass the tests

* Test to assert that polygons are unaffected by the multiplier

* Clean up and comment

* Remove accidental unused case
2024-01-18 15:56:09 -08:00
Erica Fischer 02e3bac2c0 Reduce memory consumption during tiling (#177)
* Drop duplicate geometries sooner in coalescing-as-needed

* Concatenate geometries before partial cleaning

* Spend less time with two feature representations in memory

* Further reduce in-memory duplication

* Remove another copy

* Don't keep duplicates in memory while coalescing

* Also clear out the mvt layer once it is no longer needed

* Update changelog
2023-12-22 14:20:21 -08:00
Erica Fischer 26bf08deb2 Externalizing polygon shard detection (#146)
* Start of externalizing polygon shard detection

* Completely untested external quicksort

* Add unit test for external quicksort

* Remember to clean up temporary files

* Sort and scan the vertices

* Bring over more vertex logic

* Make nodes from vertices

* Checkpoint on switching over to global shared nodes

* Do the thing

* Revert unintended change to coalesced linestring behavior

* Take shared nodes into account in early simplification

* Let it do more sorting in memory

* Fix overnoding of collinear linestrings

* Remove duplicate nodes, since only duplicate vertices now matter

* Remember to delete temporary files

* Fix out of bounds memory access below, apparently

* Fix the actual undefined behavior

* Still running out of memory in one case. Find out where.

* Forgot the conditional

* Try again to make it not run out of memory

* Update version and changelog
2023-10-03 10:53:12 -07:00
Erica Fischer f7dc7faf31 Reduce memory use of polygon shard detection (#139)
* Crunch down memory required per polygon joint

* *Actually* reduce the size of the structure

* Update version and changelog
2023-09-21 10:59:12 -07:00
Erica Fischer 390771f855 Polygon shards (#105)
* Put all of this back in geometry.cpp for conflict resolution

* Stabilize line simplification to behave the same regardless of winding

* Round instead of truncating when clipping lines

* Restore non-Wagyu polygon clipping from prior to 2fdec7d2

* Make it round, not truncate, which reverts the last commit's test diffs

* Clip in floating point, not integers, which makes no difference

* Track nodes added at tile edges during clipping

* Scale geometry up before wagyu to prevent changes from precision loss

* Actually do the shared edge detection

* Fix cases where nodes were not being added at the tile boundary

* One more place I should have rounded

* Narrow down where the discrepancy comes in

* Revert "Narrow down where the discrepancy comes in"

This reverts commit 221c4c5fc0ac9a6567e091c6a94b3d87dc8ade83.

* Another attempt to narrow it down

* The discrepancy seems to be introduced in reordering. No obvious bug

* Was still truncating instead of rounding in projection

* Also makes no difference...

* Just forget that line reversal exists for a minute

* Just reversal no coalescing

* Try clipping in integers instead of floating point

* Don't simplify after coalescing if they said no simplification

* Check whether behavior is consistent with intentional simplification

* Replace more floating point with integer

* Are these three features enough to demonstrate the problem?

* Add a few more nearby borders

* All the features that touch tile 6/16/23

* Stay in integers in line simplification

* More attempts to solve failures to simplify consistently

* Fix most of the overflow errors

* Fix known cases of integer overflow

* All the tests change again

* Pull clipping and scaling code back out into clip.cpp

* Resolve the test conflicts

* Stabilize choice of which three points to keep with different windings

* Almost right, I think!

* Fix collapse of islands to shards

* Self-intersections in the same feature don't count

* Revert "Self-intersections in the same feature don't count"

This reverts commit e04b19916e.

* Don't scale down geometry if we are going to look for shared nodes

* Fix the missing multiply that was keeping simplification from happening

* Fix one more opportunity for overflow

* Somehow I deleted this test?

* Lost this test too

* Clean up debugging printfs

* Restore code sequence from main to make it reviewable

* Remove unneeded rounding

* Update documentation

* This test is no longer useful

* Back to floating point Douglas-Peucker to fix undersimplification

* Try an older ubuntu

* Revert "Try an older ubuntu"

This reverts commit 13fefacfd7.

* Log OS info

* Remove tests that are no longer needed

* Oops, did need that one after all

* Fix the arm vs x86 discrepancy?

* Try another quantization

* Cleanup from review

* Add a test for the actual purpose of this PR
2023-09-14 16:51:38 -07:00
Erica Fischer e6d05bc317 Fix tile-join crash when trying to merge empty tilesets with --overzoom (#138)
* Fix tile-join crash when trying to merge empty tilesets with --overzoom

* Add an option not to reduce tiny polygons to dust at maxzoom

* Add test for prevention of tiny polygon reduction at maxzoom

* Change version number
2023-08-31 13:02:03 -07:00
Erica Fischer 2ec6180003 Fix "strategies" accounting for 0-length linestrings and degenerate polygons (#137)
* Dropping a 0-length feature doesn't count as dropping-as-needed

* Add an option not to reduce tiny polygons to dust at maxzoom

* Add test for prevention of tiny polygon reduction at maxzoom

* Fix accounting for tiny polygons not to include degenerate geometries

* Revert "Add test for prevention of tiny polygon reduction at maxzoom"

This reverts commit f931bbd73e.

* Revert "Add an option not to reduce tiny polygons to dust at maxzoom"

This reverts commit 03f0882bb6.

* Fix tests

* Another test that no longer has any really tiny polygons

* Oops, that broke LineString simplification

* This time for sure!

* Update changelog and version
2023-08-29 10:54:39 -07:00
Erica Fischer 6778aeac52 Add an option to extend zooms if still dropping, but with a limit (#131)
* Add an option to extend zooms if still dropping, but with a limit

* At least when to overzoom, even if not actually doing it yet

* Refactor to give tile-join access to overzoom()

* Didn't work, but *might* have worked

* OK, it did something now

* Ah, there's the bug!

* Hook up pmtiles and dirtiles as overzooming sources

* Add command line option to enable or disable overzooming

* Add (currently broken) test of overzooming in tile-join

* Slightly more abstraction for the tile-join readers

* Factor out duplicated code

* Move construction into a constructor

* More changing accessors to methods

* Reduce magic

* Start tracking a list of the tiles at maxzoom

* I think it worked?

* Add missing #include

* Fix sequence of overzoomed tiles (Y sorts backwards for TMS)

* Don't spend memory on overzooming when we aren't going to use it

* Diff rather than cmp, in the hope of figuring out this broken test

* Keep full coordinate precision if we might extend zooms

* Try a slightly different byte limit

* Make drop-densest more consistent across tile boundaries

* Also affects this test

* Does it behave any differently if it can extend forever?

* I think the discrepancy is a thread-safety problem here

* Revert "Does it behave any differently if it can extend forever?"

This reverts commit 0dff0a0acc.

* Lost this change to the test

* This time for sure!

* Revert "Also affects this test"

This reverts commit cd1f7c2e78.

* Revert "Make drop-densest more consistent across tile boundaries"

This reverts commit 563f7d2bc2.

* Revert "Try a slightly different byte limit"

This reverts commit 2e271213d6.

* Add some more explanatory comments

* Amend the join-test to detect my current bug

* Allow overzooming to complete the zoom if it ever starts

* Forgot to correct the test

* Update changelog and version

* Cleanups from code review

* Remove version number from fixture to fix test
2023-08-25 12:42:38 -07:00
Erica Fischer 4318e964e3 Add maximum pseudocluster size option for low-zoom points (#119)
* Trying to improve cluster positions

* Fix centroid calculation

* Just move points to cluster centroids; add point gap threshold

* Add maximum-point-gap option

* Fix formatting

* Rename options for clarity; add tests; update changelog

* Add missing break to switch

* Revert unneeded constructor cleanup that broke a test somehow

* Clean up comments

* Remove unused --move-points-to-cluster-centroids

* Mention tile-join change in changelog
2023-07-17 15:33:07 -07:00
Erica Fischer c1051232ad Do more of line simplification with integer coordinates (#113)
* Turn off unit, which has bit-rotted

* Instrument line simplification

* Do more work in integer space

* Update tests with simplification done with more integers

* Clean up

* Update changelog

* Eradicate calls to deprecated sprintf()

* Revert "Turn off unit, which has bit-rotted"

This reverts commit 998779313f.

* Fix catch by switching to an unsigned type
2023-07-10 15:54:21 -07:00
Erica Fischer 523ca55243 Fix bugs in --no-simplification-of-shared-nodes (#99)
* Fix bugs in --no-simplification-of-shared-nodes

* Update changelog
2023-05-08 15:31:46 -07:00
Erica FischerandAurele Nitoref 329dd86bfd Add --cluster-maxzoom option (#75)
* Add `--cluster-maxzoom` option

* Add tests, update changelog

* Update manpage

* Update tests

---------

Co-authored-by: Aurele Nitoref <aurele.nitoref@icloud.com>
2023-02-24 16:05:45 -08:00
Erica FischerandAurele Nitoref 3d714e32b0 Add point_count_abbreviated property to clustered features (#74)
* Add `point_count_abbreviated` property to clustered features

* Update tests and changelog

---------

Co-authored-by: Aurele Nitoref <aurele.nitoref@icloud.com>
2023-02-24 15:59:00 -08:00
Erica Fischer f2127fec97 Reduce thrashing during feature ingestion and tiling (#56)
* Remove the concept of "separate metadata"

This was an extra level of attribute indirection (features point
to metadata records which point to key and value strings) which was
intended to reduce the size of temporary storage for features with
large numbers of attributes that were also spread across large numbers
of tiles at maxzoom.

For other kinds of features, the extra indirection slowed things down
instead, and, especially when maxzoom guessing was being used, many more
features were having their metadata externalized than could actually
benefit from it.

* Shave a few bytes off temporary files by using more unsigned integers

* Flush stderr after logging progress

* Revert "Shave a few bytes off temporary files by using more unsigned integers"

This reverts commit eef29084ec.

* Limit the size of the string pools and trees to fit in memory

* Add missing #include

* Move the string pool and search tree from mmap to allocated memory

* Sort in allocated rather than mapped memory too

* Also use pread instead of mapping to read in the data to sort

* When the pool gets too big, switch to just the file, not memory

* Switch string pool from memory to disk when memory is 10% full

* Add to-memory versions of the serialization functions

* Crashy work in progress toward compression

* Fix the pointer bug that was causing the crash

* Serialize features into memory rather than straight to disk

* Compress individual features in the temporary files

* Don't need to store the length of the geometry

* Remove per-feature compression; move minzoom back into the object

* Start adding a stream compressor object

* Track file position within fwrite_check()

* Add compressed stream writer functions

* Pull the writing of the serialized feature out to the callers

* Starting toward compression again from a different point

* Hook up more compression functions

* Remove unused code from the other day

* Make enough deflate calls to flush out all the buffered data

* Start on decompression

* Tile number is uncompressed, tile content is compressed

* Work on alternating compressed and uncompressed in decompression

* Closer, but still doesn't work

* Sort of works

* Works until we get to concatenated tiles

* More attempts that don't work

* One bug down

* It made a tileset!

* Handle nonzero initial zooms

* Fix seeking within compressed feature streams

* Tests pass!

* Remove debug spew

* Oops: remember to delete the temporary files so they don't hang around

* Test that fails with the current compression code

* Properly account for bytes read while closing the compressed stream

* Limit the number of warnings about bad label points

* A little more armor when closing decompression

* This time for sure

* A different, less fragile, test that failed previously with compression

* Move feature stream compression to its own file

* Remove now-unused code to deserialize from a file

* Forgot to add the new files

* Remove a little debugging logging

* Add a couple of comments on what it means to be within decompression

* Fix indentation

* Update changelog. Remove stray debugging comment.
2023-02-14 12:47:40 -08:00
Erica Fischer b155b4671b It only needs to look for other small features when coalescing, not dropping (#64)
* Back out the slow search for other small features when dropping features

* Actually, *do* look for small features still when coalescing

* Update changelog
2023-01-27 12:04:36 -08:00
Erica Fischer c58a8e3b97 Round coordinates instead of truncating them (#60)
* Round coordinates instead of truncating them

* Update all the tests for coordinate rounding changes

* Curses, integer division still truncates

* Fix tests

* Also round instead of shifting when looking for no-op linetos

* Should I worry that the same change for moveto doesn't change any tests?

* Also round instead of shifting when scaling down to maxzoom resolution

* Replace another explicit shift, for origin point

* Round instead of shift when writing clipped geometries to the next zoom

* Fix low-zoom gridding and smaller-than-a-pixel checks

* Don't guess an excessively large maxzoom when there is only one feature

* Add a test for guessing the maxzoom of a single point

* Explicitly sort by index if no other order distinguishes features

* Another affected test
2023-01-27 10:57:22 -08:00
Erica FischerandBrandon Liu 3b2599f587 Add support for pmtiles output format (#55)
* add pmtiles.hpp from github.com/protomaps/PMTiles [#10]

* tippecanoe main writes pmtiles output. [#10]

* detect output format using suffix
* after mbtiles is done writing, replace with pmtiles based on map/image tables.
* add method to write_json for writing json sub-object.

* tippecanoe-decode reads pmtiles input. [#10]

* tile-join reads and writes pmtiles. [#10]

* pmtiles test suite for decode and tile-join [#10]

* add base GitHub CI action for compiling and test suite.

* update pmtiles.hpp with z>15 fix

* Fix some ordering problems with pmtiles decode

* Pmtiles should also pass the raw tiles tests

* Eradicate spaces from tileset metadata JSON fields

* Eradicate spaces from more test fixtures

* Update more tests

* Pmtiles tests pass now too

* Remove unnecessary sort (and make indent)

* Update changelog

* The allow-existing test for pmtiles needs -o, not -e

* Declare --allow-existing to be unsupported for pmtiles.

It was always a bad idea even for mbtiles.

Co-authored-by: Brandon Liu <bdon@bdon.org>
2022-12-29 12:00:13 -08:00
Erica Fischer 079929717e Avoid spending gigabytes of memory on statistics for -as-needed dropping (#50)
* Downsample indices and areas during tiling if they get too big

* Cap the indices rather than downsampling them
2022-12-20 12:54:21 -08:00
Erica Fischer f1df09f147 Generate fewer duplicate label points at high zoom levels (#42) 2022-12-02 11:02:33 -08:00
Erica Fischer 5c647cdb83 Don't preflight zoom levels for potential as-needed dropping (#40)
* Working on eliminating preflighting

* Adjust for the map/images schema change

* Avoid generating duplicate tiles with the detail reduction strategy

* Do error checking if tiles in a directory can't be written

* Update changelog and version

* Revert unintentional code reordering

* Neglected to add the exit on error here
2022-11-29 12:42:30 -08:00
Erica Fischer 3095adb467 Simplify geometry earlier when the in-memory representation of a tile gets large, to reduce peak memory usage (#38)
* Simplify geometries earlier if the in-memory tile starts getting big

* Update changelog
2022-11-22 14:57:07 -08:00
Erica Fischer 8eec2be462 Improve quality of coalescing-as-needed; add options to limit feature count (#25)
* Add an option to limit geometry vertex count

* Update docs and changelog

* Track desired feature count and geometry size in strategies

* Remove dead code for a long-forgotten inaccessible option

* Fix the option name in changelog

* Refine coalesce-smallest to only coalesce onto other small features

* Clean polygons before coalescing-as-needed

* Don't accumulate tiny polygon holes as negative dust

* Add the option not to limit the feature count at maxzoom

* Add options to limit feature count more abruptly in each tile

* Revert "Add the option not to limit the feature count at maxzoom"

This reverts commit ace173ab7a.

* Revert "Fix the option name in changelog"

This reverts commit 38dbce8405.

* Revert "Add an option to limit geometry vertex count"

This reverts commit d379fdf06b.

* Remove test for reverted option

* Remove more leftovers from geometry size limiting

* Update changelog
2022-11-17 13:09:10 -08:00
Erica Fischer 5b03b18f31 Generate label points after simplification, not before. (#27)
* Generate labels points after simplification, not before.

Previously there were some cases where dropping the smallest
polygons would never reduce the number of labels.

* Also wait until after polygon cleaning to make labels

* Remove the extra newlines after the TILES ONLY COMPLETE message
2022-10-21 13:59:36 -07:00
Erica Fischer b763625862 Add an option to generate label points in place of polygons (#20)
* Add an option to generate label points in place of polygons

* Change all these places where I said "extent" but really meant "area"

* Revert "Change all these places where I said "extent" but really meant "area""

This reverts commit 403828d2f7.

* Add --order-smallest-first and --order-largest-first options

* Use Turf's center-of-mass algorithm for polygon label points

* If the label point isn't within the polygon, find one that is

* Don't choose a label point that is too close to a border

* Try a little harder to find an optimal label point

* Checkerboard which tiles labels are generated in, to reduce adjacency

* Use a label point for the general representative point for polygons

(Skipping the iteration to find one that is as far as possible from
the borders)

This makes the labels look better in many cases (like France at z1)
but unfortunately ripples into changing the sequence of polygons in
many tests, so the diff is big.

* Revert "Use a label point for the general representative point for polygons"

This reverts commit 2261adf05e.

* Checkpoint work on spiral labels

* Clip label spirals to the feature bounds

* Fix label test

* Be careful not to place spiral labels too close to borders either

* For spiral anchors, only check tile scale, not feature size

* Update test

* Only try to find a central label point for the largest ring

* In tiny polygon dust, keep the attributes of the largest feature
2022-10-13 14:21:36 -07:00
Erica Fischer a6abb0bc30 Add the option to use a different simplification level at maxzoom (#17)
* Add an option to specify a different simplification at maxzoom

* Add test
2022-10-03 12:18:02 -07:00
Erica Fischer af1a7ed7ae Once features can't possibly fit in a tile, stop trying (#9)
* Stop adding features to a tile if it can't possibly work

* Add --integer and --fraction options to tippecanoe-decode

* Carry the strategies field from tileset metadata through tile-join

* Update changelog

* Assign different codes to different kinds of error exits
2022-09-23 20:01:13 -04:00
Erica Fischer a447dfc089 Extra coordinate precision; feature ordering; compression improvements
* Add an option to retain extra coordinate precision at maxzoom

* Make sure not to shift away the extra detail from coordinates

* Add an option to convert double-precision attributes to single

* Sort attribute values in tiles to make them compress a little better

* Slightly improve polygon simplification

By choosing a point that would be retained after simplification
to be the start/end point that always gets retained

* I regret making all of these tests involve polygons

* Add an option to specify the size of tiny polygons

* Fix accidental requiring of argument for --single-precision

* Guard against duplicate points when generating "sizes" for them

* Restore the intended behavior that tiny polygons don't get simplified

* Make the extra detail settable rather than always maximizing it

* Revert "Improve maxzoom guessing for tightly-clustered point data sources (#4)"

This reverts commit fec5e8354c.

* Add an option to prevent choosing a base zoom higher than the maxzoom

* Keep the drop rate high enough when the basezoom gets constrained

* Revert "Revert "Improve maxzoom guessing for tightly-clustered point data sources (#4)""

This reverts commit db6bc27d9e.

* Add --order-by and --order-descending options

* Accept multiple --order-by and --order-descending-by sort keys
2022-09-06 13:08:11 -07:00