TOUR.md: describe the code as it is, not as it was

Drop the before-and-after framing. Where a passage explained a design by
contrasting it with an older one, it now just explains the design, and
where it noted that something was added or removed at some point, it
describes what is there.

The two places where the history was carrying real information keep the
information without the history: the radix sort's account of two past
bugs becomes a statement of the two invariants that are easy to break
and the reason --prefer-radix-sort exists to exercise them, and the note
that the per-feature minzoom field was once thought useless becomes a
description of what deciding it up front actually buys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012tkvf5HbM7yL8u3gSv9Vxq
This commit is contained in:
Claude
2026-08-13 16:42:54 +00:00
parent 80d79c72f1
commit d8b9555905
+26 -29
View File
@@ -3,21 +3,16 @@ A tour of Tippecanoe
Tippecanoe has what sounds like a simple job: It takes geographic objects in one file format ([GeoJSON](https://geojson.org/)) and copies them into a different file format ([Mapbox Vector Tiles](https://github.com/mapbox/vector-tile-spec)). But it takes a lot of code to do that. What's really going on? Tippecanoe has what sounds like a simple job: It takes geographic objects in one file format ([GeoJSON](https://geojson.org/)) and copies them into a different file format ([Mapbox Vector Tiles](https://github.com/mapbox/vector-tile-spec)). But it takes a lot of code to do that. What's really going on?
These days it is not quite that simple even at the surface. The input can also be [CSV](https://datatracker.ietf.org/doc/html/rfc4180), [Geobuf](https://github.com/mapbox/geobuf), or [FlatGeobuf](https://flatgeobuf.org/), and the output can be an [mbtiles](https://github.com/mapbox/mbtiles-spec) file, a directory of tiles, or a [PMTiles](https://github.com/protomaps/PMTiles) archive. And Tippecanoe is now a family of programs, not one: `tile-join`, `tippecanoe-overzoom`, `tippecanoe-decode`, `tippecanoe-json-tool`, and `tippecanoe-enumerate` all operate on the tiles after the fact. But the core of it is still the same: read a lot of features, put them in an order that makes them easy to thin out, and then divide and conquer the world into tiles. Even at the surface it is not quite that simple. The input can also be [CSV](https://datatracker.ietf.org/doc/html/rfc4180), [Geobuf](https://github.com/mapbox/geobuf), or [FlatGeobuf](https://flatgeobuf.org/), and the output can be an [mbtiles](https://github.com/mapbox/mbtiles-spec) file, a directory of tiles, or a [PMTiles](https://github.com/protomaps/PMTiles) archive. And Tippecanoe is a family of programs, not one: `tile-join`, `tippecanoe-overzoom`, `tippecanoe-decode`, `tippecanoe-json-tool`, and `tippecanoe-enumerate` all operate on the tiles after the fact. But the core of it is this: read a lot of features, put them in an order that makes them easy to thin out, and then divide and conquer the world into tiles.
About this document All the links below point at [`63fcac72`](https://github.com/felt/tippecanoe/tree/63fcac725abeb841a101abfc64189edd66d1cf14) (version 2.81.0), so the line numbers stay meaningful even as the code moves.
-------------------
The first version of this was written in 2017, around [`dc86eb6b5a`](https://github.com/mapbox/tippecanoe/tree/dc86eb6b5a425c91c95666594a7c8289b60eb09d) of `mapbox/tippecanoe`, and three of its sections were never written at all. Tippecanoe has since been forked to [`protomaps/tippecanoe`](https://github.com/protomaps/tippecanoe) and then to [`felt/tippecanoe`](https://github.com/felt/tippecanoe), and has picked up quite a lot of new machinery on the way. This version has been brought up to date, and the missing sections have been filled in.
All the links below point at [`63fcac72`](https://github.com/felt/tippecanoe/tree/63fcac725abeb841a101abfc64189edd66d1cf14) (version 2.81.0), so the line numbers will stay meaningful even as the code moves. Wherever this document says something is "new," it means new since 2017, not new in any particular release. The [CHANGELOG](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/CHANGELOG.md) has the release-by-release version of the same story.
Starting up Starting up
----------- -----------
[The main function](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L3168) has the job of processing the list of options and files that the user provides. I'll go into more detail later about what all those options actually do, but the list of them is [spelled out here](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2979) in a single table. [The main function](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L3168) has the job of processing the list of options and files that the user provides. I'll go into more detail later about what all those options actually do, but the list of them is [spelled out here](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2979) in a single table.
The table used to be written out twice, once with the long names and once with the short names and their constraints, which meant the two could disagree, and a third time by hand in the usage message, which meant that could fall behind both. Now there is only the one table. The entries whose `val` is 0 and whose `flag` is null are not options at all but headings for the usage message, and [`strip_usage_headings()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.hpp#L29) removes them before the table is handed to `getopt_long()`. [`getopt_string()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.hpp#L23) derives the short-option string from the same table, and [`print_usage()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.cpp#L88) prints the message from it, so none of the three can drift apart any more. The other programs in the family use the same mechanism. Everything about the options comes from that one table, so that no two descriptions of them can disagree. The entries whose `val` is 0 and whose `flag` is null are not options at all but headings for the usage message, and [`strip_usage_headings()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.hpp#L29) removes them before the table is handed to `getopt_long()`. [`getopt_string()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.hpp#L23) derives the short-option string from it, and [`print_usage()`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/usage.cpp#L88) prints the usage message from it. The other programs in the family use the same mechanism.
Most of the options set an entry in the `additional[]` or `prevent[]` arrays, indexed by the single-character codes listed in [`options.hpp`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/options.hpp), which is why so much of the code below reads like `if (additional[A_DROP_DENSEST_AS_NEEDED])`. Most of the options set an entry in the `additional[]` or `prevent[]` arrays, indexed by the single-character codes listed in [`options.hpp`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/options.hpp), which is why so much of the code below reads like `if (additional[A_DROP_DENSEST_AS_NEEDED])`.
@@ -28,7 +23,7 @@ Creating the output tileset
[Creating an mbtiles file](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/mbtiles.cpp#L29) is a series of sqlite operations, most of which I copied from [mbutil](https://github.com/mapbox/mbutil). There is an option (`-F`) to keep going even if these operations fail, to support a geocoding use case that needed to add new, non-conflicting, tiles to an existing mbtiles file. [Creating an mbtiles file](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/mbtiles.cpp#L29) is a series of sqlite operations, most of which I copied from [mbutil](https://github.com/mapbox/mbutil). There is an option (`-F`) to keep going even if these operations fail, to support a geocoding use case that needed to add new, non-conflicting, tiles to an existing mbtiles file.
The schema is no longer the simple `metadata` and `tiles` tables that it originally was. Instead there is a `map` table that maps z/x/y to a content hash, an `images` table that maps a content hash (within a zoom level) to the tile data, and a `tiles` *view* that joins the two back together into what a consumer expects. The point is that identical tiles — which are extremely common at low zooms, and in any tileset with a lot of empty ocean — are stored only once. [`mbtiles_write_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/mbtiles.cpp#L104) hashes each tile's contents with [fnv1a](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/text.cpp#L260) to find the key. The schema is a `map` table that maps z/x/y to a content hash, an `images` table that maps a content hash (within a zoom level) to the tile data, and a `tiles` *view* that joins the two back together into what a consumer expects. The point of the indirection is that identical tiles — which are extremely common at low zooms, and in any tileset with a lot of empty ocean — are stored only once. [`mbtiles_write_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/mbtiles.cpp#L104) hashes each tile's contents with [fnv1a](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/text.cpp#L260) to find the key.
There are two other output forms. With `-e`, the output is [a directory of tiles](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/dirtiles.cpp#L28), `z/x/y.pbf`, with a `metadata.json` alongside. If the output name ends in `.pmtiles`, Tippecanoe writes an mbtiles file first and then [repacks it](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/pmtiles_file.cpp#L139) into the single-file PMTiles format at the end, which is described in [its own section](#converting-to-pmtiles) below. From the tiling code's point of view all three are the same: there is an `outdb` or an `outdir`, and it writes tiles to whichever one it has. There are two other output forms. With `-e`, the output is [a directory of tiles](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/dirtiles.cpp#L28), `z/x/y.pbf`, with a `metadata.json` alongside. If the output name ends in `.pmtiles`, Tippecanoe writes an mbtiles file first and then [repacks it](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/pmtiles_file.cpp#L139) into the single-file PMTiles format at the end, which is described in [its own section](#converting-to-pmtiles) below. From the tiling code's point of view all three are the same: there is an `outdb` or an `outdir`, and it writes tiles to whichever one it has.
@@ -48,9 +43,9 @@ The temporary files are:
* `vertex`, which records, for `--no-simplification-of-shared-nodes`, each vertex of each line or ring together with the two vertices on either side of it, so that the points where features diverge from each other can be found globally. * `vertex`, which records, for `--no-simplification-of-shared-nodes`, each vertex of each line or ring together with the two vertices on either side of it, so that the points where features diverge from each other can be found globally.
* `node`, which records the individual points that must not be simplified away. * `node`, which records the individual points that must not be simplified away.
There used to be a sixth, `meta`, which held the key/value lists of features that had enough properties that it would take a lot of space to repeat them in the `geom` file. That second level of indirection is gone: features now always reference their pooled keys and values directly, which turned out to be both simpler and smaller. Features reference their pooled keys and values directly, by their offsets into `pool`; there is no second level of indirection for features that have a lot of properties.
All of this data is in temporary files instead of in memory because it can be very large, and putting it on disk makes it possible to tile data that is too big to fit in memory. Processing will be much faster if the data does fit in memory, though. The `pool` and `tree` in particular use a hybrid: [`memfile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/memfile.cpp#L11) keeps them as an in-memory `std::string` until they get too big, at which point [`memfile_full`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/memfile.cpp#L71) switches to appending to the real file. (Writable memory maps used to be used here, but they had bad performance problems in containers, so now the data is explicitly copied in and out.) All of this data is in temporary files instead of in memory because it can be very large, and putting it on disk makes it possible to tile data that is too big to fit in memory. Processing will be much faster if the data does fit in memory, though. The `pool` and `tree` in particular use a hybrid: [`memfile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/memfile.cpp#L11) keeps them as an in-memory `std::string` until they get too big, at which point [`memfile_full`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/memfile.cpp#L71) switches to appending to the real file. Note that data is explicitly copied in and out rather than accessed through a writable memory map, which has bad performance problems in containers.
What is straightforwardly in memory, not on disk, because it is small, is [the "layermap"](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1417) for each CPU, which is a list (by name) of the tile layers and all the property keys that have been used for any features in that layer, plus the sample values and ranges that will become the tileset's [tilestats](https://github.com/mapbox/mbtiles-spec/blob/master/1.3/spec.md). This will eventually end up in the tileset metadata. What is straightforwardly in memory, not on disk, because it is small, is [the "layermap"](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1417) for each CPU, which is a list (by name) of the tile layers and all the property keys that have been used for any features in that layer, plus the sample values and ranges that will become the tileset's [tilestats](https://github.com/mapbox/mbtiles-spec/blob/master/1.3/spec.md). This will eventually end up in the tileset metadata.
@@ -81,7 +76,7 @@ Now that the temporary files are in place, it is time to start reading input, ei
3. If the user didn't ask for parallel input, it just [runs a single JSON parser](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1834) on the streaming input file. 3. If the user didn't ask for parallel input, it just [runs a single JSON parser](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1834) on the streaming input file.
One thing that has changed is that parallel reading is no longer strictly opt-in: if the stream [begins with an ASCII record separator](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1728) (0x1E), the input is [GeoJSON Text Sequences](https://datatracker.ietf.org/doc/html/rfc8142), which *does* guarantee that features are separated, so Tippecanoe turns parallel parsing on by itself and splits on the separator instead of on newlines. Parallel reading is not strictly opt-in, though: if the stream [begins with an ASCII record separator](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1728) (0x1E), the input is [GeoJSON Text Sequences](https://datatracker.ietf.org/doc/html/rfc8142), which *does* guarantee that features are separated, so Tippecanoe turns parallel parsing on by itself and splits on the separator instead of on newlines.
Parsing GeoJSON in parallel Parsing GeoJSON in parallel
--------------------------- ---------------------------
@@ -97,9 +92,9 @@ Parsing a single GeoJSON stream
To parse GeoJSON, whether from a fraction of a memory-mapped file as described immediately above, or from a regular file stream, Tippecanoe uses [a generic JSON pull-parser](https://github.com/felt/tippecanoe/tree/63fcac725abeb841a101abfc64189edd66d1cf14/jsonpull). By pull-parser, I mean that the parser returns JSON tokens and objects one at a time as they are encountered, rather than building up a full JSON object structure in memory for the entire input file and then presenting that to the caller. To parse GeoJSON, whether from a fraction of a memory-mapped file as described immediately above, or from a regular file stream, Tippecanoe uses [a generic JSON pull-parser](https://github.com/felt/tippecanoe/tree/63fcac725abeb841a101abfc64189edd66d1cf14/jsonpull). By pull-parser, I mean that the parser returns JSON tokens and objects one at a time as they are encountered, rather than building up a full JSON object structure in memory for the entire input file and then presenting that to the caller.
[The loop over those tokens](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L39) has been factored out into `geojson-loop.cpp`, because `tippecanoe-json-tool` needs the same thing. It runs until it encounters either a [bare geometry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L76) (an object whose `type` is `Point`, `LineString`, `Polygon`, `MultiPoint`, `MultiLineString`, or `MultiPolygon` and is not contained within something that could be a feature) or [a feature](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L129) (an object whose `type` is `Feature`). In either case it calls the `add_feature` method of the `json_feature_action` it was given, which for Tippecanoe proper is [the one that calls `serialize_geojson_feature`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson.cpp#L244). [The loop over those tokens](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L39) lives in `geojson-loop.cpp` rather than in `geojson.cpp`, because `tippecanoe-json-tool` needs the same thing. It runs until it encounters either a [bare geometry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L76) (an object whose `type` is `Point`, `LineString`, `Polygon`, `MultiPoint`, `MultiLineString`, or `MultiPolygon` and is not contained within something that could be a feature) or [a feature](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L129) (an object whose `type` is `Feature`). In either case it calls the `add_feature` method of the `json_feature_action` it was given, which for Tippecanoe proper is [the one that calls `serialize_geojson_feature`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson.cpp#L244).
The checks for whether something that looks like a geometry or a feature really is one have become more careful over the years: an object is not a feature or geometry of its own if it [appears inside another object's `properties`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L103), and a `GeometryCollection` inside a feature is [expanded into one feature per geometry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson.cpp#L246) sharing the same attributes. The checks for whether something that looks like a geometry or a feature really is one are fussier than they first appear: an object is not a feature or geometry of its own if it [appears inside another object's `properties`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson-loop.cpp#L103), and a `GeometryCollection` inside a feature is [expanded into one feature per geometry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geojson.cpp#L246) sharing the same attributes.
From GeoJSON feature to internal feature From GeoJSON feature to internal feature
---------------------------------------- ----------------------------------------
@@ -113,13 +108,13 @@ It then calls [`parse_coordinates`](https://github.com/felt/tippecanoe/blob/63fc
Internal representation of a feature Internal representation of a feature
------------------------------------ ------------------------------------
The [`serial_feature` structure](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L105) is the thing that every frontend produces and that every stage of tiling passes around. It has grown a lot. Roughly, its fields divide into three groups, and the comments in the header say which is which: The [`serial_feature` structure](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L105) is the thing that every frontend produces and that every stage of tiling passes around. It is a big structure, but its fields divide into three groups, and the comments in the header say which is which:
* The ones that are actually serialized to the `geom` file: the layer, segment, and sequence numbers; the geometry type `t`; the geometry itself; the feature `id`, if any; the per-feature `tippecanoe_minzoom` and `tippecanoe_maxzoom`, if any; the quadkey `index`; the `extent` (area, or a stand-in for it); the `label_point`; and the pooled `keys` and `values`. * The ones that are actually serialized to the `geom` file: the layer, segment, and sequence numbers; the geometry type `t`; the geometry itself; the feature `id`, if any; the per-feature `tippecanoe_minzoom` and `tippecanoe_maxzoom`, if any; the quadkey `index`; the `extent` (area, or a stand-in for it); the `label_point`; and the pooled `keys` and `values`.
* The ones that only exist during initial serialization: `full_keys` and `full_values`, the string forms of the attributes, which get replaced by the pooled `keys` and `values` offsets. * The ones that only exist during initial serialization: `full_keys` and `full_values`, the string forms of the attributes, which get replaced by the pooled `keys` and `values` offsets.
* The ones that only exist during tiling: the bounding box, the `dropped` state, whether the feature is polygon dust or was coalesced, the current detail and simplification level, and so on. * The ones that only exist during tiling: the bounding box, the `dropped` state, whether the feature is polygon dust or was coalesced, the current detail and simplification level, and so on.
`dropped` deserves a note, because it is not a boolean any more. It is [`FEATURE_DROPPED` (-1), `FEATURE_KEPT` (0), a positive sequence number](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L146) meaning "this is the *n*th extra feature retained by `--retain-points-multiplier` alongside its cluster's lead feature," or `INT_MAX` meaning "retained by `--preserve-multiplier-density-threshold`." Much of the dropping logic in `write_tile` is really about keeping these multiplier clusters intact. `dropped` deserves a note, because despite the name it is not a boolean. It is [`FEATURE_DROPPED` (-1), `FEATURE_KEPT` (0), a positive sequence number](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L146) meaning "this is the *n*th extra feature retained by `--retain-points-multiplier` alongside its cluster's lead feature," or `INT_MAX` meaning "retained by `--preserve-multiplier-density-threshold`." Much of the dropping logic in `write_tile` is really about keeping these multiplier clusters intact.
Attribute values are represented by [`serial_val`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L39): a type tag (one of the `mvt_value` types) plus the value as a string. Every number, integer or floating point, is `mvt_double` here and is stored in its stringified form, which is how integers too large for a double survive with their original precision. Attribute values are represented by [`serial_val`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.hpp#L39): a type tag (one of the `mvt_value` types) plus the value as a string. Every number, integer or floating point, is `mvt_double` here and is stored in its stringified form, which is how integers too large for a double survive with their original precision.
@@ -135,7 +130,7 @@ Writing a feature to the temporary files
* If `--no-simplification-of-shared-nodes` is in effect, [writes out the `vertex` and `node` records](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L507) described above. This is one of the more subtle parts of the program: for each vertex it records the two neighboring vertices, so that after a global sort it can tell the difference between a point where two features merely run alongside each other (same neighbors) and a point where they actually converge, diverge, or cross (different neighbors). Only the latter must be protected from simplification. It also unconditionally protects each ring's start point, its farthest point from the start, and the point farthest from the line between those two, so that polygons can't be simplified out of existence. * If `--no-simplification-of-shared-nodes` is in effect, [writes out the `vertex` and `node` records](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L507) described above. This is one of the more subtle parts of the program: for each vertex it records the two neighboring vertices, so that after a global sort it can tell the difference between a point where two features merely run alongside each other (same neighbors) and a point where they actually converge, diverge, or cross (different neighbors). Only the latter must be protected from simplification. It also unconditionally protects each ring's start point, its farthest point from the start, and the point farthest from the line between those two, so that polygons can't be simplified out of existence.
* If `-zg` was given, [samples the distances between the vertices within the feature](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L608), which will feed into the maxzoom guess. * If `-zg` was given, [samples the distances between the vertices within the feature](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L608), which will feed into the maxzoom guess.
* [Calculates the feature's extent](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L636): the area for polygons, and for lines the area of a circle whose diameter is the line's length. Points get theirs later, in `write_tile`, from the distance to the adjacent feature. This is what `--drop-smallest-as-needed` and `--order-largest-first` sort on. * [Calculates the feature's extent](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L636): the area for polygons, and for lines the area of a circle whose diameter is the line's length. Points get theirs later, in `write_tile`, from the distance to the adjacent feature. This is what `--drop-smallest-as-needed` and `--order-largest-first` sort on.
* [Chooses the point that the feature will be indexed by](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L686). For points this is the bounding box center as it always was, but for lines and polygons it is now an arbitrary-but-predictable vertex chosen by hashing the geometry, so that a pile of nearly-identical LineStrings that follow the same route don't all land on the same index and make `-zg` think the data is much denser than it is. Polygons being dropped or coalesced by density instead use their center of mass. * [Chooses the point that the feature will be indexed by](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L686). For points this is the bounding box center. For lines and polygons it is an arbitrary but predictable vertex, chosen by hashing the geometry, so that a pile of nearly-identical LineStrings that follow the same route don't all land on the same index and make `-zg` think the data is much denser than it is. (Hashing the geometry rather than picking at random means that features with identical geometry still get identical indexes.) Polygons being dropped or coalesced by density instead use their center of mass.
* [Generates a label point](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L734) for polygons, if `--generate-polygon-label-points` was given, using [`polygon_to_anchor`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geometry.cpp#L766), which tries to find a point that is well inside the polygon and not in one of its holes. * [Generates a label point](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L734) for polygons, if `--generate-polygon-label-points` was given, using [`polygon_to_anchor`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/geometry.cpp#L766), which tries to find a point that is well inside the polygon and not in one of its holes.
* [Decides whether the index is needed at all](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L743) — it is only serialized if some option that uses it is in effect — and [adds a layermap entry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L763) for the feature's layer. Because layers within each serialized feature are specified by number, not by name, different threads may assign different layer numbers to the same layer name, which will need to be reconciled later. * [Decides whether the index is needed at all](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L743) — it is only serialized if some option that uses it is in effect — and [adds a layermap entry](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L763) for the feature's layer. Because layers within each serialized feature are specified by number, not by name, different threads may assign different layer numbers to the same layer name, which will need to be reconciled later.
* Applies the attribute options: [`--set-attribute`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L785), [`--attribute-type` coercion, `--single-precision`, `--use-attribute-for-id`, `--exclude`, `--include`, and `--exclude-all`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L801), and [`--maximum-string-attribute-length`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L870). It also [feeds each surviving attribute into the tilestats](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L863) — unless a shell filter is in use, in which case there is no point, because the filter may change everything. * Applies the attribute options: [`--set-attribute`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L785), [`--attribute-type` coercion, `--single-precision`, `--use-attribute-for-id`, `--exclude`, `--include`, and `--exclude-all`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L801), and [`--maximum-string-attribute-length`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L870). It also [feeds each surviving attribute into the tilestats](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L863) — unless a shell filter is in use, in which case there is no point, because the filter may change everything.
@@ -160,7 +155,7 @@ Serialized representation of a feature
Serializing each feature [probably looks familiar](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L173) if you have done any serialization with [protozero](https://github.com/mapbox/protozero). There is not much sanity-checking so you have to be careful if you make changes, especially in the field that [does some bit packing](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L187) to combine the layer number and the flags for the presence of a label point, index, extent, `id`, `minzoom`, and `maxzoom` into the same number. Serializing each feature [probably looks familiar](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L173) if you have done any serialization with [protozero](https://github.com/mapbox/protozero). There is not much sanity-checking so you have to be careful if you make changes, especially in the field that [does some bit packing](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/serial.cpp#L187) to combine the layer number and the flags for the presence of a label point, index, extent, `id`, `minzoom`, and `maxzoom` into the same number.
One thing that has changed: serialization now goes into a `std::string` and the caller writes that out, rather than going straight to the file. That is what makes it possible for the same function to be used both by the frontends, writing to the per-thread `geom` files, and by [`rewrite`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L443) during tiling, writing (compressed) into the next zoom level's files. Serialization goes into a `std::string` and the caller writes that out, rather than going straight to a file. That is what makes it possible for the same function to serve both the frontends, writing to the per-thread `geom` files, and [`rewrite`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L443) during tiling, writing compressed into the next zoom level's files.
All of the primitive serialization types (integers of various sizes and signednesses) are spelled out earlier in the same `serial.cpp` file. They use `protozero`'s zigzag conversion functions to represent signed numbers as unsigned, and use its same variable-length encoding for numbers, with the high bit of each byte indicating that there are more bytes to follow. All of the primitive serialization types (integers of various sizes and signednesses) are spelled out earlier in the same `serial.cpp` file. They use `protozero`'s zigzag conversion functions to represent signed numbers as unsigned, and use its same variable-length encoding for numbers, with the high bit of each byte indicating that there are more bytes to follow.
@@ -192,12 +187,12 @@ So now [the `FILE` for each is closed](https://github.com/felt/tippecanoe/blob/6
Merging the string pool is the easiest. The `tree` files are of no more use and are simply closed and discarded. The `pool` files for each CPU [are concatenated together](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1950) into a new file, with an array to keep track of the offset into the new file where the data for each CPU begins. (Which thread some data originally came from is frequently referred to in the code as its `segment`.) Because the pool may be partly in memory and partly on disk by this point, the merge has to handle both cases. The combined pool is then [mapped back into memory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2003) so it can be accessed like an array. Merging the string pool is the easiest. The `tree` files are of no more use and are simply closed and discarded. The `pool` files for each CPU [are concatenated together](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1950) into a new file, with an array to keep track of the offset into the new file where the data for each CPU begins. (Which thread some data originally came from is frequently referred to in the code as its `segment`.) Because the pool may be partly in memory and partly on disk by this point, the merge has to handle both cases. The combined pool is then [mapped back into memory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2003) so it can be accessed like an array.
Then, if `--no-simplification-of-shared-nodes` was requested, there are two more merges that didn't exist in 2017: Then, if `--no-simplification-of-shared-nodes` was requested, there are two more merges:
* [The vertices are sorted](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2015) with [`fqsort`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/sort.cpp#L9) — a quicksort that partitions into temporary files when the data doesn't fit in memory — and then scanned. Any middle vertex that appears with *different* neighbors in two different records is a place where features diverge, so it becomes a node. * [The vertices are sorted](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2015) with [`fqsort`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/sort.cpp#L9) — a quicksort that partitions into temporary files when the data doesn't fit in memory — and then scanned. Any middle vertex that appears with *different* neighbors in two different records is a place where features diverge, so it becomes a node.
* [The nodes are sorted and deduplicated](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2080) the same way. The result is mapped into memory so that each tile can binary-search it, and, because that search would otherwise be a lot of cache misses, a [34-megabyte Bloom filter](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2075) is built alongside it so that the common "this point is not a shared node" answer can be given without touching the array at all. The nodes are keyed by quadkey so that the nodes for any one tile are adjacent in memory. * [The nodes are sorted and deduplicated](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2080) the same way. The result is mapped into memory so that each tile can binary-search it, and, because that search would otherwise be a lot of cache misses, a [34-megabyte Bloom filter](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2075) is built alongside it so that the common "this point is not a shared node" answer can be given without touching the array at all. The nodes are keyed by quadkey so that the nodes for any one tile are adjacent in memory.
This used to be done per-tile and in memory, which was correct but used far too much memory on large inputs. Doing this globally, once, is what keeps the memory in bounds: the alternative of assembling the node list separately within each tile is correct but ruinously expensive on large inputs.
Why sort the features? Why sort the features?
---------------------- ----------------------
@@ -223,12 +218,14 @@ The [top-level sort](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a10
It then [goes through each of those parts](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L890) and checks whether it would fit in memory. If it would, it splits that portion of the index up again by however many CPUs it has to work with, [sorts each of those sub-indices](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L413) in its own thread with the system `qsort`, and then [merges those sub-sub-indices](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L352) to copy the sub-geometry to the final geometry file in index order, which should be fast because it's all in memory. If one of the parts *didn't* fit in memory, then it does [another radix sort](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1054) to split it up further, based on the next bits of the quadkey. In at most a few passes of this it will have generated final sorted index and geometry files in quadkey order. It then [goes through each of those parts](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L890) and checks whether it would fit in memory. If it would, it splits that portion of the index up again by however many CPUs it has to work with, [sorts each of those sub-indices](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L413) in its own thread with the system `qsort`, and then [merges those sub-sub-indices](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L352) to copy the sub-geometry to the final geometry file in index order, which should be fast because it's all in memory. If one of the parts *didn't* fit in memory, then it does [another radix sort](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1054) to split it up further, based on the next bits of the quadkey. In at most a few passes of this it will have generated final sorted index and geometry files in quadkey order.
Two details of this were wrong for a long time and are worth pointing out as warnings. A bucket that was small enough to be written out directly rather than through the merge was being [written one byte longer than its length prefix claimed](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1014) — the `MAGIC` minzoom byte again — which desynchronized everything read from the geometry after it. And the subdivision could [recurse forever](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L746) once it ran out of files to split with, because a "split" into one bucket consumes no bits of the index and the shift by the full width of the index is undefined. Both were only reachable on inputs big enough that nobody was running them in a test, which is why `--prefer-radix-sort` exists: it [pretends there are only 8 KB of memory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1103) so that the deep recursion happens on tiny test inputs, and the test checks the radix-sorted output against the in-memory sort of the same data. Two invariants here are easy to break and hard to notice. The first is that a bucket small enough to be written out directly, rather than through the merge, must still be [written one byte shorter than the index says](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1014) and then given a fresh minzoom byte, exactly as `merge()` does — the `MAGIC` byte again. Write the original byte as well as the new one and every feature read after it comes out shifted by one. The second is that each level of subdivision must [consume at least one bit of the index](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L746), which means at least two buckets: with only one, the recursion never reaches the prefix width that stops it, and the shift that selects a bucket is by the full width of the index, which is undefined.
Both invariants only matter on inputs large enough to recurse, which is more data than a test wants to handle, so `--prefer-radix-sort` exists to reach them: it [pretends there are only 8 KB of memory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1103) so that a tiny test input goes down the deep path, and the test checks the radix-sorted output against the in-memory sort of the same data.
Guessing maxzoom Guessing maxzoom
---------------- ----------------
`-zg` did not exist in 2017; `-z` was something you had to choose yourself. Now, once the index is in order, Tippecanoe can [guess a maxzoom](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2279) from the data. Once the index is in order, Tippecanoe can [guess a maxzoom](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2279) from the data, if `-zg` was given instead of an explicit `-z`.
The idea is that a good maxzoom is one where most features are distinguishable from each other, so it walks the sorted index and accumulates the mean and standard deviation of the log of the gaps between adjacent quadkeys, using [Welford's algorithm](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2312). Distances between features are typically lognormally distributed, so the geometric mean is the right average, and [1.5 standard deviations below the mean](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2353) is taken as the distance at which features should still be distinguishable. The conversion from quadkey gaps to feet is an empirical fit; the `#if 0` block just above the calculation is the code that produced the data for it. The idea is that a good maxzoom is one where most features are distinguishable from each other, so it walks the sorted index and accumulates the mean and standard deviation of the log of the gaps between adjacent quadkeys, using [Welford's algorithm](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2312). Distances between features are typically lognormally distributed, so the geometric mean is the right average, and [1.5 standard deviations below the mean](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2353) is taken as the distance at which features should still be distinguishable. The conversion from quadkey gaps to feet is an empirical fit; the `#if 0` block just above the calculation is the code that produced the data for it.
@@ -254,7 +251,7 @@ Now that it knows how many features are in the densest tile at each zoom level,
Precalculating the dot dropping Precalculating the dot dropping
------------------------------- -------------------------------
The original version of this document had a note that the per-feature minzoom field "doesn't really make sense any more and should be removed." That turned out to be exactly backwards: it is now the mechanism that makes dot-dropping consistent from one zoom to the next, which is what keeps points from popping in and out as you zoom. The per-feature minzoom field is what makes dot-dropping consistent from one zoom to the next, which is what keeps points from popping in and out as you zoom. Deciding it up front, for all the features at once, is what makes that consistency possible: no tile has to work out for itself which features it is entitled to show.
[`calc_feature_minzoom`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L299) assigns each feature the lowest zoom at which it will appear, by keeping a running counter per zoom level that is decremented by that zoom's interval (`droprate^(basezoom - z)`) for each feature that goes by. It is called from inside the merge, as the features are being copied into index order, so it sees them in quadkey order and therefore spreads its choices evenly over space rather than picking a random subset. `--preserve-point-density-threshold` adds an override: if a feature was assigned to a high zoom but is nevertheless very far from the last feature chosen for some low zoom, it gets pushed out at that low zoom anyway, so that sparse areas of the map don't go blank. [`calc_feature_minzoom`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L299) assigns each feature the lowest zoom at which it will appear, by keeping a running counter per zoom level that is decremented by that zoom's interval (`droprate^(basezoom - z)`) for each feature that goes by. It is called from inside the merge, as the features are being copied into index order, so it sees them in quadkey order and therefore spreads its choices evenly over space rather than picking a random subset. `--preserve-point-density-threshold` adds an override: if a feature was assigned to a high zoom but is nevertheless very far from the last feature chosen for some low zoom, it gets pushed out at that low zoom anyway, so that sparse areas of the map don't go blank.
@@ -265,7 +262,7 @@ Running through the zoom levels
Now it is finally time for tiling to begin. All that will remain after tiling is complete is to calculate the final bounding box and layer metadata and write it to the tileset. Now it is finally time for tiling to begin. All that will remain after tiling is complete is to calculate the final bounding box and layer metadata and write it to the tileset.
The [outer loop of tiling](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3301) starts by running through the possible zoom levels. It normally starts at zoom 0 even if some higher `minzoom` was specified, because the entire geometry is now in a single file, not at all split up by tile. If we tried to use the same simple index-based technique as above to split it up into `minzoom` tiles, we would miss some features that are big enough to span (or be buffered into) multiple tiles. Instead, we must "divide and conquer" the low zooms to get to the high zooms correctly, even if some of them will not ultimately be written to the tileset. (Early versions of Tippecanoe did use the index instead, and it worked badly with large and buffered features.) The [outer loop of tiling](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3301) starts by running through the possible zoom levels. It normally starts at zoom 0 even if some higher `minzoom` was specified, because the entire geometry is now in a single file, not at all split up by tile. If we tried to use the same simple index-based technique as above to split it up into `minzoom` tiles, we would miss some features that are big enough to span (or be buffered into) multiple tiles. Instead, we must "divide and conquer" the low zooms to get to the high zooms correctly, even if some of them will not ultimately be written to the tileset.
The one exception is that [`choose_first_zoom`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1147) checks whether the bounding box of all the features fits within a single tile at some higher zoom, and if so starts there, since dividing and conquering an empty world is a waste of time. The one exception is that [`choose_first_zoom`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L1147) checks whether the bounding box of all the features fits within a single tile at some higher zoom, and if so starts there, since dividing and conquering an empty world is a waste of time.
@@ -281,14 +278,14 @@ Once the input files and output files have been allocated to threads, it is time
Each thread, then, [starts going through its list of input files](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3141). From each one, it reads a z/x/y tile number and [then calls the badly-named `write_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3206) to do the work of processing that tile and writing out its children to that thread's set of output files. This loop continues until the thread has exhausted all the tiles in all its input files. The loop also does a little bit of work [to track the densest tile at maxzoom](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3214) which will be used later for the map center in the tileset metadata. Each thread, then, [starts going through its list of input files](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3141). From each one, it reads a z/x/y tile number and [then calls the badly-named `write_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3206) to do the work of processing that tile and writing out its children to that thread's set of output files. This loop continues until the thread has exhausted all the tiles in all its input files. The loop also does a little bit of work [to track the densest tile at maxzoom](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3214) which will be used later for the map center in the tileset metadata.
Each tile's data in the temporary file is now preceded by an [uncompressed `estimated_complexity`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3148), the number of bytes of feature data that tile contains. It can only be known after the data has been written, so a placeholder is written first and then [rewritten with `pwrite`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2355) once the stream is finished (and likewise [for the initial zoom-0 file](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2227)). It is used by variable-depth pyramids, described below. Each tile's data in the temporary file is preceded by an [uncompressed `estimated_complexity`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3148), the number of bytes of feature data that tile contains. It can only be known after the data has been written, so a placeholder is written first and then [rewritten with `pwrite`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2355) once the stream is finished (and likewise [for the initial zoom-0 file](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/main.cpp#L2227)). It is used by variable-depth pyramids, described below.
Retrying a whole zoom level Retrying a whole zoom level
--------------------------- ---------------------------
There is a loop around the whole zoom level, not just around each tile: [`for (size_t pass = 0;; pass++)`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3392). If any tile in the zoom found that it had to raise one of the as-needed thresholds (the minimum gap, the minimum extent, the minimum drop sequence, the minimum attribute value, or gamma) to make itself fit, that new threshold is [collected from all the threads](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3506), the zoom level is [erased from the output](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3563), and the whole zoom is tiled again with the higher threshold. This is why dropping is consistent across a zoom level rather than varying tile by tile. There is a loop around the whole zoom level, not just around each tile: [`for (size_t pass = 0;; pass++)`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3392). If any tile in the zoom found that it had to raise one of the as-needed thresholds (the minimum gap, the minimum extent, the minimum drop sequence, the minimum attribute value, or gamma) to make itself fit, that new threshold is [collected from all the threads](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3506), the zoom level is [erased from the output](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3563), and the whole zoom is tiled again with the higher threshold. This is why dropping is consistent across a zoom level rather than varying tile by tile.
It used to preflight each zoom instead, to find the thresholds before writing anything, but writing the tiles and then throwing them away if necessary turned out to be faster in the common case where nothing has to be thrown away. Note that the tiles are written before it is known whether they will have to be thrown away, rather than the zoom being preflighted to find its thresholds first. That is the cheaper arrangement, because in the common case nothing has to be thrown away at all.
`--extend-zooms-if-still-dropping` is handled [here too](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3549): if the maxzoom is still dropping features, the maxzoom goes up by one and tiling continues. `--extend-zooms-if-still-dropping` is handled [here too](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3549): if the maxzoom is still dropping features, the maxzoom goes up by one and tiling continues.
@@ -331,18 +328,18 @@ Once all the features have been read, the work becomes per-layer rather than per
Then the layers are [converted into an `mvt_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2617), with the pooled keys and values turned back into tile attributes by [`decode_meta`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L235), the postfilter is run if there is one, and the tile is [encoded and compressed](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2876). Then the layers are [converted into an `mvt_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2617), with the pooled keys and values turned back into tile attributes by [`decode_meta`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L235), the postfilter is run if there is one, and the tile is [encoded and compressed](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2876).
And now the size checks, which are the reason for the whole loop. If the tile has too many features or too many bytes — extrapolated upward if features were skipped — then Tippecanoe picks the strategy it was told to use and [raises that strategy's threshold](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2913), aiming at a fraction of the current size, and goes around again. The functions that choose the new threshold ([`choose_mingap`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L754) and its siblings) work on the samples that were collected on the way through. If the threshold can't be raised any further, that is now an error rather than a silent fallback to lower detail — an infinite loop is worse than a failure. If no dropping strategy was requested at all, the detail is reduced by one and the tile is tried again from the top. And now the size checks, which are the reason for the whole loop. If the tile has too many features or too many bytes — extrapolated upward if features were skipped — then Tippecanoe picks the strategy it was told to use and [raises that strategy's threshold](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L2913), aiming at a fraction of the current size, and goes around again. The functions that choose the new threshold ([`choose_mingap`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L754) and its siblings) work on the samples that were collected on the way through. If the threshold can't be raised any further, that is an error: there is nothing left to try, and looping forever is worse than failing. If no dropping strategy was requested at all, the detail is reduced by one and the tile is tried again from the top.
If the tile does fit, it is [written to the database or directory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3048) under the `db_lock`, the strategies used are folded into the zoom's statistics, and `write_tile` returns. If the tile does fit, it is [written to the database or directory](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3048) under the `db_lock`, the strategies used are folded into the zoom's statistics, and `write_tile` returns.
Variable-depth tile pyramids Variable-depth tile pyramids
---------------------------- ----------------------------
`--generate-variable-depth-tile-pyramid` is the newest of the big features and cuts across everything above, so it gets its own section. `--generate-variable-depth-tile-pyramid` cuts across everything above, so it gets its own section.
The idea is that a tile in an empty or simple part of the map doesn't need children: if all of its features can be included at full precision, a client can overzoom that tile instead of downloading deeper ones. So when it is enabled, `write_tile` [makes an extra first attempt](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L1701) at a much higher detail (`30 - z`) with simplification turned off, and only bothers to do so if the `estimated_complexity` recorded with the tile's data suggests it might work. If everything fits at that detail and nothing was dropped, the tile is a leaf: it is [added to `skip_children`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3065) and its descendants are [skipped rather than written](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3179) at the next zoom. The idea is that a tile in an empty or simple part of the map doesn't need children: if all of its features can be included at full precision, a client can overzoom that tile instead of downloading deeper ones. So when it is enabled, `write_tile` [makes an extra first attempt](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L1701) at a much higher detail (`30 - z`) with simplification turned off, and only bothers to do so if the `estimated_complexity` recorded with the tile's data suggests it might work. If everything fits at that detail and nothing was dropped, the tile is a leaf: it is [added to `skip_children`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3065) and its descendants are [skipped rather than written](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3179) at the next zoom.
The bookkeeping around that is fiddlier than it sounds, and most of it exists because of bugs found after the fact: The bookkeeping around that is fiddlier than it sounds:
* The children's geometry is still in the stream even for a skipped tile, so [`skip_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L1659) has to read past it rather than seek. * The children's geometry is still in the stream even for a skipped tile, so [`skip_tile`](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L1659) has to read past it rather than seek.
* A zoom that later has to start dropping features [invalidates the truncation](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3180), because a leaf is only legitimate if it is complete. Those revived tiles have never written their children, so they have to do it on [the pass where the dropping started](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3191) rather than on pass 0 as usual. * A zoom that later has to start dropping features [invalidates the truncation](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3180), because a leaf is only legitimate if it is complete. Those revived tiles have never written their children, so they have to do it on [the pass where the dropping started](https://github.com/felt/tippecanoe/blob/63fcac725abeb841a101abfc64189edd66d1cf14/tile.cpp#L3191) rather than on pass 0 as usual.