Files
tippecanoe/TOUR.md
T
Claude d8b9555905 TOUR.md: describe the code as it is, not as it was
Drop the before-and-after framing. Where a passage explained a design by
contrasting it with an older one, it now just explains the design, and
where it noted that something was added or removed at some point, it
describes what is there.

The two places where the history was carrying real information keep the
information without the history: the radix sort's account of two past
bugs becomes a statement of the two invariants that are easy to break
and the reason --prefer-radix-sort exists to exercise them, and the note
that the per-feature minzoom field was once thought useless becomes a
description of what deciding it up front actually buys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012tkvf5HbM7yL8u3gSv9Vxq
2026-08-13 16:42:54 +00:00

81 KiB
Raw Blame History

A tour of Tippecanoe

Tippecanoe has what sounds like a simple job: It takes geographic objects in one file format (GeoJSON) and copies them into a different file format (Mapbox Vector Tiles). But it takes a lot of code to do that. What's really going on?

Even at the surface it is not quite that simple. The input can also be CSV, Geobuf, or FlatGeobuf, and the output can be an mbtiles file, a directory of tiles, or a PMTiles archive. And Tippecanoe is a family of programs, not one: tile-join, tippecanoe-overzoom, tippecanoe-decode, tippecanoe-json-tool, and tippecanoe-enumerate all operate on the tiles after the fact. But the core of it is this: read a lot of features, put them in an order that makes them easy to thin out, and then divide and conquer the world into tiles.

All the links below point at 63fcac72 (version 2.81.0), so the line numbers stay meaningful even as the code moves.

Starting up

The main function has the job of processing the list of options and files that the user provides. I'll go into more detail later about what all those options actually do, but the list of them is spelled out here in a single table.

Everything about the options comes from that one table, so that no two descriptions of them can disagree. The entries whose val is 0 and whose flag is null are not options at all but headings for the usage message, and strip_usage_headings() removes them before the table is handed to getopt_long(). getopt_string() derives the short-option string from it, and print_usage() prints the usage message from it. The other programs in the family use the same mechanism.

Most of the options set an entry in the additional[] or prevent[] arrays, indexed by the single-character codes listed in options.hpp, which is why so much of the code below reads like if (additional[A_DROP_DENSEST_AS_NEEDED]).

After the options are processed we reach the core operations: create the output tileset, read input into it, close it, and then, if the output name ended in .pmtiles, convert what was written into a PMTiles archive.

Creating the output tileset

Creating an mbtiles file is a series of sqlite operations, most of which I copied from mbutil. There is an option (-F) to keep going even if these operations fail, to support a geocoding use case that needed to add new, non-conflicting, tiles to an existing mbtiles file.

The schema is a map table that maps z/x/y to a content hash, an images table that maps a content hash (within a zoom level) to the tile data, and a tiles view that joins the two back together into what a consumer expects. The point of the indirection is that identical tiles — which are extremely common at low zooms, and in any tileset with a lot of empty ocean — are stored only once. mbtiles_write_tile hashes each tile's contents with fnv1a to find the key.

There are two other output forms. With -e, the output is a directory of tiles, z/x/y.pbf, with a metadata.json alongside. If the output name ends in .pmtiles, Tippecanoe writes an mbtiles file first and then repacks it into the single-file PMTiles format at the end, which is described in its own section below. From the tiling code's point of view all three are the same: there is an outdb or an outdir, and it writes tiles to whichever one it has.

Temporary files to read input into

Reading input is deceptively named because it also ultimately invokes the whole tiling process.

Its first job is to create as many sets of temporary files as your computer has CPUs, so that multiple threads can read input into these files at the same time without interfering with each other. At the end of the actual input stage, the per-thread temporary files will be merged or concatenated together so that tiling will operate on the combined results.

The temporary files are:

  • geom, which, as the name suggests, contains the feature geometry, plus other characteristics of each feature.
  • index, which is an index of the features in geom by their byte offset within the file and their quadkey-encoded location. It will be used later to merge the per-thread geom files back together.
  • pool, which contains one deduplicated instance of each property key or value that appears in any feature.
  • tree, which is a binary tree of pool entries, so that it can look up the existing pooled copy of any keys or values that are duplicated across multiple features, or add a new entry to the pool when a key or value is used for the first time.
  • vertex, which records, for --no-simplification-of-shared-nodes, each vertex of each line or ring together with the two vertices on either side of it, so that the points where features diverge from each other can be found globally.
  • node, which records the individual points that must not be simplified away.

Features reference their pooled keys and values directly, by their offsets into pool; there is no second level of indirection for features that have a lot of properties.

All of this data is in temporary files instead of in memory because it can be very large, and putting it on disk makes it possible to tile data that is too big to fit in memory. Processing will be much faster if the data does fit in memory, though. The pool and tree in particular use a hybrid: memfile keeps them as an in-memory std::string until they get too big, at which point memfile_full switches to appending to the real file. Note that data is explicitly copied in and out rather than accessed through a writable memory map, which has bad performance problems in containers.

What is straightforwardly in memory, not on disk, because it is small, is the "layermap" for each CPU, which is a list (by name) of the tile layers and all the property keys that have been used for any features in that layer, plus the sample values and ranges that will become the tileset's tilestats. This will eventually end up in the tileset metadata.

Input formats

Before reading anything, Tippecanoe makes up a layer name for each source that doesn't have one, from the last component of its filename with the recognized suffixes and any characters that can't appear in a selector trimmed off.

Then, for each source in turn, the format is chosen from the file suffix, or from the format given in a -L{"format":"…"} JSON layer specification:

GeoJSON input can also be gzipped: streamfdopen transparently wraps the file in a gzdopen if its name ends in .gz.

Whichever frontend reads the features, they all converge on the same place: they fill in a serial_feature and call serialize_feature to write it to the temporary files. Everything downstream of that point is format-independent.

Three ways to read GeoJSON

Now that the temporary files are in place, it is time to start reading input, either from stdin or from the files named on the command line. There are three ways that input files may be read, depending on whether the input is from a file or a stream and whether the user has asked for parallel input processing.

  1. If the --read-parallel option was specified, Tippecanoe first tries to map the file into memory. If this was successful, it creates several threads to parse parts of it in parallel as will be described below.

  2. If the file can't be mapped into memory (which means that it is an input stream, not a file on disk), Tippecanoe opens it as a stream. If the user has nevertheless asked for parallel processing of input, Tippecanoe starts reading the stream into another temporary file that it can map into memory. Once it has accumulated enough to be worth starting up several parser threads, it does that, and the main thread goes back to reading more streaming input into the next temporary file. This back-and-forth continues until the whole input has been consumed. It also waits for the parsers rather than reading indefinitely far ahead of them.

  3. If the user didn't ask for parallel input, it just runs a single JSON parser on the streaming input file.

Parallel reading is not strictly opt-in, though: if the stream begins with an ASCII record separator (0x1E), the input is GeoJSON Text Sequences, which does guarantee that features are separated, so Tippecanoe turns parallel parsing on by itself and splits on the separator instead of on newlines.

Parsing GeoJSON in parallel

To avoid massive digression I will first talk about parsing GeoJSON in parallel. The do_read_parallel function starts with the file already mapped into memory and knows how many CPUs can work on it simultaneously. It starts by making a guess that each CPU should have an equal-sized fraction of the file to work on, and then adjusts this guess by moving the divisions forward until each of them begins with a separator character (a newline, or a 0x1E if this is a GeoJSON text sequence).

(Note that GeoJSON itself makes no guarantee that feature objects are separated by newlines, or that newlines will never appear in the middle of a feature. This lack of a guarantee is why parallel parsing only happens if the user asks for it, or if the input announces itself as a text sequence, because it will misinterpret many correct GeoJSON inputs that are not designed for line-delimited use.)

It sets up a parameter block for each CPU pointing it to all the global settings and to the per-CPU temporary files and semi-global properties (bounding box, layermaps) and a JSON parser pointed to the portion of the file that was allocated to that CPU, and creates a thread to run the single-stream parsing described immediately below, from that JSON parser. Then it waits for all the threads to finish and is done.

Parsing a single GeoJSON stream

To parse GeoJSON, whether from a fraction of a memory-mapped file as described immediately above, or from a regular file stream, Tippecanoe uses a generic JSON pull-parser. By pull-parser, I mean that the parser returns JSON tokens and objects one at a time as they are encountered, rather than building up a full JSON object structure in memory for the entire input file and then presenting that to the caller.

The loop over those tokens lives in geojson-loop.cpp rather than in geojson.cpp, because tippecanoe-json-tool needs the same thing. It runs until it encounters either a bare geometry (an object whose type is Point, LineString, Polygon, MultiPoint, MultiLineString, or MultiPolygon and is not contained within something that could be a feature) or a feature (an object whose type is Feature). In either case it calls the add_feature method of the json_feature_action it was given, which for Tippecanoe proper is the one that calls serialize_geojson_feature.

The checks for whether something that looks like a geometry or a feature really is one are fussier than they first appear: an object is not a feature or geometry of its own if it appears inside another object's properties, and a GeometryCollection inside a feature is expanded into one feature per geometry sharing the same attributes.

From GeoJSON feature to internal feature

Turning a GeoJSON feature into a Tippecanoe feature starts with some checking: does the geometry have a coordinates array? Is the type one of the kinds of things that GeoJSON defines? It also checks the feature for a special tippecanoe object that can specify a per-feature minzoom, maxzoom, or layer name, and for the top-level feature id that is independent of the feature properties.

For each of the feature's properties it identifies the property's type and, if it is a compound JSON object that can't be natively represented as one of the attribute types that vector tiles supports, stringifies it. These are all just in memory for the moment, in the full_keys and full_values fields of the serial_feature. The keys go through a key_pool so that the many features that share an attribute name also share one copy of the string in memory.

It then calls parse_coordinates to recursively unpack the coordinates array from the geometry, and hands the whole thing to serialize_feature. See below for more about parsing and reprojecting geometry.

Internal representation of a feature

The serial_feature structure is the thing that every frontend produces and that every stage of tiling passes around. It is a big structure, but its fields divide into three groups, and the comments in the header say which is which:

  • The ones that are actually serialized to the geom file: the layer, segment, and sequence numbers; the geometry type t; the geometry itself; the feature id, if any; the per-feature tippecanoe_minzoom and tippecanoe_maxzoom, if any; the quadkey index; the extent (area, or a stand-in for it); the label_point; and the pooled keys and values.
  • The ones that only exist during initial serialization: full_keys and full_values, the string forms of the attributes, which get replaced by the pooled keys and values offsets.
  • The ones that only exist during tiling: the bounding box, the dropped state, whether the feature is polygon dust or was coalesced, the current detail and simplification level, and so on.

dropped deserves a note, because despite the name it is not a boolean. It is FEATURE_DROPPED (-1), FEATURE_KEPT (0), a positive sequence number meaning "this is the nth extra feature retained by --retain-points-multiplier alongside its cluster's lead feature," or INT_MAX meaning "retained by --preserve-multiplier-density-threshold." Much of the dropping logic in write_tile is really about keeping these multiplier clusters intact.

Attribute values are represented by serial_val: a type tag (one of the mvt_value types) plus the value as a string. Every number, integer or floating point, is mvt_double here and is stored in its stringified form, which is how integers too large for a double survive with their original precision.

Writing a feature to the temporary files

serialize_feature is where a feature from any frontend actually becomes bytes on disk, and it is where a surprising number of the command line options take effect. In order, it:

All that remains is to increase the per-CPU bounding box to encompass the feature bounding box if necessary, and increment the user-visible progress indicator.

The string pool

addpool is worth its own section, because it is on the hot path for every attribute of every feature and it has been tuned accordingly. It has three layers of defense against being slow:

  • A direct-mapped hash cache in front of the tree, per reader thread, which catches the overwhelmingly common case of the same handful of keys and values recurring over and over. (It verifies the type and the string on a hit, because a hash cache that trusted the hash would collide.)
  • A depth limit on the tree search. If a string is so deep in the tree that the search is getting expensive, it is probably unique anyway, so it is appended to the pool without being added to the tree, and future copies of it just won't be deduplicated.
  • A size limit. Once the pool and tree together exceed a fraction of physical memory, the pool switches over to appending to the file, and stops maintaining the tree at all. Deduplication is a nice-to-have; thrashing is not.

The comparison function, swizzlecmp, orders by hash first and only compares strings when the hashes match, which keeps the tree from degenerating into a list when the input is sorted.

Serialized representation of a feature

Serializing each feature probably looks familiar if you have done any serialization with protozero. There is not much sanity-checking so you have to be careful if you make changes, especially in the field that does some bit packing to combine the layer number and the flags for the presence of a label point, index, extent, id, minzoom, and maxzoom into the same number.

Serialization goes into a std::string and the caller writes that out, rather than going straight to a file. That is what makes it possible for the same function to serve both the frontends, writing to the per-thread geom files, and rewrite during tiling, writing compressed into the next zoom level's files.

All of the primitive serialization types (integers of various sizes and signednesses) are spelled out earlier in the same serial.cpp file. They use protozero's zigzag conversion functions to represent signed numbers as unsigned, and use its same variable-length encoding for numbers, with the high bit of each byte indicating that there are more bytes to follow.

Two pieces of magic are worth knowing about, because both are commented as MAGIC in the source and both will bite you if you forget them. The first is that the feature minzoom is the last byte of the serialized feature, so that the reordering pass can rewrite it in place without decoding anything. The second is that deserialize_feature checks that it consumed exactly as many bytes as the length prefix promised, which is the check that catches it when the two halves get out of step.

Internal representation of geometry

Parsing the GeoJSON coordinates array into Tippecanoe's internal representation of geometry, as mentioned above, deserves a little bit more detail.

Tippecanoe's internal geometry structure is called draw, and is composed of an operation (moveto, lineto, closepath) and an x and y coordinate, plus an auxiliary field used during line simplification to track whether a point has been determined to be necessary. A series of draw operations is called a drawvec. The draw structure uses some bitfields to keep its size down to 128 bits even with the op and necessary fields, to avoid eating a lot of unnecessary memory.

The parse_coordinates function mostly just builds up a drawvec from the GeoJSON coordinate arrays for the feature, after doing some sanity checking to make sure the arrays contain numbers and an appropriate number of elements. It calls itself recursively to parse the multiple points that appear within each LineString, the multiple rings that appear within each Polygon, and so on.

It also contains the call that reprojects the coordinates from their original projection (usually WGS84) to tile coordinates, where the world is a square, 2^32 units on each side.

There is also a weird special case that adds a closepath operation after the last ring of each Polygon, so that the following ring can be identified as an outer ring. (Tippecanoe does not otherwise use closepath, except in the last stages of writing out the vector tile format, which requires it as part of each polygon ring.) The fix_polygon function uses this special tag when it reorders the points in each ring to make sure that inner rings are oriented the opposite direction from outer rings. It also makes sure the ring is closed, with the last point duplicating the first, if it wasn't already.

Separately, in serialize_feature, there is a special case (to save space in the temporary files) that shifts each of these world coordinates down by geometry_scale, which was set back in the main function to correspond to the global minzoom, so that coordinates on disk are not stored in more detail than will eventually be used in the final tiles. The reduced coordinates will be shifted back up to world scale during each deserialization during tiling. The shift rounds rather than truncating, which matters for overzooming, and adds a COORD_OFFSET first so that features slightly off the edge of the world still shift correctly.

Merging temporary files after parsing

At this point Tippecanoe has all of the input from all of its input files parsed and copied into temporary files, but each layer's features are scattered across multiple files, depending on which CPU happened to read the feature. The files have to be consolidated before tiling can begin.

But it is important to note that back when the temporary files were being opened, each of them was opened twice, and then deleted. Each of them was opened as an integer file descriptor with mkstemp, and as a C FILE with fopen, and deleted with unlink. The two open file descriptors to each file cause it to continue to exist on disk even though it has been deleted from the temporary directory. Deleting the file before it is written should also serve as a hint to the operating system that there is no need ever to write the file's contents to physical disk if it is small enough to continue to fit in memory, in the operating system's buffer cache. (Using a deleted file for temporary storage is a traditional Unix idiom, and there is a tmpfile call in the standard library to create and delete a temporary file if you only need a single FILE reference to it.)

So now the FILE for each is closed, so each temporary file is kept alive by only one file descriptor, not two, and any buffering that was being done inside the Tippecanoe process itself, as opposed to in the operating system, will have been flushed out.

Merging the string pool is the easiest. The tree files are of no more use and are simply closed and discarded. The pool files for each CPU are concatenated together into a new file, with an array to keep track of the offset into the new file where the data for each CPU begins. (Which thread some data originally came from is frequently referred to in the code as its segment.) Because the pool may be partly in memory and partly on disk by this point, the merge has to handle both cases. The combined pool is then mapped back into memory so it can be accessed like an array.

Then, if --no-simplification-of-shared-nodes was requested, there are two more merges:

  • The vertices are sorted with fqsort — a quicksort that partitions into temporary files when the data doesn't fit in memory — and then scanned. Any middle vertex that appears with different neighbors in two different records is a place where features diverge, so it becomes a node.
  • The nodes are sorted and deduplicated the same way. The result is mapped into memory so that each tile can binary-search it, and, because that search would otherwise be a lot of cache misses, a 34-megabyte Bloom filter is built alongside it so that the common "this point is not a shared node" answer can be given without touching the array at all. The nodes are keyed by quadkey so that the nodes for any one tile are adjacent in memory.

Doing this globally, once, is what keeps the memory in bounds: the alternative of assembling the node list separately within each tile is correct but ruinously expensive on large inputs.

Why sort the features?

The geometry is going to be much more of a pain to merge, not just because it is typically so much bigger than the other files, but because we also want it in a specific order. In particular we want it to be ordered by quadkey so that

  1. when the low zooms are thinned out by dropping some fraction of features, those features are evenly distributed by location rather than clumpy,
  2. it is possible to make global estimates of the number of features in each tile, before actually generating those tiles, and
  3. within each tile, the density of features near any particular feature can be estimated by the difference between its quadkey index and the numerically adjacent feature's quadkey index.

These characteristics will be used by the -r dot dropping rate, the -Bg guessing of an appropriate base zoom level, and the -g thinning of dense features, respectively. The maxzoom guess of -zg uses the same ordering for the same reason.

We can accomplish this by sorting the index by quadkey and then copying features from the original files into a new temporary file in index order. This is the user-visible "Reordering geometry" phase.

Sorting and merging

The index itself may actually be uncomfortably large to sort. Each index entry takes 32 bytes of storage, so on a 4GB laptop, 100 million features will exhaust available memory and grind progress to a halt, sometimes crashing the system in the process. And recopying the geometry in index order after sorting may be even worse, because the original order of the geometry probably has no relationship to the index order, so if it does not fit in memory, retrieving each feature may require a disk access.

Tippecanoe's strategy, then, is to do as much of the reordering as possible in a streaming way, without taking advantage of random access. This is an old, old technique, like people would have used in the 1970s with a computer that had almost no memory but had a lot of tape decks that could read and write at full speed as long as they were going in order. It still works the same way if you read a stream from one file and split it into streams into several output files, because they can all go as fast as the physical disk I/O can go.

The top-level sort is a radix sort. It checks how many files it can have open at once, and if the answer is, say, 512, then it tries to split the index and geometry into 128 parts (512/4), based on the leading bits of each feature's quadkey.

It then goes through each of those parts and checks whether it would fit in memory. If it would, it splits that portion of the index up again by however many CPUs it has to work with, sorts each of those sub-indices in its own thread with the system qsort, and then merges those sub-sub-indices to copy the sub-geometry to the final geometry file in index order, which should be fast because it's all in memory. If one of the parts didn't fit in memory, then it does another radix sort to split it up further, based on the next bits of the quadkey. In at most a few passes of this it will have generated final sorted index and geometry files in quadkey order.

Two invariants here are easy to break and hard to notice. The first is that a bucket small enough to be written out directly, rather than through the merge, must still be written one byte shorter than the index says and then given a fresh minzoom byte, exactly as merge() does — the MAGIC byte again. Write the original byte as well as the new one and every feature read after it comes out shifted by one. The second is that each level of subdivision must consume at least one bit of the index, which means at least two buckets: with only one, the recursion never reaches the prefix width that stops it, and the shift that selects a bucket is by the full width of the index, which is undefined.

Both invariants only matter on inputs large enough to recurse, which is more data than a test wants to handle, so --prefer-radix-sort exists to reach them: it pretends there are only 8 KB of memory so that a tiny test input goes down the deep path, and the test checks the radix-sorted output against the in-memory sort of the same data.

Guessing maxzoom

Once the index is in order, Tippecanoe can guess a maxzoom from the data, if -zg was given instead of an explicit -z.

The idea is that a good maxzoom is one where most features are distinguishable from each other, so it walks the sorted index and accumulates the mean and standard deviation of the log of the gaps between adjacent quadkeys, using Welford's algorithm. Distances between features are typically lognormally distributed, so the geometric mean is the right average, and 1.5 standard deviations below the mean is taken as the distance at which features should still be distinguishable. The conversion from quadkey gaps to feet is an empirical fit; the #if 0 block just above the calculation is the code that produced the data for it.

Several things adjust the result afterward:

  • The distances within each feature, sampled back in serialize_feature, can push the maxzoom higher, because a tileset of detailed coastlines needs resolution even if there are only a few features.
  • Duplicate feature locations push it higher too, since a pile of features at exactly the same point will never be distinguishable but will be dropped at the drop rate.
  • Clustering pushes it higher, so that features eventually become unclustered.
  • A 2-million-tile budget pulls it back down, estimated from the total polygon area, so that a request that would spend a week filling in the interiors of polygons doesn't.
  • --smallest-maximum-zoom-guess sets a floor, and the detail limits set a ceiling.

The drop rate can be guessed at the same time, from a curve fitted by eye to the drop rates that looked right for a handful of real point tilesets: evenly spaced features want a large drop rate, clumpy features can get away with a small one.

Guessing base zoom and drop rate

Independently of -zg, Tippecanoe can calculate a base zoom and drop rate if requested to do so with -Bg or -rg.

It can do this because all the features that are (centered in) the same tile are now guaranteed to be contiguous in the index because it is ordered by quadkey. So Tippecanoe can calculate the densest tile in each zoom level just by running through the index in order, accumulating one count of features for each zoom level, and resetting the counter every time the tile number at that zoom level changes.

Now that it knows how many features are in the densest tile at each zoom level, it can calculate what the lowest zoom level is where the densest tile contains an acceptable number of features and what fraction of features must be dropped at lower zoom levels to make the densest tile acceptable at every zoom level. If no base zoom at or below the maxzoom will do, it works from the other direction instead, choosing a drop rate first and then the base zoom that rate implies.

Precalculating the dot dropping

The per-feature minzoom field is what makes dot-dropping consistent from one zoom to the next, which is what keeps points from popping in and out as you zoom. Deciding it up front, for all the features at once, is what makes that consistency possible: no tile has to work out for itself which features it is entitled to show.

calc_feature_minzoom assigns each feature the lowest zoom at which it will appear, by keeping a running counter per zoom level that is decremented by that zoom's interval (droprate^(basezoom - z)) for each feature that goes by. It is called from inside the merge, as the features are being copied into index order, so it sees them in quadkey order and therefore spreads its choices evenly over space rather than picking a random subset. --preserve-point-density-threshold adds an override: if a feature was assigned to a high zoom but is nevertheless very far from the last feature chosen for some low zoom, it gets pushed out at that low zoom anyway, so that sparse areas of the map don't go blank.

If the base zoom or drop rate had to be guessed, they weren't known yet when the merge ran, so there is a fixup pass that goes back and rewrites all the minzoom bytes in place, which is possible precisely because that byte is the last byte of each feature and its position is recorded in the index. --drop-denser is implemented in the same pass, by sorting a sample of features by the gap to their predecessor and assigning zooms so that the sparsest ones appear first.

Running through the zoom levels

Now it is finally time for tiling to begin. All that will remain after tiling is complete is to calculate the final bounding box and layer metadata and write it to the tileset.

The outer loop of tiling starts by running through the possible zoom levels. It normally starts at zoom 0 even if some higher minzoom was specified, because the entire geometry is now in a single file, not at all split up by tile. If we tried to use the same simple index-based technique as above to split it up into minzoom tiles, we would miss some features that are big enough to span (or be buffered into) multiple tiles. Instead, we must "divide and conquer" the low zooms to get to the high zooms correctly, even if some of them will not ultimately be written to the tileset.

The one exception is that choose_first_zoom checks whether the bounding box of all the features fits within a single tile at some higher zoom, and if so starts there, since dividing and conquering an empty world is a waste of time.

The basic strategy of tiling is that each zoom level will read through the current set of temporary files and will write out both a set of vector tiles into the tileset and a new set of internal temporary files for the next zoom level. For example, the initial single geometry file will produce the single zoom level 0 tile 0/0/0 as well as the four temporary files that will be used at the next step for tiles 1/0/0, 1/0/1, 1/1/0, and 1/1/1. Alternately, if the minzoom was 2, and we don't care about those zoom level 1 tiles at all, we can declare that the "next" zoom past 0 is 2, and the zoom level 0 processing can produce the right temporary files for tiles 2/0/0, 2/0/1, 2/0/2, 2/0/3, 2/1/0, and so on, and never do any work at all for zoom level 1.

Where it gets complicated is that we want Tippecanoe to be able to do as many of these things at the same time as possible, to take advantage of multiple CPUs. The zoom level 1 processing should be able to vectorize and subdivide all four of the zoom level 1 tiles at the same time rather than sequentially. But the ability to do this is limited by the number of CPUs, the number of files that can be open at one time, and the number of below-minzoom levels we are trying to skip over. Each CPU is going to be reading from one file and producing 4 (or 16, or 64, or maybe even 256, if we are skipping three zoom levels) new files as it tiles. Two CPUs can't write to the same output file without risking stomping on each other's work.

So traverse_zooms starts out by making a new set of temporary files for the next zoom level and calculating how many threads there are to split those files among. Then it takes all the current set of temporary files that need to be processed and allocates them approximately evenly across the threads. (The assumption is that processing time will be proportional to file size, which may not be exactly right, but is reasonably close.)

Note also the compressor wrapped around each of those output files. The feature stream for each tile in the temporary files is zlib-compressed, which reduces both the disk space used and the I/O latency, at the cost of not being able to seek within a tile's data except by starting over. (Zoom 0's input is the exception: it is the file the sort produced, and it has to stay uncompressed so that the minzoom fixup can rewrite bytes in it.)

Once the input files and output files have been allocated to threads, it is time to make a new set of parameter blocks to tell each thread what files it is consuming and what other files it is producing, and to start the thread to do the work.

Each thread, then, starts going through its list of input files. From each one, it reads a z/x/y tile number and then calls the badly-named write_tile to do the work of processing that tile and writing out its children to that thread's set of output files. This loop continues until the thread has exhausted all the tiles in all its input files. The loop also does a little bit of work to track the densest tile at maxzoom which will be used later for the map center in the tileset metadata.

Each tile's data in the temporary file is preceded by an uncompressed estimated_complexity, the number of bytes of feature data that tile contains. It can only be known after the data has been written, so a placeholder is written first and then rewritten with pwrite once the stream is finished (and likewise for the initial zoom-0 file). It is used by variable-depth pyramids, described below.

Retrying a whole zoom level

There is a loop around the whole zoom level, not just around each tile: for (size_t pass = 0;; pass++). If any tile in the zoom found that it had to raise one of the as-needed thresholds (the minimum gap, the minimum extent, the minimum drop sequence, the minimum attribute value, or gamma) to make itself fit, that new threshold is collected from all the threads, the zoom level is erased from the output, and the whole zoom is tiled again with the higher threshold. This is why dropping is consistent across a zoom level rather than varying tile by tile.

Note that the tiles are written before it is known whether they will have to be thrown away, rather than the zoom being preflighted to find its thresholds first. That is the cheaper arrangement, because in the common case nothing has to be thrown away at all.

--extend-zooms-if-still-dropping is handled here too: if the maxzoom is still dropping features, the maxzoom goes up by one and tiling continues.

The work of each tile

The first task of write_tile is to figure out whether it is actually subdividing the tile that it has been tasked with into children, grandchildren, great-grandchildren, or whatever, based on the minzoom and the number of output files it has to work with.

It then moves on to a loop that usually only runs once, from the usual tile detail down to the minimum acceptable detail. In most cases this loop will be short-circuited at the end. It will only run multiple times if the tile is too big and Tippecanoe must try again — either at a higher dropping threshold or, as a last resort, at a lower resolution. If this happens it will also have to rewind its input file back to its original position, and restart the decompressor, so that it can reprocess the data.

Then it reads features one at a time from next_feature, which does the opposite of the feature serialization described above, and rather more besides. For each feature it:

  • Fills in the gap to the previous feature, the squared planar distance that --drop-densest-as-needed and its relatives sort on. This is computed once, at zoom 0, and then carried along in the serialized feature, so that every zoom has the same idea of which features are in dense company.
  • Clips the feature to the tile plus its buffer. decode_geometry has already both undone the geometry_scale bit shifting and adjusted the offset of the geometry so that (0,0) is at the top-left corner of the tile instead of the top-left corner of the earth. Both of these will be undone again when the geometry is reserialized into the child data for the next zoom level. At zoom 0 there is a special case that duplicates features near the antimeridian onto both sides of the world.
  • Writes the feature out to the child tiles via rewrite — but only on the first pass through the detail loop, since retries must not write the children again. rewrite works out from the feature's bounding box which children it can touch, and picks the shard for each child by interleaving the low bits of x and y, being careful that all the data for any one child tile stays contiguous within one shard.
  • Applies the per-feature tippecanoe.minzoom and maxzoom and the -j feature filter. Note the ordering: the child tiles are written before these tests, so that a feature excluded from this zoom still reaches the zooms where it belongs.
  • Decides whether the feature is dropped by rate at this zoom, from the precalculated feature_minzoom, and if it is dropped, whether it should nevertheless be retained as one of the extra features of a --retain-points-multiplier cluster. The comparison uses a fractional zoom level derived from the bit-reversed index, so that the multiplier can target a specific number of features rather than only the powers of the drop rate.
  • Removes null attributes — after the filter has run, since a filter may want to test for them.

Back in write_tile, each feature that comes back is assigned to its layer and then run through the gauntlet of size-reduction strategies, in a long else if chain: gamma, -K clustering, --drop-densest-as-needed, --cluster-densest-as-needed, --coalesce-densest-as-needed, --drop-smallest-as-needed, --coalesce-smallest-as-needed, --drop-fraction-as-needed, --coalesce-fraction-as-needed, and --drop-by-attribute-as-needed. They all have the same shape: sample the relevant quantity into a vector for later, compare it against the threshold for this pass, and either keep the feature or find an already-kept feature to accumulate its attributes onto. Only the lead feature of a multiplier cluster can be dropped this way, but if it is, it drags the rest of its cluster with it.

Two more things happen per feature before it is accepted:

  • Polygons that are too small to be worth drawing at this zoom are reduced to dust by reduce_tiny_poly, which accumulates their area and occasionally emits a placeholder square so that the area is still somehow represented.
  • If the tile is already hopeless — more features than could possibly fit even at one byte each — Tippecanoe stops adding features and just counts how many it skipped, so that it can extrapolate what the real size would have been without spending the memory to find out exactly. The size and feature limits are scaled up in proportion to the number of multiplier features being carried alongside the lead features, since those are expected to be filtered out by the consumer.

The rest of tiling

Once all the features have been read, the work becomes per-layer rather than per-feature. In order:

Then the layers are converted into an mvt_tile, with the pooled keys and values turned back into tile attributes by decode_meta, the postfilter is run if there is one, and the tile is encoded and compressed.

And now the size checks, which are the reason for the whole loop. If the tile has too many features or too many bytes — extrapolated upward if features were skipped — then Tippecanoe picks the strategy it was told to use and raises that strategy's threshold, aiming at a fraction of the current size, and goes around again. The functions that choose the new threshold (choose_mingap and its siblings) work on the samples that were collected on the way through. If the threshold can't be raised any further, that is an error: there is nothing left to try, and looping forever is worse than failing. If no dropping strategy was requested at all, the detail is reduced by one and the tile is tried again from the top.

If the tile does fit, it is written to the database or directory under the db_lock, the strategies used are folded into the zoom's statistics, and write_tile returns.

Variable-depth tile pyramids

--generate-variable-depth-tile-pyramid cuts across everything above, so it gets its own section.

The idea is that a tile in an empty or simple part of the map doesn't need children: if all of its features can be included at full precision, a client can overzoom that tile instead of downloading deeper ones. So when it is enabled, write_tile makes an extra first attempt at a much higher detail (30 - z) with simplification turned off, and only bothers to do so if the estimated_complexity recorded with the tile's data suggests it might work. If everything fits at that detail and nothing was dropped, the tile is a leaf: it is added to skip_children and its descendants are skipped rather than written at the next zoom.

The bookkeeping around that is fiddlier than it sounds:

  • The children's geometry is still in the stream even for a skipped tile, so skip_tile has to read past it rather than seek.
  • A zoom that later has to start dropping features invalidates the truncation, because a leaf is only legitimate if it is complete. Those revived tiles have never written their children, so they have to do it on the pass where the dropping started rather than on pass 0 as usual.
  • A feature with an explicit tippecanoe.minzoom deeper than the current zoom blocks leafing, because it isn't in this tile and its children would never be generated, so it would appear at no zoom at all.
  • The tileset's reported maxzoom is the deepest zoom actually written, not the requested one.

Plugins

Tippecanoe can pipe each tile's contents through a shell command, either before tiling (-C, the prefilter) or after (-c, the postfilter). In both cases the interchange format is newline-delimited GeoJSON, one feature per line, so the filter can be anything from grep to a program of your own.

setup_filter does the plumbing: two pipes, a fork, and an execlp of sh -c with the tile's z, x, and y as $1, $2, and $3.

The prefilter is the more interesting of the two, because it has to interpose itself in the middle of the feature-reading loop. run_prefilter runs in its own thread — a real thread, not just a subprocess, because Tippecanoe needs to write to and read from the filter at the same time and would otherwise deadlock — calling next_feature and writing each feature out as GeoJSON in world coordinates. Meanwhile the main loop in write_tile reads the filter's output with parse_feature instead of calling next_feature itself. Note that the features handed to the prefilter have already been clipped to the tile and already had their dot-dropping decision made — whether the feature was dropped is passed through as a top-level dropped member of each GeoJSON object, alongside its index, sequence, and extent — so the filter sees what the tile would have contained.

The postfilter is simpler, because by the time it runs the tile is a finished set of mvt_layers: filter_layers writes them out as GeoJSON from a writer thread and parses the result back into layers, updating the layermap and tilestats from whatever came out, since the filter may have invented layers and attributes that were never in the input.

Both filters are skipped below the minzoom, since those tiles are only being generated to subdivide their way toward real ones.

Feature filters

Filtering through a subprocess is flexible but slow. For the common cases there is -j, which takes a filter expression and applies it in-process. The syntax is the Mapbox GL Style Specification filter syntax, extended with $type, $id, and $zoom pseudo-attributes and with an attribute-filter form that removes individual attributes rather than whole features.

eval is a straightforward recursive interpreter over the parsed JSON of the expression. It also understands a second, Felt-style expression dialect, and --unidecode-data lets string comparisons match transliterated forms so that a filter for "Zurich" can find "Zürich".

The same evaluator is used by tile-join and tippecanoe-overzoom, which is why it lives in its own file.

Writing tileset metadata

When tiling is done, read_input computes the tileset's bounds and center and writes the metadata.

The center is the center of the densest tile at maxzoom, which traverse_zooms tracked, clamped to lie within the bounds. The bounds themselves come in two forms: the ordinary one, and an antimeridian-aware one that picks whichever of the two world planes tracked during input gives the narrower box, so that a tileset covering Fiji doesn't claim to span the entire globe.

The per-thread layermaps are merged — this is where the differing per-thread layer numbering is finally reconciled, by name — and make_metadata assembles everything into a metadata struct, which is then written either into the sqlite metadata table or into a metadata.json in the output directory. Having the struct in the middle is what keeps the two output forms from drifting apart, and lets tile-join and the PMTiles writer reuse the same code.

The fields are:

  • name, description, version, type, format, minzoom, maxzoom, bounds, center, and attribution, which are the standard mbtiles metadata.
  • antimeridian_adjusted_bounds, the alternative bounds described above.
  • generator and generator_options, which record the Tippecanoe version and the entire command line, so that a tileset can say how it was made.
  • strategies, an array with one entry per zoom level recording how many features were dropped by rate, by gamma, or as-needed; how many were coalesced; how many tiles had their detail reduced; how many tiny polygons there were; how many zooms were truncated by variable-depth pyramids; and what the largest tile would have been if nothing had been dropped. If you want to know whether your tileset is losing data, this is where to look.
  • tippecanoe_decisions, which records the basezoom, drop rate, and multiplier that were actually used, which matters when they were guessed rather than specified.
  • json, a string containing two things: vector_layers, the per-layer list of attributes and their types that style editors read, and tilestats, the more detailed statistics — counts, minimums, maximums, and sample values for every attribute of every layer. The tilestats are bounded by --tile-stats-attributes-limit and friends, because for a source with hundreds of millions of features the statistics can otherwise get quite large.

Finally, mbtiles_close runs ANALYZE and closes the database.

Converting to PMTiles

If the output name ended in .pmtiles, everything above has been writing an ordinary mbtiles file, and mbtiles_map_image_to_pmtiles now converts it. It reads the tiles back out in tile-id order, writes them into the single-file PMTiles layout, and translates the metadata into the PMTiles flavor of JSON metadata. Because the mbtiles schema already deduplicates tiles by content hash, the PMTiles directory can carry that deduplication straight through.

The other programs

Tippecanoe proper is only part of what gets built. The rest of the tools in the Makefile share most of their code with it:

tile-join merges tilesets, joins attributes onto their features from a CSV or a sqlite database, filters layers and attributes, and can overzoom to a deeper maxzoom on the way. Its main loop reads the inputs in parallel through a set of readers, groups the tiles by z/x/y, hands each group to a worker thread that appends each source tile's layers into one output tile, and reassembles the metadata from the inputs'. Because it works entirely on already-tiled data, it can do in a few minutes what re-tiling from source would take hours to do.

tippecanoe-overzoom takes one or more source tiles and produces a single deeper tile from them, doing the clipping, the multiplier de-duplication, the attribute accumulation, the filtering, the simplification, and optionally the binning of points into a supplied set of polygons. The overzoom function itself lives in clip.cpp because tile-join uses it too. It converts each source feature's geometry back to world coordinates, offsets it into the destination tile, and then runs it through the same clipping code that tiling uses.

tippecanoe-decode turns tiles back into GeoJSON, which is indispensable for figuring out what actually ended up in a tileset. --stats reports the size and feature count of each layer instead of the contents.

tippecanoe-json-tool does GeoJSON manipulation outside of tiling: sorting features by an attribute (so that tile-join-style joins can be done with sort and join) and joining CSV attributes onto features.

tippecanoe-enumerate lists the z/x/y of every tile in a tileset, one per line, which is the input to shell pipelines that want to do something to each tile.

Tests

The test suite is mostly of one kind: run Tippecanoe over a small input in tests/ with some set of options, decode the result, and compare it against a checked-in expected output. This makes it easy to see the effect of a change — the diff in the expected outputs is the change — but it also means that a change to feature ordering shows up as an enormous diff, which is why a few comments in the code apologize for not fixing something because it would churn all the fixtures.

There is also unit.cpp, a small set of Catch unit tests for the pieces that are awkward to test end-to-end.