Add a test for joining with tippecanoe-json-tool

This commit is contained in:
Eric Fischer
2017-11-17 14:05:37 -08:00
parent 2b1cba0b53
commit 3f54a70459
4 changed files with 345 additions and 2 deletions
+74
View File
@@ -708,3 +708,77 @@ resolutions.
.IP \(bu 2
\fB\fC\-f\fR or \fB\fC\-\-force\fR: Decode tiles even if polygon ring order or closure problems are detected
.RE
.SH tippecanoe\-json\-tool
.PP
Extracts GeoJSON features or standalone geometries as line\-delimited JSON objects from a larger JSON file,
following the same extraction rules that Tippecanoe uses when parsing JSON.
.PP
.RS
.nf
tippecanoe\-json\-tool file.json [... file.json]
.fi
.RE
.PP
Optionally also wraps them in a FeatureCollection or GeometryCollection as appropriate.
.PP
Optionally extracts an attribute from the GeoJSON \fB\fCproperties\fR for sorting.
.PP
Optionally joins a sorted CSV of new attributes to a sorted GeoJSON file.
.PP
The reason for requiring sorting is so that it is possible to work on CSV and GeoJSON files that are larger
than can comfortably fit in memory by streaming through them in parallel, in the same way that the Unix
\fB\fCjoin\fR command does. The Unix \fB\fCsort\fR command can be used to sort large files to prepare them for joining.
.PP
The sorting interface is weird, and future version of \fB\fCtippecanoe\-json\-tool\fR will replace it with
something better.
.SS Options
.RS
.IP \(bu 2
\fB\fC\-w\fR or \fB\fC\-\-wrap\fR: Add the FeatureCollection or GeometryCollection wrapper.
.IP \(bu 2
\fB\fC\-e\fR \fIattribute\fP or \fB\fC\-\-extract=\fR\fIattribute\fP: Extract the named attribute as a prefix to each feature.
The formatting makes excessive use of \fB\fC\\u\fR quoting so that it follows JSON string rules but will still
be sorted correctly by tools that just do ASCII comparisons.
.IP \(bu 2
\fB\fC\-c\fR \fIfile.csv\fP or \fB\fC\-\-csv=\fR\fIfile.csv\fP: Join properties from the named sorted CSV file, using its first column as the join key. Geometries will be passed through even if they do not match the CSV; CSV lines that do not match a geometry will be discarded.
.RE
.SS Example
.PP
Join Census LEHD (Longitudinal Employer\-Household Dynamics \[la]https://lehd.ces.census.gov/\[ra]) employment data to a file of Census block geography
for Tippecanoe County, Indiana.
.PP
Download Census block geometry, and convert to GeoJSON:
.PP
.RS
.nf
$ curl \-L \-O https://www2.census.gov/geo/tiger/TIGER2010/TABBLOCK/2010/tl_2010_18157_tabblock10.zip
$ unzip tl_2010_18157_tabblock10.zip
$ ogr2ogr \-f GeoJSON tl_2010_18157_tabblock10.json tl_2010_18157_tabblock10.shp
.fi
.RE
.PP
Download Indiana employment data, and fix name of join key in header
.PP
.RS
.nf
$ curl \-L \-O https://lehd.ces.census.gov/data/lodes/LODES7/in/wac/in_wac_S000_JT00_2015.csv.gz
$ gzip \-dc in_wac_S000_JT00_2015.csv.gz | sed '1s/w_geocode/GEOID10/' > in_wac_S000_JT00_2015.csv
.fi
.RE
.PP
Sort GeoJSON block geometry so it is ordered by block ID. If you don't do this, you will get a
"GeoJSON file is out of sort" error.
.PP
.RS
.nf
$ tippecanoe\-json\-tool \-e GEOID10 tl_2010_18157_tabblock10.json | LC_ALL=C sort > tl_2010_18157_tabblock10.sort.json
.fi
.RE
.PP
Join block geometries to employment properties:
.PP
.RS
.nf
$ tippecanoe\-json\-tool \-c in_wac_S000_JT00_2015.csv tl_2010_18157_tabblock10.sort.json > blocks\-wac.json
.fi
.RE