mirror of
https://github.com/felt/tippecanoe.git
synced 2026-10-02 16:35:40 +02:00
The \uXXXX decoder tested `ch < 0xFFFF` before taking the three-byte UTF-8 path, so U+FFFF itself fell through to the four-byte branch and came out as F0 8F BF BF -- an overlong, and therefore invalid, encoding of a code point that fits in three bytes. check_utf8() only checks that continuation bytes look like continuation bytes, not that a sequence is the shortest form, so nothing downstream noticed: a GeoJSON attribute containing U+FFFF put invalid UTF-8 into the output tile, where a strict consumer would reject it. Since `ch` is parsed from exactly four hex digits it cannot exceed 0xFFFF on its own, so after this change the four-byte branch is reached only for a code point assembled from a surrogate pair, which is the only way to name one above the BMP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KNxyHKasyWrWcvre2yK4r