Cheap perf wins in jsonpull C++ port

Profiling tl_2022_us_county.json (sample(1) on Apple Silicon) showed
~38% of parse time in allocator work and ~14% in std::string::push_back
during string-token construction. These changes target the low-hanging
fruit from that profile:

- Pre-reserve 2 slots in json_array and 4 slots in json_hash so
  coordinate `[x, y]` pairs and typical GeoJSON property maps avoid
  the 0 -> 1 -> 2 -> 4 vector-growth chain (and the shared_ptr copies
  it incurs).
- Reuse a parser-wide std::string buffer for JSON_STRING tokens
  instead of constructing a fresh local std::string per token. The
  buffer is cleared (capacity preserved) at the start of each token
  and copied into the final json_string, so once it has grown to the
  longest string seen it stops reallocating entirely.
- std::move the freshly-created container shared_ptr into the parser
  container stack in the `[` and `{` handlers, and move it out of the
  frame on the matching `]` / `}`. Each move skips one atomic
  inc/dec round-trip per container open and close.

On a tl_2022_us_county.json benchmark (4-iter user-time mean, Apple
Silicon, /usr/bin/time):
- main baseline:                              ~8.17s
- jsonpull-cpp before these changes:          ~10.90s  (+33%)
- jsonpull-cpp with these changes:            ~9.33s   (+14%)

So this commit recovers roughly half of the post-port regression.
The remaining gap is dominated by shared_ptr atomic refcount traffic
on the parse tree and per-node heap allocations, which would require
the larger unique_ptr/arena reworks to address.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Erica Fischer
2026-05-30 18:28:11 -07:00
co-authored by Cursor
parent 29b5dff333
commit 3f526ccad7
2 changed files with 44 additions and 10 deletions
+23 -6
View File
@@ -230,7 +230,10 @@ again:
if (o == nullptr) {
return nullptr;
}
j->container_stack.push_back({o, JSON_ITEM});
// add_object already installed `o` in the parent (or the
// parser's root); moving the local copy into the frame
// avoids one shared_ptr atomic inc/dec pair per container.
j->container_stack.push_back({std::move(o), JSON_ITEM});
if (cb != nullptr) {
cb(JSON_ARRAY, j, state);
@@ -259,7 +262,9 @@ again:
}
}
json_object_ptr ret = f->container;
// Move the container out of the frame so pop_back doesn't
// drop the last reference; saves one atomic inc/dec.
json_object_ptr ret = std::move(f->container);
j->container_stack.pop_back();
return ret;
}
@@ -271,7 +276,9 @@ again:
if (o == nullptr) {
return nullptr;
}
j->container_stack.push_back({o, JSON_KEY});
// See the [ case above: move into the frame to skip a
// shared_ptr atomic inc/dec round-trip.
j->container_stack.push_back({std::move(o), JSON_KEY});
if (cb != nullptr) {
cb(JSON_HASH, j, state);
@@ -300,7 +307,8 @@ again:
}
}
json_object_ptr ret = f->container;
// See the ] case: move out to skip an atomic refcount round-trip.
json_object_ptr ret = std::move(f->container);
j->container_stack.pop_back();
return ret;
}
@@ -511,7 +519,11 @@ again:
/////////////////////////// Strings
case '"': {
std::string val;
// Reuse the parser-wide string buffer so we don't construct a
// fresh std::string (with its inevitable SSO->heap promotion
// and capacity doublings) for every JSON_STRING token.
std::string &val = j->string_buffer;
val.clear();
int surrogate = -1;
while ((c = read_wrap(j)) != EOF) {
@@ -632,7 +644,12 @@ again:
json_object_ptr s = add_object(j, JSON_STRING);
if (s != nullptr) {
s->string() = std::move(val);
// Copy (don't move) so j->string_buffer retains its
// grown capacity for the next token. The copy is a
// single right-sized allocation plus one memcpy, which
// is cheaper than the multiple capacity doublings the
// per-token std::string would otherwise incur.
s->string() = val;
}
return s;
}