From notifications at github.com Thu Oct 1 07:09:00 2026 From: notifications at github.com (leijurv) Date: Thu, 01 Oct 2026 00:09:00 -0700 Subject: [Tile-serving] [osm2pgsql-dev/osm2pgsql] CODA-style middle prototype (Issue #2505) In-Reply-To: References: Message-ID: leijurv left a comment (osm2pgsql-dev/osm2pgsql#2505) Further progress, I have pushed a "v2", which adds some more complexity but reduces the filesize, speeds up reading, and speeds up applying osmchanges, at the cost of creation a bit slower. https://github.com/leijurv/osm2pgsql/commit/24888e48c8aebefe0777bc54821b9dade4a02374 |Comparison|v1|v2| |---|---|---| |Middle size|81.32GB|70.03GB| |PBF -> middle time|~14m30s|~15m| |middle -> Carto time|~39m|~38m| |1 hour osmchange
median time (EC2)|26s|17s| |1 minute osmchange
median time (EC2)|0.41s|0.24s| Some of the changes: - `-1.56GB` In v1, every node coordinate was coded as a delta from the previous node in the way. Now, we decide between three strategies based on the way. If the way is not closed loop, then we linearly extrapolate our best guess for the next node from the delta of the last two nodes: `nodes[i-1] + (nodes[i-1] - nodes[i-2])`. The delta encoding is from there, which results in lower numbers and less entropy for zstd in the higher bits (lower delta bits and the sign bit are stored ~raw). If the way is a closed ring we just do it like v1 (delta from `nodes[i-1]`), **unless** it's a ring of 4 segments and we are on the last point, in which case we predict it's a parallelogram `fourth_guess = first + third - second` (and `delta = fourth - fourth_guess`). This works well for buildings (perhaps thanks to iD "Square"). The paralellogram predictor alone saves 0.65GB! - `-2.34GB` in the node-to-way index. Within each block of `node_id>>14`, we sort the `(node_id,way_id)` pairs by way (we can forget/ignore the actual order of the nodes in the way), then for each way we store the way_id (delta'd), the number of node_id "runs" (meaning consecutive increases, run length encoded), and then actual runs (start and length). The run length encode works well because within a way, very often node_ids are consecutive. That start used to be delta relative to `block<<14` (meaning the lowest `node_id` that could be here), but now it's delta relative to the prior run's end (even if the prior run was part of a different way). The deltas are also now split coded in v2 (like the x,y deltas), where we barrel roll the sign bit to the LSB, leave those lower bits as-is, and zstd the upper bits. - `-1.88GB` Deduplicate node locations within a way block. If a node appears in multiple different ways that have similar IDs (e.g. roads sharing nodes at junctions, buildings/landuse/etc having glued edges, etc), meaning they end up in the same `ways` block, we now refer to the prior encoded location by its node ID (skipping it implicitly in the delta stream). Node ids that appear in multiple ways that don't share `way_id>>10` are still repeated though. - `-2.8GB` Better LMDB paging. Anything over 2KB was put in overflow pages, which are made in units of 4KB, leaving half the last page wasted in expectation. Instead, store in chunks of 4080 bytes, and put the remainder in leaf pages which pack many such entries together. This removes the need for contiguous runs of pages for those overflows, so any 4KB page can be used. This is what flattens the free list growth. And, `-0.83GB` Split tail data between 2KB and 4KB in half. LMDB's page split threshold only allows up to ~2KB to live in a B+ leaf so this reduces fragmentation a little more. Filesize bloat from applying osmchanges and fragmenting the LMDB free list is also ~resolved: Image -- Reply to this email directly or view it on GitHub: https://github.com/osm2pgsql-dev/osm2pgsql/issues/2505#issuecomment-5926478592 You are receiving this because you are subscribed to this thread. Message ID: -------------- next part -------------- An HTML attachment was scrubbed... URL: