LV-CHECKPOINT-META-RECLAMATION.md
_checkpoints/meta/ growth - design and implementation planfeat(sql): add live views, branch puzpuzpuz_live_view7d6f5f4bc3a2712f7217) - section 5.1. Phase 2a (3175302fdc) - section 5.2, where
implementing it replaced the closure-summary mechanism this document originally
specified with per-segment page counts. Phase 2b (ab0bd57871) - section 5.3,
which took the closure-summary mechanism after all, for a reason the plan had
not stated. Phase 3 (d7bf14f612) - section 5.4, removed again by 524a8e833b.
Phase 4 (a8bc16da33) -
section 5.5, the catalogue's own entry retirement, which was decision 7 rather
than a planned phase. Phase 5 (ec6493a268) - section 5.6, the purge cadence,
which was the "when the sweep runs" leftover decision 7 recorded. Phase 6
(afcf762b2a) - section 5.7, the files a failed publication leaves that no
count describes, which was the final-orphan leftover section 5.6 recorded and
turned out to be a leak rather than a lag. Phase 7 (5d1dd718b6) - section 5.8,
which gave reconciliation the catalogue rule too, closing the half of that leak
a process with no cadence sweep still carried. Then 524a8e833b - section 5.9,
the removal of Phase 3 and the move of the non-seal failure injection that had
to go with it. Class A is closed as a mechanism, so is the residual it left, so
is the restart-scoped lag in collecting it, and so is the one disposition that
was neither, under both collectors rather than one. Class B is not closed and
no longer has a mechanism.
Scope decision, 2026-07-31: the retention horizon is not shipping - and as of
524a8e833b it is out of the branch. The brief was metadata GC, section 1
widened it to bound retained state too, and that widening is withdrawn - so Class
A is the deliverable and Phase 3 came back out. Tasks 1 and 2 in section 12
landed together in that commit and section 5.9 records them; task 3 followed on
2026-07-31 and section 5.10 records it, task 4 the same day - section 5.11 -
where the narrowed key set decision 5 had left open turned out to be a live wrong
answer for ROWS views and was fixed by refusing the splice, and task 5 the same
day again - section 5.12 - which gave the splice back keyed on the output key
domain, so a ROWS repair re-versions the keys its replay describes and leaves
every other key's entry as the old root wrote it. Every task in section 12 is
done. Neither task 4 nor task 5 is reclamation work. Class A itself
needs no further code. The consequence, stated
because it is the point: nothing bounds a default install's retained checkpoint
state; what the cadence and the reconciler bound is the garbage beside it. The
#6939 PR body now says both halves in two clauses rather than one.
The original three semantic questions in section 6 were traced on 2026-07-30 and
came back clear.The brief was _checkpoints/meta/. That is the right primary target - it is the
only surface with no reclamation mechanism whatsoever - but the plan below
also covers _checkpoints/data/, because the two share a root cause and the data
side needs no new machinery once that cause is addressed.
data/ is not fine today. Its garbage collection works (reference counts plus
PurgeSweep), but nothing ever releases a reference in the steady state, so its
retained-state growth is unbounded in exactly the same way meta/'s is. Phase 1
addresses both at once; Phase 2 is metadata-only.
The second correction this section originally made is withdrawn. It read that
the plan "also bounds" the two directories, and Phase 3 - the retention horizon -
was how it proposed to do that. Bounding retained state means deciding to stop
keeping live boundaries, which is a policy feature rather than a collector, and
the scope decision of 2026-07-31 (section 11) takes it back out. So this plan
reclaims garbage in both directories and bounds neither: what a generation cannot
reach is collected, and what it can reach stays for as long as the boundary
naming it does. Task 1 in section 12 removed Phase 3 (524a8e833b, section 5.9);
sections 2 through 10 are left as written, because the growth model and the
trade-offs they record are what a future retention proposal would start from.
Where those sections say Phase 3 does something, read it as what the removed
mechanism did, not as a description of the branch.
The live view's own table (partitions) is explicitly out of scope - TTL support lands separately.
Reachability from the superblock decides whether a byte is garbage or retained state. The chain is:
superblock slot (A/B)
-> timeline root ref -> timeline B+ tree
-> entry (one per sealed boundary)
-> checkpoint root
-> function roots
-> partition-map pages
-> state pages (data/)
-> anchor root
-> partition-map pages
-> row-position-delta root ref -> row-position-delta B+ tree
-> segment-directory root ref -> segment-directory B+ tree
The three direct metadata roots are versioned per generation. The newest two root versions are reachable from the two valid superblock slots, but each copy-on-write tree can reuse subtrees written by many earlier generations. Everything hanging off a timeline entry stays reachable for as long as that entry exists, and no code path removes an entry from below.
| Item | Per-seal volume | Reclaimable today |
|---|---|---|
| superseded timeline spine pages | O(log N) | yes, since Phase 2a |
| superseded segment-directory spine pages | O(log N) | yes, since Phase 2a |
metadata closure orphaned by a repair splice (LiveViewCheckpointTimelineStoreWriter.java:574) | O(live_keys) per repair | yes, since Phase 2b |
metadata closure orphaned by publishTruncate (:776) | O(dropped entries) | yes, since Phase 2b |
| row-position delta index pages superseded by a repair (section 6.5) | O(log R) per repair, where R is the number of delta breakpoints | yes, since Phase 2a |
Metadata only. data/'s equivalent was already handled: both of those sites
release their data-segment references through applyRootReferenceChanges, and
neither had a metadata counterpart. That asymmetry was the whole of Class A, and
Phase 2b closed it by giving the metadata half the same reference transaction -
the same call, on the same list, with metadata ids beside the data ones.
Order of magnitude for the steady-state part: single-digit MB/day per view.
Class A is closed. Everything meta/ and data/ still hold belongs to a
boundary the timeline names, so what remains is a retention question rather than
a garbage-collection one - which is what Phase 3's horizon moved, and what
nothing moves now that task 1 has taken it out. The catalogue's
own entries were the one exception - reachable from the newest generation but
naming nothing, once the sweep had unlinked their files - and Phase 4 retires
them. Closing a row of this table means a collector may reclaim it, which is
not the same as its doing so: until Phase 5 the collector ran once per process,
so every row above was closed in accounting and open on disk.
One kind of file this table cannot hold at all: the segments a publication renamed into place and then failed to commit. They are not reachable from the superblock, so no row of the model above prices them and no reference count decides them - their whole description is that the catalogue never held them. That put them outside both classes rather than inside Class A, which is why they survived every phase up to Phase 6, and outside the id-ceiling rule too once a later publication allocated past them. Volume is one publication's worth of segments per failure - two files in the measured case - and the population is every failed compaction or repair, since a failed seal is the one shape the old rule did catch. Retention was a third until task 1 removed the publication. Phase 6 collected them on the sweep cadence and Phase 7 at reconciliation, so both collectors now decide them; before Phase 7 a process that never swept - the cadence disabled, or a view that stopped sealing after the failure - still held them for the life of the directory.
Every retained boundary's closure, in both meta/ and data/. From
freezeFunction (LiveViewCheckpointTimelineStoreWriter.java:1217):
| Function shape | Written per seal | Growth variable |
|---|---|---|
non-ring (sum, avg, row_number, rank, first_value, ...) | fresh whole-image state page per live key, no diff against the previous boundary (:1270-1274) | seals x live_keys |
ring, under MAX_LIVE_CHUNKS = 256 | new chunks only; chunk sharing works | cumulative rows |
| ring, above that wall (>~1M live rows/key) | full ring re-image | seals x live_rows |
Plus the matching partition-map rewrite in meta/: putPartition runs for
every live key (:2313), so every leaf and therefore every internal node is
rewritten each seal.
Seal cadence is cairo.live.view.checkpoint.rows = 1_000_000 or
cairo.live.view.checkpoint.max.duration.micros = 5m, whichever comes first
(PropServerConfiguration.java:1526,1530), so >=288 seals/day for anything
continuously ingesting.
Order of magnitude at 10K live keys, non-ring: ~1 MB/seal, ~275 MB/day, ~100 GB/year, for one view. At 1M keys, ~27 GB/day. Class B is roughly two orders of magnitude larger than Class A.
The sharp edge: there is no dirty-key check. One row into one key satisfies the cadence, and the seal then re-images every live key. A view with a static key set and a one-row-per-five-minutes trickle rewrites its complete state 288x/day forever.
Phase 1 closed the sharp edge for non-ring functions: an untouched key's page and map entry are both reused, so that trickle now costs one key per seal rather than the whole key set. The rest of the table stands - the retained closure of every surviving boundary is still Class B.
Phase 3 bounded the table as a whole, but only where an operator set a horizon,
and task 1 removed it. Every row above prices one boundary; the horizon fixed how
many boundaries a generation holds, so the store's footprint became
retained_boundaries x per-boundary cost instead of seals x per-boundary cost.
It shipped at a default horizon of zero, where it fixed nothing, and nothing
replaces it: this table now stands unbounded as written.
LiveViewCheckpointAnchorRootBuilder's javadoc (:42-47) records that
LiveViewCheckpointPartitionMapWriter already drops a put whose key and value
already match, and that this is what keeps an adjacent seal proportional to the
partitions whose anchor value actually moved rather than to the map's size.
Anchors benefit because an anchor value is a small immutable long: an unchanged key produces a byte-identical payload, so the writer drops the put.
Function state does not benefit, and the reason is ordering, not a missing
mechanism. freezeStatePage (:1279) writes a fresh page at a fresh offset
before the put, so the entry payload carries a new (segmentId, offset) even
when the encoded state bytes are byte-identical. The existing elision can never
fire.
So Phase 1 is not "add dirty tracking to 150 window classes". It is: compare
against the previous boundary's page before writing, and reuse its ref when the
bytes match. The plumbing is already threaded - freezeFunction takes
previousBoundary and the ring branch already calls
previousBoundary.find(identity, stateFormatVersion, key) at :1258. Only the
non-ring branch ignores it.
meta/ and data/skipPublishedSegmentIds (:1293) and nextFreeSegmentId
(LiveViewCheckpointCompaction.java:237) both treat a candidate id as taken if
either meta/<id> or data/<id> exists. So an id names at most one file, in
exactly one of the two directories.
That means the existing catalogue can hold metadata entries with no change to
addSegment / applyRootReferenceChanges semantics and no new id space. It also
means PurgeSweep can be taught one extra path probe rather than a second sweep.
Function partition maps rewrite every leaf today because their state-page refs always change. Anchor maps do not: their equal-put elision already leaves later anchor roots pointing into partition-map subtrees written for earlier boundaries. The timeline, row-position delta index and segment directory also reuse old subtrees across generations.
Phase 1 extends this existing property to function partition maps; it does not introduce it. A naive "drop every segment older than the oldest retained boundary" rule is therefore unsafe even before Phase 1. Any metadata reclamation must account for the complete transitive closure of every surviving root. Phase 2a does that for the three direct trees by counting pages rather than closures; Phase 2b does it for the function and anchor maps - which is exactly where the cross-boundary sharing described here lives - by persisting the closure in each root and counting roots at the catalogue.
Worth inventorying because substantial pieces are reusable, even though the closure accounting itself is new structure.
| Existing | Location | Reused for |
|---|---|---|
per-segment referenceCount + retireGeneration catalogue | LiveViewCheckpointSegmentDirectoryEntry.java:47,53 | metadata entries - done, Phase 2a/2b |
addSegment(segmentId, fileLength, referenceCount) | LiveViewCheckpointSegmentDirectoryWriter.java:111 | registering metadata segments - done, Phase 2a/2b |
applyRootReferenceChanges(removed, added, generation) | :142 | boundary-metadata reference deltas - done, Phase 2b; META (tree) entries move through releaseMetadataPages instead |
purge rule (oldestValidSlotGeneration >= retireGeneration && minPinnedGeneration > retireGeneration) | LiveViewCheckpointDataStore.java:595-626 | unchanged, gained a metadata path probe - done, Phase 2a; Phase 2b needed no change to it at all; Phase 4 made it also the proof that an entry is dead; Phase 5 runs it on a cadence rather than once per process, and the rule is indifferent to how often it runs |
| persisted data-segment use-count tally inside a function root | LiveViewCheckpointFunctionRootBuilder.java:156-158, adjustSegment:193 | carries the metadata closure beside the data one - done, Phase 2b |
| copy-on-write put elision | LiveViewCheckpointPartitionMapWriter (documented at LiveViewCheckpointAnchorRootBuilder.java:42-47) | Phase 1 |
| high-side timeline truncate | LiveViewCheckpointTimelineWriter.truncateAbove, publishTruncate | template for the low-side mirror Phase 3 built and task 1 removed; the high-side half is untouched, and it is the only publication left that drops a boundary |
| generation pins, try-with-resources at all seven sites | LiveViewCheckpointGenerationPin | unchanged |
| foreign-layout-version retire path | LiveViewCheckpointLifecycle.java:54,84,403 | format migration |
| final-orphan id-ceiling scan | LiveViewCheckpointLifecycle.cleanupOrphans, purgeFinalOrphans:204 | the shape Phase 6's catalogue scan copied, and the rule it had to replace - the ceiling it compares against stops naming a file once a later publication has stepped over it. Phase 7 left it only the case the catalogue cannot answer: a directory with no valid generation, where the ceiling is zero |
| deterministic publish-stage crash injection | setTestFailureStage, TEST_FAIL_AFTER_{DATA,METADATA,SUPERBLOCK}_PUBLISH | crash tests; TEST_FAIL_AFTER_COMPACTION_METADATA_PUBLISH is the non-seal one Phases 6 and 7 need, moved there by task 2 |
Phase 2a supplied the directory's deferred self-registration, per-segment page
accounting for the three direct trees, and the metadata purge path; Phase 2b
supplied the metadata closure accounting for function and anchor maps, and found
that the release sites it needed - the repair splice and publishTruncate -
already carried it once the roots stated the closure. Phase 3 added the low-side
timeline truncate and the row-position-delta prune this table listed as absent,
and needed no new reclamation machinery at all: retiring a boundary is the same
reference transaction the truncate already ran, against the closure Phase 2b
taught the roots to state - all of which task 1 has since removed. Phase 4 added one tree operation - an entry removal
that prunes through the existing emit path - and one field on the writer to
carry the sweep's proposal to the seal that applies it. Phase 5 added no
machinery whatsoever: it calls the sweep more than once. Phase 6 added one
directory scan, held against a catalogue read the purge rule was already entitled
to make, and no state at all - it decides and acts in the same pass, which is
what keeps it correct where the deferral it replaced was not. Phase 7 added none
either: it calls Phase 6's scan from the reconciler as well, over a superblock
that method already had open.
Three mechanisms, landed in this order and in four commits, plus a fifth commit for the residual they left, a sixth for the collection lag that residual's own fix inherited, a seventh for the one disposition none of the six accounts for, and an eighth for the collector that still could not take it. Phase 1 is independent. Phase 2 lands without a retention horizon and closes existing metadata garbage; it split into 2a and 2b during implementation, for the reason section 5's Phase 2 preamble gives. Phase 3 depended on Phase 2b for metadata reclamation, and in the event needed no reclamation code of its own: retiring a boundary is the reference transaction Phase 2b had already made complete. Phase 4 was not in this plan at all - it is open decision 7, taken as code once it was the only term left growing with a view's age. Phase 5 was not either - it is the "when the sweep runs" leftover Phase 4 recorded, which had become the only thing left between the mechanisms above and the disk they were supposed to give back. Phase 6 is the final-orphan leftover Phase 5 recorded, and it is the one whose scope the plan had wrong: those files were not waiting for a restart, they were lost, because the rule naming them stops holding as soon as another publication allocates past them. Phase 7 is the leftover Phase 6 recorded, and it finished that correction: the cadence sweep collected them, the reconciler still could not, so a process that never swept lost them exactly as before. A ninth commit then took Phase 3 back out, on the scope decision of section 11 - section 5.9. Section 5.10 is not a commit at all: it is what the #6939 body now says about the whole of the above, which is the only place a reader outside this document meets it.
Closes the seal-rate multiplier on cold keys, for both data/ and meta/.
No on-disk format change. It does require reader/cache plumbing in addition to
the byte scratch buffer.
For a partitioned non-ring function (freezeFunction:1270):
previousBoundary.find(frozen.identity, frozen.stateFormatVersion, key),
as the ring branch already does.dataWriter.beginPage().LiveViewCheckpointStatePageRef
verbatim and skip the page write.putPartition then receives a byte-identical payload for an unchanged key, the
partition-map writer drops the put, and neither the leaf nor its ancestors are
rewritten.
There are two previous-boundary implementations and they need different byte access:
RootPreviousBoundary resolves metadata only today. Comparing a published
state page requires a bounded data-segment reader opened with the catalogue's
checksummed file length. Cache readers by segment id for the duration of the
seal; do not assume the data segment is already mmap'd.CapturedPreviousBoundary points into the repair capture's still-unpublished
temporary data segment. Add a comparison method on the open data writer, or
retain the encoded bytes with the captured partition. A published-segment
reader cannot open this case. A first implementation may conservatively skip
elision for an in-flight previous boundary, but the tests and sizing must say
so.The map == null scalar branch (:1225-1232) also ignores the previous root.
Extend PreviousBoundary with scalar lookup and apply the same comparison there.
It is not the wide-key growth term, but leaving it out makes the claimed
non-ring elision incomplete.
Cost: one encode and one page comparison per cold key per seal in place of one
page write. It may map several old data segments because earlier elision spreads
live refs across generations. A cheaper variant - a per-partition dirty or
version stamp on WindowFunction so the encode and read are skipped - touches
many window implementations and remains a later optimisation.
Does not apply to ring functions: their entry payload carries an advancing row count, so no payload is ever byte-identical. Chunk sharing is their analogue and already exists.
One detail that must not be missed: a reused ref keeps a reference on an older
data segment. It has to be reported to applyRootReferenceChanges in the
added list, or Phase 3 will purge a segment a live root still names. Harmless
until Phase 2b/3 exist, fatal after.
a2712f7217)Landed on puzpuzpuz_live_view as designed above, in
LiveViewCheckpointTimelineStoreWriter plus two small additions to
LiveViewCheckpointDataSegmentWriter (addressOfPage, getSegmentId). No
on-disk format change; ~300 lines of production code.
What differs from the plan, and what the implementation turned up:
RootPreviousBoundary gained a LiveViewCheckpointSegmentDirectoryReader
bound to the old directory root, so a comparison read is bounded by the
catalogue's checksummed file length exactly as a restore's is, plus an
eight-slot clock cache of LiveViewCheckpointDataSegmentReader.
CapturedPreviousBoundary reads its own unpublished segment through the open
data writer, and refuses any reference that does not name it - which cannot
happen, because the first boundary of a repair shares against nothing.map == null) arm is implemented but effectively unobservable.
A non-partitioned function's single global state moves whenever any row
arrives, and a seal requires rows, so the comparison practically never
succeeds. It is completeness, not a saving, and the tests assert only that the
path still seals and restores correctly.LiveViewCheckpointFunctionRootBuilder.build adjusts the candidate segment
use counts from the old and new entry of every mutation, whether or not the
partition-map writer then elides the put, and getReferencedSegmentIds unions
those into the root's own catalogue. A reused reference on an older data
segment therefore reaches applyRootReferenceChanges in the added list with
no new code. Phase 3 inherits that unchanged.[L, H) over runtime
state the scratch overlay has taken out of the way, so a boundary it
re-versions images the keys that appear in that range rather than the view's
whole key set. It is pre-existing behaviour, unrelated to elision, and
orthogonal to this plan - but it is worth a separate look, because a restore
from such a boundary starts from a key set narrower than the one the boundary
originally described.Costs the change adds, stated because they are real:
Tests: LiveViewCheckpointStatePageElisionTest - a trickle into one key of 24
(exactly one map entry may change per seal; data/ grows by the touched key's
pages and nothing else), a restart whose head boundary names the first seal's
segment for every cold key, a repair whose capture shares one page across the
boundaries it re-versions, and a RANGE control. The first three fail before the
change and pass after; the control passes either way. The whole
io.questdb.test.cairo.lv package (1401 tests) is green.
Phase 2 split in two while it was being implemented, along the line the accounting itself draws. Phase 2a covers everything reachable from the three roots the superblock names directly - the timeline, the row-position delta index and the segment catalogue itself. Phase 2b covers everything reachable from a timeline entry: the checkpoint root and the anchor root, function root and partition-map pages below it. The split is not cosmetic: the two halves want different mechanisms, because the first is one live version at a time and the second is one live version per surviving boundary.
The split also reorders the value. 2a closes both steady-state Class A rows and the delta-index one - the recurring per-seal leak - while 2b closes the two repair-driven rows and is the prerequisite Phase 3 actually needs. 2a is a strict prefix of 2b: the kind byte, the purge path and the deferred registration are the same machinery.
kind byte (DATA / META) to
the segment-directory leaf entry. The shared id namespace (3.2) makes a kind
field strictly unnecessary, but LiveViewCheckpointCompaction iterates the
catalogue expecting data segments only, and inferring kind from an exists()
probe makes a missing file indistinguishable from a wrong-kind entry. Pay the
byte.referenceCount is the number of its pages the selected
generation's trees still reach. A publication adds the pages each tree writer
wrote and releases the ones its path copy replaced. Zero still means "the
selected generation names nothing in this file", which is all the purge rule
reads, and the existing slot-generation and reader-pin gates keep an older
slot or pinned reader safe without adding to the count.truncateAbove, which drops subtrees without reading them; those
are walked explicitly, at a cost proportional to what is dropped.addSegment(id, bytes, pageCount, META) before publishing the new directory root.(pendingDirectorySegmentId, bytes, pages) and the next publication registers
it. This is safe because:
PurgeSweep.onEntry to build metaSegmentPath for META entries
under the identical slot-floor plus reader-pin rule. The rule needs no change.metadataBytes, dataBytes and
rowPositionDeltaBytes currently mean bytes written in the history epoch;
LiveViewCheckpointTimelineStats.getPhysicalBytes explicitly says purged
segments are not subtracted. Do not decrement one of them. Report current
live and obsolete bytes from the catalogue/sweep instead, and extend the
existing obsolete-segment metric to include META entries.SLOT_FORMAT_VERSION 2 -> 3. An older _timeline then retires
through the existing foreign-layout path
(LiveViewCheckpointLifecycle.java:403), which is free: the timeline is
derived state, so discarding it costs fast restart recovery, not correctness.
Live views are unreleased, so no real migration exists.3175302fdc)Landed on puzpuzpuz_live_view as the nine steps above, across eleven production
classes; ~700 lines of production code plus a 400-line test.
The mechanism changed, and this is the substantive deviation from the plan as
written. Steps 2-4 of the original design called for persisted transitive
closure summaries in every root and node. Prototyping showed that shape is wrong
for the direct trees on its own terms: a timeline root's closure holds one entry
per distinct segment its live pages sit in, which is O(N / leafCapacity) and
grows with the timeline, so writing the summary every seal is quadratic in the
seal count. Per-segment page counts are exact, incremental, O(path length)
per publication, and need no format change to any root or node at all - the whole
format cost of Phase 2a is the kind field and three superblock longs.
The reason the plan reached for closures is still valid, and it is what makes Phase 2b a different problem rather than more of the same. Page counts are cheap to maintain and expensive to release in bulk: retiring a boundary means decrementing the pages that die with it, and which of its pages die depends on what the neighbouring boundaries still share. For the direct trees that question never arises - one live version at a time, and the publication knows exactly which pages it replaced. For boundary metadata it is the whole problem. See Phase 2b.
What else the implementation turned up:
publish
resolves its release set with a read-only twin of applyRec; an assert
compares what the twin predicted against what the path copy actually visited,
so the duplicated descent logic cannot drift apart silently.checkpoint_data_segment_count stays data-only. It is a published
catalogue column documented as data segments, so cataloguing metadata beside
them must not redefine it. checkpoint_obsolete_segment_bytes does span both,
because both kinds wait on the same fallback slot, so a view that has sealed
twice now reports a non-zero collection lag where it reported nothing.meta/ and it wants an entry-retirement path
of its own. Not scoped here. Phase 4 (a8bc16da33, section 5.5) added that
path.Measured on the test workload (one boundary per commit, 24 further seals): the
sweep reclaimed 47 segments, meta/ grew by 49 files rather than the ~96 it
would have, and the remainder is the two boundary-metadata files per seal that
Phase 2b and Phase 3 own.
Closes the two repair-driven Class A rows, and is the prerequisite Phase 3 depends on. The unit of accounting is the checkpoint root's closure: its anchor root, function directory, function roots and every partition-map page below them.
The design question was which of the two mechanisms to use, and unlike in Phase 2a the answer was not obvious from the write cost alone:
LiveViewCheckpointFunctionRoot's existing (segmentId, useCount) list and
LiveViewCheckpointRoot's segment list to carry META ids beside the DATA ones
they already hold - the shared id namespace (3.2) means no new id space and,
for the checkpoint root, no format change at all. Releasing a boundary is then
what publishTruncate already does: decrement every id the root names. The
cost is at write time and it is real: a function root's summary holds one entry
per distinct segment its partition-map pages sit in, which after Phase 1's
elision is roughly one per cold leaf, so ~160 entries at 10K keys - about 2.5
KB per function per seal.K requires knowing which
of its pages K+1 does not share. A parallel descent of the two partition maps
that prunes wherever a child ref is identical costs what the two boundaries
differ by, which is what seal K+1 wrote - so the cost lands at retirement
rather than at seal, and only for boundaries that actually retire.The first won, and not on the cost trade this document framed it on. The
page-count variant is unsound at the site Phase 2b exists for. An
adjacent-boundary diff rests on reachability of a page being a contiguous
interval in boundary order - true while boundaries are only ever appended by a
cadence seal, because each seal's map is a path copy of the one below it. A
repair breaks it: publishRepair builds each re-versioned boundary from its
own old root rather than from the boundary before it, so boundary K+1 can
supersede a page that boundary K and boundary K+2 both still name. Diffing a
retired range against its two surviving neighbours would then release a page a
live root reaches. Reference counting is order-independent and has no such
precondition.
The write cost that argued against summaries turns out to be moot as well. A
checkpoint root already lists one entry per data segment its state pages sit in,
and Phase 1's elision already spread that over roughly one segment per seal; the
seal already applies that whole list to the catalogue on every publication.
Adding the metadata ids to the same list keeps the same shape and the same
applyRootReferenceChanges call - it does not introduce an asymptotic the seal
did not already pay.
Phase 2b also needed:
addSegment(..., META) call Phase 2a
makes - see section 5.3 on why the kind field grew a third value instead;publishTruncate (:776) and at the repair splice
(:574), which is where the Class A volume actually is - and which needed no
new code at all once the roots stated the closure, because both sites already
hand the old root's segment list to applyRootReferenceChanges;ab0bd57871)Landed on puzpuzpuz_live_view across ten production classes; ~250 lines of
production code plus ~200 lines of test. One on-disk format change to a metadata
page, and SLOT_FORMAT_VERSION 3 -> 4 so an older _timeline retires through
the existing foreign-layout path rather than meeting it.
The catalogue now keeps two kinds of metadata segment, not one. The plan said
registration would be "the same addSegment(..., META) call Phase 2a already
makes". It cannot be: a META entry's referenceCount counts pages and a boundary
entry's counts roots, and one field cannot carry both meanings unchecked. A third
kind - SEGMENT_KIND_BOUNDARY - makes the unit explicit, so
applyRootReferenceChanges refuses a page-counted segment and
releaseMetadataPages refuses a root-counted one. isMetadata() becomes "not
DATA" and decides only which directory the file lives in, which is all the purge
sweep ever asked it.
Where the closure lives:
(segmentId, useCount) list carries metadata
entries beside its data ones - the number of pages of that segment the root
reaches, its own page plus the partition-map pages below it. No format change:
the two id spaces are disjoint, so one ordered list serves both.FORMAT_VERSION 1 -> 2). It needed it for the same reason a function root
does: the equal-put elision documented at LiveViewCheckpointAnchorRootBuilder
leaves later anchor roots pointing into map pages older seals wrote.Counts are maintained from the delta of one build, never from a walk:
released + written, where written is the pages the build put in its fresh
segment and released is what the path copy took away.
The subtle half is what "released" excludes.
LiveViewCheckpointPartitionMapWriter now records the segment of every decoded
page it rewrites or drops. A page it decoded and left alone is not released,
and after Phase 1 that is the common case rather than a corner: a put whose key
and value already match makes mutate answer false, the parent keeps its
existing child reference, and the new map still names that page. Counting it
would release a page a live root reaches. The three drop sites that are not a
rewrite - an emptied child removed from its parent, a collapsed root, a map that
went empty entirely - are released explicitly, and releaseSource clears the
node's source so a second visit cannot double-count it.
What else the implementation turned up:
publishTruncate and the repair splice
already build removedSegmentIds from the old root's segment list and hand it
to applyRootReferenceChanges. Once that list carries metadata ids, both
reclaim the boundary metadata by construction. The whole of the "release the
closure" bullet is one widened list.addSegment(dataSegmentId, ..., 1)
plus dropSegmentId already worked. A written segment the root does not name
is refused rather than silently registered, because it would mean the closure
the root publishes and the files the build wrote have diverged.LiveViewCheckpointLifecycleTest
published superblocks naming a directory root it never registered or declared
pending. That is exactly what the new rule calls corruption, so the fixture now
sets the pending triple as a real publication does.Costs the change adds, stated because they are real:
Measured on a 12-seal history with 24 boundary-metadata segments live before the event: a truncate deep in the history reclaimed 14 of them, and a splice just below the head reclaimed 2. Both reclaimed nothing at all before the change.
What Phase 2b does not change: a cadence seal still leaves its boundary
metadata behind, because the boundary is live. At the test workload's shape that
is about three meta/ files per seal; Phase 3's horizon retires them, and
without one set they stay.
One thing worth recording for Phase 3, because it inverted an assumption while
the tests were being written: which disposition a correction takes is not
"deeper means more drastic". In the ROWS 3 PRECEDING view the cases use, a
correction just below the head classified a converged suffix and spliced, while
one seven boundaries down could not and truncated. Section 6.4 already says the
lower bound comes from the function's dependency rather than from checkpoint
availability; this is the same point from the other end, and it means a Phase 3
horizon test must assert which publication ran rather than infer it from how deep
the correction was.
Removed by task 1 in 524a8e833b (section 5.9). This subsection and section
5.4 are kept as the record of what the mechanism was and what implementing it
turned up, because a future retention proposal starts from here rather than from
scratch. Nothing below describes the branch as it stands.
The retention semantics are clear - see section 6 - but implementation depends on
Phase 2b for metadata reclamation. publishTruncate:711 is a close template for
the publication, with the delta-index exception below. The timeline pages a
low-side truncate drops are already released by Phase 2a's dropped-subtree walk,
which truncateBelow inherits by mirroring truncateAbove.
LiveViewCheckpointTimelineWriter.truncateBelow - the mirror of
truncateAbove:205. Keep the high suffix by page reference, path-copy the
boundary spine, promote a subtree that collapses to a single child.LiveViewCheckpointTimelineStoreWriter.publishTruncateBelow - the mirror
of publishTruncate:711. Walk the dropped range, release each entry's data
and metadata references, and decrement logicalStateBytes. The physical
byte counters remain cumulative per Phase 2a step 8.
Carry normalizedBaseSeqTxn and coveredLvSeqTxn forward untouched exactly as
publishTruncate does (:795-800) - that is what keeps the WAL purge floor
still (section 6.1).seedCursorOffset forward
rather than clearing it (a low-side truncate does not invalidate a mid-sweep
resume point). Prune the row-position delta index, but preserve its prefix
contribution: let K be the first surviving timeline key and P the sum of
every delta entry below K; drop those entries and add P to the delta at
K (inserting one when absent). Every surviving lookup then sees the same
prefix sum as before. Simply deleting breakpoints keyed to dropped boundaries
is incorrect because each difference applies to the entire later suffix.
Implement this as a low-side prune operation on
LiveViewCheckpointRowPositionDeltaWriter, path-copying the boundary spine and
folding the discarded subtree sums into K.DISPOSITION_BOUNDARY_REBUILD, an already-named and already-priced
disposition. A cost dial, not a correctness change.d7bf14f612), before task 1 removed itLanded on puzpuzpuz_live_view as the five steps above, across nine production
classes; ~450 lines of production code plus ~900 lines of test. One superblock
format change and SLOT_FORMAT_VERSION 4 -> 5, for a reason no step anticipated
(the entry count, below).
Policy: an event-time window, cairo.live.view.checkpoint.retention.micros,
default zero (disabled). Step 4 offered "the last D of event time, or the newest
K" and left the choice to a human. Event time won on cost rather than on
semantics: the floor is head.maxTimestamp - D, so deciding whether anything
retires is one O(log N) predecessor probe, while "newest K" would need the
timeline's k-th-from-the-end key and the reader exposes no such navigation. It
also matches where decision 1 says this ends up - derived from view TTL, which is
an event-time window.
The pass runs after every seal rather than on a cadence of its own. When the horizon has nothing to retire it costs that one probe; once saturated it retires one boundary per seal, which keeps the footprint flat instead of sawtoothing, at the price of one extra publication per seal. That price is real and it is the main cost this phase adds: a seal that retires something now writes a second timeline segment and a second catalogue segment.
What differs from the plan, and what the implementation turned up:
truncateBelow as the mirror of truncateAbove, and structurally it is - same
spine copy, same single-child collapse. It is not symmetric in one respect the
plan did not name: a prefix truncation changes the minimum of the straddling
child, and navigation reads the minimum a parent stores rather than the subtree
under it. A stale-low minimum is not immediately wrong - a descent into it finds
nothing and reports absence, which is the right answer - but it misclassifies
the straddle on the next truncation, so the recursion returns the new minimum
and the parent stores it. truncateAbove never had to: dropping a suffix leaves
every surviving child's minimum where it was.checkpoint_timeline_entries was nextCheckpointId, on the stated grounds that
ids are allocated from zero and monotonically. Retention breaks that: the id
counter keeps climbing while the live set stays flat, so the column would grow
without bound while reporting a number that is meant to be bounded. A
retiredCheckpointCount field makes the count exact. Note this was already
wrong before Phase 3 - publishTruncate has always dropped boundaries without
adjusting the counter - so the fix is a correction as much as an addition, and
the high-side truncate maintains the field too.START FROM, which is not a retention outcome. Both the tree
operation and the publication check it, so a mis-set horizon costs nothing.seedCursorOffset carries forward, and that is the one field where the
template is wrong. publishTruncate clears it because it discards the head
the sweep was resuming into. A retention pass drops boundaries the sweep is long
past, so clearing it would lose a resume point that is still correct. Step 3
predicted this; it is recorded here because it is the only line of
publishTruncate that could not be copied.applyRootReferenceChanges over the closure the root already states, which is
exactly what publishTruncate does to the other end of the timeline. The whole
of "release the data and metadata references" is one loop that already existed.Costs the change adds, stated because they are real:
RANGE view still plans a dependency-localized repair with nothing sealed under
the change - but a view carrying a function with no finite dependency reads its
whole history instead. Section 6.4 prices that population; the shape census that
would say how large it is has still not been done.SLOT_FORMAT_VERSION bump retires an older _timeline through the
foreign-layout path, which costs a rebuild of derived state. Free in practice,
since live views are unreleased.Measured on the test workload (one boundary per commit, a 60-second event-time
horizon at a 10-second commit spacing) between 10 seals and 30: the boundary count
held at 7 and 7, meta/ went 20 -> 24 files and data/ 10 -> 11 over the 20
further seals. Without a horizon those 20 seals add 20 boundaries and about four
files each. The residual that does move is the segment catalogue's own tree
(decision 7, closed by Phase 4) plus the one data segment a cold key keeps its
elided reference in.
What Phase 3 does not do: change the default. The horizon ships at zero, so a default install still grows exactly as it did before this phase. Open decisions 1 and 2 own that, and they are now the only thing between the mechanism and a bounded default.
Not a planned phase: this is open decision 7, taken as code once the three
mechanisms above had left it as the only term in _checkpoints/ that still grew
with a view's age. The catalogue held one entry per segment ever written, because
a purge unlinks the file and nothing removed the entry, and its own B+ tree
gained a leaf every 64 of them - the residual section 5.2 named and Phase 2b made
about 2.5x faster by cataloguing boundary segments beside the rest.
Decision 7 already stated both halves of the answer: retiring an entry is safe exactly when its file has been unlinked, which the sweep proves, and the sweep publishes no generation, so the removal has to be staged into a publication.
a8bc16da33)Landed on puzpuzpuz_live_view across six production classes; ~215 lines of
production code, over half of it comment, plus ~370 lines of test. No on-disk format change and no
SLOT_FORMAT_VERSION bump: an entry leaving a leaf is an ordinary copy-on-write
mutation of a tree whose layout is unchanged.
The hand-off is a proposal, not a transfer. PurgeSweep collects the id of
every entry it leaves with no file - the ones it unlinked in this pass and the
ones an earlier pass unlinked that no publication has carried away yet - and
PurgeResult / ReconcileResult carry the list out.
LiveViewCheckpointTimelineStoreWriter holds it per checkpoint directory between
the reconciliation that produced it and the seal that applies it. Re-proposing
the already-gone ones is what makes the whole thing crash-proof without a durable
queue: a failed publication, a BoundaryNotAboveHeadException, or a process that
dies between the two loses nothing, because the next sweep says it again.
That is not a corner case, it is the ordinary production path. CairoEngine
reconciles every checkpoint directory at boot and drops the list, because it
publishes nothing; the writer then reconciles the same directory again at its
first seal of it, and the second sweep re-proposes what the boot sweep unlinked.
Without the re-proposal, the boot sweep's work would never reach the tree.
Where the work lands:
removeSegment stages a removal like any other mutation, so the removals
path-copy once alongside the publication's registrations and releases rather
than in a pass of their own. The seal's releaseOwnPages pre-pass and the path
copy descend the same paths for a removal key as for any other, so the
assertion that holds those two descents against each other needed no change.emitNodes returns without writing at
a count of zero, so the parent keeps no child reference to it, and a parent
that loses every child empties in turn. That is the whole of the pruning: no
rebalancing pass, no merge rule, no minimum-occupancy invariant. A B+ tree
whose only deletion pattern is "the low ids die first" gets a correct shape out
of the emit path alone.truncateBelow: navigation reads the minimum a parent
stores, and a descent into a subtree whose real minimum has risen finds nothing
and reports absence, which is the right answer. Unlike the timeline's prefix
truncate, the catalogue never has to raise it, because it takes no second
operation whose straddle classification would depend on it.LiveViewCheckpointMetaSegmentWriter
gained a discard() so that publication leaves no page-less segment behind.
This is unreachable from a real publication - every one of them registers at
least the timeline segment it just wrote, so staged always holds an insert -
but it is reachable from the catalogue's own unit tests, and a null root is the
empty shape begin() already accepts, so supporting it costs less than
refusing it.Two guards keep the two units from meeting. removeSegment refuses a
still-referenced entry, whose file cannot have been unlinked, and one this
publication registers; applyRootReferenceChanges and releaseMetadataPages
refuse an entry already staged for retirement. Neither can fire from a correct
sweep - the file is gone, so nothing can reach it - which is exactly why they are
worth having: firing means the count the sweep acted on and the closure a root
publishes have diverged.
Crash safety adds no case. The removal commits with the generation carrying it, and the fallback slot keeps its own copy of the entry at a zero count, so nothing under that generation reads the missing file and a sweep over it re-proposes the retirement rather than faulting.
What Phase 4 does not change: when the sweep runs. purge() is called from
LiveViewCheckpointLifecycle.reconcile and nowhere else, and a writer reconciles
a directory once - at its first seal of it. So entry retirement follows the
sweep: one seal after a restart's reconciliation, not on a cadence of its own.
That is the same bound the segment files already have - nothing unlinks a
superseded segment mid-process either - so the catalogue is now exactly as
bounded as the directory it catalogues, and making the sweep periodic would move
both at once. Worth doing, and out of scope here. Phase 5 (ec6493a268,
section 5.6) did it, and moved both.
Measured on the test workload (one boundary per commit, 32 seals, then one reconciliation): the sweep left 57 of the catalogue's 159 entries naming unlinked files, and the seal that followed removed exactly those 57 and no others.
Not a planned phase either: this is the leftover section 5.5 names. Every
mechanism above decides what may be collected; none of them decides when
anything is. purge() had exactly one caller, and a worker reaches it once per
checkpoint directory - at its first seal of it - so a process that runs for a
week collects what its first seal could see and nothing after. The accounting was
complete and the disk did not come back.
ec6493a268)Landed on puzpuzpuz_live_view across eight production classes; ~290 lines of
production code, of which the reclamation logic is about forty - the rest is
javadoc, a result class and the five files a new config key has to touch - plus
~190 lines of test. No on-disk format change and no SLOT_FORMAT_VERSION bump:
running an existing pass more often changes nothing a file records.
cairo.live.view.checkpoint.purge.interval, counted in seals, default one.
It matches the compaction interval beside it - the same counter shape, and for
the same reason that one counts seals rather than testing a base seqTxn modulo -
and zero disables it, which restores the reconcile-only behaviour exactly.
Default-on is the deliberate part: unlike the retention horizon, whose default
decision 2 leaves to a human because a horizon costs repair reach, a sweep costs
nothing semantically. It unlinks files that nothing can reach, under a rule the
reconciler already ran on every restart.
LiveViewCheckpointTimelineStoreWriter.sweep() is the reclamation half of
reconcile on its own, without the epoch, repair-descriptor and orphan rules
that only a directory nobody has published under yet needs. It opens the
generation, refuses one some other history epoch owns, runs
LiveViewCheckpointDataStore.purge(), and stores the retirement proposal in the
per-directory map Phase 4 already gave the writer. LiveViewRefreshJob runs it
after retention and compaction, so it walks a catalogue both of them have
finished writing, and gates it on an actual seal exactly as those two are.
What the implementation turned up:
oldestValidSlotGeneration >= retireGeneration already means the
fallback slot has caught up with the retirement, so a segment retired at
generation G becomes collectable once G+1 commits and not before - whether
the sweep that notices runs then or a week later. Frequency is not a term in
the rule, which is why a cadence needs no argument about safety, only about
cost.skipPublishedSegmentIds and nextFreeSegmentId allocate from the monotonic
nextSegmentId and only skip forward over files that exist, so unlinking a
file below that ceiling never makes its id reusable. Without that property a
cadence sweep would hand out ids the catalogue still holds entries for.live_views() columns change meaning slightly.
checkpoint_data_segment_count and checkpoint_obsolete_segment_bytes are
recorded by whatever sweep last ran, so they used to read NULL until a restart
had reconciled the view. They now fill in from the first seal. That is better
observability and a visible change, and LiveViewSmokeTest states it.Costs the change adds, stated because they are real:
unlink syscalls spread across the process rather than batched at
restart. Same total, different distribution.Measured on the test workload (one boundary per commit, 32 seals, no restart and
no explicit reconciliation): the catalogue ended holding 1 zero-reference
segment whose file was still on disk, against 59 for the same run with the
cadence disabled - the one is what the fallback slot still protects, and the 59
is about two per seal that nothing was going to collect until the process ended.
meta/ ended at 69 files rather than 128, and held 20 at seal 8 rather than 32.
The 69 that remain are retained state, not garbage: with no retention horizon set
every boundary is live, and its metadata with it.
What Phase 5 does not change: the final-orphan pass.
cleanupOrphans / purgeFinalOrphans still run only at reconciliation. They
collect the final-name files a crashed or failed publication leaves above the
valid slots' id ceilings, which is a crash-recovery rule rather than a
steady-state one - but a publication that fails inside a running process leaves
the same shape, and those files still wait for a restart. Smaller than what the
cadence now collects, and out of scope here. Phase 6 (afcf762b2a, section
5.7) collected them, and found the wait was not the whole of it: the rule those
files waited for stops being able to name them once any later publication has
stepped its allocation over their ids, so they were not late but lost.
Not a planned phase either: this is the leftover section 5.6 names, and reading it against the code turned it from "collected late" into "not collected at all".
Three publications can leave a final-name file behind - retention, compaction and
repair, two of them since task 1 removed retention - and none of them re-arms a
reconciliation the way a failed append does. So nothing read the id ceiling that named their files, and the next seal's
skipPublishedSegmentIds stepped over them and raised the ceiling past them. The
id-ceiling rule then has no way to tell those files from live segments, at that
seal or at any restart after it. LiveViewCheckpointDataStore already recorded
the consequence in a javadoc - a compaction target abandoned under a fail-closed
catalogue read "leaks it once a publication has stepped over that id" - without
naming it as reachable from the other two.
afcf762b2a)Landed on puzpuzpuz_live_view across four production classes; ~145 lines of
production code, of which the pass itself is about eighty and the rest is a
result field and a test-only failure stage, plus ~165 lines of test. No on-disk
format change and no SLOT_FORMAT_VERSION bump: the pass reads a catalogue the
publication protocol already maintains.
The rule is the catalogue's silence, not an id comparison.
LiveViewCheckpointLifecycle.purgeUncataloguedSegments removes every final-name
file in meta/ or data/ that the newest durable generation neither catalogues
nor names as its pending directory segment. Since Phase 2a/2b the catalogue holds
an entry for every segment a published root can reach - data, tree metadata and
boundary metadata alike - and the pending directory segment is the one documented
exception, because a tree cannot list the file it is being written into. A file
outside both sets is reachable from nothing, whatever its id, and stays so
however far the ceiling travels. LiveViewCheckpointTimelineStoreWriter.sweep
runs it beside purge(), so the two halves of one sweep decide the fate of the
segments a generation named and of the files it never named.
What the implementation turned up:
nextSegmentId is what preserves
allocation order on its own. That is Phase 5's third bullet read the other way
round: unlinking never lowers nextSegmentId, and the id skip only ever moves
forward, so an id can never come back into circulation carrying a meaning some
root remembers. Removing the deferral is also what removes the leak - the
proposal-and-apply shape Phases 4 and 5 use would have re-introduced it,
because the ceiling can move between the two.isSegmentDurablyCatalogued already applied to an abandoned
compaction target, which is now the pass that finishes that job rather than the
one whose javadoc apologised for not finishing it..tmp belongs to whichever writer holds
it open - a repair capture spans several capture calls before it commits - so
ownership rather than reachability decides its fate. Reconciliation runs where
no writer can own one, which is where that decision belongs, and it keeps it.TEST_FAIL_AFTER_RETENTION_METADATA_PUBLISH fires only in
publishTruncateBelow, so the seal gets through and the retention pass behind
it does not. Task 2 moved it to publishCompaction when task 1 removed the
retention publication under it - section 5.9.Costs the change adds, stated because they are real:
meta/ and data/ per sweep, plus one O(log N)
catalogue probe per file found. Against the several segments a seal already
writes and the catalogue walk purge() already makes, but it is a second scan
of the same two directories.purge() just used, so a sweep opens the directory root twice.Measured on the test workload (one boundary per commit, a 60-second horizon, one
retention publication failed after its metadata publish): that publication left
exactly two files behind - its timeline path copy and its catalogue path copy -
and the next cadence sweep removed both. With the cadence
disabled the same two survived five further seals, by which point the durable
nextSegmentId ceiling had reached 84 against their ids of 47 and 48; a full
reconciliation then took meta/ from 70 files to 23 and left both of them
exactly where they were.
What Phase 6 does not change: reconciliation's own orphan rule.
cleanupOrphans still records the id ceiling and append0 still applies it,
which is what a directory whose first seal in this process follows a restart
needs, alongside the .tmp and crashed-repair rules only a reconciliation runs.
The two now overlap, with the catalogue rule strictly stronger for final names,
and unifying them would mean giving reconciliation a catalogue read it currently
does without. Worth doing, and out of scope here. Phase 7 (5d1dd718b6,
section 5.8) did it, and the read turned out to be one the reconciler already
had open.
Not a planned phase either: this is the leftover section 5.7 names, and reading it against the code turned "the two overlap" into "one of the two collectors is still running the rule that decays".
Phase 6 gave the cadence sweep a rule that holds however far the id ceiling has
travelled, and left the reconciler the one that does not. That is only a
duplication where a sweep runs. Where none does - purge.interval at zero, or a
view that stops sealing after the failed publication - reconciliation was the
only collector left, and it still compared ids against a ceiling later seals had
moved past. So the leak Phase 6 closed for a running process stayed open for
those two configurations, across every restart, for the life of the directory.
5d1dd718b6)Landed on puzpuzpuz_live_view across four production classes; ~15 lines of
production code and about 50 of javadoc, plus ~150 lines of test. No on-disk
format change and no SLOT_FORMAT_VERSION bump: it calls an existing pass from a
second place.
reconcile runs purgeUncataloguedSegments over the generation it adopts,
beside the purge() it already ran, so both halves of one sweep decide the fate
of the segments a generation named and of the files it never named - the same
pairing LiveViewCheckpointTimelineStoreWriter.sweep makes. The superblock it
needs is the one the reconciler already has open for the purge, and the catalogue
read is one hasUnregisteredRootSegment already makes a few lines above.
What the implementation turned up:
cleanupOrphans, and every catalogued id sits below
the ceiling by construction - a generation's nextSegmentId bounds every id it
ever allocated. So when the pass has run, cleanupOrphans finds nothing above
the ceiling and records a bound equal to it, which makes purgeFinalOrphans a
no-op. No signal has to say whether the catalogue answered: the directory it
left behind says it.hasUnregisteredRootSegment already
treats an unreadable catalogue as a mismatch and resets the directory. What is
left is a slot that lost the newest-generation race, where the pass refuses and
the ceiling rule still records - which is the one case the two genuinely still
overlap, and the weaker rule is the right one to keep there.Costs the change adds, stated because they are real:
meta/ and data/ per reconciliation, plus one
O(log N) catalogue probe per file found, and a second open of the directory
root beside the one purge() just used. Both are the costs section 5.7 already
priced for the cadence sweep, now also paid once per directory at boot.unlink work before the
view is available.Measured on the test workload (one boundary per commit, a 60-second horizon, one retention publication failed after its metadata publish, the purge cadence disabled): the two files that publication left survived five further seals - by which point the durable ceiling had stepped over both - and the reconciliation that followed removed them, where before the change it left them exactly where they were.
What Phase 7 does not change: the .tmp and crashed-repair rules. Both are
still reconciliation's alone, and both are ownership questions rather than
reachability ones - a .tmp belongs to whichever writer holds it open, and a
repair descriptor to a repair that may still be running. Reconciliation runs
where no writer can own either, which is where those decisions belong. The
catalogue rule has nothing to say about them.
524a8e833b)Not a phase: this is the scope decision of section 11 taken as code, plus the
regression guard that had to move with it. Landed on puzpuzpuz_live_view across
fourteen production classes; ~1,180 lines of production code and test removed
against ~180 added, of which the additions are almost entirely the two orphan
cases' new fixture. No on-disk format change and no SLOT_FORMAT_VERSION bump -
deliberately, and that is the one non-obvious part.
What came out, and it is exactly what task 1 listed: truncateBelow and its
recursion and result pool, pruneBelow with its read-only probe and subtree
release, the retainSuffix helper both leaned on in each of the two node classes,
publishTruncateBelow and RetentionResult, maybeTrimCheckpointTimeline and its
call in the refresh worker's post-seal maintenance block, and
cairo.live.view.checkpoint.retention.micros end to end. The tests went with the
code: LiveViewCheckpointRetentionHorizonTest whole, the retention cases in the
timeline and delta tree suites, and the fuzz arm.
What stayed, for the reason task 1 gave: retiredCheckpointCount and the
format version carrying it. checkpoint_timeline_entries was already wrong before
Phase 3 - publishTruncate has always dropped boundaries without adjusting
nextCheckpointId - so the field is a correction to a pre-existing bug that
happened to arrive with retention, and the high-side truncate maintains it. The
javadoc on both now names only the high-side truncate, since it is the only
publication left that retires a boundary at all.
Task 2 did not go where the plan sent it, and the reason is a property of the
workload rather than of compaction. The plan said to add the equivalent failure
stage to publishCompaction and point the two orphan cases at it. The stage moved
as planned - TEST_FAIL_AFTER_COMPACTION_METADATA_PUBLISH - but the cases could
not simply keep their old fixture, because compaction refuses a segment whose live
bytes equal its file length, and in the suite's ROWS 3 PRECEDING workload every
data segment holds exactly one state page: the commit touches one key, Phase 1's
elision writes only that key's page, and a one-page segment is either wholly live
or wholly dead. No candidate ever qualifies, so compact returned
Result.NOTHING and the injection never fired.
The two cases therefore build the history compaction needs, which is the recipe
LiveViewCheckpointCompactionTest already proved: a RANGE 30 SECOND PRECEDING
view whose ring shares chunk pages across boundaries, forty in-order seals two
keys wide, then three pairs of overlapping corrections deep in the history. Each
repair re-versions some of the roots naming a shared chunk while their neighbours
keep naming it, which is what leaves a segment part live and part dead. Measured:
three candidate segments at 114 live bytes of 342, against none at all before the
corrections.
Why not a repair, which is the other non-seal publication. A failed
publishRepair returns null and LiveViewRefreshJob answers by retiring the whole
checkpoint timeline - the durable output has moved under every root it holds - so
the orphans go with it and there is nothing left to collect. Compaction is the only
non-seal publication whose failure the refresh worker absorbs and leaves the
published generation byte-identical, which is what makes it the only usable
injection point. Section 8's Phase 6 entry recorded a failed compaction as not
covered; this converts that line, and leaves the repair half uncovered for that
reason.
Both cases were re-checked red-before/green-after against a build with
purgeUncataloguedSegments stubbed out, which is the guard task 2 exists to
preserve. The io.questdb.test.cairo.lv package is green at 1419 tests.
Not a commit: the deliverable is the #6939 body, which no branch state records. Four edits, all confined to what this plan owns.
The two clauses task 3 specifies became a Tradeoffs bullet - "The checkpoint
store collects its garbage by default, and bounds nothing" - carrying the garbage
half (what is collected, on which cadence, under which key, and that
purge.interval=0 is the one configuration where a long-running process holds
superseded segments for its whole life) and the retained half (every sealed
boundary stays reachable, no live-view setting bounds it, no upstream setting
sizes it). It prices the per-seal cost as what a seal actually wrote - Phase 1's
elision for non-ring state, chunk sharing for ring state - and names the seal
cadence the defaults imply, so "a large constant, not a bound" is a number a
reader can check rather than a claim.
The old line was wrong in a second way this plan had not named. Decision 4
recorded that "operators size retention upstream" was false because no upstream
setting bounds _checkpoints/. Writing the replacement turned up that it is also
false for the live view's own table, which is the thing the bullet was actually
about: a base-table TTL evicting partitions is freeze-and-continue for the view,
so the derived rows stay. The out-of-scope entry now states both, and points at
the Tradeoffs bullet for the checkpoint store rather than folding the two into one
sentence.
The "In V1 scope" checkpoint-timeline bullet carries the mechanism, because that is where a reader meets the store: the sweep's cadence and its second caller, the three populations it collects, and one sentence saying it bounds nothing that a generation can still reach. Stating collection where the timeline is described and the bound where the tradeoffs are is what keeps the two from reading as one claim.
The test plan gained the reclamation line it had none of. The suite is
LiveViewCheckpointMetadataReclamationTest, LiveViewCheckpointSegmentDirectoryTest
and LiveViewCheckpointStatePageElisionTest; the entry names what each case
measures, that every reclamation case is red before the change it guards - with
the ring-function elision control the stated exception - and that each ends on a
restart and a from-base recompute at a zero refresh-fault count. That last clause
was checked against the code rather than copied from section 8: twelve cases, and
assertViewMatchesRecompute is called thirteen times.
What did not change: the CodeRabbit-generated summary at the foot of the body still lists "retention" among the configurable knobs. It is auto-generated and regenerates from the diff, so editing it by hand is noise.
One operational note, because it cost a cycle: gh pr edit --body-file fails
against this repository with the projects-classic GraphQL deprecation error and
writes nothing. gh api --method PATCH repos/questdb/questdb/pulls/6939 -F body=@<file> applies the same edit and returns the stored length.
Not reclamation work at all - task 4 is the narrowed key set a localized repair leaves behind, which Phase 1 only made visible. Its three steps were reproduce, classify, route, and the classification came back bug, so the routing decision had to be taken as code rather than as a note.
Step 1 reproduced it directly, and it is wider than decision 5 described. A
localized ROWS 3 PRECEDING rebuild over an eight-key view whose cold keys sit
below the replay floor left three re-versioned boundaries naming one key
each, against eight in the boundaries either side of them. That is the missing-key
half decision 5 named. The other half only shows in the argument: L comes from
a discovery that walks back far enough to warm up the output key domain
Q - the keys with a row in [R, H) - so a key with a row in [L, R) and none
above R is carried by the replay and re-imaged from the rows that happened to
fall inside the interval. It survives in the root with a truncated history rather
than disappearing from it, which no key-count assertion can see.
Step 2 classified it, and the classifying evidence is a wrong answer. A
second correction one boundary above the first takes DISPOSITION_RESUME_FROM_ANCHOR
2.0 where the recompute says 4.0. The live view's own rows, not a
metric.Only the ROWS shape is affected, and the control says so. A RANGE frame and
an anchor segment both expire by time: what a function holds at a row at or above
R came from rows at or above L, so the replay reconstructs every key, and a
key it never carried holds nothing there either - which is exactly what an absent
key restores as. The RANGE control narrows the same three boundaries, takes the
same resume off one of them, and lands on the recompute. Section 6.4's table said
half of this already, from the runtime's side: a ROWS or anchored function "cannot
be localized behind an EOF bound" because "a key with no row at or above R keeps
state a replay from L never sees". The finite-H rule answers that for the
runtime by restoring the scratch overlay over it. Nothing restores a published
root, and that is the whole of the bug. The anchor is the arm that argument groups
with ROWS and the data does not: the recognized anchor is a calendar-period floor
of the designated timestamp, so a key that skipped the segment carries a
strictly older anchor value and its next row resets it - which is also why the
anchor root needs no key filter once task 5 gives the function roots one.
Step 3 routed it: the gate lands in #6939, the splice does not come back yet.
LiveViewCheckpointRepairPlan.isReplayStateKeyComplete() is false for any
localization a ROWS arm took part in, and LiveViewRefreshJob reads it beside the
finite H when it decides whether to open a RepairCapture. A ROWS repair
therefore takes the disposition a localized repair with no converged suffix
already took: truncateOrRetireTimelineOnO3 keeps every root below R, drops the
rest, and the post-replay seal puts a head back at the frontier. That combination
Costs the change adds, stated because they are real:
R stand, so a later resume can still anchor under the
correction, but the boundaries between R and the frontier go and one head
replaces them. That is a resume-reach and restart-recovery cost, and it is the
price of not publishing a root that describes a key set the view does not have.LiveViewCheckpointTimelineRepairTest now assert the
truncate through one shared helper, LiveViewCheckpointSoakTest's ROWS arm
expects no compaction redirect where the RANGE arm still requires one (a
boundary dropped whole fragments no segment), and Phase 1's repair-capture
elision case moved to an anchored view - the shape that is still both
whole-state and spliceable - so the capture-side comparison keeps its coverage.The #6939 body carries it in three places, on the same day and through the
gh api --method PATCH route section 5.10 records. The checkpoint-timeline bullet
now splits what a repair does with the timeline by the dependency that bounded it,
where it used to say a repair "re-versions only the roots inside" its interval
flatly; a Tradeoffs bullet states the cost, that a ROWS-framed view loses the
boundaries between the correction floor and the frontier and so anchors lower on
the next correction; and the test plan gains the suite, its red-before numbers and
the RANGE control that makes the rule a dependency question rather than a blanket
refusal. The localization claim itself did not move: a ROWS repair still reads and
re-emits its interval and nothing else.
What the fix is not. Task 4 said to size it "against capture and the scratch
overlay". Neither is where it went. RepairCapture freezes what it is handed, and
the overlay is the reason the runtime survives what the roots do not; the defect is
that the caller handed a capture a replay that could not describe every key. The
precise repair - keep the splice and re-version only the keys of Q, leaving every
other key's entry as the old root wrote it - needs Q itself at publication time,
and LiveViewCheckpointRowsBounds holds it only as a count on the walk path.
That is task 5.
7d6f5f4bc3)Landed on puzpuzpuz_live_view across six production classes; ~150 lines of
production code and about as much javadoc, plus ~250 lines of test. No on-disk
format change and no SLOT_FORMAT_VERSION bump: which keys a root re-versions is
a property of the publication rather than of the layout, and a root that keeps an
old entry keeps it by the page reference the copy-on-write already used.
The rule landed at the freeze rather than at the build, which is not where task 5
sent it. The plan said removeMissingPartitions removes only keys in Q and a
frozen key outside Q keeps the old entry - two filters. One suffices, because a
key outside Q need not be imaged at all: freezeFunction skips it before it
encodes a page, so the root's put loop never sees it and only the removal learns the
domain. That is strictly better than the two-filter shape it replaced - a repair over
a wide key set now writes pages for the keys of Q rather than for every live key,
where before task 4 it wrote all of them and published a wrong answer for most.
The encoding is what the plan did not price, and it is most of the work. The
discovery keys its scans by the reader's table-local SYMBOL integer - deliberately,
because the ids are stable for exactly the scope one repair plans in and an integer
key is what the scans are fast on. A window function's partition map keys the same
column by its resolved string, because a live view's partition-by sinks rewrite
SYMBOL that way (segment-local ids would collide across refresh cycles). The two
never had to agree before, and comparing them at publication time would have matched
nothing at all - every key would have read as outside Q, leaving every re-versioned
root byte-identical to the one it replaced. So:
LiveViewCheckpointRowsPlan carries a second projector - the same column
filter with writeSymbolAsString set, and STRING in place of SYMBOL in its key
types. A plan with no SYMBOL key column shares one sink, and an expression-keyed
plan always did: a SYMBOL key function is already written through its resolved
string, because the integers a function hands out index its own map rather than
the reader's.LiveViewCheckpointRowsBounds fills a second map through it, once per distinct
key rather than once per row, so the forward and backward scans keep the encoding
they were built for and only a key joining Q pays for the string.collectOutputKeys then writes each key with LiveViewSnapshotKeyCodec off that
map's record - the identical call freezeFunction makes off the window function's
own map record, at the same start index and over the same key types. The two sides
agree by construction rather than by inspection.The gate has two ways to be satisfied now, and the first keeps its meaning.
isReplayStateKeyComplete() still means the replay reconstructs every live key, and
a RANGE or anchored view still earns the splice with it alone.
LiveViewCheckpointRepairPlan.getOutputKeyDomain() is the second way: a ROWS
localization publishes the set instead of the guarantee, and LiveViewRefreshJob
opens a capture on either.
One refinement to what task 5 asked for. It said a discovery stopped by either
scan budget knows an incomplete Q. Only the forward budgets do. Q is collected by
the forward pass, and both of its stops - the row budget and the key budget - also
leave H at end-of-frame, so a fragment cannot reach a publication under any
configuration. A backward stop leaves Q whole and drops L to S, which is the
floor a replay reconstructs every key from, so it costs the repair its localization
below rather than its key domain. isOutputKeyDomainComplete() therefore reports the
forward pass alone, and collectOutputKeys refuses rather than returning a fragment.
Costs the change adds, stated because they are real:
RecordSink per compiled view, for a view whose ROWS plan has a
SYMBOL partition-key column. Views with none, and expression-keyed views, share the
one sink they already had.Q during the discovery, and one set copy of
Q per repair capture. Both are bounded by
cairo.live.view.checkpoint.repair.scan.max.keys (default 100,000) and paid once
per repair rather than per row.Q writes strictly less than it did before task 4.logicalStateBytes stops moving for a key-filtered splice. The freeze counts
the keys it imaged, which is a share of the boundary rather than the whole of it,
and recomputing the whole would need the old root's per-key state sizes in the
freeze's own units - a ring entry's logical size is the row stream it holds rather
than the pages it stores. So a re-versioned boundary keeps the total it already had.
That is a stat rather than a reference count: nothing reads it to decide what a
generation reaches.Measured on the eight-key ROWS fixture task 4 built: the correction that used to
re-version three boundaries down to one key each now leaves all forty boundaries
standing and every one of them naming all eight keys, and the resume off one of them
still lands on the recompute. On the twenty-four-key elision fixture, a repair
crossing nine boundaries images the four keys of Q and shares a page across the
boundaries none of them was touched at, where before task 4 it imaged all
twenty-four.
The #6939 body carries it in the three places task 4 wrote into, through the
gh api --method PATCH route section 5.10 records. The checkpoint-timeline bullet
now says what the ROWS arm does with the timeline rather than what it cannot do -
the publication is handed the keys the replay describes and keeps every other key's
entry - and names the one case that still rebuilds, a discovery a scan budget left
short of the whole set. The Tradeoffs bullet task 4 added is gone, and the counter
this change does move took its place: checkpoint_timeline_logical_bytes stops
tracking what a ROWS repair re-images, stated as an observability figure rather than
as a reference count, and stated beside the physical counters it does not affect. The
test plan entry gains the carried-but-truncated case and the encoding cases, because
the encoding is the part a reader cannot infer from the rule.
The original three semantic questions came back clear: the WAL purge floor is independent of the oldest boundary, no boundary has distinguished anchor status, and the below-horizon repair fallback already exists. The later implementation review found a separate row-position-delta preservation requirement, resolved in section 6.5 and Phase 3 step 3.
LiveViewCheckpointSuperblock.select():527-542 computes
walPurgeFloor = min(normalizedBaseSeqTxn) across the two valid A/B slots.
normalizedBaseSeqTxn is a plain scalar field in the 176-byte slot
(SLOT_NORMALIZED_BASE_SEQTXN_OFFSET), supplied by the caller as an append
parameter (LiveViewCheckpointTimelineStoreWriter.java:134). Nothing derives it
from the timeline tree, so no boundary - oldest or otherwise - feeds it.
publishTruncate already demonstrates the safe pattern: it assigns
generation, nextSegmentId, the byte counters and the two root refs, and
deliberately leaves the watermarks alone, with the comment "The base and
live-view watermarks carry forward unchanged: this publication moves no
coordinate" (:795-800). A publishTruncateBelow that follows the template
inherits the property. No new hazard, and no verification burden beyond
mirroring the existing code.
Anchors are per-boundary, not global: each checkpoint root carries its own
anchorRootRef (LiveViewCheckpointTimelineStoreReader.java:515-517, :599-607),
holding the anchor map as of that boundary's timestamp. No entry has special
status.
truncateAbove's javadoc line "the tail roots go, the long-term anchors stay"
(LiveViewCheckpointTimelineWriter.java:51-57) is descriptive of which end that
operation preserves, not a structural dependency on the oldest entry. Dropping
the low end costs resume reach, not correctness.
It is a named disposition, not an error path: DISPOSITION_BOUNDARY_REBUILD
(LiveViewCheckpointRepairPlan.java:288), documented as "The residual O(view age)
fallback: no sealed anchor sits below the change, the trigger carries no timestamp
to search with, or the apply-ahead range cannot be classified."
The lookup already reports absence rather than raising, on three separate routes
(LiveViewRefreshJob.TimelineAnchorSource.findAnchorBelow:9128-9166):
predecessorLvRowPosition returns LONG_NULL when the generation holds no
boundary below the correction (documented at
LiveViewCheckpointTimelineStoreReader.java:368-371), which maps to falsecatch (Throwable) logs "live view checkpoint timeline holds no resume
anchor" and returns falseThe class javadoc states the intent outright: "A view with no readable timeline - never sealed, retired by an earlier repair, or corrupt - reports no anchor rather than raising, so the plan takes the rebuild it would take for a change below every boundary." A retention horizon produces exactly that condition, which is already exercised.
This came out of tracing 6.3 and it materially lowers the cost of a horizon.
LiveViewCheckpointRepairPlan's javadoc (:88-155) documents three localization
shapes, none of which consults the checkpoint store for its lower bound:
| Shape | Lower bound L | Needs a sealed anchor? |
|---|---|---|
RANGE W PRECEDING | R - W, key-independent arithmetic | no |
ROWS N PRECEDING | discovered by RowsBoundSource over the pinned snapshot | no, but needs a provably insert-only change set and a FINITE high bound |
| anchored | the anchor segment walls holding R and changeMaxTs | no |
Quoting :89-96: "The L/R split is what makes a correction older than every
sealed anchor local. A boundary rebuild has no anchor to restore from, but a
bounded RANGE W PRECEDING ... CURRENT ROW view needs none ... The finite
dependency, rather than checkpoint availability, provides the lower bound."
The plan then prices the two dispositions rather than preferring availability
(:156-168): it derives the rebuild bounds even when an anchor is available,
prices both intervals through ScanCostSource against the same pinned snapshot,
and takes the cheaper. And it names the case where keeping old anchors actively
loses: "the anchor a cadence leaves just below an old correction is exactly the
anchor whose resume replays every row above it."
So the horizon's cost is not "deep corrections get expensive". It is narrower:
H: a correction below the
horizon costs nothing extra. The dependency supplies the floor.EOF: the union
sinks to EOF, the floors collapse to the START FROM boundary S, and the
repair reads the whole view history. This is the only population that pays, and
it pays only on corrections below the horizon.That is the input open decision 2 was missing.
One correction to this table, from writing the Phase 3 tests (section 5.4). The
population named above was originally "an unbounded-preceding aggregate, or a ROWS
arm whose high bound comes back EOF". The first half does not exist: a bare
unbounded window is rejected at CREATE LIVE VIEW, which requires an ANCHOR, and
an anchored window localizes off the anchor wall. What does not localize is a ROWS
dependency over a DEDUP base - the change set cannot be proven insert-only, so
the discovered bound is unavailable. That is the shape the census in decision 2
should be looking for.
LiveViewCheckpointRowPositionDeltaWriter's only mutation is suffixAdd
(:123) - no truncate, no prune. It is written on the repair path only
(LiveViewCheckpointTimelineStoreWriter.java:628-641) and accounted as
rowPositionDeltaBytes, documented as a share of metadataBytes
(LiveViewCheckpointSuperblock.java:154-159).
It is another append-only metadata B+ tree with no reclamation, repair-driven
rather than seal-driven. Each suffixAdd creates one difference breakpoint and
path-copies O(log R) pages, rather than rewriting the logical suffix. It is
smaller than the per-seal surfaces but still unbounded on an O3-heavy workload.
Phase 2a counts its pages like the other two direct trees, so a repair's
superseded delta pages are now reclaimed, and Phase 3's pruneBelow prunes the
index itself when low boundaries retire.
Note the asymmetry this creates against the publishTruncate template: that
method carries rowPositionDeltaRootRef forward unchanged because "dropping the
suffix moves no surviving prefix key's cumulative recovery position", and clears
seedCursorOffset because "a truncate leaves no mid-sweep resume point behind".
A low-side truncate carries the seed offset forward, but it cannot merely delete
delta entries keyed to dropped boundaries. effectivePosition adds the prefix
sum of all earlier differences, so an old breakpoint still contributes to every
surviving suffix key. The prune must fold the discarded prefix sum into the first
surviving key as Phase 3 step 3 specifies.
Metadata reference counts live in the segment-directory spine, which is published copy-on-write and named by the superblock slot, so ordinary reference deltas commit atomically with the generation exactly as the data counts already do. The existing fsync order - data pages, then metadata segments, then the superblock that names them - remains unchanged.
Phase 2a nevertheless adds one crash-safety case: deferred self-registration of the selected directory segment. Reconciliation must accept that exactly the currently selected directory-root segment is absent from its own catalogue; any other reachable unregistered metadata segment is corruption. The next publication registers that segment before publishing its successor, and a crash before the successor superblock commits leaves only final-name orphans above the valid slots' id ceilings, which the existing orphan sweep removes.
As shipped, the exception is named rather than inferred: the slot carries
pendingDirectorySegmentId and bounded slot validation rejects a half-set
triple, so a reconciler can check the rule instead of deducing it from what the
catalogue lacks. Because the field is part of the atomically published slot, a
crash cannot leave the pending registration and the root it describes
disagreeing: either both generations' worth of state commits or neither does.
Phase 2b checks that rule, for the three roots the superblock names directly. A generation naming a metadata segment its own catalogue does not hold - the pending directory segment excepted - resets the checkpoint directory and rebuilds from the base table, which is the disposition a foreign layout already takes and costs a rebuild rather than correctness. The boundary half of the rule holds by construction now that a boundary's segments are registered by the publication that writes them, but it is not checked: proving it needs one partition-map walk per surviving boundary, which is the sweep the whole accounting exists to avoid.
A segment may be unlinked only after the superblock generation that stopped
referencing it is durable. PurgeSweep's existing
oldestValidSlotGeneration gate already expresses this; metadata entries inherit
it unchanged.
Phase 1 is crash-neutral - it only decides whether to write a page or reuse a ref, both of which are captured in the same atomic publish.
Phase 4 adds no case either, and the reason is worth stating because the removal looks like it should. An entry retirement commits with the generation carrying it, and the fallback slot keeps its own copy of the entry - at a zero count, so no root of that generation reads the missing file, and a sweep over it re-proposes the retirement rather than faulting. The proposal itself is deliberately not durable: the sweep re-derives it from what the catalogue and the directory disagree about, so a crash between the sweep and the publication that would have applied it costs one more sweep and nothing else.
Phase 5 adds none either, and for the reason that makes it a small change: a sweep commits nothing. It unlinks files the purge rule proves unreachable and leaves a proposal in memory, so a crash mid-sweep leaves a directory the next sweep re-derives the same answer over, and a crash between the sweep and the seal that would have applied its proposal costs one more sweep. The rule the sweep applies is generation-scoped, not sweep-scoped, so running it more often cannot make it collect anything it would not have collected once.
Phase 6 adds no case, and the reason is worth stating because it removes one
rather than adding it. The id-ceiling rule needs the deferral - a final-name file
above the ceiling goes only after a new slot commits above it, so a crash between
the decision and the removal leaves the file for the next reconciliation to
decide on again. The catalogue rule needs no deferral, because a file the newest
durable generation does not catalogue is unreachable from every root of that
generation whatever happens next, and the monotonic nextSegmentId preserves
allocation order without help: unlinking never lowers it, and the id skip only
ever moves forward, so no id comes back into circulation carrying a meaning some
root remembers. A crash mid-pass therefore leaves a directory the next sweep
re-derives the same answer over. What the pass will not do is act on a partial
answer: a catalogue it cannot read, a superblock naming no directory root, or a
slot that lost the newest-generation race all leave every file where it is.
Phase 7 adds no case for the same reason, from a caller that already had it. A reconciliation publishes no generation either, so the pass commits nothing there that it does not commit on a sweep, and a crash mid-reconciliation leaves a directory the next reconciliation re-derives the same answer over. The one question it raises that a sweep does not - whether an adopted generation is enough evidence to unlink under - is answered by the invariant the catalogue already maintains: the newest generation's catalogue holds an entry for every segment any valid slot can reach, because an entry retires only after the purge rule proved the fallback slot had passed the retirement generation.
Phase 3 adds no case of its own. A retention pass is one A/B swap over an unchanged fsync order, so a crash before it commits leaves the pre-retention generation whole and the segments it wrote as final-name orphans above the valid slots' id ceilings, which the existing orphan sweep removes. It differs from the high-side truncate in needing no repair marker beside it: that publication deliberately leaves the superblock naming a head it discarded, while a retention pass moves no coordinate at all - the head, both watermarks, the checkpoint id counter and the seed resume point all carry forward - so there is no window in which the committed generation describes something the runtime cannot restore.
The review recorded that no test asserts any bound on meta/ file count. That is
the headline gap, and the correction it also recorded matters:
LiveViewCheckpointCompactionTest.purgeCycle calls reconcile directly four
times, but that is not why nothing catches this - purge() is a data-segment
sweep by construction and would never delete metadata however often it ran.
New tests, per phase:
Phase 1 - all landed in LiveViewCheckpointStatePageElisionTest
data/ and meta/ byte growth is proportional to touched keys, not to
N x live_keys. Red before the change. Done, stated structurally rather
than as a byte threshold: exactly one map entry may change per seal, the run
writes live_keys + seals distinct pages, and data/ grows by exactly the
touched key's pages. The metadata half follows from the entry being
byte-identical, which is what the partition-map writer's own tested elision
keys on.Phase 2a - all landed in LiveViewCheckpointMetadataReclamationTest except
the format case, which went to LiveViewCheckpointLifecycleTest
meta/ file count stays bounded. This is the test the
review asked for. Done, stated as two measurement points rather than an
absolute: 32 seals against 8, with the live metadata segment count allowed to
move by at most 2 (see section 5.2 on why it is not flat) and meta/ growth
required to be under four files per seal. Red before the change: the sweep
reclaimed nothing.TEST_FAIL_AFTER_METADATA_PUBLISH, which is the stage that leaves
orphan metadata segments; the other two stages are covered by
LiveViewCheckpointTimelineSealTest, which now checks the catalogue by kind._timeline retires and rebuilds rather than erroring. Done
in LiveViewCheckpointLifecycleTest.testSupersededTimelineVersionResets...,
which stamps the previous magic and version rather than a later one, so it
covers the migration direction this bump actually creates.Phase 2b - all landed in LiveViewCheckpointMetadataReclamationTest
LiveViewCheckpointRetentionHorizonTest states it as the boundary count and the
file counts both staying flat while the view keeps sealing.publishTruncate orphans get reclaimed (the two Class A
entries with O(live_keys) volume). Done, one case each, told apart by the
property that distinguishes the two publications from the outside: a splice
preserves every logical key and a truncate drops a suffix of them. Both go red
when the sweep is made to skip boundary segments.Phase 3 - landed in LiveViewCheckpointRetentionHorizonTest, plus tree-level
cases in LiveViewCheckpointTimelineTest / LiveViewCheckpointRowPositionDeltaTest
and one arm in LiveViewFuzzTest. All removed with the mechanism by task 1
(524a8e833b); kept here as the record of what a retention proposal would have to
cover again.
data/ and meta/ both shrink; the view
still restores from the newest boundary after restart. Done, stated as
flatness rather than as a shrink: the boundary count is measured at 10 seals and
again at 30 and must be identical, and meta/ and data/ may each grow by at
most a quarter of a file per further seal against the four a seal writes. The
residual is the catalogue's own tree, which retired no entry when this case was
written; Phase 4 makes it retire them, and the bound the case states is loose
enough to hold either way.DISPOSITION_BOUNDARY_REBUILD and
lands on the same answer as a from-scratch recompute. Done, and the case had
to change shape to be honest. A bare unbounded window is rejected at CREATE
(live views require an ANCHOR), and a ROWS N PRECEDING view over an ordinary
base localizes below the horizon - it reports BOUNDARY_REBUILD with
DENIAL_NONE, which is a localized rebuild and exactly section 6.4's point.
The genuinely denied case needed a DEDUP base, which is what makes a change set
unprovable as insert-only and leaves the ROWS dependency with no floor. That is
the population that pays, and the case asserts a non-NONE denial rather than
the disposition alone.DENIAL_NONE, against a view with nothing sealed under the change.effectivePosition lookup on every surviving boundary returns the same value
it returned before the prune. Include multiple discarded breakpoints, an
existing breakpoint at the first survivor, and a first survivor with no prior
breakpoint, so prefix folding is non-vacuous. Done at the tree level, where
the invariant can be stated directly: each case snapshots prefixSum at every
key at or above the floor, prunes, and requires each to be unchanged. All three
shapes are covered, plus the tree-goes-empty case and randomized rounds.metadataBytes, dataBytes and rowPositionDeltaBytes remain cumulative
across reclamation; live/obsolete catalogue metrics and actual directory bytes
reflect the shrink instead. Done for the cumulative half and for
logicalStateBytes, which is the counter that does shed retired boundaries.seedCursorOffset survives a low-side truncate taken mid-sweep, and the sweep
resumes from it. Partly done. The carry-forward is asserted directly, by
publishing a generation carrying a resume point and reading it back after the
retention pass. That a sweep then resumes from it is not covered: driving a
seed sweep to a mid-sweep seal and a retention pass in the same case is a
fixture this suite does not have.LiveViewFuzzTest shape. Done -
LiveViewFuzzTest.testFuzzCheckpointRetentionHorizon, one boundary per row
against a horizon of at most five seconds of event time, so most boundaries
retire while the fuzz is still ingesting.Phase 4 - four cases in LiveViewCheckpointSegmentDirectoryTest and one in
LiveViewCheckpointMetadataReclamationTest
LiveViewCheckpointMetadataReclamationTest.testTheCatalogueRetiresTheEntriesOfSweptSegments,
stated as an identity rather than as a bound: it names the exact set the sweep
left dead, requires every one of them to be gone afterwards, requires the
catalogue to have shrunk, and requires every survivor to name a file. Red
before the change, on the first of those.TreeMap oracle, at node capacity three so a retirement empties leaves a
neighbouring insert just split into. Done, 120 rounds, checked through a
freshly bound reader per round so a corrupted reused subtree cannot survive to
the end of the walk.Phase 5 - two cases in LiveViewCheckpointMetadataReclamationTest, plus one
column-value update in LiveViewSmokeTest
testTheCadenceSweepCollectsWithoutARestart, stated as three things the
reconcile-only build cannot produce: at least one file present at seal 8 is
gone by seal 32, the catalogue ends with at most a quarter of a seal's worth of
zero-reference segments whose files are still there (measured: 1, against 59),
and the entries the last sweep proposed are all gone after one further seal.
Red before the change on the first of those.testTheCadenceSweepStaysOffAtIntervalZero, which also fixes the reconcile-only
numbers the case above is measured against.live_views() reports the collection columns from the first seal rather than
from the first restart. Done in
LiveViewSmokeTest.testLiveViewsCatalogueExposesCheckpointTimelineColumns,
which asserted the pre-change NULLs directly.Phase 6 - two cases in LiveViewCheckpointMetadataReclamationTest
testTheCadenceSweepCollectsTheOrphansOfAFailedPublication, which fails a
compaction after its metadata publish, names the exact set it left
uncatalogued, seals past it, and requires every one of them to be gone and
nothing uncatalogued to remain. Red before the change, on that set surviving.
Task 2 (524a8e833b) moved it there from a failed retention publication and
gave it the fragmentable history compaction needs - see section 5.9.nextSegmentId ceiling has moved past every one of them. Done -
testTheIdCeilingRuleCannotReachTheOrphansOfAFailedPublication, which asserts
the ceiling has stepped over each id before asserting the file is still there,
so a run where the seals happened not to step over them fails as
inconclusive rather than passing. Phase 7 changed its tail and its name: the
full reconciliation it ends with now collects them, so the case is
testAReconciliationCollectsTheOrphansTheIdCeilingRuleCannotReach and states
both halves - which rule cannot name them, and which one does.TEST_FAIL_AFTER_COMPACTION_METADATA_PUBLISH stage, which fires only in
publishCompaction: every other injection fires in a seal, and a failed seal
re-arms the reconciliation, which is exactly the case that was never broken.
Compaction is also the only non-seal publication whose failure the refresh
worker absorbs - a failed repair answers with a whole-timeline retire, which
takes the orphans with it - so it is the only usable injection point rather than
merely a convenient one.Phase 7 - one reshaped case in LiveViewCheckpointMetadataReclamationTest
and two in LiveViewCheckpointLifecycleTest
testAReconciliationCollectsTheOrphansTheIdCeilingRuleCannotReach
above, which is red before the change on exactly that step: it still proves the
seals stepped the ceiling over every orphan first, so the collection cannot be
the old rule firing.testOrphanCleanupCollectsWhatNoGenerationCatalogues, which needed a fixture
that publishes a real catalogue over two generations rather than the empty
leaf the suite used: a referenced data entry survives, so does the previous
generation's directory segment at a zero count against a retirement the
fallback slot has not reached, and so does the pending directory segment the
superblock names. Four uncatalogued final names and two temporaries go. Red
before the change, on the removal count.testOrphanCleanupDefersWhereNoGenerationCanAnswer, a directory with no
_timeline at all: the ceiling is zero, the bound is recorded, nothing final is
removed until a publication allocates above it, and the publication's own
segment - which sits above the bound, as a seal's does - survives the sweep that
follows. Passes before the change too, which is the point: that path is
untouched.Task 4 - three cases in LiveViewCheckpointRepairKeyCoverageTest, plus five
existing cases reshaped (section 5.11)
testARowsRepairLeavesEveryBoundaryNamingEveryKey, stated as the whole
timeline in one string, because both halves matter: the roots at or below R
survive untouched, everything above them is dropped rather than re-versioned, and
one head stands at the frontier. Red before the change, on three boundaries
coming back naming one key of eight.testAResumeAfterARowsRepairKeepsEveryKeysHistory, which corrects twice: the
first correction re-versions the boundaries under the second, and the second
resumes off one of them. Red before the change on the recompute oracle, at
2.0 against 4.0, which is the cold key's own row summing without its history.testARangeRepairStillReVersionsItsBoundaries, the same
workload over a frame that expires by time: it still splices, its roots still
come back naming one key, and the resume off one of them still matches the
recompute. Passes before and after.Task 5 - one new case and one reshaped one in
LiveViewCheckpointRepairKeyCoverageTest, two new ones in
LiveViewCheckpointRowsBoundsTest, one in LiveViewCheckpointRepairPlanTest, and
the five cases task 4 reshaped restored to asserting the splice
testARowsRepairLeavesEveryBoundaryNamingEveryKey now states the whole timeline as
forty boundaries at eight keys each, so it is red twice against the two earlier
states: the boundaries above R were dropped outright by task 4, and before that
they came back naming one key.testARowsRepairKeepsTheHistoryOfAKeyItsReplayOnlyPartlyCarried, whose warm key
has two rows below L and two inside [L, R) and none inside the replacement
interval, so the replay carries a fragment of its history rather than none of it.
A fix that only stopped the removals would publish that fragment, and the resume
off one of the re-versioned boundaries then reads it as the whole of the key's past.testOutputKeyDomainLeavesInTheCheckpointKeyEncoding asserts the collected keys of
a SYMBOL-partitioned view are the resolved strings a partition map holds and
explicitly not the four-byte symbol ids the scans key by, and that a LONG-keyed view
testOutputKeyDomainRefusesTheFragmentABudgetLeaves, which drives the key budget to
one and requires collectOutputKeys to refuse rather than answer.testRowsDependencyPublishesTheKeysItsReplayDescribes, which pairs a ROWS
localization carrying a domain against a RANGE one that carries the guarantee
instead, and a budget-stopped one that carries neither.LiveViewCheckpointTimelineRepairTest assert their per-entry root identities again
and the shared truncate helper is gone, LiveViewCheckpointSoakTest's ROWS arm
requires the compaction redirect again, and the repair-capture elision case is back
on a ROWS view. That last one needed a new fixture: the sharing it measures is now
between two boundaries at which a key inside Q was untouched, so its trickle
round-robins over four keys instead of feeding one, and its history is deep enough
that the bounded rebuild prices cheaper than a resume.Q and whose live key set coincide, which is the dense
workload the whole task exists for. Every case here deliberately has keys outside
Q, because that is where the rule is load-bearing; the coinciding case is the one
that behaves exactly as it did before task 4.Not covered, and worth naming: nothing asserts the purge half end to end under a concurrent reader pin - the generation gates are shared across the publications and are covered there, but a case that pins a generation across a publication and proves nothing it reaches was unlinked would be the direct statement. Nor does anything drive a sweep concurrently with the seal that consumes its proposal, because nothing can: both run on the refresh worker, under the same serialization the reconciliation sweep always had.
| Phase | Scope | Notes |
|---|---|---|
| 1 | landed: 2 production classes, ~300 lines, plus a 560-line test | estimate held; no format change; both the published and the in-flight previous page are covered |
| 2a | landed: 11 production classes, ~700 lines, plus a 400-line test | came in well under the "larger than 600 lines" estimate, because per-segment page counts removed the closure summaries that were supposed to dominate it; format cost is one leaf field and three superblock longs |
| 2b | landed: 10 production classes, ~250 lines, plus ~200 lines of test | came in under the estimate because the release sites needed no code: once a root states its closure, publishTruncate and the repair splice reclaim it through the reference transaction they already ran. Format cost is one field on the anchor root and a third value for the catalogue's kind |
| 3 | landed, then removed: 9 production classes, ~450 lines, plus ~900 lines of test | came in under the estimate for the same reason 2b did: retiring a boundary is the reference transaction publishTruncate already ran. Format cost is one superblock field, needed to keep the published entry count honest rather than by the plan's design - which is why that field stays behind after task 1 took the rest out |
| 4 | landed: 6 production classes, ~215 lines, plus ~370 lines of test | never estimated - it was decision 7 rather than a phase. Small because a B+ tree whose only deletion pattern is "the low ids die first" gets a correct shape out of the existing emit path, so there is no rebalancing rule and no format cost at all |
| 5 | landed: 8 production classes, ~290 lines, plus ~190 lines of test | never estimated either - it was the leftover Phase 4 recorded. The reclamation logic is about forty lines; the rest is javadoc, a result class and the five files a new config key touches. No format cost, and no new rule: the purge rule is indifferent to how often it runs |
| 6 | landed: 4 production classes, ~145 lines, plus ~165 lines of test | never estimated either - it was the leftover Phase 5 recorded. The pass is about eighty lines, and it is small because the catalogue already states what it needs: the reachability rule was written by Phase 2a/2b, and this only reads it from the other side. No format cost, and no new state - it decides and acts in one pass |
| 7 | landed: 4 production classes, ~15 lines of code and ~50 of javadoc, plus ~150 lines of test | never estimated either - it was the leftover Phase 6 recorded. The smallest phase by a wide margin: the pass, the superblock it needs and the catalogue read all existed, and the two rules separate themselves by ordering rather than by a flag. Most of the work is in the test fixture, which had to learn to publish a real catalogue |
| tasks 1+2 | landed: 14 production classes, ~1,180 lines of code and test removed against ~180 added | never estimated - it is the scope decision rather than a phase. Almost all of the addition is the fixture the two orphan cases needed once their injection point moved to compaction, which accepts only a part-dead segment and so needs ring sharing and overlapping corrections that the suite's other cases do not produce |
| task 3 | landed: no code at all - four edits to the #6939 body | never estimated - it is the deliverable's description rather than the deliverable. Sized here only to record that the reclamation work now has a public statement, and that it needed a Tradeoffs bullet rather than a corrected sentence, because the garbage half and the retained half are two claims and the old line collapsed them into one |
| task 4 | landed: 3 production classes, ~15 lines of code and ~45 of javadoc, plus ~250 lines of new test and five existing cases reshaped | never estimated - it is pre-existing repair behaviour rather than reclamation work. The code is one predicate on the plan and one term on the gate that opens a repair capture; the weight is in the tests, because proving the defect needs two corrections and a control, and because disabling the ROWS splice moved every case that asserted it |
| task 5 | landed: 6 production classes, ~150 lines of code and about as much javadoc, plus ~250 lines of test and task 4's five cases restored | never estimated either. Larger than task 4 and for a reason the task statement did not name: the publication filter itself is one skip in the freeze and one guard in the removal, and the rest is getting Q into the encoding a partition map keys by - a second projector on the plan, a second map in the discovery, and the collection that writes one through LiveViewSnapshotKeyCodec off the other |
LiveViewCheckpointCompaction repacks live state pages into a fresh data
segment, touches metadata only via an exists() probe at :240, and
publishCompaction then writes new root, timeline and directory segments and
adds to metadataBytes. Setting the interval non-zero does not bound meta/
at all.O(log N) path copy into a full spine rewrite, which is the write amplification
Phase 1 exists to remove.WindowFunction interface as the way to get Phase 1.
Rejected for now in favour of the byte comparison, which stays inside the
checkpoint storage layer rather than changing every window implementation.
Worth revisiting as an optimisation once Phase 1 proves the shape.O(N / leafCapacity) and grows with the timeline, so writing the summary every
seal makes the metadata cost quadratic in the seal count - reintroducing, in a
smaller constant, the growth Phase 1 exists to remove. Per-segment page counts
give the same answer incrementally and need no format change. The summary shape
is not rejected for boundary metadata, where the closure is bounded by the key
set and bulk release is the operation that matters - and that is what Phase 2b
took, for the soundness reason section 5's Phase 2b preamble states rather than
for the cost.O(seals x live_keys) pages - so a sweep is
proportional to the whole history. It is the same objection as metadata
compaction, one level up.Settled on 2026-07-31, and it changes what this plan delivers. The brief was
metadata GC - reclaiming the garbage in _checkpoints/meta/ - and the plan
widened it in section 1 to bound retained state as well, because Class A alone
leaves the store growing. That widening is withdrawn: the retention horizon is
not shipping. The deliverable is Class A, which Phases 2a, 2b, 4, 5, 6 and 7
close and which is on by default. Task 1 in section 12 removed Phase 3
(524a8e833b, section 5.9), and decisions 1 and 2 below went with it. Section 1's
correction to the scope stands as a record of why the horizon was proposed, not as
something this plan does.
cairo.live.view.checkpoint.retention.micros, because the
floor it implies is one O(log N) probe while a boundary count would need
navigation the timeline reader does not expose (section 5.4). The end state
this decision named - deriving that window from the view's TTL once TTL lands,
so the checkpoint horizon never sits below data the view still retains - has no
horizon to apply to once task 1 lands. It comes back only if checkpoint
retention is proposed again as its own change.ROWS N PRECEDING view over an
ordinary base still localizes below the horizon. What did not localize was the
same view over a DEDUP base, where the change set cannot be proven
insert-only. So the census is narrower than "unbounded-preceding aggregates":
it is closer to "views over a DEDUP base, or with a ROWS arm whose high bound
comes back EOF". It has still not been done.a2712f7217 on puzpuzpuz_live_view, and Phases 2a, 2b and 3 followed it on
the same branch. None of them needed data-page access outside the checkpoint
storage layer._checkpoints/. Phase 3 made the first
clause conditional rather than false, and removing Phase 3 makes it flatly true
again for retained state: with no horizon there is no live-view config key that
bounds the checkpoint store. What did not revert with it is everything below,
which is why the honest line is two clauses rather than one - the garbage half
is bounded and on by default, and the retained half is not bounded at all.
Phase 4 closed the residual that used to
qualify even that - the segment catalogue's own tree - so nothing is left that
grows with the view's age as opposed to with what it holds.
Phase 5 closed the last qualifier after that: unlinking used to wait for a
restart, so a long-running process held every superseded segment whether a
horizon was set or not, and now it collects on a seal cadence
(cairo.live.view.checkpoint.purge.interval, default one, zero to disable).
Phase 6 closed the term that was not a qualifier at all: the files a failed
compaction or repair publication renamed into place were never
collected, in this process or any later one, because the rule naming them
compares against a ceiling the next seal moves past. The same cadence sweep
now removes them under the catalogue's own reachability rule, and Phase 7 gave
reconciliation the same rule, so the sentence no longer needs the qualifier
"with the cadence on": a process that never sweeps collects them at its next
restart. Nothing in _checkpoints/ is left that a running process cannot
reclaim, and nothing that a restart cannot either; what a restart still owns
alone is the .tmp and crashed-repair rules, which are ownership questions
rather than reachability ones.a8bc16da33, section 5.5. The answer was the
one the decision stated: the sweep proves an entry dead by unlinking its file,
and since it publishes no generation of its own, it hands the ids to the next
cadence seal, which removes them in the same path copy it was already making.
What the implementation added to that was the failure discipline - the sweep
re-proposes every entry whose file is already gone, so the hand-off needs no
durability - and the observation that no rebalancing rule is needed, because a
node that empties simply writes no page and its parent keeps no reference to
it. ec6493a268, section 5.6. It did move both at once, and
needed no new machinery to - the hand-off Phase 4 built for reconciliation is
the one a cadence sweep uses, and the purge rule that decides what may go is
generation-scoped rather than sweep-scoped, so calling it more often collects
sooner and nothing else.[L, K] - and a ROWS repair now truncates the timeline at R instead of
re-versioning boundaries it cannot describe. Q: a ROWS repair re-versions the entries of the keys its replay
describes and leaves every other key's entry as the old root wrote it, so the
blanket refusal is gone and the correctness it bought stands.All five are done. Tasks 1 and 2 came from the scope decision at the head of
section 11 and landed together in 524a8e833b; task 3 followed it on 2026-07-31
(section 5.10), task 4 on the same day (section 5.11) and task 5 the same day again
(section 5.12) - and the last two are task 4's own lineage rather than anything the
scope decision named. Neither is reclamation work: task 4 was pre-existing repair
behaviour that Phase 1 made visible, and task 5 is the performance half of the fix
task 4 landed for correctness.
524a8e833b)Landed as written; section 5.9 records what shipped. The list below is kept as the statement of scope it was checked against.
Why: the deliverable is metadata GC. Retention is a policy mechanism that decides to stop keeping live boundaries, which is a different feature, and it ships disabled - so removing it costs a default install nothing and takes ~450 lines of production code, a config key and a public surface out of #6939.
d7bf14f612 is the commit, but do not expect git revert to apply: Phases 4, 5,
6 and 7 all landed on top of it and two of them reference it.
Remove:
LiveViewCheckpointTimelineWriter.truncateBelow and its recursion, result pool
and javadoc (~161 lines; the low-side mirror of truncateAbove, which stays).LiveViewCheckpointRowPositionDeltaWriter's low-side prune (~221 lines) and the
LiveViewCheckpointRowPositionDeltaNode / LiveViewCheckpointTimelineNode
helpers Phase 3 added for it (~73 lines across the two).LiveViewCheckpointTimelineStoreWriter.publishTruncateBelow and its
RetentionResult, less the retiredCheckpointCount maintenance - see below.LiveViewRefreshJob.maybeTrimCheckpointTimeline and its call in the
if (sealed) maintenance block, where retention currently runs ahead of
compaction and the sweep. The comment there explaining that ordering needs
rewriting for the two passes that remain.PropertyKey.CAIRO_LIVE_VIEW_CHECKPOINT_RETENTION_MICROS,
the PropServerConfiguration read, CairoConfiguration.getLiveViewCheckpointRetentionMicros,
the wrapper and default implementations, the server.conf template entry and
the ServerMainTest line that counts it.LiveViewCheckpointRetentionHorizonTest (627 lines), the retention cases in
LiveViewCheckpointTimelineTest (268) and LiveViewCheckpointRowPositionDeltaTest
(174), and LiveViewFuzzTest.testFuzzCheckpointRetentionHorizon (18).Keep, and this is the one piece that looks removable and is not:
LiveViewCheckpointSuperblock.retiredCheckpointCount and the
SLOT_FORMAT_VERSION bump that carries it. Section 5.4 records that the counter
was already wrong before Phase 3 - publishTruncate has always dropped
boundaries without adjusting nextCheckpointId, so checkpoint_timeline_entries
over-reported after a high-side truncate - and the high-side truncate maintains
the field too (LiveViewCheckpointTimelineStoreWriter.java:893, against the
retention site at :1081). Removing the field would reintroduce that bug. Live
views are unreleased, so leaving the format version where it is costs nothing and
avoids a second bump.
Acceptance: no retention symbol left in core/src/main outside the WAL and
base-table senses of the word; checkpoint_timeline_entries still correct across
a high-side truncate; the io.questdb.test.cairo.lv package green. All three
met - the surviving retention mentions are the WAL floor and the base table,
publishTruncate still maintains retiredCheckpointCount, and the package is
green at 1419 tests.
524a8e833b)Blocked task 1 from being complete. It is a regression guard on the GC work that is shipping, so it landed in the same commit.
Phase 6 exists for the files a publication renames into place and then fails to
commit. Proving that needs a failure in a publication that is not a seal: a
failed seal re-arms the reconciliation that reads the id ceiling, which is
exactly the case that was never broken. The only such injection point today is
TEST_FAIL_AFTER_RETENTION_METADATA_PUBLISH
(LiveViewCheckpointTimelineStoreWriter.java:94, fired at :1068), and it fires
inside the retention publication task 1 removed.
Two cases depend on it -
LiveViewCheckpointMetadataReclamationTest.testTheCadenceSweepCollectsTheOrphansOfAFailedPublication
and testAReconciliationCollectsTheOrphansTheIdCeilingRuleCannotReach - and both
are red-before/green-after guards, one for Phase 6 and one for Phase 7.
Do: add the equivalent stage to publishCompaction and point both cases at
it. Compaction is the natural replacement - it is one of the three publications
section 5.7 names as producing this shape, it re-arms no reconciliation either,
and cairo.live.view.checkpoint.compaction.interval already drives it from a
test. Section 8's Phase 6 entry currently records "the same shape from a failed
compaction" as not covered; this converts that line rather than adding to it.
What it took, beyond that. The stage moved as written, but the compaction
interval could not drive it and the cases could not keep their fixture:
compaction refuses a segment whose live bytes equal its file length, and the
suite's ROWS 3 PRECEDING workload writes exactly one state page per data
segment, so no candidate ever qualified and the injection never fired. Both cases
now build the ring-shared, correction-fragmented history compaction needs. The
choice also turned out to be forced rather than natural: a failed repair makes
LiveViewRefreshJob retire the whole timeline, taking the orphans with it, so
compaction is the only non-seal publication whose failure leaves anything to
collect. Section 5.9 carries both findings.
Acceptance: both cases still fail against a build with Phase 6's pass disabled
and pass with it, and section 8's Phase 6 and 7 entries are updated to name the
new stage. Both met, the first re-checked directly by stubbing
purgeUncataloguedSegments out.
Landed as written; section 5.10 records what shipped and the one thing writing it turned up. The statement of scope below is what it was checked against.
Decision 4 in section 11 carries the wording history. Task 1 has landed, which makes the current draft wrong again: with retention removed there is no live-view config key that bounds the checkpoint store, so the accurate statement is two clauses rather than one.
cairo.live.view.checkpoint.purge.interval,
default one, zero to disable) and again at every reconciliation. Nothing in
_checkpoints/ that a generation cannot reach survives a running process or a
restart.Both clauses are in the body, as a Tradeoffs bullet, with the mechanism half also stated where the checkpoint timeline is described and the reclamation suite added to the test plan. Section 5.10 has the detail.
Landed as the three steps below; section 5.11 records what shipped and the two things the reproduction turned up that the steps did not anticipate - that the defect also reaches keys the replay did carry, and that it produces a wrong answer rather than a debatable shape. The statement of scope is kept as what it was checked against.
Acceptance, and where it came out: step 1 reproduced the narrowing directly
(three boundaries naming one key of eight); step 2 classified it as a bug, on a
live view returning 2.0 where the recompute says 4.0; step 3 routed the gate
into #6939, because the wrong answer is reachable from code #6939 introduces, and
routed the splice's return to task 5. The io.questdb.test.cairo.lv package is
green.
This is decision 5, tracked as work rather than as a question, per the scope discussion on 2026-07-31. It is pre-existing repair behaviour, unrelated to reclamation - Phase 1 only made it visible - so it neither blocks nor is blocked by tasks 1 to 3, and it may well belong outside #6939 entirely. Establishing that is step 3.
The observation (section 5.1, found while writing the Phase 1 capture test):
RepairCapture.capture (LiveViewCheckpointTimelineStoreWriter.java:2678)
freezes the runtime's window state as the new root version of each boundary the
replay crosses. A localized repair replays [L, H) over runtime state the scratch
overlay has taken out of the way, so the state it freezes covers the keys that
range carried, not the key set the boundary originally described. A restore from
such a re-versioned boundary therefore comes back with fewer keys than the
boundary it replaced.
Steps:
[L, H) touches one key, and
compare the restored key set against the one the original boundary described.
The Phase 1 capture case in LiveViewCheckpointStatePageElisionTest is the
nearest existing fixture, but it asserts page sharing rather than key coverage.capture and the scratch
overlay and decide whether it lands in #6939 or its own PR. If it is intended,
the deliverable is javadoc on RepairCapture.capture and on the restore path,
plus a test that pins the narrowing so it cannot change silently.What is already known, and bounds the urgency: nothing in this plan reshapes a boundary. Phase 3 would have shrunk the set a restore could fall back to - fewer intact neighbours under a narrowed boundary - but task 1 removed it, so the population that meets a narrowed boundary is whatever it was before this work started. That reading of the urgency was wrong, and step 2 is what corrected it: the population is every ROWS view that takes two out-of-order corrections near each other, which needs no retention horizon and no restart.
Landed as the rule below; section 5.12 records what shipped and the two things implementing it turned up that the statement of scope did not anticipate - that the two sides encode a SYMBOL partition key differently and had to be made to agree before anything could be compared, and that filtering at the freeze rather than at the build makes one filter do the work of the two below. The statement of scope is kept as what it was checked against.
Task 4's gate is correct and blunt: no ROWS view re-versions a boundary now,
including the ones whose every key sits in Q, which is the common shape for a
dense workload. The cost is a timeline truncated at R on every out-of-order row,
so a later resume anchors lower and a restart recovers from a fresher but lonelier
head.
The rule that would restore it, derived in section 5.11 and stated here as the
thing to implement: a re-versioned boundary must take the replay's entry for a key
in Q and keep the old root's entry for every key outside it. Both halves are
load-bearing. A key outside Q that the replay never carried is dropped today -
that is the missing-key half - and a key outside Q that the replay did carry
is re-imaged from a truncated history, so a fix that only stopped the removals
would still publish a wrong page. A, the affected key set, is the tighter
statement of the same rule - a key the change did not touch has the state the old
root recorded, whatever the replay saw - and A is a subset of Q, so either
answers.
What it needs:
Q out of LiveViewCheckpointRowsBounds, as keys rather than as a count. The
discovery holds the key domain in its own Map keyed by the plan's partition-by
columns, and outputKeys exists only on the indexed-symbol path - "populated
only while the indexed seek is available". The walk path keeps the count alone.LiveViewSnapshotKeyCodec, off the window function's own map record), so the
two sides are comparable at publication time.publishRepair, and applied in
LiveViewCheckpointTimelineStoreWriter.buildRoot: removeMissingPartitions
removes only keys in Q, and a frozen key outside Q that the old root already
holds keeps the old entry rather than being put.Q, and isReplayStateKeyComplete() must
stay false there.Acceptance: LiveViewCheckpointRepairKeyCoverageTest's two ROWS cases stay
green with the splice back on rather than with it refused - the boundaries the
repair crosses survive naming all eight keys - plus a case for the carried-but-
truncated key the current gate makes unreachable, and the five cases task 4
reshaped go back to asserting the splice. All four met, and two cases were added
beyond them: the encoding the two sides have to agree on is asserted directly rather
than only through its consequences, and so is the refusal of a budget-truncated Q.
The io.questdb.test.cairo.lv package is green.