From 1e43f29eeca0ec49166933a5226d7436b9ac6d75 Mon Sep 17 00:00:00 2001 From: Leo Vasanko Date: Wed, 2 Sep 2026 16:14:54 +0000 Subject: [PATCH] docs: add database rotation design --- docs/rotation.md | 261 +++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 261 insertions(+) create mode 100644 docs/rotation.md diff --git a/docs/rotation.md b/docs/rotation.md new file mode 100644 index 0000000..d7756ab --- /dev/null +++ b/docs/rotation.md @@ -0,0 +1,261 @@ +# Database Rotation + +Goal: bound the on-disk history of a kanta database to a configurable retention +window (e.g. the last 30 days) by *rotating* the database file: the old content +is copied to a timestamped sibling file and the main file is truncated and +rewritten in place with only the retained history plus fresh snapshots. Normal +operation stays append-only under the exclusive lock; rotation is the only +operation that rewrites the file. + +## Current facts the design must respect + +- The writer holds an exclusive `flock` on the file from `open()` until + `close()` (`kanta/filelock.py`). No other process can safely touch the file + while a writer has it open. +- Records are append-only frames. Each `ChangeRecord` carries `ts` (record time) + and `m` (modification time); snapshots carry `ts`, `v` (schema version) and + `state` (`kanta/structs.py`). +- Replay reads the whole file, then starts from the **last snapshot** + (`framer.scan_last_snapshot`, `serialization/base.py:replay`). Anything before + the last snapshot is already logically dead. +- Snapshot state is validated in tooling: replayed state must equal snapshot + state (`kanta/replaylog.py`). A snapshot is therefore a consistency + checkpoint, not just an accelerator. +- `BinFramer` checksums are **offset-keyed** (checksum includes the absolute + `record_offset`). A binary frame copied to a different byte offset is + corrupted. `LineFramer` (JSONL) has no checksums. +- There is **no fsync/fdatasync** anywhere; durability currently relies on the + OS page cache. Rotation must not make this worse, and should fix it for the + rotation path at minimum. +- Migrations run on open, after replay, against the snapshot/replay version. + A snapshot records the version it was written at, so "db already migrated" + survives in the snapshot even if the migrations produced no change records. + +## Rotated file naming + +The history that aged out is preserved at: + +``` +{stem}@{ISO-8601 timestamp}.kantadb +``` + +- `{stem}` is the original filename with its extension stripped + (`Path(filename).stem`). +- The timestamp is the **ts of the last record dropped by the rotation** (see + step 4 — the leading snapshot of the rewritten main file carries the same + ts), not the current time. The name tells you exactly which point in history + the rotated file ends at. Use a filesystem-safe rendering (e.g. + `2026-09-02T15-24-57` — no `:` characters, which are awkward on some + filesystems). +- The rotated name always ends in `.kantadb`, regardless of the original + extension. Users may name their databases with no extension, `.kantadb`, or + anything else (`.db`, …). Since the rotated name is derived from the *stem*, + all of these work uniformly: `data` → `data@2026-09-02T15-24-57.kantadb`, + `data.kantadb` → `data@….kantadb`, `data.db` → `data@….kantadb`. +- Rotated files live in the same directory. +- Collision: if a rotated file with the same name already exists (rotation + rerun over identical history — should be prevented by the eligibility check + below, but be defensive), append a disambiguating suffix rather than + overwriting. + +## Why in-place rewrite (and not rename-and-recreate) + +An earlier draft renamed the locked file away and created a fresh file at the +main path. That opens a race: between the rename and the creation of the new +file, a second instance can open the (now missing) main path with `O_CREAT`, +acquire its own lock on the fresh inode, and bootstrap an empty database. The +rotating instance then cannot lock the path it needs, and two divergent +databases exist. `flock` is attached to the open file description (inode), not +the path — renaming never blocks a newcomer. + +Instead, rotation **never renames or unlinks the main file and never releases +its lock**: + +- Unix: `ftruncate(fd, 0)` on the open, locked fd is unaffected by the flock + and does not affect it. Subsequent writes use `lseek(fd, 0, SEEK_END)` + + `os.write` (`filelock.py:226`), which work identically after a truncate, so + append-mode operation continues unchanged. +- Windows: this is also the *more* portable option — the DB is opened with + `FILE_SHARE_READ` only (`filelock.py:240`), so renaming the locked file would + fail outright on Windows. In-place rewrite only needs `SetFilePointer(0)` + + `SetEndOfFile` on a handle we own. +- The main path therefore exists and remains locked throughout; a second + instance opening it at any moment gets either the old content or the new, + never a missing or half-created file, and never its own lock. + +The only new capability `LockedFile` needs is a `replace_content(data)` method +(seek 0, truncate, write, fsync) implemented per platform. + +## When to rotate: at open time, not at runtime + +Rotation happens **inside `Kanta.open()`, after acquiring the lock, before +replay**, gated by a retention option (see Configuration). Rationale: + +- The lock is already held and no background flush loop is running yet, so the + file is quiescent — no in-flight `pending_changes`, no concurrent snapshots. +- Runtime rotation would have to fence the background writer, drain the queue, + and prove no record lands in the file after the cutoff was computed. That + is a second synchronization protocol for a rare operation; not worth it. +- Open-time rotation also means rotation never races with `request_snapshot()` + or migration snapshot writes, which all happen under the same open() sequence. + +Consequence: a database that is never reopened never rotates. Document this; +for long-running services, rotation takes effect on the next restart. + +## Rotation algorithm (under the exclusive lock) + +Let `cutoff = now - retention`. Steps 1–3 operate on the bytes already read +into memory by `open_and_read`; no second disk read is needed. + +1. **Check eligibility.** Skip rotation when there is nothing to do: + - The file contains **no change records older than `cutoff`** — the + retention window already covers all history. + - The file contains **no change records at all** (snapshot-only file). + Opening a long-untouched database may legitimately rotate it down to a + single snapshot (that *is* the intended purge), but once a file has been + reduced to just a snapshot, rotating it again would be a pure no-op + rewrite. Treat "no change records" as "already fully rotated" and skip. + +2. **Find the replay base.** Replay normally starts at the most recent + snapshot, but that snapshot's `ts` is likely newer than `cutoff` — replaying + from it would silently drop history we intend to keep. Instead, scan + **backwards from the end of file**, collecting snapshots newest-first, and + pick the oldest snapshot `S` whose `ts <= cutoff` (i.e. walk back past + snapshots until one covers the required range, or until start of file). If + no such snapshot exists, `S` is "start of file" and the retained range is + replayed from the empty initial state. + - For `LineFramer` this is a reverse scan for `\nSNAPSHOT ` lines. + - For `BinFramer` frames are forward-scannable only; keep the forward scan + but record every snapshot position, then pick from the collected list. + +3. **Replay and validate.** Replay from `S` (or start of file) forward to end + of file, keeping every record with `ts >= cutoff`. At **every** snapshot + encountered after `S`, validate that the replayed state equals the snapshot + state; a mismatch means the history is corrupt or the chosen base is wrong — + abort rotation (leave the original file untouched) and surface the error. + The last snapshot in the file must always validate; if even that fails, + rotation must not proceed. + - Records with `ts < cutoff` are applied to the replay (they are needed to + reach the cutoff state) but not retained in the output. + - Remember `cutoff_end`: the byte offset in the original content just after + the last record with `ts < cutoff` (frame-boundary aligned). The rotated + file will be truncated to this length in step 6. + +4. **Copy the original aside.** `shutil.copy2(main_path, rotated_path)` — + no lock needed on the copy, and no temporary name: the content is written + directly to its final `{stem}@{ts}.kantadb` name. `copy2` preserves + metadata and, on filesystems with copy-on-write (btrfs, XFS with reflinks, + APFS, …), performs a cheap reflink copy instead of duplicating data; it is + also generally faster than re-writing the same bytes from memory. The + original bytes remain readable from the locked fd if the copy fails, so a + failure here simply aborts rotation. + +5. **Rewrite the main file in place.** On the locked fd: seek to 0, truncate + to 0, write the new content, `fdatasync`. The new content is, in order: + 1. A **snapshot of the state at the cutoff** — the replayed state after + applying all records with `ts < cutoff`, stamped with the **current + schema version**. Its `ts` is the **ts of the last pre-cutoff record** + (not the rotation time), and this is exactly the timestamp used in the + rotated filename. This snapshot is the new replay base and carries the + version forward so migrations are not re-run; it is always written. + 2. The retained change records (`ts >= cutoff`), **recreated record by + record** — no internal snapshots are carried over, even if the original + file had many in the retained range. + 3. A **final snapshot** of the state after the last retained record, + stamped with the current schema version — written **only if** there were + retained change records (and, in line with the existing snapshot policy + in `kanta/snapshot.py`, only when a meaningful number of changes + accumulated; a handful of trailing changes need not force one). If no + records survived the cutoff, the new file consists of the single leading + snapshot and nothing else — this is the steady state for databases whose + history has fully aged out, and the eligibility check in step 1 prevents + re-rotating such files. + +6. **Trim the rotated copy.** Truncate `{stem}@{ts}.kantadb` to `cutoff_end` + bytes, so it contains **only the dropped history** and does not duplicate + the records retained in the main file. The cut is at a frame boundary, so + the rotated file remains a valid, replayable database on its own (it is a + prefix of a valid log). This truncation happens only after step 5's fsync, + so until then the rotated file still holds the complete original content as + a crash-recovery anchor. + +7. **Continue normal open.** Replay/migrations proceed on the same locked fd. + Because the leading snapshot carries the current version, migrations run + exactly as they would have against the old content. + +Failure rule: any error before step 5 leaves the main file byte-identical +(only an extra copy exists). A crash during step 5 may leave the main file +torn, but the rotated copy still holds the complete original content +(truncated only after the main file is durable) — recovery is copying it back. +After step 6 the split is complete and both files are consistent. + +## Verbatim copy or rewrite? + +**Rewrite (re-frame), not verbatim copy**, for all records written to the main +file: + +- `BinFramer` checksums include `record_offset`, so a verbatim byte copy to a + new offset is unreadable. Binary records must be re-framed at their new + offsets regardless. +- Rewriting also normalizes encoding drift and lets us drop the redundant + intermediate snapshots the original file accumulated: none of them are + carried over — the new file contains only the leading cutoff snapshot, the + recreated change records, and (conditionally) the final snapshot. + +The rotated copy is the one place where verbatim bytes are used — a raw +`copy2` plus a frame-aligned tail truncation — which is safe precisely because +it preserves original offsets (the truncated prefix keeps every frame at its +original `record_offset`, so binary checksums stay valid). + +## Configuration + +Add keyword options to `Kanta(...)` (`kanta/kanta.py`), surfaced through +`open()`: + +- `retention: timedelta | None = None` — history window to keep. `None` + (default) disables rotation entirely; current behavior is unchanged. +- `rotate_keep: int = 3` (optional, later) — how many rotated backups to + retain; older ones are pruned at rotation time. + +Rotation uses `impl.now()` so the `@Kanta.clock` test clock controls it, same +as record timestamps. + +## Integrity checklist + +- Rotation runs under the exclusive lock, before the background writer starts. +- The main path is never renamed, unlinked, or unlocked during rotation; no + bootstrap race with a second instance is possible. +- Replay base is chosen by walking snapshots backwards until the retained range + is covered; replay is validated against every snapshot in range. +- A leading cutoff snapshot (ts = last pre-cutoff record, current schema + version) is always written; a final snapshot is written only when warranted + by retained changes. +- The full original content sits at `{stem}@{ts}.kantadb` before the main file + is touched, and is only trimmed to the dropped-history prefix after the + rewritten main file is `fdatasync`ed. +- Rotated files are never deleted by the rotation itself. +- Any validation failure aborts rotation with the original file intact. +- Files with no change records (already reduced to a snapshot) are never + re-rotated. + +## Testing notes + +- Use the test clock (`tests/test_clock.py`) to age records past the cutoff. +- Cover both framers: JSONL rotation and BinFramer rotation (assert the + rewritten binary file passes checksum validation and replays identically, + and that the truncated rotated prefix still passes checksum validation). +- Assert state equality before/after rotation, version continuity (no + re-migration), correct behavior when no snapshot precedes the cutoff, when + the newest snapshot is already older than the cutoff, and when retention + covers everything (no-op). +- Naming: databases named `x`, `x.kantadb`, and `x.db` all rotate to + `x@{ts}.kantadb`; the timestamp equals the last dropped record's ts and the + leading snapshot's ts. +- Assert the rotated file ends exactly at the last dropped record's frame + boundary (no overlap with the retained history in the main file). +- No-change files: a snapshot-only database opened with retention set is left + untouched (no copy, no rewrite). +- Aged-out database: all history older than the cutoff → new file contains + exactly one snapshot; opening it again performs no rotation. +- Concurrency: while one instance rotates, a second instance opening the main + path must fail with the normal "already locked" error at every stage.