Files
kanta/docs/database.md
T
2026-06-12 19:15:46 +00:00

3.8 KiB

Kanta Database Format and Design Principles

This document describes the on-disk format and design principles of Kanta. It is intentionally focused on the current standalone package behavior.

Core Principles

  1. Append-only durability
  • State changes are persisted as appended JSON lines.
  • Existing lines are never edited in place.
  1. Differential persistence
  • Kanta stores diffs (patches), not full state, for normal writes.
  • This keeps write volume small and preserves a clear change history.
  1. Deterministic replay
  • Current state is reconstructed by replaying log records in order.
  • Snapshot records accelerate replay while preserving deterministic results.
  1. Transactional in-memory writes
  • Application code mutates in-memory data inside kanta.transaction(...).
  • On success, Kanta computes and queues a diff record.
  • On failure, in-memory data is rolled back.
  1. Explicit schema evolution
  • Schema migration functions are versioned (migrate_vN).
  • Migrations run at open time and advance the stored version.

On-Disk Record Types

Kanta uses a newline-delimited stream where each line is either a change record or a snapshot record.

Change record

One JSON object per line:

{"ts":"2026-06-10T02:55:00Z","a":"update","v":5,"u":"user-id","m":"2026-06-10T02:55:00Z","diff":{"users":{"alice":{"age":31}}}}

Fields:

  • ts: UTC timestamp of the record.
  • a: action name.
  • v: schema version after this change.
  • u: optional actor identifier.
  • m: optional domain modification timestamp.
  • diff: jsondiff patch payload.

Snapshot record

Snapshot lines are prefixed with SNAPSHOT , followed by JSON:

SNAPSHOT {"ts":"2026-06-10T00:00:00Z","v":5,"state":{"users":{}},"m":"2026-06-10T00:00:00Z"}

Fields:

  • ts: snapshot creation time.
  • v: schema version represented by the snapshot.
  • state: full state dictionary.
  • m: optional domain modification timestamp.

Replay Model

  1. Find the last snapshot in the file, if present.
  2. Initialize replay state from snapshot state (or {} if none).
  3. Replay subsequent change records in order using patch application.
  4. The final replay state becomes in-memory kanta.data.

This model provides fast startup for large logs while retaining append-only history.

Serialization Semantics

  • In-memory data is defined by an application msgspec.Struct type.
  • Kanta round-trips through plain builtins for persistence and diffing.
  • Dict keys are serialized as strings (str_keys=True) for stable JSON form.
  • Normalization changes introduced by struct decode/encode are logged as migrate:msgspec when they produce a diff.

Transaction Semantics

  • kanta.transaction(action=...) captures a pre-transaction snapshot dict.
  • On success:
    • compute diff between previous builtins and current builtins,
    • queue a ChangeRecord if non-empty.
  • On exception:
    • restore in-memory data from snapshot,
    • re-raise the exception.

Nested transactions are rejected.

Flush and Lifecycle

  • Writes are queued in memory.
  • kanta.flush() appends queued records to disk.
  • A background async task can flush periodically.
  • kanta.close() performs final flush and releases file resources.
  • async with Kanta(...) guarantees open/close lifecycle management.

Migrations

  • Migration source is configured on Kanta(...) via migrations=.
  • Accepted values:
    • imported module object,
    • import path string.
  • Migrations mutate replayed dict state in-place and return the new version.

Safety Invariants

  • Any detected out-of-transaction mutation is treated as a fatal consistency violation.
  • Flush failures mark the instance as failed and trigger shutdown behavior.
  • Object identity of kanta.data is preserved across rollback when possible, minimizing stale-reference hazards for callers.