3.8 KiB
3.8 KiB
Kanta Database Format and Design Principles
This document describes the on-disk format and design principles of Kanta. It is intentionally focused on the current standalone package behavior.
Core Principles
- Append-only durability
- State changes are persisted as appended JSON lines.
- Existing lines are never edited in place.
- Differential persistence
- Kanta stores diffs (patches), not full state, for normal writes.
- This keeps write volume small and preserves a clear change history.
- Deterministic replay
- Current state is reconstructed by replaying log records in order.
- Snapshot records accelerate replay while preserving deterministic results.
- Transactional in-memory writes
- Application code mutates in-memory data inside
kanta.transaction(...). - On success, Kanta computes and queues a diff record.
- On failure, in-memory data is rolled back.
- Explicit schema evolution
- Schema migration functions are versioned (
migrate_vN). - Migrations run at open time and advance the stored version.
On-Disk Record Types
Kanta uses a newline-delimited stream where each line is either a change record or a snapshot record.
Change record
One JSON object per line:
{"ts":"2026-06-10T02:55:00Z","a":"update","v":5,"u":"user-id","m":"2026-06-10T02:55:00Z","diff":{"users":{"alice":{"age":31}}}}
Fields:
ts: UTC timestamp of the record.a: action name.v: schema version after this change.u: optional actor identifier.m: optional domain modification timestamp.diff: jsondiff patch payload.
Snapshot record
Snapshot lines are prefixed with SNAPSHOT , followed by JSON:
SNAPSHOT {"ts":"2026-06-10T00:00:00Z","v":5,"state":{"users":{}},"m":"2026-06-10T00:00:00Z"}
Fields:
ts: snapshot creation time.v: schema version represented by the snapshot.state: full state dictionary.m: optional domain modification timestamp.
Replay Model
- Find the last snapshot in the file, if present.
- Initialize replay state from snapshot state (or
{}if none). - Replay subsequent change records in order using patch application.
- The final replay state becomes in-memory
kanta.data.
This model provides fast startup for large logs while retaining append-only history.
Serialization Semantics
- In-memory data is defined by an application
msgspec.Structtype. - Kanta round-trips through plain builtins for persistence and diffing.
- Dict keys are serialized as strings (
str_keys=True) for stable JSON form. - Normalization changes introduced by struct decode/encode are logged as
migrate:msgspecwhen they produce a diff.
Transaction Semantics
kanta.transaction(action=...)captures a pre-transaction snapshot dict.- On success:
- compute diff between previous builtins and current builtins,
- queue a
ChangeRecordif non-empty.
- On exception:
- restore in-memory data from snapshot,
- re-raise the exception.
Nested transactions are rejected.
Flush and Lifecycle
- Writes are queued in memory.
kanta.flush()appends queued records to disk.- A background async task can flush periodically.
kanta.close()performs final flush and releases file resources.async with Kanta(...)guarantees open/close lifecycle management.
Migrations
- Migration source is configured on
Kanta(...)viamigrations=. - Accepted values:
- imported module object,
- import path string.
- Migrations mutate replayed dict state in-place and return the new version.
Safety Invariants
- Any detected out-of-transaction mutation is treated as a fatal consistency violation.
- Flush failures mark the instance as failed and trigger shutdown behavior.
- Object identity of
kanta.datais preserved across rollback when possible, minimizing stale-reference hazards for callers.