Mark phase-2 localization storage implemented; list chunks.py in AGENTS.md
This commit is contained in:
@@ -6,7 +6,7 @@ parameter or the `Accept-Language` header.
|
||||
- **Phase 1 (implemented):** negotiation, URL scheme, caching, rendering
|
||||
plumbing. Translations are consumed through a stub interface; the database
|
||||
still holds only the original language.
|
||||
- **Phase 2 (final plan):** gettext-style fragment storage in the
|
||||
- **Phase 2 (implemented):** gettext-style fragment storage in the
|
||||
database — machine-translated chunks plus user override patches, assembled
|
||||
at render time. Storage details in `docs/migrate.md`.
|
||||
|
||||
@@ -66,7 +66,7 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
|
||||
- The markdown typographer (SmartyPants) is English-centric; per-language
|
||||
typographer options are a possible follow-up, not blocking.
|
||||
|
||||
## Phase 2: fragment-based translation storage (draft)
|
||||
## Phase 2: fragment-based translation storage (implemented)
|
||||
|
||||
Phase 1 assumed whole-page translated Markdown delivered from outside. The
|
||||
refined model is gettext-style: an article has **one primary version** (its
|
||||
@@ -170,7 +170,7 @@ Full storage design and the `migrate_v3` restructuring live in
|
||||
- Article paths are stored and keyed **without leading slashes**
|
||||
(`"docs/setup"`, front page `""`); slashes are added only in hrefs.
|
||||
|
||||
### Render pipeline (replaces the phase-1 `get_translation` stub)
|
||||
### Render pipeline (the phase-1 `get_translation` stub, now real)
|
||||
|
||||
```python
|
||||
def get_translation(path, lang, data) -> Translation | None:
|
||||
|
||||
+16
-1
@@ -1,6 +1,6 @@
|
||||
# migrate_v3: content-addressed chunk storage
|
||||
|
||||
Status: **final plan**. `migrate_v3` restructures how article text and
|
||||
Status: **implemented**. `migrate_v3` restructures how article text and
|
||||
translations are stored, motivated by the localization model in
|
||||
`docs/localization.md` (phase 2). Since it is a full migration, it is free to
|
||||
break the current `Node.content: str | None` layout.
|
||||
@@ -132,6 +132,21 @@ Chunking must be deterministic and shared with render/save, so
|
||||
`pagerite/chunks.py`) and are imported by both `migrations.py` and
|
||||
`views.py`/`app.py`.
|
||||
|
||||
## Implementation notes (deviations from the plan above)
|
||||
|
||||
- Chunking lives in `pagerite/chunks.py`; hashing uses the `blake3` package
|
||||
(already a dependency) with a 16-byte digest (`hexdigest(16)`).
|
||||
- `Translation.titles` stayed keyed by node path (phase-1 shape, views
|
||||
untouched): `get_translation` builds it by walking the menu with the same
|
||||
per-title `trans[f"{chunk_key(node.title)}:{lang}"]` lookups.
|
||||
- Insert hunks anchor on the whole preceding block (not just its tail) —
|
||||
a stronger, simpler search context.
|
||||
- `make_patch` diffs with `SequenceMatcher(autojunk=False)` so patches are
|
||||
deterministic (popular lines like blank separators never become junk).
|
||||
- Step 3's path normalization is a no-op in practice: the only path-keyed
|
||||
store (`patches`) starts empty at v3; analytics paths live outside the
|
||||
kantadb. The code still strips leading slashes defensively.
|
||||
|
||||
## Garbage collection (later, manual or idle-time)
|
||||
|
||||
Orphaned entries accumulate: chunks no longer referenced by any
|
||||
|
||||
Reference in New Issue
Block a user