- Server replicates a ?lang= override onto the navigation links it renders (nav, sidebar, cards, brand), so clicks and prefetches stay in the chosen language even without JS; link_lang is part of the ETag and body cache key (query and header renders of the same language differ in their links). - pagerite.js drops the Accept-Language header hack: the remembered language rides internal fetches as ?lang= instead (added when a link lacks one), the page cache keys on path+query, and history/address bar keep the pretty query-less URL. - Canonical names the actually served language (plain URL for the original, ?lang= for translations); hreflang alternates are site-wide from translate_langs, identical on every page: x-default (the plain autodetecting URL) first, then every language explicitly, default included, emitted right after canonical before the social tags.
13 KiB
Localization
Pages are served in the visitor's language based on a ?lang= query
parameter or the Accept-Language header.
- Phase 1 (implemented): negotiation, URL scheme, caching, rendering plumbing. Translations are consumed through a stub interface; the database still holds only the original language.
- Phase 2 (implemented): gettext-style fragment storage in the
database — machine-translated chunks plus user override patches, assembled
at render time. Storage details in
docs/migrate.md.
Phase 1: negotiation and URLs
Language selection
Deliberately simple — q-values are ignored:
- All known
Accept-Languageimplementations send the header in order of preference, so we parse it as an ordered list and never reorder. - Selection rule (
select_languageinpagerite/i18n.py):- If
?lang=<tag>is present, use it (if a translation exists; otherwise fall through to header logic). - If the article's original language appears anywhere in the header list,
use the original. Rationale: an AI translation is strictly worse
than the original for anyone who has that language configured at all
(e.g.
fi-FI, fi, en-US, engets English, not machine-translated Finnish). - Otherwise walk the header list in order and use the first language for which a translation exists.
- Fall back to the original.
- If
Region tags normalize to their base subtag (fi-FI → fi).
URLs: pretty for users, indexable for search engines
- Canonical URLs stay pretty (
/some-page). Each language version is addressable as/some-page?lang=fiso search engines can index them. <link rel="canonical">names the actually served language: the plain URL when serving the original (for SEO the non-query URL means the article's own language),?lang=xxwhen serving a translation — however the language was arrived at (query or header).<link rel="alternate" hreflang="…">entries follow the canonical directly (before the social meta tags) and are the same set on every page — the site-wide configured languages (translate_langs, which the translator works to fill in):x-defaultfirst, pointing at the plain autodetecting URL, then every language explicitly with?lang=, the default language included.- The override sticks for the session of clicks: a page requested with
?lang=replicates the query onto the navigation links it renders (nav, sidebar, cards, brand — in-article links are content and stay as authored), so plain clicks and no-JS navigation keep the language. pagerite.js additionally strips the query from the address bar viahistory.replaceState(pretty, shareable URLs), remembers the language, and adds it to every internal fetch that lacks one (preloads, fetch-navigations, history traversals); history entries stay query-less. - A full page refresh or a shared link resets to automatic selection (header only). This gives a clean one-time override without cookies.
Response correctness
- Content responses carry
Vary: accept-language(added to the existingaccept-encodingvary). _cached_bodyand the page ETag include the selected language (not the raw header, which would blow up the cache key space) and the replicated link language: a?lang=firender and a header-selected Finnish render of the same page differ in their navigation links, so they are cached as separate variants.<html lang="…">reflects the served language.
Rendering
- The translated Markdown goes through the same
markdown.renderpipeline. - Navigation/sidebar titles come from the translation's title map, with per-node fallback to the original title (a partially translated tree must still render).
- Fixed UI strings ("Not Found" etc.) and the editor UI stay English for now.
- The markdown typographer (SmartyPants) is English-centric; per-language typographer options are a possible follow-up, not blocking.
Phase 2: fragment-based translation storage (implemented)
Phase 1 assumed whole-page translated Markdown delivered from outside. The
refined model is gettext-style: an article has one primary version (its
content, in its own language) plus, per target language, machine
fragments (translated chunks of Markdown) and user patches (minimal
editor overrides). Both are stored in the database and assembled into the
served Markdown at render time.
The scenario this must handle
- Article written in English.
- Machine-translated into Spanish → fragments stored.
- Editor fixes one Spanish paragraph and changes a link elsewhere to point at a Spanish resource → user patch hunks stored.
- English article edited → the edited chunk's key changes; its Spanish fragment no longer matches.
- Page requested before the machine translation refreshes → served as a hybrid: old fragments for unchanged chunks, plain English for the edited chunk. User patches are attempted against this hybrid, best effort, each hunk independently: the text fix is stale (its search text no longer exists) and silently skipped; the link change still applies even though the link sits in the now-English paragraph.
- Machine translation refreshes → full Spanish again, with both patch hunks applying.
Chunks
chunk_markdown(markdown) splits the source into block-level chunks —
blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
list blocks, tables, HTML blocks. A chunk's identity is its source text,
gettext-msgid style:
chunk_key = blake3(normalize(chunk_text)).digest(9) # bytes; base64 at the JSON level
(normalize: strip trailing whitespace per line, collapse surrounding blank
lines — so whitespace-only source edits don't invalidate translations.)
Consequences:
- Editing the English source invalidates exactly the edited chunks; all other fragments keep applying. Stale fragments are simply never referenced again and can be garbage-collected lazily (or left; they are tiny).
- No explicit "source version" bookkeeping is needed — staleness falls out of the keys.
User patches
Editors always edit full Markdown in the existing editor UX — never
fragments. When editing a translated view (?lang=es), the editor is loaded
with the current hybrid Markdown; on save, the server computes a minimal
diff against that hybrid and stores it as a patch:
class Patch(msgspec.Struct, omit_defaults=True):
"""One editing session's overrides, applied independently per hunk."""
hunks: list[tuple[str, str]] = [] # (search, replace) on hybrid Markdown
Hunks are produced from difflib.SequenceMatcher on the hybrid vs. the
edited text at block granularity: each replace/delete/insert opcode
becomes one (search, replace) pair, with the preceding block's tail as
left context for insert (pure inserts have empty search context otherwise).
Application is dead simple:
def apply_patch(hybrid: str, patch: Patch) -> str:
for search, replace in patch.hunks:
if search and search in hybrid:
hybrid = hybrid.replace(search, replace, 1)
# missing search text = stale hunk -> silently skipped
return hybrid
Per-hunk independence is the robustness property from the scenario: a stale text fix does not block a still-valid link change. Patches are stored as an ordered list and applied in order.
Storage
Full storage design and the migrate_v3 restructuring live in
docs/migrate.md. The short version, as it concerns this document:
- Originals and translations are content-addressed text chunks in flat
stores:
Data.chunks: dict[bytes, str]andData.trans: dict[bytes, dict[str, str]](chunk hash → lang → text) — path-independent, so repeated paragraphs and menu titles are translated once and article moves touch nothing.Node.chunks: list[bytes]gives each article its order. Nodegainslanguage: str = "", inherited down the tree likebanner(empty = nearest ancestor, front page last, site defaultenfinal).select_languageand<html lang>use the resolved value instead of the globalORIGINAL_LANGUAGEconstant.- Known weakness: changing a page's (or subtree's)
languageafter translations exist mis-keys everything — translations are keyed by source chunks, so old entries silently stop matching and user patches (searching for old-hybrid text) mostly go stale. That is acceptable: the orphaned data is harmless and translations regenerate. We do not migrate translations across a language change.
- Known weakness: changing a page's (or subtree's)
- Article paths are stored and keyed without leading slashes
(
"docs/setup", front page""); slashes are added only in hrefs.
Render pipeline (the phase-1 get_translation stub, now real)
def get_translation(path, lang, data) -> Translation | None:
if lang not in node.langs:
return None
hybrid = "\n\n".join(
chunks[h] if h in node.no_trans else trans.get(h, {}).get(lang, chunks[h])
for h in node.chunks
)
for patch in data.patches.get(f"{path}:{lang}", []):
hybrid = apply_patch(hybrid, patch)
return Translation(markdown=hybrid, titles=title_map(data, lang))
- Availability is an article-level index:
node.langs: dict[lang, True], maintained by the translation writers (translator job, patch saves) in the same transaction as their data writes — rendering and language selection never probe thetransstore chunk by chunk. A stale key is benign (the "translation" just renders as the original). titlesfor nav/sidebar/cards: each node's translated title istrans.get(hash(node.title), {}).get(lang)with per-node fallback — one dict lookup per nav item at render time.- Cache invalidation: writes to
chunks/trans/patches(translator, editor saves) call_invalidate_pages(), same as content writes.
Editor flow
GETof page Markdown for editing with alangparameter returns the hybrid (not the raw original) when the article has that language.PUT/WS save withlangdoes not touchnode.chunks; it diffs against the hybrid that was served and appends aPatch. (Serve a hybrid generation token with the editor payload so a save based on a stale hybrid can be rebased or rejected — simplest: recompute the diff against the current hybrid and accept best-effort, matching the patch philosophy.)- Saving the primary-language version re-chunks the submitted Markdown and
updates
Data.chunks/node.chunks— only genuinely new text lands in the kanta change diff (see docs/migrate.md).
Translator service API
An external machine-translation service connects over WebSocket at
/_translate/{key} — deliberately not under /_api: the SSO
forward-auth does not cover that route, and the key in the path is the
access control. The key is Data.translate_key, generated once at startup
and surfaced to the admin in GET /_api/settings as translate_key. A
wrong or empty key rejects the handshake (close-before-accept → HTTP 403).
Frames are JSON-encoded tagged msgspec structs (pagerite/translate.py;
bytes fields ride as base64):
{"type": "hello", "langs": [...]}— client greeting announcing its capabilities: the language codes its model can produce (normalized to base subtags;en/empty dropped).{"type": "job", "lang", "key", "text", "path", "kind"}— server push: ONE fragment to translate (an article title or a chunk).{"type": "result", "lang", "key", "text"}— client reply: the translation of the connection's current job, matching it by (lang, key).
Which languages get translated is server-configured:
Data.translate_langs (presence-key dict, read/set via /_api/settings
as translate_langs; no editing UI yet). The dispatcher offers a
connection jobs only in wanted ∩ capable; a connection without overlap
simply stays idle.
Dispatch semantics (all in app.py):
- One job at a time per connection — the next job is sent only after
the current one's result. Clients wanting parallelism open multiple
connections (e.g. several
scripts/translator.pyinstances). - Pending work is derived from the
transstore (translate.pending_items) minus the items in flight on any connection, so a disconnect requeues that connection's in-flight item and it is offered to any free capable connection. - Dispatch re-runs on every relevant event: Hello, result, disconnect and
content change (
_invalidate_pages()schedules it, so the pass runs after the writing transaction commits). - A result with no job in flight, a mismatched (lang, key), a duplicate hello, or any malformed frame closes the socket with a protocol error.
Results are stored into trans in one transaction and set
node.langs[lang] on every article they touch (shared chunks make several
pages gain a language from one fragment). Unknown keys are stored anyway
and re-storing overwrites — results are idempotent.
Explicitly out of scope for phase 2
- The machine translation itself: the API above moves fragments in and out;
the translating is external.
scripts/translator.pyis the reference client (Seed-X-PPO-7B only — its 28 languages are the ceiling). - Garbage collection of orphaned chunks/translations (see docs/migrate.md).
- sitemap.xml per-language entries; translated UI chrome; per-language typographer options; multi-locale date/number formatting.