diff --git a/AGENTS.md b/AGENTS.md index 925cb99..baa3920 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,6 +37,8 @@ Pagerite is a CMS. See `docs` for the full design and implementation details. Ke - `assets/` — base CSS, Pygments styles, fonts. - `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test). - `scripts/translator.py` — Seed-X translator service client for the `/_translate/{key}` socket (reference client, runs in its own uv env via PEP 723); stays connected full time, unloads the model after 60 s idle and reloads on the next job. +- `scripts/llm_translator.py` — instruct-LLM translator service client (docs/llm-translation.md): speaks the `markdown`/`article` job modes against an OpenAI Chat Completions endpoint or ollama's native `/api/chat` (its `/v1` ignores `think: false`); all LLM specifics (prompts, sampling, generation caps) live here, not in pagerite. +- `scripts/import_translation.py` — import a human-made whole-article translation file into the fragment store (same `align_article` validation as article-mode results; run with the server stopped). Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed). diff --git a/docs/llm-translation.md b/docs/llm-translation.md index c5dce5c..e3e289c 100644 --- a/docs/llm-translation.md +++ b/docs/llm-translation.md @@ -1,5 +1,14 @@ # Whole-article and scoped LLM translation +**Status: implemented** (protocol modes and validation in +`pagerite/translate.py`, reference client `scripts/llm_translator.py`, +human-translation import `scripts/import_translation.py`; wire-level docs +in docs/localization.md). One deviation from the text below: ollama's +OpenAI-compatible `/v1/chat/completions` silently ignores `think: false` +(verified on 0.34.2), so the client's `api` config selects ollama's native +`/api/chat` for ollama backends; the OpenAI shape serves llama.cpp and +hosted APIs. + Design for augmenting the fragment-based machine translation (docs/localization.md) with general-purpose instruct LLMs that understand Markdown natively — as opposed to pure text-to-text models like Seed-X. diff --git a/docs/localization.md b/docs/localization.md index 9908e28..26ddf30 100644 --- a/docs/localization.md +++ b/docs/localization.md @@ -337,15 +337,18 @@ transaction `user`. Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`; `bytes` fields ride as base64): -- `{"type": "hello", "langs": [...]}` — client greeting announcing its - **capabilities**: the language codes its model can produce (normalized - to base subtags; `en`/empty dropped). -- `{"type": "job", "lang", "key", "texts", "path", "kind", "contexts"}` — - server push: ONE fragment to translate (an article title or a chunk), as - a list of **prose segments** (see Segmentation below). `contexts` is - parallel to `texts` ("" = none): the surround to translate the segment - in — for clients that translate better with context (see below). - Contexts are not part of the result. +- `{"type": "hello", "langs": [...], "model", "modes"}` — client greeting + announcing its **capabilities**: the language codes its model can produce + (normalized to base subtags; `en`/empty dropped). `model` is a free-form + model string (logging only); `modes` lists the job granularities the + client accepts (default `["segments"]`, see Job modes below). +- `{"type": "job", "lang", "key", "texts", "path", "kind", "mode", + "contexts"}` — server push: ONE fragment to translate (an article title + or a chunk). In the default `segments` mode `texts` is a list of **prose + segments** (see Segmentation below) and `contexts` is parallel to `texts` + ("" = none): the surround to translate the segment in — for clients that + translate better with context (see below). Contexts are not part of the + result. See Job modes for the other modes. - `{"type": "result", "lang", "key", "texts"}` — client reply: the segments translated, same order and count, matching its job by (lang, key). @@ -387,6 +390,55 @@ Results are stored into `trans` in one transaction and set pages gain a language from one fragment). Unknown keys are stored anyway and re-storing overwrites — results are idempotent. +#### Job modes: segments, markdown, article + +Instruct LLMs understand Markdown natively, so for them the segmentation +round trip below is unnecessary scaffolding (docs/llm-translation.md for +the design and the model trial evidence). `Hello.modes` announces which +job granularities a connection accepts; routing is per connection and per +mode, so a mixed fleet (a Seed-X instance, a local qwen, an API-backed +client) shares the work by capability. The validation skip-list is +mode-scoped — `(lang, key, mode)` — so a fragment one model rejects stays +offerable to clients of another approach. + +- **`segments`** (default when a client omits `modes`) — the protocol as + described so far: `Job.texts` carries prose segments, `Result.texts` + returns them, the server splices by offset. +- **`markdown`** — one fragment as full Markdown: `Job.texts` carries a + single element, the chunk (a title crosses as plain text, with the + article's opening as its context as today). `Job.contexts` carries the + previous and next block of the **served hybrid** in the target language + (machine translation with user patches applied, "" where none), so human + corrections propagate into fresh translations as terminology/tone + reference; contexts are never part of the result. The result must + re-chunk to exactly one block with the source's anchor constructs (link + and image destinations, `{...}` placeholders) intact (`clean_block`), + then stores to `Data.trans` as usual. +- **`article`** — a whole page at once, offered only to article-capable + connections and only while a page is *mostly* pending (a new article or + a full refresh; steady-state edit follow-up stays scoped jobs). The + job's key is the page's first chunk; `Job.texts` carries the full + original Markdown. The result is decomposed per chunk + (`align_article`): non-translatable blocks (code fences, container + fences, raw HTML — everything `needs_translation` rejects) must appear + verbatim and in order and anchor the alignment; regions between anchors + pair positionally, a region whose block count changed stores nothing + (its chunks stay pending and fall back to scoped jobs), and a paired + block whose destinations/placeholders did not survive likewise. + +`scripts/llm_translator.py` is the reference markdown+article client +(instruct LLMs via an OpenAI Chat Completions endpoint or ollama's native +API); `scripts/translator.py` (Seed-X) is untouched and announces +`["segments"]` implicitly. + +**Importing human-made full translations:** `align_article` doubles as an +import path — `scripts/import_translation.py PATH LANG FILE.md` (run with +the server stopped) decomposes a pasted whole-article translation (e.g. +from ChatGPT) into proper `Data.trans` fragments with the same validation, +so later source edits invalidate and re-translate per chunk rather than +letting the translation editor's one monolithic patch go stale hunk by +hunk. + #### Segmentation Fragments cross the wire as **prose segments** (`pagerite/segments.py`): the @@ -423,10 +475,12 @@ of the block it splices into, closing fence included — segments are inline prose, so `pure_prose` alone cannot see this) — each returned segment must parse as pure prose with no block-starting line or blank line, or the whole result -is dropped and logged, and the (lang, key) -pair is skipped for the rest of the server run (generation is -near-deterministic, so an immediate retry would re-fail; the fragment stays -pending and gets another chance on restart or `DELETE /_api/translations`). +is dropped and logged, and the (lang, key, mode) +combination is skipped for the rest of the server run (generation is +near-deterministic per model, so an immediate retry in the same mode would +re-fail; the fragment stays +pending and gets another chance on restart, in another mode, or on +`DELETE /_api/translations`). `Data.trans` therefore only ever holds clean translated Markdown. Link- and formatting-carrying blocks are the one place a segment is not