Segments, not sentinels: prose-only translation wire protocol
- pagerite/segments.py replaces masking.py: a fragment is parsed with the project's own markdown-it (markdown.make_md(verbatim=True), byte-identical tokens) and split into pure-prose segments with source spans; only the segments plus per-segment context surrounds cross the wire (Job.texts / contexts / Result.texts) and translations splice back by offset — markup can no longer break, it never leaves the server. Count/empty/non-prose results are rejected and skipped for the run. - Localization machinery out of app.py: the translator dispatcher (clients, job pipeline, validation skip-list) moves into translate.Dispatcher; translated-edit recording moves into i18n (add_patch, set_title_translation, clear_translations). app.py keeps only the routes. - Structure editor localized: flag strip switches the language titles are shown/edited in (GET /_api/pages?lang= flags translated rows, originals dimmed); retitling in a translation writes a per-language title fragment via StructureOp.lang — slugs, order and hierarchy stay language-independent. - Localization tab: refresh-all button (DELETE /_api/translations) drops machine translations, keeps user patches and clears the skip-list so the dispatcher re-translates everything. - PageEditor always opens in the primary language; editor socket gets reconnect/doc-mismatch logging. Reference translator: per-segment calls with context prompts, deterministic punctuation matching and the "<" markup-bleed cut.
This commit is contained in:
@@ -13,8 +13,9 @@ Pagerite is a CMS. See `docs` for the full design and implementation details. Ke
|
||||
- `app.py` — FastAPI app and route registration.
|
||||
- `data.py` — msgspec Structs for the kanta database.
|
||||
- `chunks.py` — block-level Markdown chunking and content-hash keys for the chunk stores (docs/migrate.md).
|
||||
- `i18n.py` — language selection, translation assembly (chunks + patches).
|
||||
- `translate.py` — translator service protocol (msgspec structs) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md).
|
||||
- `i18n.py` — language selection, translation assembly (chunks + patches) and translated-edit recording (user patches, per-language title overrides, refresh).
|
||||
- `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); app.py only registers the route.
|
||||
- `segments.py` — the translation round trip: fragments split into pure-prose wire segments (via markdown.make_md's verbatim parser) and translations spliced back by source offset (docs/localization.md).
|
||||
- `migrations.py` — kanta migrations (`migrate_vN`); ALL schema/storage upgrades live here (raw state dict before struct decoding), never in the app lifespan: v1 moves legacy in-db file blobs to the on-disk store and rebuilds the legacy flat `pages` as the menu tree, v2 rewrites `/_f/{hash}.ext` image links to the extension-less form, backfills AVIF/WebP/JPEG derivatives on disk and drops the obsolete `version` field.
|
||||
- `markdown.py` — markdown-it-py renderer.
|
||||
- `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`).
|
||||
|
||||
+103
-28
@@ -194,7 +194,7 @@ Full storage design and the `migrate_v3` restructuring live in
|
||||
### Render pipeline (the phase-1 `get_translation` stub, now real)
|
||||
|
||||
```python
|
||||
def get_translation(path, lang, data) -> Translation | None:
|
||||
def get_translation(data, path, lang) -> Translation | None:
|
||||
if lang not in node.langs:
|
||||
return None
|
||||
hybrid = "\n\n".join(
|
||||
@@ -222,8 +222,9 @@ def get_translation(path, lang, data) -> Translation | None:
|
||||
The page editor has a language picker (flag + name; the same
|
||||
country-flag-icons set as the analytics visitor cells) listing the primary
|
||||
language and the union of the page's translations (`node.langs`) and the
|
||||
site-wide `translate_langs`. It opens in the language the page was served
|
||||
in (`<html lang>`). A note under the toolbar states the blast radius:
|
||||
site-wide `translate_langs`. It always opens in the primary language, even
|
||||
when the page itself was served in a translation. A note under the toolbar
|
||||
states the blast radius:
|
||||
edits to the primary language re-chunk the original (invalidating the
|
||||
affected translation fragments everywhere); edits to a translation stay
|
||||
local to that language.
|
||||
@@ -253,6 +254,17 @@ local to that language.
|
||||
updates `Data.chunks` / `node.chunks` — only genuinely new text lands in
|
||||
the kanta change diff (see docs/migrate.md).
|
||||
|
||||
The **structure editor** has the same flag strip for titles. The tree it
|
||||
lists (`GET /_api/pages?lang=`) comes back with per-language titles where a
|
||||
translation exists (`translated` marks those rows; untranslated rows show
|
||||
the original title, dimmed). Retitling in a non-primary language posts the
|
||||
structure op with a `lang` and writes a per-language title fragment in
|
||||
`Data.trans` (keyed by the original title's chunk hash, exactly like a
|
||||
machine title translation — a user edit simply overwrites it); sending the
|
||||
original's text drops the override. The structure itself — slugs,
|
||||
hierarchy, order — is language-independent, so pending rows, slug edits,
|
||||
drag-and-drop and deletes work identically in every language.
|
||||
|
||||
### Translator service API
|
||||
|
||||
An external machine-translation service connects over WebSocket at
|
||||
@@ -275,11 +287,14 @@ Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
|
||||
- `{"type": "hello", "langs": [...]}` — client greeting announcing its
|
||||
**capabilities**: the language codes its model can produce (normalized
|
||||
to base subtags; `en`/empty dropped).
|
||||
- `{"type": "job", "lang", "key", "text", "path", "kind"}` — server push:
|
||||
ONE fragment to translate (an article title or a chunk), its text
|
||||
**masked** (see Masking below).
|
||||
- `{"type": "result", "lang", "key", "text"}` — client reply: the
|
||||
translation of the connection's current job, matching it by (lang, key).
|
||||
- `{"type": "job", "lang", "key", "texts", "path", "kind", "contexts"}` —
|
||||
server push: ONE fragment to translate (an article title or a chunk), as
|
||||
a list of **prose segments** (see Segmentation below). `contexts` is
|
||||
parallel to `texts` ("" = none): the surround to translate the segment
|
||||
in — for clients that translate better with context (see below).
|
||||
Contexts are not part of the result.
|
||||
- `{"type": "result", "lang", "key", "texts"}` — client reply: the
|
||||
segments translated, same order and count, matching its job by (lang, key).
|
||||
|
||||
Which languages get translated is **server-configured**:
|
||||
`Data.translate_langs` (presence-key dict, bootstrapped to Spanish and
|
||||
@@ -289,7 +304,15 @@ The dispatcher offers a
|
||||
connection jobs only in `wanted ∩ capable`; a connection without overlap
|
||||
simply stays idle.
|
||||
|
||||
Dispatch semantics (all in app.py):
|
||||
`DELETE /_api/translations` (the localization tab's "refresh all
|
||||
translations" button) drops every machine translation (`Data.trans`) and
|
||||
rebuilds the availability index (`node.langs`) from the surviving user
|
||||
patches, so the dispatcher re-translates everything from scratch; the
|
||||
run's validation skip-list is cleared with it, giving rejected fragments
|
||||
another chance.
|
||||
|
||||
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; app.py only
|
||||
registers the route):
|
||||
|
||||
- **One job at a time per connection** — the next job is sent only after
|
||||
the current one's result. Clients wanting parallelism open multiple
|
||||
@@ -309,27 +332,79 @@ Results are stored into `trans` in one transaction and set
|
||||
pages gain a language from one fragment). Unknown keys are stored anyway
|
||||
and re-storing overwrites — results are idempotent.
|
||||
|
||||
#### Masking
|
||||
#### Segmentation
|
||||
|
||||
Fragments cross the wire **masked** (`pagerite/masking.py`): spans the model
|
||||
must copy byte-identically are replaced with numbered `⟦N⟧` sentinels before
|
||||
dispatch and restored by number from the result. Masked: code spans,
|
||||
container-fence names, link and image *destinations* (link text, alt text
|
||||
and captions stay visible for translation), reference and footnote labels,
|
||||
`{...}` spans (placeholders like `{dates}` as well as attrs), inline HTML
|
||||
tags and bare URLs. Markdown punctuation (`*`, `|`, `[]()`, `:::`) is not
|
||||
masked — it carries no lexical content and models preserve it. Chunks with
|
||||
no prose left after masking (a lone `{dates}`, container fences, pure
|
||||
code/HTML) are never dispatched at all (`needs_translation`); every language
|
||||
renders them from the original chunk.
|
||||
Fragments cross the wire as **prose segments** (`pagerite/segments.py`): the
|
||||
fragment is parsed with the project's own markdown-it setup
|
||||
(`markdown.make_md(verbatim=True)` — all extensions, but no typographer or
|
||||
tasklist label wrapping, so token text stays byte-identical to the source)
|
||||
and split into the runs a model may touch: paragraph/heading/table-cell text
|
||||
(merged across soft line breaks), link text, image alt texts and captions,
|
||||
footnote bodies. Everything else never leaves the server: code spans and
|
||||
fences, URLs and autolinks, link/image *destinations*, `{...}` spans
|
||||
(placeholders like `{dates}` as well as attrs), reference and footnote
|
||||
labels, container fences, GFM alert markers (`[!NOTE]`), raw HTML — and all
|
||||
markup punctuation (`*`, `|`, `[]()`, `:::`), which is a run boundary.
|
||||
Chunks with no segments (a lone `{dates}`, container fences, pure
|
||||
code/HTML) are never dispatched at all (`needs_translation`); every
|
||||
language renders them from the original chunk. Each segment is accompanied
|
||||
by a context string (a segment carved out of a larger block carries the
|
||||
block's plain text; a whole-block segment carries "") — context is a
|
||||
prompt aid only, never spliced into the result.
|
||||
|
||||
A result is accepted only if every sentinel survived exactly once, in any
|
||||
order (translations legitimately reorder spans). A mangled result is dropped
|
||||
and logged, and the (lang, key) pair is skipped for the rest of the server
|
||||
run — generation is near-deterministic, so an immediate retry would re-fail
|
||||
the same way; the fragment stays pending and gets another chance on restart
|
||||
or a model/masking change. `Data.trans` therefore only ever holds clean,
|
||||
unmasked text.
|
||||
Reassembly is offset splicing, not text the model produced: each segment's
|
||||
source span was located at dispatch (sequential search; a run that is not a
|
||||
verbatim source substring — entity-decoded text, backslash escapes — is
|
||||
skipped and stays in the original language), and the returned translations
|
||||
are swapped in by offset. Markup corruption is therefore impossible by
|
||||
construction; the failure modes that remain are a wrong segment count, an
|
||||
empty segment, or markup injected INTO a segment (a `<br>` in a title
|
||||
translation would splice live HTML) — each returned segment must parse as
|
||||
pure prose, or the whole result is dropped and logged, and the (lang, key)
|
||||
pair is skipped for the rest of the server run (generation is
|
||||
near-deterministic, so an immediate retry would re-fail; the fragment stays
|
||||
pending and gets another chance on restart or `DELETE /_api/translations`).
|
||||
`Data.trans` therefore only ever holds clean translated Markdown.
|
||||
|
||||
The trade-off: segments splice back at fixed positions, so a translation
|
||||
cannot move a link or image within a sentence — word order around inline
|
||||
markup follows the original. That is the price for never feeding the model
|
||||
markup (an earlier sentinel-masking design let the model see and mangle
|
||||
exactly that punctuation: Seed-X turned `![` into `¡¡…!!`).
|
||||
|
||||
Punctuation is the translator's own job: Seed-X tends to "finish" short
|
||||
labels (titles, nav items) with a comma or period the source never had.
|
||||
Prompt wording is NOT the fix — a punctuation-instruction clause made
|
||||
Seed-X slip into its `[COT]` reasoning mode (minutes-long generations with
|
||||
reasoning text in the output, observed for Chinese). The reference client
|
||||
enforces punctuation deterministically instead (`match_punctuation` in
|
||||
scripts/translator.py): a translation of a segment without terminal
|
||||
punctuation gets any added trailing marks (and a newly opened Spanish ¡/¿)
|
||||
stripped before the result goes back.
|
||||
|
||||
The same client-side enforcement covers markup bleed as a CLASS, not per
|
||||
artifact: `<` is the prose/markup boundary on the wire and never appears in
|
||||
a segment in either direction. Source pieces containing `<` are never
|
||||
dispatched (they stay in the original language — segments.py), and the
|
||||
reference client cuts the model's output at the first `<`
|
||||
(scripts/translator.py) — echoed language tags, stray `<br>`s and any
|
||||
future variant are one handled case. (The cut is post-decode, not a
|
||||
generation stop string: Seed-X opens every generation with its `<s>`
|
||||
framing token, which would trip a `<` stop immediately.)
|
||||
|
||||
Short fragments get more than a bare prompt: each segment may carry its
|
||||
surround in `Job.contexts` — a title carries the article's opening prose
|
||||
(its own block is just the title word), a segment carved out of a larger
|
||||
block (a link text, a partial run) carries the block's plain text, and a
|
||||
whole-block segment (a plain paragraph) is self-contextualizing and carries
|
||||
"". The reference client translates segment and surround together, stops
|
||||
generation at the blank line separating them, and keeps the segment's own
|
||||
part of the output (its line resp. paragraph; a hard-break `␣␣\n` separator
|
||||
works too). If the model merged them (no separator, or an empty first
|
||||
part), it falls back to translating the segment alone. The surround fixes
|
||||
context-free readings ("About" as "approximately" — with the opening it
|
||||
becomes "Tietoa"/"Acerca de"; "here" as "就在这里" → the idiomatic
|
||||
"点击这里") and, as a side effect, most stray trailing punctuation.
|
||||
|
||||
### Explicitly out of scope for phase 2
|
||||
|
||||
|
||||
@@ -4,7 +4,9 @@
|
||||
// language is configured per site hierarchy, not here. Flag clicks toggle
|
||||
// and save immediately; the settings round-trip re-reads the payload, so
|
||||
// this tab only ever changes translate_langs. The settings write's
|
||||
// invalidation hook kicks the translation dispatcher.
|
||||
// invalidation hook kicks the translation dispatcher. The refresh button
|
||||
// drops all machine translations (user patches are kept), making the
|
||||
// dispatcher re-translate everything.
|
||||
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
|
||||
import { TRANSLATABLE, flagFor, langName } from './langs'
|
||||
import { dropPageCache } from './swapdoc'
|
||||
@@ -75,6 +77,24 @@ async function toggle(code) {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
}
|
||||
|
||||
// Delete all machine translations server-side; the dispatcher re-fills
|
||||
// them (a connected translator starts getting jobs right away). User
|
||||
// patches survive — they are edits, not machine output.
|
||||
const refreshing = ref(false)
|
||||
async function refresh() {
|
||||
if (refreshing.value) return
|
||||
refreshing.value = true
|
||||
try {
|
||||
const res = await fetch('/_api/translations', { method: 'DELETE' })
|
||||
saveError.value = res.ok ? '' : '⚠️ translations could not be refreshed'
|
||||
if (res.ok) dropPageCache()
|
||||
} catch {
|
||||
saveError.value = '⚠️ translations could not be refreshed'
|
||||
} finally {
|
||||
refreshing.value = false
|
||||
}
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
@@ -100,6 +120,22 @@ async function toggle(code) {
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">translations</span>
|
||||
<small class="muted">deleting re-translates everything; user edits are kept</small>
|
||||
</div>
|
||||
<button
|
||||
type="button"
|
||||
class="refresh-btn"
|
||||
:disabled="refreshing"
|
||||
title="delete all machine translations and let the translator re-fill them"
|
||||
@click="refresh"
|
||||
>
|
||||
{{ refreshing ? 'refreshing…' : 'refresh all translations' }}
|
||||
</button>
|
||||
</section>
|
||||
|
||||
<section v-if="keyUrls.length" class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">translator service</span>
|
||||
@@ -207,4 +243,27 @@ async function toggle(code) {
|
||||
.key-row code {
|
||||
user-select: all;
|
||||
}
|
||||
|
||||
.refresh-btn {
|
||||
align-self: flex-start;
|
||||
margin-bottom: 0.2rem;
|
||||
padding: 0.3rem 0.8rem;
|
||||
font: inherit;
|
||||
font-size: 0.85rem;
|
||||
color: var(--muted);
|
||||
background: none;
|
||||
border: 1px solid var(--line);
|
||||
border-radius: 5px;
|
||||
cursor: pointer;
|
||||
}
|
||||
|
||||
.refresh-btn:hover:not(:disabled) {
|
||||
color: var(--text);
|
||||
border-color: var(--muted);
|
||||
}
|
||||
|
||||
.refresh-btn:disabled {
|
||||
opacity: 0.5;
|
||||
cursor: default;
|
||||
}
|
||||
</style>
|
||||
|
||||
+27
-10
@@ -14,8 +14,8 @@
|
||||
// stashing unsaved text per path and language (stashes) so returning to the
|
||||
// page restores the working draft; stashes clear on save and on real reload.
|
||||
//
|
||||
// Languages: the editor starts in the language the page was served in and
|
||||
// the toolbar picker (flags, like the analytics visitor cells) switches
|
||||
// Languages: the editor always starts in the primary language and the
|
||||
// toolbar picker (flags, like the analytics visitor cells) switches
|
||||
// between the primary language and its translations. A translation is
|
||||
// edited as its effective (hybrid) Markdown; the hybrid the session
|
||||
// started from is kept as a shadow copy (shadowBase) and sent along at
|
||||
@@ -47,10 +47,10 @@ const saveError = ref('')
|
||||
const editorEl = ref(null)
|
||||
const fileInput = ref(null)
|
||||
|
||||
// The language being edited: "" = the primary language. Starts as the
|
||||
// language this page was served in (<html lang>); the server normalizes
|
||||
// the primary to "" in its doc reply.
|
||||
const lang = ref(document.documentElement.lang || '')
|
||||
// The language being edited: "" = the primary language, where the editor
|
||||
// always starts (a served translation does not follow it into the editor;
|
||||
// the picker switches). The server normalizes the primary to "" anyway.
|
||||
const lang = ref('')
|
||||
const primaryLang = ref('en')
|
||||
const pageLangs = ref([]) // translations this page has
|
||||
const siteLangs = ref([]) // site-wide configured target languages
|
||||
@@ -733,6 +733,14 @@ function onMessage(ev) {
|
||||
requestRender()
|
||||
// A section pen's target line survives the open/path-switch here.
|
||||
consumePendingLine()
|
||||
} else if (msg.type === 'doc') {
|
||||
// A doc that answered neither path nor language of the current session
|
||||
// (a late reply to a pre-switch open) — visible because it leaves the
|
||||
// editor empty when it's the only doc that ever arrives.
|
||||
console.warn(
|
||||
'[pagerite] doc dropped:', msg.path, msg.lang || '(primary)',
|
||||
'— editor is on', path.value, lang.value || '(primary)',
|
||||
)
|
||||
} else if (msg.type === 'html' && msg.path === path.value) {
|
||||
previewIntoArticle(msg.html, msg.multicol)
|
||||
} else if (msg.type === 'saved') {
|
||||
@@ -941,11 +949,16 @@ function connect() {
|
||||
`${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`,
|
||||
)
|
||||
ws.onmessage = onMessage
|
||||
ws.onerror = (ev) => {
|
||||
console.error('[pagerite] editor socket error', ev)
|
||||
}
|
||||
ws.onopen = () => {
|
||||
reconnectDelay = 2000
|
||||
if (everConnected) {
|
||||
// Reconnected: local text is authoritative — don't re-open (that
|
||||
// would clobber the editor), just resync preview and pending saves.
|
||||
if (everConnected && docLoaded) {
|
||||
// Reconnected with a document loaded: local text is authoritative —
|
||||
// don't re-open (that would clobber the editor), just resync preview
|
||||
// and pending saves. Without a doc (the disconnect came first) fall
|
||||
// through to a normal open, or the editor would stay empty forever.
|
||||
requestRender()
|
||||
if (pendingSave) send(pendingSave)
|
||||
} else {
|
||||
@@ -953,7 +966,11 @@ function connect() {
|
||||
}
|
||||
everConnected = true
|
||||
}
|
||||
ws.onclose = () => {
|
||||
ws.onclose = (ev) => {
|
||||
// 1006 = abnormal (e.g. the dev proxy refused/dropped the upgrade);
|
||||
// worth seeing since a dead socket before the first doc bricks the
|
||||
// editor until this retry loop lands one.
|
||||
console.warn('[pagerite] editor socket closed:', ev.code, ev.reason || '')
|
||||
clearTimeout(reconnectTimer)
|
||||
reconnectTimer = setTimeout(() => {
|
||||
connect()
|
||||
|
||||
@@ -7,9 +7,16 @@
|
||||
// real — a label with a title and slug, with content (landing page) or
|
||||
// without (category whose URL renders a placeholder page). The front page
|
||||
// is a top-level row with an empty slug, not the parent of the others.
|
||||
import { inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
|
||||
//
|
||||
// Languages: the flag strip switches which language the TITLES are shown
|
||||
// and edited in (rows without a translation show the original, dimmed).
|
||||
// Translated title edits write a per-language fragment (POST /_api/structure
|
||||
// with lang); the structure itself — slugs, order, hierarchy — is
|
||||
// language-independent and always edits the same tree.
|
||||
import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
|
||||
import StructureTree from './StructureTree.vue'
|
||||
import { slugify } from './slugify'
|
||||
import { flagFor, langName } from './langs'
|
||||
import { dropPageCache, loadPlain } from './swapdoc'
|
||||
|
||||
const props = defineProps({
|
||||
@@ -23,6 +30,33 @@ const path = ref('')
|
||||
const saveError = ref('')
|
||||
const tree = ref([])
|
||||
|
||||
// The language the tree's titles are shown and edited in: "" = primary.
|
||||
const lang = ref('')
|
||||
const primaryLang = ref('en')
|
||||
const siteLangs = ref([])
|
||||
|
||||
// The strip's options: the primary language first, then the configured
|
||||
// translation targets (the lang tab manages that set).
|
||||
const langOptions = computed(() =>
|
||||
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
|
||||
.map((code) => ({
|
||||
tag: code === primaryLang.value ? '' : code,
|
||||
code,
|
||||
name: langName(code),
|
||||
flag: flagFor(code),
|
||||
primary: code === primaryLang.value,
|
||||
})),
|
||||
)
|
||||
const currentLang = computed(
|
||||
() => langOptions.value.find((o) => o.tag === lang.value) ?? langOptions.value[0],
|
||||
)
|
||||
|
||||
function switchLang(tag) {
|
||||
if (tag === lang.value) return
|
||||
lang.value = tag
|
||||
refreshPages()
|
||||
}
|
||||
|
||||
function normPath(p) {
|
||||
return p.trim().replace(/^\/+|\/+$/g, '')
|
||||
}
|
||||
@@ -154,7 +188,8 @@ async function commitPending() {
|
||||
// --- Site structure tree (drag-and-drop ordering/moving) ----------------
|
||||
async function refreshPages() {
|
||||
try {
|
||||
tree.value = await (await fetch('/_api/pages')).json()
|
||||
const q = lang.value ? `?lang=${lang.value}` : ''
|
||||
tree.value = await (await fetch(`/_api/pages${q}`)).json()
|
||||
} catch { /* list stays stale; not fatal */ }
|
||||
}
|
||||
|
||||
@@ -207,13 +242,15 @@ async function onReorder(parentPath, list, evt) {
|
||||
}
|
||||
|
||||
// Inline title/slug editing: rows are always editable. Title saves while
|
||||
// typing (debounced); the slug commits on blur/Enter, since it renames
|
||||
// the path (moving the whole subtree with it).
|
||||
// typing (debounced) — in the selected language (a translation writes a
|
||||
// title fragment, the primary language the original); the slug commits on
|
||||
// blur/Enter, since it renames the path (moving the whole subtree with it).
|
||||
// Slugs are language-independent.
|
||||
function onTitleInput(node, ev) {
|
||||
const title = ev.target.value.trim()
|
||||
if (!title || title === node.title) return
|
||||
debounce(`title:${node.path}`, async () => {
|
||||
await postStructure({ path: node.path, title })
|
||||
await postStructure({ path: node.path, title, lang: lang.value })
|
||||
})
|
||||
}
|
||||
|
||||
@@ -272,6 +309,11 @@ onMounted(() => {
|
||||
path.value = normPath(props.pagePath)
|
||||
refreshPages()
|
||||
addEventListener('pagerite:editor-shown', onEditorShown)
|
||||
// The language strip: site primary + configured targets.
|
||||
fetch('/_api/settings').then((r) => r.json()).then((s) => {
|
||||
primaryLang.value = s.primary_lang || 'en'
|
||||
siteLangs.value = s.translate_langs || []
|
||||
}).catch(() => { /* no strip */ })
|
||||
})
|
||||
|
||||
onUnmounted(() => {
|
||||
@@ -283,8 +325,29 @@ onUnmounted(() => {
|
||||
<template>
|
||||
<div class="structure-editor">
|
||||
<div v-if="saveError">{{ saveError }}</div>
|
||||
<div v-if="langOptions.length > 1" class="block lang-block">
|
||||
<div class="lang-strip">
|
||||
<button
|
||||
v-for="o in langOptions"
|
||||
:key="o.code"
|
||||
type="button"
|
||||
class="lang-flag"
|
||||
:class="{ active: o.tag === lang }"
|
||||
:title="o.primary
|
||||
? `${o.name} — the primary language; title edits affect all translations`
|
||||
: `${o.name} — title edits affect only this language`"
|
||||
@click="switchLang(o.tag)"
|
||||
>
|
||||
<span class="flag" v-html="o.flag" />
|
||||
</button>
|
||||
</div>
|
||||
<small v-if="lang" class="muted">
|
||||
viewing {{ currentLang.name }} titles — dimmed rows are untranslated
|
||||
(shown in the primary language); slugs never translate
|
||||
</small>
|
||||
</div>
|
||||
<section class="block structure">
|
||||
<StructureTree :nodes="tree" />
|
||||
<StructureTree :nodes="tree" :lang="lang" />
|
||||
</section>
|
||||
</div>
|
||||
</template>
|
||||
@@ -309,4 +372,52 @@ onUnmounted(() => {
|
||||
overflow-y: auto;
|
||||
min-height: 0;
|
||||
}
|
||||
|
||||
/* Language strip: the same flag chips as the PageEditor picker /
|
||||
localization tab; active = the language titles are shown/edited in. */
|
||||
.lang-strip {
|
||||
display: flex;
|
||||
gap: 0.4rem;
|
||||
}
|
||||
|
||||
.lang-flag {
|
||||
padding: 2px;
|
||||
background: none;
|
||||
border: 2px solid transparent;
|
||||
border-radius: 5px;
|
||||
cursor: pointer;
|
||||
opacity: 0.45;
|
||||
filter: grayscale(0.8);
|
||||
transition: opacity 0.15s, filter 0.15s, border-color 0.15s;
|
||||
}
|
||||
|
||||
.lang-flag:hover {
|
||||
opacity: 0.85;
|
||||
filter: none;
|
||||
}
|
||||
|
||||
.lang-flag.active {
|
||||
opacity: 1;
|
||||
filter: none;
|
||||
border-color: var(--accent);
|
||||
}
|
||||
|
||||
.flag {
|
||||
display: inline-flex;
|
||||
width: 18px;
|
||||
height: 12px;
|
||||
border-radius: 2px;
|
||||
overflow: hidden;
|
||||
border: 1px solid var(--line);
|
||||
}
|
||||
|
||||
.flag :deep(svg) {
|
||||
width: 100%;
|
||||
height: 100%;
|
||||
display: block;
|
||||
}
|
||||
|
||||
.muted {
|
||||
color: var(--muted);
|
||||
}
|
||||
</style>
|
||||
|
||||
@@ -1,7 +1,10 @@
|
||||
<script setup>
|
||||
// Recursive site-structure tree with drag-and-drop ordering (vue-draggable).
|
||||
// Nodes come from the server (GET /_api/pages via StructureEditor.vue) as
|
||||
// {slug, path, title, order, published, has_content, children}.
|
||||
// {slug, path, title, translated, order, published, has_content, children}.
|
||||
// With a `lang` prop (StructureEditor's language strip) the titles shown
|
||||
// are that language's; `translated` marks rows with an actual translation
|
||||
// (untranslated rows show the original title, dimmed).
|
||||
// Every node is real: a label whose title and slug are always editable
|
||||
// inline — the title saves while typing (and focusing it opens the page),
|
||||
// the slug commits on blur/Enter since it renames the path, moving the
|
||||
@@ -30,6 +33,10 @@ const props = defineProps({
|
||||
nodes: { type: Array, required: true },
|
||||
parentPath: { type: String, default: '' },
|
||||
depth: { type: Number, default: 0 },
|
||||
// StructureEditor's selected language ('' = original). Only used for the
|
||||
// untranslated-title styling here; the fetch and title edits live in the
|
||||
// parent (handlers.titleInput posts the lang with the op).
|
||||
lang: { type: String, default: '' },
|
||||
})
|
||||
|
||||
const handlers = inject('structureHandlers')
|
||||
@@ -126,8 +133,11 @@ function onEnd() {
|
||||
<template v-else>
|
||||
<input
|
||||
class="edit title-edit"
|
||||
:class="{ untranslated: lang && !element.translated }"
|
||||
:value="element.title"
|
||||
title="Label in the navigation — saves while typing; click opens the page"
|
||||
:title="lang && !element.translated
|
||||
? 'No translation yet — showing the original; typing creates the translated title'
|
||||
: 'Label in the navigation — saves while typing; click opens the page'"
|
||||
@input="handlers.titleInput(element, $event)"
|
||||
@focus="handlers.open(element.path)"
|
||||
/>
|
||||
@@ -158,6 +168,7 @@ function onEnd() {
|
||||
:nodes="element.children"
|
||||
:parent-path="element.path"
|
||||
:depth="depth + 1"
|
||||
:lang="lang"
|
||||
/>
|
||||
</div>
|
||||
</template>
|
||||
@@ -274,6 +285,13 @@ body.tree-dragging .treelist {
|
||||
cursor: text;
|
||||
}
|
||||
|
||||
/* With a language selected (StructureEditor's strip), rows without an
|
||||
actual translation show the original title dimmed and italic. */
|
||||
.title-edit.untranslated {
|
||||
color: var(--muted);
|
||||
font-style: italic;
|
||||
}
|
||||
|
||||
.slug-edit {
|
||||
font-family: var(--font-code);
|
||||
}
|
||||
|
||||
+71
-200
@@ -49,7 +49,7 @@ from mediapreview import dispatch
|
||||
from pydantic import BaseModel
|
||||
from zstandard import ZstdCompressor
|
||||
|
||||
from pagerite import analytics, i18n, masking, seed, translate, views
|
||||
from pagerite import analytics, i18n, seed, translate, views
|
||||
from pagerite.__main__ import DEVMODE
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.data import (
|
||||
@@ -411,7 +411,7 @@ def _render_html(kind: str, path: str, base_url: str, lang: str = i18n.ORIGINAL_
|
||||
if kind == "page":
|
||||
# A selected language without an actual translation renders the
|
||||
# original (translation is None = English; see docs/localization.md).
|
||||
translation = i18n.get_translation(path, lang, data) if lang != i18n.ORIGINAL_LANGUAGE else None
|
||||
translation = i18n.get_translation(data, path, lang) if lang != i18n.ORIGINAL_LANGUAGE else None
|
||||
return views.render_page(data.menu, data, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, base_url, transition=data.transition, lang=lang, translation=translation, link_lang=link_lang)
|
||||
if kind == "category":
|
||||
# A category has no Markdown of its own; only the title map
|
||||
@@ -439,7 +439,7 @@ def _invalidate_pages() -> None:
|
||||
global _render_gen
|
||||
_render_gen += 1
|
||||
_cached_body.cache_clear()
|
||||
_schedule_translation_dispatch()
|
||||
dispatcher.schedule()
|
||||
|
||||
|
||||
@lru_cache(maxsize=128)
|
||||
@@ -516,11 +516,21 @@ class PageIn(BaseModel):
|
||||
|
||||
|
||||
@app.get("/_api/pages")
|
||||
async def list_pages() -> list[dict]:
|
||||
async def list_pages(lang: str | None = None) -> list[dict]:
|
||||
"""The site tree for the structure editor (all nodes, drafts included).
|
||||
|
||||
Nested by slug; each node carries its full path, menu order and flags.
|
||||
With a ``?lang=`` translation, titles come out in that language where a
|
||||
translation exists (``translated`` flags it; the row still falls back
|
||||
to the original title otherwise) — the structure itself (slugs, order,
|
||||
hierarchy) is language-independent.
|
||||
"""
|
||||
tag = i18n.base_tag(lang or "")
|
||||
titles = (
|
||||
i18n.title_map(data, tag)
|
||||
if tag and tag != i18n.ORIGINAL_LANGUAGE
|
||||
else {}
|
||||
)
|
||||
|
||||
def dump(nodes: dict[str, Node], prefix: str) -> list[dict]:
|
||||
out = []
|
||||
@@ -529,7 +539,8 @@ async def list_pages() -> list[dict]:
|
||||
out.append({
|
||||
"slug": slug,
|
||||
"path": path,
|
||||
"title": node.title,
|
||||
"title": titles.get(path) or node.title,
|
||||
"translated": path in titles,
|
||||
"order": node.order,
|
||||
"published": node.published,
|
||||
"has_content": node.chunks is not None,
|
||||
@@ -565,14 +576,9 @@ async def save_page(path: str, page: PageIn, lang: str | None = None) -> None:
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
patch = i18n.make_patch(
|
||||
i18n.hybrid_markdown(data, node, path, lang), page.markdown
|
||||
)
|
||||
if patch.hunks:
|
||||
with kanta.transaction("save translation", extra=path):
|
||||
# Patches alone make the translated version exist.
|
||||
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
|
||||
node.langs[lang] = True
|
||||
if i18n.add_patch(data, node, path, lang, page.markdown):
|
||||
_invalidate_pages()
|
||||
return
|
||||
with kanta.transaction("save page", extra=path):
|
||||
@@ -596,12 +602,19 @@ class StructureOp(BaseModel):
|
||||
just the top-level node with slug "": renaming it away leaves no front
|
||||
page ("/" then redirects to the first nav item), and any childless
|
||||
top-level node can take the empty slug to become the front page.
|
||||
|
||||
With `lang` (a translation, not the primary language) a `title` edit
|
||||
writes a per-language title fragment instead of the original — the same
|
||||
storage as machine title translations (docs/localization.md); sending
|
||||
the original's text removes the override. Structural fields are not
|
||||
combinable with a translated title edit.
|
||||
"""
|
||||
|
||||
path: str
|
||||
order: float | None = None
|
||||
move_to: str | None = None
|
||||
title: str | None = None
|
||||
lang: str | None = None
|
||||
|
||||
|
||||
@app.post("/_api/structure", status_code=204)
|
||||
@@ -612,6 +625,14 @@ async def update_structure(op: StructureOp) -> None:
|
||||
if chain is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
node = chain[-1]
|
||||
lang = i18n.base_tag(op.lang or "")
|
||||
if op.title is not None and lang and lang != i18n.ORIGINAL_LANGUAGE:
|
||||
# Translated title (i18n.set_title_translation): original title,
|
||||
# slugs and hierarchy stay untouched.
|
||||
with kanta.transaction("translate title", extra=path):
|
||||
if i18n.set_title_translation(data, node, lang, op.title):
|
||||
_invalidate_pages()
|
||||
return
|
||||
target = op.move_to.strip("/") if op.move_to is not None else None
|
||||
if target is not None and target != path:
|
||||
_check_reserved(target)
|
||||
@@ -699,6 +720,22 @@ async def put_settings(settings: SettingsIn) -> None:
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@app.delete("/_api/translations", status_code=204)
|
||||
async def delete_translations() -> None:
|
||||
"""Drop all machine translations (Data.trans) so the dispatcher
|
||||
re-translates everything from scratch (a "refresh translations" action:
|
||||
the invalidation hook re-offers every fragment to connected
|
||||
translators). User patches are kept; the availability index
|
||||
(node.langs) is rebuilt from them — patches alone still make a language
|
||||
exist on a page."""
|
||||
with kanta.transaction("refresh translations"):
|
||||
i18n.clear_translations(data)
|
||||
_invalidate_pages()
|
||||
# Fragments rejected this run (segment validation) stay skipped no
|
||||
# longer: a refresh is precisely the "another chance" for them.
|
||||
dispatcher.validation_failures.clear()
|
||||
|
||||
|
||||
@app.put("/_api/settings/favicon")
|
||||
async def put_favicon(request: Request) -> dict[str, str]:
|
||||
"""Upload a favicon into the content-addressed store and activate it.
|
||||
@@ -1013,98 +1050,9 @@ async def delete_page(path: str) -> None:
|
||||
|
||||
|
||||
# WebSocket API for external translation services (not under /_api: it is keyed
|
||||
# with Data.translate_keys instead of the SSO forward-auth). The server is a
|
||||
# dispatcher: one single-item job at a time per connection, offered in the
|
||||
# intersection of the wanted languages (Data.translate_langs) and the
|
||||
# connection's announced capabilities. Results are matched to content by
|
||||
# chunk key alone.
|
||||
class _TranslatorState:
|
||||
"""One connected translator socket: the language codes it announced as
|
||||
capabilities (Hello) and the (lang, chunk-key) job currently in flight
|
||||
on it, with the mask spans to restore into its Result
|
||||
(pagerite/masking.py) — one at a time, the next is sent only after its
|
||||
Result.
|
||||
|
||||
Per-connection only: in-flight lives solely here, so on disconnect the
|
||||
item simply becomes pending again and is re-offered to any free capable
|
||||
connection."""
|
||||
|
||||
def __init__(self, capable: set[str]) -> None:
|
||||
self.capable = capable
|
||||
self.inflight: tuple[str, bytes] | None = None
|
||||
self.spans: list[str] = [] # mask spans of the in-flight job
|
||||
|
||||
|
||||
#: Connected translator sockets and their per-connection state.
|
||||
_translator_clients: dict[WebSocket, _TranslatorState] = {}
|
||||
|
||||
#: (lang, chunk key) of fragments whose result failed sentinel validation
|
||||
#: (masking.unmask): skipped on later dispatches this run — generation is
|
||||
#: near-deterministic, so an immediate retry would just re-fail.
|
||||
_mask_failures: set[tuple[str, bytes]] = set()
|
||||
|
||||
|
||||
def _schedule_translation_dispatch() -> None:
|
||||
"""Schedule a dispatch pass, if any translator is connected.
|
||||
|
||||
The content-change hook is _invalidate_pages (sync, called inside
|
||||
transactions): the task first runs once the current coroutine awaits
|
||||
again, i.e. after the transaction has committed. No-op without a
|
||||
running loop (CLI use).
|
||||
"""
|
||||
if not _translator_clients:
|
||||
return
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return
|
||||
asyncio.create_task(_dispatch_translations())
|
||||
|
||||
|
||||
async def _dispatch_translations() -> None:
|
||||
"""Offer one pending item to every free capable connection.
|
||||
|
||||
Runs on every relevant event: Hello, Result, disconnect and content
|
||||
change (via _invalidate_pages). A connection with no wanted ∩ capable
|
||||
overlap simply stays idle. Pending is derived from the trans store
|
||||
(translate.pending_items) minus the items in flight on any connection.
|
||||
"""
|
||||
wanted = {
|
||||
tag
|
||||
for lang in data.translate_langs
|
||||
if (tag := i18n.base_tag(lang)) and tag != i18n.ORIGINAL_LANGUAGE
|
||||
}
|
||||
if not wanted:
|
||||
return
|
||||
for ws, state in list(_translator_clients.items()):
|
||||
if state.inflight is not None:
|
||||
continue
|
||||
langs = wanted & state.capable
|
||||
if not langs:
|
||||
continue
|
||||
inflight = {s.inflight for s in _translator_clients.values() if s.inflight}
|
||||
job = None
|
||||
spans: list[str] = []
|
||||
for lang in sorted(langs):
|
||||
for item in translate.pending_items(data, lang):
|
||||
if (lang, item.key) in inflight or (lang, item.key) in _mask_failures:
|
||||
continue
|
||||
masked, spans = masking.mask(item.text)
|
||||
job = translate.Job(
|
||||
lang=lang, key=item.key, text=masked,
|
||||
path=item.path, kind=item.kind,
|
||||
)
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is None:
|
||||
continue
|
||||
state.inflight = (job.lang, job.key) # before the await: no double-assign
|
||||
state.spans = spans
|
||||
try:
|
||||
await ws.send_text(msgspec.json.encode(job).decode())
|
||||
except Exception: # send failed: the receive loop cleans up
|
||||
_translator_clients.pop(ws, None)
|
||||
# with Data.translate_keys instead of the SSO forward-auth). The dispatcher —
|
||||
# protocol, connected clients and the job pipeline — lives in translate.py.
|
||||
dispatcher = translate.Dispatcher(data, kanta, _invalidate_pages)
|
||||
|
||||
|
||||
@app.websocket("/_translate/{clientkey}")
|
||||
@@ -1114,76 +1062,9 @@ async def translate_ws(ws: WebSocket, clientkey: str) -> None:
|
||||
Deliberately NOT under /_api/: the external forward-auth is skipped;
|
||||
the server-generated client key in the path is the access control
|
||||
(``Data.translate_keys``: key -> display name; the first is generated
|
||||
at bootstrap, all are shown in the admin's
|
||||
/_api/settings). A wrong/empty key rejects the handshake — closing
|
||||
before accept makes Starlette answer HTTP 403.
|
||||
|
||||
Protocol (JSON frames, msgspec structs in translate.py): the client
|
||||
opens with Hello(langs) announcing its CAPABILITIES — the language
|
||||
codes its model can produce (normalized to base subtags; "en"/empty
|
||||
dropped). The dispatcher sends one Job(lang, key, text, path, kind)
|
||||
at a time and waits for the matching Result(lang, key, text) before
|
||||
offering the next. A Result without an in-flight job or with a
|
||||
different (lang, key), a duplicate Hello, or any malformed frame
|
||||
closes the socket with a protocol error.
|
||||
at bootstrap, all are shown in the admin's /_api/settings).
|
||||
"""
|
||||
if clientkey not in data.translate_keys:
|
||||
await ws.close(code=1008) # policy violation; pre-accept = HTTP 403
|
||||
return
|
||||
await ws.accept()
|
||||
state: _TranslatorState | None = None
|
||||
try:
|
||||
while True:
|
||||
raw = await ws.receive_text()
|
||||
try:
|
||||
msg = msgspec.json.decode(raw.encode(), type=translate.ClientMsg)
|
||||
except msgspec.DecodeError:
|
||||
await ws.close(code=1002) # protocol error
|
||||
return
|
||||
if isinstance(msg, translate.Hello):
|
||||
if state is not None: # one Hello per connection
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
state = _TranslatorState({
|
||||
tag
|
||||
for lang in msg.langs
|
||||
if (tag := i18n.base_tag(lang)) and tag != i18n.ORIGINAL_LANGUAGE
|
||||
})
|
||||
_translator_clients[ws] = state
|
||||
_schedule_translation_dispatch()
|
||||
else: # translate.Result
|
||||
lang = i18n.base_tag(msg.lang)
|
||||
if (
|
||||
state is None # results before Hello
|
||||
or state.inflight is None # no job in flight
|
||||
or (lang, msg.key) != state.inflight # wrong job
|
||||
):
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
text = masking.unmask(msg.text, state.spans)
|
||||
state.inflight = None
|
||||
state.spans = []
|
||||
if text is None:
|
||||
# The model mangled the sentinels: drop the result and
|
||||
# skip the fragment for this run (it stays pending; a
|
||||
# restart or a masking/prompt change gets another chance).
|
||||
_mask_failures.add((lang, msg.key))
|
||||
print(f"[{lang}] result for chunk {msg.key.hex()} rejected: sentinels mangled")
|
||||
_schedule_translation_dispatch()
|
||||
continue
|
||||
with kanta.transaction("translator results", user=clientkey, extra=lang):
|
||||
paths = translate.store_results(
|
||||
data, lang, [translate.TransResult(key=msg.key, text=text)]
|
||||
)
|
||||
_invalidate_pages() # schedules the next dispatch
|
||||
if paths:
|
||||
print(f"[{lang}] now available for {len(paths)} page(s): {', '.join(sorted(paths))}")
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
finally:
|
||||
if _translator_clients.pop(ws, None) is not None:
|
||||
# The in-flight item (if any) is pending again; offer it around.
|
||||
_schedule_translation_dispatch()
|
||||
await dispatcher.handle_ws(ws, clientkey)
|
||||
|
||||
|
||||
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
|
||||
@@ -1678,35 +1559,25 @@ async def editor_ws(ws: WebSocket) -> None:
|
||||
else:
|
||||
node = old if old is not None else _ensure(data.menu, path)
|
||||
if translated:
|
||||
# Diff the editor's shadow base (the hybrid the
|
||||
# user started editing, sent along as "base";
|
||||
# absent: the current hybrid) against the
|
||||
# submitted text and append a Patch; node.chunks
|
||||
# and the original-language fields stay
|
||||
# untouched.
|
||||
# node.chunks and the original-language fields
|
||||
# stay untouched: the markdown diff (against the
|
||||
# editor's shadow "base" — the hybrid it started
|
||||
# from; absent: the current hybrid) is appended
|
||||
# as a Patch, a changed title becomes a
|
||||
# per-language title override (i18n).
|
||||
changed = False
|
||||
if "markdown" in msg:
|
||||
base = msg.get("base")
|
||||
if not isinstance(base, str):
|
||||
base = i18n.hybrid_markdown(data, node, path, lang)
|
||||
patch = i18n.make_patch(base, msg["markdown"])
|
||||
if patch.hunks:
|
||||
# Patches alone make the translated
|
||||
# version exist.
|
||||
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
|
||||
node.langs[lang] = True
|
||||
_invalidate_pages()
|
||||
if "title" in msg and node.title:
|
||||
# A changed title becomes a fragment in
|
||||
# Data.trans, keyed by the original title's
|
||||
# chunk hash — same storage as machine
|
||||
# title translations.
|
||||
effective = (
|
||||
i18n.title_map(data, lang).get(path) or node.title
|
||||
changed = i18n.add_patch(
|
||||
data, node, path, lang, msg["markdown"],
|
||||
base=base if isinstance(base, str) else None,
|
||||
)
|
||||
if msg["title"] != effective:
|
||||
key = i18n.chunk_key(node.title)
|
||||
data.trans.setdefault(key, {})[lang] = msg["title"]
|
||||
node.langs[lang] = True
|
||||
if "title" in msg and node.title:
|
||||
changed = (
|
||||
i18n.set_title_translation(data, node, lang, msg["title"])
|
||||
or changed
|
||||
)
|
||||
if changed:
|
||||
_invalidate_pages()
|
||||
else:
|
||||
if "markdown" in msg:
|
||||
@@ -1850,7 +1721,7 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda l: l in node.langs,
|
||||
lambda tag: tag in node.langs,
|
||||
)
|
||||
# A ?lang= override is replicated onto the page's navigation links
|
||||
# (link_lang), so clicks and prefetches stay in the chosen language.
|
||||
@@ -1893,7 +1764,7 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda l: l in subtree_langs,
|
||||
lambda tag: tag in subtree_langs,
|
||||
)
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
if _is_trackable_path(path):
|
||||
|
||||
+2
-2
@@ -11,7 +11,7 @@ import re
|
||||
|
||||
import blake3
|
||||
|
||||
from pagerite.masking import has_prose
|
||||
from pagerite.segments import has_prose
|
||||
|
||||
#: Fenced code block opener/closer: up to 3 spaces indent, then 3+
|
||||
#: backticks or tildes (CommonMark).
|
||||
@@ -126,7 +126,7 @@ def chunk_key(text: str) -> bytes:
|
||||
|
||||
def needs_translation(chunk: str) -> bool:
|
||||
"""False for chunks without prose: pure code fences, HTML blocks, and
|
||||
anything whose masked form (pagerite/masking.py) has no letters left —
|
||||
anything that yields no translatable segments (pagerite/segments.py) —
|
||||
container fences, lone {placeholders}, reference definitions.
|
||||
|
||||
These are inherently no-translate (docs/migrate.md): derived from the
|
||||
|
||||
+66
-1
@@ -37,6 +37,14 @@ def base_tag(tag: str) -> str:
|
||||
return tag.strip().lower().partition("-")[0]
|
||||
|
||||
|
||||
def translation_tag(lang: str | None) -> str:
|
||||
"""The normalized translation selector: the base subtag of ``lang``, or
|
||||
"" when it is absent or the original language (the original is never a
|
||||
translation target)."""
|
||||
tag = base_tag(lang or "")
|
||||
return tag if tag != ORIGINAL_LANGUAGE else ""
|
||||
|
||||
|
||||
def parse_accept_language(header: str) -> list[str]:
|
||||
"""Accept-Language header as an ordered, deduped list of base subtags.
|
||||
|
||||
@@ -141,6 +149,63 @@ def hybrid_markdown(data: Data, node: Node, path: str, lang: str) -> str:
|
||||
return hybrid
|
||||
|
||||
|
||||
def add_patch(
|
||||
data: Data, node: Node, path: str, lang: str, edited: str, base: str | None = None
|
||||
) -> bool:
|
||||
"""Record a translated-view edit as a user Patch: the minimal diff of
|
||||
``edited`` against ``base`` (default: the currently served hybrid),
|
||||
appended to the language's patch list. Patches alone make the
|
||||
translated version exist, so ``node.langs`` is set. Returns True when
|
||||
a patch was stored. Pure data ops — the caller wraps in a transaction
|
||||
and invalidates."""
|
||||
patch = make_patch(base if base is not None else hybrid_markdown(data, node, path, lang), edited)
|
||||
if not patch.hunks:
|
||||
return False
|
||||
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
|
||||
node.langs[lang] = True
|
||||
return True
|
||||
|
||||
|
||||
def set_title_translation(data: Data, node: Node, lang: str, title: str) -> bool:
|
||||
"""Record (or drop) a per-language title override: a fragment in
|
||||
``Data.trans`` keyed by the ORIGINAL title's chunk hash — the same
|
||||
storage machine title translations use, overriding them. Sending the
|
||||
original's text drops the override. Returns True when anything changed.
|
||||
Pure data ops — the caller wraps in a transaction and invalidates."""
|
||||
key = chunk_key(node.title)
|
||||
current = data.trans.get(key, {}).get(lang)
|
||||
if title == node.title:
|
||||
if current is None:
|
||||
return False
|
||||
del data.trans[key][lang]
|
||||
return True
|
||||
if current == title:
|
||||
return False
|
||||
data.trans.setdefault(key, {})[lang] = title
|
||||
node.langs[lang] = True
|
||||
return True
|
||||
|
||||
|
||||
def clear_translations(data: Data) -> None:
|
||||
"""Drop all machine translations (``Data.trans``) and rebuild the
|
||||
availability index (``node.langs``) from the surviving user patches —
|
||||
patches alone make a language exist on a page. Pure data ops — the
|
||||
caller wraps in a transaction and invalidates."""
|
||||
data.trans.clear()
|
||||
patch_langs: dict[str, set[str]] = {}
|
||||
for key in data.patches:
|
||||
path, _, lang = key.rpartition(":")
|
||||
patch_langs.setdefault(path, set()).add(lang)
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str) -> None:
|
||||
for slug, node in nodes.items():
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
node.langs = {lang: True for lang in patch_langs.get(path, ())}
|
||||
walk(node.children, path)
|
||||
|
||||
walk(data.menu, "")
|
||||
|
||||
|
||||
def title_map(data: Data, lang: str) -> dict[str, str]:
|
||||
"""path -> translated title for every node that has one.
|
||||
|
||||
@@ -174,7 +239,7 @@ def subtree_languages(node: Node) -> set[str]:
|
||||
return langs
|
||||
|
||||
|
||||
def get_translation(path: str, lang: str, data: Data) -> Translation | None:
|
||||
def get_translation(data: Data, path: str, lang: str) -> Translation | None:
|
||||
"""The translation of the page at ``path`` for ``lang``, or None.
|
||||
|
||||
None when the page does not exist or is not available in ``lang``:
|
||||
|
||||
+27
-15
@@ -447,14 +447,22 @@ def _heading_ids(state) -> None:
|
||||
wrap(i, token, f"#{hid}")
|
||||
|
||||
|
||||
md = (
|
||||
def make_md(*, verbatim: bool = False) -> MarkdownIt:
|
||||
"""A fully configured parser. The module-level ``md`` (below) is the
|
||||
render instance; ``verbatim=True`` builds the segmentation instance for
|
||||
segments.py, where token text must stay byte-identical to the source so
|
||||
prose spans can be spliced back by offset: no typographer (quotes and
|
||||
dashes stay straight), no tasklist label wrapping (the item text stays
|
||||
a plain text token), and soft line breaks (wrapped prose merges into
|
||||
one segment instead of splitting at hardbreaks)."""
|
||||
parser = (
|
||||
MarkdownIt(
|
||||
"default",
|
||||
{
|
||||
"html": True,
|
||||
"highlight": _highlight,
|
||||
"typographer": True,
|
||||
"breaks": True,
|
||||
"typographer": not verbatim,
|
||||
"breaks": not verbatim,
|
||||
},
|
||||
)
|
||||
.use(attrs_plugin)
|
||||
@@ -462,24 +470,28 @@ md = (
|
||||
.use(container_plugin, "block", validate=_container_validate)
|
||||
.use(footnote_plugin)
|
||||
.use(deflist_plugin)
|
||||
# label_after: the item text is wrapped in <label for> after the
|
||||
# checkbox, so clicking the text toggles it.
|
||||
.use(tasklists_plugin, enabled=True, label=True, label_after=True)
|
||||
# label wrapping (render) puts the item text inside the checkbox
|
||||
# <label> html_inline; without it the text stays a plain token.
|
||||
.use(tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim)
|
||||
.use(gfm_autolink_plugin)
|
||||
.use(sub_plugin)
|
||||
.use(superscript_plugin)
|
||||
)
|
||||
md.add_render_rule("image", _image_rule)
|
||||
md.add_render_rule("fence", _fence_rule)
|
||||
parser.add_render_rule("image", _image_rule)
|
||||
parser.add_render_rule("fence", _fence_rule)
|
||||
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
|
||||
md.options["alerts"] = True
|
||||
parser.options["alerts"] = True
|
||||
# Block attrs must be stripped before the typographer curlifies their quotes.
|
||||
md.core.ruler.before("replacements", "block_attrs", _block_attrs)
|
||||
md.core.ruler.push("container_attrs", _container_attrs)
|
||||
md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
|
||||
md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
|
||||
md.core.ruler.push("shorten_autolinks", _shorten_autolinks)
|
||||
md.core.ruler.push("heading_ids", _heading_ids)
|
||||
parser.core.ruler.before("replacements", "block_attrs", _block_attrs)
|
||||
parser.core.ruler.push("container_attrs", _container_attrs)
|
||||
parser.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
|
||||
parser.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
|
||||
parser.core.ruler.push("shorten_autolinks", _shorten_autolinks)
|
||||
parser.core.ruler.push("heading_ids", _heading_ids)
|
||||
return parser
|
||||
|
||||
|
||||
md = make_md()
|
||||
|
||||
|
||||
# Text-length thresholds (visible characters, code blocks excluded) for the
|
||||
|
||||
@@ -1,170 +0,0 @@
|
||||
"""Masking of non-translatable spans for the machine-translation round trip.
|
||||
|
||||
A translator model must copy technical spans (code, URLs, {placeholders},
|
||||
attrs, footnote and link labels, container names, HTML tags) byte-identically
|
||||
while translating the prose around them — and small models translate anything
|
||||
that looks like a word (a {dates} placeholder once came back as
|
||||
{päivämäärät}). So before a fragment is dispatched, each such span is
|
||||
replaced with a numbered sentinel (``⟦1⟧``, ``⟦2⟧``, ...) — the model only
|
||||
ever sees prose — and on the way back the sentinels are restored by number
|
||||
(``unmask``). A result whose sentinels did not all survive — missing,
|
||||
duplicated or out of range — is rejected and the fragment stays pending.
|
||||
|
||||
Punctuation structure (*, |, [], (), :::) is not masked: it carries no
|
||||
lexical content and models preserve it. Link and image text — including alt
|
||||
text and captions — stays visible for translation; only the destination is
|
||||
masked. Rule order matters: earlier rules consume syntax later ones would
|
||||
misread, and no rule may match across or inside an already emitted sentinel
|
||||
(the container-fence rule runs before the brace rule for that reason).
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
#: A masked span marker: the span's 1-based number in brackets that never
|
||||
#: appear in content and are atomic enough for a model to copy verbatim.
|
||||
#: unmask() validates survival, so a model that mangles them only loses its
|
||||
#: own result.
|
||||
_SENTINEL = re.compile(r"⟦(\d+)⟧")
|
||||
|
||||
#: URL-ish span: an <angle-bracketed> destination, or a whitespace-free run
|
||||
#: allowing one level of balanced parens (Wikipedia-style).
|
||||
_URLISH = r"<[^<>\n]*>|[^\s()]*(?:\([^()\n]*\)[^\s()]*)*"
|
||||
|
||||
#: Inline code: matching backtick runs, whole span masked. The content may
|
||||
#: not cross a paragraph break, so a stray backtick cannot swallow the rest
|
||||
#: of the chunk.
|
||||
_CODE = re.compile(r"(`+)((?:(?!\n\n).)+?)\1(?!`)", re.DOTALL)
|
||||
|
||||
#: Container fence line (`:::: aside {.x}`): the name-and-attrs tail is
|
||||
#: masked; a bare `:::` has nothing to mask. Runs before the brace rule so
|
||||
#: fence-line attrs are masked together with the name.
|
||||
_FENCE = re.compile(r"^( {0,3}:{3,})[ \t]*(\S[^\n]*)", re.MULTILINE)
|
||||
|
||||
#: Link/image destination: `[text](url "title")` -> `[text](⟦N⟧ "title")`.
|
||||
_DEST = re.compile(r"(\]\(\s*)(" + _URLISH + r")")
|
||||
|
||||
#: Autolinks and inline HTML (<http://...>, <b>, <!-- ... -->, <? ... ?>).
|
||||
#: A `<` followed by whitespace (a prose "a < b") is not matched.
|
||||
_TAG = re.compile(r"<[A-Za-z/!?][^<>\n]*>")
|
||||
|
||||
#: Footnote definition `[^label]: text...` — label masked; the text after
|
||||
#: the colon is prose.
|
||||
_FOOTDEF = re.compile(r"^( {0,3}\[\^)([^\]\n]+)(\]:)", re.MULTILINE)
|
||||
|
||||
#: Reference-style link definition `[label]: url "title"` — label and
|
||||
#: destination masked, title stays visible.
|
||||
_LINKDEF = re.compile(r"^( {0,3}\[)(?!\^)([^\]\n]+)(\]:[ \t]*)(" + _URLISH + r")", re.MULTILINE)
|
||||
|
||||
#: Footnote reference `[^label]` ((?!:) — definitions are _FOOTDEF's).
|
||||
_FOOTREF = re.compile(r"\[\^([^\]\n]+)\](?!:)")
|
||||
|
||||
#: Reference-style link usage `[text][label]` — the label.
|
||||
_REFPAIR = re.compile(r"(\][ \t]?\[)([^\]\n]+)(\])")
|
||||
|
||||
#: Any {...} span: {placeholders} and {#id .class} attrs alike.
|
||||
_BRACES = re.compile(r"\{[^{}\n]*\}")
|
||||
|
||||
#: Bare URLs in prose (GFM autolinks); trailing sentence punctuation stays
|
||||
#: outside the mask.
|
||||
_BARE_URL = re.compile(r"(?<![\w/])(?:https?://|www\.)[^\s<>()\[\]]+")
|
||||
|
||||
#: Any Unicode letter (digits and underscore are not prose).
|
||||
_LETTER = re.compile(r"[^\W\d_]")
|
||||
|
||||
|
||||
def mask(text: str) -> tuple[str, list[str]]:
|
||||
"""Replace every non-translatable span with a ⟦N⟧ sentinel; return the
|
||||
masked text and the original spans in sentinel order."""
|
||||
spans: list[str] = []
|
||||
|
||||
def emit(original: str) -> str:
|
||||
if not original:
|
||||
return original
|
||||
spans.append(original)
|
||||
return f"⟦{len(spans)}⟧"
|
||||
|
||||
for sub in (_code, _fence, _dest, _tag, _footdef, _linkdef, _footref,
|
||||
_refpair, _braces, _bare_url):
|
||||
text = sub(text, emit)
|
||||
return text, spans
|
||||
|
||||
|
||||
def unmask(text: str, spans: list[str]) -> str | None:
|
||||
"""Restore the masked spans into a translated fragment; None when the
|
||||
sentinels did not all survive intact (missing, duplicated or out of
|
||||
range) — the caller drops the result and the fragment stays pending.
|
||||
|
||||
Order is not checked: translations legitimately reorder spans.
|
||||
"""
|
||||
if not spans:
|
||||
return text
|
||||
counts: dict[int, int] = {}
|
||||
|
||||
def repl(m: re.Match) -> str:
|
||||
n = int(m.group(1))
|
||||
counts[n] = counts.get(n, 0) + 1
|
||||
return spans[n - 1] if 0 < n <= len(spans) else m.group(0)
|
||||
|
||||
restored = _SENTINEL.sub(repl, text)
|
||||
if counts != dict.fromkeys(range(1, len(spans) + 1), 1):
|
||||
return None
|
||||
return restored
|
||||
|
||||
|
||||
def has_prose(text: str) -> bool:
|
||||
"""True when the masked form still contains a letter — i.e. there is
|
||||
something for a translator to translate. Chunks that are all markup,
|
||||
placeholders or code (a lone {dates}, container fences, reference
|
||||
definitions) have no business reaching the model: every language
|
||||
renders them from the original chunk."""
|
||||
return bool(_LETTER.search(mask(text)[0]))
|
||||
|
||||
|
||||
def _code(text: str, emit) -> str:
|
||||
return _CODE.sub(lambda m: emit(m.group(0)), text)
|
||||
|
||||
|
||||
def _fence(text: str, emit) -> str:
|
||||
return _FENCE.sub(lambda m: m.group(1) + emit(m.group(2)), text)
|
||||
|
||||
|
||||
def _dest(text: str, emit) -> str:
|
||||
return _DEST.sub(lambda m: m.group(1) + emit(m.group(2)), text)
|
||||
|
||||
|
||||
def _tag(text: str, emit) -> str:
|
||||
return _TAG.sub(lambda m: emit(m.group(0)), text)
|
||||
|
||||
|
||||
def _footdef(text: str, emit) -> str:
|
||||
return _FOOTDEF.sub(lambda m: m.group(1) + emit(m.group(2)) + m.group(3), text)
|
||||
|
||||
|
||||
def _linkdef(text: str, emit) -> str:
|
||||
return _LINKDEF.sub(
|
||||
lambda m: m.group(1) + emit(m.group(2)) + m.group(3) + emit(m.group(4)), text
|
||||
)
|
||||
|
||||
|
||||
def _footref(text: str, emit) -> str:
|
||||
return _FOOTREF.sub(lambda m: "[^" + emit(m.group(1)) + "]", text)
|
||||
|
||||
|
||||
def _refpair(text: str, emit) -> str:
|
||||
return _REFPAIR.sub(lambda m: m.group(1) + emit(m.group(2)) + m.group(3), text)
|
||||
|
||||
|
||||
def _braces(text: str, emit) -> str:
|
||||
return _BRACES.sub(lambda m: emit(m.group(0)), text)
|
||||
|
||||
|
||||
def _bare_url(text: str, emit) -> str:
|
||||
def repl(m: re.Match) -> str:
|
||||
core = m.group(0)
|
||||
tail = ""
|
||||
while core and core[-1] in ".,;:!?":
|
||||
tail = core[-1] + tail
|
||||
core = core[:-1]
|
||||
return emit(core) + tail
|
||||
|
||||
return _BARE_URL.sub(repl, text)
|
||||
+2
-2
@@ -114,7 +114,7 @@ Headings from `##` down organize the article. On pages with at least three of th
|
||||
> and a blank `>` line starts a new paragraph.
|
||||
|
||||
> [!NOTE]
|
||||
> GitHub-style alerts — NOTE, TIP, IMPORTANT, WARNING, CAUTION —
|
||||
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
|
||||
> render as callout boxes.
|
||||
```
|
||||
|
||||
@@ -122,7 +122,7 @@ Headings from `##` down organize the article. On pages with at least three of th
|
||||
> and a blank `>` line starts a new paragraph.
|
||||
|
||||
> [!NOTE]
|
||||
> GitHub-style alerts — NOTE, TIP, IMPORTANT, WARNING, CAUTION —
|
||||
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
|
||||
> render as callout boxes.
|
||||
|
||||
## Code
|
||||
|
||||
@@ -0,0 +1,257 @@
|
||||
"""Segmented translation round trip: prose out, translations back in.
|
||||
|
||||
A translator model mangles anything that is not plain prose — sentinels get
|
||||
renumbered, ``![`` becomes sentence punctuation, stray ``<br>`` tags appear.
|
||||
So the model is never shown any of it: a fragment (a Markdown chunk or a
|
||||
node title) is parsed with the project's own markdown-it setup
|
||||
(``markdown.make_md(verbatim=True)`` — extensions included, so container,
|
||||
attrs, footnote and tasklist syntax never leaks into text tokens) and split
|
||||
into **prose segments**: the merged text runs, plus image alt texts and
|
||||
link/image titles. Only those cross the wire, as a plain list of strings
|
||||
(Job.texts / Result.texts in translate.py) — accompanied, per segment, by
|
||||
a CONTEXT (Job.contexts): a segment carved out of a larger block (a link
|
||||
text, a partial run) carries the block's plain text, so the model sees the
|
||||
sentence it lives in; whole-block segments are self-contextualizing and
|
||||
carry "". Title fragments carry the article's opening instead (assigned by
|
||||
the dispatcher from TransItem.context).
|
||||
|
||||
Reassembly is server-side offset splicing, not text the model produced:
|
||||
each segment's source span was located at dispatch (``split``), and
|
||||
``join`` swaps in the translations. Markup therefore cannot break — it
|
||||
never left the server. A returned segment must still be pure prose itself
|
||||
(the model could inject markup INTO a segment); anything else — count
|
||||
mismatch, empty segment, markup tokens — rejects the whole result and the
|
||||
fragment stays pending.
|
||||
|
||||
Translations legitimately reorder markup within a sentence... but segments
|
||||
splice back at fixed positions, so a link or image stays where the original
|
||||
put it. That is the accepted trade-off for never feeding the model markup
|
||||
(docs/localization.md).
|
||||
|
||||
Locating is best effort: a run that is not a verbatim source substring
|
||||
(entity-decoded text, backslash escapes) is skipped — it simply stays in
|
||||
the original language. So is any piece containing "<": "<" is the
|
||||
prose/markup boundary on the wire — translators cut their output there,
|
||||
so such pieces could not survive the round trip.
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
from pagerite.markdown import make_md
|
||||
|
||||
#: The segmentation parser: the project's own markdown-it, verbatim flavor
|
||||
#: (see make_md). Never used for rendering.
|
||||
_MD = make_md(verbatim=True)
|
||||
|
||||
#: Any Unicode letter (digits and underscore are not prose).
|
||||
_LETTER = re.compile(r"[^\W\d_]")
|
||||
|
||||
#: A GFM alert marker ([!NOTE] etc.) at the start of a blockquote's first
|
||||
#: paragraph: syntax, not prose — stripped from the first segment.
|
||||
_ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
|
||||
|
||||
#: Any {...} span: {placeholders} and attrs that ended up inside prose
|
||||
#: (inline attrs are consumed by the parser; a lone {dates} is not).
|
||||
_BRACES = re.compile(r"\{[^{}\n]*\}")
|
||||
|
||||
|
||||
def _runs(children: list) -> list[str]:
|
||||
"""Prose runs of an inline token's children, in order.
|
||||
|
||||
Text tokens merge across soft breaks into one run; every markup token
|
||||
(emphasis, links, code, images, HTML, footnote refs, hard breaks) is a
|
||||
run boundary. Link and image *text* is prose; autolink text (the URL
|
||||
itself) is not. Image tokens contribute their alt-text children and
|
||||
their title attribute.
|
||||
"""
|
||||
runs: list[str] = []
|
||||
cur: list[str] = []
|
||||
|
||||
def flush() -> None:
|
||||
if cur:
|
||||
s = "".join(cur)
|
||||
cur.clear()
|
||||
if _LETTER.search(s):
|
||||
runs.append(s)
|
||||
|
||||
skip = 0 # inside an autolink (its text is the URL — not prose)
|
||||
for t in children:
|
||||
if skip:
|
||||
if t.type == "link_close":
|
||||
skip -= 1
|
||||
continue
|
||||
if t.type == "text":
|
||||
cur.append(t.content)
|
||||
elif t.type == "softbreak":
|
||||
cur.append("\n")
|
||||
elif t.type == "link_open" and t.markup == "autolink":
|
||||
flush()
|
||||
skip = 1
|
||||
elif t.type == "image":
|
||||
flush()
|
||||
if t.children:
|
||||
runs.extend(_runs(t.children))
|
||||
title = t.attrGet("title")
|
||||
if title and _LETTER.search(title):
|
||||
runs.append(title)
|
||||
else:
|
||||
flush()
|
||||
if t.children:
|
||||
runs.extend(_runs(t.children))
|
||||
flush()
|
||||
return runs
|
||||
|
||||
|
||||
def _block_text(children: list) -> str:
|
||||
"""The block's text as a reader sees it: text runs and link texts
|
||||
merged (softbreaks as newlines); image alts, autolink URLs, code and
|
||||
other markup content excluded. Used as the translation CONTEXT for
|
||||
segments carved out of the block (link texts, partial runs): a lone
|
||||
word translates differently than the same word inside its sentence."""
|
||||
parts: list[str] = []
|
||||
skip = 0 # inside an autolink (its text is the URL)
|
||||
for t in children:
|
||||
if skip:
|
||||
if t.type == "link_close":
|
||||
skip -= 1
|
||||
continue
|
||||
if t.type == "text":
|
||||
parts.append(t.content)
|
||||
elif t.type == "softbreak":
|
||||
parts.append("\n")
|
||||
elif t.type == "link_open" and t.markup == "autolink":
|
||||
skip = 1
|
||||
elif t.type == "image":
|
||||
continue
|
||||
elif t.children:
|
||||
parts.append(_block_text(t.children))
|
||||
return "".join(parts)
|
||||
|
||||
|
||||
def _locate(source: str, needle: str, cursor: int) -> int:
|
||||
"""The needle's offset in source at/after cursor, -1 when absent.
|
||||
|
||||
An occurrence preceded by a backslash is an escaped character, not the
|
||||
token's source: keep looking (failing that, the run is skipped — it
|
||||
stays in the original language).
|
||||
"""
|
||||
pos = source.find(needle, cursor)
|
||||
while pos > 0 and source[pos - 1] == "\\":
|
||||
pos = source.find(needle, pos + 1)
|
||||
return pos
|
||||
|
||||
|
||||
def split(text: str) -> tuple[list[tuple[int, int]], list[str], list[str]]:
|
||||
"""Split a fragment into (spans, segments, contexts): prose segments to
|
||||
translate, their byte offsets in ``text`` for splicing the translations
|
||||
back, and per-segment translation context.
|
||||
|
||||
Segments containing {...} spans are carved further — the braces stay
|
||||
out of the wire text. A run that cannot be located verbatim in the
|
||||
source contributes no segment. A segment's context is its block's plain
|
||||
text when the segment was carved OUT of a larger block (a link text, a
|
||||
partial run); a segment that IS the whole block (a plain paragraph, a
|
||||
heading) is self-contextualizing and gets "".
|
||||
"""
|
||||
spans: list[tuple[int, int]] = []
|
||||
segments: list[str] = []
|
||||
contexts: list[str] = []
|
||||
cursor = 0
|
||||
blockquote_fresh = 0 # blockquote depth whose first inline is upcoming
|
||||
|
||||
def emit(run: str, at: int, ctx: str) -> None:
|
||||
"""Carve {...} spans out of the located run; emit the prose pieces,
|
||||
stripped — padding whitespace stays in the template, off the wire.
|
||||
Pieces containing "<" are never emitted: translators cut output at
|
||||
the first "<" (the prose/markup boundary, scripts/translator.py),
|
||||
so such a piece could not survive the round trip — it stays in the
|
||||
original language instead."""
|
||||
pieces = []
|
||||
pos = 0
|
||||
for m in _BRACES.finditer(run):
|
||||
pieces.append((pos, m.start()))
|
||||
pos = m.end()
|
||||
pieces.append((pos, len(run)))
|
||||
for p0, p1 in pieces:
|
||||
raw = run[p0:p1]
|
||||
piece = raw.strip()
|
||||
if _LETTER.search(piece) and "<" not in piece:
|
||||
start = at + p0 + (len(raw) - len(raw.lstrip()))
|
||||
spans.append((start, start + len(piece)))
|
||||
segments.append(piece)
|
||||
contexts.append(ctx)
|
||||
|
||||
tokens = _MD.parse(text)
|
||||
for t in tokens:
|
||||
if t.type == "blockquote_open":
|
||||
blockquote_fresh += 1
|
||||
elif t.type == "blockquote_close":
|
||||
blockquote_fresh -= 1
|
||||
elif t.type == "inline":
|
||||
kids = t.children or []
|
||||
runs = _runs(kids)
|
||||
block = _block_text(kids).strip()
|
||||
if blockquote_fresh:
|
||||
# An alert marker ([!NOTE]) leading the blockquote's first
|
||||
# paragraph is syntax; strip it from the segment. (Only the
|
||||
# first inline of the blockquote can carry it — the flag
|
||||
# clears on the first inline seen.)
|
||||
blockquote_fresh = 0
|
||||
if runs:
|
||||
run = _ALERT.sub("", runs[0], count=1)
|
||||
if _LETTER.search(run):
|
||||
runs[0] = run
|
||||
else:
|
||||
runs.pop(0)
|
||||
for run in runs:
|
||||
ctx = block if block and run.strip() != block else ""
|
||||
pos = _locate(text, run, cursor)
|
||||
if pos != -1:
|
||||
emit(run, pos, ctx)
|
||||
cursor = pos + len(run)
|
||||
elif "\n" in run:
|
||||
# Indented continuation lines etc. break the verbatim
|
||||
# match: locate each line separately instead.
|
||||
for part in run.split("\n"):
|
||||
if not _LETTER.search(part):
|
||||
continue
|
||||
pos = _locate(text, part, cursor)
|
||||
if pos != -1:
|
||||
emit(part, pos, ctx)
|
||||
cursor = pos + len(part)
|
||||
return spans, segments, contexts
|
||||
|
||||
|
||||
def pure_prose(text: str) -> bool:
|
||||
"""True when the text parses as nothing but prose (text and softbreak
|
||||
tokens) — the acceptance test for a translated segment: the model may
|
||||
not return markup of its own (a `<br>` here would splice live HTML into
|
||||
the fragment)."""
|
||||
children = _MD.parseInline(text)[0].children or []
|
||||
return all(t.type in ("text", "softbreak") for t in children)
|
||||
|
||||
|
||||
def join(original: str, spans: list[tuple[int, int]], texts: list[str]) -> str | None:
|
||||
"""Splice translated segments back into the original fragment; None on
|
||||
any validation failure (count mismatch, empty or non-prose segment) —
|
||||
the caller drops the result and the fragment stays pending."""
|
||||
if len(texts) != len(spans):
|
||||
return None
|
||||
out: list[str] = []
|
||||
cursor = 0
|
||||
for (start, end), translation in zip(spans, texts):
|
||||
if not translation.strip() or not pure_prose(translation):
|
||||
return None
|
||||
out.append(original[cursor:start])
|
||||
out.append(translation)
|
||||
cursor = end
|
||||
out.append(original[cursor:])
|
||||
return "".join(out)
|
||||
|
||||
|
||||
def has_prose(text: str) -> bool:
|
||||
"""True when the fragment yields at least one translatable segment.
|
||||
Chunks that are all markup, code, placeholders or reference definitions
|
||||
have no business reaching the model: every language renders them from
|
||||
the original chunk."""
|
||||
return bool(split(text)[1])
|
||||
+239
-17
@@ -1,28 +1,35 @@
|
||||
"""Translator service protocol and its transport-independent core.
|
||||
"""Translator service protocol, dispatcher and its transport-independent core.
|
||||
|
||||
The external machine-translation service connects over WebSocket
|
||||
(``/_translate/<key>``, see app.py) and exchanges JSON frames decoded into
|
||||
the tagged msgspec structs below (``bytes`` fields ride as base64 — no
|
||||
manual encoding anywhere). This module holds the message structs plus the
|
||||
shared computations around the WS handler: which fragments are pending for
|
||||
a language (``pending_items``), storing a result (``store_results``) and
|
||||
the startup URL listing (``log_service_urls``).
|
||||
The dispatcher itself (one job at a time per connection, wanted ∩ capable
|
||||
language matching, requeue on disconnect) lives in app.py.
|
||||
(``/_translate/<key>``, the route itself is in app.py) and exchanges JSON
|
||||
frames decoded into the tagged msgspec structs below (``bytes`` fields ride
|
||||
as base64 — no manual encoding anywhere). This module holds everything
|
||||
else: the message structs, the connected-client dispatcher (``Dispatcher``
|
||||
— one job at a time per connection, wanted ∩ capable language matching,
|
||||
requeue on disconnect), which fragments are pending for a language
|
||||
(``pending_items``), storing a result (``store_results``) and the startup
|
||||
URL listing (``log_service_urls``).
|
||||
|
||||
Fragments cross the wire masked: non-translatable spans (code, URLs,
|
||||
{placeholders}, tags, ...) are numbered ⟦N⟧ sentinels in ``Job.text``,
|
||||
restored and validated before storage (``pagerite/masking.py``).
|
||||
Fragments cross the wire as **prose segments**: the model only ever
|
||||
receives plain text runs (Job.texts) plus per-segment context surrounds
|
||||
(Job.contexts) and returns their translations (Result.texts, same order);
|
||||
markup never leaves the server — reassembly is offset splicing
|
||||
(``pagerite/segments.py``).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import os
|
||||
|
||||
import msgspec
|
||||
from fastapi import WebSocket, WebSocketDisconnect
|
||||
from kanta import Kanta
|
||||
|
||||
from pagerite import i18n
|
||||
from pagerite.__main__ import DEFAULT_PORT
|
||||
from pagerite.chunks import chunk_key, needs_translation
|
||||
from pagerite.data import Data, Node, sorted_nodes
|
||||
from pagerite.segments import join, split
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -42,6 +49,9 @@ class TransItem(msgspec.Struct):
|
||||
text: str
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Title jobs only: the article's opening prose, so the model sees the
|
||||
#: title as a heading in context, not a lone sentence.
|
||||
context: str = ""
|
||||
|
||||
|
||||
class Job(msgspec.Struct, tag="job"):
|
||||
@@ -53,9 +63,20 @@ class Job(msgspec.Struct, tag="job"):
|
||||
|
||||
lang: str
|
||||
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
|
||||
text: str #: masked Markdown (pagerite/masking.py)
|
||||
#: The fragment's prose segments (pagerite/segments.py): plain text
|
||||
#: runs only — no markup, URLs, code or placeholders ever cross the
|
||||
#: wire. Translate each element independently.
|
||||
texts: list[str]
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Per segment (parallel to texts; "" = none): the surround to
|
||||
#: translate it in — a carved-out segment (link text, partial run)
|
||||
#: carries its block's plain text, a title the article's opening.
|
||||
#: Reference client behavior (scripts/translator.py): translate
|
||||
#: segment+context together, keep the segment's part (its own line /
|
||||
#: paragraph); fall back to the segment alone when the output holds no
|
||||
#: separator. Contexts are not part of the result.
|
||||
contexts: list[str] = msgspec.field(default_factory=list)
|
||||
|
||||
|
||||
class TransResult(msgspec.Struct):
|
||||
@@ -71,7 +92,9 @@ class Result(msgspec.Struct, tag="result"):
|
||||
|
||||
lang: str
|
||||
key: bytes
|
||||
text: str
|
||||
#: The job's segments, translated, same order and count. Each must be
|
||||
#: pure prose — the server rejects the result otherwise.
|
||||
texts: list[str]
|
||||
|
||||
|
||||
#: Union of the client -> server frames (the "type" tag selects).
|
||||
@@ -90,18 +113,29 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
|
||||
items: list[TransItem] = []
|
||||
seen: set[bytes] = set()
|
||||
|
||||
def emit(key: bytes, text: str, path: str, kind: str) -> None:
|
||||
def emit(key: bytes, text: str, path: str, kind: str, context: str = "") -> None:
|
||||
if key in seen or lang in data.trans.get(key, {}):
|
||||
return
|
||||
seen.add(key)
|
||||
items.append(TransItem(key=key, text=text, path=path, kind=kind))
|
||||
items.append(TransItem(key=key, text=text, path=path, kind=kind, context=context))
|
||||
|
||||
def opening(node: Node) -> str:
|
||||
"""The article's opening prose (first segment, capped): the title
|
||||
job's context — a lone word like "About" reads as a heading on top
|
||||
of an article, not as a sentence. Empty when there's no prose."""
|
||||
for h in node.chunks or ():
|
||||
text = data.chunks.get(h)
|
||||
if text and (segs := split(text)[1]):
|
||||
return segs[0][:400]
|
||||
return ""
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str) -> None:
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
if node.chunks is not None:
|
||||
if node.title:
|
||||
emit(chunk_key(node.title), node.title, path, "title")
|
||||
emit(chunk_key(node.title), node.title, path, "title",
|
||||
context=opening(node))
|
||||
for h in node.chunks:
|
||||
text = data.chunks.get(h)
|
||||
if (
|
||||
@@ -160,3 +194,191 @@ def log_service_urls(keys: dict[str, str], hostname: str) -> None:
|
||||
base = f"ws://localhost:{port}" if hostname == "localhost" else f"wss://{hostname}"
|
||||
urls = ", ".join(f"{base}/_translate/{key} ({name})" for key, name in keys.items())
|
||||
logger.info("Translator %s", urls)
|
||||
|
||||
|
||||
class _Connection:
|
||||
"""One connected translator socket: the language codes it announced as
|
||||
capabilities (Hello) and the (lang, chunk-key) job currently in flight
|
||||
on it, with the segment spans to splice its Result into
|
||||
(pagerite/segments.py) — one at a time, the next is sent only after its
|
||||
Result.
|
||||
|
||||
Per-connection only: in-flight lives solely here, so on disconnect the
|
||||
item simply becomes pending again and is re-offered to any free capable
|
||||
connection."""
|
||||
|
||||
def __init__(self, capable: set[str]) -> None:
|
||||
self.capable = capable
|
||||
self.inflight: tuple[str, bytes] | None = None
|
||||
#: Source spans of the in-flight job's segments (splice offsets).
|
||||
self.spans: list[tuple[int, int]] = []
|
||||
self.original: str = "" # its full source text (for the splicing)
|
||||
|
||||
|
||||
class Dispatcher:
|
||||
"""The translator dispatcher: connected client sockets and the job
|
||||
pipeline (docs/localization.md).
|
||||
|
||||
One single-item job at a time per connection, offered in the
|
||||
intersection of the wanted languages (``Data.translate_langs``) and the
|
||||
connection's announced capabilities. Pending work is derived from the
|
||||
``trans`` store (``pending_items``) minus the items in flight on any
|
||||
connection, so a dropped connection's in-flight item is simply
|
||||
re-offered. Results are matched to content by chunk key alone. A
|
||||
(lang, key) whose Result fails segment validation is skipped for the
|
||||
rest of the run — generation is near-deterministic, so an immediate
|
||||
retry would just re-fail.
|
||||
"""
|
||||
|
||||
def __init__(self, data: Data, db: Kanta, invalidate) -> None:
|
||||
self.data = data
|
||||
self.db = db
|
||||
#: Sync content-change hook (app._invalidate_pages), called inside
|
||||
#: transactions; schedules the next dispatch pass.
|
||||
self.invalidate = invalidate
|
||||
#: Connected translator sockets and their per-connection state.
|
||||
self.clients: dict[WebSocket, _Connection] = {}
|
||||
#: (lang, chunk key) of fragments whose result failed validation
|
||||
#: (segment count, empty or non-prose segments, segments.py) this run.
|
||||
self.validation_failures: set[tuple[str, bytes]] = set()
|
||||
|
||||
def schedule(self) -> None:
|
||||
"""Schedule a dispatch pass, if any translator is connected.
|
||||
|
||||
The invalidate hook is sync and called inside transactions: the
|
||||
task first runs once the current coroutine awaits again, i.e. after
|
||||
the transaction has committed. No-op without a running loop (CLI
|
||||
use)."""
|
||||
if not self.clients:
|
||||
return
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return
|
||||
asyncio.create_task(self._dispatch())
|
||||
|
||||
async def _dispatch(self) -> None:
|
||||
"""Offer one pending item to every free capable connection."""
|
||||
wanted = {
|
||||
tag
|
||||
for lang in self.data.translate_langs
|
||||
if (tag := i18n.translation_tag(lang))
|
||||
}
|
||||
if not wanted:
|
||||
return
|
||||
for ws, state in list(self.clients.items()):
|
||||
if state.inflight is not None:
|
||||
continue
|
||||
langs = wanted & state.capable
|
||||
if not langs:
|
||||
continue
|
||||
inflight = {s.inflight for s in self.clients.values() if s.inflight}
|
||||
job = None
|
||||
spans: list[tuple[int, int]] = []
|
||||
original = ""
|
||||
for lang in sorted(langs):
|
||||
for item in pending_items(self.data, lang):
|
||||
if (lang, item.key) in inflight or (lang, item.key) in self.validation_failures:
|
||||
continue
|
||||
spans, texts, contexts = split(item.text)
|
||||
if not texts:
|
||||
continue # prose that could not be located for splicing
|
||||
original = item.text
|
||||
if item.kind == "title" and item.context:
|
||||
# A title's surround is the article's opening prose
|
||||
# (TransItem.context), not its own one-word block.
|
||||
contexts = [item.context] * len(texts)
|
||||
job = Job(
|
||||
lang=lang, key=item.key, texts=texts,
|
||||
path=item.path, kind=item.kind, contexts=contexts,
|
||||
)
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is None:
|
||||
continue
|
||||
state.inflight = (job.lang, job.key) # before the await: no double-assign
|
||||
state.spans = spans
|
||||
state.original = original
|
||||
try:
|
||||
await ws.send_text(msgspec.json.encode(job).decode())
|
||||
except Exception: # send failed: the receive loop cleans up
|
||||
self.clients.pop(ws, None)
|
||||
|
||||
async def handle_ws(self, ws: WebSocket, clientkey: str) -> None:
|
||||
"""The /_translate/<key> channel (docs/localization.md).
|
||||
|
||||
A wrong/empty key rejects the handshake (closing before accept
|
||||
makes Starlette answer HTTP 403). Protocol (JSON frames): the
|
||||
client opens with Hello(langs) announcing its CAPABILITIES — the
|
||||
language codes its model can produce (normalized to translation
|
||||
tags; "en"/empty dropped) — then answers each Job with its
|
||||
Result(lang, key, texts). A Result without an in-flight job or with
|
||||
a different (lang, key), a duplicate Hello, or any malformed frame
|
||||
closes the socket with a protocol error.
|
||||
"""
|
||||
if clientkey not in self.data.translate_keys:
|
||||
await ws.close(code=1008) # policy violation; pre-accept = HTTP 403
|
||||
return
|
||||
await ws.accept()
|
||||
state: _Connection | None = None
|
||||
try:
|
||||
while True:
|
||||
raw = await ws.receive_text()
|
||||
try:
|
||||
msg = msgspec.json.decode(raw.encode(), type=ClientMsg)
|
||||
except msgspec.DecodeError:
|
||||
await ws.close(code=1002) # protocol error
|
||||
return
|
||||
if isinstance(msg, Hello):
|
||||
if state is not None: # one Hello per connection
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
state = _Connection({
|
||||
tag for lang in msg.langs if (tag := i18n.translation_tag(lang))
|
||||
})
|
||||
self.clients[ws] = state
|
||||
self.schedule()
|
||||
else: # Result
|
||||
lang = i18n.translation_tag(msg.lang)
|
||||
if (
|
||||
state is None # results before Hello
|
||||
or state.inflight is None # no job in flight
|
||||
or (lang, msg.key) != state.inflight # wrong job
|
||||
):
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
texts, spans, original = msg.texts, state.spans, state.original
|
||||
state.inflight = None
|
||||
state.spans = []
|
||||
state.original = ""
|
||||
text = join(original, spans, texts) if len(texts) == len(spans) else None
|
||||
if text is None:
|
||||
# The model broke the segment contract (count
|
||||
# mismatch, empty or non-prose segment): drop the
|
||||
# result and skip the fragment for this run (it
|
||||
# stays pending; a restart, a refresh or a model
|
||||
# change gets another chance).
|
||||
self.validation_failures.add((lang, msg.key))
|
||||
logger.warning(
|
||||
"[%s] result for chunk %s rejected: invalid segments",
|
||||
lang, msg.key.hex(),
|
||||
)
|
||||
self.schedule()
|
||||
continue
|
||||
with self.db.transaction("translator results", user=clientkey, extra=lang):
|
||||
paths = store_results(
|
||||
self.data, lang, [TransResult(key=msg.key, text=text)]
|
||||
)
|
||||
self.invalidate() # schedules the next dispatch
|
||||
if paths:
|
||||
logger.info(
|
||||
"[%s] now available for %d page(s): %s",
|
||||
lang, len(paths), ", ".join(sorted(paths)),
|
||||
)
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
finally:
|
||||
if self.clients.pop(ws, None) is not None:
|
||||
# The in-flight item (if any) is pending again; offer it around.
|
||||
self.schedule()
|
||||
|
||||
+2
-2
@@ -956,8 +956,8 @@ def render_page(
|
||||
canonical = url if lang == i18n.ORIGINAL_LANGUAGE else f"{url}?lang={lang}"
|
||||
if data.translate_langs:
|
||||
alternates = [("x-default", url)] + [
|
||||
(l, f"{url}?lang={l}")
|
||||
for l in [i18n.ORIGINAL_LANGUAGE, *sorted(data.translate_langs)]
|
||||
(tag, f"{url}?lang={tag}")
|
||||
for tag in [i18n.ORIGINAL_LANGUAGE, *sorted(data.translate_langs)]
|
||||
]
|
||||
return str(
|
||||
_layout(
|
||||
|
||||
+120
-16
@@ -56,10 +56,59 @@ SEED_X_TAGS = {
|
||||
}
|
||||
SEED_X_NAMES = {v: k for k, v in SEED_X_TAGS.items()}
|
||||
|
||||
#: The fragments are masked Markdown (pagerite/masking.py: ⟦N⟧ sentinels
|
||||
#: stand in for code, URLs, placeholders...); Seed-X has no system prompt,
|
||||
#: so the instruction goes in-line.
|
||||
NOTE = ", preserving all Markdown formatting and keeping every ⟦N⟧ token exactly unchanged"
|
||||
#: The fragments arrive as prose segments (pagerite/segments.py): plain
|
||||
#: text runs only — no markup, URLs, code or placeholders. The wire
|
||||
#: invariant is that segments are PURE PROSE, and one character marks the
|
||||
#: boundary both ways: "<" never appears in a segment. Sources containing
|
||||
#: it are never dispatched (pagerite/segments.py keeps them in the
|
||||
#: original language); the model's output is cut at the first "<" — one
|
||||
#: rule that covers the whole class of markup bleed (an echoed <lang> tag,
|
||||
#: a "<br>", ...) instead of a pattern per artifact. (Generation-level
|
||||
#: stop strings can't do this job: the model's <s> framing token would
|
||||
#: trip a "<" stop at the first token; skip_special_tokens strips the
|
||||
#: framing at decode.)
|
||||
#:
|
||||
#: Two kinds cross the wire (Job.kind), each with its own prompt template:
|
||||
#: titles get told they ARE titles (a lone word otherwise invites
|
||||
#: context-free readings — "About" as "approximately"). Any segment may
|
||||
#: carry its surround in Job.contexts (a title: the article's opening; a
|
||||
#: carved-out segment like a link text: its block's plain text) and is then
|
||||
#: translated together with that surround (seed_x_chunk). No punctuation
|
||||
#: clause, on purpose: Seed-X handles trailing-punctuation instructions by
|
||||
#: slipping into its [COT] reasoning mode (observed for Chinese:
|
||||
#: minutes-long generations, reasoning text in the output) —
|
||||
#: match_punctuation handles stray punctuation deterministically instead.
|
||||
PROMPTS = {
|
||||
"chunk": "Translate the following {source_lang} text into {target_lang}:\n{text} <{tag}>",
|
||||
"title": "Translate the following {source_lang} title into {target_lang}:\n{text} <{tag}>",
|
||||
# Title with the article's opening as context (Job.contexts): the model
|
||||
# translates both; generation stops at the blank line separating them,
|
||||
# and the segment's own part of the output is the translation. No
|
||||
# separator in the output (the model merged them) → seed_x_chunk falls
|
||||
# back to the plain kind template.
|
||||
"title+context": "Translate the following {source_lang} title and the beginning of its article "
|
||||
"into {target_lang}:\n{text}\n\n{context} <{tag}>",
|
||||
# A segment carved out of a larger block (link text, partial run) with
|
||||
# its sentence as context — same mechanics as title+context.
|
||||
"chunk+context": "Translate the following {source_lang} text into {target_lang}:\n"
|
||||
"{text}\n\n{context} <{tag}>",
|
||||
}
|
||||
TERMINAL_PUNCT = ".,!?:;…。,!?;:、"
|
||||
|
||||
|
||||
def match_punctuation(source: str, translated: str) -> str:
|
||||
"""Drop terminal punctuation the model added.
|
||||
|
||||
When the source segment ends without terminal punctuation, the
|
||||
translation must not gain any either. A leading Spanish ¡/¿ only pairs
|
||||
with a terminal !/?, so it goes with it.
|
||||
"""
|
||||
if not source or source[-1] in TERMINAL_PUNCT:
|
||||
return translated
|
||||
trimmed = translated.rstrip(TERMINAL_PUNCT)
|
||||
if trimmed and trimmed[0] in "¡¿":
|
||||
trimmed = trimmed[1:].lstrip()
|
||||
return trimmed
|
||||
|
||||
|
||||
# The wire structs below duplicate pagerite/translate.py: this script runs
|
||||
@@ -79,9 +128,15 @@ class Job(msgspec.Struct, tag="job"):
|
||||
|
||||
lang: str
|
||||
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
|
||||
text: str #: masked (pagerite/masking.py): the ⟦N⟧ tokens must survive verbatim
|
||||
#: The fragment's prose segments: plain text runs only, no markup —
|
||||
#: translate each element independently (pagerite/segments.py).
|
||||
texts: list[str]
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Per segment (parallel to texts; "" = none): the surround to
|
||||
#: translate it in — a link text carries its sentence, a title the
|
||||
#: article's opening. See seed_x_chunk for how they are used.
|
||||
contexts: list[str] = msgspec.field(default_factory=list)
|
||||
|
||||
|
||||
class Result(msgspec.Struct, tag="result"):
|
||||
@@ -90,7 +145,7 @@ class Result(msgspec.Struct, tag="result"):
|
||||
|
||||
lang: str
|
||||
key: bytes
|
||||
text: str
|
||||
texts: list[str] #: the job's segments translated, same order and count
|
||||
|
||||
|
||||
def load_seed_x():
|
||||
@@ -101,29 +156,78 @@ def load_seed_x():
|
||||
return tokenizer, model
|
||||
|
||||
|
||||
def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str, source_lang: str = "English",
|
||||
note: str = ""):
|
||||
"""Translate one segment; returns (translation, output_tokens, generation_seconds)."""
|
||||
def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str,
|
||||
kind: str = "chunk", context: str = "", source_lang: str = "English"):
|
||||
"""Translate one segment; returns (translation, output_tokens, generation_seconds).
|
||||
|
||||
With context, the segment is translated together with its surround (a
|
||||
link text with its sentence, a title with the article's opening), and
|
||||
the segment's own part of the output is the translation: its own line
|
||||
for a single-line source (a single-line segment's translation never
|
||||
contains a line break — generation stops at the blank line separating
|
||||
the two), its own paragraph for a multi-line one (softbreak-merged
|
||||
lines keep single newlines, the separator is the blank line). If the
|
||||
model merged them — no separator, or an empty first part — fall back to
|
||||
translating the segment alone; the wasted tokens are counted either
|
||||
way.
|
||||
"""
|
||||
# No chat template on this model; the trailing language tag is required (trans/ style prompt).
|
||||
prompt = f"Translate the following {source_lang} text into {target_lang}{note}:\n{text} <{tag}>"
|
||||
template = PROMPTS.get(f"{kind}+context" if context else kind, PROMPTS["chunk"])
|
||||
prompt = template.format(source_lang=source_lang, target_lang=target_lang,
|
||||
text=text, tag=tag, context=context)
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||
t0 = time.monotonic()
|
||||
out = model.generate(**inputs, max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]), do_sample=False)
|
||||
# The only stop string is the context separator. "<" must NOT be one:
|
||||
# stopping works on the raw output, which always starts with the
|
||||
# model's <s> framing token. skip_special_tokens strips <s>/</s> at
|
||||
# decode; the post-decode cut at the first "<" then enforces the wire
|
||||
# invariant (prose only) against markup bleed.
|
||||
kwargs = {"stop_strings": ["\n\n"], "tokenizer": tokenizer} if context else {}
|
||||
out = model.generate(**inputs, max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]),
|
||||
do_sample=False, **kwargs)
|
||||
dt = time.monotonic() - t0
|
||||
n = out.shape[1] - inputs.input_ids.shape[1]
|
||||
return tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True).strip(), n, dt
|
||||
decoded = tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
|
||||
translated = decoded.partition("<")[0]
|
||||
if not context:
|
||||
return translated.strip(), n, dt
|
||||
if "\n" in text:
|
||||
# Multi-line segment: its translation keeps single newlines; the
|
||||
# blank line is the separator from the context translation.
|
||||
sep = "\n\n" in translated
|
||||
out = translated.split("\n\n", 1)[0] if sep else ""
|
||||
else:
|
||||
out, sep, _ = translated.partition("\n")
|
||||
if not sep:
|
||||
out = ""
|
||||
out = out.strip()
|
||||
if out:
|
||||
return out, n, dt
|
||||
# The model merged segment and context (no separator, or an empty first
|
||||
# part): retry without the context.
|
||||
again, n2, dt2 = seed_x_chunk(tokenizer, model, text, target_lang, tag,
|
||||
kind=kind, source_lang=source_lang)
|
||||
return again, n + n2, dt + dt2
|
||||
|
||||
|
||||
async def do_job(ws, job: Job, tokenizer, model) -> None:
|
||||
"""Translate the job's one fragment and send the result back."""
|
||||
"""Translate the job's segments (one model call each) and send them back."""
|
||||
lang_name = SEED_X_NAMES[job.lang].capitalize()
|
||||
# Deliberately blocking: nothing else needs the loop while the job is
|
||||
# being answered, and the reconnect loop recovers a dropped connection
|
||||
# (the in-flight item is simply re-offered).
|
||||
text, tokens, dt = seed_x_chunk(tokenizer, model, job.text, lang_name, job.lang, note=NOTE)
|
||||
print(f"[{job.lang} {job.kind} {job.path or '/'}: "
|
||||
texts = []
|
||||
tokens = dt = 0
|
||||
for i, text in enumerate(job.texts):
|
||||
ctx = job.contexts[i] if i < len(job.contexts) else ""
|
||||
translated, n, t = seed_x_chunk(tokenizer, model, text, lang_name, job.lang,
|
||||
kind=job.kind, context=ctx)
|
||||
texts.append(match_punctuation(text, translated))
|
||||
tokens += n
|
||||
dt += t
|
||||
print(f"[{job.lang} {job.kind} {job.path or '/'}: {len(texts)} segments, "
|
||||
f"{tokens} tokens in {dt:.1f}s = {tokens / dt:.1f} tok/s]", file=sys.stderr)
|
||||
await ws.send(msgspec.json.encode(Result(lang=job.lang, key=job.key, text=text)).decode())
|
||||
await ws.send(msgspec.json.encode(Result(lang=job.lang, key=job.key, texts=texts)).decode())
|
||||
|
||||
|
||||
async def serve(url: str, tokenizer, model) -> None:
|
||||
|
||||
Reference in New Issue
Block a user