Compare commits

...
6 Commits
Author SHA1 Message Date
LeoVasanko b57b7060ec Keep container fence lines out of prose chunks
A closing ::: glued to a paragraph (no blank line before it) rode inside
the prose chunk and crossed to the translator as part of the text run;
when the model dropped it, validation passed and the splice lost the
fence — the rest of the page rendered inside the container (seen in the
Spanish translation). Container fence lines (::: openers and closers
alike) are now always their own prose-free chunk, never reaching the
translator. Affected pages re-chunk on next save and re-translate under
the new hashes, repairing themselves.
2026-09-04 19:04:24 +00:00
LeoVasanko 3f27a0a292 Fix unstyled editor language selector, order all language menus logically
LangSelect's scoped CSS landed on the shared store chunk, whose stylesheet
the editor never loaded (only the public selector path injected it), so the
editor's language selector rendered unstyled on untranslated pages. Collect
editor stylesheets from the entry's imported chunks too (same traversal as
the langselect assets).

Also: hreflang alternates now skip languages disabled site-wide, and all
selectors (page editor, structure tab, public selector) order languages the
same way — primary first, then the lang tab's geographic grouping.
2026-09-04 18:52:22 +00:00
LeoVasanko b30d909a23 Don't extract GeoIP .mmdb.gz on filesystem, only in RAM. 2026-09-04 18:30:55 +00:00
LeoVasanko 13fecd2118 Hreflang alternates on category placeholder pages too 2026-09-04 18:12:48 +00:00
LeoVasanko 11a138e19f Translate category-label titles, not just page titles 2026-09-04 18:09:40 +00:00
LeoVasanko ebd5911a38 Store DBIP database in folder where the program is ran, not where it is installed. 2026-09-04 18:04:30 +00:00
11 changed files with 172 additions and 86 deletions
+4 -3
View File
@@ -73,10 +73,11 @@ Each `Client` record (shared by every event, keyed by hash):
A reverse-DNS lookup is attempted for each new client and the result, when
available, is stored as `host`; local/reserved/multicast addresses are
skipped. If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present
in the repository root, it is loaded at startup and used to look up
in the working directory, it is loaded at startup and used to look up
`country`/`city`. These lookups run in background tasks after the event is
stored, so WebSocket message handling is never delayed. The decompressed
`dbip-*.mmdb` file is kept in the repository root and ignored by git. The
stored, so WebSocket message handling is never delayed. Only the downloaded
`.mmdb.gz` is kept on disk (in the working directory, ignored by git); it is
decompressed into RAM when opened. The
CLI flag `--dbip` (`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP at startup (in the app lifespan,
before the MMDB is opened), skipping the download when the local database is
+8 -2
View File
@@ -106,7 +106,9 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
language like content pages, but over the **subtree's** combined
availability (`subtree_languages`) — they have no chunks of their own;
the heading, navigation and card text localize from the title map and
the target articles' translations.
the target articles' translations. Their hreflang alternates are
computed exactly like a content page's (a translated title counts as
availability, so the language selector is offered there too).
- Card descriptions and cover picks run on the target article's hybrid
Markdown where that page is available in the served language, with
per-card fallback to the original.
@@ -144,7 +146,11 @@ served Markdown at render time.
`chunk_markdown(markdown)` splits the source into block-level chunks —
blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
list blocks, tables, HTML blocks. A chunk's identity is its **source text**,
list blocks, tables, HTML blocks. Container fence lines (`::: name` openers
and `:::` closers) are always their own chunk, blank lines or not — folded
into a prose chunk the closer would cross to the translator as part of the
text, where the model can drop it (the rest of the page then renders inside
the container). A chunk's identity is its **source text**,
gettext-msgid style:
```python
+16 -11
View File
@@ -9,23 +9,28 @@
// store and re-renders it).
import { computed } from 'vue'
import LangSelect from './LangSelect.vue'
import { flagFor, langName } from './langs'
import { flagFor, langName, langSort } from './langs'
import { useStore } from './store'
const store = useStore()
// The "(primary)" marker is admin-panel information; the public selector
// lists plain languages.
const options = computed(() =>
store.langAlternates.map((a) => ({
tag: a.tag,
code: a.tag,
name: langName(a.tag),
flag: flagFor(a.tag),
primary: false,
})),
)
// lists plain languages. Order: the primary language first, then the rest
// in the lang tab's geographic grouping (./langs langSort) — the head's
// hreflang order is just alphabetical.
const primaryTag = computed(() => store.langAlternates.find((a) => a.primary)?.tag ?? '')
const options = computed(() => {
const rest = langSort(
store.langAlternates.map((a) => a.tag).filter((t) => t !== primaryTag.value),
)
return [primaryTag.value, ...rest].filter(Boolean).map((tag) => ({
tag,
code: tag,
name: langName(tag),
flag: flagFor(tag),
primary: false,
}))
})
// The explicit pick, else the served language (header-autodetected pages
// may have neither), else the primary.
const model = computed(() => store.lang || store.servedLang || primaryTag.value)
+7 -5
View File
@@ -33,7 +33,7 @@ import { keymap } from '@codemirror/view'
import { indentWithTab } from '@codemirror/commands'
import { markdown } from '@codemirror/lang-markdown'
import { cmHighlight, cmTheme } from './cmtheme'
import { flagFor, langName } from './langs'
import { flagFor, langName, langSort } from './langs'
import { editorLang, pagePrimary } from './editorLang'
import LangSelect from './LangSelect.vue'
import ConnNote from './ConnNote.vue'
@@ -121,11 +121,13 @@ function normPath(p) {
// localization settings tab).
// The picker's options: the primary language first, then the union of the
// page's translations and the site-wide configured targets, sorted.
// page's translations and the site-wide configured targets in the lang
// tab's geographic grouping (./langs langSort).
const langOptions = computed(() => {
const others = [...new Set([...siteLangs.value, ...pageLangs.value])]
.filter((l) => l && l !== primaryLang.value)
.sort()
const others = langSort(
[...new Set([...siteLangs.value, ...pageLangs.value])]
.filter((l) => l && l !== primaryLang.value),
)
return [primaryLang.value, ...others].map((code) => ({
tag: code === primaryLang.value ? '' : code,
code,
+5 -4
View File
@@ -19,7 +19,7 @@ import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, wa
import StructureTree from './StructureTree.vue'
import LangSelect from './LangSelect.vue'
import { slugify } from './slugify'
import { flagFor, langName } from './langs'
import { flagFor, langName, langSort } from './langs'
import { editorLang, pagePrimary } from './editorLang'
import { dropPageCache, loadPlain } from './swapdoc'
@@ -40,9 +40,10 @@ const primaryLang = ref('en')
const siteLangs = ref([])
// The strip's options: the primary language first, then the configured
// translation targets (the lang tab manages that set).
// translation targets (the lang tab manages that set) in the lang tab's
// geographic grouping (./langs langSort).
const langOptions = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
[primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
.map((code) => ({
tag: code === primaryLang.value ? '' : code,
code,
@@ -63,7 +64,7 @@ watch(lang, () => refreshPages())
// dropdown lists "inherit" first (naming what it resolves to), then every
// site language. Setting it on a section covers its whole subtree.
const rowLangChoices = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
[primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
.map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })),
)
function rowLangOptions(el) {
+14
View File
@@ -30,6 +30,20 @@ export const LANG_GROUPS = [
const displayNames = new Intl.DisplayNames(['en'], { type: 'language' })
// Consistent menu ordering for language selectors: the geographic/cultural
// grouping above (similar languages sit together, and it does not vary with
// the display language the way alphabetical-by-name would). Tags outside
// the groups trail, ordered by tag. The primary language is not special
// here — callers put it first themselves.
const groupOrder = new Map(LANG_GROUPS.flat().map((c, i) => [c, i]))
export function langSort(codes) {
return [...codes].sort(
(a, b) =>
(groupOrder.get(a) ?? groupOrder.size) - (groupOrder.get(b) ?? groupOrder.size)
|| a.localeCompare(b),
)
}
// English display name for a language tag ("fi" -> "Finnish").
export function langName(tag) {
try {
+19 -3
View File
@@ -17,6 +17,13 @@ from pagerite.segments import has_prose
#: backticks or tildes (CommonMark).
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
#: A container fence line (mdit-py-plugins container): the "::: aside"
#: opener and the ":::" closer alike. Always its own block, even with no
#: blank line around it: folded into a prose paragraph it would cross to
#: the translator as part of the text run, where the model can drop it —
#: the rest of the page then renders inside the container.
_CONTAINER = re.compile(r"^ {0,3}:{3,}(?:[ \t]|$)")
#: HTML block openers that may span blank lines (CommonMark types 1-5:
#: script/pre/style/textarea, comments, processing instructions,
#: declarations, CDATA) with their closing condition. Other HTML blocks
@@ -54,9 +61,11 @@ def chunk_markdown(markdown: str) -> list[str]:
Blocks are separated by blank lines; fenced code blocks and the
multi-line HTML blocks (comments, script/pre/style, CDATA...) are
kept atomic, even across blank lines, and end at their closing
condition. Chunks carry no surrounding blank lines and no trailing
newline; rejoining with ``join_chunks`` reproduces the source modulo
blank-line normalization.
condition. Container fence lines (:::, open and close alike) are
always their own block, blank lines or not (see _CONTAINER). Chunks
carry no surrounding blank lines and no trailing newline; rejoining
with ``join_chunks`` reproduces the source modulo blank-line
normalization.
"""
chunks: list[str] = []
buf: list[str] = []
@@ -91,6 +100,13 @@ def chunk_markdown(markdown: str) -> list[str]:
fence = m.group(1)
buf.append(line)
continue
if _CONTAINER.match(line):
# Container fence lines (open and close alike) are their own
# block — never part of a prose chunk (see _CONTAINER).
flush()
buf.append(line)
flush()
continue
if not buf:
for open_re, close_re in _HTML_ATOMIC:
if open_re.match(line):
+1
View File
@@ -132,6 +132,7 @@ def _render_html(
data.theme,
data.favicon,
data.brand_html,
base_url,
transition=data.transition,
lang=lang,
translation=translation,
+26 -25
View File
@@ -3,18 +3,19 @@
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
enriched in background tasks with reverse DNS (cached PTR lookups) and the
DB-IP city MMDB (``GeoIP``, decompressed and opened once at startup);
DB-IP city MMDB (``GeoIP``, decompressed into RAM and opened once at
startup);
external referrers get their favicon fetched and stored content-hashed.
Snapshot broadcasts to connected admin sockets are debounced.
"""
import asyncio
import gzip
import io
import ipaddress
import logging
import os
import re
import shutil
import socket
from datetime import date
from functools import lru_cache
@@ -44,8 +45,9 @@ _analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None
# Repository root from this file's location (pagerite/tracking.py -> ..).
_REPO_ROOT = Path(__file__).resolve().parent.parent
# DB-IP databases persist in the working directory (one download serves all
# sites run from it). Not the package directory: reinstalls/upgrades wipe it.
_DBIP_DIR = Path.cwd()
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
@@ -60,7 +62,7 @@ def _download_dbip() -> None:
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
for p in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
logger.info("DB-IP database is current (%s), skipping download", existing[-1])
@@ -68,7 +70,7 @@ def _download_dbip() -> None:
for month in months:
url = DBIP_URL.format(month=month)
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
target = _DBIP_DIR / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
logger.info("Downloading %s", url)
try:
@@ -93,7 +95,7 @@ def _download_dbip() -> None:
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
for old in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
logger.info("DB-IP database updated to %s", target.name)
@@ -102,15 +104,19 @@ def _download_dbip() -> None:
def _geoip_db_path() -> Path | None:
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present.
"""Find a DB-IP MMDB in the working directory: the ``.mmdb.gz`` download
is canonical (decompressed into RAM at open); a plain ``.mmdb`` left over
from older versions is still usable, and removed once the matching ``.gz``
is present so it does not linger on disk. Returns None if none is present.
"""
mmdb = sorted(_REPO_ROOT.glob("dbip-*.mmdb"))
gz = sorted(_DBIP_DIR.glob("dbip-*.mmdb.gz"))
if gz:
for stale in _DBIP_DIR.glob("dbip-*.mmdb"):
stale.unlink()
return gz[0]
mmdb = sorted(_DBIP_DIR.glob("dbip-*.mmdb"))
if mmdb:
return mmdb[0]
gz = sorted(_REPO_ROOT.glob("dbip-*.mmdb.gz"))
if gz:
return gz[0]
return None
@@ -123,27 +129,22 @@ class GeoIP:
def __init__(self) -> None:
self._reader: object | None = None
def _decompress(self, source: Path, target: Path) -> None:
if target.exists():
return
tmp = target.with_suffix(target.suffix + ".tmp")
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
shutil.copyfileobj(src, dst)
os.replace(tmp, target)
def _load(self) -> None:
if self._reader is not None:
return
source = _geoip_db_path()
if source is None:
return
if source.suffix == ".gz":
target = source.with_suffix("")
self._decompress(source, target)
source = target
try:
import maxminddb
if source.suffix == ".gz":
# Only the .gz is kept on disk; the database is decompressed
# into RAM (MODE_FD makes the pure-Python Reader .read() the
# buffer — never mmap — and bypasses the C extension).
buf = io.BytesIO(gzip.decompress(source.read_bytes()))
self._reader = maxminddb.open_database(buf, maxminddb.MODE_FD)
else:
self._reader = maxminddb.open_database(str(source))
except Exception:
pass
+8 -5
View File
@@ -101,7 +101,8 @@ ClientMsg = Hello | Result
def pending_items(data: Data, lang: str) -> list[TransItem]:
"""Fragments of the site still untranslated for ``lang``, deduped by key.
Every page node (published or not) contributes its title and each chunk
Every node (published or not, pages and pure category labels alike)
contributes its title; pages also contribute each chunk
that needs translation (``needs_translation``), is not editor-flagged
no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang``
yet. Content-addressed text (shared paragraphs, repeated titles) appears
@@ -133,8 +134,10 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
path = f"{prefix}/{slug}" if prefix else slug
# An article whose primary language IS the target needs no
# translation into it — skip its title and chunks entirely.
# Category labels (chunks is None) contribute only their title:
# it is their nav-menu label.
node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang:
if node_lang != lang:
if node.title:
emit(
chunk_key(node.title),
@@ -143,7 +146,7 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
"title",
context=opening(node),
)
for h in node.chunks:
for h in node.chunks or ():
text = data.chunks.get(h)
if (
text is not None
@@ -178,8 +181,8 @@ def store_results(data: Data, lang: str, items: list[TransResult]) -> list[str]:
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang:
keys = set(node.chunks)
if node_lang != lang:
keys = set(node.chunks or ())
if node.title:
keys.add(chunk_key(node.title))
if keys & stored:
+63 -27
View File
@@ -263,20 +263,31 @@ def _transition_css_url(transition: str) -> str | None:
def _editor_css_url(vite_url: str | None) -> str | None:
"""URL for the editor-specific stylesheet (Vue component styles).
"""URLs (comma-joined) for the editor-specific stylesheets (Vue
component styles).
This is linked by the public-page edit pen so the editor styles are
loaded before the editor JS dynamic-import resolves.
loaded before the editor JS dynamic-import resolves. Component styles
can land on shared chunks rather than the entry's own stylesheet —
LangSelect's ride on the shared store chunk, as it is also used by the
on-demand public language selector — so collect the stylesheets of the
entry and its imported chunks (the same traversal _langselect_assets
does).
"""
if vite_url:
return None
manifest = _manifest()
entry = manifest["src/main.js"]
base = manifest.get(_BASE_CSS_KEY, {}).get("file")
for css in entry.get("css", []):
if css != base:
return f"/{css}"
return None
stylesheets, seen = [], set()
queue = ["src/main.js"]
for key in queue: # grows with imported chunks
if key in seen:
continue
seen.add(key)
entry = manifest[key]
stylesheets += [f"/{css}" for css in entry.get("css", []) if css != base]
queue += entry.get("imports", [])
return ",".join(stylesheets) or None
def _inline_asset(url: str) -> str:
@@ -1021,6 +1032,42 @@ def _social_meta(
}
def _language_urls(
data: Data,
path: str,
node: Node,
lang: str,
original: str,
base_url: str,
) -> tuple[str, list[tuple[str, str]]]:
"""(canonical, hreflang alternates) for a page (docs/localization.md).
The canonical names the actually served language — the plain URL for
the original (for SEO the non-query URL means the article's language),
?lang= for a translation — regardless of how the language was arrived
at (query or header). The alternates list the languages the page is
actually available in (``node.langs``; a category label's title counts
as its content): x-default first (the plain, autodetecting URL), then
every available language — the original again by its plain URL,
translations by ?lang=. The public language selector keys off these.
("", []) without a base_url.
"""
if not base_url:
return "", []
url = f"{base_url}/{path}"
canonical = url if lang == original else f"{url}?lang={lang}"
alternates = []
if data.translate_langs:
# Only languages the page actually has AND that are still enabled
# site-wide (a disabled target stops being advertised).
enabled = {original, *data.translate_langs}
alternates = [("x-default", url)] + [
(tag, url if tag == original else f"{url}?lang={tag}")
for tag in sorted({original, *node.langs} & enabled)
]
return canonical, alternates
def render_page(
menu: dict[str, Node],
data: Data,
@@ -1050,24 +1097,7 @@ def render_page(
title = _title(path.rpartition("/")[2], node, translation, path)
main = page_content(menu, data, path, translation, link_lang, lang)
social = _social_meta(node, path, title, str(main), brand, base_url)
# Canonical/hreflang URLs (docs/localization.md): the canonical names
# the actually served language — the plain URL for the original (for
# SEO the non-query URL means the article's language), ?lang= for a
# translation — regardless of how the language was arrived at (query
# or header). The alternates list the languages the page is actually
# available in: x-default first (the plain, autodetecting URL), then
# every available language — the original again by its plain URL,
# translations by ?lang=. The public language selector keys off these.
canonical = ""
alternates = []
if base_url:
url = f"{base_url}/{path}"
canonical = url if lang == original else f"{url}?lang={lang}"
if data.translate_langs:
alternates = [("x-default", url)] + [
(tag, url if tag == original else f"{url}?lang={tag}")
for tag in sorted({original, *node.langs})
]
canonical, alternates = _language_urls(data, path, node, lang, original, base_url)
return str(
_layout(
*_page_assets(),
@@ -1100,6 +1130,7 @@ def render_category(
theme: str = "",
favicon: str = "",
brand_html: str = "",
base_url: str = "",
transition: str = "cube",
lang: str = i18n.ORIGINAL_LANGUAGE,
translation: Translation | None = None,
@@ -1115,12 +1146,16 @@ def render_category(
With a translation (titles only — the category has no Markdown) the
heading, navigation and card text localize per target article
(docs/localization.md); ``link_lang`` replicates the ?lang= override
onto the navigation links as on content pages.
onto the navigation links as on content pages. The hreflang alternates
are computed as on content pages — a translated title makes the
language available here too.
"""
node = resolve(menu, path)[-1]
original = i18n.primary_lang(menu, path)
if translation is None:
lang = i18n.primary_lang(menu, path)
lang = original
title = _title(path.rpartition("/")[2], node, translation, path)
_, alternates = _language_urls(data, path, node, lang, original, base_url)
doc = E.article
with doc:
doc.h1(title)
@@ -1137,6 +1172,7 @@ def render_category(
transition,
favicon,
lang=lang,
alternates=alternates,
)(
Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html, link_lang),