Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3a4745389e | ||
|
|
4eea3ae3af | ||
|
|
69583fb9fe | ||
|
|
b57b7060ec | ||
|
|
3f27a0a292 | ||
|
|
b30d909a23 | ||
|
|
13fecd2118 | ||
|
|
11a138e19f | ||
|
|
ebd5911a38 | ||
|
|
db57125953 | ||
|
|
986e28c220 | ||
|
|
33a4a76364 | ||
|
|
9fb4b5a681 | ||
|
|
0c1349b037 | ||
|
|
54f8c8e09b | ||
|
|
78f4ddb2f0 | ||
|
|
e6a0c57446 | ||
|
|
d6d07db2c5 | ||
|
|
38af57218a | ||
|
|
eb2e8f8273 | ||
|
|
792b9e7aa9 | ||
|
|
4c6ab3dde6 | ||
|
|
9d70f17587 | ||
|
|
029bfe105e | ||
|
|
62031fd5dd | ||
|
|
13cf716bb1 | ||
|
|
cbd50cfece | ||
|
|
85ef968cd6 | ||
|
|
7631c9f0a3 | ||
|
|
d7d03754d1 | ||
|
|
65fec6c519 | ||
|
|
dc55445ae0 | ||
|
|
b133ad6dd6 | ||
|
|
4413c7efdf | ||
|
|
2f533eaf09 | ||
|
|
1d46843c76 | ||
|
|
d3e2196c83 | ||
|
|
b6e6e46cfb | ||
|
|
7a4544731d |
@@ -10,19 +10,33 @@ Please instead ask the user to see from dev tools what you need, e.g. to look up
|
||||
Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes:
|
||||
|
||||
- `pagerite/` — Python backend package (hatchling build target).
|
||||
- `app.py` — FastAPI app and route registration.
|
||||
- `app.py` — thin FastAPI assembly: lifespan, `FastAPI(...)`, router includes (route ordering: api/tracking/files routers, then `frontend.route(app, "/")`, then the pages catch-all last).
|
||||
- `state.py` — shared core, no routes: env-derived site constants, `data`/`kanta`, `analytics_store`, the fastapi-vue `frontend`, the render cache (`_html_response`, `_invalidate_pages`), the translator `dispatcher`, slug helpers, `@kanta.bootstrap` hooks.
|
||||
- `files.py` — `FileStore` (content-addressed, RAM-cached), image derivative helpers (`store_image`), file routes (`/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, favicon settings).
|
||||
- `api.py` — editor REST + WS: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, `/_translate/{key}`.
|
||||
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetch, `/_ws`, `/_api/ws/analytics`, the `/_a` page (docs/analytics.md).
|
||||
- `pages.py` — public content pages: `/`, `/sitemap.xml`, `/robots.txt`, the `/{path:path}` catch-all.
|
||||
- `data.py` — msgspec Structs for the kanta database.
|
||||
- `chunks.py` — block-level Markdown chunking and content-hash keys for the chunk stores (docs/migrate.md).
|
||||
- `i18n.py` — language selection, translation assembly (chunks + patches) and translated-edit recording (user patches, per-language title overrides, refresh).
|
||||
- `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); api.py only registers the route.
|
||||
- `segments.py` — the translation round trip: fragments split into pure-prose wire segments (via markdown.make_md's verbatim parser; link- and formatting-carrying blocks stay whole, link/formatted texts inline, Markdown stripped) and translations spliced back by source offset, link/formatting markdown re-inserted at weight-mapped positions (docs/localization.md).
|
||||
- `migrations.py` — kanta migrations (`migrate_vN`); ALL schema/storage upgrades live here (raw state dict before struct decoding), never in the app lifespan: v1 moves legacy in-db file blobs to the on-disk store and rebuilds the legacy flat `pages` as the menu tree, v2 rewrites `/_f/{hash}.ext` image links to the extension-less form, backfills AVIF/WebP/JPEG derivatives on disk and drops the obsolete `version` field.
|
||||
- `markdown.py` — markdown-it-py renderer.
|
||||
- `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`).
|
||||
- `seed.py` — demo content, written only on first database creation.
|
||||
- `analytics.py` — visit analytics collection (see `docs/analytics.md`).
|
||||
- `analytics.py` — visit analytics collection (see `docs/analytics.md`). UA formatting/bot detection comes from the **uarite** package.
|
||||
- `frontend/src/` — Vue editor and public-page JS entries.
|
||||
- `main.js` — Vue editor app entry.
|
||||
- `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`).
|
||||
- `langselect-main.js` + `LangSelector.vue` — public language selector, imported on demand by pagerite.js on pages with more than one hreflang alternate (the editors' `LangSelect` flag dropdown).
|
||||
- `store.js` — the shared Pinia store (`useStore`, id `pagerite`) for cross-bundle UI state.
|
||||
- `pagerite.js` — public page entry.
|
||||
- `editorLang.js` + `LangSelect.vue` — the editor shell's shared language selection and its selector component (page + structure tabs; drives the page preview while the panel is open, via `swapdoc.setLangOverride`).
|
||||
- `reconnect.js` — shared WebSocket pacing for all sockets (staggered connect slots, stuck-CONNECTING watchdog, exponential backoff): bursts and rapid retries trip the browser's WebSocket throttling.
|
||||
- `assets/` — base CSS, Pygments styles, fonts.
|
||||
- `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test).
|
||||
- `scripts/translator.py` — Seed-X translator service client for the `/_translate/{key}` socket (reference client, runs in its own uv env via PEP 723); stays connected full time, unloads the model after 60 s idle and reloads on the next job.
|
||||
|
||||
Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed).
|
||||
|
||||
@@ -52,6 +66,6 @@ Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed).
|
||||
## Conventions
|
||||
|
||||
- Keep dependencies minimal; add via `uv add` and mention it.
|
||||
- The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
|
||||
- No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session.
|
||||
- The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` (backend redirect to the configured site icon). Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
|
||||
- No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. The one keyed exception is `/_translate/{key}` (translator service; `Data.translate_keys`, see docs/localization.md).
|
||||
- Update the relevant MarkDown files when architecture, tooling, or conventions change.
|
||||
|
||||
+211
-164
@@ -1,23 +1,104 @@
|
||||
# Analytics
|
||||
|
||||
Server-side visit analytics. Data lives in a plain JSON file — a msgspec
|
||||
Struct dumped to disk — separate from the kanta content database, path from
|
||||
`PAGERITE_ANALYTICS` (default: `analytics.json` in the per-site data
|
||||
directory, e.g. `localhost/analytics.json`).
|
||||
Server-side visit analytics built on a **raw access-log-style event store**.
|
||||
Data lives in a plain JSON file — a msgspec Struct dumped to disk — separate
|
||||
from the kanta content database, path from `PAGERITE_ANALYTICS` (default:
|
||||
`analytics.json` in the per-site data directory, e.g. `localhost/analytics.json`).
|
||||
|
||||
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
|
||||
`CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map,
|
||||
atomic JSON persistence).
|
||||
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
|
||||
the `/_ws` activity WebSocket, and `WebSocket /_api/ws/analytics`
|
||||
(admin-gated like every `/_api` endpoint).
|
||||
- `pagerite/analytics.py` — data model (`Analytics`, `Get`, `Msg`, `Client`,
|
||||
`Favicon`), the `Store` (raw log + atomic JSON persistence) and
|
||||
`Store.display()`, where **all** classification happens.
|
||||
- `pagerite/pages.py` — records every served document as one raw GET line
|
||||
(`_record_get`, in `pagerite/tracking.py`) with its true HTTP status.
|
||||
- `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
|
||||
`WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
|
||||
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
|
||||
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
|
||||
normal site layout on the `/_a` analytics page.
|
||||
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
|
||||
into `#analytics-app` inside `#main`.
|
||||
|
||||
## What is collected
|
||||
## Raw records
|
||||
|
||||
The store is deliberately close to an access log: two append-only lists plus
|
||||
shared metadata. **Nothing is classified when recorded** — whether a client
|
||||
turns out to be a reader, a crawler or a scanner is decided by
|
||||
`Store.display()` from the raw events, so the stored data survives any future
|
||||
change to the classification rules.
|
||||
|
||||
Each `Get` record (one per served document):
|
||||
|
||||
- `t` — timestamp of the request,
|
||||
- `path` — full request path, query string included (e.g. `/.env?x=1`),
|
||||
- `status` — the true HTTP status of the response (200, or 404 for a category
|
||||
placeholder or a missing page),
|
||||
- `ref` — external https origin of the `Referer`, `""` for direct/internal
|
||||
(same-origin referers are dropped by the recorder),
|
||||
- `pre` — true for idle-time link preloads from pagerite.js
|
||||
(`x-pagerite-preload` header): never counted as a view, crawler hit or
|
||||
abuse — recorded only so a navigation later served from the in-memory page
|
||||
cache (which issues no GET at all) can be attributed this GET's status,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`.
|
||||
|
||||
304 revalidation responses return before recording and are not logged.
|
||||
|
||||
Each `Msg` record (one per pagerite.js activity message over `/_ws`):
|
||||
|
||||
- `t` — timestamp,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `fr` — path of the page the activity happened on (`""` for the initial
|
||||
load),
|
||||
- `to` — navigation target (validated at record time: internal slug path or
|
||||
external https URL; anything else is dropped — sanitation, not
|
||||
classification),
|
||||
- `read` — active seconds spent on `fr` since the previous report.
|
||||
|
||||
Each `Client` record (shared by every event, keyed by hash):
|
||||
|
||||
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
|
||||
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
|
||||
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
|
||||
- `country` — two-letter country code. Initially derived from the
|
||||
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
|
||||
when a database is available,
|
||||
- `city` — city name from the DB-IP MMDB lookup, when available,
|
||||
- `ua` — raw `User-Agent` string,
|
||||
- `hide` — true for admin clients (`hide` message field): everything this
|
||||
client ever did is recorded but excluded from every statistic and from the
|
||||
viewer payload. This is the one flag set at record time — it is a client
|
||||
property, not a classification.
|
||||
|
||||
The viewer payload adds one display-time field to each client, never
|
||||
persisted (stored records keep the default and old data always follows the
|
||||
current uarite version):
|
||||
|
||||
- `uarite` — the `uarite.UA` dataclass from parsing the raw UA
|
||||
(`pretty`/`engine`/`os`/`provider`/`kind`/`url`): the crawler name for
|
||||
bots,
|
||||
with a category suffix only where a provider runs crawlers of more than
|
||||
one kind (`GPTBot (AI)` vs `OAI-SearchBot (search)`, `Googlebot (search)`
|
||||
vs `Google-Extended (AI)`; single-kind providers stay plain: `Facebook`,
|
||||
`WhatsApp`), `Browser/major OS` on the desktop, the device where that is
|
||||
the relevant information (iPhone reports its iOS version, Android phones
|
||||
their model instead of the OS), otherwise the raw string; `url` is the
|
||||
crawler's info page when uarite knows one (rendered as a 🔗 link after the
|
||||
pretty UA in the viewer), `kind` drives the bot classification.
|
||||
|
||||
A reverse-DNS lookup is attempted for each new client and the result, when
|
||||
available, is stored as `host`; local/reserved/multicast addresses are
|
||||
skipped. If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present
|
||||
in the working directory, it is loaded at startup and used to look up
|
||||
`country`/`city`. These lookups run in background tasks after the event is
|
||||
stored, so WebSocket message handling is never delayed. Only the downloaded
|
||||
`.mmdb.gz` is kept on disk (in the working directory, ignored by git); it is
|
||||
decompressed into RAM when opened. The
|
||||
CLI flag `--dbip` (`uv run pagerite --dbip`) downloads the latest
|
||||
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP at startup (in the app lifespan,
|
||||
before the MMDB is opened), skipping the download when the local database is
|
||||
already current and removing older versions after an update; without the flag
|
||||
only an existing file is used.
|
||||
|
||||
## What the client sends
|
||||
|
||||
The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
|
||||
whole browsing session and sends activity messages over it — JSON text
|
||||
@@ -26,36 +107,20 @@ frames matching the server's `Ping` msgspec struct with the fields `fr`
|
||||
since the last report) and `hide`; falsy fields are omitted. One channel
|
||||
follows the session, so the activity of a visit stays tied together, and
|
||||
while the user is active the accumulated reading time is flushed every few
|
||||
seconds: the trail times are cumulative, so a disconnection simply leaves
|
||||
the last reported time in place (no close beacon). After 5 minutes without
|
||||
seconds: the times are incremental, so a disconnection simply leaves the
|
||||
last reported time in place (no close beacon). After 5 minutes without
|
||||
any activity the client closes the socket itself — a sleeping browser tab
|
||||
would lose it anyway — and the next activity reconnects as a fresh session;
|
||||
reconnects are attempted only on user activity, with an exponential backoff
|
||||
between attempts so a failing endpoint is never hammered. Idle-time link preloads
|
||||
would lose it anyway — and the next activity reconnects; reconnects are
|
||||
attempted only on user activity, with an exponential backoff between
|
||||
attempts so a failing endpoint is never hammered. Idle-time link preloads
|
||||
stay plain `fetch()` calls so the browser may cache the responses; the
|
||||
WebSocket reports actual navigations and active time spent on a page.
|
||||
|
||||
- **Initial page load**: only `to` — the loaded path — is sent, never `fr`
|
||||
(an `fr` equal to `to` would log a bogus self-transition when a session
|
||||
already exists, e.g. a second tab). This message is what starts
|
||||
the visit and counts the entry page view — the document GET alone records
|
||||
nothing, so bots never register (admin browsing does register, but
|
||||
flagged `hide`; see **Admins** below). JS-running crawlers
|
||||
(Googlebot, GoogleOther, Applebot, ...) do connect and report, but their
|
||||
User-Agent gives them away: messages whose UA matches `_is_bot_ua`
|
||||
(anything calling
|
||||
itself a "bot", plus known exceptions such as GoogleOther) are ignored
|
||||
server-side, and their document GETs land in the crawler list instead.
|
||||
No source-IP verification is done: a spoofed bot UA merely lands in the
|
||||
crawler stats, and scanners that probe telltale paths are caught by the
|
||||
abuse rules regardless. Reloads are not
|
||||
already exists, e.g. a second tab). Reloads are not
|
||||
visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
|
||||
refresh neither counts a second view nor logs a self-transition. The GET
|
||||
handler stashes a cross-origin https `Referer` (origin part only —
|
||||
unavailable to JS once the page has loaded) and any
|
||||
`utm_*` query parameters in in-memory IP tables, consumed by the first
|
||||
message that
|
||||
starts the visit; internal or absent referers never touch the referer table.
|
||||
refresh neither counts a second view nor logs a self-transition.
|
||||
- **Internal fetch-navigations**: `to` is the target path, sent only after
|
||||
the swap actually happened (a failed swap falls back to a full load,
|
||||
whose initial message counts the view instead — no gap, no double count).
|
||||
@@ -67,28 +132,19 @@ WebSocket reports actual navigations and active time spent on a page.
|
||||
- **Excluded**: back/forward (popstate) navigations, navigating *to* the
|
||||
analytics page (`/_a` — its GET is untracked, and the server cannot
|
||||
record it as a navigation target anyway), and everything while the user has
|
||||
the editor
|
||||
open (`body.editing`). Admin noise, not visits. Navigating *away* from
|
||||
`/_a` does report: the fetch-navigation already GET-ed the target page
|
||||
without the preload header, and without the message that GET would flush to
|
||||
the crawler list.
|
||||
the editor open (`body.editing`). Admin noise, not visits. Navigating
|
||||
*away* from `/_a` does report.
|
||||
- **Admins**: when SSO is in use and the session is known to be an admin,
|
||||
the client still reports but adds `hide`. The activity is recorded as
|
||||
usual (navigations and all), but the `hide` flag is set on the **client
|
||||
record** — so it covers everything that client ever did: visits and
|
||||
crawler hits from before the login included. Hidden clients never appear
|
||||
in the viewer payload: `Store.display()` drops their visits, crawler
|
||||
hits, abuse hits and metadata, and computes every aggregate (site visits,
|
||||
page views, transitions) from the visible visits only, so nothing needs
|
||||
to be reversed or redacted. Pending crawler hits from a hidden client
|
||||
are discarded when they expire, so admin browsing never lands in the
|
||||
crawler list either. With no auth proxy
|
||||
record** — so it covers everything that client ever did, including the
|
||||
time before the login. Hidden clients never appear in the viewer payload:
|
||||
`Store.display()` drops their events and metadata, and computes every
|
||||
aggregate (site visits, page views, transitions) from the visible visits
|
||||
only, so nothing needs to be reversed or redacted. With no auth proxy
|
||||
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
|
||||
is recorded.
|
||||
- The server validates `to`: internal paths must be valid slug paths
|
||||
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
|
||||
https origin and accepted only when the client sent exactly that.
|
||||
- **External-site favicons**: for every external https origin seen as a visit
|
||||
- **External-site favicons**: for every external https origin seen as a GET
|
||||
referer or an exit link, the server fetches `{origin}/favicon.ico` in a
|
||||
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
|
||||
SVG is sniffed from the body when served without an image type) and stores
|
||||
@@ -96,112 +152,104 @@ WebSocket reports actual navigations and active time spent on a page.
|
||||
extension matching the actual MIME). The origin → file name mapping is
|
||||
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
|
||||
recorded too and retried only after 7 days. Fetches are scheduled after
|
||||
each activity message and once at startup, which backfills icons for already-recorded
|
||||
data. The viewer payload carries `favicons` (origin → `/_f/...` path),
|
||||
and the viewer shows the icon wherever an external site is mentioned:
|
||||
referer/exit trail links in the visit table and the source/exit pills of
|
||||
the transition map (UTM-attributed source nodes without an https origin
|
||||
stay text-only).
|
||||
- **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
|
||||
`User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
|
||||
the first 6 bytes identify a shared `Client` record. The `Client` stores
|
||||
the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
|
||||
`country` from the language-region subtag, and asynchronously-filled
|
||||
`country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
|
||||
crawler hits and abuse hits all reference this record by its hash, so
|
||||
client metadata is stored once instead of repeated per event.
|
||||
- The visitor IP is stored in the `Client`. A reverse-DNS lookup is
|
||||
attempted for each new client and the result, when available, is stored as
|
||||
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
|
||||
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
|
||||
root, it is loaded at startup and used to look up `country`/`city`. These
|
||||
lookups run in background tasks after the event is stored, so WebSocket
|
||||
message handling is never delayed. The decompressed `dbip-*.mmdb` file is kept in
|
||||
the repository root and ignored by git. The CLI flag `--dbip`
|
||||
(`uv run pagerite --dbip`) downloads the latest
|
||||
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
|
||||
skipping the download when the local database is already current and
|
||||
removing older versions after an update; without the flag only an existing
|
||||
file is used.
|
||||
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
|
||||
hit — except idle-time link preloads from pagerite.js, which carry an
|
||||
`x-pagerite-preload` header and are not tracked at all (the navigation
|
||||
message sent when the user actually navigates to a preloaded page does
|
||||
the counting; forging
|
||||
the header only hides a GET from the crawler stats, the path-based abuse
|
||||
classification is unaffected). If a message
|
||||
from the same client arrives within 10 seconds the hit is discarded;
|
||||
otherwise it is written to `crawlers` — unless the client is hidden
|
||||
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as
|
||||
visits or views. The `Accept-Language` header is stored on the shared
|
||||
`Client` immediately; reverse-DNS host names and DB-IP geoip
|
||||
country/city are filled in asynchronously, just like for real visits. In
|
||||
the analytics viewer, crawler hits are grouped by client hash and shown as
|
||||
a trail of internal pages that crawler visited; the crawler table lists
|
||||
the most recent crawler first, with the most active as a tie-breaker.
|
||||
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
|
||||
starting with a dot (`/.env`, `/.git/config`) or ending in `.php` —
|
||||
classifies the source IP as abuse immediately, and ten plain 404s from one
|
||||
IP do too. Classification reclassifies history: all earlier crawler hits
|
||||
from that IP (persisted and pending) move to the `abuse` list, so a
|
||||
random-UA scanner no longer pollutes the crawler stats of the legitimate
|
||||
bot it impersonates. Once classified, every document GET and 404 from the
|
||||
IP is recorded as an abuse hit with the full request path (query string
|
||||
included), and its activity messages are ignored. The classified IP set (`abuse_ips`)
|
||||
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
|
||||
viewer, abuse hits are grouped by IP (never by client/UA — scanners
|
||||
randomize theirs) in a separate "Abuse" table. Identical paths are
|
||||
collapsed into one entry with their hit count; flagged paths that
|
||||
triggered classification are lifted to the top, followed by other 404s and
|
||||
then document GETs from the abuser. Raw User-Agent strings are shown one
|
||||
per line with their occurrence counts, and the full lists are click-to-copy.
|
||||
each activity message and once at startup, which backfills icons for
|
||||
already-recorded data. The viewer payload carries `favicons` (origin →
|
||||
`/_f/...` path), and the viewer shows the icon wherever an external site
|
||||
is mentioned: referer/exit trail links in the visit table and the
|
||||
source/exit pills of the transition map (UTM-attributed source nodes
|
||||
without an https origin stay text-only).
|
||||
|
||||
## Visits and sessions
|
||||
## Display-time classification
|
||||
|
||||
There are no cookies. A visit is tied together by a client hash — the first
|
||||
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
|
||||
/64 network), the raw `User-Agent` string and the extracted
|
||||
`Accept-Language` tag. The first message from a client hash starts a new
|
||||
visit; subsequent messages extend it. Messages arriving with no known session
|
||||
(server restart) start a fresh visit from the first message — treated as
|
||||
missing data rather than dropped. The client-hash → visit map and the IP →
|
||||
entry-referer/UTM tables are in-memory only; client metadata is stored in
|
||||
`Analytics.clients` keyed by the client hash.
|
||||
`Store.display(in_menu)` derives the viewer payload from the raw events on
|
||||
every (debounced) broadcast — O(n log n) over the log, cheap enough for a
|
||||
small CMS. `in_menu(path)` resolves a path against the current menu (passed
|
||||
in from `tracking.py`, which owns the content database import) so 404
|
||||
responses for real menu nodes — category placeholders — are not mistaken
|
||||
for misses.
|
||||
|
||||
Each `Client` record:
|
||||
- **Visits and sessions**: a client's messages are grouped into visits
|
||||
chronologically; a new visit starts after 30 minutes of inactivity
|
||||
(`_SESSION_GAP`). A fresh page load with an already-open visit (second
|
||||
tab) extends it, logging a `(direct)` transition. The visit's trail holds
|
||||
first-seen targets in order; `read` updates accumulate active seconds on
|
||||
the trail item matching `fr`. Each trail item's HTTP status comes from
|
||||
the client's latest GET for that path — preloads included, which is what
|
||||
allows 404 pages to render red in the viewer even when the navigation
|
||||
itself was served from the page cache. The entry page's referer and
|
||||
`utm_*` tags come from the GET that loaded it (within 10 s before the
|
||||
first message).
|
||||
- **Crawler hits**: a document GET no activity message matched within
|
||||
`_CRAWLER_TIMEOUT` (10 s) is a crawler hit — plain bots that only fetch
|
||||
documents never register as visits. JS-running crawlers (Googlebot,
|
||||
GoogleOther, Applebot, ...) do connect and send messages, but their UA
|
||||
gives them away (`_is_bot_ua`, backed by `uarite.uaparse` — which
|
||||
also knows the disguised ones: facebookexternalhit, Google-Extended,
|
||||
WhatsApp, ...): their messages are ignored at display
|
||||
time, so their GETs never match and land in the crawler list too. Real-
|
||||
browser bots whose UA does not match are caught by engagement: a visit
|
||||
whose total reported reading time is under 5 seconds (`_MIN_VISIT_READ`;
|
||||
durations are client-provided and trusted — such bots report 0–2 s) is
|
||||
reclassified as crawler hits, one per internal trail page, and counts in
|
||||
no visit aggregate. No source-IP verification is done: a spoofed bot UA
|
||||
merely lands in the crawler stats, and scanners that probe telltale paths
|
||||
are caught by the abuse rules regardless. In the viewer, crawler hits are
|
||||
grouped by client hash and shown as a trail of pages, preceded by the
|
||||
referer when there is one (rendered with its favicon like visit
|
||||
referers). The crawler table lists the most recent crawler first, with
|
||||
the most active as a tie-breaker.
|
||||
- **Abuse (scanner) hits**: a 404 on a telltale path — an empty URL segment
|
||||
(`//foo` — no real client generates those), any segment starting with a
|
||||
dot (`/.env`, `/.git/config`) or ending in `.php` — classifies the source
|
||||
IP as abuse, and ten plain 404s within one hour (`_ABUSE_404_WINDOW`) on
|
||||
paths that don't resolve to a menu node do too. Two exemptions keep
|
||||
legitimate traffic out: RFC 8615 well-known URIs (`/.well-known/…` —
|
||||
browsers and services probe them, e.g. Chrome's devtools fetch of
|
||||
`appspecific/com.chrome.devtools.json`) are never telltale and never
|
||||
count toward the threshold, and category placeholders return 404 but are
|
||||
real menu nodes, so they never count either. The window keeps a
|
||||
long-time reader's slowly accumulating misses from ever crossing the
|
||||
threshold — scanners spray in bursts. Hidden (admin) clients never
|
||||
trigger classification: editing means visiting not-found pages, since
|
||||
that is where the create pen lives. Once an IP is classified, **all** its document GETs are shown in the abuse list —
|
||||
including any that arrived before classification, since the raw log keeps
|
||||
everything — and its activity messages are ignored. In the viewer, abuse
|
||||
hits are grouped by IP (never by client/UA — scanners randomize theirs)
|
||||
in a separate "Abuse" table, split by the recorded status: the 404 probes
|
||||
("paths abused" — flagged paths that triggered classification first, then
|
||||
other 404s, shown verbatim with query strings) versus the real articles
|
||||
the abuser actually read ("articles read" — the 200 document GETs,
|
||||
rendered as trail links like the visitor and crawler tables, query string
|
||||
stripped). Raw User-Agent strings are shown one per line with their
|
||||
occurrence counts, and the full lists are click-to-copy.
|
||||
|
||||
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
|
||||
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
|
||||
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
|
||||
- `country` — two-letter country code. Initially derived from the
|
||||
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
|
||||
when a database is available,
|
||||
- `city` — city name from the DB-IP MMDB lookup, when available,
|
||||
- `ua` — raw `User-Agent` string,
|
||||
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
|
||||
parsable, otherwise the raw string,
|
||||
- `hide` — true for admin clients (`hide` message field): all their visits,
|
||||
crawler hits and abuse hits are recorded but excluded from every
|
||||
statistic and from the viewer payload.
|
||||
In the visitor and crawler tables, internal paths that returned a 404 status
|
||||
are shown in red and the link title includes the status code, so it is easy
|
||||
to tell misses from real pages at a glance.
|
||||
|
||||
Each `Visit` record:
|
||||
## Derived shapes (the viewer payload)
|
||||
|
||||
- `start` — timestamp of the first event,
|
||||
The `Display` payload contains the derived `visits`, `crawlers` and `abuse`
|
||||
rows (structs `Visit`/`Nav`/`TrailItem`, `CrawlerHit`, `AbuseHit` — display
|
||||
DTOs only, never persisted), the visible `clients`, the fetched `favicons`,
|
||||
and the aggregates below.
|
||||
|
||||
Each derived `Visit`:
|
||||
|
||||
- `start` — timestamp of the first activity,
|
||||
- `entry` — first page (path) seen,
|
||||
- `referer` — external https origin of the initial load, `""` for direct,
|
||||
- `referer` — external https origin of the entry GET, `""` for direct,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `trail` — the entry page and everything seen afterwards, keyed by the
|
||||
timestamp of first sight (insertion order = first-seen order). Each item
|
||||
holds `to` (page path or external exit URL), the accumulated active
|
||||
reading time in seconds (`read`) and the most recent HTTP status seen
|
||||
for the target (`status`). Re-visiting an already seen target updates
|
||||
its item instead of appending.
|
||||
- `navs` — every navigation message (`fr`, `to`), keyed by its timestamp,
|
||||
repeats included. The aggregates are computed from this log at display
|
||||
time.
|
||||
for the target (`status`),
|
||||
- `navs` — every navigation (`fr`, `to`), keyed by its timestamp, repeats
|
||||
included. The aggregates are computed from this log,
|
||||
- `utm` — `utm_*` query parameters from the landing URL, as a dict.
|
||||
|
||||
Each `CrawlerHit` record:
|
||||
Each derived `CrawlerHit`:
|
||||
|
||||
- `start` — timestamp of the document GET,
|
||||
- `entry` — page path requested,
|
||||
@@ -211,35 +259,31 @@ Each `CrawlerHit` record:
|
||||
- `status` — HTTP status of the served response (200 for a real page, 404
|
||||
for a category placeholder or missing page).
|
||||
|
||||
Each `AbuseHit` record:
|
||||
Each derived `AbuseHit`:
|
||||
|
||||
- `start` — timestamp of the request,
|
||||
- `path` — full request path including the query string (e.g. `/.env?x=1`),
|
||||
- `path` — full request path including the query string,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `flag` — true for the path that triggered abuse classification (telltale
|
||||
path or the 404 that crossed the threshold),
|
||||
- `is_404` — true for 404 responses, false for document GETs from the
|
||||
abuser.
|
||||
- `flag` — true for the paths that triggered abuse classification (telltale
|
||||
paths, or the 404 that crossed the threshold),
|
||||
- `is_404` — true for 404 responses, false for real (200) document GETs.
|
||||
|
||||
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
|
||||
are grouped by IP alone (resolved from the referenced `Client`). In the
|
||||
Abuse table identical paths are collapsed with their counts; flagged paths
|
||||
that triggered classification are lifted to the top, followed by other 404s
|
||||
and then document GETs from the abuser. Within each category paths are
|
||||
sorted by count descending, then by their earliest hit.
|
||||
|
||||
In the visitor and crawler tables, internal paths that returned a 404 status
|
||||
are shown in red and the link title includes the status code, so it is easy
|
||||
to tell misses from real pages at a glance.
|
||||
are grouped by IP alone (resolved from the referenced `Client`). In the
|
||||
Abuse table identical requests (same path and status class) are collapsed
|
||||
with their counts — a path's 404 probes and its later 200 reads never
|
||||
merge. Within each list paths are sorted by count descending, then by their
|
||||
earliest hit.
|
||||
|
||||
## Aggregates
|
||||
|
||||
Aggregates are **not stored**; they are computed at display time by
|
||||
`Store.display()` from the visit records (entry + `navs` log), skipping
|
||||
hidden clients' visits. This is what allows a client to become hidden after
|
||||
navigations were already logged: no counts need reversing. The computed
|
||||
shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
|
||||
`crawlers`, `abuse` and `clients`):
|
||||
`Store.display()` from the derived visits (entry + `navs` log), skipping
|
||||
hidden clients and short visits reclassified as crawler hits. This is
|
||||
what allows a client to become hidden after navigations were already
|
||||
logged: no counts need reversing. The computed shapes, part of the
|
||||
WebSocket payload (`Display` struct alongside `visits`, `crawlers`, `abuse`
|
||||
and `clients`):
|
||||
|
||||
- `transitions`: time series of page transitions, sparse nested dict
|
||||
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
|
||||
@@ -253,13 +297,16 @@ shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
|
||||
5-minute bucketing.
|
||||
|
||||
Sparseness keeps quiet sites small; dropping old data is a matter of deleting
|
||||
list entries (`visits` is a plain append-only list).
|
||||
list entries (`gets`/`msgs` are plain append-only lists).
|
||||
|
||||
## Persistence
|
||||
|
||||
The whole `Analytics` struct is JSON-encoded and written atomically
|
||||
(temp file + rename) on every recorded event. Traffic on a small CMS makes
|
||||
this cheap enough; batching can be added later without changing the format.
|
||||
A file written by the pre-redesign schema (stored `visits`/`crawlers`/`abuse`
|
||||
lists) is not convertible; it is renamed to `analytics.json.bak-legacy` and
|
||||
recording starts fresh.
|
||||
|
||||
## Viewing
|
||||
|
||||
|
||||
+12
-4
@@ -4,13 +4,21 @@ The Python backend lives in `pagerite/`.
|
||||
|
||||
## `app.py`
|
||||
|
||||
The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything defined earlier wins. The one exception is the content catch-all `/{path:path}`, registered AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
|
||||
Thin FastAPI assembly: lifespan (open the kanta database, load the file store, the frontend build and GeoIP), the `FastAPI(...)` instance with built-in API docs disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content, the `server` header middleware, and router includes. The routes themselves live in specialized modules:
|
||||
|
||||
- `state.py` — shared core, no routes: the environment-derived site constants (`HOSTNAME`, `SITE_URL`, `DB_PATH`, `FILES_DIR`, image/favicon tunables), the `data` root and its `kanta` handle (`Kanta(..., migrations="pagerite.migrations")`), the `analytics_store`, the fastapi-vue `frontend`, the page render cache and `_html_response`, the translator `dispatcher`, the slug charset helpers, and the `@kanta.bootstrap` hooks (demo seed, translator defaults).
|
||||
- `files.py` — the `FileStore` and image derivative helpers, and the file routes: `/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, the favicon settings endpoints.
|
||||
- `api.py` — the editor REST API and WebSockets: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, and the translator channel `/_translate/{clientkey}`.
|
||||
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetching, debounced broadcasts, the `/_ws` activity socket, the admin stream `/_api/ws/analytics`, and the `/_a` viewer page.
|
||||
- `pages.py` — the public content pages: `/`, `/sitemap.xml`, `/robots.txt` and the `/{path:path}` catch-all.
|
||||
|
||||
Route ordering is load-bearing and lives in `app.py`: the api/tracking/files routers are included BEFORE `frontend.route(app, "/")` is called — fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything registered earlier wins. The content catch-all `/{path:path}` is included AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
|
||||
|
||||
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
|
||||
|
||||
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding — and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
|
||||
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `state.py`'s `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding — and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
|
||||
|
||||
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore`: content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
|
||||
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore` (in `files.py`): content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
|
||||
|
||||
## `data.py`
|
||||
|
||||
@@ -34,4 +42,4 @@ Any page with published children — a category page — lists them as a card gr
|
||||
|
||||
## `seed.py`
|
||||
|
||||
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `app.py`.
|
||||
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `state.py`.
|
||||
|
||||
@@ -10,11 +10,11 @@ The site structure is stored in the kanta database managed by `pagerite/data.py`
|
||||
|
||||
Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `pages` (pre-tree databases) migrates into `menu` via `migrate_v1`. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`.
|
||||
|
||||
Every content/settings write calls `_invalidate_pages()` in app.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
|
||||
Every content/settings write calls `_invalidate_pages()` in state.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
|
||||
|
||||
## Files
|
||||
|
||||
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in app.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
|
||||
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in files.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
|
||||
|
||||
## Banners
|
||||
|
||||
|
||||
@@ -19,8 +19,8 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
|
||||
|
||||
- Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
|
||||
- **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
|
||||
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's left (the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
|
||||
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `{.right}` — `{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone left of the text on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
|
||||
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's start edge (left in LTR, right in RTL — the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
|
||||
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `{.right}` — `{.right}`, `{.left}` float at 30% of the text column, to its end/start edge following the text direction (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone at the text's start edge on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
|
||||
|
||||
## Page structure and navigation
|
||||
|
||||
|
||||
+8
-3
@@ -4,17 +4,22 @@ The Vue editor is a single tabbed `EditorShell.vue` mounted in a host div create
|
||||
|
||||
## Tabs
|
||||
|
||||
The shell hosts four kept-alive tabs (ordered site-wide first — site, structure — then, after a visual break, the per-page tabs — article, banner):
|
||||
The shell hosts five kept-alive tabs (ordered site-wide first — site, structure, localization — then, after a visual break, the per-page tabs — article, banner):
|
||||
|
||||
- `PageEditor.vue` — CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`, previewing into the visible article; editor and article scrolls are linked piecewise-linearly, keyed on the section anchors' `data-line` (markdown source line the backend stamps on top-level anchored h1/h2s): the page follows the cursor (fractional, wrap-aware, scrolling only when the cursor's page position leaves the viewport, with an edge margin), the editor follows page scroll with a progress-based viewport anchor, applied instantly (the window keeps scrolling normally while any editor is open — the panel is fixed to the viewport's left edge, its top tracking the banner's bottom edge until the banner scrolls away — and the panel scrolls internally); anchored h2s carry their own edit pens that open the editor scrolled to that section; a format bar offers Markdown helpers — bold/italic/code/link/table/image upload (always block-level on a fresh blank-separated line of its own — a cursor on a non-empty line, e.g. inside an existing image tag, inserts after that line, never into it; always with an empty `""` caption, cursor inside the quotes), toggling fences (` ``` ` code blocks and `::: aside` containers share the same machinery: clicked inside one they remove it and select the content, otherwise they wrap the selection or the cursor's line, keeping it selected), and `.left`/`.right`/`.wide`/`.margin` placement toggles plus `.small`/`.large`/`.huge` text-size toggles (brace attributes on the block at the cursor, mutually exclusive within each group; on `:::` containers a placement class replaces the container name instead — `::: aside` → `::: margin`), with Ctrl/Cmd-B/I/S bindings — for the hard-to-remember syntax. Edits content and title only, never the path.
|
||||
- `BannerEditor.vue` — per-page banner HTML + banner design selector, previewed into `#page-banner`.
|
||||
- `SiteEditor.vue` — site brand + optional custom brand HTML with image/video upload + theme selector + page-transition selector + font picker + favicon upload — clicking the preview tile picks a new one — + site-wide custom CSS, CSS injected into `<head id="pagerite-user">`.
|
||||
- `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row.
|
||||
- `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row, plus a per-row flag dropdown setting the page's primary language (`Node.language`, inherited by the subtree).
|
||||
- `LocalizationEditor.vue` — the site-wide translation settings: target languages as a flag grid (toggles, grouped in geographic rows; see docs/localization.md), the refresh-all-translations button, and the translator service WebSocket URL(s) to connect `scripts/translator.py` to.
|
||||
|
||||
Media uploads everywhere use the image icon buttons (pasting into the editor works too). The article, banner and site-settings pens are shorthands that open the shell on the matching tab; once open, clicking a pen switches tabs (and retargets the editors to the current page) instead of closing/remounting. The close button in the tab bar closes the shell (deliberately NOT Escape — it fired too easily by accident); tabs have no close buttons of their own. Closing only HIDES the shell — the Vue app stays mounted, so page-editor state (unsaved text included) survives until a real page reload; the editor always follows the URL, so fetch-navigating with the shell open (or before re-opening it) retargets it to the new page — unsaved text is stashed per path for the session and restored when returning, cleared on save. Saving there is explicit (Ctrl+S) and refreshes the page regions in place. Admin panels never reload the page.
|
||||
|
||||
In-place page re-rendering shared by the banner/site/structure tabs lives in `swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic regions, replaceState). It also exports `dropPageCache`, which the editor tabs call after any save that can alter the rendered HTML of other pages (theme, headings, structure, banners, site brand/CSS, favicon). Dropping the cache while editing avoids re-fetching every page immediately; the public runtime re-preloads visible links once the editor panel closes.
|
||||
|
||||
The page and structure tabs share one language selector: `LangSelect.vue` (small flag + dropdown) v-modeled on the shell-wide selection in `editorLang.js` (`''` = primary). While the panel is open that selection overrides the page's normal language preferences: EditorShell calls `swapdoc.setLangOverride`, which pins every `loadPlain` fetch (`?lang=`, the primary by its own code) and pagerite.js's own fetches/prefetches (`pagerite:session-lang`), until the panel closes and the override clears.
|
||||
|
||||
All WebSockets (page/banner editors, analytics view, the pagerite.js activity channel) pace their connections through `reconnect.js`: new sockets are created a staggered slot apart (a page load opens Vite's HMR socket plus several of ours at the same moment, and such bursts — like rapid retries — trip the browser's WebSocket throttling, leaving every socket to the host "pending" for minutes), a watchdog closes sockets stuck CONNECTING so they reschedule instead of hanging forever, and retries follow an exponential backoff with jitter that only a healthy connection resets. While a socket is connecting or waiting to reconnect the panel says so (`ConnNote.vue`), and the CodeMirror editors stay locked until their document arrives (typing before the doc accept would be clobbered by it).
|
||||
|
||||
## Saving behavior
|
||||
|
||||
Everything saves immediately as you edit (brand/title/CSS debounced, slug on commit since it renames the path), theme change swaps the stylesheet in place, tree rows navigate in place without transitions when focused, and the front page is a root-only row whose empty slug is editable like any other. Saves that can affect other pages drop the prefetch cache; the cache is rebuilt when the editor panel closes so navigation stays instant.
|
||||
@@ -25,6 +30,6 @@ Dropping ON the lower part of a row moves the page under that row (the child lis
|
||||
|
||||
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
|
||||
|
||||
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme, banner-design and transition CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-transition`/`#pagerite-user` elements to restore the canonical order (base < theme < design < transition < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
|
||||
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js`, `src/pagerite.js` and `src/analytics-main.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme, banner-design and transition CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-transition`/`#pagerite-user` elements to restore the canonical order (base < theme < design < transition < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
|
||||
|
||||
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
|
||||
|
||||
@@ -26,6 +26,7 @@ Vite builds ES-module `.js` outputs; in dev the backend links them as `<script t
|
||||
|
||||
All site data lives under `<hostname>/` in the cwd — `content.kantadb`,
|
||||
`analytics.json` and `files/` — where `<hostname>` is the CLI's first
|
||||
positional argument (default `localhost`, exported as `PAGERITE_HOSTNAME`;
|
||||
positional argument (default `localhost`, passed to the app as JSON in
|
||||
`PAGERITE_CONFIG`, see `pagerite/config.py`;
|
||||
`PAGERITE_DB`/`PAGERITE_ANALYTICS`/`PAGERITE_FILES` override individual
|
||||
paths). gitignored. Do not delete it without asking.
|
||||
|
||||
@@ -0,0 +1,517 @@
|
||||
# Localization
|
||||
|
||||
Pages are served in the visitor's language based on a `?lang=` query
|
||||
parameter or the `Accept-Language` header.
|
||||
|
||||
- **Phase 1 (implemented):** negotiation, URL scheme, caching, rendering
|
||||
plumbing. Translations are consumed through a stub interface; the database
|
||||
still holds only the original language.
|
||||
- **Phase 2 (implemented):** gettext-style fragment storage in the
|
||||
database — machine-translated chunks plus user override patches, assembled
|
||||
at render time. Storage details in `docs/migrate.md`.
|
||||
|
||||
## Phase 1: negotiation and URLs
|
||||
|
||||
### The primary language
|
||||
|
||||
Each article has a primary (original) language: `Node.language`, inherited
|
||||
down the tree like `banner` — "" = the nearest ancestor's, the front page
|
||||
last (it doubles as the site default), with `en` as the final fallback
|
||||
(`ORIGINAL_LANGUAGE`, `primary_lang()` in `pagerite/i18n.py`). It is
|
||||
configured per row in the structure editor. Everything per-article keys
|
||||
off the resolved value: language selection, `<html lang>`, canonical URLs,
|
||||
what counts as a translation, and the translation targets (a node's own
|
||||
primary is never one — so the target set may include the site default, and
|
||||
a page in another language can be translated into it).
|
||||
|
||||
### Language selection
|
||||
|
||||
Deliberately simple — **q-values are ignored**:
|
||||
|
||||
- All known `Accept-Language` implementations send the header **in order of
|
||||
preference**, so we parse it as an ordered list and never reorder.
|
||||
- Selection rule (`select_language` in `pagerite/i18n.py`):
|
||||
1. If `?lang=<tag>` is present, use it (if a translation exists; otherwise
|
||||
fall through to header logic).
|
||||
2. If the article's original language appears anywhere in the header list,
|
||||
use the **original**. Rationale: an AI translation is strictly worse
|
||||
than the original for anyone who has that language configured at all
|
||||
(e.g. `fi-FI, fi, en-US, en` gets English, not machine-translated
|
||||
Finnish).
|
||||
3. Otherwise walk the header list in order and use the first language for
|
||||
which a translation exists.
|
||||
4. Fall back to the original.
|
||||
|
||||
Region tags normalize to their base subtag (`fi-FI` → `fi`).
|
||||
|
||||
### URLs: pretty for users, indexable for search engines
|
||||
|
||||
- Canonical URLs stay pretty (`/some-page`). Each language version is
|
||||
addressable as `/some-page?lang=fi` so search engines can index them.
|
||||
- `<link rel="canonical">` names the **actually served language**: the plain
|
||||
URL when serving the original (for SEO the non-query URL means the
|
||||
article's own language), `?lang=xx` when serving a translation — however
|
||||
the language was arrived at (query or header).
|
||||
- `<link rel="alternate" hreflang="…">` entries follow the canonical
|
||||
directly (before the social meta tags) and list the languages the page
|
||||
is **actually available in**: `x-default` first, pointing at the plain
|
||||
autodetecting URL, then every available language — the original again by
|
||||
its plain URL, translations by `?lang=`. The public language selector
|
||||
keys off these: pagerite.js mounts the editors' flag dropdown in the
|
||||
top-right corner when the head advertises x-default plus more than one
|
||||
language, loading its bundle (Vue + the flag SVG set) on demand.
|
||||
- The override sticks for the session of clicks: a page requested with
|
||||
`?lang=` replicates the query onto the navigation links it renders (nav,
|
||||
sidebar, cards, brand — in-article links are content and stay as
|
||||
authored), so plain clicks and no-JS navigation keep the language.
|
||||
pagerite.js additionally strips the query from the address bar via
|
||||
`history.replaceState` (pretty, shareable URLs), remembers the language,
|
||||
and adds it to every internal fetch that lacks one (preloads,
|
||||
fetch-navigations, history traversals); history entries stay query-less.
|
||||
- The public selector's pick is the same override, pure JS state
|
||||
(`pagerite:set-session-lang`): the session language changes and the page
|
||||
swaps in place — no `?lang=` in the address bar, no reload. The choice is
|
||||
linked with the editor panel's language dropdown both ways; closing the
|
||||
panel keeps the chosen language instead of reverting.
|
||||
- A full page refresh or a shared link resets to automatic selection (header
|
||||
only). This gives a clean one-time override without cookies.
|
||||
|
||||
### Response correctness
|
||||
|
||||
- Content responses carry `Vary: accept-language` (added to the existing
|
||||
`accept-encoding` vary).
|
||||
- `_cached_body` and the page ETag include the **selected language** (not the
|
||||
raw header, which would blow up the cache key space) and the **replicated
|
||||
link language**: a `?lang=fi` render and a header-selected Finnish render
|
||||
of the same page differ in their navigation links, so they are cached as
|
||||
separate variants.
|
||||
- `<html lang="…">` reflects the served language, and an RTL language
|
||||
(`i18n.RTL_LANGUAGES` — ar, fa, he, ur) also sets `dir="rtl"` on `<html>`
|
||||
(the editor panel carries its own `lang="en" dir="ltr"` so it stays LTR).
|
||||
Client-side page swaps (fetch navigation in pagerite.js, editor re-renders
|
||||
in swapdoc.js) copy both attributes from the fetched document, so a hot
|
||||
switch into or out of an RTL page flips the layout without a reload.
|
||||
|
||||
### Rendering
|
||||
|
||||
- The translated Markdown goes through the same `markdown.render` pipeline.
|
||||
- Section anchors (`#hash` ids on h1/h2 headings) stay in the original
|
||||
language: render(anchors_from=...) pins the translated render's heading
|
||||
ids to the original text's slugs, matched by heading position, so links
|
||||
to sections don't break across languages.
|
||||
- Navigation/sidebar titles come from the translation's title map, with
|
||||
per-node fallback to the original title (a partially translated tree must
|
||||
still render).
|
||||
- Category placeholder pages (the 404s for content-less labels) select a
|
||||
language like content pages, but over the **subtree's** combined
|
||||
availability (`subtree_languages`) — they have no chunks of their own;
|
||||
the heading, navigation and card text localize from the title map and
|
||||
the target articles' translations. Their hreflang alternates are
|
||||
computed exactly like a content page's (a translated title counts as
|
||||
availability, so the language selector is offered there too).
|
||||
- Card descriptions and cover picks run on the target article's hybrid
|
||||
Markdown where that page is available in the served language, with
|
||||
per-card fallback to the original.
|
||||
- Fixed UI strings ("Not Found" etc.) and the editor UI stay English for now.
|
||||
- The markdown typographer (SmartyPants) is English-centric; per-language
|
||||
typographer options are a possible follow-up, not blocking.
|
||||
|
||||
## Phase 2: fragment-based translation storage (implemented)
|
||||
|
||||
Phase 1 assumed whole-page translated Markdown delivered from outside. The
|
||||
refined model is gettext-style: an article has **one primary version** (its
|
||||
`content`, in its own language) plus, per target language, **machine
|
||||
fragments** (translated chunks of Markdown) and **user patches** (minimal
|
||||
editor overrides). Both are stored in the database and assembled into the
|
||||
served Markdown at render time.
|
||||
|
||||
### The scenario this must handle
|
||||
|
||||
1. Article written in English.
|
||||
2. Machine-translated into Spanish → fragments stored.
|
||||
3. Editor fixes one Spanish paragraph and changes a link elsewhere to point
|
||||
at a Spanish resource → user patch hunks stored.
|
||||
4. English article edited → the edited chunk's key changes; its Spanish
|
||||
fragment no longer matches.
|
||||
5. Page requested before the machine translation refreshes → served as a
|
||||
**hybrid**: old fragments for unchanged chunks, plain English for the
|
||||
edited chunk. User patches are attempted against this hybrid, best effort,
|
||||
each hunk independently: the text fix is stale (its search text no longer
|
||||
exists) and silently skipped; the link change still applies even though
|
||||
the link sits in the now-English paragraph.
|
||||
6. Machine translation refreshes → full Spanish again, with both patch hunks
|
||||
applying.
|
||||
|
||||
### Chunks
|
||||
|
||||
`chunk_markdown(markdown)` splits the source into block-level chunks —
|
||||
blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
|
||||
list blocks, tables, HTML blocks. Container fence lines (`::: name` openers
|
||||
and `:::` closers) are always their own chunk, blank lines or not — folded
|
||||
into a prose chunk the closer would cross to the translator as part of the
|
||||
text, where the model can drop it (the rest of the page then renders inside
|
||||
the container). A chunk's identity is its **source text**,
|
||||
gettext-msgid style:
|
||||
|
||||
```python
|
||||
chunk_key = blake3(normalize(chunk_text)).digest(9) # bytes; base64 at the JSON level
|
||||
```
|
||||
|
||||
(`normalize`: strip trailing whitespace per line, collapse surrounding blank
|
||||
lines — so whitespace-only source edits don't invalidate translations.)
|
||||
|
||||
Consequences:
|
||||
|
||||
- Editing the English source invalidates exactly the edited chunks; all
|
||||
other fragments keep applying. Stale fragments are simply never referenced
|
||||
again and can be garbage-collected lazily (or left; they are tiny).
|
||||
- No explicit "source version" bookkeeping is needed — staleness falls out
|
||||
of the keys.
|
||||
|
||||
### User patches
|
||||
|
||||
Editors always edit **full Markdown** in the existing editor UX — never
|
||||
fragments. When editing a translated view (`?lang=es`), the editor is loaded
|
||||
with the *current hybrid Markdown*; on save, the server computes a minimal
|
||||
diff against that hybrid and stores it as a patch:
|
||||
|
||||
```python
|
||||
class Patch(msgspec.Struct, omit_defaults=True):
|
||||
"""One editing session's overrides, applied independently per hunk."""
|
||||
|
||||
hunks: list[tuple[str, str]] = [] # (search, replace) on hybrid Markdown
|
||||
```
|
||||
|
||||
Hunks are produced from `difflib.SequenceMatcher` on the hybrid vs. the
|
||||
edited text at block granularity: each `replace`/`delete`/`insert` opcode
|
||||
becomes one `(search, replace)` pair, with the preceding block's tail as
|
||||
left context for `insert` (pure inserts have empty search context otherwise).
|
||||
Application is dead simple:
|
||||
|
||||
```python
|
||||
def apply_patch(hybrid: str, patch: Patch) -> str:
|
||||
for search, replace in patch.hunks:
|
||||
if search and search in hybrid:
|
||||
hybrid = hybrid.replace(search, replace, 1)
|
||||
# missing search text = stale hunk -> silently skipped
|
||||
return hybrid
|
||||
```
|
||||
|
||||
Per-hunk independence is the robustness property from the scenario: a stale
|
||||
text fix does not block a still-valid link change. Patches are stored as an
|
||||
ordered list and applied in order.
|
||||
|
||||
### Storage
|
||||
|
||||
Full storage design and the `migrate_v3` restructuring live in
|
||||
`docs/migrate.md`. The short version, as it concerns this document:
|
||||
|
||||
- Originals **and** translations are content-addressed text chunks in flat
|
||||
stores: `Data.chunks: dict[bytes, str]` and
|
||||
`Data.trans: dict[bytes, dict[str, str]]` (chunk hash → lang → text) —
|
||||
path-independent, so repeated paragraphs and menu titles are translated
|
||||
once and article moves touch nothing. `Node.chunks: list[bytes]` gives
|
||||
each article its order.
|
||||
- `Node` gains **`language: str = ""`**, inherited down the tree like
|
||||
`banner` (empty = nearest ancestor, front page last, site default `en`
|
||||
final). `select_language` and `<html lang>` use the resolved value instead
|
||||
of the global `ORIGINAL_LANGUAGE` constant.
|
||||
- **Known weakness:** changing a page's (or subtree's) `language` after
|
||||
translations exist mis-keys everything — translations are keyed by
|
||||
*source* chunks, so old entries silently stop matching and user patches
|
||||
(searching for old-hybrid text) mostly go stale. That is acceptable:
|
||||
the orphaned data is harmless and translations regenerate. We do not
|
||||
migrate translations across a language change.
|
||||
- Article paths are stored and keyed **without leading slashes**
|
||||
(`"docs/setup"`, front page `""`); slashes are added only in hrefs.
|
||||
|
||||
### Render pipeline (the phase-1 `get_translation` stub, now real)
|
||||
|
||||
```python
|
||||
def get_translation(data, path, lang) -> Translation | None:
|
||||
if lang not in node.langs:
|
||||
return None
|
||||
hybrid = "\n\n".join(
|
||||
chunks[h] if h in node.no_trans else trans.get(h, {}).get(lang, chunks[h])
|
||||
for h in node.chunks
|
||||
)
|
||||
for patch in data.patches.get(f"{path}:{lang}", []):
|
||||
hybrid = apply_patch(hybrid, patch)
|
||||
return Translation(markdown=hybrid, titles=title_map(data, lang))
|
||||
```
|
||||
|
||||
- Availability is an article-level index: `node.langs: dict[lang, True]`,
|
||||
maintained by the translation writers (translator job, patch saves) in the
|
||||
same transaction as their data writes — rendering and language selection
|
||||
never probe the `trans` store chunk by chunk. A stale key is benign (the
|
||||
"translation" just renders as the original).
|
||||
- `titles` for nav/sidebar/cards: each node's translated title is
|
||||
`trans.get(hash(node.title), {}).get(lang)` with per-node fallback — one
|
||||
dict lookup per nav item at render time.
|
||||
- Cache invalidation: writes to `chunks` / `trans` / `patches` (translator,
|
||||
editor saves) call `_invalidate_pages()`, same as content writes.
|
||||
|
||||
### Editor flow
|
||||
|
||||
The page and structure editors share one language selector (`LangSelect.vue`:
|
||||
a small flag button opening a dropdown; the same country-flag-icons set as
|
||||
the analytics visitor cells), v-modeled on one shell-wide selection
|
||||
(`editorLang.js`, `''` = the primary language). The page editor lists the
|
||||
page's own primary language (`Node.language`, resolved through the
|
||||
hierarchy and echoed in the WS doc as `primary_lang`) plus the union of
|
||||
the page's translations (`node.langs`) and
|
||||
the site-wide `translate_langs`; it always opens in the primary language,
|
||||
even when the page itself was served in a translation. A note under the
|
||||
toolbar states the blast radius:
|
||||
edits to the primary language re-chunk the original (invalidating the
|
||||
affected translation fragments everywhere); edits to a translation stay
|
||||
local to that language.
|
||||
|
||||
While the editor panel is open, its language selection **overrides the
|
||||
normal language preferences** for the page preview: EditorShell pins every
|
||||
in-place re-render and pagerite.js fetch/prefetch to it (`?lang=` — a
|
||||
primary selection pins by the current page's own resolved primary, which
|
||||
`select_language` honors), and closing the panel restores the normal
|
||||
preferences.
|
||||
|
||||
- WS `open` with a `lang` returns the effective **hybrid** Markdown and
|
||||
title for that language (ungated by `node.langs` — a language without
|
||||
any fragments yet starts from the original text), plus the language
|
||||
metadata (`lang`, `primary_lang`, `langs`, `translate_langs`).
|
||||
- The editor keeps a **shadow copy** of the Markdown it opened. WS `save`
|
||||
with `lang` sends it as `base`; the server diffs `base` → submitted text
|
||||
(`make_patch`) and appends a `Patch`. Diffing against the shadow (rather
|
||||
than the current hybrid) keeps hunks correct when the original or the
|
||||
machine translation moved under an open editor; application against the
|
||||
then-current hybrid stays best-effort per hunk, as designed.
|
||||
- A changed **title** on a translated save becomes a fragment in
|
||||
`Data.trans` keyed by the original title's chunk hash — the same storage
|
||||
as machine title translations. An untouched title field (holding the
|
||||
served translation) is not sent, so saving never freezes a stale machine
|
||||
title into an override.
|
||||
- Saving never deletes; a translation additionally cannot be emptied (that
|
||||
would render as a blank page in that language).
|
||||
- The live preview renders the version being edited, whichever language
|
||||
the page itself was loaded in (the render is just the edited Markdown +
|
||||
title). A translated save keeps that preview in place — re-fetching the
|
||||
page would come back in the header-selected language.
|
||||
- Saving the primary-language version re-chunks the submitted Markdown and
|
||||
updates `Data.chunks` / `node.chunks` — only genuinely new text lands in
|
||||
the kanta change diff (see docs/migrate.md).
|
||||
|
||||
The **structure editor** selects from the same languages with the same
|
||||
`LangSelect` (the selection is shared — switching in either tab switches
|
||||
both, and the preview). It is also where a page's **primary language** is
|
||||
configured: each row carries a small flag dropdown (the resolved flag,
|
||||
dimmed while inherited) that sets `Node.language` via a structure op —
|
||||
'' = inherit, so setting it on a section covers the whole subtree. The
|
||||
tree it
|
||||
lists (`GET /_api/pages?lang=`) comes back with per-language titles where a
|
||||
translation exists (`translated` marks those rows; untranslated rows show
|
||||
the original title, dimmed). Retitling in a non-primary language posts the
|
||||
structure op with a `lang` and writes a per-language title fragment in
|
||||
`Data.trans` (keyed by the original title's chunk hash, exactly like a
|
||||
machine title translation — a user edit simply overwrites it); sending the
|
||||
original's text drops the override. The structure itself — slugs,
|
||||
hierarchy, order — is language-independent, so pending rows, slug edits,
|
||||
drag-and-drop and deletes work identically in every language.
|
||||
|
||||
### Translator service API
|
||||
|
||||
An external machine-translation service connects over WebSocket at
|
||||
`/_translate/{key}` — deliberately **not** under `/_api`: the SSO
|
||||
forward-auth does not cover that route, and the key in the path is the
|
||||
access control. Keys live in `Data.translate_keys` (key -> display name) —
|
||||
12 lowercase alphanumeric characters each; the first is generated at
|
||||
database bootstrap, further ones are managed in the editor's lang tab
|
||||
(add/rename/delete ride the `PUT /_api/settings` round-trip; the name is
|
||||
an inline display label only). The full WS URL(s) are printed in the
|
||||
startup log (`ws://localhost:{port}/_translate/{key}` locally,
|
||||
`wss://{hostname}/_translate/{key}` on a public hostname) and shown in the
|
||||
lang tab as click-to-copy links; the keys are also surfaced in
|
||||
`GET /_api/settings` as `translate_keys`. An unknown or empty key rejects
|
||||
the handshake (close-before-accept → HTTP 403). Transactions storing results record the connecting key as the kanta
|
||||
transaction `user`.
|
||||
|
||||
Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
|
||||
`bytes` fields ride as base64):
|
||||
|
||||
- `{"type": "hello", "langs": [...]}` — client greeting announcing its
|
||||
**capabilities**: the language codes its model can produce (normalized
|
||||
to base subtags; `en`/empty dropped).
|
||||
- `{"type": "job", "lang", "key", "texts", "path", "kind", "contexts"}` —
|
||||
server push: ONE fragment to translate (an article title or a chunk), as
|
||||
a list of **prose segments** (see Segmentation below). `contexts` is
|
||||
parallel to `texts` ("" = none): the surround to translate the segment
|
||||
in — for clients that translate better with context (see below).
|
||||
Contexts are not part of the result.
|
||||
- `{"type": "result", "lang", "key", "texts"}` — client reply: the
|
||||
segments translated, same order and count, matching its job by (lang, key).
|
||||
|
||||
Which languages get translated is **server-configured**:
|
||||
`Data.translate_langs` (presence-key dict, bootstrapped to Spanish and
|
||||
Chinese — edited in the editor shell's localization tab, whose flag grid
|
||||
lists every language including English, or set via `/_api/settings` as
|
||||
`translate_langs`). A target equal to an article's own primary language is
|
||||
skipped per article (its original already is that language), so the set
|
||||
may freely contain the site default. The dispatcher offers a
|
||||
connection jobs only in `wanted ∩ capable`; a connection without overlap
|
||||
simply stays idle.
|
||||
|
||||
`DELETE /_api/translations` (the localization tab's "refresh all
|
||||
translations" button) drops every machine translation (`Data.trans`) and
|
||||
rebuilds the availability index (`node.langs`) from the surviving user
|
||||
patches, so the dispatcher re-translates everything from scratch; the
|
||||
run's validation skip-list is cleared with it, giving rejected fragments
|
||||
another chance.
|
||||
|
||||
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; api.py only
|
||||
registers the route):
|
||||
|
||||
- **One job at a time per connection** — the next job is sent only after
|
||||
the current one's result. Clients wanting parallelism open multiple
|
||||
connections (e.g. several `scripts/translator.py` instances).
|
||||
- Pending work is derived from the `trans` store
|
||||
(`translate.pending_items`) minus the items in flight on any connection,
|
||||
so a **disconnect requeues** that connection's in-flight item and it is
|
||||
offered to any free capable connection.
|
||||
- Dispatch re-runs on every relevant event: Hello, result, disconnect and
|
||||
content change (`_invalidate_pages()` schedules it, so the pass runs
|
||||
after the writing transaction commits).
|
||||
- A result with no job in flight, a mismatched (lang, key), a duplicate
|
||||
hello, or any malformed frame closes the socket with a protocol error.
|
||||
|
||||
Results are stored into `trans` in one transaction and set
|
||||
`node.langs[lang]` on every article they touch (shared chunks make several
|
||||
pages gain a language from one fragment). Unknown keys are stored anyway
|
||||
and re-storing overwrites — results are idempotent.
|
||||
|
||||
#### Segmentation
|
||||
|
||||
Fragments cross the wire as **prose segments** (`pagerite/segments.py`): the
|
||||
fragment is parsed with the project's own markdown-it setup
|
||||
(`markdown.make_md(verbatim=True)` — all extensions, but no typographer or
|
||||
tasklist label wrapping, so token text stays byte-identical to the source)
|
||||
and split into the runs a model may touch: paragraph/heading/table-cell text
|
||||
(merged across soft line breaks), image alt texts and captions, footnote
|
||||
bodies. A block of plain text, inline **links and paired text formatting**
|
||||
(strong/em/s) **stays whole** — link and formatted texts cross inline, in
|
||||
sentence context, with the Markdown stripped (see below). Everything else
|
||||
never leaves the server: code spans and
|
||||
fences, URLs and autolinks, link/image *destinations*, `{...}` spans
|
||||
(placeholders like `{dates}` as well as attrs), reference and footnote
|
||||
labels, container fences, GFM alert markers (`[!NOTE]`), raw HTML — and the
|
||||
remaining markup punctuation (`|`, `:::`), which is a run boundary.
|
||||
Chunks with no segments (a lone `{dates}`, container fences, pure
|
||||
code/HTML) are never dispatched at all (`needs_translation`); every
|
||||
language renders them from the original chunk. Each segment is accompanied
|
||||
by a context string (a segment carved out of a larger block carries the
|
||||
block's plain text; a whole-block segment carries "") — context is a
|
||||
prompt aid only, never spliced into the result.
|
||||
|
||||
Reassembly is offset splicing, not text the model produced: each segment's
|
||||
source span was located at dispatch (sequential search; a run that is not a
|
||||
verbatim source substring — entity-decoded text, backslash escapes — is
|
||||
skipped and stays in the original language), and the returned translations
|
||||
are swapped in by offset. Markup corruption is therefore impossible by
|
||||
construction; the failure modes that remain are a wrong segment count, an
|
||||
empty segment, markup injected INTO a segment (a `<br>` in a title
|
||||
translation would splice live HTML), or a line that would start a new
|
||||
block where the segment lands (a ``` or ::: fence line would eat the rest
|
||||
of the block it splices into, closing fence included — segments are
|
||||
inline prose, so `pure_prose` alone cannot see this) — each returned
|
||||
segment must parse as
|
||||
pure prose with no block-starting line or blank line, or the whole result
|
||||
is dropped and logged, and the (lang, key)
|
||||
pair is skipped for the rest of the server run (generation is
|
||||
near-deterministic, so an immediate retry would re-fail; the fragment stays
|
||||
pending and gets another chance on restart or `DELETE /_api/translations`).
|
||||
`Data.trans` therefore only ever holds clean translated Markdown.
|
||||
|
||||
Link- and formatting-carrying blocks are the one place a segment is not
|
||||
spliced verbatim: a label translated apart from its sentence comes back
|
||||
grammatically incompatible with it (case government, particles, word
|
||||
order), and shown the Markdown the model mangles it (Seed-X dropped the
|
||||
`**` and the glued-on colon in `**Pagerite**: …`), so the block crosses
|
||||
whole — all Markdown stripped — and the server re-inserts the link and
|
||||
formatting syntax into the translated block. The boundaries are found by
|
||||
**text processing alone** —
|
||||
markers on the wire are hopeless (an earlier sentinel-masking design let
|
||||
the model see and mangle exactly that punctuation: Seed-X renumbered the
|
||||
tokens and turned `` or `**`. Blocks mixing in any other inline
|
||||
markup (code spans, images, raw HTML) don't qualify and still split into
|
||||
runs at those boundaries.
|
||||
|
||||
Punctuation is the translator's own job: Seed-X tends to "finish" short
|
||||
labels (titles, nav items) with a comma or period the source never had.
|
||||
Prompt wording is NOT the fix — a punctuation-instruction clause made
|
||||
Seed-X slip into its `[COT]` reasoning mode (minutes-long generations with
|
||||
reasoning text in the output, observed for Chinese). The reference client
|
||||
enforces punctuation deterministically instead (`match_punctuation` in
|
||||
scripts/translator.py): a translation of a segment without terminal
|
||||
punctuation gets any added trailing marks (and a newly opened Spanish ¡/¿)
|
||||
stripped before the result goes back.
|
||||
|
||||
The same client-side enforcement covers markup bleed as a CLASS, not per
|
||||
artifact: `<` is the prose/markup boundary on the wire and never appears in
|
||||
a segment in either direction. A literal `<` in the source text (`<1MB` is
|
||||
text, not markup — a tag needs a letter or `/!?`) crosses encoded as the
|
||||
fullwidth `<` and is decoded on return, before the result is validated and
|
||||
spliced (segments.py) — the wire itself still never carries `<`, and the
|
||||
reference client cuts the model's output at the first `<`
|
||||
(scripts/translator.py) — echoed language tags, stray `<br>`s and any
|
||||
future variant are one handled case. (The cut is post-decode, not a
|
||||
generation stop string: Seed-X opens every generation with its `<s>`
|
||||
framing token, which would trip a `<` stop immediately.)
|
||||
|
||||
Server-side, a second layer covers what the inline parser cannot: ASCII
|
||||
punctuation that is plain prose on the wire but Markdown syntax in the
|
||||
splice context — quotes (a translated `"` would close the quoted image
|
||||
title it lands in), brackets (alt texts, re-inserted link texts), `|` in
|
||||
table rows, `\` escapes. Rather than rejecting such results, `join` swaps
|
||||
them for Unicode look-alikes before splicing (`_NEUTRAL` in
|
||||
segments.py — curly quotes, fullwidth brackets; the renderer's
|
||||
typographer curls straight quotes anyway).
|
||||
|
||||
Short fragments get more than a bare prompt: each segment may carry its
|
||||
surround in `Job.contexts` — a title carries the article's opening prose
|
||||
(its own block is just the title word), a segment carved out of a larger
|
||||
block (a partial run; a link text whose block didn't qualify for the
|
||||
whole-block treatment) carries the block's plain text, and a
|
||||
whole-block segment (a plain paragraph) is self-contextualizing and carries
|
||||
"". The reference client translates segment and surround together, stops
|
||||
generation at the blank line separating them, and keeps the segment's own
|
||||
part of the output (its line resp. paragraph; a hard-break `␣␣\n` separator
|
||||
works too). If the model merged them (no separator, or an empty first
|
||||
part), it falls back to translating the segment alone. The surround fixes
|
||||
context-free readings ("About" as "approximately" — with the opening it
|
||||
becomes "Tietoa"/"Acerca de"; "here" as "就在这里" → the idiomatic
|
||||
"点击这里") and, as a side effect, most stray trailing punctuation.
|
||||
|
||||
### Explicitly out of scope for phase 2
|
||||
|
||||
- The machine translation itself: the API above moves fragments in and out;
|
||||
the translating is external. `scripts/translator.py` is the reference
|
||||
client (Seed-X-PPO-7B only — its 28 languages are the ceiling).
|
||||
- Garbage collection of orphaned chunks/translations (see docs/migrate.md).
|
||||
- sitemap.xml per-language entries; translated UI chrome; per-language
|
||||
typographer options; multi-locale date/number formatting.
|
||||
+178
@@ -0,0 +1,178 @@
|
||||
# migrate_v3: content-addressed chunk storage
|
||||
|
||||
Status: **implemented**. `migrate_v3` restructures how article text and
|
||||
translations are stored, motivated by the localization model in
|
||||
`docs/localization.md` (phase 2). Since it is a full migration, it is free to
|
||||
break the current `Node.content: str | None` layout.
|
||||
|
||||
## Goals
|
||||
|
||||
- **Minimal change diffs.** kanta persists change diffs; editing one
|
||||
paragraph of a long article must not rewrite the whole article string, and
|
||||
a translation refresh must touch only the re-translated chunks.
|
||||
- **Fast, simple lookup.** Everything heavy lives in flat
|
||||
`dict[hash, content]` stores; ordering lives in `list[hash]`. No large
|
||||
nested structures, no deep paths.
|
||||
- **Path-independent text.** Chunks and their translations are keyed by
|
||||
content hash, not by article path — the same paragraph (or menu title)
|
||||
appearing in several articles is stored and translated once. Moving or
|
||||
renaming an article touches nothing.
|
||||
|
||||
## Design (chosen: global content-addressed stores)
|
||||
|
||||
Original articles are *also* stored as chunks; everything — originals and
|
||||
translations — lives in flat hash-keyed dicts. Costs accepted: rendering does
|
||||
one dict lookup per chunk (trivial), orphaned hashes need occasional garbage
|
||||
collection, and the editor save path re-chunks server-side (it already
|
||||
diffs). The rejected alternatives: per-article nested `LangVersion`
|
||||
structures (churn, duplication, whole-string originals) and a hybrid with
|
||||
whole originals plus global translations (keeps the worst change-diff
|
||||
property).
|
||||
|
||||
## Target layout
|
||||
|
||||
```python
|
||||
class Node(msgspec.Struct, omit_defaults=True):
|
||||
...
|
||||
#: Replaces `content: str | None`. None = pure category label;
|
||||
#: a list (possibly empty) = a page, as ordered chunk hashes.
|
||||
chunks: list[bytes] | None = None
|
||||
#: Primary language of the article (BCP-47 base tag). "" = inherit
|
||||
#: (nearest ancestor, front page last, site default "en" final).
|
||||
language: str = ""
|
||||
#: Chunk hashes the editor marked "do not translate" (always served
|
||||
#: from the original). Presence-keys, value always True.
|
||||
no_trans: dict[bytes, True] = {}
|
||||
#: Languages this article is available in (besides its primary
|
||||
#: language). Presence-keys, value always True — rendering, language
|
||||
#: selection and hreflang alternates read this set instead of probing
|
||||
#: the trans store chunk by chunk. Maintained by the writers (see
|
||||
#: "Language index maintenance" below).
|
||||
langs: dict[str, True] = {}
|
||||
|
||||
class Data(msgspec.Struct):
|
||||
...
|
||||
#: API keys gating the translator service WebSocket (/_translate/{key}):
|
||||
#: key -> display name; the first is generated at bootstrap (state.py).
|
||||
translate_keys: dict[str, str] = {}
|
||||
#: Wanted target languages for the translator service (presence-keys);
|
||||
#: jobs are offered only in these ∩ a connection's capabilities.
|
||||
translate_langs: dict[str, True] = {}
|
||||
#: All original-language text, content-addressed: blake3(normalized)
|
||||
#: digest[:9] -> Markdown chunk. Shared by every article. Keys are
|
||||
#: bytes; kanta/msgspec base64-encode them at the JSON level.
|
||||
chunks: dict[bytes, str] = {}
|
||||
#: Machine translations: chunk hash -> lang -> translated Markdown
|
||||
#: (nested, not tuple keys: msgspec's JSON serializer rejects them).
|
||||
#: Also used for node titles (hash of the title text).
|
||||
trans: dict[bytes, dict[str, str]] = {}
|
||||
#: User override patches per article and language:
|
||||
#: f"{path}:{lang}" -> ordered patches (see localization.md).
|
||||
patches: dict[str, list[Patch]] = {}
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- **Article paths never carry a leading slash** in the DB or in lookup keys
|
||||
(`"docs/setup"`, front page `""`); the leading slash is added only when
|
||||
building hrefs. `migrate_v3` audits existing stored paths (translation
|
||||
keys, analytics references, any path-valued fields) and normalizes them.
|
||||
- **Titles are chunks too**, by hash only: the nav renderer looks up
|
||||
`trans.get(hash(node.title), {}).get(lang)`. No separate title storage;
|
||||
editing a title invalidates its translations automatically.
|
||||
- **Per-hunk options** live in two places: *inherent* options are derived at
|
||||
chunking time (code fences, HTML blocks and prose-free chunks are
|
||||
no-translate without storing anything — `needs_translation`, see
|
||||
docs/localization.md "Masking"); *editor-set* flags are `node.no_trans`
|
||||
(keyed by chunk
|
||||
hash, so a heavy edit silently drops the flag — acceptable and
|
||||
self-healing).
|
||||
- **Patch payloads stay inline** in `Patch.hunks` — patches are small by
|
||||
construction (minimal server-computed diffs). If a pathological case shows
|
||||
up, hunks can be hash-stored later without schema pain.
|
||||
|
||||
## Language index maintenance (`node.langs`)
|
||||
|
||||
`node.langs` is a denormalized index over the `trans`/`patches` stores so
|
||||
that article rendering, `select_language`'s availability check, and hreflang
|
||||
alternate links never enumerate chunks. It is written by whoever writes
|
||||
translation data, in the same transaction:
|
||||
|
||||
- **Translator service:** the WebSocket API at `/_translate/{key}` (see
|
||||
docs/localization.md) offers pending fragments (titles + translatable
|
||||
chunks lacking an entry for the language) as single-item jobs — one at
|
||||
a time per connection, in `Data.translate_langs` ∩ the connection's
|
||||
announced capabilities — and receives the matching result; storing it
|
||||
writes the `trans[h][lang]` entry, sets `node.langs[lang] = True` on
|
||||
every article that gained one and invalidates the page cache — all in
|
||||
one transaction.
|
||||
- **Translated-view save:** appending the first patch for `f"{path}:{lang}"`
|
||||
sets `node.langs[lang] = True` (patches alone make the version exist).
|
||||
- **Removals:** deleting a patch or GC'ing translations re-derives the key:
|
||||
keep `lang` if any `trans` entry for the article's current chunks/title or
|
||||
any patch remains, otherwise drop it. Stale `langs` keys are benign (an
|
||||
advertised language that renders as the original), so removal can lag.
|
||||
|
||||
## Render / save pipeline (summary)
|
||||
|
||||
- **Render:** `text = "\n\n".join(chunks[h] for h in node.chunks)` for the
|
||||
original; for language `L` (only ever attempted when `L in node.langs`),
|
||||
per chunk `trans.get(h, {}).get(L)` unless missing or `h in node.no_trans`,
|
||||
falling back to `chunks[h]`; then apply `patches.get(f"{path}:{L}", [])`
|
||||
in order (per-hunk, best effort); then `markdown.render` as today. All of
|
||||
this assembles the `Translation` the phase-1 plumbing already consumes.
|
||||
- **Availability:** `node.langs` is the availability index; `?lang=`
|
||||
handling uses exactly this set. (hreflang alternates are site-wide from
|
||||
`translate_langs` instead — see docs/localization.md.)
|
||||
- **Save (primary language):** server re-chunks the submitted Markdown,
|
||||
inserts new hashes into `Data.chunks`, replaces `node.chunks`. Unchanged
|
||||
chunks keep their hashes — only genuinely new text lands in the diff.
|
||||
- **Save (translated view):** diff against the served hybrid, append a
|
||||
`Patch` under `patches[f"{path}:{lang}"]`; `node.chunks` untouched.
|
||||
- **Invalidate:** any write to `chunks` / `trans` / `patches` calls
|
||||
`_invalidate_pages()`.
|
||||
|
||||
## migrate_v3 steps
|
||||
|
||||
1. Walk `menu`; for every node with a string `content`:
|
||||
`chunks = chunk_markdown(content)`; write each into the new `chunks`
|
||||
store; replace the field with the hash list (`None` stays `None`).
|
||||
2. Initialize empty `chunks` / `trans` / `patches` stores.
|
||||
3. Normalize stored paths: strip leading slashes anywhere paths are keys or
|
||||
values.
|
||||
4. `language`, `no_trans` and `langs` need nothing — struct defaults cover
|
||||
them (`langs` starts empty; the translator job fills it as translations
|
||||
land).
|
||||
|
||||
Chunking must be deterministic and shared with render/save, so
|
||||
`chunk_markdown` + `chunk_key` live in `pagerite/i18n.py` (or a small
|
||||
`pagerite/chunks.py`) and are imported by both `migrations.py` and
|
||||
`views.py`/`state.py`.
|
||||
|
||||
## Implementation notes (deviations from the plan above)
|
||||
|
||||
- Chunking lives in `pagerite/chunks.py`; hashing uses the `blake3` package
|
||||
(already a dependency), truncated to a 9-byte `bytes` digest (kanta's
|
||||
JSON persistence base64-encodes bytes keys to 12-char strings).
|
||||
- `trans` is keyed `hash -> lang -> text` (nested dict), not by
|
||||
`f"{hash}:{lang}"` tuples: msgspec's JSON serializer only supports
|
||||
str-like/number-like dict keys, and kanta persists as JSON lines.
|
||||
- `Translation.titles` stayed keyed by node path (phase-1 shape, views
|
||||
untouched): `get_translation` builds it by walking the menu with the same
|
||||
per-title `trans.get(chunk_key(node.title), {}).get(lang)` lookups.
|
||||
- Insert hunks anchor on the whole preceding block (not just its tail) —
|
||||
a stronger, simpler search context.
|
||||
- `make_patch` diffs with `SequenceMatcher(autojunk=False)` so patches are
|
||||
deterministic (popular lines like blank separators never become junk).
|
||||
- Step 3's path normalization is a no-op in practice: the only path-keyed
|
||||
store (`patches`) starts empty at v3; analytics paths live outside the
|
||||
kantadb. The code still strips leading slashes defensively.
|
||||
|
||||
## Garbage collection (later, manual or idle-time)
|
||||
|
||||
Orphaned entries accumulate: chunks no longer referenced by any
|
||||
`node.chunks`/`node.title`, translations whose chunk hash is orphaned, patch
|
||||
hunks that never match. All are harmless (never read). A GC pass is a single
|
||||
tree walk collecting live hashes, then deleting the rest from `chunks` and
|
||||
`trans`; patches whose every hunk is stale get pruned. Not part of
|
||||
migrate_v3.
|
||||
@@ -37,3 +37,6 @@ __screenshots__/
|
||||
|
||||
# Playwright browser downloads (if ever installed locally)
|
||||
.pw-browsers/
|
||||
|
||||
# npm project config (audit/fund off: the audit endpoint stalls installs)
|
||||
!.npmrc
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
audit=false
|
||||
fund=false
|
||||
@@ -20,6 +20,7 @@
|
||||
"codemirror": "^6.0.2",
|
||||
"country-flag-icons": "^1.6.20",
|
||||
"overlayscrollbars": "^2.16.0",
|
||||
"pinia": "^4.0.3",
|
||||
"transliteration": "^2.6.1",
|
||||
"vue": "^3.5.26",
|
||||
"vuedraggable": "^4.1.0"
|
||||
|
||||
@@ -27,6 +27,8 @@ import VisitorCell from './VisitorCell.vue'
|
||||
import TransitionGraph from './TransitionGraph.vue'
|
||||
import VisitorCharts from './VisitorCharts.vue'
|
||||
import { VIEW_W } from './analytics/chart.js'
|
||||
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
|
||||
import ConnNote from './ConnNote.vue'
|
||||
|
||||
// Same centering margin as the charts, so the totals row's left edge
|
||||
// aligns with the chart svg above the natural width.
|
||||
@@ -40,8 +42,20 @@ const error = ref('')
|
||||
const now = ref(Date.now())
|
||||
let ws = null
|
||||
let reconnectTimeout = null
|
||||
let connectWatchdog = null
|
||||
const reconnects = reconnectPolicy()
|
||||
let timeInterval = null
|
||||
|
||||
// The panel is live data over its socket: while it is connecting or waiting
|
||||
// to reconnect, say so (ConnNote) instead of showing a silent stale view.
|
||||
const conn = ref('connecting') // connecting | open | waiting
|
||||
const retryIn = ref(0)
|
||||
const connNote = computed(() =>
|
||||
conn.value === 'connecting' ? 'connecting to the server…'
|
||||
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
|
||||
: '',
|
||||
)
|
||||
|
||||
// The initial range comes from the URL hash (shareable links); without one,
|
||||
// it is derived from the first analytics snapshot: day when the recorded
|
||||
// history is shorter than 24 h, week otherwise.
|
||||
@@ -51,9 +65,16 @@ let rangePinned = Boolean(RANGES[hashRange])
|
||||
|
||||
function connectAnalytics() {
|
||||
if (ws) return
|
||||
conn.value = 'connecting'
|
||||
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
|
||||
ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
|
||||
ws.onopen = () => { error.value = '' }
|
||||
clearTimeout(connectWatchdog)
|
||||
connectWatchdog = watchConnecting(ws, 'analytics')
|
||||
ws.onopen = () => {
|
||||
conn.value = 'open'
|
||||
reconnects.opened()
|
||||
error.value = ''
|
||||
}
|
||||
ws.onmessage = (event) => {
|
||||
try {
|
||||
data.value = JSON.parse(event.data)
|
||||
@@ -75,12 +96,19 @@ function connectAnalytics() {
|
||||
}
|
||||
ws.onclose = () => {
|
||||
ws = null
|
||||
reconnectTimeout = setTimeout(connectAnalytics, 2000)
|
||||
// The policy paces the retry: doubling backoff with jitter, reset only
|
||||
// by a healthy connection — a fixed rapid loop trips the browser's
|
||||
// WebSocket throttling (all sockets then sit "pending" for minutes).
|
||||
const wait = reconnects.closed()
|
||||
retryIn.value = Math.max(1, Math.round(wait / 1000))
|
||||
conn.value = 'waiting'
|
||||
reconnectTimeout = setTimeout(connectAnalytics, wait)
|
||||
}
|
||||
}
|
||||
|
||||
onMounted(async () => {
|
||||
connectAnalytics()
|
||||
// The first connection takes a staggered slot (see ./reconnect).
|
||||
reconnectTimeout = setTimeout(connectAnalytics, socketSlot())
|
||||
now.value = Date.now()
|
||||
timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
|
||||
// The site tree for the transition map (all pages in menu order). Not
|
||||
@@ -93,6 +121,7 @@ onMounted(async () => {
|
||||
|
||||
onUnmounted(() => {
|
||||
if (reconnectTimeout) clearTimeout(reconnectTimeout)
|
||||
if (connectWatchdog) clearTimeout(connectWatchdog)
|
||||
if (timeInterval) clearInterval(timeInterval)
|
||||
if (ws) {
|
||||
ws.onclose = null
|
||||
@@ -134,12 +163,13 @@ const favicons = computed(() => data.value?.favicons || {})
|
||||
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
|
||||
const crawlers = computed(() => rangeData.value?.crawlers || [])
|
||||
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
|
||||
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, now.value))
|
||||
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, pageTree.value, now.value))
|
||||
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<div class="analytics-view">
|
||||
<!-- Untranslated admin dashboard: always LTR, like the editor panel. -->
|
||||
<div class="analytics-view" lang="en" dir="ltr">
|
||||
<div class="analytics-panel">
|
||||
<header>
|
||||
<h1>Analytics</h1>
|
||||
@@ -151,6 +181,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
</nav>
|
||||
<a href="/" class="close" title="home">✕</a>
|
||||
</header>
|
||||
<ConnNote :text="connNote" />
|
||||
<p v-if="error" class="error">⚠️ {{ error }}</p>
|
||||
<p v-else-if="!data" class="loading">loading…</p>
|
||||
<template v-else>
|
||||
@@ -187,6 +218,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
:ip-display="v.ipDisplay"
|
||||
:ua="v.ua"
|
||||
:ua-raw="v.uaRaw"
|
||||
:ua-url="v.uaUrl"
|
||||
:country="v.country"
|
||||
:city="v.city"
|
||||
:lang="v.lang"
|
||||
@@ -214,6 +246,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
<tbody>
|
||||
<tr v-for="(c, i) in crawlerRows" :key="i">
|
||||
<td class="trail">
|
||||
<TrailLink v-if="c.refererStep" :step="c.refererStep" :favicons="favicons" @close="$emit('close')" />
|
||||
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
|
||||
</td>
|
||||
<VisitorCell
|
||||
@@ -221,6 +254,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
:ip-display="c.ipDisplay"
|
||||
:ua="c.ua"
|
||||
:ua-raw="c.uaRaw"
|
||||
:ua-url="c.uaUrl"
|
||||
:country="c.country"
|
||||
:city="c.city"
|
||||
:lang="c.lang"
|
||||
@@ -240,6 +274,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
<thead>
|
||||
<tr>
|
||||
<th>paths abused</th>
|
||||
<th>articles read</th>
|
||||
<th>visitor</th>
|
||||
<th class="last-seen">last seen</th>
|
||||
</tr>
|
||||
@@ -256,11 +291,18 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
|
||||
</div>
|
||||
</td>
|
||||
<td class="trail clickable-list"
|
||||
@click="copyList(a.allArticles, $event)">
|
||||
<TrailLink v-for="(s, si) in a.articles" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
|
||||
<small v-if="!a.articles.length" class="muted">—</small>
|
||||
</td>
|
||||
<VisitorCell
|
||||
:ip="a.ip"
|
||||
:ip-display="a.ipDisplay"
|
||||
:ua="a.ua"
|
||||
:ua-raw="a.uaRaw"
|
||||
:ua-url="a.uaUrl"
|
||||
:ua-raws="a.uaRaws"
|
||||
:country="a.country"
|
||||
:city="a.city"
|
||||
:lang="a.lang"
|
||||
@@ -468,22 +510,6 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
.visit-table .clickable-list,
|
||||
.visit-table .last-seen {
|
||||
cursor: pointer;
|
||||
position: relative;
|
||||
}
|
||||
|
||||
.visit-table :deep(.copy-popup) {
|
||||
position: absolute;
|
||||
bottom: calc(100% + 0.25rem);
|
||||
left: 50%;
|
||||
transform: translateX(-50%);
|
||||
padding: 0.15rem 0.4rem;
|
||||
background: var(--text, CanvasText);
|
||||
color: var(--bg, Canvas);
|
||||
border-radius: 0.25rem;
|
||||
font-size: 0.75rem;
|
||||
white-space: nowrap;
|
||||
pointer-events: none;
|
||||
z-index: 10;
|
||||
}
|
||||
|
||||
.crawler-top-uas {
|
||||
|
||||
@@ -3,11 +3,13 @@
|
||||
// the real #page-banner region. Close and tab switching live in EditorShell.
|
||||
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
|
||||
import { EditorView, basicSetup } from 'codemirror'
|
||||
import { EditorState } from '@codemirror/state'
|
||||
import { Compartment, EditorState } from '@codemirror/state'
|
||||
import { keymap } from '@codemirror/view'
|
||||
import { indentWithTab } from '@codemirror/commands'
|
||||
import { html } from '@codemirror/lang-html'
|
||||
import { cmHighlight, cmTheme } from './cmtheme'
|
||||
import ConnNote from './ConnNote.vue'
|
||||
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
|
||||
import { dropPageCache, loadPlain, runScripts } from './swapdoc'
|
||||
|
||||
const props = defineProps({
|
||||
@@ -25,9 +27,22 @@ const bannerEl = ref(null)
|
||||
let ws = null
|
||||
let pendingSave = null
|
||||
let reconnectTimer = null
|
||||
let reconnectDelay = 2000
|
||||
const MAX_RECONNECT_DELAY = 16000
|
||||
let connectWatchdog = null
|
||||
// Reconnection pacing lives in ./reconnect (shared with the other sockets).
|
||||
const reconnects = reconnectPolicy()
|
||||
let everConnected = false
|
||||
// Connection state drives the note at the top (ConnNote), and locks input
|
||||
// until the banner's document has arrived (typing before it would be
|
||||
// clobbered by the doc accept).
|
||||
const conn = ref('connecting') // connecting | open | waiting
|
||||
const retryIn = ref(0)
|
||||
const docReady = ref(false)
|
||||
const editable = new Compartment()
|
||||
const connNote = computed(() =>
|
||||
conn.value === 'connecting' ? 'connecting to the server…'
|
||||
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
|
||||
: docReady.value ? '' : 'loading the banner…',
|
||||
)
|
||||
let view = null // CodeMirror for the banner HTML
|
||||
let syncing = false // set while replacing the document programmatically
|
||||
|
||||
@@ -90,6 +105,9 @@ function save() {
|
||||
|
||||
function openPath(p) {
|
||||
path.value = p
|
||||
// Lock input until the doc arrives (typing would be clobbered by it).
|
||||
docReady.value = false
|
||||
view?.dispatch({ effects: editable.reconfigure(EditorView.editable.of(false)) })
|
||||
send({ type: 'open', path: p })
|
||||
}
|
||||
watch(() => props.pagePath, (p) => { openPath(normPath(p)) })
|
||||
@@ -209,6 +227,8 @@ function onMessage(ev) {
|
||||
const msg = JSON.parse(ev.data)
|
||||
if (msg.type === 'doc' && msg.path === path.value) {
|
||||
setDocument(msg.banner ?? '')
|
||||
docReady.value = true
|
||||
view.dispatch({ effects: editable.reconfigure(EditorView.editable.of(true)) })
|
||||
bannerDesign.value = msg.banner_design ?? null
|
||||
bannerDesignFrom.value = msg.banner_design_from ?? null
|
||||
bannerDesignInherited.value = msg.banner_design_inherited ?? ''
|
||||
@@ -234,12 +254,22 @@ function onKeydown(ev) {
|
||||
}
|
||||
|
||||
function connect() {
|
||||
clearTimeout(reconnectTimer)
|
||||
conn.value = 'connecting'
|
||||
if (ws) {
|
||||
// Replacing a stale socket: detach its handlers so its close is silent.
|
||||
ws.onopen = ws.onmessage = ws.onclose = ws.onerror = null
|
||||
if (ws.readyState !== WebSocket.CLOSED) ws.close()
|
||||
}
|
||||
ws = new WebSocket(
|
||||
`${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`,
|
||||
)
|
||||
ws.onmessage = onMessage
|
||||
clearTimeout(connectWatchdog)
|
||||
connectWatchdog = watchConnecting(ws, 'banner')
|
||||
ws.onopen = () => {
|
||||
reconnectDelay = 2000
|
||||
conn.value = 'open'
|
||||
reconnects.opened()
|
||||
if (everConnected) {
|
||||
if (pendingSave) send(pendingSave)
|
||||
} else {
|
||||
@@ -248,16 +278,19 @@ function connect() {
|
||||
everConnected = true
|
||||
}
|
||||
ws.onclose = () => {
|
||||
clearTimeout(reconnectTimer)
|
||||
reconnectTimer = setTimeout(() => {
|
||||
connect()
|
||||
reconnectDelay = Math.min(reconnectDelay * 2, MAX_RECONNECT_DELAY)
|
||||
}, reconnectDelay)
|
||||
// The wait is the policy's: doubling backoff with jitter (./reconnect),
|
||||
// reset only by a healthy connection — rapid retries trip the browser's
|
||||
// WebSocket throttling (sockets stuck "pending" for minutes).
|
||||
const wait = reconnects.closed()
|
||||
retryIn.value = Math.max(1, Math.round(wait / 1000))
|
||||
conn.value = 'waiting'
|
||||
reconnectTimer = setTimeout(connect, wait)
|
||||
}
|
||||
}
|
||||
|
||||
onMounted(async () => {
|
||||
connect()
|
||||
// The first connection takes a staggered slot (see ./reconnect).
|
||||
reconnectTimer = setTimeout(connect, socketSlot())
|
||||
view = new EditorView({
|
||||
state: EditorState.create({
|
||||
doc: '',
|
||||
@@ -269,6 +302,8 @@ onMounted(async () => {
|
||||
cmTheme,
|
||||
cmHighlight,
|
||||
EditorView.lineWrapping,
|
||||
// Locked until the banner's document arrives (docReady/ConnNote).
|
||||
editable.of(EditorView.editable.of(false)),
|
||||
EditorView.updateListener.of((u) => {
|
||||
if (u.docChanged && !syncing) {
|
||||
banner.value = view.state.doc.toString()
|
||||
@@ -286,6 +321,7 @@ onMounted(async () => {
|
||||
|
||||
onUnmounted(() => {
|
||||
clearTimeout(reconnectTimer)
|
||||
clearTimeout(connectWatchdog)
|
||||
for (const t of Object.values(timers)) clearTimeout(t)
|
||||
if (ws) {
|
||||
ws.onclose = null // intentional close, no reconnect
|
||||
@@ -300,6 +336,7 @@ onUnmounted(() => {
|
||||
<template>
|
||||
<div class="banner-editor">
|
||||
<div v-if="saveError">{{ saveError }}</div>
|
||||
<ConnNote :text="connNote" />
|
||||
|
||||
<section class="block" @paste="onBannerPaste">
|
||||
<div class="block-head">
|
||||
@@ -366,14 +403,6 @@ onUnmounted(() => {
|
||||
margin-left: auto;
|
||||
padding: 0 0.2rem;
|
||||
font-size: 1rem;
|
||||
background: none;
|
||||
border: none;
|
||||
cursor: pointer;
|
||||
opacity: 0.7;
|
||||
}
|
||||
|
||||
.block-head .icon-btn:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
/* The banner design selector stays compact; the upload button is pushed
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
<script setup>
|
||||
// Connection-state note for the WebSocket-backed panels (page/banner
|
||||
// editors, analytics view): while the socket is connecting or waiting to
|
||||
// reconnect the panel cannot load or save, and this says so. An empty
|
||||
// text hides the note.
|
||||
defineProps({ text: { type: String, default: '' } })
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<div v-if="text" class="conn-note" role="status">{{ text }}</div>
|
||||
</template>
|
||||
|
||||
<style scoped>
|
||||
.conn-note {
|
||||
padding: 0.2rem 1rem;
|
||||
border-bottom: 1px solid var(--line);
|
||||
background: var(--surface);
|
||||
color: var(--muted);
|
||||
font-size: 0.8rem;
|
||||
}
|
||||
</style>
|
||||
@@ -1,5 +1,5 @@
|
||||
<script setup>
|
||||
// Tabbed shell for the four admin editors. The individual pens are shorthands
|
||||
// Tabbed shell for the five admin editors. The individual pens are shorthands
|
||||
// that open the shell on a given tab; once open, tabs switch instantly without
|
||||
// closing the panel. Tabs are kept alive so switching preserves state.
|
||||
import { onMounted, onUnmounted, provide, ref, watch } from 'vue'
|
||||
@@ -7,6 +7,9 @@ import PageEditor from './PageEditor.vue'
|
||||
import BannerEditor from './BannerEditor.vue'
|
||||
import SiteEditor from './SiteEditor.vue'
|
||||
import StructureEditor from './StructureEditor.vue'
|
||||
import LocalizationEditor from './LocalizationEditor.vue'
|
||||
import { editorLang, pagePrimary } from './editorLang'
|
||||
import { loadPlain, setLangOverride } from './swapdoc'
|
||||
|
||||
const props = defineProps({
|
||||
pagePath: { type: String, default: '' },
|
||||
@@ -17,11 +20,46 @@ const emit = defineEmits(['close'])
|
||||
const currentPath = ref(props.pagePath)
|
||||
const activeMode = ref(props.initialMode)
|
||||
|
||||
// Tab order: site-wide settings first (site, structure), then — after a
|
||||
// visual break — the per-page editors (article, banner).
|
||||
// The shared language selection (./editorLang, v-modeled by the tabs'
|
||||
// LangSelects) is linked to the whole-page language: while the shell is
|
||||
// open it drives the page preview (overrides ?lang= / Accept-Language),
|
||||
// and closing keeps the pick as the session language. The primary
|
||||
// selection pins by the CURRENT PAGE's own primary language (pages may
|
||||
// differ — Node.language is inherited down the tree).
|
||||
let pinned = false
|
||||
function pinPreviewLang() {
|
||||
pinned = true
|
||||
// '' pagePrimary = not yet learned: pin 'en', the server's final fallback
|
||||
// (i18n.ORIGINAL_LANGUAGE).
|
||||
setLangOverride(editorLang.value || pagePrimary.value || 'en')
|
||||
loadPlain(currentPath.value)
|
||||
}
|
||||
// Opening the panel must not switch the page's language: adopt the
|
||||
// session's chosen language (public selector / earlier pick) once, then
|
||||
// pin. Runs only on (re)open — after that the selection is the user's.
|
||||
function openShell() {
|
||||
const session = window.__pageriteLang
|
||||
if (!editorLang.value && session && session !== (pagePrimary.value || 'en'))
|
||||
editorLang.value = session
|
||||
pinPreviewLang()
|
||||
}
|
||||
function unpinPreviewLang() {
|
||||
if (!pinned) return
|
||||
pinned = false
|
||||
setLangOverride(null)
|
||||
loadPlain(currentPath.value)
|
||||
}
|
||||
watch(editorLang, () => { if (pinned) pinPreviewLang() })
|
||||
// The page's primary may be (re)learned while pinned on it (doc accept,
|
||||
// tree refresh, a language change on the row) — re-pin with the new code.
|
||||
watch(pagePrimary, () => { if (pinned && !editorLang.value) pinPreviewLang() })
|
||||
|
||||
// Tab order: site-wide settings first (site, structure, localization), then
|
||||
// — after a visual break — the per-page editors (article, banner).
|
||||
const MODES = [
|
||||
{ key: 'site', label: 'site', component: SiteEditor },
|
||||
{ key: 'structure', label: 'structure', component: StructureEditor },
|
||||
{ key: 'localization', label: 'lang', component: LocalizationEditor },
|
||||
{ key: 'page', label: 'article', component: PageEditor, breakBefore: true },
|
||||
{ key: 'banner', label: 'banner', component: BannerEditor },
|
||||
]
|
||||
@@ -60,15 +98,27 @@ function onSwitchEvent(ev) {
|
||||
onMounted(() => {
|
||||
document.body.dataset.editorMode = activeMode.value
|
||||
addEventListener('pagerite:switch-editor', onSwitchEvent)
|
||||
addEventListener('pagerite:editor-shown', openShell)
|
||||
addEventListener('pagerite:editor-hidden', unpinPreviewLang)
|
||||
// The shell mounts visible (openEditor), so open immediately. The site
|
||||
// default primary language comes from the settings — it only fills the
|
||||
// unknown; the page/structure tabs refine pagePrimary per page as they
|
||||
// learn it (their knowledge is strictly better).
|
||||
openShell()
|
||||
fetch('/_api/settings').then((r) => r.json()).then((s) => {
|
||||
if (!pagePrimary.value) pagePrimary.value = s.primary_lang || 'en'
|
||||
}).catch(() => { /* keep the fallback */ })
|
||||
})
|
||||
|
||||
onUnmounted(() => {
|
||||
removeEventListener('pagerite:switch-editor', onSwitchEvent)
|
||||
removeEventListener('pagerite:editor-shown', openShell)
|
||||
removeEventListener('pagerite:editor-hidden', unpinPreviewLang)
|
||||
})
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<div class="editor-root overlay">
|
||||
<div class="editor-root overlay" lang="en" dir="ltr">
|
||||
<header class="editor-tabs">
|
||||
<template v-for="m in MODES" :key="m.key">
|
||||
<span v-if="m.breakBefore" class="tab-break" />
|
||||
|
||||
@@ -0,0 +1,168 @@
|
||||
<script setup>
|
||||
// The editor shell's one language selector (page + structure tabs): a small
|
||||
// flag button opening a clean dropdown, v-modeled on the shared editorLang
|
||||
// ('' = the primary language). The lang tab's flag grid is a different
|
||||
// control (toggles, not a select) and stays as it is.
|
||||
import { computed, nextTick, ref } from 'vue'
|
||||
import { usePopup } from './dropdown'
|
||||
|
||||
const props = defineProps({
|
||||
modelValue: { type: String, default: '' },
|
||||
options: { type: Array, required: true }, // [{tag, code, name, flag, primary}]
|
||||
title: { type: String, default: '' }, // toggle-button tooltip override
|
||||
})
|
||||
const emit = defineEmits(['update:modelValue'])
|
||||
|
||||
const open = ref(false)
|
||||
const root = ref(null)
|
||||
const toggleBtn = ref(null)
|
||||
const pop = ref(null)
|
||||
const popStyle = ref({})
|
||||
// Closes on outside click / Escape (./dropdown), not on mouseleave.
|
||||
usePopup(open, root)
|
||||
const current = computed(
|
||||
() => props.options.find((o) => o.tag === props.modelValue) ?? props.options[0],
|
||||
)
|
||||
|
||||
function toggle() {
|
||||
open.value = !open.value
|
||||
if (open.value) {
|
||||
// Position: fixed so the popup overflows the scrolling editor panel
|
||||
// onto the page area instead of being clipped by it.
|
||||
const r = toggleBtn.value.getBoundingClientRect()
|
||||
popStyle.value = { top: `${r.bottom + 2}px`, left: `${r.left}px` }
|
||||
// A toggle mounted near the right window edge (the public page
|
||||
// selector sits top-right) opens the popup flush against that edge.
|
||||
nextTick(() => {
|
||||
const p = pop.value?.getBoundingClientRect()
|
||||
if (p && p.right > innerWidth - 4) {
|
||||
popStyle.value = {
|
||||
...popStyle.value,
|
||||
left: `${Math.max(4, innerWidth - 4 - p.width)}px`,
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
function select(tag) {
|
||||
emit('update:modelValue', tag)
|
||||
open.value = false
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<span v-if="options.length > 1" ref="root" class="lang-select">
|
||||
<button
|
||||
ref="toggleBtn"
|
||||
type="button"
|
||||
class="lang-current"
|
||||
:class="{ open }"
|
||||
:title="title || (current
|
||||
? `language: ${current.name}${current.primary ? ' (primary)' : ''}`
|
||||
: '')"
|
||||
@click="toggle"
|
||||
><span v-if="current?.flag" class="flag" v-html="current.flag" /></button>
|
||||
<span v-if="open" ref="pop" class="lang-pop" :style="popStyle">
|
||||
<button
|
||||
v-for="o in options"
|
||||
:key="o.code"
|
||||
type="button"
|
||||
:class="{ active: o.tag === modelValue }"
|
||||
:title="o.primary ? `${o.name} — the primary language` : `${o.name} — translation`"
|
||||
@click="select(o.tag)"
|
||||
><span v-if="o.flag" class="flag" v-html="o.flag" /> {{ o.name }}<small v-if="o.primary"> (primary)</small></button>
|
||||
</span>
|
||||
</span>
|
||||
</template>
|
||||
|
||||
<style scoped>
|
||||
.lang-select {
|
||||
position: relative;
|
||||
display: flex;
|
||||
}
|
||||
|
||||
/* The closed state is just the small flag — no button chrome at all, on
|
||||
hover either (it sits among borderless emoji-icon buttons); like them it
|
||||
rests dimmed and brightens on hover. */
|
||||
.lang-current {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
padding: 2px;
|
||||
background: none;
|
||||
border: none;
|
||||
border-radius: 4px;
|
||||
cursor: pointer;
|
||||
opacity: 0.7;
|
||||
}
|
||||
|
||||
.lang-current:hover,
|
||||
.lang-current.open {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
/* The dropdown matches the page's existing popups (.picker-pop look).
|
||||
Fixed-positioned (anchored to the toggle's viewport rect on open) so it
|
||||
is not clipped by the editor panel's scrolling overflow. */
|
||||
.lang-pop {
|
||||
position: fixed;
|
||||
z-index: 20;
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
align-items: stretch;
|
||||
gap: 0.15rem;
|
||||
padding: 0.3rem;
|
||||
background: var(--bg);
|
||||
border: 1px solid var(--line);
|
||||
border-radius: 6px;
|
||||
box-shadow: 0 4px 16px #0004;
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
.lang-pop button {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.4rem;
|
||||
padding: 0.15rem 0.4rem;
|
||||
font: inherit;
|
||||
font-size: 0.9rem;
|
||||
text-align: left;
|
||||
color: var(--text);
|
||||
background: none;
|
||||
border: none;
|
||||
border-radius: 4px;
|
||||
cursor: pointer;
|
||||
}
|
||||
|
||||
.lang-pop button:hover {
|
||||
background: var(--surface);
|
||||
}
|
||||
|
||||
.lang-pop button.active {
|
||||
color: var(--accent);
|
||||
}
|
||||
|
||||
.lang-pop small {
|
||||
color: var(--muted);
|
||||
}
|
||||
|
||||
/* em-sized so the chip matches the surrounding text/icon size in each
|
||||
context; the hairline border delineates white-flagged countries (not
|
||||
button chrome). */
|
||||
.flag {
|
||||
display: inline-flex;
|
||||
width: 1.5em;
|
||||
height: 1em;
|
||||
flex: 0 0 auto;
|
||||
border-radius: 2px;
|
||||
overflow: hidden;
|
||||
border: 1px solid var(--line);
|
||||
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
|
||||
}
|
||||
|
||||
.flag :deep(svg) {
|
||||
width: 100%;
|
||||
height: 100%;
|
||||
display: block;
|
||||
}
|
||||
</style>
|
||||
@@ -0,0 +1,46 @@
|
||||
<script setup>
|
||||
// The public page's language selector: the editors' flag dropdown
|
||||
// (LangSelect) as the first item of the banner's corner container, fed
|
||||
// from the shared store (pagerite.js sets the page's hreflang alternates
|
||||
// and served language per navigation). It binds the same store.lang the
|
||||
// editor's dropdown binds, so both always show the same selection. A pick
|
||||
// also dispatches pagerite:set-session-lang — pagerite.js swaps the page
|
||||
// in place when the editor is closed (open, the editor reacts to the
|
||||
// store and re-renders it).
|
||||
import { computed } from 'vue'
|
||||
import LangSelect from './LangSelect.vue'
|
||||
import { flagFor, langName, langSort } from './langs'
|
||||
import { useStore } from './store'
|
||||
|
||||
const store = useStore()
|
||||
|
||||
// The "(primary)" marker is admin-panel information; the public selector
|
||||
// lists plain languages. Order: the primary language first, then the rest
|
||||
// in the lang tab's geographic grouping (./langs langSort) — the head's
|
||||
// hreflang order is just alphabetical.
|
||||
const primaryTag = computed(() => store.langAlternates.find((a) => a.primary)?.tag ?? '')
|
||||
const options = computed(() => {
|
||||
const rest = langSort(
|
||||
store.langAlternates.map((a) => a.tag).filter((t) => t !== primaryTag.value),
|
||||
)
|
||||
return [primaryTag.value, ...rest].filter(Boolean).map((tag) => ({
|
||||
tag,
|
||||
code: tag,
|
||||
name: langName(tag),
|
||||
flag: flagFor(tag),
|
||||
primary: false,
|
||||
}))
|
||||
})
|
||||
// The explicit pick, else the served language (header-autodetected pages
|
||||
// may have neither), else the primary.
|
||||
const model = computed(() => store.lang || store.servedLang || primaryTag.value)
|
||||
|
||||
function go(tag) {
|
||||
store.lang = tag === primaryTag.value ? '' : tag
|
||||
dispatchEvent(new CustomEvent('pagerite:set-session-lang', { detail: { lang: tag } }))
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<LangSelect :model-value="model" :options="options" @update:model-value="go" />
|
||||
</template>
|
||||
@@ -0,0 +1,410 @@
|
||||
<script setup>
|
||||
// Lang tab: the site-wide translation target languages (translate_langs)
|
||||
// and the translator service keys (translate_keys) with their WebSocket
|
||||
// URLs. ALL languages are listed, English included — a page whose primary
|
||||
// language (Node.language, configured per row in the structure tab,
|
||||
// inherited down the hierarchy) differs can be translated INTO any other.
|
||||
// Flag clicks toggle and save immediately; the settings round-trip
|
||||
// re-reads the payload, so this tab only ever changes translate_langs. The
|
||||
// settings write's invalidation hook kicks the translation dispatcher. The
|
||||
// refresh button drops all machine translations (user patches are kept),
|
||||
// making the dispatcher re-translate everything. Translator keys are
|
||||
// managed inline (➕ add, name edit, ✕ delete); new keys are generated
|
||||
// here in the server's format and everything rides the settings
|
||||
// round-trip. Clicking a key copies its full URL (following ws:// would
|
||||
// fail).
|
||||
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
|
||||
import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs'
|
||||
import { copyList } from './analytics/format.js'
|
||||
import { dropPageCache } from './swapdoc'
|
||||
|
||||
defineProps({ pagePath: { type: String, default: '' } })
|
||||
// close/path-change are wired by EditorShell; this tab never emits them.
|
||||
defineEmits(['close', 'pathChange'])
|
||||
|
||||
const saveError = ref('')
|
||||
const selected = ref(new Set())
|
||||
const keyUrls = ref([])
|
||||
|
||||
// Full WebSocket URL for a key. New keys are generated right here: 12
|
||||
// lowercase alphanumerics, the server-side format (state._KEY_ALPHABET).
|
||||
const wsUrl = (key) =>
|
||||
`${location.origin.replace(/^http/, 'ws')}/_translate/${key}`
|
||||
const KEY_ALPHABET = 'abcdefghijklmnopqrstuvwxyz0123456789'
|
||||
const newKey = () =>
|
||||
[...crypto.getRandomValues(new Uint8Array(12))]
|
||||
.map((b) => KEY_ALPHABET[b % KEY_ALPHABET.length])
|
||||
.join('')
|
||||
|
||||
// The toggleable targets: every translatable language, laid out in
|
||||
// geographic/cultural groups (one row each) rather than alphabetized —
|
||||
// related languages sit together (a node's own primary is excluded per
|
||||
// article, server-side). Any code missing from LANG_GROUPS trails as an
|
||||
// extra row.
|
||||
const groups = computed(() => {
|
||||
const tile = (code) => ({ code, name: langName(code), flag: flagFor(code) })
|
||||
const rows = LANG_GROUPS.map((g) => g.filter((c) => c in TRANSLATABLE).map(tile))
|
||||
const covered = new Set(LANG_GROUPS.flat())
|
||||
const rest = Object.keys(TRANSLATABLE).filter((c) => !covered.has(c)).map(tile)
|
||||
if (rest.length) rows.push(rest)
|
||||
return rows.filter((r) => r.length)
|
||||
})
|
||||
|
||||
function updateWindowTitle() {
|
||||
document.title = 'lang 🖊️'
|
||||
}
|
||||
|
||||
onActivated(updateWindowTitle)
|
||||
|
||||
// The shell stays mounted while hidden: when it is re-shown with this tab
|
||||
// active, restore the window title.
|
||||
function onEditorShown() {
|
||||
if (document.body.dataset.editorMode === 'localization') updateWindowTitle()
|
||||
}
|
||||
|
||||
onMounted(async () => {
|
||||
addEventListener('pagerite:editor-shown', onEditorShown)
|
||||
try {
|
||||
const s = await (await fetch('/_api/settings')).json()
|
||||
selected.value = new Set(s.translate_langs || [])
|
||||
keyUrls.value = Object.entries(s.translate_keys || {})
|
||||
.map(([key, name]) => ({ key, name, url: wsUrl(key) }))
|
||||
} catch { /* keep defaults */ }
|
||||
})
|
||||
|
||||
onUnmounted(() => removeEventListener('pagerite:editor-shown', onEditorShown))
|
||||
|
||||
async function toggle(code) {
|
||||
const next = new Set(selected.value)
|
||||
if (next.has(code)) next.delete(code)
|
||||
else next.add(code)
|
||||
selected.value = next
|
||||
try {
|
||||
const s = await (await fetch('/_api/settings')).json()
|
||||
const res = await fetch('/_api/settings', {
|
||||
method: 'PUT',
|
||||
headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ ...s, translate_langs: [...next] }),
|
||||
})
|
||||
if (res.ok) {
|
||||
saveError.value = ''
|
||||
dropPageCache()
|
||||
} else {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
} catch {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
}
|
||||
|
||||
// Delete all machine translations server-side; the dispatcher re-fills
|
||||
// them (a connected translator starts getting jobs right away). User
|
||||
// patches survive — they are edits, not machine output.
|
||||
const refreshing = ref(false)
|
||||
async function refresh() {
|
||||
if (refreshing.value) return
|
||||
refreshing.value = true
|
||||
try {
|
||||
const res = await fetch('/_api/translations', { method: 'DELETE' })
|
||||
saveError.value = res.ok ? '' : '⚠️ translations could not be refreshed'
|
||||
if (res.ok) dropPageCache()
|
||||
} catch {
|
||||
saveError.value = '⚠️ translations could not be refreshed'
|
||||
} finally {
|
||||
refreshing.value = false
|
||||
}
|
||||
}
|
||||
|
||||
// Key management rides the settings round-trip, like toggle() above:
|
||||
// mutate keyUrls, then PUT the whole settings payload with the new
|
||||
// translate_keys. ➕ adds a fresh unnamed key, names save on every
|
||||
// keystroke (@input — spamming the server is fine), ✕ deletes without
|
||||
// confirmation.
|
||||
async function saveKeys() {
|
||||
try {
|
||||
const s = await (await fetch('/_api/settings')).json()
|
||||
const res = await fetch('/_api/settings', {
|
||||
method: 'PUT',
|
||||
headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
...s,
|
||||
translate_keys: Object.fromEntries(keyUrls.value.map((k) => [k.key, k.name])),
|
||||
}),
|
||||
})
|
||||
saveError.value = res.ok ? '' : '⚠️ changes could not be saved'
|
||||
} catch {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
}
|
||||
|
||||
function addKey() {
|
||||
const key = newKey()
|
||||
keyUrls.value.push({ key, name: '', url: wsUrl(key) })
|
||||
saveKeys()
|
||||
}
|
||||
|
||||
function removeKey(k) {
|
||||
keyUrls.value = keyUrls.value.filter((x) => x.key !== k.key)
|
||||
saveKeys()
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<div class="localization-editor">
|
||||
<div v-if="saveError">{{ saveError }}</div>
|
||||
|
||||
<section class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">languages</span>
|
||||
</div>
|
||||
<div class="flags">
|
||||
<div v-for="(row, ri) in groups" :key="ri" class="flag-row">
|
||||
<button
|
||||
v-for="o in row"
|
||||
:key="o.code"
|
||||
type="button"
|
||||
class="flag-tile"
|
||||
:class="{ selected: selected.has(o.code) }"
|
||||
:title="`${o.name} (${o.code})`"
|
||||
@click="toggle(o.code)"
|
||||
>
|
||||
<span class="flag" v-html="o.flag" />
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
<section class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">Translator API</span>
|
||||
</div>
|
||||
<div v-for="k in keyUrls" :key="k.key" class="key-row">
|
||||
<a
|
||||
:href="k.url"
|
||||
class="key-link"
|
||||
title="click to copy the URL"
|
||||
@click.prevent="copyList(k.url, $event)"
|
||||
>{{ k.key }}</a>
|
||||
<input
|
||||
v-model="k.name"
|
||||
type="text"
|
||||
class="edit key-name"
|
||||
title="display name"
|
||||
@input="saveKeys()"
|
||||
>
|
||||
<button type="button" class="act del" title="delete key" @click="removeKey(k)">✕</button>
|
||||
</div>
|
||||
<div class="add-row">
|
||||
<button type="button" class="add" title="new translator key" @click="addKey()">➕ API key</button>
|
||||
</div>
|
||||
<p><small class="muted">AI translator agents can connect with the API keys to do machine translations to your selected languages. Click the button below to delete all translations and start over. User edits are kept.</small></p>
|
||||
<div class="refresh-row">
|
||||
<button
|
||||
type="button"
|
||||
class="refresh-btn"
|
||||
:disabled="refreshing"
|
||||
@click="refresh"
|
||||
>
|
||||
{{ refreshing ? 'Reseting…' : 'Reset' }}
|
||||
</button>
|
||||
</div>
|
||||
</section>
|
||||
</div>
|
||||
</template>
|
||||
|
||||
<style scoped>
|
||||
.localization-editor {
|
||||
overflow-y: auto;
|
||||
background: var(--surface);
|
||||
}
|
||||
|
||||
.block {
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
gap: 0.4rem;
|
||||
padding: 0.5rem 1rem;
|
||||
border-bottom: 1px solid var(--line);
|
||||
background: var(--surface);
|
||||
}
|
||||
|
||||
.block-head {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.6rem;
|
||||
}
|
||||
|
||||
.field-label {
|
||||
color: var(--muted);
|
||||
font-size: 0.85rem;
|
||||
}
|
||||
|
||||
.muted {
|
||||
color: var(--muted);
|
||||
}
|
||||
|
||||
/* Flag grid: one geographic group per row. Deselected flags sit dimmed and
|
||||
grayed; a click brings one to full color (selected = a translation
|
||||
target) — the shading alone carries the state, no outline. */
|
||||
.flags {
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
gap: 0.4rem;
|
||||
padding: 0.2rem 0;
|
||||
}
|
||||
|
||||
.flag-row {
|
||||
display: flex;
|
||||
flex-wrap: wrap;
|
||||
gap: 0.5rem;
|
||||
}
|
||||
|
||||
.flag-tile {
|
||||
padding: 3px;
|
||||
background: none;
|
||||
border: 2px solid transparent;
|
||||
border-radius: 5px;
|
||||
cursor: pointer;
|
||||
opacity: 0.4;
|
||||
filter: grayscale(0.8);
|
||||
transition: opacity 0.15s, filter 0.15s, border-color 0.15s;
|
||||
}
|
||||
|
||||
.flag-tile:hover {
|
||||
opacity: 0.8;
|
||||
filter: none;
|
||||
}
|
||||
|
||||
.flag-tile.selected {
|
||||
opacity: 1;
|
||||
filter: none;
|
||||
}
|
||||
|
||||
/* Same flag chips as the PageEditor language picker / analytics cells. */
|
||||
.flag {
|
||||
display: inline-flex;
|
||||
width: 18px;
|
||||
height: 12px;
|
||||
flex: 0 0 auto;
|
||||
border-radius: 2px;
|
||||
overflow: hidden;
|
||||
border: 1px solid var(--line);
|
||||
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
|
||||
}
|
||||
|
||||
.flag-tile .flag {
|
||||
width: 36px;
|
||||
height: 24px;
|
||||
}
|
||||
|
||||
.flag :deep(svg) {
|
||||
width: 100%;
|
||||
height: 100%;
|
||||
display: block;
|
||||
}
|
||||
|
||||
.key-row {
|
||||
display: flex;
|
||||
align-items: baseline;
|
||||
gap: 0.6rem;
|
||||
}
|
||||
|
||||
/* Real links (handy for right-click/drag) showing just the key, but the
|
||||
click copies the full URL instead of following — ws:// would fail to
|
||||
navigate. Normal text color, not link-styled; position: relative
|
||||
anchors the "Copied!" popup (analytics/format.js). */
|
||||
.key-link {
|
||||
position: relative;
|
||||
color: var(--text);
|
||||
font-family: var(--font-code);
|
||||
user-select: all;
|
||||
}
|
||||
|
||||
.refresh-row {
|
||||
display: flex;
|
||||
align-items: baseline;
|
||||
gap: 0.6rem;
|
||||
}
|
||||
|
||||
/* Name input / ✕ / ➕ follow the structure tab's conventions: inputs stay
|
||||
borderless until interacted with, glyph buttons redden / solidify on
|
||||
hover. */
|
||||
.key-name {
|
||||
flex: 0 0 9rem;
|
||||
}
|
||||
|
||||
.edit {
|
||||
font: inherit;
|
||||
font-size: 0.85rem;
|
||||
padding: 0.1rem 0.4rem;
|
||||
background: transparent;
|
||||
color: var(--text);
|
||||
border: 1px solid transparent;
|
||||
border-radius: 4px;
|
||||
min-width: 0;
|
||||
}
|
||||
|
||||
.edit:hover {
|
||||
border-color: var(--line);
|
||||
}
|
||||
|
||||
.edit:focus {
|
||||
background: var(--bg);
|
||||
border-color: var(--accent);
|
||||
outline: none;
|
||||
}
|
||||
|
||||
.act {
|
||||
padding: 0 0.25rem;
|
||||
background: none;
|
||||
border: none;
|
||||
color: var(--muted);
|
||||
font-size: 0.8rem;
|
||||
cursor: pointer;
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
.del:hover {
|
||||
color: #e06c75;
|
||||
}
|
||||
|
||||
.add-row {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
}
|
||||
|
||||
.add {
|
||||
padding: 0 0.3rem;
|
||||
background: none;
|
||||
border: none;
|
||||
font-size: 0.9rem;
|
||||
cursor: pointer;
|
||||
opacity: 0.5;
|
||||
}
|
||||
|
||||
.add:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
.refresh-btn {
|
||||
align-self: flex-start;
|
||||
margin-bottom: 0.2rem;
|
||||
padding: 0.3rem 0.8rem;
|
||||
font: inherit;
|
||||
font-size: 0.85rem;
|
||||
color: var(--muted);
|
||||
background: none;
|
||||
border: 1px solid var(--line);
|
||||
border-radius: 5px;
|
||||
cursor: pointer;
|
||||
}
|
||||
|
||||
.refresh-btn:hover:not(:disabled) {
|
||||
color: var(--text);
|
||||
border-color: var(--muted);
|
||||
}
|
||||
|
||||
.refresh-btn:disabled {
|
||||
opacity: 0.5;
|
||||
cursor: default;
|
||||
}
|
||||
</style>
|
||||
+272
-49
@@ -11,15 +11,33 @@
|
||||
// and refreshes the page regions in place — never a reload — so the editor
|
||||
// state (unsaved text included) also survives closing the shell. The editor
|
||||
// always follows the URL: navigating away retargets it to the new page,
|
||||
// stashing unsaved text per path (unsavedStash) so returning to the page
|
||||
// restores the working draft; stashes clear on save and on real reload.
|
||||
import { onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
|
||||
// stashing unsaved text per path and language (stashes) so returning to the
|
||||
// page restores the working draft; stashes clear on save and on real reload.
|
||||
//
|
||||
// Languages: the editor always starts in the primary language; the
|
||||
// toolbar's LangSelect (shared with the structure tab via ./editorLang)
|
||||
// switches between the primary language and its translations. A translation is
|
||||
// edited as its effective (hybrid) Markdown; the hybrid the session
|
||||
// started from is kept as a shadow copy (shadowBase) and sent along at
|
||||
// save time, so the server diffs the user's changes only and stores them
|
||||
// as a patch — edits to a translation never touch the original, while
|
||||
// edits to the primary language re-chunk the original (and thereby
|
||||
// invalidate the affected translation fragments). The live preview always
|
||||
// renders the version being edited, whichever language the page itself
|
||||
// was loaded in.
|
||||
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
|
||||
import { usePopup } from './dropdown'
|
||||
import { EditorView, basicSetup } from 'codemirror'
|
||||
import { EditorState } from '@codemirror/state'
|
||||
import { Compartment, EditorState } from '@codemirror/state'
|
||||
import { keymap } from '@codemirror/view'
|
||||
import { indentWithTab } from '@codemirror/commands'
|
||||
import { markdown } from '@codemirror/lang-markdown'
|
||||
import { cmHighlight, cmTheme } from './cmtheme'
|
||||
import { flagFor, langName, langSort } from './langs'
|
||||
import { editorLang, pagePrimary } from './editorLang'
|
||||
import LangSelect from './LangSelect.vue'
|
||||
import ConnNote from './ConnNote.vue'
|
||||
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
|
||||
import { dropPageCache, loadPlain } from './swapdoc'
|
||||
|
||||
const props = defineProps({
|
||||
@@ -34,14 +52,49 @@ const saveError = ref('')
|
||||
const editorEl = ref(null)
|
||||
const fileInput = ref(null)
|
||||
|
||||
// The language being edited: "" = the primary language, where the editor
|
||||
// always starts (a served translation does not follow it into the editor;
|
||||
// the picker switches). The server normalizes the primary to "" anyway.
|
||||
// Shared with the other tabs (./editorLang) — one selection for the whole
|
||||
// shell, and for the page preview while the shell is open.
|
||||
const lang = editorLang
|
||||
const primaryLang = ref('en')
|
||||
const pageLangs = ref([]) // translations this page has
|
||||
const siteLangs = ref([]) // site-wide configured target languages
|
||||
// The shadow copy: the Markdown this editing session started from, sent as
|
||||
// "base" on translated saves so the server diffs the user's changes only.
|
||||
let shadowBase = ''
|
||||
let titleTouched = false
|
||||
// The (path, lang) the editor's current content came from: set when a doc
|
||||
// is accepted, and the test for whether a reconnected socket must re-open
|
||||
// (a dropped open would otherwise leave the editor empty/stale forever).
|
||||
let sessionDoc = null
|
||||
|
||||
// Connection state drives the note above the editor (ConnNote), and locks
|
||||
// input until the page's document has arrived: typing before the accept
|
||||
// would be clobbered by it. While merely DISconnected the editor stays
|
||||
// editable — text stashes and pending saves flush on reconnect.
|
||||
const conn = ref('connecting') // connecting | open | waiting
|
||||
const retryIn = ref(0)
|
||||
const docReady = ref(false)
|
||||
const editable = new Compartment()
|
||||
const connNote = computed(() =>
|
||||
conn.value === 'connecting' ? 'connecting to the server…'
|
||||
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
|
||||
: docReady.value ? '' : 'loading the page…',
|
||||
)
|
||||
|
||||
let ws = null
|
||||
let view = null
|
||||
let savedResolve = null
|
||||
let pendingSave = null
|
||||
let reconnectTimer = null
|
||||
let reconnectDelay = 2000
|
||||
const MAX_RECONNECT_DELAY = 16000
|
||||
let everConnected = false
|
||||
let connectWatchdog = null
|
||||
// Reconnection pacing lives in ./reconnect (shared with the other sockets):
|
||||
// doubling backoff with jitter, reset only by a healthy connection — rapid
|
||||
// retries trip the browser's WebSocket throttling (sockets stuck "pending"
|
||||
// for minutes), which is what kept hollowing out the editor.
|
||||
const reconnects = reconnectPolicy()
|
||||
const dirty = ref(false) // unsaved text exists (drives the 💾 button)
|
||||
let syncingScroll = false
|
||||
|
||||
@@ -63,6 +116,41 @@ function normPath(p) {
|
||||
return p.trim().replace(/^\/+|\/+$/g, '')
|
||||
}
|
||||
|
||||
// --- Language picker -------------------------------------------------------
|
||||
// Flag icons and display names come from ./langs (shared with the
|
||||
// localization settings tab).
|
||||
|
||||
// The picker's options: the primary language first, then the union of the
|
||||
// page's translations and the site-wide configured targets in the lang
|
||||
// tab's geographic grouping (./langs langSort).
|
||||
const langOptions = computed(() => {
|
||||
const others = langSort(
|
||||
[...new Set([...siteLangs.value, ...pageLangs.value])]
|
||||
.filter((l) => l && l !== primaryLang.value),
|
||||
)
|
||||
return [primaryLang.value, ...others].map((code) => ({
|
||||
tag: code === primaryLang.value ? '' : code,
|
||||
code,
|
||||
name: langName(code),
|
||||
flag: flagFor(code),
|
||||
primary: code === primaryLang.value,
|
||||
}))
|
||||
})
|
||||
|
||||
const currentLang = computed(
|
||||
() => langOptions.value.find((o) => o.tag === lang.value)
|
||||
?? { tag: '', code: lang.value || primaryLang.value, name: langName(lang.value || primaryLang.value), flag: flagFor(lang.value || primaryLang.value), primary: !lang.value },
|
||||
)
|
||||
|
||||
// The picker's selection is the shell-wide shared language (./editorLang):
|
||||
// a change stashes the working text under the PREVIOUS language (the view
|
||||
// still holds that doc) and opens the current page in the new one.
|
||||
watch(lang, (tag, prev) => {
|
||||
if (!view) return
|
||||
stashAs(prev || '')
|
||||
openPath(path.value)
|
||||
})
|
||||
|
||||
function pageLabel() {
|
||||
return title.value.trim() || ('/' + (path.value || ''))
|
||||
}
|
||||
@@ -96,11 +184,17 @@ function save() {
|
||||
// never moves the page.
|
||||
const markdown = view.state.doc.toString()
|
||||
if (markdown.trim() === '') {
|
||||
if (lang.value) {
|
||||
// Emptying a translation would render it as a blank page; deleting
|
||||
// pages is a primary-language action.
|
||||
saveError.value = '⚠️ a translation cannot be emptied'
|
||||
return Promise.resolve()
|
||||
}
|
||||
// Empty text means delete — an explicit choice made here, in the page
|
||||
// editor; the save APIs (REST PUT / WS save) never delete on empty.
|
||||
return fetch(`/_api/pages/${path.value}`, { method: 'DELETE' }).then((res) => {
|
||||
saveError.value = res.ok ? '' : '⚠️ changes could not be saved'
|
||||
if (res.ok) unsavedStash.delete(path.value)
|
||||
if (res.ok) stashes.delete(stashKey(path.value, lang.value))
|
||||
})
|
||||
}
|
||||
const msg = {
|
||||
@@ -110,6 +204,15 @@ function save() {
|
||||
markdown,
|
||||
published: published.value,
|
||||
}
|
||||
if (lang.value) {
|
||||
msg.lang = lang.value
|
||||
// The shadow copy this session started from: the server diffs base →
|
||||
// markdown and stores only the user's changes as a patch.
|
||||
msg.base = shadowBase
|
||||
// An untouched title field is not sent: it holds the served
|
||||
// translation, which a save must not freeze into an override fragment.
|
||||
if (!titleTouched) delete msg.title
|
||||
}
|
||||
pendingSave = msg
|
||||
send(msg)
|
||||
return new Promise((resolve) => { savedResolve = resolve })
|
||||
@@ -120,9 +223,12 @@ async function saveAndRefresh() {
|
||||
dirty.value = false
|
||||
// Refresh the page regions from the server so nav/sidebar changes apply
|
||||
// (never a reload: the editor keeps its state). Drop the prefetch cache
|
||||
// first: heading/title changes affect navigation on every page.
|
||||
// first: heading/title changes affect navigation on every page. A
|
||||
// translation save keeps the preview as-is — it already shows the saved
|
||||
// text, and loadPlain would swap the article to the header-selected
|
||||
// language's render.
|
||||
dropPageCache()
|
||||
loadPlain(path.value)
|
||||
if (!lang.value) loadPlain(path.value)
|
||||
}
|
||||
|
||||
function close() {
|
||||
@@ -518,8 +624,15 @@ const TABLE_MAX_ROWS = 6
|
||||
// Class pickers: popup listing the block class toggles (placement ↔︎,
|
||||
// text size AA), closed after applying. The block's current class of the
|
||||
// group is marked; choosing "normal" (or the current class) removes it.
|
||||
// All popups share the close behavior of ./dropdown (outside click /
|
||||
// Escape; never mouseleave).
|
||||
const classPicker = ref(null) // 'place' | 'size' | null
|
||||
const activeClasses = ref(new Set())
|
||||
const placeRoot = ref(null)
|
||||
const sizeRoot = ref(null)
|
||||
const tableRoot = ref(null)
|
||||
usePopup(classPicker, computed(() => (classPicker.value === 'place' ? placeRoot : sizeRoot).value))
|
||||
usePopup(tablePicker, tableRoot)
|
||||
|
||||
function openClassPicker(which) {
|
||||
classPicker.value = classPicker.value === which ? null : which
|
||||
@@ -572,18 +685,34 @@ function insertTable(cols, rows) {
|
||||
view.focus()
|
||||
}
|
||||
|
||||
// Unsaved edits survive navigation within the session: leaving a page
|
||||
// stashes its working text here, returning restores it (the server doc
|
||||
// still arrives, for title/published and as the base underneath).
|
||||
// Entries clear on save and on real reload (the shell is in-memory only).
|
||||
const unsavedStash = new Map()
|
||||
// Unsaved edits survive navigation within the session: leaving a page (or
|
||||
// switching the language) stashes its working text and shadow base here,
|
||||
// returning restores them (the server doc still arrives, for
|
||||
// title/published and language metadata). Entries clear on save and on
|
||||
// real reload (the shell is in-memory only).
|
||||
const stashes = new Map()
|
||||
const stashKey = (p, l) => `${p}|${l}`
|
||||
|
||||
function stashAs(l) {
|
||||
if (dirty.value && path.value) {
|
||||
stashes.set(stashKey(path.value, l), {
|
||||
text: view.state.doc.toString(),
|
||||
base: shadowBase,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
function stashCurrent() {
|
||||
stashAs(lang.value)
|
||||
}
|
||||
|
||||
function openPath(p) {
|
||||
if (dirty.value && path.value && p !== path.value) {
|
||||
unsavedStash.set(path.value, view.state.doc.toString())
|
||||
}
|
||||
if (p !== path.value) stashCurrent()
|
||||
path.value = p
|
||||
send({ type: 'open', path: p })
|
||||
// Lock input until the doc arrives (typing would be clobbered by it).
|
||||
docReady.value = false
|
||||
view?.dispatch({ effects: editable.reconfigure(EditorView.editable.of(false)) })
|
||||
send({ type: 'open', path: p, lang: lang.value })
|
||||
}
|
||||
|
||||
function setDocument(text, preserveSelection = false) {
|
||||
@@ -624,26 +753,57 @@ function previewIntoArticle(html, multicol) {
|
||||
|
||||
function onMessage(ev) {
|
||||
const msg = JSON.parse(ev.data)
|
||||
if (msg.type === 'doc' && msg.path === path.value) {
|
||||
// The doc must answer the current path and language; before the first
|
||||
// doc the language is not yet server-normalized (the primary arrives as
|
||||
// "" while the picker may have started from an explicit code), so the
|
||||
// first doc is accepted on path alone and adopts the echoed language.
|
||||
if (msg.type === 'doc' && msg.path === path.value
|
||||
&& (!sessionDoc || (msg.lang || '') === lang.value)) {
|
||||
title.value = msg.title
|
||||
published.value = msg.published
|
||||
primaryLang.value = msg.primary_lang || 'en'
|
||||
pagePrimary.value = primaryLang.value // the page's own — the shell pins the preview by it
|
||||
pageLangs.value = msg.langs || []
|
||||
siteLangs.value = msg.translate_langs || []
|
||||
lang.value = msg.lang || ''
|
||||
sessionDoc = { path: msg.path, lang: lang.value }
|
||||
docReady.value = true
|
||||
view.dispatch({ effects: editable.reconfigure(EditorView.editable.of(true)) })
|
||||
titleTouched = false
|
||||
// Restore stashed unsaved edits over the server doc when returning
|
||||
// to a page left dirty.
|
||||
const stashed = unsavedStash.get(msg.path)
|
||||
setDocument(stashed ?? msg.markdown)
|
||||
const stashed = stashes.get(stashKey(msg.path, lang.value))
|
||||
setDocument(stashed ? stashed.text : msg.markdown)
|
||||
shadowBase = stashed ? stashed.base : msg.markdown
|
||||
dirty.value = stashed != null
|
||||
requestRender()
|
||||
// A section pen's target line survives the open/path-switch here.
|
||||
consumePendingLine()
|
||||
} else if (msg.type === 'doc') {
|
||||
// A doc that answered neither path nor language of the current session
|
||||
// (a late reply to a pre-switch open) — visible because it leaves the
|
||||
// editor empty when it's the only doc that ever arrives.
|
||||
console.warn(
|
||||
'[pagerite] doc dropped:', msg.path, msg.lang || '(primary)',
|
||||
'— editor is on', path.value, lang.value || '(primary)',
|
||||
)
|
||||
} else if (msg.type === 'html' && msg.path === path.value) {
|
||||
previewIntoArticle(msg.html, msg.multicol)
|
||||
} else if (msg.type === 'saved') {
|
||||
saveError.value = ''
|
||||
pendingSave = null
|
||||
dirty.value = false
|
||||
unsavedStash.delete(path.value)
|
||||
savedResolve?.()
|
||||
savedResolve = null
|
||||
// A late "saved" for a save sent before a path/language switch must not
|
||||
// clear the current session's state.
|
||||
if (pendingSave && pendingSave.path === path.value
|
||||
&& (pendingSave.lang || '') === lang.value) {
|
||||
saveError.value = ''
|
||||
// The saved text becomes the shadow base for further saves.
|
||||
shadowBase = pendingSave.markdown
|
||||
pendingSave = null
|
||||
dirty.value = false
|
||||
titleTouched = false
|
||||
stashes.delete(stashKey(path.value, lang.value))
|
||||
savedResolve?.()
|
||||
savedResolve = null
|
||||
}
|
||||
} else if (msg.type === 'error') {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
@@ -831,33 +991,54 @@ function consumePendingLine() {
|
||||
}
|
||||
|
||||
function connect() {
|
||||
clearTimeout(reconnectTimer)
|
||||
conn.value = 'connecting'
|
||||
if (ws) {
|
||||
// Replacing a stale socket: detach its handlers so its close is silent.
|
||||
ws.onopen = ws.onmessage = ws.onclose = ws.onerror = null
|
||||
if (ws.readyState !== WebSocket.CLOSED) ws.close()
|
||||
}
|
||||
ws = new WebSocket(
|
||||
`${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`,
|
||||
)
|
||||
ws.onmessage = onMessage
|
||||
ws.onerror = (ev) => {
|
||||
console.error('[pagerite] editor socket error', ev)
|
||||
}
|
||||
clearTimeout(connectWatchdog)
|
||||
connectWatchdog = watchConnecting(ws, 'editor')
|
||||
ws.onopen = () => {
|
||||
reconnectDelay = 2000
|
||||
if (everConnected) {
|
||||
conn.value = 'open'
|
||||
reconnects.opened()
|
||||
if (sessionDoc && sessionDoc.path === path.value && sessionDoc.lang === lang.value) {
|
||||
// Reconnected: local text is authoritative — don't re-open (that
|
||||
// would clobber the editor), just resync preview and pending saves.
|
||||
requestRender()
|
||||
if (pendingSave) send(pendingSave)
|
||||
} else {
|
||||
// No doc behind the current page+language (first connect, or its
|
||||
// open went down with a previous socket): open fresh, or the editor
|
||||
// would stay empty forever.
|
||||
openPath(normPath(props.pagePath))
|
||||
}
|
||||
everConnected = true
|
||||
}
|
||||
ws.onclose = () => {
|
||||
clearTimeout(reconnectTimer)
|
||||
reconnectTimer = setTimeout(() => {
|
||||
connect()
|
||||
reconnectDelay = Math.min(reconnectDelay * 2, MAX_RECONNECT_DELAY)
|
||||
}, reconnectDelay)
|
||||
ws.onclose = (ev) => {
|
||||
// 1006 = abnormal (e.g. the dev proxy refused/dropped the upgrade, or
|
||||
// the connecting watchdog fired); worth seeing since a dead socket
|
||||
// before the first doc bricks the editor until a retry lands one.
|
||||
// The wait is the policy's: doubling backoff with jitter (./reconnect).
|
||||
console.warn('[pagerite] editor socket closed:', ev.code, ev.reason || '')
|
||||
const wait = reconnects.closed()
|
||||
retryIn.value = Math.max(1, Math.round(wait / 1000))
|
||||
conn.value = 'waiting'
|
||||
reconnectTimer = setTimeout(connect, wait)
|
||||
}
|
||||
}
|
||||
|
||||
onMounted(() => {
|
||||
connect()
|
||||
// The first connection takes a staggered slot (see ./reconnect): page
|
||||
// load opens several sockets at once, and the burst trips throttling.
|
||||
reconnectTimer = setTimeout(connect, socketSlot())
|
||||
updateWindowTitle()
|
||||
|
||||
view = new EditorView({
|
||||
@@ -871,6 +1052,8 @@ onMounted(() => {
|
||||
cmTheme,
|
||||
cmHighlight,
|
||||
EditorView.lineWrapping, // Markdown lines are long: soft-wrap them
|
||||
// Locked until the page's document arrives (docReady/ConnNote).
|
||||
editable.of(EditorView.editable.of(false)),
|
||||
EditorView.updateListener.of((u) => {
|
||||
if (u.docChanged) requestRender()
|
||||
// Cursor moves (typing included) drive the page scroll sync.
|
||||
@@ -912,6 +1095,7 @@ onMounted(() => {
|
||||
|
||||
onUnmounted(() => {
|
||||
clearTimeout(reconnectTimer)
|
||||
clearTimeout(connectWatchdog)
|
||||
if (ws) {
|
||||
ws.onclose = null // intentional close, no reconnect
|
||||
ws.close()
|
||||
@@ -929,11 +1113,12 @@ onUnmounted(() => {
|
||||
<template>
|
||||
<div class="page-editor">
|
||||
<header class="toolbar">
|
||||
<LangSelect v-model="lang" :options="langOptions" />
|
||||
<label class="title-field">
|
||||
<span class="field-label">title</span>
|
||||
<input v-model="title" class="title" @input="requestRender" />
|
||||
<input v-model="title" class="title" @input="requestRender(); titleTouched = true" />
|
||||
</label>
|
||||
<label><input v-model="published" type="checkbox" /> published</label>
|
||||
<label :title="lang ? 'translations follow the original page’s published state' : ''"><input v-model="published" type="checkbox" :disabled="!!lang" /> published</label>
|
||||
<input
|
||||
ref="fileInput"
|
||||
type="file"
|
||||
@@ -949,18 +1134,43 @@ onUnmounted(() => {
|
||||
@click="saveAndRefresh"
|
||||
>💾</button>
|
||||
</header>
|
||||
<ConnNote :text="connNote" />
|
||||
<div v-if="langOptions.length > 1" class="lang-note">
|
||||
<template v-if="lang">
|
||||
{{ currentLang.name }} translation — edits affect only this language.
|
||||
</template>
|
||||
<template v-else>
|
||||
{{ currentLang.name }} is the primary language — edits here affect all translations.
|
||||
</template>
|
||||
</div>
|
||||
<div class="format-bar">
|
||||
<button type="button" class="code-btn" title="code — inline wrap, or a fenced block for line-spanning selections; click again to unwrap" @click="insertCode"><code></></code></button>
|
||||
<button type="button" title="link (toggle: click inside a link to unwrap it)" @click="insertLink">🔗︎</button>
|
||||
<button
|
||||
type="button"
|
||||
title="table"
|
||||
:class="{ active: tablePicker }"
|
||||
@click="tablePicker = !tablePicker"
|
||||
>⊞</button>
|
||||
<span class="picker" ref="tableRoot">
|
||||
<button
|
||||
type="button"
|
||||
title="table"
|
||||
:class="{ active: tablePicker }"
|
||||
@click="tablePicker = !tablePicker"
|
||||
>⊞</button>
|
||||
<div v-if="tablePicker" class="table-picker" @mouseleave="tableSize = { cols: 0, rows: 0 }">
|
||||
<div class="tp-grid" :style="{ gridTemplateColumns: `repeat(${TABLE_MAX_COLS}, 1fr)` }">
|
||||
<button
|
||||
v-for="n in TABLE_MAX_COLS * TABLE_MAX_ROWS"
|
||||
:key="n"
|
||||
type="button"
|
||||
class="tp-cell"
|
||||
:class="{ on: tableSize.cols >= (n - 1) % TABLE_MAX_COLS + 1 && tableSize.rows >= Math.floor((n - 1) / TABLE_MAX_COLS) + 1 }"
|
||||
@mouseenter="tableSize = { cols: (n - 1) % TABLE_MAX_COLS + 1, rows: Math.floor((n - 1) / TABLE_MAX_COLS) + 1 }"
|
||||
@click="insertTable(tableSize.cols, tableSize.rows)"
|
||||
/>
|
||||
</div>
|
||||
<div class="tp-size">{{ tableSize.cols || '–' }} × {{ tableSize.rows || '–' }}</div>
|
||||
</div>
|
||||
</span>
|
||||
<button type="button" title="insert image (upload) — pasting works too" @click="fileInput.click()">🖼︎</button>
|
||||
<button type="button" title="aside box (::: aside) — wraps the selection or the cursor's line; clicked inside one, removes it" @click="insertAside">◧</button>
|
||||
<span class="picker">
|
||||
<span class="picker" ref="placeRoot">
|
||||
<button
|
||||
type="button"
|
||||
title="block placement class"
|
||||
@@ -982,7 +1192,7 @@ onUnmounted(() => {
|
||||
</span>
|
||||
<button type="button" title="bold" @click="wrapInline('**')"><b>B</b></button>
|
||||
<button type="button" title="italic" @click="wrapInline('*')"><i>i</i></button>
|
||||
<span class="picker">
|
||||
<span class="picker" ref="sizeRoot">
|
||||
<button
|
||||
type="button"
|
||||
title="text size class"
|
||||
@@ -1089,6 +1299,19 @@ onUnmounted(() => {
|
||||
cursor: default;
|
||||
}
|
||||
|
||||
/* The language selector (toolbar, left) is LangSelect.vue — its styles
|
||||
live there. */
|
||||
|
||||
/* Why a language is highlighted: primary edits fan out to translations,
|
||||
translation edits stay local to that language. */
|
||||
.lang-note {
|
||||
padding: 0.2rem 1rem;
|
||||
border-bottom: 1px solid var(--line);
|
||||
background: var(--surface);
|
||||
color: var(--muted);
|
||||
font-size: 0.8rem;
|
||||
}
|
||||
|
||||
/* Markdown helpers: plain icon buttons under the toolbar. */
|
||||
.format-bar {
|
||||
position: relative;
|
||||
@@ -1171,7 +1394,7 @@ onUnmounted(() => {
|
||||
.table-picker {
|
||||
position: absolute;
|
||||
top: 100%;
|
||||
left: 6.5rem;
|
||||
left: 0;
|
||||
z-index: 20;
|
||||
padding: 0.5rem;
|
||||
background: var(--bg);
|
||||
|
||||
@@ -782,14 +782,6 @@ onUnmounted(() => {
|
||||
margin-left: auto;
|
||||
padding: 0 0.2rem;
|
||||
font-size: 1rem;
|
||||
background: none;
|
||||
border: none;
|
||||
cursor: pointer;
|
||||
opacity: 0.7;
|
||||
}
|
||||
|
||||
.block-head .icon-btn:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
.text-input {
|
||||
|
||||
@@ -7,9 +7,20 @@
|
||||
// real — a label with a title and slug, with content (landing page) or
|
||||
// without (category whose URL renders a placeholder page). The front page
|
||||
// is a top-level row with an empty slug, not the parent of the others.
|
||||
import { inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
|
||||
//
|
||||
// Languages: the LangSelect switches which language the TITLES are shown
|
||||
// and edited in (rows without a translation show the original, dimmed) —
|
||||
// the selection is shared shell-wide (./editorLang) with the page editor
|
||||
// and the page preview. Translated title edits write a per-language
|
||||
// fragment (POST /_api/structure with lang); the structure itself —
|
||||
// slugs, order, hierarchy — is language-independent and always edits the
|
||||
// same tree.
|
||||
import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
|
||||
import StructureTree from './StructureTree.vue'
|
||||
import LangSelect from './LangSelect.vue'
|
||||
import { slugify } from './slugify'
|
||||
import { flagFor, langName, langSort } from './langs'
|
||||
import { editorLang, pagePrimary } from './editorLang'
|
||||
import { dropPageCache, loadPlain } from './swapdoc'
|
||||
|
||||
const props = defineProps({
|
||||
@@ -23,6 +34,51 @@ const path = ref('')
|
||||
const saveError = ref('')
|
||||
const tree = ref([])
|
||||
|
||||
// The language the tree's titles are shown and edited in: "" = primary.
|
||||
const lang = editorLang
|
||||
const primaryLang = ref('en')
|
||||
const siteLangs = ref([])
|
||||
|
||||
// The strip's options: the primary language first, then the configured
|
||||
// translation targets (the lang tab manages that set) in the lang tab's
|
||||
// geographic grouping (./langs langSort).
|
||||
const langOptions = computed(() =>
|
||||
[primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
|
||||
.map((code) => ({
|
||||
tag: code === primaryLang.value ? '' : code,
|
||||
code,
|
||||
name: langName(code),
|
||||
flag: flagFor(code),
|
||||
primary: code === primaryLang.value,
|
||||
})),
|
||||
)
|
||||
const currentLang = computed(
|
||||
() => langOptions.value.find((o) => o.tag === lang.value) ?? langOptions.value[0],
|
||||
)
|
||||
|
||||
// The selection is shared (./editorLang): a change re-fetches the tree's
|
||||
// titles in it (and EditorShell swaps the page preview into it).
|
||||
watch(lang, () => refreshPages())
|
||||
|
||||
// Per-row primary language (Node.language, '' = inherit): the row's
|
||||
// dropdown lists "inherit" first (naming what it resolves to), then every
|
||||
// site language. Setting it on a section covers its whole subtree.
|
||||
const rowLangChoices = computed(() =>
|
||||
[primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
|
||||
.map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })),
|
||||
)
|
||||
function rowLangOptions(el) {
|
||||
const resolved = el.primary || primaryLang.value
|
||||
return [
|
||||
{ tag: '', code: '_inherit', name: `inherit (${langName(resolved)})`, flag: flagFor(resolved), primary: false },
|
||||
...rowLangChoices.value,
|
||||
]
|
||||
}
|
||||
|
||||
async function setLanguage(node, tag) {
|
||||
await postStructure({ path: node.path, language: tag })
|
||||
}
|
||||
|
||||
function normPath(p) {
|
||||
return p.trim().replace(/^\/+|\/+$/g, '')
|
||||
}
|
||||
@@ -152,9 +208,22 @@ async function commitPending() {
|
||||
}
|
||||
|
||||
// --- Site structure tree (drag-and-drop ordering/moving) ----------------
|
||||
function findNode(nodes, p) {
|
||||
for (const n of nodes) {
|
||||
if (n.path === p) return n
|
||||
const found = findNode(n.children, p)
|
||||
if (found) return found
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
async function refreshPages() {
|
||||
try {
|
||||
tree.value = await (await fetch('/_api/pages')).json()
|
||||
const q = lang.value ? `?lang=${lang.value}` : ''
|
||||
tree.value = await (await fetch(`/_api/pages${q}`)).json()
|
||||
// The tree carries each node's resolved primary language: publish the
|
||||
// current page's (the shell pins the preview by it on '' selection).
|
||||
pagePrimary.value = findNode(tree.value, path.value)?.primary || 'en'
|
||||
} catch { /* list stays stale; not fatal */ }
|
||||
}
|
||||
|
||||
@@ -207,13 +276,15 @@ async function onReorder(parentPath, list, evt) {
|
||||
}
|
||||
|
||||
// Inline title/slug editing: rows are always editable. Title saves while
|
||||
// typing (debounced); the slug commits on blur/Enter, since it renames
|
||||
// the path (moving the whole subtree with it).
|
||||
// typing (debounced) — in the selected language (a translation writes a
|
||||
// title fragment, the primary language the original); the slug commits on
|
||||
// blur/Enter, since it renames the path (moving the whole subtree with it).
|
||||
// Slugs are language-independent.
|
||||
function onTitleInput(node, ev) {
|
||||
const title = ev.target.value.trim()
|
||||
if (!title || title === node.title) return
|
||||
debounce(`title:${node.path}`, async () => {
|
||||
await postStructure({ path: node.path, title })
|
||||
await postStructure({ path: node.path, title, lang: lang.value })
|
||||
})
|
||||
}
|
||||
|
||||
@@ -266,12 +337,19 @@ provide('structureHandlers', {
|
||||
commitPending,
|
||||
discardPending,
|
||||
newPage,
|
||||
langOptions: rowLangOptions,
|
||||
setLanguage,
|
||||
})
|
||||
|
||||
onMounted(() => {
|
||||
path.value = normPath(props.pagePath)
|
||||
refreshPages()
|
||||
addEventListener('pagerite:editor-shown', onEditorShown)
|
||||
// The language strip: site primary + configured targets.
|
||||
fetch('/_api/settings').then((r) => r.json()).then((s) => {
|
||||
primaryLang.value = s.primary_lang || 'en'
|
||||
siteLangs.value = s.translate_langs || []
|
||||
}).catch(() => { /* no strip */ })
|
||||
})
|
||||
|
||||
onUnmounted(() => {
|
||||
@@ -283,8 +361,15 @@ onUnmounted(() => {
|
||||
<template>
|
||||
<div class="structure-editor">
|
||||
<div v-if="saveError">{{ saveError }}</div>
|
||||
<div v-if="langOptions.length > 1" class="block lang-block">
|
||||
<div><LangSelect v-model="lang" :options="langOptions" /></div>
|
||||
<small v-if="lang" class="muted">
|
||||
viewing {{ currentLang.name }} titles — dimmed rows are untranslated
|
||||
(shown in the primary language); slugs never translate
|
||||
</small>
|
||||
</div>
|
||||
<section class="block structure">
|
||||
<StructureTree :nodes="tree" />
|
||||
<StructureTree :nodes="tree" :lang="lang" />
|
||||
</section>
|
||||
</div>
|
||||
</template>
|
||||
@@ -309,4 +394,10 @@ onUnmounted(() => {
|
||||
overflow-y: auto;
|
||||
min-height: 0;
|
||||
}
|
||||
|
||||
/* The language selector is LangSelect.vue — its styles live there. */
|
||||
|
||||
.muted {
|
||||
color: var(--muted);
|
||||
}
|
||||
</style>
|
||||
|
||||
@@ -1,7 +1,13 @@
|
||||
<script setup>
|
||||
// Recursive site-structure tree with drag-and-drop ordering (vue-draggable).
|
||||
// Nodes come from the server (GET /_api/pages via StructureEditor.vue) as
|
||||
// {slug, path, title, order, published, has_content, children}.
|
||||
// {slug, path, title, translated, order, published, has_content, language,
|
||||
// primary, children}. The row's flag (LangSelect) sets the node's primary
|
||||
// language (language; '' = inherit — dimmed, showing the resolved flag);
|
||||
// the setting covers the whole subtree.
|
||||
// With a `lang` prop (StructureEditor's language strip) the titles shown
|
||||
// are that language's; `translated` marks rows with an actual translation
|
||||
// (untranslated rows show the original title, dimmed).
|
||||
// Every node is real: a label whose title and slug are always editable
|
||||
// inline — the title saves while typing (and focusing it opens the page),
|
||||
// the slug commits on blur/Enter since it renames the path, moving the
|
||||
@@ -24,12 +30,17 @@
|
||||
import { inject } from 'vue'
|
||||
import draggable from 'vuedraggable'
|
||||
import { slugify } from './slugify'
|
||||
import LangSelect from './LangSelect.vue'
|
||||
|
||||
defineOptions({ name: 'StructureTree' })
|
||||
const props = defineProps({
|
||||
nodes: { type: Array, required: true },
|
||||
parentPath: { type: String, default: '' },
|
||||
depth: { type: Number, default: 0 },
|
||||
// StructureEditor's selected language ('' = original). Only used for the
|
||||
// untranslated-title styling here; the fetch and title edits live in the
|
||||
// parent (handlers.titleInput posts the lang with the op).
|
||||
lang: { type: String, default: '' },
|
||||
})
|
||||
|
||||
const handlers = inject('structureHandlers')
|
||||
@@ -126,8 +137,11 @@ function onEnd() {
|
||||
<template v-else>
|
||||
<input
|
||||
class="edit title-edit"
|
||||
:class="{ untranslated: lang && !element.translated }"
|
||||
:value="element.title"
|
||||
title="Label in the navigation — saves while typing; click opens the page"
|
||||
:title="lang && !element.translated
|
||||
? 'No translation yet — showing the original; typing creates the translated title'
|
||||
: 'Label in the navigation — saves while typing; click opens the page'"
|
||||
@input="handlers.titleInput(element, $event)"
|
||||
@focus="handlers.open(element.path)"
|
||||
/>
|
||||
@@ -140,6 +154,17 @@ function onEnd() {
|
||||
@change="handlers.commitSlug(element, $event)"
|
||||
/>
|
||||
<span class="acts">
|
||||
<span
|
||||
class="row-lang"
|
||||
:class="{ inherited: !element.language }"
|
||||
><LangSelect
|
||||
:model-value="element.language"
|
||||
:options="handlers.langOptions(element)"
|
||||
:title="element.language
|
||||
? `primary language: set on this page (subtree inherits)`
|
||||
: `primary language: inherited — set it here (subtree inherits)`"
|
||||
@update:model-value="handlers.setLanguage(element, $event)"
|
||||
/></span>
|
||||
<span v-if="!element.published" class="draft">draft</span>
|
||||
<button
|
||||
v-if="element.has_content || !element.children.length"
|
||||
@@ -158,6 +183,7 @@ function onEnd() {
|
||||
:nodes="element.children"
|
||||
:parent-path="element.path"
|
||||
:depth="depth + 1"
|
||||
:lang="lang"
|
||||
/>
|
||||
</div>
|
||||
</template>
|
||||
@@ -223,7 +249,7 @@ body.tree-dragging .treelist {
|
||||
level, not across levels). */
|
||||
.row {
|
||||
display: grid;
|
||||
grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem 5rem;
|
||||
grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem auto;
|
||||
align-items: baseline;
|
||||
gap: 0.35rem;
|
||||
/* Vertical spacing widens the drop zones: the exposed top strip is the
|
||||
@@ -274,6 +300,13 @@ body.tree-dragging .treelist {
|
||||
cursor: text;
|
||||
}
|
||||
|
||||
/* With a language selected (StructureEditor's strip), rows without an
|
||||
actual translation show the original title dimmed and italic. */
|
||||
.title-edit.untranslated {
|
||||
color: var(--muted);
|
||||
font-style: italic;
|
||||
}
|
||||
|
||||
.slug-edit {
|
||||
font-family: var(--font-code);
|
||||
}
|
||||
@@ -285,6 +318,20 @@ body.tree-dragging .treelist {
|
||||
justify-content: end;
|
||||
}
|
||||
|
||||
/* Row language selector (LangSelect): the effective primary language's
|
||||
flag; dimmed while the setting is inherited rather than set on the row. */
|
||||
.row-lang {
|
||||
display: inline-flex;
|
||||
}
|
||||
|
||||
.row-lang.inherited :deep(.lang-current) {
|
||||
opacity: 0.45;
|
||||
}
|
||||
|
||||
.row-lang.inherited:hover :deep(.lang-current) {
|
||||
opacity: 0.85;
|
||||
}
|
||||
|
||||
.draft {
|
||||
color: var(--muted);
|
||||
font-size: 0.75rem;
|
||||
|
||||
@@ -4,15 +4,20 @@
|
||||
// Clicking the IP copies the full address to the clipboard.
|
||||
// ``variantCount`` overrides the UA line to warn when multiple client
|
||||
// fingerprints share the same IP (e.g. a scanner rotating UAs).
|
||||
// Clicking the UA line copies the raw UA(s) to the clipboard, one per line
|
||||
// (``uaRaws`` carries every variation for multi-client IPs).
|
||||
import { computed } from 'vue'
|
||||
import * as flagSvgs from 'country-flag-icons/string/3x2'
|
||||
import { copyIp, formatLang } from './analytics/format.js'
|
||||
import { copyIp, copyList, formatLang } from './analytics/format.js'
|
||||
import { langName } from './langs.js'
|
||||
|
||||
const props = defineProps({
|
||||
ip: { type: String, default: '' },
|
||||
ipDisplay: { type: String, default: '—' },
|
||||
ua: { type: String, default: '' },
|
||||
uaRaw: { type: String, default: '' },
|
||||
uaRaws: { type: String, default: '' },
|
||||
uaUrl: { type: String, default: '' },
|
||||
country: { type: String, default: '' },
|
||||
city: { type: String, default: '' },
|
||||
lang: { type: String, default: '' },
|
||||
@@ -26,6 +31,7 @@ const hasCity = computed(() => !!(props.city && props.city !== '—'))
|
||||
const hasLocale = computed(() => hasCountry.value || hasCity.value)
|
||||
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
|
||||
const showLang = computed(() => langValue.value && langValue.value !== '—')
|
||||
const uaCopy = computed(() => props.uaRaws || props.uaRaw)
|
||||
|
||||
function flagSvg(code) {
|
||||
return flagSvgs[code?.toUpperCase()] || ''
|
||||
@@ -58,10 +64,16 @@ function countryName(code) {
|
||||
</div>
|
||||
<div class="visitor-row">
|
||||
<div class="ua-line">
|
||||
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
|
||||
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
|
||||
<small v-if="variantCount > 1" class="muted variant-hint clickable-ip"
|
||||
:title="uaCopy"
|
||||
@click="copyList(uaCopy, $event)">{{ variantCount }} client variations</small>
|
||||
<small v-else class="muted clickable-ip" :title="uaRaw"
|
||||
@click="copyList(uaCopy, $event)">{{ ua || '—' }}</small><a v-if="uaUrl && variantCount <= 1"
|
||||
class="ua-link icon-btn" :href="uaUrl"
|
||||
target="_blank" rel="noopener noreferrer"
|
||||
@click.stop>🔗</a>
|
||||
</div>
|
||||
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
|
||||
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted" :title="langName(lang)">{{ langValue }}</small></div>
|
||||
</div>
|
||||
</div>
|
||||
</td>
|
||||
@@ -121,6 +133,13 @@ function countryName(code) {
|
||||
text-align: left;
|
||||
}
|
||||
|
||||
.ua-link {
|
||||
text-decoration: none;
|
||||
font-size: 0.75em;
|
||||
margin-left: 0.2em;
|
||||
vertical-align: middle;
|
||||
}
|
||||
|
||||
.locale-lang {
|
||||
flex: 0 0 auto;
|
||||
overflow: hidden;
|
||||
|
||||
@@ -75,25 +75,33 @@ const viewChart = computed(() => buildChart(viewSeries.value, now.value))
|
||||
<svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`"
|
||||
:style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }"
|
||||
role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)">
|
||||
<!-- Clip the plot curves to the chart area: past-week overlays can
|
||||
run far above the autoscaled y range, and the svg itself is
|
||||
overflow: visible for the axis labels. -->
|
||||
<clipPath :id="`plot-${c.ylabel}`">
|
||||
<rect x="0" y="0" :width="CHART_W" :height="CHART_H" />
|
||||
</clipPath>
|
||||
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
|
||||
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
|
||||
<template v-for="t in c.chart.xticks" :key="'t' + t.x">
|
||||
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
|
||||
class="minor vertical" />
|
||||
</template>
|
||||
<template v-if="c.chart.bars">
|
||||
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
|
||||
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
|
||||
<path :d="c.chart.skyline" class="line" />
|
||||
</template>
|
||||
<template v-else>
|
||||
<!-- Oldest overlay weeks first so the current week paints on top. -->
|
||||
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
|
||||
<path v-if="s.area" :d="s.area" class="area" />
|
||||
<path :d="s.line" class="line" :class="{ past: s.past }"
|
||||
:style="{ opacity: s.opacity }" />
|
||||
<g :clip-path="`url(#plot-${c.ylabel})`">
|
||||
<template v-if="c.chart.bars">
|
||||
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
|
||||
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
|
||||
<path :d="c.chart.skyline" class="line" />
|
||||
</template>
|
||||
</template>
|
||||
<template v-else>
|
||||
<!-- Oldest overlay weeks first so the current week paints on top. -->
|
||||
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
|
||||
<path v-if="s.area" :d="s.area" class="area" />
|
||||
<path :d="s.line" class="line" :class="{ past: s.past }"
|
||||
:style="{ opacity: s.opacity }" />
|
||||
</template>
|
||||
</template>
|
||||
</g>
|
||||
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
|
||||
class="axis" />
|
||||
<text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y"
|
||||
|
||||
@@ -25,32 +25,31 @@ export const hostIP = (ip) => {
|
||||
}
|
||||
}
|
||||
|
||||
function showCopiedFeedback(el) {
|
||||
if (!el || typeof document === 'undefined') return
|
||||
function showCopiedFeedback(el, event) {
|
||||
if (typeof document === 'undefined') return
|
||||
const popup = document.createElement('span')
|
||||
popup.textContent = 'Copied!'
|
||||
popup.className = 'copy-popup'
|
||||
// Fixed to the viewport at the click point: table cells clip absolute
|
||||
// popups with their overflow: hidden ellipsis styling.
|
||||
const x = event?.clientX ?? 0
|
||||
const y = event?.clientY ?? 0
|
||||
popup.style.cssText =
|
||||
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
|
||||
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
|
||||
`position:fixed;left:${x}px;top:${y}px;` +
|
||||
'transform:translate(-50%, calc(-100% - 0.5rem));padding:0.15rem 0.4rem;' +
|
||||
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
|
||||
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
|
||||
'pointer-events:none;z-index:10;'
|
||||
el.classList.add('has-copy-popup')
|
||||
el.appendChild(popup)
|
||||
setTimeout(() => {
|
||||
popup.remove()
|
||||
el.classList.remove('has-copy-popup')
|
||||
}, 1200)
|
||||
'pointer-events:none;z-index:100;'
|
||||
document.body.appendChild(popup)
|
||||
setTimeout(() => popup.remove(), 1200)
|
||||
}
|
||||
|
||||
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
|
||||
export async function copyIp(ip, event) {
|
||||
if (!ip) return
|
||||
const el = event?.currentTarget
|
||||
try {
|
||||
await navigator.clipboard.writeText(ip)
|
||||
showCopiedFeedback(el)
|
||||
showCopiedFeedback(event?.currentTarget, event)
|
||||
} catch {
|
||||
/* ignore */
|
||||
}
|
||||
@@ -59,10 +58,9 @@ export async function copyIp(ip, event) {
|
||||
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
|
||||
export async function copyList(text, event) {
|
||||
if (!text) return
|
||||
const el = event?.currentTarget
|
||||
try {
|
||||
await navigator.clipboard.writeText(text)
|
||||
showCopiedFeedback(el)
|
||||
showCopiedFeedback(event?.currentTarget, event)
|
||||
} catch {
|
||||
/* ignore */
|
||||
}
|
||||
@@ -342,7 +340,7 @@ export function countCrawlerUas(crawlers, clients) {
|
||||
const counts = {}
|
||||
for (const c of crawlers || []) {
|
||||
const client = (clients || {})[c.client] || {}
|
||||
const value = client.ua_pretty || client.ua || '(no UA)'
|
||||
const value = client.uarite?.pretty || client.ua || '(no UA)'
|
||||
counts[value] = (counts[value] || 0) + 1
|
||||
}
|
||||
return Object.entries(counts).sort((a, b) => b[1] - a[1])
|
||||
@@ -370,7 +368,9 @@ export function mainDomain(host, limit = 24) {
|
||||
/**
|
||||
* Group raw crawler hits by client hash and format each group as a row showing
|
||||
* every internal page that crawler visited. Rows are sorted by most recent hit
|
||||
* first, with total hits as a tie-breaker.
|
||||
* first, with total hits as a tie-breaker. The group's ``refererStep`` is the
|
||||
* latest external referer seen for the crawler — spiders often advertise
|
||||
* their own site there — rendered with its favicon like visit referers.
|
||||
* ``clients`` maps client hashes to client records.
|
||||
*/
|
||||
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
|
||||
@@ -382,10 +382,12 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
|
||||
clientHash: c.client,
|
||||
client,
|
||||
lastStart: 0,
|
||||
referer: '',
|
||||
pages: new Map(),
|
||||
}
|
||||
const start = new Date(c.start).getTime()
|
||||
if (start > g.lastStart) g.lastStart = start
|
||||
if (c.referer) g.referer = c.referer
|
||||
if (c.entry?.startsWith('/')) {
|
||||
const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
|
||||
existing.count += 1
|
||||
@@ -410,14 +412,16 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
|
||||
lastSeen: formatWhen(g.lastStart, now),
|
||||
lastSeenIso: formatWhenIso(g.lastStart),
|
||||
lastSeenLocal: formatWhenLocal(g.lastStart),
|
||||
refererStep: stepOf(g.referer, titles),
|
||||
pages: [...g.pages.entries()]
|
||||
.sort((a, b) => b[1].count - a[1].count)
|
||||
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
|
||||
ip: client.ip || '',
|
||||
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
|
||||
isHost,
|
||||
ua: client.ua_pretty || client.ua || '—',
|
||||
ua: client.uarite?.pretty || client.ua || '—',
|
||||
uaRaw: client.ua || '',
|
||||
uaUrl: client.uarite?.url || '',
|
||||
lang: client.lang || '—',
|
||||
langDisplay: formatLang(client.lang),
|
||||
country: client.country || '—',
|
||||
@@ -429,17 +433,22 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
|
||||
|
||||
/**
|
||||
* Group abuse hits by IP and format each group as a row with the full paths
|
||||
* probed. Identical paths are collapsed into one entry with their hit count.
|
||||
* Flagged paths (the ones that triggered abuse classification) are lifted to
|
||||
* the top, followed by other 404s, then document GETs from the abuser. Within
|
||||
* each category paths are sorted by count descending, then earliest first.
|
||||
* probed. Identical requests (same path and status class) are collapsed
|
||||
* into one entry with their hit count; a path's 404 probes and its real
|
||||
* (200) reads never merge.
|
||||
* The paths split into two lists: ``paths`` holds the 404 probes (flagged
|
||||
* paths — the ones that triggered abuse classification — first, then other
|
||||
* 404s) shown verbatim, query string included, and ``articles`` holds the
|
||||
* real (200) document GETs as trail steps resolved against the page tree
|
||||
* (query string stripped), rendered like the visitor/crawler trails. Within
|
||||
* each list paths are sorted by count descending, then earliest first.
|
||||
* Rows are sorted by most recent hit first. Visitor metadata comes from the
|
||||
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
|
||||
* how many distinct client variations the IP produced. Paths are shown
|
||||
* verbatim (query string included), not resolved against the page tree.
|
||||
* how many distinct client variations the IP produced.
|
||||
* ``clients`` maps client hashes to client records.
|
||||
*/
|
||||
export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||
export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
|
||||
const titles = buildTitleMap(pageTree)
|
||||
const groups = new Map()
|
||||
for (const a of abuse || []) {
|
||||
const client = (clients || {})[a.client] || {}
|
||||
@@ -457,7 +466,11 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||
g.lastClient = a.client
|
||||
}
|
||||
const path = a.path || ''
|
||||
const existing = g.pathCounts.get(path) || {
|
||||
// Collapse identical requests, but never merge a path's 404 probes with
|
||||
// its real (200) reads — a page probed while missing and later created
|
||||
// must show up in both columns, not flip to "articles read".
|
||||
const key = `${a.is_404 ? '4' : '2'}${path}`
|
||||
const existing = g.pathCounts.get(key) || {
|
||||
path,
|
||||
count: 0,
|
||||
firstStart: start,
|
||||
@@ -467,8 +480,7 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||
existing.count += 1
|
||||
if (start < existing.firstStart) existing.firstStart = start
|
||||
if (a.flag) existing.flag = true
|
||||
if (!a.is_404) existing.is_404 = false
|
||||
g.pathCounts.set(path, existing)
|
||||
g.pathCounts.set(key, existing)
|
||||
g.clientHashes.add(a.client)
|
||||
groups.set(ip, g)
|
||||
}
|
||||
@@ -481,16 +493,24 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||
.sort((a, b) => b.lastStart - a.lastStart)
|
||||
.slice(0, 10)
|
||||
.map((g) => {
|
||||
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2)
|
||||
const paths = [...g.pathCounts.values()].sort(
|
||||
(a, b) =>
|
||||
pathCategory(a) - pathCategory(b) ||
|
||||
b.count - a.count ||
|
||||
a.firstStart - b.firstStart,
|
||||
)
|
||||
const all = [...g.pathCounts.values()]
|
||||
const byCount = (a, b) => b.count - a.count || a.firstStart - b.firstStart
|
||||
const paths = all
|
||||
.filter((p) => p.flag || p.is_404)
|
||||
.sort((a, b) => (a.flag ? 0 : 1) - (b.flag ? 0 : 1) || byCount(a, b))
|
||||
const articles = all.filter((p) => !p.flag && !p.is_404).sort(byCount)
|
||||
const pathList = (list) =>
|
||||
list.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path)).join('\n')
|
||||
const client = (clients || {})[g.lastClient] || {}
|
||||
const host = client.host || ''
|
||||
const isHost = !!host
|
||||
const uaRaws = [
|
||||
...new Set(
|
||||
[...g.clientHashes]
|
||||
.map((h) => (clients || {})[h]?.ua)
|
||||
.filter(Boolean),
|
||||
),
|
||||
].join('\n')
|
||||
return {
|
||||
lastSeen: formatWhen(g.lastStart, now),
|
||||
lastSeenIso: formatWhenIso(g.lastStart),
|
||||
@@ -501,15 +521,22 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||
flag: p.flag,
|
||||
is_404: p.is_404,
|
||||
})),
|
||||
allPaths: paths
|
||||
.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path))
|
||||
.join('\n'),
|
||||
allPaths: pathList(paths),
|
||||
articles: articles
|
||||
.map((p) => {
|
||||
const step = stepOf(p.path.split('?')[0], titles)
|
||||
return step ? { ...step, count: p.count } : null
|
||||
})
|
||||
.filter(Boolean),
|
||||
allArticles: pathList(articles),
|
||||
clientCount: g.clientHashes.size,
|
||||
ip: client.ip || g.ip,
|
||||
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
|
||||
isHost,
|
||||
ua: client.ua_pretty || client.ua || '—',
|
||||
ua: client.uarite?.pretty || client.ua || '—',
|
||||
uaRaw: client.ua || '',
|
||||
uaUrl: client.uarite?.url || '',
|
||||
uaRaws,
|
||||
lang: client.lang || '—',
|
||||
langDisplay: formatLang(client.lang),
|
||||
country: client.country || '—',
|
||||
@@ -563,8 +590,9 @@ export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
|
||||
lang: dash(client.lang),
|
||||
country: dash(client.country),
|
||||
city: dash(client.city),
|
||||
ua: client.ua_pretty || client.ua || '—',
|
||||
ua: client.uarite?.pretty || client.ua || '—',
|
||||
uaRaw: client.ua || '',
|
||||
uaUrl: client.uarite?.url || '',
|
||||
utm: utm || '—',
|
||||
utmTitle,
|
||||
}
|
||||
|
||||
+143
-123
@@ -288,8 +288,10 @@ body {
|
||||
|
||||
/* Sidebar + main row. A symmetric grid: the article column is sized by the
|
||||
viewport alone (never by content), with equally sized flexible gutters
|
||||
on both sides. The sidebar sits in the left gutter, so it appearing or
|
||||
disappearing never shifts the article; the right gutter balances it.
|
||||
on both sides. The sidebar sits in the start gutter (grid columns are
|
||||
flow-relative: on RTL pages the whole composition mirrors), so it
|
||||
appearing or disappearing never shifts the article; the end gutter
|
||||
balances it.
|
||||
The outer tracks are minmax(0, 1fr) — a plain 1fr has an `auto` minimum,
|
||||
which let the 12rem sidebar expand its track at narrow widths and push
|
||||
the article off-center; now the sidebar overlays the gutter edge instead
|
||||
@@ -305,16 +307,16 @@ body {
|
||||
content length — code excluded) lift the 78rem cap: main takes the full
|
||||
width and the article composes itself inside it — fluid, bounded text
|
||||
lanes with the surplus left vacant (see the article layout rules
|
||||
below). The left track — main always sits in column 2 — collapses to
|
||||
below). The start track — main always sits in column 2 — collapses to
|
||||
zero when the page has no sidebar; the sidebar then simply overlays
|
||||
the vacant zone, as it does on single-column pages. */
|
||||
body:has(.multicol) #content {
|
||||
grid-template-columns: 0 minmax(0, 1fr);
|
||||
}
|
||||
|
||||
/* With a sidebar the left lane gets its own track at every width, so the
|
||||
/* With a sidebar the start lane gets its own track at every width, so the
|
||||
sidebar never overlaps the article and the article leans on the
|
||||
viewport's right edge (surplus extends the lane). The lane is flexible:
|
||||
viewport's end edge (surplus extends the lane). The lane is flexible:
|
||||
12rem when space is tight, growing up to 150% (18rem) once the viewport
|
||||
exceeds the article's 88.5rem (86rem + main's side padding). The --lane
|
||||
variable doubles as the measure for the margin boxes and the .wide
|
||||
@@ -398,7 +400,7 @@ body.editing #sidebar {
|
||||
|
||||
#sidebar {
|
||||
grid-column: 1;
|
||||
/* Pinned to the page's left edge (not the article's) and kept in view
|
||||
/* Pinned to the page's start edge (not the article's) and kept in view
|
||||
while scrolling. Translucent + blurred rather than an opaque box, so
|
||||
full-bleed .wide images can pass underneath without a hard edge. */
|
||||
justify-self: start;
|
||||
@@ -409,8 +411,11 @@ body.editing #sidebar {
|
||||
width: 12rem;
|
||||
max-height: 100vh;
|
||||
overflow-y: auto;
|
||||
padding: 1rem 1rem 1rem 1.25rem;
|
||||
border-radius: 0 0 0.5rem 0;
|
||||
/* The extra inline-start padding (the viewport-edge side, mirroring with
|
||||
the direction) matches main's side padding. */
|
||||
padding: 1rem;
|
||||
padding-inline: 1.25rem 1rem;
|
||||
border-end-start-radius: 0.5rem;
|
||||
background: color-mix(var(--bg) 75%, transparent);
|
||||
backdrop-filter: blur(0.5rem);
|
||||
}
|
||||
@@ -437,7 +442,7 @@ body.editing #sidebar {
|
||||
#sidebar ul ul li::before {
|
||||
content: "🔹";
|
||||
display: inline-block;
|
||||
margin-left: -1.3em;
|
||||
margin-inline-start: -1.3em;
|
||||
width: 1.3em;
|
||||
}
|
||||
|
||||
@@ -578,7 +583,7 @@ main {
|
||||
background: color-mix(var(--bg) 30%, transparent);
|
||||
color: var(--muted);
|
||||
font-size: 0.95rem;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
hyphens: none;
|
||||
}
|
||||
|
||||
@@ -679,19 +684,19 @@ article * + h6 {
|
||||
so the gap becomes the figure's margin plus the full --list-indent. */
|
||||
article :is(ul, ol:not([type])) {
|
||||
list-style: none;
|
||||
padding-left: 0;
|
||||
padding-inline-start: 0;
|
||||
--list-indent: 2em;
|
||||
display: flow-root;
|
||||
}
|
||||
|
||||
article :is(ul, ol:not([type])) > li {
|
||||
padding-left: var(--list-indent);
|
||||
padding-inline-start: var(--list-indent);
|
||||
}
|
||||
|
||||
article ul li::before {
|
||||
content: "🔹";
|
||||
display: inline-block;
|
||||
margin-left: calc(-1 * var(--list-indent));
|
||||
margin-inline-start: calc(-1 * var(--list-indent));
|
||||
width: var(--list-indent);
|
||||
text-align: center;
|
||||
}
|
||||
@@ -713,13 +718,13 @@ article ol:not([type]) > li::before {
|
||||
content: counter(item) ".";
|
||||
color: var(--muted);
|
||||
display: inline-block;
|
||||
margin-left: calc(-1 * var(--list-indent));
|
||||
margin-inline-start: calc(-1 * var(--list-indent));
|
||||
width: var(--list-indent);
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
}
|
||||
|
||||
/* Task lists: real clickable checkboxes. The checkbox stands in for the
|
||||
list marker — taken out of flow, left-aligned in the indent box and
|
||||
list marker — taken out of flow, start-aligned in the indent box and
|
||||
centered on the first line's middle, so the item text starts at the
|
||||
same edge as every other list item's. */
|
||||
article .task-list-item {
|
||||
@@ -732,7 +737,7 @@ article .task-list-item::before {
|
||||
|
||||
article .task-list-item-checkbox {
|
||||
position: absolute;
|
||||
left: 0;
|
||||
inset-inline-start: 0;
|
||||
top: calc(0.5lh - .15ex);
|
||||
translate: 0 -50%;
|
||||
font-size: inherit;
|
||||
@@ -751,6 +756,20 @@ article {
|
||||
position: relative;
|
||||
}
|
||||
|
||||
/* Emoji/symbol icon buttons and links: dim until hovered. */
|
||||
.icon-btn {
|
||||
padding: 0;
|
||||
font: inherit;
|
||||
background: none;
|
||||
border: none;
|
||||
cursor: pointer;
|
||||
opacity: 0.7;
|
||||
}
|
||||
|
||||
.icon-btn:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
.edit-link {
|
||||
position: absolute;
|
||||
top: 0.2rem;
|
||||
@@ -758,12 +777,6 @@ article {
|
||||
left: -2.2rem;
|
||||
z-index: 2;
|
||||
/* stay above full-bleed .wide images */
|
||||
font: inherit;
|
||||
background: none;
|
||||
border: none;
|
||||
padding: 0;
|
||||
cursor: pointer;
|
||||
opacity: 0.7;
|
||||
text-shadow: 0 0 0.1em black;
|
||||
}
|
||||
|
||||
@@ -772,7 +785,7 @@ article h1 .edit-link {
|
||||
position: static;
|
||||
font-size: 1.1rem;
|
||||
vertical-align: 0.3em;
|
||||
margin-left: 0.4rem;
|
||||
margin-inline-start: 0.4rem;
|
||||
}
|
||||
|
||||
/* Section pens sit at the end of anchored h2s, dimmer than the page pen
|
||||
@@ -781,28 +794,18 @@ article h2 .edit-section {
|
||||
position: static;
|
||||
font-size: 0.85rem;
|
||||
vertical-align: 0.35em;
|
||||
margin-left: 0.4rem;
|
||||
margin-inline-start: 0.4rem;
|
||||
opacity: 0.35;
|
||||
}
|
||||
|
||||
.edit-link:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
/* Login/profile links injected by pagerite.js when Paskia SSO is in use.
|
||||
They live inside the .editor-pens flex container in the banner's top-right
|
||||
corner and inherit its reset; keep only their opacity/text-shadow tweaks. */
|
||||
corner and inherit its reset; keep only their text-shadow tweak. */
|
||||
.editor-pens a.login-link,
|
||||
.editor-pens a.profile-link {
|
||||
opacity: 0.7;
|
||||
text-shadow: 0 0 0.1em black;
|
||||
}
|
||||
|
||||
.editor-pens a.login-link:hover,
|
||||
.editor-pens a.profile-link:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
article p,
|
||||
article li,
|
||||
article dd {
|
||||
@@ -811,10 +814,10 @@ article dd {
|
||||
}
|
||||
|
||||
/* The long-article composition (.multicol): a fluid but bounded text
|
||||
lane with a 16rem side zone at the article's left, centered in main —
|
||||
surplus width becomes vacant space, never endless text (with a sidebar
|
||||
the article leans right instead and the sidebar's track is the left
|
||||
lane; see below). (The backend render splits the body into .colseg
|
||||
lane with a 16rem side zone at the article's start side, centered in
|
||||
main — surplus width becomes vacant space, never endless text (with a
|
||||
sidebar the article leans to the end edge instead and the sidebar's
|
||||
track is the start lane; see below). (The backend render splits the body into .colseg
|
||||
segments separated by full-width h2s and .wide elements, tags
|
||||
text-heavy segments of several paragraphs .cols — a ::: nocols
|
||||
container opts its section out, and column-filling paragraphs are
|
||||
@@ -832,7 +835,7 @@ article.multicol {
|
||||
@container (min-width: 45rem) {
|
||||
/* The side zone (not on phones): lane content indents 16rem; margin
|
||||
boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures)
|
||||
are taken out of flow and placed against the article's left edge —
|
||||
are taken out of flow and placed against the article's start edge —
|
||||
the same region the nav sidebar overlays. The boxes stay in the
|
||||
column segment at their anchor point (the backend render no longer
|
||||
splits segments around them); absolute positioning off the article —
|
||||
@@ -842,14 +845,14 @@ article.multicol {
|
||||
article.multicol>.colseg,
|
||||
article.multicol>h1,
|
||||
article.multicol>h2 {
|
||||
margin-left: 16rem;
|
||||
margin-inline-start: 16rem;
|
||||
}
|
||||
|
||||
article.multicol .margin,
|
||||
article.multicol .aside,
|
||||
article.multicol figure:has(.margin) {
|
||||
article.multicol figure.margin {
|
||||
position: absolute;
|
||||
left: 0;
|
||||
inset-inline-start: 0;
|
||||
width: 14rem;
|
||||
max-width: none;
|
||||
margin: 0.3rem 0 0;
|
||||
@@ -870,11 +873,11 @@ article.multicol {
|
||||
}
|
||||
}
|
||||
|
||||
/* With a sidebar, the sidebar's 12rem track IS the left lane at every
|
||||
/* With a sidebar, the sidebar's 12rem track IS the start lane at every
|
||||
width (see #content): no in-article zone, the text lane runs fluid (up
|
||||
to 86rem) and leans on main's right edge — surplus width extends the
|
||||
left lane instead of balancing out on the right — and margin boxes
|
||||
hang into the lane off the article's left border, sliding under the
|
||||
to 86rem) and leans on main's end edge — surplus width extends the
|
||||
start lane instead of balancing out at the end — and margin boxes
|
||||
hang into the lane off the article's start border, sliding under the
|
||||
translucent sticky nav, which only ever occupies its top. (Not below
|
||||
48rem: there the sidebar becomes a link strip above the article and
|
||||
there is no lane to fall into.) */
|
||||
@@ -884,8 +887,8 @@ article.multicol {
|
||||
margin-inline: auto 0;
|
||||
}
|
||||
|
||||
/* The sidebar fills the flexible lane (its left side stays on the
|
||||
viewport's left edge, growing rightward). */
|
||||
/* The sidebar fills the flexible lane (its start side stays on the
|
||||
viewport's start edge, growing toward the article). */
|
||||
body:has(#sidebar):has(.multicol):not(.editing) #sidebar {
|
||||
width: 100%;
|
||||
}
|
||||
@@ -893,30 +896,30 @@ article.multicol {
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.colseg,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h1,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h2 {
|
||||
margin-left: 0;
|
||||
margin-inline-start: 0;
|
||||
}
|
||||
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .margin,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .aside,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure:has(.margin) {
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure.margin {
|
||||
position: absolute;
|
||||
/* Attached to the article's left border (1.25rem gap), hanging into
|
||||
the left lane and growing leftward with it: 12rem when the lane is
|
||||
/* Attached to the article's start border (1.25rem gap), hanging into
|
||||
the start lane and growing with it: 12rem when the lane is
|
||||
tight, up to 150% (18rem) when the track or the surplus has room
|
||||
(100cqw - 100% is the surplus left of the right-leaning article).
|
||||
(100cqw - 100% is the surplus beside the end-leaning article).
|
||||
The lane (track + main's padding) always guarantees the room. */
|
||||
--box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem);
|
||||
width: var(--box-w);
|
||||
max-width: none;
|
||||
left: calc(-1.25rem - var(--box-w));
|
||||
inset-inline-start: calc(-1.25rem - var(--box-w));
|
||||
margin: 0.3rem 0 0;
|
||||
}
|
||||
}
|
||||
|
||||
/* A shrink-wrapped figure (explicit image width) centers in the plain
|
||||
layout; inside a column the centering looks adrift — left-align.
|
||||
Floated figures keep their own margins (the text gap). */
|
||||
.multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), :has(.margin)) {
|
||||
layout; inside a column the centering looks adrift — align to the start
|
||||
edge. Floated figures keep their own margins (the text gap). */
|
||||
.multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), .margin) {
|
||||
margin-inline: 0;
|
||||
}
|
||||
|
||||
@@ -988,13 +991,15 @@ article a:hover {
|
||||
|
||||
/* Blockquotes: spacing comes from the blockquote itself (bottom-only like
|
||||
everything else in articles); inner paragraphs keep only the gap between
|
||||
them. The negative left margin pushes the bar out past the text edge, so
|
||||
them. The negative start margin pushes the bar out past the text edge, so
|
||||
quoted text aligns with the surrounding paragraphs — same trick as code
|
||||
blocks. */
|
||||
blockquote {
|
||||
margin: 0 0 1rem -0.5rem;
|
||||
padding: 0 0 0 0.25rem;
|
||||
border-left: 0.25rem solid var(--accent2);
|
||||
margin: 0 0 1rem;
|
||||
margin-inline-start: -0.5rem;
|
||||
padding: 0;
|
||||
padding-inline-start: 0.25rem;
|
||||
border-inline-start: 0.25rem solid var(--accent2);
|
||||
color: var(--muted);
|
||||
}
|
||||
|
||||
@@ -1009,17 +1014,19 @@ blockquote p + p {
|
||||
/* Admonitions (markdown !!! note/warning/...) and GitHub-style alerts
|
||||
(> [!NOTE] ...): a lightweight callout in the blockquote idiom — accent
|
||||
bar and a faint wash, recolored per type, with a type emoji on the
|
||||
title. The negative left margin pushes bar and wash out past the text
|
||||
title. The negative start margin pushes bar and wash out past the text
|
||||
edge so the inner text aligns with surrounding paragraphs — same trick
|
||||
as blockquotes and code blocks (margin-left = border + padding-left).
|
||||
Bottom-only margins like everything else in articles; inner paragraphs
|
||||
carry no margins of their own. */
|
||||
as blockquotes and code blocks (margin-inline-start = border +
|
||||
padding-inline-start). Bottom-only margins like everything else in
|
||||
articles; inner paragraphs carry no margins of their own. */
|
||||
.admonition,
|
||||
.markdown-alert {
|
||||
margin: 0 0 1rem -1.15rem;
|
||||
margin: 0 0 1rem;
|
||||
margin-inline-start: -1.15rem;
|
||||
padding: 0.4rem 0.9rem;
|
||||
border-left: 0.25rem solid var(--admonition-color, var(--accent));
|
||||
border-radius: 0 0.3rem 0.3rem 0;
|
||||
border-inline-start: 0.25rem solid var(--admonition-color, var(--accent));
|
||||
border-start-end-radius: 0.3rem;
|
||||
border-end-end-radius: 0.3rem;
|
||||
background: color-mix(var(--admonition-color, var(--accent)) 7%, transparent);
|
||||
}
|
||||
|
||||
@@ -1037,7 +1044,7 @@ blockquote p + p {
|
||||
|
||||
.admonition-title::before,
|
||||
.markdown-alert-title::before {
|
||||
padding-right: 0.35em;
|
||||
padding-inline-end: 0.35em;
|
||||
}
|
||||
|
||||
.admonition.note .admonition-title::before,
|
||||
@@ -1094,22 +1101,24 @@ blockquote p + p {
|
||||
}
|
||||
|
||||
/* Side boxes: ::: aside is a muted floated box; {.margin} / ::: margin
|
||||
is a plainer margin note, and figures take {.margin} like {.left}.
|
||||
is a plainer margin note, and figures take the class directly on the
|
||||
<figure> (the renderer moves it off the img), floating like {.left}.
|
||||
Where the layout has room for a side zone — multicol pages, the
|
||||
sidebar's track, the wide single-column gutter (see the article
|
||||
section and the figure rules below) — the boxes are taken out of flow
|
||||
and absolutely positioned into it, off the article's left border, each
|
||||
and absolutely positioned into it, off the article's start border, each
|
||||
at the vertical spot where it occurs in the text (boxes occurring
|
||||
closer together than their heights may overlap — keep them apart);
|
||||
otherwise they stay in-column left floats (consecutive floats stack
|
||||
via clear: left). Headings already clear floats, so in-column boxes
|
||||
never bleed into the next section. */
|
||||
otherwise they stay in-column start floats (consecutive floats stack
|
||||
via clear: inline-start). Headings already clear floats, so in-column
|
||||
boxes never bleed into the next section. */
|
||||
.aside {
|
||||
float: left;
|
||||
clear: left;
|
||||
float: inline-start;
|
||||
clear: inline-start;
|
||||
width: 30%;
|
||||
max-width: 20rem;
|
||||
margin: 0.3rem 1.2rem 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1.2rem;
|
||||
padding: 0.6rem 0.9rem;
|
||||
font-size: 0.9rem;
|
||||
color: var(--muted);
|
||||
@@ -1139,11 +1148,12 @@ blockquote p + p {
|
||||
}
|
||||
|
||||
.margin {
|
||||
float: left;
|
||||
clear: left;
|
||||
float: inline-start;
|
||||
clear: inline-start;
|
||||
width: 30%;
|
||||
max-width: 20rem;
|
||||
margin: 0.3rem 1.2rem 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1.2rem;
|
||||
font-size: 0.9rem;
|
||||
color: var(--muted);
|
||||
}
|
||||
@@ -1152,10 +1162,10 @@ pre {
|
||||
overflow-x: auto;
|
||||
padding: 0.5rem 0.8rem;
|
||||
/* Code text aligns with the surrounding paragraphs: the box extends
|
||||
past them by its own padding. Themes that add a left border must
|
||||
extend margin-left by the border width to keep this alignment. */
|
||||
margin-left: -0.8rem;
|
||||
margin-right: -0.8rem;
|
||||
past them by its own padding. Themes that add a leading border must
|
||||
extend margin-inline-start by the border width to keep this
|
||||
alignment. */
|
||||
margin-inline: -0.8rem;
|
||||
background: var(--code-bg);
|
||||
border-radius: 4px;
|
||||
position: relative;
|
||||
@@ -1195,7 +1205,7 @@ p code {
|
||||
}
|
||||
|
||||
code:not(pre code):first-child {
|
||||
padding-left: 0;
|
||||
padding-inline-start: 0;
|
||||
}
|
||||
|
||||
/* Click-to-copy button (added by pagerite.js) */
|
||||
@@ -1240,7 +1250,7 @@ td {
|
||||
}
|
||||
|
||||
th {
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
background: linear-gradient(180deg,
|
||||
var(--table-head-a, color-mix(var(--accent) 10%, var(--surface))),
|
||||
var(--table-head-b, color-mix(var(--accent) 18%, var(--surface))));
|
||||
@@ -1287,7 +1297,8 @@ dd {
|
||||
Markdown images standing alone in a paragraph render as a block
|
||||
<figure> (with <figcaption> when the image has a title); the
|
||||
brace-attribute positioning class ({.left}, {.right}, {.wide}) lives
|
||||
on the img inside, but only the figure is ever positioned, so the
|
||||
on the img inside — except {.margin}, which the renderer moves onto
|
||||
the figure itself — but only the figure is ever positioned, so the
|
||||
caption stays below the image. Raw <img> HTML written by the author
|
||||
stays inline and unstyled beyond these defaults. */
|
||||
img {
|
||||
@@ -1311,20 +1322,24 @@ figure img:not([width]) {
|
||||
width: 100%;
|
||||
}
|
||||
|
||||
/* Floated figures: {.right} / {.left}, defaulting to 30% of the column
|
||||
and capped at half of it. */
|
||||
/* Floated figures: {.right} / {.left} float to the text column's end/start
|
||||
edge — the class names are author-facing and fixed, but the sides follow
|
||||
the text direction (in RTL, .left floats right) — defaulting to 30% of
|
||||
the column and capped at half of it. */
|
||||
figure:has(.right) {
|
||||
float: right;
|
||||
float: inline-end;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 0 1rem 1em;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-start: 1em;
|
||||
}
|
||||
|
||||
figure:has(.left) {
|
||||
float: left;
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* The same floats for other blocks: ::: left / ::: right containers
|
||||
@@ -1332,17 +1347,19 @@ figure:has(.left) {
|
||||
tables all take the class directly ({.right} at the end of a
|
||||
paragraph's last line, a trailing {.left} line after a fence, ...). */
|
||||
:is(div, p, pre, blockquote, table).right {
|
||||
float: right;
|
||||
float: inline-end;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 0 1rem 1em;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-start: 1em;
|
||||
}
|
||||
|
||||
:is(div, p, pre, blockquote, table).left {
|
||||
float: left;
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* An image with an explicit width attribute shrink-wraps instead: the
|
||||
@@ -1353,13 +1370,16 @@ figure:has(img[width]) {
|
||||
width: fit-content;
|
||||
}
|
||||
|
||||
/* {.margin} figures float left like {.left} ones — until they fall into
|
||||
the side zone (see the composition rules up in the article section). */
|
||||
figure:has(.margin) {
|
||||
float: left;
|
||||
/* {.margin} figures (the class moves onto the figure wrapper at render —
|
||||
see markdown.py) float to the start edge like {.left} ones — until
|
||||
they fall into the side zone (see the composition rules up in the
|
||||
article section). */
|
||||
figure.margin {
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* Click-to-enlarge (pagerite.js): article figure images open in a
|
||||
@@ -1436,28 +1456,28 @@ article figure img {
|
||||
}
|
||||
}
|
||||
|
||||
/* Wide single-column pages: margin boxes lean into the vacant left
|
||||
/* Wide single-column pages: margin boxes lean into the vacant start-side
|
||||
gutter instead (below 104rem the gutter cannot hold the box, and while
|
||||
editing the docked panel reshapes the gutters — in both they stay
|
||||
plain floats). Out of flow like on multicol pages: the box hangs off
|
||||
the article's left border, growing with the gutter up to 150% (18rem),
|
||||
its right side 1.25rem off the border. */
|
||||
the article's start border, growing with the gutter up to 150% (18rem),
|
||||
its end side 1.25rem off the border. */
|
||||
@media (min-width: 104rem) {
|
||||
body:not(.editing):not(:has(.multicol)) article .margin,
|
||||
body:not(.editing):not(:has(.multicol)) article .aside,
|
||||
body:not(.editing):not(:has(.multicol)) article figure:has(.margin) {
|
||||
body:not(.editing):not(:has(.multicol)) article figure.margin {
|
||||
position: absolute;
|
||||
--box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem);
|
||||
width: var(--box-w);
|
||||
max-width: none;
|
||||
left: calc(-1.25rem - var(--box-w));
|
||||
inset-inline-start: calc(-1.25rem - var(--box-w));
|
||||
margin: 0.3rem 0 0;
|
||||
}
|
||||
}
|
||||
|
||||
/* In the wide symmetric gutters (where the sidebar overlays the flexible
|
||||
left gutter rather than a reserved track) the sidebar flexes with the
|
||||
gutter up to 150% — its left side stays on the viewport's edge. */
|
||||
start gutter rather than a reserved track) the sidebar flexes with the
|
||||
gutter up to 150% — its start side stays on the viewport's edge. */
|
||||
@media (min-width: 102rem) {
|
||||
body:not(.editing) #sidebar {
|
||||
width: min(18rem, 100%);
|
||||
@@ -1506,7 +1526,7 @@ body.editing pre.wide {
|
||||
|
||||
/* Narrow single-column pages with a sidebar: below 102rem the symmetric
|
||||
gutters can no longer both hold the sidebar, so #content reserves it
|
||||
with a flexible left track (see the matching media query below) and the
|
||||
with a flexible start track (see the matching media query below) and the
|
||||
article always starts at the lane's width (+ main's 1.25rem padding) —
|
||||
the bleed margin measures off --lane. Scoped by :has(#sidebar) since
|
||||
the sidebar element is omitted entirely on pages without
|
||||
@@ -1533,9 +1553,9 @@ body:has(.multicol) pre.wide {
|
||||
}
|
||||
|
||||
/* Multicol with a sidebar track (≥48rem, see #content): main starts at
|
||||
the flexible lane's width and the article leans right, so the bleed
|
||||
extends left past the surplus and the lane to the true viewport edge —
|
||||
sliding under the translucent sidebar — and right past main's
|
||||
the flexible lane's width and the article leans to the end edge, so the
|
||||
bleed extends past the surplus and the lane to the true viewport start
|
||||
edge — sliding under the translucent sidebar — and past main's end
|
||||
padding. */
|
||||
@media (min-width: 48rem) {
|
||||
body:has(#sidebar):has(.multicol):not(.editing) figure:has(.wide),
|
||||
@@ -1553,7 +1573,7 @@ figcaption {
|
||||
hyphens: auto;
|
||||
-webkit-hyphens: auto;
|
||||
text-wrap: pretty;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
/* Never let a long caption stretch a shrink-to-fit figure wider than the
|
||||
image; the caption wraps at the figure's width instead. */
|
||||
width: 0;
|
||||
@@ -1579,7 +1599,7 @@ article h2 {
|
||||
|
||||
/* Narrow windows with a sidebar: below 102rem the symmetric gutters can no
|
||||
longer both hold the 12rem sidebar, so reserve its space with a flexible
|
||||
left track instead of letting it overlap the article (multicol pages use
|
||||
start track instead of letting it overlap the article (multicol pages use
|
||||
the same track at every width — see the #content rules above; their
|
||||
higher-specificity rule wins there). The lane is 12rem when space is
|
||||
tight, growing up to 150% (18rem) once the viewport exceeds the
|
||||
@@ -1668,7 +1688,7 @@ article h2 {
|
||||
|
||||
figure:has(.right),
|
||||
figure:has(.left),
|
||||
figure:has(.margin) {
|
||||
figure.margin {
|
||||
float: none;
|
||||
width: 100%;
|
||||
max-width: none;
|
||||
@@ -1699,10 +1719,10 @@ article h2 {
|
||||
width: fit-content;
|
||||
}
|
||||
|
||||
/* The single-column sidebar .wide margins assume a left sidebar column;
|
||||
with the sidebar on top the article is viewport-wide and the plain
|
||||
centered bleed applies again. (Multicol pages need no override: their
|
||||
cqw bleed is exact at any width.) */
|
||||
/* The single-column sidebar .wide margins assume a start-side sidebar
|
||||
column; with the sidebar on top the article is viewport-wide and the
|
||||
plain centered bleed applies again. (Multicol pages need no override:
|
||||
their cqw bleed is exact at any width.) */
|
||||
body:has(#sidebar):not(.editing):not(:has(.multicol)) figure:has(.wide) {
|
||||
margin-inline: calc(50% - 50vw);
|
||||
}
|
||||
@@ -1743,7 +1763,7 @@ article h2 {
|
||||
.dateline {
|
||||
color: var(--muted);
|
||||
font-size: 0.85rem;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
}
|
||||
|
||||
.footnotes {
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
// Shared popup open-state behavior: while `open` (a ref, truthy = open)
|
||||
// is set, a pointerdown outside `root` (a template ref covering both the
|
||||
// toggle button and the popup) or Escape resets it to null. One logic for
|
||||
// every dropdown (LangSelect, the page editor's class/table pickers), so
|
||||
// they can't drift apart.
|
||||
import { onBeforeUnmount, watch } from 'vue'
|
||||
|
||||
export function usePopup(open, root) {
|
||||
let off = null
|
||||
const stop = watch(open, (v) => {
|
||||
off?.()
|
||||
off = null
|
||||
if (!v) return
|
||||
const down = (ev) => { if (!root.value?.contains(ev.target)) open.value = null }
|
||||
const key = (ev) => { if (ev.key === 'Escape') open.value = null }
|
||||
addEventListener('pointerdown', down, true)
|
||||
addEventListener('keydown', key)
|
||||
off = () => {
|
||||
removeEventListener('pointerdown', down, true)
|
||||
removeEventListener('keydown', key)
|
||||
}
|
||||
})
|
||||
onBeforeUnmount(() => { off?.(); stop() })
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
// The editor shell's shared language selection ('' = the primary language):
|
||||
// backed by the app-wide store (./store), so the editor tabs' LangSelects
|
||||
// and the public corner selector bind the same value. Linked to the
|
||||
// whole-page language: while the panel is open it drives the page preview
|
||||
// (EditorShell applies it as the fetch-time language override, swapdoc).
|
||||
import { computed, ref } from 'vue'
|
||||
import { pinia, useStore } from './store'
|
||||
|
||||
export const editorLang = computed({
|
||||
get: () => useStore(pinia).lang,
|
||||
set: (v) => { useStore(pinia).lang = v },
|
||||
})
|
||||
|
||||
// The CURRENT PAGE's primary language ('' = not yet learned): the shell's
|
||||
// settings fetch fills it with the site default; the page/structure tabs
|
||||
// then refine it per page (doc accept / tree rows — strictly better
|
||||
// sources, so they overwrite freely while the settings fetch only fills
|
||||
// the unknown). EditorShell pins the preview by it when the selection is
|
||||
// '' (the primary).
|
||||
export const pagePrimary = ref('')
|
||||
@@ -0,0 +1,71 @@
|
||||
// Language helpers shared by the editors (the PageEditor language picker,
|
||||
// the localization settings tab). Flags come from the country-flag-icons
|
||||
// set, same as the analytics visitor cells.
|
||||
import * as flagSvgs from 'country-flag-icons/string/3x2'
|
||||
|
||||
// The Seed-X reference translator's languages (scripts/translator.py) — the
|
||||
// translation-target ceiling — each mapped to the language's home country
|
||||
// flag (England for English, Portugal for Portuguese — not the most
|
||||
// populous variant). Internal tags are the bare 2-letter base subtags; the
|
||||
// translator decides the variant. A variant tag (en-US, pt-BR) is still a
|
||||
// valid explicit selection for a future translator that distinguishes them
|
||||
// — flagFor shows its own region then.
|
||||
export const TRANSLATABLE = {
|
||||
ar: 'EG', cs: 'CZ', da: 'DK', de: 'DE', el: 'GR', en: 'GB', es: 'ES',
|
||||
fa: 'IR', fi: 'FI', fr: 'FR', hu: 'HU', id: 'ID', it: 'IT', ja: 'JP',
|
||||
ko: 'KR', ms: 'MY', nl: 'NL', no: 'NO', pl: 'PL', pt: 'PT', ro: 'RO',
|
||||
ru: 'RU', sv: 'SE', th: 'TH', tr: 'TR', uk: 'UA', vi: 'VN', zh: 'CN',
|
||||
}
|
||||
|
||||
// The languages in geographic/cultural groups (the lang tab's flag grid
|
||||
// lays them out one group per row, in this order): English with the
|
||||
// Nordics, then Western/Central and Eastern Europe, Southern Europe with
|
||||
// the Middle East, and Asia.
|
||||
export const LANG_GROUPS = [
|
||||
['en', 'nl', 'da', 'no', 'sv', 'fi', 'ru'],
|
||||
['fr', 'de', 'pl', 'cs', 'hu', 'ro', 'uk'],
|
||||
['es', 'pt', 'it', 'el', 'tr', 'ar', 'fa'],
|
||||
['zh', 'ja', 'ko', 'vi', 'th', 'id', 'ms'],
|
||||
]
|
||||
|
||||
const displayNames = new Intl.DisplayNames(['en'], { type: 'language' })
|
||||
|
||||
// Consistent menu ordering for language selectors: the geographic/cultural
|
||||
// grouping above (similar languages sit together, and it does not vary with
|
||||
// the display language the way alphabetical-by-name would). Tags outside
|
||||
// the groups trail, ordered by tag. The primary language is not special
|
||||
// here — callers put it first themselves.
|
||||
const groupOrder = new Map(LANG_GROUPS.flat().map((c, i) => [c, i]))
|
||||
export function langSort(codes) {
|
||||
return [...codes].sort(
|
||||
(a, b) =>
|
||||
(groupOrder.get(a) ?? groupOrder.size) - (groupOrder.get(b) ?? groupOrder.size)
|
||||
|| a.localeCompare(b),
|
||||
)
|
||||
}
|
||||
|
||||
// English display name for a language tag ("fi" -> "Finnish").
|
||||
export function langName(tag) {
|
||||
try {
|
||||
return displayNames.of(tag) || tag
|
||||
} catch {
|
||||
return tag
|
||||
}
|
||||
}
|
||||
|
||||
// Flag SVG string for a language tag: an explicit region variant (en-US)
|
||||
// gets its own region's flag; a bare base tag maps to the language's home
|
||||
// country (en → GB, pt → PT); languages outside the list fall back to the
|
||||
// tag's most likely region.
|
||||
export function flagFor(tag) {
|
||||
tag = tag || ''
|
||||
if (!tag.includes('-')) {
|
||||
const country = TRANSLATABLE[tag.split('-')[0].toLowerCase()]
|
||||
if (country) return flagSvgs[country] || ''
|
||||
}
|
||||
try {
|
||||
return flagSvgs[new Intl.Locale(tag).maximize().region] || ''
|
||||
} catch {
|
||||
return ''
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
// Public language-selector entry: imported on demand by pagerite.js on
|
||||
// pages advertising more than one language in their hreflang alternates.
|
||||
// Vue, Pinia and the flag SVG set live in this chunk only — untranslated
|
||||
// pages never pay for them. The selector's state lives in the shared
|
||||
// store (./store), not the DOM: the corner container is rebuilt freely
|
||||
// and ensureMounted re-mounts from the store.
|
||||
import { createApp } from 'vue'
|
||||
import LangSelector from './LangSelector.vue'
|
||||
import { pinia, useStore } from './store'
|
||||
|
||||
let app = null
|
||||
|
||||
function store() {
|
||||
return useStore(pinia)
|
||||
}
|
||||
|
||||
// The current page's languages (called on every navigation).
|
||||
export function setLanguages(alternates, current) {
|
||||
Object.assign(store(), {
|
||||
langAlternates: alternates,
|
||||
servedLang: current,
|
||||
langSelectorActive: true,
|
||||
})
|
||||
}
|
||||
|
||||
// The current page is single-language: the selector goes away.
|
||||
export function hide() {
|
||||
store().langSelectorActive = false
|
||||
app?.unmount()
|
||||
app = null
|
||||
}
|
||||
|
||||
// Mount the selector as the container's first item; re-mount when its
|
||||
// element went away with a container rebuild (a live app updates from the
|
||||
// store reactively).
|
||||
export function ensureMounted(host) {
|
||||
if (!store().langSelectorActive || !host) {
|
||||
app?.unmount()
|
||||
app = null
|
||||
return
|
||||
}
|
||||
if (app && host.contains(app._container)) return
|
||||
app?.unmount()
|
||||
const el = document.createElement('div')
|
||||
el.id = 'lang-selector'
|
||||
host.prepend(el)
|
||||
app = createApp(LangSelector)
|
||||
app.use(pinia)
|
||||
app.mount(el)
|
||||
}
|
||||
+187
-47
@@ -6,6 +6,7 @@
|
||||
// support from the article itself and are re-applied after each swap.
|
||||
import { OverlayScrollbars } from "overlayscrollbars";
|
||||
import "overlayscrollbars/overlayscrollbars.css";
|
||||
import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
|
||||
|
||||
(() => {
|
||||
// Overlay scrollbars: the native kind reserves a strip of layout (or
|
||||
@@ -34,6 +35,59 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
});
|
||||
}
|
||||
|
||||
// --- Language override (?lang=) ---------------------------------------
|
||||
// /page?lang=fi serves a translated, indexable version (each language is
|
||||
// its own canonical). The chosen language sticks for the session of
|
||||
// clicks: the server replicates ?lang= onto the navigation links it
|
||||
// renders (nav, sidebar, cards — in-article links are content and stay
|
||||
// as authored), and pageUrl adds it to internal fetches that lack one.
|
||||
// The address bar keeps the pretty URL: the query is stripped on load
|
||||
// and never pushed into history. A full refresh or a shared link resets
|
||||
// to automatic selection (the browser's own Accept-Language — every
|
||||
// plain fetch carries it by default). See docs/localization.md.
|
||||
const langParam = new URL(location.href).searchParams.get("lang");
|
||||
if (langParam) {
|
||||
const url = new URL(location.href);
|
||||
url.searchParams.delete("lang");
|
||||
history.replaceState(history.state, "", url);
|
||||
}
|
||||
// The session language: the user's explicit pick (initial ?lang=, public
|
||||
// selector, editor dropdown) is kept in chosenLang; while the editor is
|
||||
// open its selection overrides it (swapdoc.setLangOverride), and closing
|
||||
// falls back to chosenLang. JS state only — pretty URLs, no reloads.
|
||||
// window.__pageriteLang is the pin for swapdoc.loadPlain's fetches.
|
||||
let chosenLang = langParam;
|
||||
let sessionLang = langParam;
|
||||
window.__pageriteLang = sessionLang;
|
||||
addEventListener("pagerite:session-lang", (ev) => {
|
||||
if (ev.detail?.lang) chosenLang = ev.detail.lang;
|
||||
sessionLang = ev.detail?.lang || chosenLang;
|
||||
window.__pageriteLang = sessionLang;
|
||||
});
|
||||
// An internal URL as fetched: carries the session's ?lang= unless the
|
||||
// link already pins a language of its own. With no ?lang= on the initial
|
||||
// load nothing is ever added.
|
||||
const pageUrl = (url) => {
|
||||
const u = new URL(url, location.href);
|
||||
if (sessionLang && u.origin === location.origin && !u.searchParams.has("lang")) {
|
||||
u.searchParams.set("lang", sessionLang);
|
||||
}
|
||||
return u;
|
||||
};
|
||||
// The in-memory page cache is keyed by path + query: the same pathname
|
||||
// holds different HTML for each language version.
|
||||
const rawKey = (url) => {
|
||||
const u = new URL(url, location.href);
|
||||
return u.pathname + u.search;
|
||||
};
|
||||
const cacheKey = (url) => rawKey(pageUrl(url));
|
||||
// What goes into the address bar and history: the pretty URL, no ?lang=.
|
||||
const prettyUrl = (url) => {
|
||||
const u = pageUrl(url);
|
||||
u.searchParams.delete("lang");
|
||||
return u;
|
||||
};
|
||||
|
||||
// Regions every page has. #sidebar is NOT among them: it is omitted
|
||||
// entirely when the section has no sub-navigation, and handled below.
|
||||
const REGIONS = ["page-banner", "nav", "main"];
|
||||
@@ -78,16 +132,16 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
if (line != null) {
|
||||
// Section pen on an anchored h2: opens the page editor at the
|
||||
// section's markdown source line (data-line, from the backend).
|
||||
btn.className = "edit-link edit-section";
|
||||
btn.className = "edit-link edit-section icon-btn";
|
||||
btn.title = "edit section";
|
||||
btn.textContent = "🖊️";
|
||||
btn.dataset.editorLine = line;
|
||||
} else if (mode === "page") {
|
||||
btn.className = "edit-link edit-page";
|
||||
btn.className = "edit-link edit-page icon-btn";
|
||||
btn.title = "edit page";
|
||||
btn.textContent = "🖊️";
|
||||
} else {
|
||||
btn.className = "edit-link site-edit-link";
|
||||
btn.className = "edit-link site-edit-link icon-btn";
|
||||
btn.title = "site settings";
|
||||
btn.textContent = "⚙️";
|
||||
}
|
||||
@@ -111,13 +165,29 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
|
||||
function makeAuthLink(admin) {
|
||||
const a = document.createElement("a");
|
||||
a.className = admin ? "profile-link" : "login-link";
|
||||
a.className = (admin ? "profile-link" : "login-link") + " icon-btn";
|
||||
a.href = "/auth/";
|
||||
a.title = admin ? "profile" : "log in";
|
||||
a.textContent = admin ? "\u{1F510}" : "\u{1F511}";
|
||||
return a;
|
||||
}
|
||||
|
||||
// The banner top-right corner container: the language selector (first
|
||||
// item) plus the admin pens and auth links. renderAuthUi rebuilds it from
|
||||
// scratch; the selector's state lives in the shared store, not the DOM,
|
||||
// so the langselect bundle re-mounts it into the fresh container.
|
||||
function pensContainer() {
|
||||
let pens = document.querySelector(".editor-pens");
|
||||
if (!pens) {
|
||||
const banner = document.getElementById("page-banner");
|
||||
if (!banner) return null;
|
||||
pens = document.createElement("div");
|
||||
pens.className = "editor-pens";
|
||||
banner.after(pens);
|
||||
}
|
||||
return pens;
|
||||
}
|
||||
|
||||
function removePens() {
|
||||
document.querySelectorAll(".editor-pens, #main article button.edit-link")
|
||||
.forEach((el) => el.remove());
|
||||
@@ -129,32 +199,33 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
// pens that may have been injected while the browser cache made us look
|
||||
// authenticated.
|
||||
removePens();
|
||||
if (!authReady) return;
|
||||
|
||||
// Editing is open for admins and, as a dev/no-proxy fallback, when no
|
||||
// Paskia SSO is detected at all.
|
||||
const canEdit = isAdmin || !ssoAvailable;
|
||||
// The analytics page is a read-only dashboard: editing pens and the side
|
||||
// panel do not apply there. Login/logout links are still useful.
|
||||
const onAnalytics = currentPath === "/_a";
|
||||
const banner = document.getElementById("page-banner");
|
||||
if (banner) {
|
||||
const pens = document.createElement("div");
|
||||
pens.className = "editor-pens";
|
||||
if (canEdit && !onAnalytics) {
|
||||
// Analytics viewer is now a normal page at /_a.
|
||||
const a = document.createElement("a");
|
||||
a.className = "edit-link analytics-link";
|
||||
a.href = "/_a";
|
||||
a.title = "analytics";
|
||||
a.textContent = "📊";
|
||||
pens.append(a);
|
||||
pens.append(makePen("site"));
|
||||
if (authReady) {
|
||||
// Editing is open for admins and, as a dev/no-proxy fallback, when no
|
||||
// Paskia SSO is detected at all.
|
||||
const canEdit = isAdmin || !ssoAvailable;
|
||||
// The analytics page is a read-only dashboard: editing pens and the side
|
||||
// panel do not apply there. Login/logout links are still useful.
|
||||
const onAnalytics = currentPath === "/_a";
|
||||
if (document.getElementById("page-banner")) {
|
||||
const pens = pensContainer();
|
||||
if (canEdit && !onAnalytics) {
|
||||
// Analytics viewer is now a normal page at /_a.
|
||||
const a = document.createElement("a");
|
||||
a.className = "edit-link analytics-link icon-btn";
|
||||
a.href = "/_a";
|
||||
a.title = "analytics";
|
||||
a.textContent = "📊";
|
||||
pens.append(a);
|
||||
pens.append(makePen("site"));
|
||||
}
|
||||
if (ssoAvailable) pens.append(makeAuthLink(isAdmin));
|
||||
if (!pens.firstElementChild) pens.remove();
|
||||
}
|
||||
if (ssoAvailable) pens.append(makeAuthLink(isAdmin));
|
||||
banner.after(pens);
|
||||
if (canEdit && !onAnalytics) injectPagePen();
|
||||
}
|
||||
if (canEdit && !onAnalytics) injectPagePen();
|
||||
// Re-mount the selector into the fresh container (no-op until the
|
||||
// bundle has been loaded once).
|
||||
langselectMod?.ensureMounted(document.querySelector(".editor-pens"));
|
||||
}
|
||||
|
||||
async function setupAuth() {
|
||||
@@ -352,9 +423,15 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
// received it as the document (re-fetching would be redundant, and
|
||||
// browser heuristics may send it without if-none-match, defeating the
|
||||
// conditional request); it enters the cache when navigated to.
|
||||
const pageCache = new Map(); // pathname -> HTML text
|
||||
const pageCache = new Map(); // rawKey/cacheKey(url) -> HTML text
|
||||
addEventListener("pagerite:page-fetched", (ev) => {
|
||||
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html);
|
||||
// Key by the URL as announced, exactly as the editor fetched it: a
|
||||
// copy pinned to a language (?lang=) caches under its own key, where
|
||||
// navigation with the same session language finds it.
|
||||
pageCache.set(rawKey(ev.detail.url), ev.detail.html);
|
||||
// Editor-driven swaps don't go through load(): re-evaluate the
|
||||
// language selector from the fresh copy too.
|
||||
mountLangselect(new DOMParser().parseFromString(ev.detail.html, "text/html"));
|
||||
});
|
||||
|
||||
// Editors mutate site-wide state (theme, structure, headings, banners),
|
||||
@@ -373,21 +450,22 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
});
|
||||
|
||||
function preload() {
|
||||
const urls = new Set();
|
||||
const urls = new Map(); // cache key -> URL, deduped (hashes collapse)
|
||||
for (const a of document.querySelectorAll(
|
||||
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
|
||||
)) {
|
||||
urls.add(a.pathname);
|
||||
const u = pageUrl(a.href);
|
||||
urls.set(rawKey(u), u);
|
||||
}
|
||||
for (const url of urls) {
|
||||
if (pageCache.has(url)) continue;
|
||||
for (const [key, u] of urls) {
|
||||
if (pageCache.has(key)) continue;
|
||||
// x-pagerite-preload: idle cache warm-up, not a page view — the
|
||||
// server excludes these GETs from analytics (the navigation message
|
||||
// sent on actual navigation does the counting).
|
||||
fetch(url, { headers: { "x-pagerite-preload": "1" } })
|
||||
fetch(u, { headers: { "x-pagerite-preload": "1" } })
|
||||
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
|
||||
? r.text() : ""))
|
||||
.then((html) => { if (html) pageCache.set(url, html); })
|
||||
.then((html) => { if (html) pageCache.set(key, html); })
|
||||
.catch(() => {});
|
||||
}
|
||||
}
|
||||
@@ -506,8 +584,12 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
// failing in the background.
|
||||
let ws = null;
|
||||
const wsQueue = [];
|
||||
let wsReconnectMs = 1000;
|
||||
let wsNotBefore = 0;
|
||||
const wsPolicy = reconnectPolicy({ min: 1000 });
|
||||
// The first attempt is staggered too: page load opens several sockets at
|
||||
// once (Vite's HMR socket, the editors), and the burst trips the browser's
|
||||
// WebSocket throttling (sockets then sit "pending" for minutes).
|
||||
let wsNotBefore = Date.now() + socketSlot();
|
||||
let wsWatchdog = null;
|
||||
|
||||
function activityWs() {
|
||||
if (ws || Date.now() < wsNotBefore) return;
|
||||
@@ -518,15 +600,16 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
} catch {
|
||||
return;
|
||||
}
|
||||
clearTimeout(wsWatchdog);
|
||||
wsWatchdog = watchConnecting(ws, "activity");
|
||||
ws.onopen = () => {
|
||||
wsReconnectMs = 1000;
|
||||
wsPolicy.opened();
|
||||
for (const msg of wsQueue.splice(0)) ws.send(JSON.stringify(msg));
|
||||
};
|
||||
ws.onclose = () => {
|
||||
ws = null;
|
||||
// No timer here: the next user activity retries, after the backoff.
|
||||
wsNotBefore = Date.now() + wsReconnectMs;
|
||||
wsReconnectMs = Math.min(wsReconnectMs * 2, 30_000);
|
||||
wsNotBefore = Date.now() + wsPolicy.closed();
|
||||
};
|
||||
ws.onerror = () => ws.close();
|
||||
}
|
||||
@@ -684,6 +767,56 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
}
|
||||
}
|
||||
|
||||
// --- Public language selector ------------------------------------------
|
||||
// Pages translated into more than one language advertise it via hreflang
|
||||
// alternates (x-default + one link per language). Those pages get the
|
||||
// editors' flag dropdown as the first item of the corner container; its
|
||||
// bundle (Vue + the flag SVG set) loads on demand. Re-evaluated from the
|
||||
// fresh document on every swap (the head's own alternates stay stale).
|
||||
let langselectMod = null;
|
||||
async function mountLangselect(doc) {
|
||||
const links = [...doc.head.querySelectorAll('link[rel="alternate"][hreflang]')];
|
||||
const dflt = links.find((l) => l.hreflang === "x-default");
|
||||
const langs = links.filter((l) => l.hreflang && l.hreflang !== "x-default");
|
||||
if (!dflt || langs.length <= 1) return langselectMod?.hide();
|
||||
try {
|
||||
langselectMod ??= await import(/* @vite-ignore */ assets["pagerite:langselect-src"]);
|
||||
for (const css of (assets["pagerite:langselect-css"] || "").split(",")) {
|
||||
if (css && !document.querySelector(`link[href="${css}"]`)) {
|
||||
const link = document.createElement("link");
|
||||
link.rel = "stylesheet";
|
||||
link.href = css;
|
||||
link.dataset.pagerite = "langselect-css";
|
||||
document.head.append(link);
|
||||
}
|
||||
}
|
||||
langselectMod.setLanguages(
|
||||
// The original's alternate is the plain URL — x-default's href —
|
||||
// which also marks it as the primary option.
|
||||
langs.map((l) => ({ tag: l.hreflang, href: l.href, primary: l.href === dflt.href })),
|
||||
doc.documentElement.lang,
|
||||
);
|
||||
langselectMod.ensureMounted(pensContainer());
|
||||
} catch (e) {
|
||||
console.error("language selector mount failed:", e);
|
||||
}
|
||||
}
|
||||
|
||||
// The selector's pick (LangSelector dispatches this): make it the
|
||||
// session language and swap the page in place. With the editor open the
|
||||
// pick already landed in the shared store — the editor's watch re-renders
|
||||
// the page itself, so there is nothing to do here.
|
||||
addEventListener("pagerite:set-session-lang", async (ev) => {
|
||||
const tag = ev.detail?.lang;
|
||||
if (!tag || tag === sessionLang) return;
|
||||
if (document.body.classList.contains("editing")) return;
|
||||
chosenLang = sessionLang = tag;
|
||||
window.__pageriteLang = tag;
|
||||
const y = scrollY; // a language switch is not a navigation: keep scroll
|
||||
await load(currentPath, false);
|
||||
scrollTo(0, y);
|
||||
});
|
||||
|
||||
// --- Fetch navigation ------------------------------------------------
|
||||
async function load(url, push = true, back = false) {
|
||||
// Navigating with the editor open closes it; unsaved edits are lost
|
||||
@@ -697,12 +830,12 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
teardownAnalytics();
|
||||
let doc;
|
||||
let finalUrl = url;
|
||||
const cached = !editing && pageCache.get(new URL(url, location.href).pathname);
|
||||
const cached = !editing && pageCache.get(cacheKey(url));
|
||||
if (cached) {
|
||||
doc = new DOMParser().parseFromString(cached, "text/html");
|
||||
} else {
|
||||
try {
|
||||
const res = await fetch(url);
|
||||
const res = await fetch(pageUrl(url));
|
||||
const type = res.headers.get("content-type") || "";
|
||||
if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
|
||||
// Reflect any redirect the server issued.
|
||||
@@ -710,15 +843,15 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
const html = await res.text();
|
||||
// Populate the cache too, so returning here (back/forward, or a
|
||||
// self-link in the nav) is served from memory.
|
||||
pageCache.set(new URL(finalUrl, location.href).pathname, html);
|
||||
pageCache.set(cacheKey(finalUrl), html);
|
||||
doc = new DOMParser().parseFromString(html, "text/html");
|
||||
} catch {
|
||||
location.href = url; // fall back to a normal navigation
|
||||
location.href = pageUrl(url); // fall back to a normal navigation
|
||||
return false;
|
||||
}
|
||||
}
|
||||
if (REGIONS.some((id) => !doc.getElementById(id))) {
|
||||
location.href = url;
|
||||
location.href = pageUrl(url);
|
||||
return false;
|
||||
}
|
||||
const doit = () => {
|
||||
@@ -774,12 +907,18 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
// stylesheet after the server-rendered tag.
|
||||
const userStyle = document.getElementById("pagerite-user");
|
||||
if (userStyle) document.head.appendChild(userStyle);
|
||||
// The served language rides on <html> (lang + dir, rtl for e.g.
|
||||
// Arabic) — follow the swapped page (the editor panel carries its
|
||||
// own lang="en" dir="ltr", so it is unaffected).
|
||||
document.documentElement.lang = doc.documentElement.lang;
|
||||
document.documentElement.dir = doc.documentElement.dir;
|
||||
document.title = doc.title;
|
||||
// Banners may contain scripts (canvas etc.), content pages may too.
|
||||
runScripts(document.getElementById("page-banner"));
|
||||
runScripts(document.getElementById("main"));
|
||||
applyEffects();
|
||||
mountAnalytics(doc);
|
||||
mountLangselect(doc);
|
||||
};
|
||||
// Rotating cube page transition (styles injected as #pagerite-transition
|
||||
// from the selected design's transition.css, e.g. themes/cube/);
|
||||
@@ -798,7 +937,7 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
doit();
|
||||
}
|
||||
currentPath = new URL(finalUrl, location.href).pathname;
|
||||
if (push) history.pushState({ idx: ++historyIdx }, "", finalUrl);
|
||||
if (push) history.pushState({ idx: ++historyIdx }, "", prettyUrl(finalUrl));
|
||||
// The open editor follows the URL: retarget the per-page tabs to the
|
||||
// navigated-to page (unsaved text of the previous page is discarded —
|
||||
// the article it previewed into is gone).
|
||||
@@ -1013,4 +1152,5 @@ import "overlayscrollbars/overlayscrollbars.css";
|
||||
setupAuth();
|
||||
applyEffects();
|
||||
mountAnalytics(document);
|
||||
mountLangselect(document);
|
||||
})();
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
// Shared reconnect policy for the WebSockets (page/banner editors,
|
||||
// analytics view, the activity channel). Two things trip a browser's
|
||||
// WebSocket throttling, after which every socket to the host sits
|
||||
// "pending" (never opens, never closes) for minutes:
|
||||
//
|
||||
// 1. A burst of simultaneous attempts — page load opens Vite's HMR
|
||||
// socket plus several of ours at the same moment, and every refresh
|
||||
// repeats the burst. socketSlot() spaces new sockets out.
|
||||
// 2. Too-frequent retries — so failed attempts back off exponentially
|
||||
// (a few seconds, doubling to half a minute), reset only after a
|
||||
// connection stayed open long enough to count as healthy. A socket
|
||||
// that closes right after opening must NOT reset the backoff.
|
||||
export function reconnectPolicy({ min = 2000, max = 30000, healthyAfter = 30000 } = {}) {
|
||||
let delay = min
|
||||
let openedAt = 0
|
||||
return {
|
||||
// Stamp a socket that just opened.
|
||||
opened() {
|
||||
openedAt = Date.now()
|
||||
},
|
||||
// The socket closed: the wait before the next attempt (up to 50%
|
||||
// jitter; the base doubles per failure). A healthy streak resets it.
|
||||
closed() {
|
||||
if (openedAt && Date.now() - openedAt >= healthyAfter) delay = min
|
||||
openedAt = 0
|
||||
const wait = Math.round(delay * (1 + Math.random() * 0.5))
|
||||
delay = Math.min(delay * 2, max)
|
||||
return wait
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// Sockets created at the same moment (page load: Vite's HMR socket plus
|
||||
// ours) read as one burst to the browser's throttling. Space new sockets
|
||||
// out: each call reserves a slot a beat after the previous one.
|
||||
let nextSlot = 0
|
||||
export function socketSlot() {
|
||||
const now = Date.now()
|
||||
const wait = Math.max(0, nextSlot - now)
|
||||
nextSlot = Math.max(now, nextSlot) + 300
|
||||
return wait
|
||||
}
|
||||
|
||||
// A socket still CONNECTING after this long counts as a failed attempt:
|
||||
// browser throttling leaves sockets "pending" (no open, no close) for
|
||||
// minutes, and without a watchdog the app would wait on one forever (the
|
||||
// recurring empty editor). Closing it fires onclose, which reschedules
|
||||
// through the policy's backoff — it never reconnects aggressively itself.
|
||||
export function watchConnecting(ws, label) {
|
||||
return setTimeout(() => {
|
||||
if (ws.readyState === WebSocket.CONNECTING) {
|
||||
console.warn(`[pagerite] ${label} socket stuck connecting — closing it, retrying with backoff`)
|
||||
ws.close()
|
||||
}
|
||||
}, 10_000)
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
// The app's shared Pinia store — cross-bundle UI state lives here. Every
|
||||
// entry chunk imports its own copy of this module, so the Pinia instance
|
||||
// is parked on window (Vue itself is a shared chunk, so reactivity works
|
||||
// across the copies). Pass `pinia` explicitly when calling useStore
|
||||
// outside a component (module code, no active instance).
|
||||
import { createPinia, defineStore } from 'pinia'
|
||||
|
||||
export const pinia = (window.__pageritePinia ??= createPinia())
|
||||
|
||||
export const useStore = defineStore('pagerite', {
|
||||
state: () => ({
|
||||
// The ONE language selection, v-modeled by both dropdowns (editor
|
||||
// tabs, public corner selector): '' = no explicit pick (the page's
|
||||
// primary / autodetect), else a concrete tag. A pick from either
|
||||
// dropdown is visible to everyone immediately.
|
||||
lang: '',
|
||||
// The language the current page was actually served in (set by
|
||||
// pagerite.js per navigation) — the selector's highlight fallback
|
||||
// when there is no explicit pick.
|
||||
servedLang: '',
|
||||
// The public selector's page data: hreflang alternates
|
||||
// ([{tag, href, primary}]) and whether to show at all.
|
||||
langAlternates: [],
|
||||
langSelectorActive: false,
|
||||
}),
|
||||
})
|
||||
+31
-3
@@ -11,6 +11,20 @@ export function dropPageCache() {
|
||||
dispatchEvent(new CustomEvent('pagerite:drop-page-cache'))
|
||||
}
|
||||
|
||||
// The editor's language override (set by EditorShell): while the panel is
|
||||
// open, its language selection wins over the normal preferences — every
|
||||
// in-place re-render asks for that language explicitly, and pagerite.js
|
||||
// applies it to its own fetches and prefetches (pagerite:session-lang).
|
||||
// The primary selection pins by its code: ?lang=<primary> selects the
|
||||
// original explicitly (i18n.select_language). Panel closed, the session's
|
||||
// chosen language (window.__pageriteLang) takes over — the pick stays.
|
||||
let overrideLang = null // the ?lang= value in force, null = the session's
|
||||
|
||||
export function setLangOverride(queryLang) {
|
||||
overrideLang = queryLang || null
|
||||
dispatchEvent(new CustomEvent('pagerite:session-lang', { detail: { lang: overrideLang } }))
|
||||
}
|
||||
|
||||
export function runScripts(root) {
|
||||
// Scripts injected via innerHTML do not execute; re-create them.
|
||||
if (!root) return
|
||||
@@ -101,6 +115,11 @@ function swapRegions(doc) {
|
||||
}
|
||||
anchor = imported
|
||||
}
|
||||
// The served language rides on <html> (lang + dir, rtl for e.g. Arabic):
|
||||
// follow the swapped page. The editor panel carries its own lang="en"
|
||||
// dir="ltr", so it is unaffected.
|
||||
document.documentElement.lang = doc.documentElement.lang
|
||||
document.documentElement.dir = doc.documentElement.dir
|
||||
// The editor keeps its own title while open; only inherit the server title
|
||||
// when navigating outside the editor (e.g. fetch-navigation swaps).
|
||||
if (!document.body.classList.contains('editing')) {
|
||||
@@ -111,13 +130,16 @@ function swapRegions(doc) {
|
||||
// Fetch /p, swap its regions into the live page and replaceState to it.
|
||||
// Returns the final URL (after redirects), or null when the fetch did not
|
||||
// yield a page. Category and missing URLs render a placeholder 404 page —
|
||||
// fine to swap in (new pages are created by editing them).
|
||||
// fine to swap in (new pages are created by editing them). The fetch pins
|
||||
// the editor's language override, or — panel closed — the session's chosen
|
||||
// language (window.__pageriteLang).
|
||||
export async function loadPlain(p) {
|
||||
let doc
|
||||
let finalUrl = `/${p}`
|
||||
let html
|
||||
try {
|
||||
const res = await fetch(finalUrl)
|
||||
const pin = overrideLang || window.__pageriteLang
|
||||
const res = await fetch(pin ? `${finalUrl}?lang=${pin}` : finalUrl)
|
||||
const type = res.headers.get('content-type') || ''
|
||||
if (!type.includes('text/html')) return null
|
||||
if (res.redirected) finalUrl = res.url
|
||||
@@ -126,10 +148,16 @@ export async function loadPlain(p) {
|
||||
} catch { return null }
|
||||
if (!doc.getElementById('main')) return null
|
||||
swapRegions(doc)
|
||||
history.replaceState(history.state, '', finalUrl)
|
||||
// The address bar keeps the pretty URL: a language query is a fetch
|
||||
// detail, never shown (pagerite.js's initial ?lang= works the same).
|
||||
const pretty = new URL(finalUrl, location.href)
|
||||
pretty.searchParams.delete('lang')
|
||||
history.replaceState(history.state, '', pretty)
|
||||
runScripts(document.getElementById('page-banner'))
|
||||
runScripts(document.getElementById('main'))
|
||||
// Keep pagerite.js's in-memory page cache in sync with the fresh copy.
|
||||
// The URL is announced as fetched: a language-pinned copy caches under
|
||||
// its own ?lang= key, where navigation with the same pin finds it.
|
||||
dispatchEvent(new CustomEvent('pagerite:page-fetched', { detail: { url: finalUrl, html } }))
|
||||
dispatchEvent(new CustomEvent('pagerite:preview')) // re-inject + re-tuck the edit pens
|
||||
return finalUrl
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
* Configures Vite for FastAPI backend integration:
|
||||
* - Proxies /api/* requests to the FastAPI backend
|
||||
* - Builds to the Python module's frontend-build directory
|
||||
* - Disables Vite's screen clearing on startup
|
||||
*
|
||||
* Options:
|
||||
* paths - Array of paths to proxy (default: ["/api"])
|
||||
@@ -26,6 +27,7 @@ export default function fastapiVue({ paths = ["/api"] } = {}) {
|
||||
return {
|
||||
name: "vite-plugin-fastapi-pagerite",
|
||||
config: () => ({
|
||||
clearScreen: false,
|
||||
server: { proxy },
|
||||
build: {
|
||||
outDir: "../pagerite/frontend-build",
|
||||
|
||||
@@ -16,7 +16,7 @@ const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
|
||||
// https://vite.dev/config/
|
||||
export default defineConfig({
|
||||
plugins: [
|
||||
fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_fonts", "/_a"] }),
|
||||
fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_fonts", "/_a", "/_ws", "/_translate"] }),
|
||||
vue(),
|
||||
vueDevTools(),
|
||||
],
|
||||
@@ -38,7 +38,7 @@ export default defineConfig({
|
||||
chunkSizeWarningLimit: 1200,
|
||||
// Mirror the URL space in the build output: hashed files land under
|
||||
// frontend-build/_assets/ and the Frontend serves the build directory
|
||||
// at the site root (frontend/public/favicon.ico -> /favicon.ico).
|
||||
// at the site root.
|
||||
manifest: true,
|
||||
assetsDir: '_assets',
|
||||
rollupOptions: {
|
||||
@@ -50,6 +50,7 @@ export default defineConfig({
|
||||
main: fileURLToPath(new URL('./src/main.js', import.meta.url)),
|
||||
pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)),
|
||||
analytics: fileURLToPath(new URL('./src/analytics-main.js', import.meta.url)),
|
||||
langselect: fileURLToPath(new URL('./src/langselect-main.js', import.meta.url)),
|
||||
// Only the base CSS is built; theme/banner-design stylesheets live
|
||||
// in pagerite/themes/{name}/ and are served by the backend as-is.
|
||||
pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)),
|
||||
|
||||
+21
-70
@@ -1,74 +1,17 @@
|
||||
"""Command-line entry point for running the backend server."""
|
||||
|
||||
import argparse
|
||||
import gzip
|
||||
import os
|
||||
import sys
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
|
||||
import httpx
|
||||
import msgspec
|
||||
from fastapi_vue import server
|
||||
|
||||
from pagerite.config import Config
|
||||
|
||||
DEFAULT_PORT = 8100
|
||||
DEVMODE = os.getenv("PAGERITE_DEV") == "1"
|
||||
|
||||
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
|
||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
|
||||
|
||||
|
||||
def _download_dbip() -> None:
|
||||
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
|
||||
today = date.today()
|
||||
months = [f"{today:%Y-%m}"]
|
||||
# The current month's file may not be published yet; fall back to last month.
|
||||
prev = (today.replace(day=1) - date.resolution).replace(day=1)
|
||||
months.append(f"{prev:%Y-%m}")
|
||||
|
||||
existing = sorted(
|
||||
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
|
||||
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
|
||||
)
|
||||
if existing and existing[-1] >= months[0]:
|
||||
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download")
|
||||
return
|
||||
|
||||
for month in months:
|
||||
url = DBIP_URL.format(month=month)
|
||||
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
|
||||
tmp = target.with_suffix(".mmdb.gz.tmp")
|
||||
print(f"pagerite: downloading {url}")
|
||||
try:
|
||||
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
|
||||
if r.status_code == 404:
|
||||
continue
|
||||
r.raise_for_status()
|
||||
with open(tmp, "wb") as f:
|
||||
for chunk in r.iter_bytes():
|
||||
f.write(chunk)
|
||||
except httpx.HTTPError as e:
|
||||
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
# Verify it is actually gzip data before installing it.
|
||||
try:
|
||||
with gzip.open(tmp, "rb") as f:
|
||||
f.read(1)
|
||||
except OSError:
|
||||
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
os.replace(tmp, target)
|
||||
# Drop older databases so the app never picks up a stale one.
|
||||
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
|
||||
if old.name != target.name:
|
||||
old.unlink()
|
||||
print(f"pagerite: DB-IP database updated to {target.name}")
|
||||
return
|
||||
print("pagerite: could not download a DB-IP database", file=sys.stderr)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
"""Run the backend server with optional arguments."""
|
||||
@@ -77,9 +20,11 @@ def main() -> None:
|
||||
"hostname",
|
||||
nargs="?",
|
||||
default="localhost",
|
||||
help=("Public hostname of the site; names the data directory "
|
||||
"<hostname>/{content.kantadb, analytics.json, files} under the "
|
||||
"cwd (default: localhost)."),
|
||||
help=(
|
||||
"Public hostname of the site; names the data directory "
|
||||
"<hostname>/{content.kantadb, analytics.json, files} under the "
|
||||
"cwd (default: localhost)."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"-l",
|
||||
@@ -93,18 +38,24 @@ def main() -> None:
|
||||
help="Download/update the DB-IP city lite database before starting.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
# Export the hostname before pagerite.app is imported: it derives the
|
||||
# data directory and public origin from it at import time.
|
||||
os.environ["PAGERITE_HOSTNAME"] = args.hostname
|
||||
if args.dbip:
|
||||
_download_dbip()
|
||||
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
|
||||
# Hand configuration to the app as JSON in PAGERITE_CONFIG; it must be
|
||||
# set before pagerite.app is imported, as state.py reads it at import
|
||||
# time (data directory, public origin).
|
||||
os.environ["PAGERITE_CONFIG"] = msgspec.json.encode(
|
||||
Config(hostname=args.hostname, dbip=args.dbip)
|
||||
).decode()
|
||||
run_args: dict = {}
|
||||
if args.hostname != "localhost":
|
||||
# A public site sits behind TLS on its hostname; show that URL in the
|
||||
# startup box instead of the local listen address.
|
||||
run_args["startup_box"] = f"{{Name}} {{version}}\nhttps://{args.hostname}"
|
||||
server.run(
|
||||
"pagerite.app:app",
|
||||
listen=args.listen,
|
||||
default_port=DEFAULT_PORT,
|
||||
server_header=False,
|
||||
**dev,
|
||||
reload=Path(__file__).parent if DEVMODE else False,
|
||||
**run_args,
|
||||
)
|
||||
|
||||
|
||||
|
||||
+543
-500
File diff suppressed because it is too large
Load Diff
+654
@@ -0,0 +1,654 @@
|
||||
"""Editor REST API and WebSocket sessions.
|
||||
|
||||
The management endpoints behind the SSO forward-auth gate: the site tree
|
||||
(``/_api/pages``), structure operations (``/_api/structure``), site-wide
|
||||
settings (``/_api/settings``), task-list toggles (``/_api/toggle-task``),
|
||||
the translations refresh (``/_api/translations``), the editor session
|
||||
socket (``/_api/ws/editor``), and the translator service channel
|
||||
(``/_translate/{clientkey}`` — deliberately NOT under ``/_api``: the
|
||||
server-generated key in the path is the access control).
|
||||
"""
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import (
|
||||
APIRouter,
|
||||
HTTPException,
|
||||
Request,
|
||||
WebSocket,
|
||||
WebSocketDisconnect,
|
||||
)
|
||||
from pydantic import BaseModel
|
||||
|
||||
from pagerite import i18n, views
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.data import (
|
||||
Node,
|
||||
append_order,
|
||||
find_slot,
|
||||
node_markdown,
|
||||
resolve,
|
||||
sorted_nodes,
|
||||
)
|
||||
from pagerite.markdown import render, toggle_task
|
||||
from pagerite.state import (
|
||||
_check_reserved,
|
||||
_ensure,
|
||||
_invalidate_pages,
|
||||
_remove_page,
|
||||
data,
|
||||
dispatcher,
|
||||
kanta,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
class PageIn(BaseModel):
|
||||
"""Payload for creating or replacing a page."""
|
||||
|
||||
title: str
|
||||
markdown: str
|
||||
published: bool = True
|
||||
banner: str | None = None # None keeps the existing banner
|
||||
|
||||
|
||||
@router.get("/_api/pages")
|
||||
async def list_pages(lang: str | None = None) -> list[dict]:
|
||||
"""The site tree for the structure editor (all nodes, drafts included).
|
||||
|
||||
Nested by slug; each node carries its full path, menu order, flags and
|
||||
language settings (``language`` is the node's own primary-language
|
||||
setting, "" = inherit; ``primary`` is the resolved effective one).
|
||||
With a ``?lang=`` translation, titles come out in that language where a
|
||||
translation exists (``translated`` flags it — true trivially for rows
|
||||
whose primary language IS the selected one; other rows fall back to
|
||||
the original title, dimmed) — the structure itself (slugs, order,
|
||||
hierarchy) is language-independent.
|
||||
"""
|
||||
tag = i18n.base_tag(lang or "")
|
||||
titles = i18n.title_map(data, tag) if tag else {}
|
||||
|
||||
def dump(nodes: dict[str, Node], prefix: str, inherited: str) -> list[dict]:
|
||||
out = []
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
primary = node.language or inherited
|
||||
out.append(
|
||||
{
|
||||
"slug": slug,
|
||||
"path": path,
|
||||
"title": titles.get(path) or node.title,
|
||||
"translated": path in titles or (bool(tag) and primary == tag),
|
||||
"order": node.order,
|
||||
"published": node.published,
|
||||
"has_content": node.chunks is not None,
|
||||
"language": node.language,
|
||||
"primary": primary,
|
||||
"children": dump(node.children, path, primary),
|
||||
}
|
||||
)
|
||||
return out
|
||||
|
||||
return dump(data.menu, "", i18n.ORIGINAL_LANGUAGE)
|
||||
|
||||
|
||||
@router.put("/_api/pages/{path:path}", status_code=204)
|
||||
async def save_page(
|
||||
path: str, page: PageIn, request: Request, lang: str | None = None
|
||||
) -> None:
|
||||
"""Create or replace the page at a slug path ("" or "/" = front page).
|
||||
|
||||
Missing ancestors are created as content-less category labels. Giving
|
||||
a category markdown turns it into a landing page. Empty markdown (after
|
||||
stripping) creates an empty page that renders with just its title —
|
||||
saving never deletes; use DELETE to remove a page (the page editor
|
||||
issues DELETE when you save empty text).
|
||||
|
||||
With a ``?lang=`` query (a translation, not the primary language) the
|
||||
save is a translated-view edit (docs/localization.md): the markdown is
|
||||
diffed against the currently served hybrid and the minimal diff is
|
||||
appended as a Patch under ``patches[f"{path}:{lang}"]`` — node.chunks
|
||||
and the original-language fields (title, published, banner) stay
|
||||
untouched.
|
||||
"""
|
||||
path = path.strip("/")
|
||||
_check_reserved(path)
|
||||
lang = i18n.base_tag(lang or "")
|
||||
if lang and lang != i18n.primary_lang(data.menu, path):
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
with kanta.transaction(
|
||||
f"page:{lang}", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
# Patches alone make the translated version exist.
|
||||
if i18n.add_patch(data, node, path, lang, page.markdown):
|
||||
_invalidate_pages()
|
||||
return
|
||||
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
|
||||
node = _ensure(data.menu, path)
|
||||
node.title = page.title
|
||||
node.chunks = store_chunks(data.chunks, page.markdown)
|
||||
node.published = page.published
|
||||
if page.banner is not None:
|
||||
node.banner = page.banner
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.delete("/_api/pages/{path:path}", status_code=204)
|
||||
async def delete_page(path: str, request: Request) -> None:
|
||||
"""Delete a node by slug path.
|
||||
|
||||
A category (node with children) loses only its landing page and stays
|
||||
as a content-less label; a childless node is removed entirely.
|
||||
"""
|
||||
path = path.strip("/")
|
||||
_check_reserved(path)
|
||||
with kanta.transaction(
|
||||
"page:delete", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
if not _remove_page(data.menu, path):
|
||||
raise HTTPException(404, "no such page")
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
class StructureOp(BaseModel):
|
||||
"""Rearrange the site tree: reorder, move/rename or retitle a node.
|
||||
|
||||
`order` is a fresh fractional key computed client-side from the node's
|
||||
new siblings (a value halfway between them); all other items keep
|
||||
theirs. `move_to` is the full target path — the parent must exist and
|
||||
the new slug be free. Moves carry the whole subtree. The front page is
|
||||
just the top-level node with slug "": renaming it away leaves no front
|
||||
page ("/" then redirects to the first nav item), and any childless
|
||||
top-level node can take the empty slug to become the front page.
|
||||
|
||||
With `lang` (a translation, not the node's primary language) a `title`
|
||||
edit writes a per-language title fragment instead of the original — the
|
||||
same storage as machine title translations (docs/localization.md);
|
||||
sending the original's text removes the override. Structural fields are
|
||||
not combinable with a translated title edit.
|
||||
|
||||
`language` sets the node's primary language (a BCP-47 base tag; "" =
|
||||
inherit from the nearest ancestor, the front page last, site default
|
||||
"en" final — Node.language), inherited by the whole subtree.
|
||||
"""
|
||||
|
||||
path: str
|
||||
order: float | None = None
|
||||
move_to: str | None = None
|
||||
title: str | None = None
|
||||
lang: str | None = None
|
||||
language: str | None = None
|
||||
|
||||
|
||||
@router.post("/_api/structure", status_code=204)
|
||||
async def update_structure(op: StructureOp, request: Request) -> None:
|
||||
"""Apply one structure operation (see StructureOp)."""
|
||||
path = op.path.strip("/")
|
||||
chain = resolve(data.menu, path)
|
||||
if chain is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
node = chain[-1]
|
||||
lang = i18n.base_tag(op.lang or "")
|
||||
if op.language is not None:
|
||||
# Primary-language setting (inherited by the subtree): reselects
|
||||
# what "the original" means for the node — its language is part of
|
||||
# every render, so a change invalidates everywhere.
|
||||
language = i18n.base_tag(op.language)
|
||||
with kanta.transaction(
|
||||
"page:language", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
if language != node.language:
|
||||
node.language = language
|
||||
_invalidate_pages()
|
||||
return
|
||||
if op.title is not None and lang and lang != i18n.primary_lang(data.menu, path):
|
||||
# Translated title (i18n.set_title_translation): original title,
|
||||
# slugs and hierarchy stay untouched.
|
||||
with kanta.transaction(
|
||||
f"page:{lang}:title", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
if i18n.set_title_translation(data, node, lang, op.title):
|
||||
_invalidate_pages()
|
||||
return
|
||||
target = op.move_to.strip("/") if op.move_to is not None else None
|
||||
if target is not None and target != path:
|
||||
_check_reserved(target)
|
||||
if path and target.startswith(f"{path}/"):
|
||||
raise HTTPException(400, "cannot move a page under itself")
|
||||
slot = find_slot(data.menu, target)
|
||||
if slot is None:
|
||||
raise HTTPException(404, "target parent does not exist")
|
||||
tnodes, tslug = slot
|
||||
if tslug in tnodes:
|
||||
raise HTTPException(400, "target path exists")
|
||||
if not tslug and node.children:
|
||||
raise HTTPException(400, "the front page cannot have children")
|
||||
# One structure call can combine a title set, a move/rename and a
|
||||
# reorder; the action names the most significant of them.
|
||||
action = (
|
||||
"page:slug"
|
||||
if target is not None and target != path
|
||||
else "page:title"
|
||||
if op.title is not None
|
||||
else "structure:reorder"
|
||||
)
|
||||
with kanta.transaction(action, user=request.headers.get("remote-user"), extra=path):
|
||||
if op.title is not None:
|
||||
node.title = op.title
|
||||
if target is not None and target != path:
|
||||
snodes, sslug = find_slot(data.menu, path)
|
||||
del snodes[sslug]
|
||||
# A pure rename (same parent) keeps its position; only a move
|
||||
# to another level appends at the end (unless an order came
|
||||
# with the drop).
|
||||
same_level = path.rpartition("/")[0] == target.rpartition("/")[0]
|
||||
node.order = (
|
||||
op.order
|
||||
if op.order is not None
|
||||
else node.order
|
||||
if same_level
|
||||
else append_order(tnodes)
|
||||
)
|
||||
tnodes[tslug] = node
|
||||
elif op.order is not None:
|
||||
node.order = op.order
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.get("/_api/settings")
|
||||
async def get_settings() -> dict:
|
||||
"""Site-wide settings (brand, theme, custom CSS and favicon URL), plus
|
||||
the themes, banner designs and user fonts available on disk for the
|
||||
selectors, the translator service keys and the wanted translation
|
||||
languages (for the /_translate socket)."""
|
||||
return {
|
||||
"brand": data.brand,
|
||||
"brand_html": data.brand_html,
|
||||
"theme": data.theme,
|
||||
"custom_css": data.custom_css,
|
||||
"favicon": f"/_f/{data.favicon}" if data.favicon else "",
|
||||
"themes": views._theme_info(),
|
||||
"banner_designs": views._banner_design_names(),
|
||||
"fonts": views._user_fonts(),
|
||||
"transition": data.transition,
|
||||
"transitions": views._transition_names(),
|
||||
"translate_keys": data.translate_keys,
|
||||
# The site default primary language: the front page's resolved
|
||||
# setting (every page may override it, inherited down the tree).
|
||||
"primary_lang": i18n.primary_lang(data.menu, ""),
|
||||
"translate_langs": sorted(data.translate_langs),
|
||||
}
|
||||
|
||||
|
||||
class SettingsIn(BaseModel):
|
||||
"""Payload for updating site-wide settings."""
|
||||
|
||||
brand: str
|
||||
theme: str
|
||||
custom_css: str
|
||||
brand_html: str = ""
|
||||
transition: str = "cube"
|
||||
translate_langs: list[str] | None = None # None keeps the current set
|
||||
translate_keys: dict[str, str] | None = None # None keeps the current keys
|
||||
|
||||
|
||||
@router.put("/_api/settings", status_code=204)
|
||||
async def put_settings(settings: SettingsIn, request: Request) -> None:
|
||||
"""Update site-wide settings; invalidates cached pages and ETags."""
|
||||
with kanta.transaction("settings", user=request.headers.get("remote-user")):
|
||||
data.brand = settings.brand
|
||||
data.brand_html = settings.brand_html
|
||||
data.theme = settings.theme
|
||||
data.custom_css = settings.custom_css
|
||||
data.transition = settings.transition
|
||||
if settings.translate_langs is not None:
|
||||
# Any language may be a target — including the site default
|
||||
# (an article in another language can be translated INTO it);
|
||||
# a node's own primary is excluded per article, not here.
|
||||
data.translate_langs = {
|
||||
tag: True
|
||||
for lang in settings.translate_langs
|
||||
if (tag := i18n.base_tag(lang))
|
||||
}
|
||||
if settings.translate_keys is not None:
|
||||
data.translate_keys = settings.translate_keys
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.delete("/_api/translations", status_code=204)
|
||||
async def delete_translations(request: Request) -> None:
|
||||
"""Drop all machine translations (Data.trans) so the dispatcher
|
||||
re-translates everything from scratch (a translate:reset action:
|
||||
the invalidation hook re-offers every fragment to connected
|
||||
translators). User patches are kept; the availability index
|
||||
(node.langs) is rebuilt from them — patches alone still make a language
|
||||
exist on a page."""
|
||||
with kanta.transaction("translate:reset", user=request.headers.get("remote-user")):
|
||||
i18n.clear_translations(data)
|
||||
_invalidate_pages()
|
||||
# Fragments rejected this run (segment validation) stay skipped no
|
||||
# longer: a refresh is precisely the "another chance" for them.
|
||||
dispatcher.reset_validation_failures()
|
||||
|
||||
|
||||
class ToggleTaskIn(BaseModel):
|
||||
"""Payload for toggling one task-list checkbox."""
|
||||
|
||||
path: str
|
||||
index: int
|
||||
markdown: str | None = None
|
||||
|
||||
|
||||
@router.post("/_api/toggle-task")
|
||||
async def toggle_task_endpoint(body: ToggleTaskIn, request: Request) -> dict[str, str]:
|
||||
"""Toggle the Nth task-list checkbox in a page's Markdown source.
|
||||
|
||||
If ``markdown`` is provided the source is left untouched and the toggled
|
||||
Markdown is returned (used while the page editor is open, so the live
|
||||
CodeMirror document can be updated). Otherwise the stored page at
|
||||
``path`` is read, toggled, and saved.
|
||||
"""
|
||||
path = body.path.strip("/")
|
||||
_check_reserved(path)
|
||||
if body.markdown is not None:
|
||||
new_markdown = toggle_task(body.markdown, body.index)
|
||||
if new_markdown is None:
|
||||
raise HTTPException(400, "invalid task index")
|
||||
return {"markdown": new_markdown}
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
|
||||
if new_markdown is None:
|
||||
raise HTTPException(400, "invalid task index")
|
||||
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
|
||||
# Re-chunk like any save: only the chunk containing the toggled
|
||||
# checkbox gets a new hash, the rest keep theirs.
|
||||
node.chunks = store_chunks(data.chunks, new_markdown)
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
return {"markdown": new_markdown}
|
||||
|
||||
|
||||
# WebSocket API for external translation services (not under /_api: it is keyed
|
||||
# with Data.translate_keys instead of the SSO forward-auth). The dispatcher —
|
||||
# protocol, connected clients and the job pipeline — lives in translate.py.
|
||||
@router.websocket("/_translate/{clientkey}")
|
||||
async def translate_ws(ws: WebSocket, clientkey: str) -> None:
|
||||
"""Translator service channel (docs/localization.md).
|
||||
|
||||
Deliberately NOT under /_api/: the external forward-auth is skipped;
|
||||
the server-generated client key in the path is the access control
|
||||
(``Data.translate_keys``: key -> display name; the first is generated
|
||||
at bootstrap, all are shown in the admin's /_api/settings).
|
||||
"""
|
||||
await dispatcher.handle_ws(ws, clientkey)
|
||||
|
||||
|
||||
@router.websocket("/_api/ws/editor")
|
||||
async def editor_ws(ws: WebSocket) -> None:
|
||||
"""Editor session: open pages, render previews, save — over one socket.
|
||||
|
||||
Stateless protocol (each message carries the path):
|
||||
<- {"type": "open", "path", "lang"?}
|
||||
-> {"type": "doc", "path", "exists", "title", "markdown", "published",
|
||||
"banner", "banner_design", "lang", "primary_lang", "langs",
|
||||
"translate_langs"}
|
||||
<- {"type": "render", "path", "markdown"}
|
||||
-> {"type": "html", "path", "html"}
|
||||
<- {"type": "save", "path", "title"?, "markdown"?, "published"?,
|
||||
"banner"?, "banner_design"?, "move_from"?, "lang"?, "base"?}
|
||||
(absent fields keep their old values; move_from: rename/move a
|
||||
page, subtree included)
|
||||
-> {"type": "saved", "path"} | {"type": "error", "detail"}
|
||||
|
||||
With "lang" (a translation, not the primary language), open returns the
|
||||
effective hybrid Markdown and title for that language plus the language
|
||||
metadata the picker's UI needs; save diffs the submitted Markdown
|
||||
against "base" (the editor's shadow copy of the hybrid it started from
|
||||
— absent: the current hybrid) and stores it as a user Patch, and a
|
||||
changed title becomes a fragment in Data.trans — node.chunks and the
|
||||
other fields stay untouched (docs/localization.md).
|
||||
"""
|
||||
await ws.accept()
|
||||
try:
|
||||
while True:
|
||||
msg = await ws.receive_json()
|
||||
path = msg.get("path", "").strip("/")
|
||||
try:
|
||||
_check_reserved(path)
|
||||
except HTTPException:
|
||||
await ws.send_json({"type": "error", "detail": "reserved path"})
|
||||
continue
|
||||
match msg.get("type"):
|
||||
case "open":
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
# The article's primary language: its own setting,
|
||||
# inherited down the tree ("en" final fallback).
|
||||
node_lang = i18n.primary_lang(data.menu, path)
|
||||
lang = i18n.base_tag(str(msg.get("lang") or ""))
|
||||
if lang == node_lang:
|
||||
lang = ""
|
||||
markdown = ""
|
||||
title = node.title if node else ""
|
||||
if node is not None:
|
||||
markdown = node_markdown(data, node) or ""
|
||||
if lang and node.chunks is not None:
|
||||
# Translation view: the effective (hybrid)
|
||||
# Markdown and title for that language —
|
||||
# machine fragments + user patches over the
|
||||
# original (docs/localization.md editor flow).
|
||||
markdown = i18n.hybrid_markdown(data, node, path, lang)
|
||||
title = i18n.title_map(data, lang).get(path) or title
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "doc",
|
||||
"path": path,
|
||||
"exists": node is not None,
|
||||
"title": title,
|
||||
"markdown": markdown,
|
||||
"published": node.published if node else True,
|
||||
"banner": node.banner if node else "",
|
||||
# Own banner design setting: null = inherit,
|
||||
# "" = none, otherwise a design name.
|
||||
"banner_design": node.banner_design if node else None,
|
||||
# Which node's banner applies here ("" = front page,
|
||||
# null = default artwork); the site editor shows it
|
||||
# as the banner field's placeholder.
|
||||
"banner_from": views.banner_source(data.menu, path),
|
||||
# Which node's banner-design setting would apply on
|
||||
# inherit ("" = front page, null = the active
|
||||
# theme's default) and what design that resolves to.
|
||||
"banner_design_from": (
|
||||
src := views.banner_design_source(
|
||||
data.menu, path, data.theme
|
||||
)
|
||||
),
|
||||
"banner_design_inherited": (
|
||||
views.banner_design(data.menu, src, data.theme)
|
||||
if src is not None
|
||||
else views.theme_banner_design(data.theme)
|
||||
),
|
||||
# Language context for the editor's picker: the
|
||||
# language this Markdown represents ("" = primary),
|
||||
# the page's own primary language, the translations
|
||||
# this page already has, and the site-wide
|
||||
# configured target languages.
|
||||
"lang": lang,
|
||||
"primary_lang": node_lang,
|
||||
"langs": sorted(node.langs) if node else [],
|
||||
"translate_langs": sorted(data.translate_langs),
|
||||
}
|
||||
)
|
||||
case "render":
|
||||
markdown = msg.get("markdown", "")
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
rendered = render(
|
||||
markdown,
|
||||
path,
|
||||
node.created if node else None,
|
||||
node.modified if node else None,
|
||||
# The title is injected as h1 when the markdown has
|
||||
# none; the editor's title field edits live-preview.
|
||||
title=msg.get("title") or (node.title if node else ""),
|
||||
# Pin section anchors to the original language so the
|
||||
# preview of a translation matches the served page
|
||||
# (no-op when the previewed markdown is the original).
|
||||
anchors_from=(
|
||||
(node_markdown(data, node) or "", node.title)
|
||||
if node
|
||||
else None
|
||||
),
|
||||
)
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "html",
|
||||
"path": path,
|
||||
"html": rendered.html,
|
||||
# Column-layout flag: the preview toggles the
|
||||
# article's .multicol class and swaps in the
|
||||
# segmented (.colseg/.cols) article html.
|
||||
"multicol": rendered.multicol,
|
||||
}
|
||||
)
|
||||
case "save":
|
||||
move_from = (msg.get("move_from") or path).strip("/")
|
||||
lang = i18n.base_tag(str(msg.get("lang") or ""))
|
||||
translated = bool(
|
||||
lang and lang != i18n.primary_lang(data.menu, move_from)
|
||||
)
|
||||
try:
|
||||
_check_reserved(move_from)
|
||||
except HTTPException:
|
||||
await ws.send_json({"type": "error", "detail": "reserved path"})
|
||||
continue
|
||||
old_chain = resolve(data.menu, move_from)
|
||||
old = old_chain[-1] if old_chain else None
|
||||
if old is None and move_from != path:
|
||||
move_from = path # nothing to carry over; plain save
|
||||
if move_from != path:
|
||||
# Rename/move: detach the node (subtree included)
|
||||
# and attach it at the new path. The target slug
|
||||
# must be free and the front page childless.
|
||||
if move_from and path.startswith(f"{move_from}/"):
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "cannot move a page under itself",
|
||||
}
|
||||
)
|
||||
continue
|
||||
tslug = path.rpartition("/")[2]
|
||||
if not tslug and old.children:
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "the front page cannot have children",
|
||||
}
|
||||
)
|
||||
continue
|
||||
tchain = resolve(data.menu, path)
|
||||
if tchain is not None:
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "target path exists",
|
||||
}
|
||||
)
|
||||
continue
|
||||
if translated and (
|
||||
move_from != path or old is None or old.chunks is None
|
||||
):
|
||||
# A translated-view save patches an existing
|
||||
# original; it cannot create or move pages.
|
||||
await ws.send_json({"type": "error", "detail": "no such page"})
|
||||
continue
|
||||
if translated and "markdown" in msg and not msg["markdown"].strip():
|
||||
# Saving never deletes; an emptied translation would
|
||||
# render as a blank page in that language.
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "a translation cannot be emptied",
|
||||
}
|
||||
)
|
||||
continue
|
||||
with kanta.transaction(
|
||||
f"page:{lang}" if translated else "page",
|
||||
user=ws.headers.get("remote-user"),
|
||||
extra=path,
|
||||
):
|
||||
if move_from != path:
|
||||
same_menu = (
|
||||
move_from.rpartition("/")[0] == path.rpartition("/")[0]
|
||||
)
|
||||
snodes, sslug = find_slot(data.menu, move_from)
|
||||
node = snodes.pop(sslug)
|
||||
parent = path.rpartition("/")[0]
|
||||
if parent:
|
||||
_ensure(data.menu, parent)
|
||||
tnodes, tslug = find_slot(data.menu, path)
|
||||
node.order = (
|
||||
node.order if same_menu else append_order(tnodes)
|
||||
)
|
||||
tnodes[tslug] = node
|
||||
else:
|
||||
node = old if old is not None else _ensure(data.menu, path)
|
||||
if translated:
|
||||
# node.chunks and the original-language fields
|
||||
# stay untouched: the markdown diff (against the
|
||||
# editor's shadow "base" — the hybrid it started
|
||||
# from; absent: the current hybrid) is appended
|
||||
# as a Patch, a changed title becomes a
|
||||
# per-language title override (i18n).
|
||||
changed = False
|
||||
if "markdown" in msg:
|
||||
base = msg.get("base")
|
||||
changed = i18n.add_patch(
|
||||
data,
|
||||
node,
|
||||
path,
|
||||
lang,
|
||||
msg["markdown"],
|
||||
base=base if isinstance(base, str) else None,
|
||||
)
|
||||
if "title" in msg and node.title:
|
||||
changed = (
|
||||
i18n.set_title_translation(
|
||||
data, node, lang, msg["title"]
|
||||
)
|
||||
or changed
|
||||
)
|
||||
if changed:
|
||||
_invalidate_pages()
|
||||
else:
|
||||
if "markdown" in msg:
|
||||
# Saving never deletes; empty markdown is an
|
||||
# empty page. Deletion is an explicit choice
|
||||
# by the page editor (REST DELETE).
|
||||
node.chunks = store_chunks(data.chunks, msg["markdown"])
|
||||
if "title" in msg:
|
||||
node.title = msg["title"]
|
||||
if "published" in msg:
|
||||
node.published = bool(msg["published"])
|
||||
if "banner" in msg:
|
||||
node.banner = msg["banner"]
|
||||
if "banner_design" in msg:
|
||||
node.banner_design = msg["banner_design"]
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
await ws.send_json({"type": "saved", "path": path})
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
+79
-1513
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,180 @@
|
||||
"""Block-level Markdown chunking for content-addressed storage.
|
||||
|
||||
A page's Markdown is split into deterministic block-level chunks, each
|
||||
stored once under its content hash in ``Data.chunks`` (docs/migrate.md).
|
||||
Shared by the render/save pipeline (app.py, views.py, i18n.py) and the
|
||||
schema migration (migrations.py), so a chunk's key is stable no matter
|
||||
where the split happens.
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
import blake3
|
||||
|
||||
from pagerite.segments import has_prose
|
||||
|
||||
#: Fenced code block opener/closer: up to 3 spaces indent, then 3+
|
||||
#: backticks or tildes (CommonMark).
|
||||
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
|
||||
|
||||
#: A container fence line (mdit-py-plugins container): the "::: aside"
|
||||
#: opener and the ":::" closer alike. Always its own block, even with no
|
||||
#: blank line around it: folded into a prose paragraph it would cross to
|
||||
#: the translator as part of the text run, where the model can drop it —
|
||||
#: the rest of the page then renders inside the container.
|
||||
_CONTAINER = re.compile(r"^ {0,3}:{3,}(?:[ \t]|$)")
|
||||
|
||||
#: HTML block openers that may span blank lines (CommonMark types 1-5:
|
||||
#: script/pre/style/textarea, comments, processing instructions,
|
||||
#: declarations, CDATA) with their closing condition. Other HTML blocks
|
||||
#: end at the first blank line, which the generic blank-line split
|
||||
#: already does.
|
||||
_HTML_ATOMIC = (
|
||||
(
|
||||
re.compile(r"^ {0,3}<(?:script|pre|style|textarea)(?:\s|>|$)", re.I),
|
||||
re.compile(r"</(?:script|pre|style|textarea)\s*>", re.I),
|
||||
),
|
||||
(re.compile(r"^ {0,3}<!--"), re.compile(r"-->")),
|
||||
(re.compile(r"^ {0,3}<\?"), re.compile(r"\?>")),
|
||||
(re.compile(r"^ {0,3}<!\[CDATA\["), re.compile(r"\]\]>")),
|
||||
(re.compile(r"^ {0,3}<![A-Za-z]"), re.compile(r">")),
|
||||
)
|
||||
|
||||
#: First line of a generic HTML block (a block-level tag).
|
||||
_HTML_TAG = re.compile(r"^ {0,3}</?[A-Za-z][^>]*>")
|
||||
|
||||
|
||||
def _fence_close(line: str, opener: str) -> bool:
|
||||
"""True when ``line`` closes a code fence opened by ``opener``: the
|
||||
same marker char, at least as many, and nothing else on the line."""
|
||||
stripped = line.strip()
|
||||
return (
|
||||
len(stripped) >= len(opener)
|
||||
and stripped[0] == opener[0]
|
||||
and set(stripped) == {opener[0]}
|
||||
)
|
||||
|
||||
|
||||
def chunk_markdown(markdown: str) -> list[str]:
|
||||
"""Split Markdown into block-level chunks, deterministically.
|
||||
|
||||
Blocks are separated by blank lines; fenced code blocks and the
|
||||
multi-line HTML blocks (comments, script/pre/style, CDATA...) are
|
||||
kept atomic, even across blank lines, and end at their closing
|
||||
condition. Container fence lines (:::, open and close alike) are
|
||||
always their own block, blank lines or not (see _CONTAINER). Chunks
|
||||
carry no surrounding blank lines and no trailing newline; rejoining
|
||||
with ``join_chunks`` reproduces the source modulo blank-line
|
||||
normalization.
|
||||
"""
|
||||
chunks: list[str] = []
|
||||
buf: list[str] = []
|
||||
fence = "" # opener marker of the code fence we are in ("" = outside)
|
||||
html_end: re.Pattern | None = None # closes the atomic HTML block we are in
|
||||
|
||||
def flush() -> None:
|
||||
text = "\n".join(buf).strip("\n")
|
||||
if text.strip():
|
||||
chunks.append(text)
|
||||
buf.clear()
|
||||
|
||||
for line in markdown.split("\n"):
|
||||
if fence:
|
||||
buf.append(line)
|
||||
if _fence_close(line, fence):
|
||||
fence = ""
|
||||
flush()
|
||||
continue
|
||||
if html_end is not None:
|
||||
buf.append(line)
|
||||
if html_end.search(line):
|
||||
html_end = None
|
||||
flush()
|
||||
continue
|
||||
if not line.strip():
|
||||
flush()
|
||||
continue
|
||||
if m := _FENCE_OPEN.match(line):
|
||||
# Fences interrupt paragraphs (CommonMark): start a new block.
|
||||
flush()
|
||||
fence = m.group(1)
|
||||
buf.append(line)
|
||||
continue
|
||||
if _CONTAINER.match(line):
|
||||
# Container fence lines (open and close alike) are their own
|
||||
# block — never part of a prose chunk (see _CONTAINER).
|
||||
flush()
|
||||
buf.append(line)
|
||||
flush()
|
||||
continue
|
||||
if not buf:
|
||||
for open_re, close_re in _HTML_ATOMIC:
|
||||
if open_re.match(line):
|
||||
buf.append(line)
|
||||
if close_re.search(line): # opens and closes on one line
|
||||
flush()
|
||||
else:
|
||||
html_end = close_re
|
||||
break
|
||||
else:
|
||||
buf.append(line)
|
||||
continue
|
||||
buf.append(line)
|
||||
flush() # an unterminated fence/HTML block runs to EOF, kept as code/HTML
|
||||
return chunks
|
||||
|
||||
|
||||
def _normalize(text: str) -> str:
|
||||
"""Whitespace-insensitive chunk identity: strip trailing whitespace
|
||||
per line and collapse surrounding blank lines, so whitespace-only
|
||||
source edits don't invalidate translations."""
|
||||
return "\n".join(line.rstrip() for line in text.split("\n")).strip("\n")
|
||||
|
||||
|
||||
def chunk_key(text: str) -> bytes:
|
||||
"""Content key of a chunk: the first 9 bytes of the blake3 digest of
|
||||
the normalized text (72 bits — a site's chunk count stays far below
|
||||
the birthday bound), using the same hasher as app.py's file store.
|
||||
|
||||
Keys are bytes: kanta/msgspec base64-encode them at the JSON
|
||||
persistence level, so the raw database dicts carry 12-char strings.
|
||||
"""
|
||||
return blake3.blake3(_normalize(text).encode()).digest(9)
|
||||
|
||||
|
||||
def needs_translation(chunk: str) -> bool:
|
||||
"""False for chunks without prose: pure code fences, HTML blocks, and
|
||||
anything that yields no translatable segments (pagerite/segments.py) —
|
||||
container fences, lone {placeholders}, reference definitions.
|
||||
|
||||
These are inherently no-translate (docs/migrate.md): derived from the
|
||||
chunk text itself, nothing is stored. Every language renders them from
|
||||
the original chunk via the hybrid fallback.
|
||||
"""
|
||||
if _FENCE_OPEN.match(chunk):
|
||||
return False
|
||||
first = chunk.split("\n", 1)[0]
|
||||
if any(open_re.match(first) for open_re, _ in _HTML_ATOMIC):
|
||||
return False
|
||||
if _HTML_TAG.match(first):
|
||||
return False
|
||||
return has_prose(chunk)
|
||||
|
||||
|
||||
def join_chunks(chunks: list[str]) -> str:
|
||||
"""The stored page form of chunks: blocks joined by a blank line,
|
||||
with a trailing newline ("" for no chunks)."""
|
||||
return "\n\n".join(chunks) + "\n" if chunks else ""
|
||||
|
||||
|
||||
def store_chunks(store: dict[bytes, str], markdown: str) -> list[bytes]:
|
||||
"""Chunk ``markdown`` into ``store`` (hash -> text); return the ordered
|
||||
hashes. Unchanged chunks keep their hashes, so only genuinely new text
|
||||
lands in the kanta change diff. First writer wins: variants sharing a
|
||||
key differ only in insignificant whitespace (see chunk_key)."""
|
||||
hashes = []
|
||||
for chunk in chunk_markdown(markdown):
|
||||
key = chunk_key(chunk)
|
||||
store.setdefault(key, chunk)
|
||||
hashes.append(key)
|
||||
return hashes
|
||||
@@ -0,0 +1,28 @@
|
||||
"""CLI → app configuration, passed as JSON in the ``PAGERITE_CONFIG`` env var.
|
||||
|
||||
Kept dependency-free (msgspec only) so ``__main__`` can build and serialize
|
||||
the config before any app module is imported, and the app side parses the
|
||||
same struct back. Import-time safe: nothing here reads the environment
|
||||
until ``load()`` is called.
|
||||
"""
|
||||
|
||||
import os
|
||||
|
||||
import msgspec
|
||||
|
||||
|
||||
class Config(msgspec.Struct):
|
||||
"""Configuration passed from the CLI entry point to the app."""
|
||||
|
||||
#: Public hostname of the site; names the per-site data directory
|
||||
#: ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
|
||||
hostname: str = "localhost"
|
||||
#: Download/update the DB-IP city lite database at startup (--dbip).
|
||||
dbip: bool = False
|
||||
|
||||
|
||||
def load() -> Config:
|
||||
"""Parse ``PAGERITE_CONFIG``, or the defaults when unset."""
|
||||
if raw := os.getenv("PAGERITE_CONFIG"):
|
||||
return msgspec.json.decode(raw.encode(), type=Config)
|
||||
return Config()
|
||||
+67
-7
@@ -2,8 +2,9 @@
|
||||
|
||||
The site structure is a tree of Nodes. Every node is a menu label with a
|
||||
configurable title and slug (its key in the parent's ``children``); the
|
||||
URL path is the chain of slugs from the top level. ``content`` is the
|
||||
node's Markdown page, or None for a pure category label, whose URL renders
|
||||
URL path is the chain of slugs from the top level. ``chunks`` is the
|
||||
node's Markdown page as ordered content-hash keys into ``Data.chunks``
|
||||
(docs/migrate.md), or None for a pure category label, whose URL renders
|
||||
a placeholder page while nav links point at its first child.
|
||||
"""
|
||||
|
||||
@@ -11,6 +12,16 @@ from datetime import UTC, datetime
|
||||
|
||||
import msgspec
|
||||
|
||||
from pagerite.chunks import join_chunks
|
||||
|
||||
|
||||
class Patch(msgspec.Struct, omit_defaults=True):
|
||||
"""One editing session's overrides on a translated view, applied
|
||||
independently per hunk (docs/localization.md)."""
|
||||
|
||||
#: (search, replace) pairs on the served hybrid Markdown.
|
||||
hunks: list[tuple[str, str]] = []
|
||||
|
||||
|
||||
class Node(msgspec.Struct, omit_defaults=True):
|
||||
"""One item of the site hierarchy.
|
||||
@@ -28,9 +39,21 @@ class Node(msgspec.Struct, omit_defaults=True):
|
||||
|
||||
title: str = ""
|
||||
order: float = 0
|
||||
#: Markdown source of the node's page; None = pure category label
|
||||
#: (its URL renders a placeholder page).
|
||||
content: str | None = None
|
||||
#: Ordered chunk hashes (9-byte keys into ``Data.chunks``); None =
|
||||
#: pure category label (its URL renders a placeholder page), a list
|
||||
#: (possibly empty) = a page.
|
||||
chunks: list[bytes] | None = None
|
||||
#: Primary language of the article (BCP-47 base tag). "" = inherit
|
||||
#: (nearest ancestor, front page last, site default "en" final).
|
||||
language: str = ""
|
||||
#: Chunk hashes the editor marked "do not translate" (always served
|
||||
#: from the original). Presence-keys, value always True.
|
||||
no_trans: dict[bytes, bool] = {}
|
||||
#: Languages this article is available in (besides its primary
|
||||
#: language). Presence-keys, value always True — the availability
|
||||
#: index for rendering and language selection; maintained by whoever
|
||||
#: writes translation data (docs/migrate.md).
|
||||
langs: dict[str, bool] = {}
|
||||
#: Raw HTML for the header banner (img, styled div, canvas+script...),
|
||||
#: rendered after the banner design's artwork so author code always
|
||||
#: wins over the design's own styles.
|
||||
@@ -75,9 +98,46 @@ class Data(msgspec.Struct):
|
||||
#: Trusted author content; not sanitized.
|
||||
custom_css: str = ""
|
||||
#: Favicon: content-addressed file name (served at "/_f/{name}"),
|
||||
#: linked as <link rel="icon"> on every page. Empty = the build's
|
||||
#: /favicon.ico.
|
||||
#: linked as <link rel="icon"> on every page; /favicon.ico redirects
|
||||
#: to it. Empty = no icon (and /favicon.ico 404s).
|
||||
favicon: str = ""
|
||||
#: API keys gating the translator service WebSocket (/_translate/{key};
|
||||
#: the external forward-auth does not cover that route): key -> display
|
||||
#: name. Keys are 12 lowercase alphanumeric characters; the first is
|
||||
#: generated at database bootstrap, more are managed in the editor
|
||||
#: shell's lang tab (via /_api/settings).
|
||||
translate_keys: dict[str, str] = {}
|
||||
#: Wanted target languages for the translator service (presence-keys,
|
||||
#: value always True). The dispatcher offers jobs only in the
|
||||
#: intersection of these and a connection's announced capabilities.
|
||||
#: Bootstrapped to es+zh; edited in the editor shell's localization
|
||||
#: tab (or via /_api/settings).
|
||||
translate_langs: dict[str, bool] = {}
|
||||
#: All original-language page text, content-addressed:
|
||||
#: chunk_key (9 bytes; base64 at the JSON level) -> Markdown chunk.
|
||||
#: Shared by every article.
|
||||
chunks: dict[bytes, str] = {}
|
||||
#: Machine translations: chunk hash -> lang -> translated Markdown
|
||||
#: (a nested dict rather than tuple keys, which msgspec's JSON
|
||||
#: serializer does not support). Also used for node titles (hash of
|
||||
#: the title text).
|
||||
trans: dict[bytes, dict[str, str]] = {}
|
||||
#: User override patches per article and language:
|
||||
#: f"{path}:{lang}" -> ordered patches (paths without leading slash).
|
||||
patches: dict[str, list[Patch]] = {}
|
||||
|
||||
|
||||
def node_markdown(data: Data, node: Node) -> str | None:
|
||||
"""The node's original Markdown assembled from the chunk store.
|
||||
|
||||
None for category labels (chunks is None); an empty page gives "".
|
||||
Hashes missing from the store (shouldn't happen) are skipped.
|
||||
"""
|
||||
if node.chunks is None:
|
||||
return None
|
||||
return join_chunks(
|
||||
[t for h in node.chunks if (t := data.chunks.get(h)) is not None]
|
||||
)
|
||||
|
||||
|
||||
def prettify(slug: str) -> str:
|
||||
|
||||
@@ -0,0 +1,391 @@
|
||||
"""Content-addressed file store, image derivatives, and file routes.
|
||||
|
||||
``FileStore`` keeps uploads, seed assets and fetched favicons on disk under
|
||||
hash-prefixed names, fully cached in RAM (uncompressed plus a zstd copy
|
||||
when compression shrinks the body), served immutable at ``/_f/``. Raster
|
||||
images and SVGs are recompressed into AVIF/WebP/JPEG derivatives
|
||||
(``store_image`` and helpers); the untouched original is kept alongside as
|
||||
``<hash>.orig<ext>`` (never served). Routes: upload/delete under
|
||||
``/_api/files``, the favicon settings endpoints, the /favicon.ico
|
||||
redirect to the configured icon, the ``/_f/`` server with
|
||||
Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import mimetypes
|
||||
import tempfile
|
||||
from contextlib import suppress
|
||||
from pathlib import Path
|
||||
|
||||
import blake3
|
||||
from fastapi import APIRouter, HTTPException, Request
|
||||
from fastapi.responses import RedirectResponse, Response
|
||||
from mediapreview import dispatch
|
||||
|
||||
from pagerite import views
|
||||
from pagerite.state import (
|
||||
FAVICON_MAXSIZE,
|
||||
FILES_DIR,
|
||||
IMAGE_JPG_QUALITY,
|
||||
IMAGE_MAXSIZE,
|
||||
IMAGE_QUALITY,
|
||||
IMAGE_WEBP_QUALITY,
|
||||
_invalidate_pages,
|
||||
_zstd,
|
||||
data,
|
||||
kanta,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# mediapreview logs pyvips noise ("VipsForeignSaveJpegTarget argument strip is
|
||||
# deprecated", "threadpool completed with N workers") at INFO; keep warnings.
|
||||
logging.getLogger("mediapreview").setLevel(logging.WARNING)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
class FileStore:
|
||||
"""Content-addressed files on disk, fully cached in RAM.
|
||||
|
||||
Every file is kept in RAM uncompressed and zstd-compressed (the
|
||||
compressed copy only when it actually shrinks the body), so ``/_f``
|
||||
serves both encodings without touching disk or re-compressing.
|
||||
"""
|
||||
|
||||
def __init__(self, path: Path) -> None:
|
||||
self.path = path
|
||||
#: name -> (uncompressed body, zstd body or None)
|
||||
self._cache: dict[str, tuple[bytes, bytes | None]] = {}
|
||||
|
||||
@staticmethod
|
||||
def _entry(body: bytes) -> tuple[bytes, bytes | None]:
|
||||
compressed = _zstd.compress(body)
|
||||
return body, compressed if len(compressed) < len(body) else None
|
||||
|
||||
def load(self) -> None:
|
||||
"""Read every stored file into the RAM cache (startup)."""
|
||||
try:
|
||||
entries = sorted(self.path.iterdir())
|
||||
except FileNotFoundError:
|
||||
return
|
||||
for f in entries:
|
||||
if f.is_file() and not f.name.startswith("."):
|
||||
self._cache.setdefault(f.name, self._entry(f.read_bytes()))
|
||||
|
||||
def get(self, name: str) -> tuple[bytes, bytes | None] | None:
|
||||
return self._cache.get(name)
|
||||
|
||||
def put(self, name: str, body: bytes) -> None:
|
||||
"""Store ``body`` under ``name`` on disk and in the RAM cache."""
|
||||
if name in self._cache:
|
||||
return
|
||||
self.path.mkdir(parents=True, exist_ok=True)
|
||||
(self.path / name).write_bytes(body)
|
||||
self._cache[name] = self._entry(body)
|
||||
|
||||
def delete(self, name: str) -> None:
|
||||
"""Delete a file plus its derivatives/original counterparts, if any.
|
||||
|
||||
An image upload is stored as a group sharing the hash prefix
|
||||
(``<hash>.orig.<ext>`` + ``<hash>.avif/.webp/.jpg``); deleting any
|
||||
of the names removes them all.
|
||||
"""
|
||||
stem = name.partition(".")[0]
|
||||
for key in [k for k in self._cache if k.partition(".")[0] == stem]:
|
||||
self._cache.pop(key, None)
|
||||
with suppress(FileNotFoundError):
|
||||
(self.path / key).unlink()
|
||||
|
||||
def __contains__(self, name: str) -> bool:
|
||||
return name in self._cache
|
||||
|
||||
|
||||
file_store = FileStore(FILES_DIR)
|
||||
|
||||
|
||||
def _ext(orig: str) -> str:
|
||||
"""Sanitized lowercase extension (with dot) of an original file name."""
|
||||
return "".join(c for c in Path(orig).suffix.lower() if c.isalnum() or c == ".")
|
||||
|
||||
|
||||
def _hash_name(body: bytes, orig: str) -> str:
|
||||
"""Content-addressed file name: blake3 hash prefix + original extension."""
|
||||
return blake3.blake3(body).hexdigest()[:12] + _ext(orig)
|
||||
|
||||
|
||||
def _to_avif(body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE) -> bytes | None:
|
||||
"""Recompress an image body to a thumbnailed AVIF via mediapreview's
|
||||
dispatch (pyvips for common formats, ffmpeg for HEIC/HEIF/AVIF), or
|
||||
None if the body is not a decodable image (stored as-is by the caller).
|
||||
Dispatch needs a real file for format routing, so the body goes
|
||||
through a temp file.
|
||||
"""
|
||||
with tempfile.NamedTemporaryFile(suffix=ext) as tmp:
|
||||
tmp.write(body)
|
||||
tmp.flush()
|
||||
try:
|
||||
avif, _resp = dispatch(
|
||||
Path(tmp.name),
|
||||
quality=IMAGE_QUALITY,
|
||||
maxsize=maxsize,
|
||||
maxzoom=1,
|
||||
)
|
||||
except Exception:
|
||||
return None
|
||||
return avif
|
||||
|
||||
|
||||
def _svg_to_png(body: bytes, maxsize: int) -> bytes | None:
|
||||
"""Rasterize an SVG to PNG via pyvips, scaled so the long side is
|
||||
``maxsize`` — SVGs often carry no meaningful intrinsic resolution, so
|
||||
we rasterize at full image size rather than the tiny nominal one."""
|
||||
import pyvips
|
||||
|
||||
try:
|
||||
img = pyvips.Image.new_from_buffer(body, "")
|
||||
scale = (
|
||||
maxsize / max(img.width, img.height)
|
||||
if img.width and img.height
|
||||
else maxsize
|
||||
)
|
||||
if scale != 1:
|
||||
img = pyvips.Image.new_from_buffer(body, "", scale=scale)
|
||||
return img.write_to_buffer(".png")
|
||||
except pyvips.Error:
|
||||
return None
|
||||
|
||||
|
||||
def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes:
|
||||
"""Re-encode the AVIF derivative into a fallback format (WebP/JPEG)
|
||||
via pyvips. JPEG has no alpha, so it is flattened onto white;
|
||||
``strip`` keeps metadata (EXIF) out of the fallbacks."""
|
||||
import pyvips
|
||||
|
||||
img = pyvips.Image.new_from_buffer(avif, "")
|
||||
if suffix == ".jpg" and img.hasalpha():
|
||||
img = img.flatten(background=[255, 255, 255])
|
||||
return img.write_to_buffer(suffix, Q=quality, strip=True)
|
||||
|
||||
|
||||
def _image_derivatives(
|
||||
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE
|
||||
) -> dict[str, bytes] | None:
|
||||
"""The served variants of an uploaded image: ``avif`` (primary,
|
||||
thumbnailed to ``maxsize``) plus ``webp`` and ``jpg`` fallbacks
|
||||
re-encoded from it. SVGs are rasterized first (they are vector, so
|
||||
the raster replaces nothing — the .svg itself stays servable).
|
||||
Returns None for non-decodable content (stored as-is by the caller).
|
||||
"""
|
||||
if ext == ".svg":
|
||||
png = _svg_to_png(body, maxsize)
|
||||
if png is None:
|
||||
return None
|
||||
body, ext = png, ".png"
|
||||
avif = _to_avif(body, ext, maxsize)
|
||||
if avif is None:
|
||||
return None
|
||||
return {
|
||||
"avif": avif,
|
||||
"webp": _avif_to_format(avif, ".webp", IMAGE_WEBP_QUALITY),
|
||||
"jpg": _avif_to_format(avif, ".jpg", IMAGE_JPG_QUALITY),
|
||||
}
|
||||
|
||||
|
||||
def store_image(
|
||||
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE, *, derive: bool = True
|
||||
) -> str:
|
||||
"""Store an image body content-addressed and return its file name.
|
||||
|
||||
Decodable images get AVIF/WebP/JPEG derivatives thumbnailed to
|
||||
``maxsize``; the original is kept as ``<hash>.orig<ext>`` (SVG
|
||||
originals as ``<hash>.svg``, still servable) and the bare ``<hash>``
|
||||
name is returned (the server negotiates the format by Accept header).
|
||||
Anything else — undecodable content, or ``derive=False`` (GIFs, whose
|
||||
animation recompression would lose) — is stored as-is and returned with
|
||||
its extension. Blocking (pyvips/ffmpeg); call via ``asyncio.to_thread``
|
||||
from async code.
|
||||
"""
|
||||
digest = blake3.blake3(body).hexdigest()[:12]
|
||||
derivatives = _image_derivatives(body, ext, maxsize) if derive else None
|
||||
if derivatives is None: # store the body as-is
|
||||
file_store.put(digest + ext, body)
|
||||
return digest + ext
|
||||
file_store.put(f"{digest}.svg" if ext == ".svg" else f"{digest}.orig{ext}", body)
|
||||
for fmt, variant in derivatives.items():
|
||||
file_store.put(f"{digest}.{fmt}", variant)
|
||||
return digest
|
||||
|
||||
|
||||
@router.put("/_api/files/{name}")
|
||||
async def upload_file(name: str, request: Request) -> dict[str, str]:
|
||||
"""Store an upload (image, video...) in the content-addressed store.
|
||||
|
||||
The stored name is a blake3 hash prefix + the original extension,
|
||||
served immutable at "/_f/{name}"; returns {"path": "/_f/..."}.
|
||||
|
||||
Raster images and SVGs are recompressed (SVGs rasterized) into AVIF
|
||||
(primary) plus WebP and JPEG fallbacks: the original goes to
|
||||
``<hash>.orig<ext>`` (kept for reprocessing, never served — it may
|
||||
carry EXIF data; SVG originals stay servable as ``<hash>.svg`` since
|
||||
vector carries no EXIF) and pages link the bare ``/_f/<hash>``, the
|
||||
server picking the format from the request's Accept header. GIFs are
|
||||
stored as-is (animation would be lost), as is other non-decodable
|
||||
content.
|
||||
"""
|
||||
if "/" in name or name in {".", ".."}:
|
||||
raise HTTPException(400, "bad file name")
|
||||
body = await request.body()
|
||||
if not body:
|
||||
raise HTTPException(400, "empty file")
|
||||
ext = _ext(name)
|
||||
stored = await asyncio.to_thread(store_image, body, ext, derive=ext != ".gif")
|
||||
return {"path": f"/_f/{stored}"}
|
||||
|
||||
|
||||
@router.delete("/_api/files/{name}", status_code=204)
|
||||
async def delete_file(name: str) -> None:
|
||||
"""Remove a file from the content-addressed store (no refcounting:
|
||||
other pages referencing the same content will 404)."""
|
||||
if name not in file_store:
|
||||
raise HTTPException(404, "no such file")
|
||||
file_store.delete(name)
|
||||
|
||||
|
||||
@router.get("/favicon.ico", include_in_schema=False)
|
||||
async def favicon_ico() -> Response:
|
||||
"""The conventional /favicon.ico: redirect to the configured site icon.
|
||||
|
||||
Browsers request this path on their own (tabs, bookmarks, feeds and
|
||||
other non-HTML contexts) regardless of the <link rel="icon"> pages
|
||||
carry. Redirect to the icon's store URL, which negotiates the format
|
||||
and caches immutably; 404 when no custom icon is configured.
|
||||
"""
|
||||
if not data.favicon:
|
||||
raise HTTPException(404)
|
||||
return RedirectResponse(f"/_f/{data.favicon}")
|
||||
|
||||
|
||||
@router.put("/_api/settings/favicon")
|
||||
async def put_favicon(request: Request) -> dict[str, str]:
|
||||
"""Upload a favicon into the content-addressed store and activate it.
|
||||
|
||||
Raw image body (ico/png/svg...). Decodable images are thumbnailed to
|
||||
FAVICON_MAXSIZE (192px — browsers scale down from there themselves)
|
||||
and stored as AVIF/WebP/JPEG derivatives linked extension-less; SVG
|
||||
originals also stay servable under their ``.svg`` name. Undecodable
|
||||
bodies are stored as-is. Pages link it as <link rel="icon">. Returns
|
||||
{"path": "/_f/..."}.
|
||||
"""
|
||||
body = await request.body()
|
||||
if not body:
|
||||
raise HTTPException(400, "empty file")
|
||||
ext = _ext(request.headers.get("x-filename", "favicon.ico"))
|
||||
stored = await asyncio.to_thread(store_image, body, ext, FAVICON_MAXSIZE)
|
||||
with kanta.transaction("settings", user=request.headers.get("remote-user")):
|
||||
data.favicon = stored
|
||||
_invalidate_pages()
|
||||
return {"path": f"/_f/{stored}"}
|
||||
|
||||
|
||||
@router.delete("/_api/settings/favicon", status_code=204)
|
||||
async def delete_favicon(request: Request) -> None:
|
||||
"""Clear the custom favicon (/favicon.ico goes back to 404, pages drop
|
||||
the <link rel="icon">).
|
||||
|
||||
The blob stays in the content-addressed store; only the reference goes.
|
||||
"""
|
||||
with kanta.transaction("settings", user=request.headers.get("remote-user")):
|
||||
data.favicon = ""
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
async def _serve_user_file(path: Path | None, request: Request) -> Response:
|
||||
"""Serve a user-asset file resolved on disk, with mtime etag.
|
||||
|
||||
Read from disk on every request (etag by mtime+size): user assets are
|
||||
never built or content-hashed, so edits on disk show on the next page
|
||||
load, in prod as well as dev.
|
||||
"""
|
||||
if path is None:
|
||||
raise HTTPException(404)
|
||||
stat = path.stat()
|
||||
etag = f'"{stat.st_mtime_ns:x}-{stat.st_size:x}"'
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
|
||||
return Response(
|
||||
path.read_bytes(),
|
||||
media_type=mime,
|
||||
headers={"etag": etag, "cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/_themes/{name}/{filename}")
|
||||
async def theme_file(name: str, filename: str, request: Request) -> Response:
|
||||
"""Serve a theme/banner-design file, resolved across views.THEME_DIRS.
|
||||
|
||||
Stylesheets plus any extra assets the CSS references (like summer's
|
||||
grass.svg).
|
||||
"""
|
||||
return await _serve_user_file(views.theme_file(name, filename), request)
|
||||
|
||||
|
||||
@router.get("/_fonts/{name}/{filename}")
|
||||
async def user_font_file(name: str, filename: str, request: Request) -> Response:
|
||||
"""Serve a user font file, resolved across views.FONT_DIRS.
|
||||
|
||||
The folder's font.css (@font-face rules + --font-{name} stack variable)
|
||||
is linked on every page; the woff2 files it references come from here.
|
||||
"""
|
||||
return await _serve_user_file(views.font_file(name, filename), request)
|
||||
|
||||
|
||||
@router.get("/_f/{name}")
|
||||
async def stored_file(name: str, request: Request) -> Response:
|
||||
"""Serve a file from the content-addressed store (immutable: the name
|
||||
is its own hash, so cache forever). Bodies are served from the RAM
|
||||
cache, zstd-compressed when the client accepts it and compression
|
||||
actually shrank the file.
|
||||
|
||||
A bare ``/_f/{hash}`` (no extension, how pages link uploaded images)
|
||||
content-negotiates between the stored derivatives: a format is served
|
||||
only when the Accept header lists it explicitly — ``image/avif`` →
|
||||
AVIF, ``image/webp`` → WebP, anything else (including ``image/*`` and
|
||||
``*/*``) → JPEG. An explicit extension pins the format. ``.orig.``
|
||||
originals are internal (they may carry EXIF data) and never served."""
|
||||
if ".orig." in name:
|
||||
raise HTTPException(404)
|
||||
etag = name
|
||||
vary = ""
|
||||
entry = file_store.get(name)
|
||||
if entry is None and "." not in name:
|
||||
# Extension-less image link: negotiate avif/webp/jpg by Accept.
|
||||
vary = "accept"
|
||||
accept = request.headers.get("accept", "")
|
||||
if "image/avif" in accept:
|
||||
order = ("avif", "webp", "jpg")
|
||||
elif "image/webp" in accept:
|
||||
order = ("webp", "jpg", "avif")
|
||||
else:
|
||||
order = ("jpg", "webp", "avif")
|
||||
for ext in order:
|
||||
etag = f"{name}.{ext}"
|
||||
entry = file_store.get(etag)
|
||||
if entry is not None:
|
||||
break
|
||||
if entry is None:
|
||||
raise HTTPException(404)
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
body, compressed = entry
|
||||
headers = {"etag": etag, "cache-control": "public, max-age=31536000, immutable"}
|
||||
if compressed is not None and "zstd" in request.headers.get("accept-encoding", ""):
|
||||
headers["content-encoding"] = "zstd"
|
||||
vary = f"{vary}, accept-encoding".lstrip(", ")
|
||||
body = compressed
|
||||
if vary:
|
||||
headers["vary"] = vary
|
||||
mime = mimetypes.guess_type(etag)[0] or "application/octet-stream"
|
||||
return Response(body, media_type=mime, headers=headers)
|
||||
@@ -0,0 +1,282 @@
|
||||
"""Localization: language selection, translation storage and assembly.
|
||||
|
||||
See docs/localization.md and docs/migrate.md. Each article's primary
|
||||
language is ``Node.language``, inherited down the hierarchy (front page =
|
||||
site default, ORIGINAL_LANGUAGE as the final fallback). The database holds
|
||||
the original language as content-addressed chunks (``Data.chunks``); per
|
||||
target language there are machine-translated fragments (``Data.trans``)
|
||||
and user override patches (``Data.patches``), assembled into the served
|
||||
Markdown at render time, with per-node fallback to the original titles.
|
||||
"""
|
||||
|
||||
from collections.abc import Callable
|
||||
from difflib import SequenceMatcher
|
||||
|
||||
import msgspec
|
||||
|
||||
from pagerite.chunks import chunk_key, chunk_markdown, join_chunks
|
||||
from pagerite.data import Data, Node, Patch, resolve
|
||||
|
||||
#: Final fallback for a page's primary language when neither it nor any
|
||||
#: ancestor (up to the front page) sets one (Node.language, "" = inherit).
|
||||
ORIGINAL_LANGUAGE = "en"
|
||||
|
||||
#: Languages written right-to-left; pages served in one get dir="rtl" on
|
||||
#: <html> (views._layout).
|
||||
RTL_LANGUAGES = frozenset({"ar", "fa", "he", "ur"})
|
||||
|
||||
|
||||
def primary_lang(menu: dict[str, Node], path: str) -> str:
|
||||
"""The primary language of the article at ``path``: its own
|
||||
``language`` setting, else the nearest ancestor's (the front page
|
||||
last — it doubles as the site default), falling back to
|
||||
ORIGINAL_LANGUAGE. Missing tail segments (a page being created)
|
||||
resolve to the nearest existing ancestor."""
|
||||
p = path.strip("/")
|
||||
while True:
|
||||
chain = resolve(menu, p)
|
||||
if chain:
|
||||
for node in reversed(chain):
|
||||
if node.language:
|
||||
return node.language
|
||||
if not p:
|
||||
return ORIGINAL_LANGUAGE
|
||||
p = p.rpartition("/")[0]
|
||||
|
||||
|
||||
class Translation(msgspec.Struct, omit_defaults=True):
|
||||
"""Translated content for one page and language.
|
||||
|
||||
``markdown`` is the translated page source in the same format as the
|
||||
original (None = keep the original Markdown); ``titles`` maps node paths
|
||||
(top-level slug, then slash-joined) to translated navigation titles, so a
|
||||
partially translated tree still renders with per-node English fallback.
|
||||
"""
|
||||
|
||||
markdown: str | None = None
|
||||
titles: dict[str, str] = {}
|
||||
|
||||
|
||||
def base_tag(tag: str) -> str:
|
||||
"""The lowercase base subtag of a language tag (fi-FI -> fi)."""
|
||||
return tag.strip().lower().partition("-")[0]
|
||||
|
||||
|
||||
def parse_accept_language(header: str) -> list[str]:
|
||||
"""Accept-Language header as an ordered, deduped list of base subtags.
|
||||
|
||||
q-values are deliberately ignored: all known implementations send the
|
||||
header in order of preference. Region tags normalize to their base
|
||||
subtag (fi-FI -> fi); "*" and empties are dropped.
|
||||
"""
|
||||
langs = []
|
||||
for part in header.split(","):
|
||||
tag = base_tag(part.split(";", 1)[0])
|
||||
if tag and tag != "*" and tag not in langs:
|
||||
langs.append(tag)
|
||||
return langs
|
||||
|
||||
|
||||
def select_language(
|
||||
query_lang: str | None,
|
||||
accept_language: str | None,
|
||||
is_available: Callable[[str], bool],
|
||||
original: str = ORIGINAL_LANGUAGE,
|
||||
) -> str:
|
||||
"""The language to serve (see docs/localization.md).
|
||||
|
||||
1. ``?lang=`` wins when a translation exists for it (otherwise falls
|
||||
through to the header logic).
|
||||
2. The original language anywhere in the header list wins — an AI
|
||||
translation is strictly worse than the original for anyone who has
|
||||
English configured at all.
|
||||
3. Otherwise the first header language with an available translation.
|
||||
4. Fall back to the original.
|
||||
"""
|
||||
if query_lang:
|
||||
tag = base_tag(query_lang)
|
||||
if tag == original or (tag and is_available(tag)):
|
||||
return tag
|
||||
langs = parse_accept_language(accept_language or "")
|
||||
if original in langs:
|
||||
return original
|
||||
for lang in langs:
|
||||
if lang != original and is_available(lang):
|
||||
return lang
|
||||
return original
|
||||
|
||||
|
||||
def apply_patch(hybrid: str, patch: Patch) -> str:
|
||||
"""Apply one patch to the hybrid Markdown, best effort, each hunk
|
||||
independently: a hunk whose search text no longer exists is stale and
|
||||
silently skipped (docs/localization.md)."""
|
||||
for search, replace in patch.hunks:
|
||||
if search and search in hybrid:
|
||||
hybrid = hybrid.replace(search, replace, 1)
|
||||
return hybrid
|
||||
|
||||
|
||||
def make_patch(base: str, edited: str) -> Patch:
|
||||
"""The minimal diff of ``edited`` against the served ``base`` hybrid as
|
||||
(search, replace) hunks at block granularity (docs/localization.md).
|
||||
|
||||
Blocks are the chunk_markdown split, so hunks align with translation
|
||||
units and code fences never straddle a hunk boundary. Pure inserts
|
||||
anchor on the preceding block (an empty search would never match);
|
||||
inserts at the very top anchor on the first block. autojunk is off:
|
||||
the diff must be deterministic, and pages are small.
|
||||
"""
|
||||
a, b = chunk_markdown(base), chunk_markdown(edited)
|
||||
hunks: list[tuple[str, str]] = []
|
||||
for tag, i1, i2, j1, j2 in SequenceMatcher(
|
||||
None, a, b, autojunk=False
|
||||
).get_opcodes():
|
||||
if tag == "equal":
|
||||
continue
|
||||
search = "\n\n".join(a[i1:i2])
|
||||
replace = "\n\n".join(b[j1:j2])
|
||||
if tag == "insert":
|
||||
if i1:
|
||||
search = a[i1 - 1]
|
||||
replace = f"{a[i1 - 1]}\n\n{replace}"
|
||||
elif a:
|
||||
search = a[0]
|
||||
replace = f"{replace}\n\n{a[0]}"
|
||||
# else: base is empty — the hunk is inert (empty search is
|
||||
# skipped by apply_patch); saving a translation of an empty
|
||||
# page records nothing applicable.
|
||||
hunks.append((search, replace))
|
||||
return Patch(hunks=hunks)
|
||||
|
||||
|
||||
def hybrid_markdown(data: Data, node: Node, path: str, lang: str) -> str:
|
||||
"""The served Markdown for ``lang``: per chunk the translation from
|
||||
``Data.trans``, unless missing or marked no-translate (fallback to the
|
||||
original chunk), then the language's user patches applied in order.
|
||||
|
||||
Not gated on ``node.langs`` (get_translation is the gated view): the
|
||||
editor save path diffs against this even for a language's first patch.
|
||||
"""
|
||||
hybrid = join_chunks(
|
||||
[
|
||||
data.chunks.get(h, "")
|
||||
if h in node.no_trans
|
||||
else data.trans.get(h, {}).get(lang) or data.chunks.get(h, "")
|
||||
for h in node.chunks or []
|
||||
]
|
||||
)
|
||||
for patch in data.patches.get(f"{path}:{lang}", []):
|
||||
hybrid = apply_patch(hybrid, patch)
|
||||
return hybrid
|
||||
|
||||
|
||||
def add_patch(
|
||||
data: Data, node: Node, path: str, lang: str, edited: str, base: str | None = None
|
||||
) -> bool:
|
||||
"""Record a translated-view edit as a user Patch: the minimal diff of
|
||||
``edited`` against ``base`` (default: the currently served hybrid),
|
||||
appended to the language's patch list. Patches alone make the
|
||||
translated version exist, so ``node.langs`` is set. Returns True when
|
||||
a patch was stored. Pure data ops — the caller wraps in a transaction
|
||||
and invalidates."""
|
||||
patch = make_patch(
|
||||
base if base is not None else hybrid_markdown(data, node, path, lang), edited
|
||||
)
|
||||
if not patch.hunks:
|
||||
return False
|
||||
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
|
||||
node.langs[lang] = True
|
||||
return True
|
||||
|
||||
|
||||
def set_title_translation(data: Data, node: Node, lang: str, title: str) -> bool:
|
||||
"""Record (or drop) a per-language title override: a fragment in
|
||||
``Data.trans`` keyed by the ORIGINAL title's chunk hash — the same
|
||||
storage machine title translations use, overriding them. Sending the
|
||||
original's text drops the override. Returns True when anything changed.
|
||||
Pure data ops — the caller wraps in a transaction and invalidates."""
|
||||
key = chunk_key(node.title)
|
||||
current = data.trans.get(key, {}).get(lang)
|
||||
if title == node.title:
|
||||
if current is None:
|
||||
return False
|
||||
del data.trans[key][lang]
|
||||
return True
|
||||
if current == title:
|
||||
return False
|
||||
data.trans.setdefault(key, {})[lang] = title
|
||||
node.langs[lang] = True
|
||||
return True
|
||||
|
||||
|
||||
def clear_translations(data: Data) -> None:
|
||||
"""Drop all machine translations (``Data.trans``) and rebuild the
|
||||
availability index (``node.langs``) from the surviving user patches —
|
||||
patches alone make a language exist on a page. Pure data ops — the
|
||||
caller wraps in a transaction and invalidates."""
|
||||
data.trans.clear()
|
||||
patch_langs: dict[str, set[str]] = {}
|
||||
for key in data.patches:
|
||||
path, _, lang = key.rpartition(":")
|
||||
patch_langs.setdefault(path, set()).add(lang)
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str) -> None:
|
||||
for slug, node in nodes.items():
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
node.langs = {lang: True for lang in patch_langs.get(path, ())}
|
||||
walk(node.children, path)
|
||||
|
||||
walk(data.menu, "")
|
||||
|
||||
|
||||
def title_map(data: Data, lang: str) -> dict[str, str]:
|
||||
"""path -> translated title for every node that has one.
|
||||
|
||||
Titles are chunks too (docs/migrate.md): keyed by the hash of the
|
||||
title text, so editing a title invalidates its translations. Nodes
|
||||
without an entry fall back to their original title in views — as do
|
||||
nodes whose primary language IS ``lang`` (their original title already
|
||||
is in that language).
|
||||
"""
|
||||
titles = {}
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
|
||||
for slug, node in nodes.items():
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
node_lang = node.language or inherited
|
||||
if node.title and node_lang != lang:
|
||||
t = data.trans.get(chunk_key(node.title), {}).get(lang)
|
||||
if t:
|
||||
titles[path] = t
|
||||
walk(node.children, path, node_lang)
|
||||
|
||||
walk(data.menu, "", ORIGINAL_LANGUAGE)
|
||||
return titles
|
||||
|
||||
|
||||
def subtree_languages(node: Node) -> set[str]:
|
||||
"""Languages available anywhere in the node's subtree (the union of the
|
||||
``langs`` indexes). Category placeholder pages select their language
|
||||
from this: they have no chunks of their own, but their title,
|
||||
navigation and card text localize wherever a translation exists."""
|
||||
langs = set(node.langs)
|
||||
for child in node.children.values():
|
||||
langs |= subtree_languages(child)
|
||||
return langs
|
||||
|
||||
|
||||
def get_translation(data: Data, path: str, lang: str) -> Translation | None:
|
||||
"""The translation of the page at ``path`` for ``lang``, or None.
|
||||
|
||||
None when the page does not exist or is not available in ``lang``:
|
||||
``node.langs`` is the availability index (a stale key is benign — the
|
||||
"translation" then just renders as the original).
|
||||
"""
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None or lang not in node.langs:
|
||||
return None
|
||||
return Translation(
|
||||
markdown=hybrid_markdown(data, node, path, lang),
|
||||
titles=title_map(data, lang),
|
||||
)
|
||||
+121
-43
@@ -156,14 +156,27 @@ def _image_rule(
|
||||
page = env.get("page_path", "")
|
||||
token.attrs["src"] = f"/{page}/{src}" if page else f"/{src}"
|
||||
token.attrs["alt"] = self.renderInlineAsText(token.children, options, env)
|
||||
img = self.renderToken(tokens, idx, options, env)
|
||||
if len(tokens) == 1:
|
||||
# The only inline content of its paragraph: render as a block
|
||||
# figure, captioned when titled. (The <p> wrapper is dropped by
|
||||
# _unwrap_lone_figures below.)
|
||||
# _unwrap_lone_figures below.) {.margin} positions the whole
|
||||
# figure, so it moves from the img onto the figure wrapper — left
|
||||
# on the img, the margin-breakout CSS would pull the image out of
|
||||
# the figure (and mostly off-screen), leaving the caption behind.
|
||||
classes = (token.attrs.get("class") or "").split()
|
||||
figure_class = ""
|
||||
if "margin" in classes:
|
||||
classes.remove("margin")
|
||||
if classes:
|
||||
token.attrs["class"] = " ".join(classes)
|
||||
else:
|
||||
del token.attrs["class"]
|
||||
figure_class = ' class="margin"'
|
||||
img = self.renderToken(tokens, idx, options, env)
|
||||
title = token.attrs.get("title")
|
||||
caption = f"<figcaption>{escapeHtml(title)}</figcaption>" if title else ""
|
||||
return f"<figure>{img}{caption}</figure>"
|
||||
return f"<figure{figure_class}>{img}{caption}</figure>"
|
||||
img = self.renderToken(tokens, idx, options, env)
|
||||
# Inline with other content: a plain inline image.
|
||||
return img
|
||||
|
||||
@@ -389,7 +402,9 @@ def _heading_ids(state) -> None:
|
||||
its self-link is ``href=""`` (back to the top of the page). An
|
||||
author-set `{#id}` always wins; auto ids slugify the heading text
|
||||
(python-slugify, mirroring the editor's slugify.js) and dedupe with
|
||||
-2/-3 suffixes per render. Headings that already contain a link are
|
||||
-2/-3 suffixes per render — unless env["anchor_ids"] presets them, as
|
||||
render(anchors_from=...) does for translated pages so section URLs
|
||||
stay in the original language. Headings that already contain a link are
|
||||
``data-line`` records the heading's markdown source line (0-based, after
|
||||
undoing the render(title=...) injection offset via ``env``) — the page
|
||||
editor uses it for section pens and piecewise-linear scroll sync.
|
||||
@@ -422,20 +437,33 @@ def _heading_ids(state) -> None:
|
||||
heads = [
|
||||
(i, token)
|
||||
for i, token in enumerate(tokens)
|
||||
if token.type == "heading_open" and token.tag in ("h1", "h2") and token.level == 0 and i != first_h1
|
||||
if token.type == "heading_open"
|
||||
and token.tag in ("h1", "h2")
|
||||
and token.level == 0
|
||||
and i != first_h1
|
||||
]
|
||||
if len(heads) < ANCHOR_MIN_HEADINGS:
|
||||
return
|
||||
seen: set[str] = set()
|
||||
for i, token in heads:
|
||||
preset = state.env.get("anchor_ids")
|
||||
for k, (i, token) in enumerate(heads):
|
||||
inline = tokens[i + 1]
|
||||
hid = token.attrGet("id")
|
||||
if not isinstance(hid, str) or not hid:
|
||||
# Slug the visible text, not the raw markdown (`## [a](url)`).
|
||||
text = "".join(
|
||||
c.content for c in inline.children if c.type in ("text", "code_inline")
|
||||
)
|
||||
base = slugify(text) or "section"
|
||||
if preset is not None and k < len(preset):
|
||||
# Translated render: the original language's slug, matched
|
||||
# by heading position (a translation never adds, removes or
|
||||
# reorders headings; a patched one that does falls back to
|
||||
# slugging its own text past the end of the list).
|
||||
base = preset[k]
|
||||
else:
|
||||
# Slug the visible text, not the raw markdown (`## [a](url)`).
|
||||
text = "".join(
|
||||
c.content
|
||||
for c in inline.children
|
||||
if c.type in ("text", "code_inline")
|
||||
)
|
||||
base = slugify(text) or "section"
|
||||
hid, n = base, 2
|
||||
while hid in seen:
|
||||
hid = f"{base}-{n}"
|
||||
@@ -447,39 +475,83 @@ def _heading_ids(state) -> None:
|
||||
wrap(i, token, f"#{hid}")
|
||||
|
||||
|
||||
md = (
|
||||
MarkdownIt(
|
||||
"default",
|
||||
{
|
||||
"html": True,
|
||||
"highlight": _highlight,
|
||||
"typographer": True,
|
||||
"breaks": True,
|
||||
},
|
||||
def anchor_ids(text: str, title: str | None = None) -> list[str]:
|
||||
"""The section anchor ids of text, in heading order.
|
||||
|
||||
render(anchors_from=...) feeds these to _heading_ids via
|
||||
env["anchor_ids"], pinning a translated render's anchors to the
|
||||
original language's slugs. The selection mirrors _heading_ids exactly
|
||||
(the same md instance assigns the ids during this parse, author-set
|
||||
{#id} included as-is); the in-body title h1 is excluded.
|
||||
"""
|
||||
if title and not has_h1(text):
|
||||
text = f"# {title}\n\n{text}"
|
||||
tokens = md.parse(text, {"page_path": ""})
|
||||
first_h1 = next(
|
||||
(
|
||||
i
|
||||
for i, t in enumerate(tokens)
|
||||
if t.type == "heading_open" and t.tag == "h1" and t.level == 0
|
||||
),
|
||||
None,
|
||||
)
|
||||
.use(attrs_plugin)
|
||||
.use(admon_plugin)
|
||||
.use(container_plugin, "block", validate=_container_validate)
|
||||
.use(footnote_plugin)
|
||||
.use(deflist_plugin)
|
||||
# label_after: the item text is wrapped in <label for> after the
|
||||
# checkbox, so clicking the text toggles it.
|
||||
.use(tasklists_plugin, enabled=True, label=True, label_after=True)
|
||||
.use(gfm_autolink_plugin)
|
||||
.use(sub_plugin)
|
||||
.use(superscript_plugin)
|
||||
)
|
||||
md.add_render_rule("image", _image_rule)
|
||||
md.add_render_rule("fence", _fence_rule)
|
||||
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
|
||||
md.options["alerts"] = True
|
||||
# Block attrs must be stripped before the typographer curlifies their quotes.
|
||||
md.core.ruler.before("replacements", "block_attrs", _block_attrs)
|
||||
md.core.ruler.push("container_attrs", _container_attrs)
|
||||
md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
|
||||
md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
|
||||
md.core.ruler.push("shorten_autolinks", _shorten_autolinks)
|
||||
md.core.ruler.push("heading_ids", _heading_ids)
|
||||
return [
|
||||
t.attrGet("id")
|
||||
for i, t in enumerate(tokens)
|
||||
if t.type == "heading_open"
|
||||
and t.tag in ("h1", "h2")
|
||||
and t.level == 0
|
||||
and i != first_h1
|
||||
]
|
||||
|
||||
|
||||
def make_md(*, verbatim: bool = False) -> MarkdownIt:
|
||||
"""A fully configured parser. The module-level ``md`` (below) is the
|
||||
render instance; ``verbatim=True`` builds the segmentation instance for
|
||||
segments.py, where token text must stay byte-identical to the source so
|
||||
prose spans can be spliced back by offset: no typographer (quotes and
|
||||
dashes stay straight), no tasklist label wrapping (the item text stays
|
||||
a plain text token), and soft line breaks (wrapped prose merges into
|
||||
one segment instead of splitting at hardbreaks)."""
|
||||
parser = (
|
||||
MarkdownIt(
|
||||
"default",
|
||||
{
|
||||
"html": True,
|
||||
"highlight": _highlight,
|
||||
"typographer": not verbatim,
|
||||
"breaks": not verbatim,
|
||||
},
|
||||
)
|
||||
.use(attrs_plugin)
|
||||
.use(admon_plugin)
|
||||
.use(container_plugin, "block", validate=_container_validate)
|
||||
.use(footnote_plugin)
|
||||
.use(deflist_plugin)
|
||||
# label wrapping (render) puts the item text inside the checkbox
|
||||
# <label> html_inline; without it the text stays a plain token.
|
||||
.use(
|
||||
tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim
|
||||
)
|
||||
.use(gfm_autolink_plugin)
|
||||
.use(sub_plugin)
|
||||
.use(superscript_plugin)
|
||||
)
|
||||
parser.add_render_rule("image", _image_rule)
|
||||
parser.add_render_rule("fence", _fence_rule)
|
||||
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
|
||||
parser.options["alerts"] = True
|
||||
# Block attrs must be stripped before the typographer curlifies their quotes.
|
||||
parser.core.ruler.before("replacements", "block_attrs", _block_attrs)
|
||||
parser.core.ruler.push("container_attrs", _container_attrs)
|
||||
parser.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
|
||||
parser.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
|
||||
parser.core.ruler.push("shorten_autolinks", _shorten_autolinks)
|
||||
parser.core.ruler.push("heading_ids", _heading_ids)
|
||||
return parser
|
||||
|
||||
|
||||
md = make_md()
|
||||
|
||||
|
||||
# Text-length thresholds (visible characters, code blocks excluded) for the
|
||||
@@ -588,12 +660,16 @@ def render(
|
||||
created: datetime | None = None,
|
||||
modified: datetime | None = None,
|
||||
title: str | None = None,
|
||||
anchors_from: tuple[str, str] | None = None,
|
||||
) -> Rendered:
|
||||
"""Render Markdown text to the article body's HTML and layout flags.
|
||||
|
||||
``title`` injects a ``# {title}`` line at the top when the markdown has
|
||||
no h1 of its own, so the implicit page title goes through the exact
|
||||
same pipeline as an explicit one (first-h1 anchor treatment included).
|
||||
``anchors_from`` is the (markdown, title) of the ORIGINAL language when
|
||||
rendering a translation: section anchors are pinned to its slugs so
|
||||
localized pages keep the original #hash URLs.
|
||||
|
||||
The top-level blocks are grouped into column segments: boundary blocks
|
||||
(h1/h2 headings, .wide — see _is_boundary) are rendered bare, the runs
|
||||
@@ -611,6 +687,8 @@ def render(
|
||||
right after the article's h1.
|
||||
"""
|
||||
env = {"page_path": page_path, "line_offset": 0}
|
||||
if anchors_from is not None:
|
||||
env["anchor_ids"] = anchor_ids(*anchors_from)
|
||||
if title and not has_h1(text):
|
||||
text = f"# {title}\n\n{text}"
|
||||
# The injected title shifts source lines by two; _heading_ids
|
||||
|
||||
+57
-12
@@ -6,15 +6,16 @@ omitted) before it is decoded into ``Data`` structs, and runs exactly once
|
||||
per database based on its recorded version.
|
||||
|
||||
All storage/schema upgrades live here — including on-disk file work, which
|
||||
runs through app.py's file store (imported lazily: app.py owns the store
|
||||
and passes this module to Kanta; at migration time, during lifespan
|
||||
``kanta.open()``, the app module is fully loaded).
|
||||
runs through files.py's file store (imported lazily: files.py owns the store
|
||||
and state.py passes this module to Kanta; at migration time, during lifespan
|
||||
``kanta.open()``, both modules are fully loaded).
|
||||
"""
|
||||
|
||||
import base64
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
from pagerite.chunks import chunk_key, chunk_markdown
|
||||
from pagerite.data import prettify
|
||||
|
||||
|
||||
@@ -24,7 +25,7 @@ def _append_order(nodes: dict) -> float:
|
||||
|
||||
|
||||
def _ensure(menu: dict, path: str) -> dict:
|
||||
"""Raw-dict equivalent of app._ensure: the node dict at ``path``,
|
||||
"""Raw-dict equivalent of state._ensure: the node dict at ``path``,
|
||||
creating it and any missing ancestors (content-less category labels)
|
||||
appended at the end of their level."""
|
||||
nodes = menu
|
||||
@@ -43,7 +44,7 @@ def migrate_v1(d: dict) -> None:
|
||||
and rebuild the legacy flat page store (``pages``) as the menu tree."""
|
||||
files = d.pop("files", None)
|
||||
if files:
|
||||
from pagerite.app import file_store
|
||||
from pagerite.files import file_store
|
||||
|
||||
for name, body in files.items():
|
||||
if isinstance(body, str): # JSON-level bytes are base64 strings
|
||||
@@ -73,9 +74,16 @@ def _backfill_derivatives() -> None:
|
||||
AVIF, and SVGs no raster variants at all). WebP/JPEG are re-encoded
|
||||
from an existing AVIF when available, everything else from the
|
||||
original (SVGs rasterized first)."""
|
||||
from pagerite import app
|
||||
from pagerite.files import (
|
||||
IMAGE_MAXSIZE,
|
||||
IMAGE_WEBP_QUALITY,
|
||||
IMAGE_JPG_QUALITY,
|
||||
_avif_to_format,
|
||||
_svg_to_png,
|
||||
_to_avif,
|
||||
file_store,
|
||||
)
|
||||
|
||||
file_store = app.file_store
|
||||
try:
|
||||
paths = [f for f in file_store.path.iterdir() if f.is_file()]
|
||||
except FileNotFoundError:
|
||||
@@ -95,22 +103,22 @@ def _backfill_derivatives() -> None:
|
||||
ext = source.suffix
|
||||
body = source.read_bytes()
|
||||
if ext == ".svg":
|
||||
png = app._svg_to_png(body, app.IMAGE_MAXSIZE)
|
||||
png = _svg_to_png(body, IMAGE_MAXSIZE)
|
||||
if png is None:
|
||||
continue
|
||||
body, ext = png, ".png"
|
||||
converted = app._to_avif(body, ext)
|
||||
converted = _to_avif(body, ext)
|
||||
if converted is None:
|
||||
continue
|
||||
file_store.put(f"{digest}.avif", converted)
|
||||
avif = file_store.get(f"{digest}.avif")
|
||||
for fmt, quality in (
|
||||
("webp", app.IMAGE_WEBP_QUALITY),
|
||||
("jpg", app.IMAGE_JPG_QUALITY),
|
||||
("webp", IMAGE_WEBP_QUALITY),
|
||||
("jpg", IMAGE_JPG_QUALITY),
|
||||
):
|
||||
if f"{digest}.{fmt}" not in names:
|
||||
file_store.put(
|
||||
f"{digest}.{fmt}", app._avif_to_format(avif[0], f".{fmt}", quality)
|
||||
f"{digest}.{fmt}", _avif_to_format(avif[0], f".{fmt}", quality)
|
||||
)
|
||||
|
||||
|
||||
@@ -131,3 +139,40 @@ def migrate_v2(d: dict) -> None:
|
||||
walk(d.get("menu") or {})
|
||||
d.pop("version", None)
|
||||
_backfill_derivatives()
|
||||
|
||||
|
||||
def migrate_v3(d: dict) -> None:
|
||||
"""Content-addressed chunk storage (docs/migrate.md): split every
|
||||
node's string ``content`` into block chunks stored once per content
|
||||
hash in the new ``chunks`` store; the node keeps the ordered hash
|
||||
list as ``chunks`` (an absent content stays absent, i.e. None = a
|
||||
pure category label; "" chunks to an empty list = an empty page).
|
||||
|
||||
Chunk keys are 9-byte blake3 digests; at this raw JSON level they are
|
||||
base64 strings (decoding into the structs restores ``bytes`` keys).
|
||||
``trans``/``patches`` start empty; the translator job fills them and
|
||||
maintains the ``langs`` index as translations land. ``language``,
|
||||
``no_trans`` and ``langs`` need nothing — struct defaults cover them.
|
||||
"""
|
||||
store = d.setdefault("chunks", {})
|
||||
d.setdefault("trans", {})
|
||||
patches = d.setdefault("patches", {})
|
||||
|
||||
def walk(nodes: dict) -> None:
|
||||
for node in nodes.values():
|
||||
content = node.pop("content", None)
|
||||
if isinstance(content, str):
|
||||
hashes = []
|
||||
for chunk in chunk_markdown(content):
|
||||
key = base64.b64encode(chunk_key(chunk)).decode()
|
||||
store.setdefault(key, chunk)
|
||||
hashes.append(key)
|
||||
node["chunks"] = hashes
|
||||
walk(node.get("children") or {})
|
||||
|
||||
walk(d.get("menu") or {})
|
||||
# Article paths never carry a leading slash in keys (docs/migrate.md).
|
||||
# The only path-keyed store starts empty here, so this is defensive
|
||||
# for databases that went through a downgrade/upgrade cycle.
|
||||
for key in [k for k in patches if k.startswith("/")]:
|
||||
patches[key.lstrip("/")] = patches.pop(key)
|
||||
|
||||
@@ -0,0 +1,229 @@
|
||||
"""Public content pages: front page, sitemap, robots, and the catch-all.
|
||||
|
||||
``GET /{path:path}`` resolves a slug path against the menu tree and renders
|
||||
the page (or a category placeholder, or 404); it must be registered AFTER
|
||||
the fastapi-vue asset routes so built frontend files win over content slugs
|
||||
(see app.py). Every served document is recorded raw in analytics (one
|
||||
access-log line with its true HTTP status; classification happens at
|
||||
display time — see pagerite/analytics.py).
|
||||
"""
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from email.utils import format_datetime
|
||||
from xml.sax.saxutils import escape as xml_escape
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Request
|
||||
from fastapi.responses import RedirectResponse, Response
|
||||
|
||||
from pagerite import i18n, state
|
||||
from pagerite.data import Node, resolve, sorted_nodes
|
||||
from pagerite.state import (
|
||||
SITE_URL,
|
||||
_html_response,
|
||||
_is_reserved,
|
||||
data,
|
||||
)
|
||||
from pagerite.tracking import _record_get
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
def _http_date(dt: datetime) -> str:
|
||||
"""RFC 7231 date for the Last-Modified header."""
|
||||
return format_datetime(dt.astimezone(UTC), usegmt=True)
|
||||
|
||||
|
||||
def _is_trackable_path(path: str) -> bool:
|
||||
"""Content URLs only: skip auth endpoints and reserved/machinery paths."""
|
||||
if not path:
|
||||
return True
|
||||
if path == "auth" or path.startswith("auth/"):
|
||||
return False
|
||||
return not _is_reserved(path)
|
||||
|
||||
|
||||
@router.get("/")
|
||||
async def front_page(request: Request) -> Response:
|
||||
"""Render the front page (slug path "")."""
|
||||
return await show_page(request, "")
|
||||
|
||||
|
||||
@router.get("/sitemap.xml")
|
||||
async def sitemap(request: Request) -> Response:
|
||||
"""Dynamically generate a sitemap of all published article pages."""
|
||||
base = SITE_URL or str(request.base_url).rstrip("/")
|
||||
entries: list[tuple[str, datetime, int]] = []
|
||||
|
||||
def walk(
|
||||
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
|
||||
) -> None:
|
||||
first_content_slug = next(
|
||||
(
|
||||
slug
|
||||
for slug, node in sorted_nodes(nodes)
|
||||
if node.published and node.chunks is not None
|
||||
),
|
||||
None,
|
||||
)
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
depth = path.count("/") if path else 0
|
||||
if (
|
||||
not parent_has_content
|
||||
and slug == first_content_slug
|
||||
and node.published
|
||||
and node.chunks is not None
|
||||
and depth > 0
|
||||
):
|
||||
depth -= 1
|
||||
if node.published and node.chunks is not None:
|
||||
entries.append((path, node.modified, depth))
|
||||
if node.children:
|
||||
walk(node.children, path, node.chunks is not None)
|
||||
|
||||
walk(data.menu, "")
|
||||
|
||||
def priority(depth: int) -> float:
|
||||
return max(0.1, 1.0 - depth * 0.2)
|
||||
|
||||
lines = [
|
||||
'<?xml version="1.0" encoding="UTF-8"?>',
|
||||
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
|
||||
]
|
||||
for path, modified, depth in entries:
|
||||
loc = xml_escape(f"{base}/{path}" if path else base)
|
||||
lastmod = (
|
||||
modified.astimezone(UTC)
|
||||
.replace(microsecond=0)
|
||||
.isoformat()
|
||||
.replace("+00:00", "Z")
|
||||
)
|
||||
lines.append(
|
||||
f" <url>"
|
||||
f"<loc>{loc}</loc>"
|
||||
f"<lastmod>{lastmod}</lastmod>"
|
||||
f"<priority>{priority(depth):.1f}</priority>"
|
||||
f"</url>"
|
||||
)
|
||||
lines.append("</urlset>")
|
||||
|
||||
return Response(
|
||||
"\n".join(lines),
|
||||
media_type="application/xml",
|
||||
headers={"cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/robots.txt")
|
||||
async def robots_txt(request: Request) -> Response:
|
||||
"""Allow content crawling, keep the SSO login (/auth/) and the
|
||||
admin-gated API (/_api) out of search results, and point crawlers at
|
||||
the sitemap."""
|
||||
base = SITE_URL or str(request.base_url).rstrip("/")
|
||||
body = f"User-agent: *\nAllow: /\nDisallow: /auth/\nDisallow: /_api\nSitemap: {base}/sitemap.xml\n"
|
||||
return Response(
|
||||
body,
|
||||
media_type="text/plain",
|
||||
headers={"cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{path:path}", response_model=None)
|
||||
async def show_page(request: Request, path: str) -> Response:
|
||||
"""Render the content page at a slug path, or 404.
|
||||
|
||||
A node without content is a category label: its URL renders a
|
||||
placeholder page (nav links point straight at its first child).
|
||||
"""
|
||||
path = path.strip("/")
|
||||
accept_language = request.headers.get("accept-language", "")
|
||||
if path and _is_reserved(path):
|
||||
# Invalid slug shape: not a content URL, let FastAPI return its
|
||||
# built-in 404 instead of rendering an editable article page.
|
||||
# Recorded like any other GET: telltale scanner paths (dotpaths
|
||||
# like /.env, *.php) classify the IP as abuse at display time.
|
||||
_record_get(request, status=404)
|
||||
raise HTTPException(404)
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is not None and node.published and node.chunks is not None:
|
||||
# Language selection (docs/localization.md): ?lang= wins when a
|
||||
# translation exists, else header logic. Analytics keep the raw
|
||||
# Accept-Language header regardless of the selection.
|
||||
query_lang = request.query_params.get("lang")
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda tag: tag in node.langs,
|
||||
original=i18n.primary_lang(data.menu, path),
|
||||
)
|
||||
# A ?lang= override is replicated onto the page's navigation links
|
||||
# (link_lang), so clicks and prefetches stay in the chosen language.
|
||||
# Query and header-selected renders of the same language differ in
|
||||
# their links, so link_lang is part of the ETag and body cache key.
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
# no-cache forbids serving a stored page without revalidation
|
||||
# (browsers would otherwise cache heuristically and serve stale
|
||||
# pages, e.g. after a theme change). In-session speed instead comes
|
||||
# from pagerite.js's in-memory page cache (preload everything, never
|
||||
# fetch on navigation); the ETag just makes those one-time preload
|
||||
# fetches and any revalidation cheap.
|
||||
etag = f'"{path}@{node.modified.timestamp()}g{state._render_gen}l{lang}q{link_lang}"'
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
if _is_trackable_path(path):
|
||||
_record_get(request)
|
||||
return _html_response(
|
||||
request,
|
||||
"page",
|
||||
path,
|
||||
headers={
|
||||
"etag": etag,
|
||||
"last-modified": _http_date(node.modified),
|
||||
"cache-control": "no-cache",
|
||||
},
|
||||
lang=lang,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if node is not None and node.published and node.chunks is None:
|
||||
# Category label without a landing page: placeholder with the pen
|
||||
# to create it (404 — no page here, but the node is real).
|
||||
# Language selection as on content pages, but over the whole
|
||||
# subtree's availability: the category has no chunks of its own —
|
||||
# its heading, the navigation and the cards' text localize from
|
||||
# the title map and the target articles' translations.
|
||||
query_lang = request.query_params.get("lang")
|
||||
subtree_langs = i18n.subtree_languages(node)
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda tag: tag in subtree_langs,
|
||||
original=i18n.primary_lang(data.menu, path),
|
||||
)
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
if _is_trackable_path(path):
|
||||
_record_get(request, status=404)
|
||||
return _html_response(
|
||||
request,
|
||||
"category",
|
||||
path,
|
||||
404,
|
||||
headers={
|
||||
"last-modified": _http_date(node.modified),
|
||||
"cache-control": "no-cache",
|
||||
},
|
||||
lang=lang,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if node is None and not path:
|
||||
# No front page (no top-level node with slug ""): "/" opens the
|
||||
# first item of the navigation instead.
|
||||
for slug, item in sorted_nodes(data.menu):
|
||||
if item.published:
|
||||
return RedirectResponse(f"/{slug}")
|
||||
if _is_trackable_path(path):
|
||||
_record_get(request, status=404)
|
||||
return _html_response(request, "not-found", path, 404)
|
||||
+3
-2
@@ -21,6 +21,7 @@ ASSETS = Path(__file__).with_name("seed-assets")
|
||||
def _asset(name: str) -> bytes:
|
||||
return (ASSETS / name).read_bytes()
|
||||
|
||||
|
||||
WELCOME = """\
|
||||
Welcome to your new **Pagerite** site. Everything you see is a page written in Markdown, served from a pretty URL, and editable right here in the browser.
|
||||
|
||||
@@ -114,7 +115,7 @@ Headings from `##` down organize the article. On pages with at least three of th
|
||||
> and a blank `>` line starts a new paragraph.
|
||||
|
||||
> [!NOTE]
|
||||
> GitHub-style alerts — NOTE, TIP, IMPORTANT, WARNING, CAUTION —
|
||||
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
|
||||
> render as callout boxes.
|
||||
```
|
||||
|
||||
@@ -122,7 +123,7 @@ Headings from `##` down organize the article. On pages with at least three of th
|
||||
> and a blank `>` line starts a new paragraph.
|
||||
|
||||
> [!NOTE]
|
||||
> GitHub-style alerts — NOTE, TIP, IMPORTANT, WARNING, CAUTION —
|
||||
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
|
||||
> render as callout boxes.
|
||||
|
||||
## Code
|
||||
|
||||
@@ -0,0 +1,700 @@
|
||||
"""Segmented translation round trip: prose out, translations back in.
|
||||
|
||||
A translator model mangles anything that is not plain prose — sentinels get
|
||||
renumbered, ``![`` becomes sentence punctuation, stray ``<br>`` tags appear.
|
||||
So the model is never shown any of it: a fragment (a Markdown chunk or a
|
||||
node title) is parsed with the project's own markdown-it setup
|
||||
(``markdown.make_md(verbatim=True)`` — extensions included, so container,
|
||||
attrs, footnote and tasklist syntax never leaks into text tokens) and split
|
||||
into **prose segments**: the merged text runs, plus image alt texts and
|
||||
link/image titles. Only those cross the wire, as a plain list of strings
|
||||
(Job.texts / Result.texts in translate.py) — accompanied, per segment, by
|
||||
a CONTEXT (Job.contexts): a segment carved out of a larger block (a link
|
||||
text, a partial run) carries the block's plain text, so the model sees the
|
||||
sentence it lives in; whole-block segments are self-contextualizing and
|
||||
carry "". Title fragments carry the article's opening instead (assigned by
|
||||
the dispatcher from TransItem.context).
|
||||
|
||||
Reassembly is server-side offset splicing, not text the model produced:
|
||||
each segment's source span was located at dispatch (``split``), and
|
||||
``join`` swaps in the translations. Markup therefore cannot break — it
|
||||
never left the server. A returned segment must still be pure prose itself
|
||||
(the model could inject markup INTO a segment); anything else — count
|
||||
mismatch, empty segment, markup tokens, a line that would start a new
|
||||
block (a ``` or ::: fence would eat the rest of the block it lands in) —
|
||||
rejects the whole result and the
|
||||
fragment stays pending. Punctuation that is prose on the wire but syntax
|
||||
in the splice context (quotes in a title attribute, brackets in an alt
|
||||
text, "|" in a table row) is not worth a rejection either: it is swapped
|
||||
for Unicode look-alikes (``_NEUTRAL``) before splicing.
|
||||
|
||||
A block of plain text, prose links and paired text formatting
|
||||
(strong/em/s) crosses as ONE segment — link texts and formatted text
|
||||
inline, in sentence context, with the Markdown stripped (the model
|
||||
mangles it: sentinels get renumbered, ``**`` gets dropped or moved) —
|
||||
because a label translated apart from its sentence comes back
|
||||
grammatically incompatible with it (case government, particles, word
|
||||
order). ``join`` re-inserts the link/formatting markdown into the
|
||||
translated block at fuzzily matched positions (``_place_marks``): no
|
||||
markers on the wire, the boundaries are found by aligning the mark's
|
||||
source words to the translation's words by form similarity (``_find_mark``
|
||||
— inflection, dropped articles and reordering tolerated), with the
|
||||
source/translation weight ratio as fallback (the CJK path, where
|
||||
cross-script form similarity is nil). Placement is approximate: better a
|
||||
coherent sentence with a slightly shifted link than separately translated
|
||||
snippets that don't fit together. Blocks with any other inline markup
|
||||
(code, images, HTML) still split into runs at those boundaries.
|
||||
|
||||
Locating is best effort: a run that is not a verbatim source substring
|
||||
(entity-decoded text, backslash escapes) is skipped — it simply stays in
|
||||
the original language. A literal "<" in prose ("<1MB") is text, not
|
||||
markup, but cannot cross as-is — "<" is the prose/markup boundary on the
|
||||
wire, translators cut their output there — so it crosses encoded as the
|
||||
fullwidth "<" (``_encode``) and ``join`` decodes it back before
|
||||
validating and splicing.
|
||||
"""
|
||||
|
||||
import bisect
|
||||
import difflib
|
||||
import re
|
||||
from typing import NamedTuple
|
||||
|
||||
from pagerite.markdown import make_md
|
||||
|
||||
#: The segmentation parser: the project's own markdown-it, verbatim flavor
|
||||
#: (see make_md). Never used for rendering.
|
||||
_MD = make_md(verbatim=True)
|
||||
|
||||
#: Any Unicode letter (digits and underscore are not prose).
|
||||
_LETTER = re.compile(r"[^\W\d_]")
|
||||
|
||||
#: A GFM alert marker ([!NOTE] etc.) at the start of a blockquote's first
|
||||
#: paragraph: syntax, not prose — stripped from the first segment.
|
||||
_ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
|
||||
|
||||
#: Any {...} span: {placeholders} and attrs that ended up inside prose
|
||||
#: (inline attrs are consumed by the parser; a lone {dates} is not).
|
||||
_BRACES = re.compile(r"\{[^{}\n]*\}")
|
||||
|
||||
|
||||
def _encode(text: str) -> str:
|
||||
"""Wire form of a segment or context: a literal "<" as fullwidth "<".
|
||||
|
||||
A "<" in prose is text, not markup ("<1MB" — a tag needs a letter or
|
||||
/!?), but "<" is the prose/markup boundary on the wire (translators
|
||||
cut output at the first "<", scripts/translator.py), so it cannot
|
||||
cross as-is. join decodes it back before the pure_prose check and
|
||||
splicing — anything tag-like the model may have formed around it is
|
||||
still rejected there.
|
||||
"""
|
||||
return text.replace("<", "<")
|
||||
|
||||
#: ASCII punctuation that is plain prose to the inline parser (so
|
||||
#: pure_prose cannot catch it) but Markdown SYNTAX in a splice context:
|
||||
#: quotes close a quoted image/link title, brackets the [...] of alt and
|
||||
#: re-inserted link texts, "|" splits a table row, and "\" escapes the
|
||||
#: character after it (a trailing one eats a title's closing quote).
|
||||
#: Neutralized to Unicode look-alikes (join), which Markdown treats as
|
||||
#: plain text everywhere — the quotes are curled the way typographer=True
|
||||
#: renders them anyway.
|
||||
_NEUTRAL = str.maketrans(
|
||||
{
|
||||
'"': "”",
|
||||
"'": "’",
|
||||
"[": "[",
|
||||
"]": "]",
|
||||
"\\": "\",
|
||||
"|": "│",
|
||||
}
|
||||
)
|
||||
|
||||
#: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]",
|
||||
#: "[]" or a bare "]" (shortcut reference); the destination may nest one
|
||||
#: level of parens. Best effort — a mis-scan fails the span-reconstruction
|
||||
#: check in _linked_block and the block falls back to per-run segments.
|
||||
_LINK_TAIL = re.compile(r"\](?:\((?:\\.|[^()\\]|\([^()]*\))*\)|\[(?:\\.|[^\]])*\])?")
|
||||
|
||||
#: Weight units for mapping link boundaries from source to translation:
|
||||
#: a word counts 1 and so does every single CJK ideograph (kana runs count
|
||||
#: as one) — CJK has no spaces to count words by. Punctuation and
|
||||
#: whitespace count nothing, so mapped boundaries always land on unit
|
||||
#: starts.
|
||||
_UNIT = re.compile(
|
||||
r"[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]" # CJK ideographs: one unit each
|
||||
r"|[\u3040-\u309f\u30a0-\u30ff]+" # kana runs: one unit each
|
||||
r"|\w+" # anything else word-like (Latin, Cyrillic, Hangul, digits)
|
||||
)
|
||||
|
||||
|
||||
class Mark(NamedTuple):
|
||||
"""One inline link or paired formatting (strong/em/s) inside a
|
||||
whole-block segment: the source weight (unit count, see _UNIT) at the
|
||||
inner text's start and end (fallback for mapping the boundaries into
|
||||
the translation when fuzzy word alignment finds nothing, _find_mark),
|
||||
the exact source syntax around the text ("[" / "](url)", "**" / "**",
|
||||
...) and the source text itself — the words fuzzy alignment looks for,
|
||||
and the fallback when the mapped slice comes out empty (better an
|
||||
untranslated label than a broken "[](url)")."""
|
||||
|
||||
w_start: int
|
||||
w_end: int
|
||||
pre: str
|
||||
post: str
|
||||
inner: str
|
||||
|
||||
|
||||
class Span(NamedTuple):
|
||||
"""A segment's source span in the fragment: offsets for splicing the
|
||||
translation back, the segment's source weight and the links to
|
||||
re-insert into its translation (empty = a plain prose segment)."""
|
||||
|
||||
start: int
|
||||
end: int
|
||||
weight: int
|
||||
marks: list[Mark]
|
||||
|
||||
|
||||
def _weight(text: str) -> int:
|
||||
"""The text's weight in translation-mapping units (see _UNIT)."""
|
||||
return len(_UNIT.findall(text))
|
||||
|
||||
|
||||
def _runs(children: list) -> list[str]:
|
||||
"""Prose runs of an inline token's children, in order.
|
||||
|
||||
Text tokens merge across soft breaks into one run; every markup token
|
||||
(emphasis, links, code, images, HTML, footnote refs, hard breaks) is a
|
||||
run boundary. Link and image *text* is prose; autolink text (the URL
|
||||
itself) is not. Image tokens contribute their alt-text children and
|
||||
their title attribute.
|
||||
"""
|
||||
runs: list[str] = []
|
||||
cur: list[str] = []
|
||||
|
||||
def flush() -> None:
|
||||
if cur:
|
||||
s = "".join(cur)
|
||||
cur.clear()
|
||||
if _LETTER.search(s):
|
||||
runs.append(s)
|
||||
|
||||
skip = 0 # inside an autolink (its text is the URL — not prose)
|
||||
for t in children:
|
||||
if skip:
|
||||
if t.type == "link_close":
|
||||
skip -= 1
|
||||
continue
|
||||
if t.type == "text":
|
||||
cur.append(t.content)
|
||||
elif t.type == "softbreak":
|
||||
cur.append("\n")
|
||||
elif t.type == "link_open" and t.markup == "autolink":
|
||||
flush()
|
||||
skip = 1
|
||||
elif t.type == "image":
|
||||
flush()
|
||||
if t.children:
|
||||
runs.extend(_runs(t.children))
|
||||
title = t.attrGet("title")
|
||||
if title and _LETTER.search(title):
|
||||
runs.append(title)
|
||||
else:
|
||||
flush()
|
||||
if t.children:
|
||||
runs.extend(_runs(t.children))
|
||||
flush()
|
||||
return runs
|
||||
|
||||
|
||||
def _block_text(children: list) -> str:
|
||||
"""The block's text as a reader sees it: text runs and link texts
|
||||
merged (softbreaks as newlines); image alts, autolink URLs, code and
|
||||
other markup content excluded. Used as the translation CONTEXT for
|
||||
segments carved out of the block (link texts, partial runs): a lone
|
||||
word translates differently than the same word inside its sentence."""
|
||||
parts: list[str] = []
|
||||
skip = 0 # inside an autolink (its text is the URL)
|
||||
for t in children:
|
||||
if skip:
|
||||
if t.type == "link_close":
|
||||
skip -= 1
|
||||
continue
|
||||
if t.type == "text":
|
||||
parts.append(t.content)
|
||||
elif t.type == "softbreak":
|
||||
parts.append("\n")
|
||||
elif t.type == "link_open" and t.markup == "autolink":
|
||||
skip = 1
|
||||
elif t.type == "image":
|
||||
continue
|
||||
elif t.children:
|
||||
parts.append(_block_text(t.children))
|
||||
return "".join(parts)
|
||||
|
||||
|
||||
def _locate(source: str, needle: str, cursor: int) -> int:
|
||||
"""The needle's offset in source at/after cursor, -1 when absent.
|
||||
|
||||
An occurrence preceded by a backslash is an escaped character, not the
|
||||
token's source: keep looking (failing that, the run is skipped — it
|
||||
stays in the original language).
|
||||
"""
|
||||
pos = source.find(needle, cursor)
|
||||
while pos > 0 and source[pos - 1] == "\\":
|
||||
pos = source.find(needle, pos + 1)
|
||||
return pos
|
||||
|
||||
|
||||
def _linked_block(
|
||||
source: str, kids: list, cursor: int, strip_alert: bool
|
||||
) -> tuple[Span, str] | None:
|
||||
"""A whole-block segment for an inline of plain text, prose links and
|
||||
paired text formatting (strong/em/s): (Span, wire text) with the links
|
||||
and formatting as marks, or None when the block has any other shape —
|
||||
the caller then falls back to per-run segments.
|
||||
|
||||
The block crosses the wire as one prose piece, link texts and formatted
|
||||
text inline (the model is never shown any Markdown — it mangles it),
|
||||
so a translation that inflects or reorders around them stays coherent;
|
||||
join re-inserts the link/formatting syntax at weight-mapped positions.
|
||||
The source span is located piece by piece and verified by
|
||||
reconstruction; anything not byte-exact (entities, escapes, an odd
|
||||
link tail) bails to the fallback.
|
||||
"""
|
||||
pieces: list[
|
||||
tuple[str, str]
|
||||
] = [] # (text, mark): "" plain, "link", else the delimiter
|
||||
buf: list[str] = [] # current plain piece
|
||||
link: list[str] | None = None # current mark's text parts
|
||||
mark_kind = "" # the current mark's opener ("link" or the delimiter)
|
||||
for tok in kids:
|
||||
if tok.type in ("link_open", "strong_open", "em_open", "s_open"):
|
||||
if link is not None or tok.markup == "autolink":
|
||||
return None
|
||||
if buf:
|
||||
pieces.append(("".join(buf), ""))
|
||||
buf = []
|
||||
link = []
|
||||
mark_kind = "link" if tok.type == "link_open" else tok.markup
|
||||
elif tok.type in ("link_close", "strong_close", "em_close", "s_close"):
|
||||
if (
|
||||
link is None
|
||||
or ("link" if tok.type == "link_close" else tok.markup) != mark_kind
|
||||
):
|
||||
return None
|
||||
inner = "".join(link)
|
||||
if not _LETTER.search(inner):
|
||||
return None
|
||||
pieces.append((inner, mark_kind))
|
||||
link = None
|
||||
elif tok.type in ("text", "softbreak"):
|
||||
(link if link is not None else buf).append(
|
||||
"\n" if tok.type == "softbreak" else tok.content
|
||||
)
|
||||
else: # code, images, HTML, footnote refs: run boundaries
|
||||
return None
|
||||
if link is not None:
|
||||
return None # unbalanced (the parser should not do this)
|
||||
if buf:
|
||||
pieces.append(("".join(buf), ""))
|
||||
if not any(mark for _, mark in pieces):
|
||||
return None
|
||||
if strip_alert and pieces and not pieces[0][1]:
|
||||
# A GFM alert marker leading the blockquote's first paragraph is
|
||||
# syntax; strip it from the wire text (it stays out of the span).
|
||||
first = _ALERT.sub("", pieces[0][0], count=1)
|
||||
if first.strip():
|
||||
pieces[0] = (first, "")
|
||||
else:
|
||||
pieces.pop(0)
|
||||
if not pieces:
|
||||
return None
|
||||
raw = "".join(text for text, _ in pieces)
|
||||
lead = len(raw) - len(raw.lstrip())
|
||||
wire = raw.strip()
|
||||
if not _LETTER.search(wire) or _BRACES.search(wire):
|
||||
return None
|
||||
# Locate each piece verbatim, in order; the source slices between the
|
||||
# located pieces are then the link syntax, exact by construction.
|
||||
located: list[tuple[int, int]] = []
|
||||
pos = cursor
|
||||
for text_, _ in pieces:
|
||||
at = _locate(source, text_, pos)
|
||||
if at == -1:
|
||||
return None
|
||||
located.append((at, at + len(text_)))
|
||||
pos = at + len(text_)
|
||||
span_start, span_end = located[0][0], located[-1][1]
|
||||
marks: list[Mark] = []
|
||||
offset = 0 # raw (pre-strip) plain-text offset of the current piece
|
||||
for i, ((text_, kind), (s, e)) in enumerate(zip(pieces, located)):
|
||||
if not kind:
|
||||
offset += len(text_)
|
||||
continue
|
||||
# The syntax around the text: the gap between pieces goes to the
|
||||
# mark on its left as post (so between two marks the whole "](u)["
|
||||
# or "**" is the first's post); a block-leading mark takes its
|
||||
# opener in front of its text ("[" or the delimiter), a
|
||||
# block-trailing one the scanned link tail or the close delimiter.
|
||||
if i == 0:
|
||||
opener = "[" if kind == "link" else kind
|
||||
if s < len(opener) or source[s - len(opener) : s] != opener:
|
||||
return None
|
||||
pre, span_start = opener, s - len(opener)
|
||||
elif pieces[i - 1][1]:
|
||||
pre = "" # the previous mark's post covers the whole gap
|
||||
else:
|
||||
pre = source[located[i - 1][1] : s]
|
||||
if i + 1 < len(pieces):
|
||||
post = source[e : located[i + 1][0]]
|
||||
elif kind == "link":
|
||||
m = _LINK_TAIL.match(source, e)
|
||||
if m is None:
|
||||
return None
|
||||
post, span_end = m.group(), m.end()
|
||||
else:
|
||||
if source[e : e + len(kind)] != kind:
|
||||
return None
|
||||
post, span_end = kind, e + len(kind)
|
||||
ps = min(max(offset - lead, 0), len(wire))
|
||||
pe = min(max(offset + len(text_) - lead, 0), len(wire))
|
||||
if pe <= ps:
|
||||
return None
|
||||
marks.append(
|
||||
Mark(_weight(wire[:ps]), _weight(wire[:pe]), pre, post, wire[ps:pe])
|
||||
)
|
||||
offset += len(text_)
|
||||
# Verify: the marks must reconstruct the source span exactly (the only
|
||||
# real risk is the guessed tail of a trailing link).
|
||||
rec: list[str] = []
|
||||
mi = 0
|
||||
for text_, kind in pieces:
|
||||
if kind:
|
||||
mark = marks[mi]
|
||||
mi += 1
|
||||
rec += [mark.pre, text_, mark.post]
|
||||
else:
|
||||
rec.append(text_)
|
||||
if source[span_start:span_end] != "".join(rec):
|
||||
return None
|
||||
return Span(span_start, span_end, _weight(wire), marks), _encode(wire)
|
||||
|
||||
|
||||
def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
"""Split a fragment into (spans, segments, contexts): prose segments to
|
||||
translate, their source spans in ``text`` for splicing the translations
|
||||
back, and per-segment translation context.
|
||||
|
||||
A block of plain text, prose links and paired formatting (strong/em/s)
|
||||
becomes ONE segment (link/formatted text inline, in context, Markdown
|
||||
stripped), the links and formatting recorded as marks on its Span for
|
||||
weight-mapped re-insertion in join. Other blocks split into text runs
|
||||
at markup boundaries; runs containing {...} spans are carved further —
|
||||
the braces stay out of the wire text. A run that cannot be located
|
||||
verbatim in the source contributes no segment. A segment's context is
|
||||
its block's plain text when the segment was carved OUT of a larger
|
||||
block (a partial run); a segment that IS the whole block (a plain
|
||||
paragraph, a heading, a linked block) is self-contextualizing and gets
|
||||
"".
|
||||
"""
|
||||
spans: list[Span] = []
|
||||
segments: list[str] = []
|
||||
contexts: list[str] = []
|
||||
cursor = 0
|
||||
blockquote_fresh = 0 # blockquote depth whose first inline is upcoming
|
||||
|
||||
def emit(run: str, at: int, ctx: str) -> None:
|
||||
"""Carve {...} spans out of the located run; emit the prose pieces,
|
||||
stripped — padding whitespace stays in the template, off the wire.
|
||||
A literal "<" crosses encoded (``_encode``): it is text, not
|
||||
markup, but the wire keeps "<" as the prose/markup boundary."""
|
||||
pieces = []
|
||||
pos = 0
|
||||
for m in _BRACES.finditer(run):
|
||||
pieces.append((pos, m.start()))
|
||||
pos = m.end()
|
||||
pieces.append((pos, len(run)))
|
||||
for p0, p1 in pieces:
|
||||
raw = run[p0:p1]
|
||||
piece = raw.strip()
|
||||
if _LETTER.search(piece):
|
||||
start = at + p0 + (len(raw) - len(raw.lstrip()))
|
||||
spans.append(Span(start, start + len(piece), 0, []))
|
||||
segments.append(_encode(piece))
|
||||
contexts.append(ctx)
|
||||
|
||||
tokens = _MD.parse(text)
|
||||
for t in tokens:
|
||||
if t.type == "blockquote_open":
|
||||
blockquote_fresh += 1
|
||||
elif t.type == "blockquote_close":
|
||||
blockquote_fresh -= 1
|
||||
elif t.type == "inline":
|
||||
kids = t.children or []
|
||||
# An alert marker ([!NOTE]) leading a blockquote's first
|
||||
# paragraph is syntax; both paths strip it. (Only the first
|
||||
# inline of the blockquote can carry it — the flag clears on
|
||||
# the first inline seen.)
|
||||
alert = bool(blockquote_fresh)
|
||||
blockquote_fresh = 0
|
||||
linked = _linked_block(text, kids, cursor, strip_alert=alert)
|
||||
if linked is not None:
|
||||
span, wire = linked
|
||||
spans.append(span)
|
||||
segments.append(wire)
|
||||
contexts.append("")
|
||||
cursor = span.end
|
||||
continue
|
||||
runs = _runs(kids)
|
||||
block = _encode(_block_text(kids).strip())
|
||||
if alert and runs:
|
||||
run = _ALERT.sub("", runs[0], count=1)
|
||||
if _LETTER.search(run):
|
||||
runs[0] = run
|
||||
else:
|
||||
runs.pop(0)
|
||||
for run in runs:
|
||||
ctx = block if block and _encode(run.strip()) != block else ""
|
||||
pos = _locate(text, run, cursor)
|
||||
if pos != -1:
|
||||
emit(run, pos, ctx)
|
||||
cursor = pos + len(run)
|
||||
elif "\n" in run:
|
||||
# Indented continuation lines etc. break the verbatim
|
||||
# match: locate each line separately instead.
|
||||
for part in run.split("\n"):
|
||||
if not _LETTER.search(part):
|
||||
continue
|
||||
pos = _locate(text, part, cursor)
|
||||
if pos != -1:
|
||||
emit(part, pos, ctx)
|
||||
cursor = pos + len(part)
|
||||
return spans, segments, contexts
|
||||
|
||||
|
||||
#: Block-level Markdown a translation must not introduce: a segment is
|
||||
#: spliced INSIDE a block of the fragment, so a line starting a heading,
|
||||
#: quote, list, code/container fence or a setext/thematic-break underline
|
||||
#: would break the fragment's block structure — a ``` or ::: line eats the
|
||||
#: rest of the fence it lands in, closing fence included. pure_prose only
|
||||
#: parses inline and lets such lines through as softbreak prose, so join
|
||||
#: rejects them here. Blank lines split the host block and are rejected
|
||||
#: too (a faithful translation of a single block has none).
|
||||
_BLOCK = re.compile(
|
||||
r"^[ \t]*(?:#{1,6}(?:[ \t]|$)|>[ \t]?|(?:[-+*]|\d{1,9}[.)])[ \t]|`{3,}|~{3,}|:{3,}(?:[ \t]|$)"
|
||||
r"|-(?:[ \t]*-){2,}[ \t]*$|=[ =]*$|_(?:[ \t]*_){2,}[ \t]*$)",
|
||||
re.M,
|
||||
)
|
||||
_BLANK = re.compile(r"\n[ \t]*\n")
|
||||
|
||||
|
||||
def pure_prose(text: str) -> bool:
|
||||
"""True when the text parses as nothing but prose (text and softbreak
|
||||
tokens) — the acceptance test for a translated segment: the model may
|
||||
not return markup of its own (a `<br>` here would splice live HTML into
|
||||
the fragment)."""
|
||||
children = _MD.parseInline(text)[0].children or []
|
||||
return all(t.type in ("text", "softbreak") for t in children)
|
||||
|
||||
|
||||
def _word_sim(a: str, b: str) -> float:
|
||||
"""How likely two words are the same term across a translation, 0..1.
|
||||
|
||||
A case-folded exact match is 1; otherwise the better of the sequence
|
||||
ratio and the shared-prefix ratio — inflection and derivational change
|
||||
mostly move the ending ("banana" -> "banaanilla") or drop an article or
|
||||
preposition around it. Case-folded so capitalization differences across
|
||||
languages don't hide a term, with a small bonus when BOTH sides are
|
||||
capitalized: a mid-sentence capital on both sides is likely the same
|
||||
name (capitalization conventions differ per language, so its absence
|
||||
proves nothing).
|
||||
"""
|
||||
bonus = 0.1 if a[:1].isupper() and b[:1].isupper() else 0.0
|
||||
a, b = a.casefold(), b.casefold()
|
||||
if a == b:
|
||||
return 1.0
|
||||
prefix = 0
|
||||
for ca, cb in zip(a, b):
|
||||
if ca != cb:
|
||||
break
|
||||
prefix += 1
|
||||
sim = max(
|
||||
difflib.SequenceMatcher(None, a, b).ratio(),
|
||||
prefix / max(len(a), len(b)),
|
||||
)
|
||||
return min(1.0, sim + bonus)
|
||||
|
||||
|
||||
#: Alignment costs for _find_mark: skipping a translation word (an article
|
||||
#: or preposition the target language added) is cheap, skipping a source
|
||||
#: word (one the translation dropped) costs more — a mark whose words
|
||||
#: mostly vanished is no match at all. Every matched pair pays _MATCH, so
|
||||
#: aligning a word to a lookalike-nothing (similarity below _MATCH) is
|
||||
#: worse than skipping it.
|
||||
_GAP_T = 0.25
|
||||
_GAP_S = 0.6
|
||||
_MATCH = 0.3
|
||||
|
||||
|
||||
def _find_mark(
|
||||
src: list[str], units: list[re.Match], start: int
|
||||
) -> tuple[int, int] | None:
|
||||
"""Locate a mark's source words in the translation's units (from unit
|
||||
index ``start`` on), as the (start, end) unit-index span of the best
|
||||
fuzzy alignment; None when no alignment is convincing (the caller falls
|
||||
back to the weight ratio).
|
||||
|
||||
Word-for-word alignment with skips (_word_sim per pair, _GAP_T/_GAP_S
|
||||
per skipped word): reordering is handled by the search itself, an added
|
||||
or dropped article/preposition by the skip penalties. Accepted only
|
||||
with an anchor — one pair of similarity >= 0.7 — and a decent average,
|
||||
so a fully reworded label doesn't snap onto chance lookalikes.
|
||||
"""
|
||||
tgt = [u.group() for u in units[start:]]
|
||||
n, m = len(src), len(tgt)
|
||||
if not n or not m:
|
||||
return None
|
||||
# dp[i][j]: best score aligning src[:i] to tgt[:j]; a free tail (the
|
||||
# answer is the best dp[n][j] over j) keeps trailing words costless.
|
||||
dp = [[0.0] * (m + 1) for _ in range(n + 1)]
|
||||
back: list[list[tuple[int, int]]] = [[(0, 0)] * (m + 1) for _ in range(n + 1)]
|
||||
for i in range(1, n + 1):
|
||||
dp[i][0] = dp[i - 1][0] - _GAP_S
|
||||
back[i][0] = (i - 1, 0)
|
||||
for j in range(1, m + 1):
|
||||
options = [
|
||||
(
|
||||
dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH,
|
||||
(i - 1, j - 1),
|
||||
),
|
||||
(dp[i][j - 1] - _GAP_T, (i, j - 1)),
|
||||
(dp[i - 1][j] - _GAP_S, (i - 1, j)),
|
||||
]
|
||||
dp[i][j], back[i][j] = max(options, key=lambda o: o[0])
|
||||
j_end = max(range(m + 1), key=lambda j: dp[n][j])
|
||||
pairs: list[tuple[int, int]] = [] # matched (source, target) indices
|
||||
i, j = n, j_end
|
||||
while i > 0:
|
||||
pi, pj = back[i][j]
|
||||
if (pi, pj) == (i - 1, j - 1):
|
||||
pairs.append((i - 1, j - 1))
|
||||
i, j = pi, pj
|
||||
if not pairs:
|
||||
return None
|
||||
pairs.reverse() # backtracking collected them last-first
|
||||
sims = [_word_sim(src[a], tgt[t]) for a, t in pairs]
|
||||
# Weak pairs at the span's ends are not part of the label (a declined
|
||||
# neighbor the DP matched for a pittance) — trim them off.
|
||||
while len(sims) > 1 and sims[0] < 0.5:
|
||||
pairs.pop(0)
|
||||
sims.pop(0)
|
||||
while len(sims) > 1 and sims[-1] < 0.5:
|
||||
pairs.pop()
|
||||
sims.pop()
|
||||
if max(sims) < 0.7 or sum(sims) / len(sims) < 0.45:
|
||||
return None
|
||||
return start + pairs[0][1], start + pairs[-1][1] + 1
|
||||
|
||||
|
||||
def _place_marks(translation: str, weight: int, marks: list[Mark]) -> str | None:
|
||||
"""Re-insert a whole-block segment's links into its translation.
|
||||
|
||||
Each mark's boundaries are found by fuzzy word-form alignment
|
||||
(_find_mark): the mark's source words are matched against the
|
||||
translation's units by form similarity — no markers on the wire
|
||||
(sentinels never survived the model), no assumption that word order or
|
||||
count survived either. Slicing exactly at unit boundaries keeps the
|
||||
whitespace between the mark and its neighbors in the plain text, where
|
||||
it belongs. A mark with no convincing alignment falls back to its
|
||||
source weight ratio (units before the boundary / total applied to the
|
||||
translation's units) — the pre-fuzz heuristic, still the CJK path,
|
||||
where form similarity across scripts is nil. A boundary landing empty
|
||||
degrades to the source link text: better an untranslated label than a
|
||||
broken "[](url)". None when the translation has no units to map onto
|
||||
(the caller rejects the result).
|
||||
"""
|
||||
units = list(_UNIT.finditer(translation))
|
||||
total = len(units)
|
||||
if not total or not weight:
|
||||
return None
|
||||
starts = [u.start() for u in units]
|
||||
bounds = starts + [len(translation)]
|
||||
out: list[str] = []
|
||||
cur = 0 # char cursor: never before the previous mark's end
|
||||
ucur = 0 # unit cursor, the same monotonicity in unit indices
|
||||
for mark in marks:
|
||||
found = _find_mark(_UNIT.findall(mark.inner), units, ucur)
|
||||
if found is not None:
|
||||
u1, u2 = found
|
||||
x1, x2 = units[u1].start(), units[u2 - 1].end()
|
||||
else:
|
||||
x1 = bounds[min(round(mark.w_start / weight * total), total)]
|
||||
x2 = bounds[min(round(mark.w_end / weight * total), total)]
|
||||
x1 = max(x1, cur)
|
||||
x2 = max(x2, x1)
|
||||
# The slice ends at the next unit's start, so the whitespace
|
||||
# and punctuation before that unit is inside it — but it
|
||||
# belongs BETWEEN the mark and the following word, not in the
|
||||
# inner text: end the inner text at its last unit and leave
|
||||
# the rest for the following slice.
|
||||
raw = translation[x1:x2]
|
||||
inner_units = list(_UNIT.finditer(raw))
|
||||
x2 = x1 + inner_units[-1].end() if inner_units else x1
|
||||
inner = translation[x1:x2].strip() or mark.inner
|
||||
out += [translation[cur:x1], mark.pre, inner, mark.post]
|
||||
cur = x2
|
||||
ucur = bisect.bisect_left(starts, x2)
|
||||
out.append(translation[cur:])
|
||||
return "".join(out)
|
||||
|
||||
|
||||
def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
|
||||
"""Splice translated segments back into the original fragment; None on
|
||||
any validation failure (count mismatch, empty, non-prose or
|
||||
block-structure segment) — the caller drops the result and the fragment
|
||||
stays pending. Segments with marks (a block that crossed as one piece)
|
||||
get their links re-inserted at weight-mapped positions after the prose
|
||||
check.
|
||||
|
||||
Markdown-significant ASCII punctuation that pure_prose cannot see
|
||||
(plain text inline, syntax in the splice context — quoted titles, alt
|
||||
and link texts, table rows) is neutralized to Unicode look-alikes
|
||||
(``_NEUTRAL``) before splicing and mark placement (the swap is
|
||||
char-for-char, so unit alignment is unaffected); lines that would
|
||||
start a new block (a heading, a ``` or ::: fence — they would eat the
|
||||
rest of the block/fence they land in) reject the result outright
|
||||
(``_BLOCK``, ``_BLANK``)."""
|
||||
if len(texts) != len(spans):
|
||||
return None
|
||||
out: list[str] = []
|
||||
cursor = 0
|
||||
for span, translation in zip(spans, texts):
|
||||
# Decode the wire form ("<" back to "<") first: pure_prose then
|
||||
# validates exactly what gets spliced — a "<" the model formed
|
||||
# into anything tag-like is markup and rejects the result.
|
||||
translation = translation.replace("<", "<")
|
||||
if (
|
||||
not translation.strip()
|
||||
or not pure_prose(translation)
|
||||
or _BLOCK.search(translation)
|
||||
or _BLANK.search(translation.strip())
|
||||
):
|
||||
return None
|
||||
translation = translation.translate(_NEUTRAL)
|
||||
if span.marks:
|
||||
translation = _place_marks(translation, span.weight, span.marks)
|
||||
if translation is None:
|
||||
return None
|
||||
out.append(original[cursor : span.start])
|
||||
out.append(translation)
|
||||
cursor = span.end
|
||||
out.append(original[cursor:])
|
||||
return "".join(out)
|
||||
|
||||
|
||||
def has_prose(text: str) -> bool:
|
||||
"""True when the fragment yields at least one translatable segment.
|
||||
Chunks that are all markup, code, placeholders or reference definitions
|
||||
have no business reaching the model: every language renders them from
|
||||
the original chunk."""
|
||||
return bool(split(text)[1])
|
||||
@@ -0,0 +1,363 @@
|
||||
"""Shared core: site constants, the kanta database, and the render cache.
|
||||
|
||||
Everything the route modules (files, api, tracking, pages) need that is not
|
||||
a route itself: environment-derived paths and tunables, the ``Data`` root
|
||||
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
|
||||
store, the page render cache
|
||||
(``_render_html``/``_cached_body``/``_html_response`` plus the
|
||||
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
|
||||
content/settings write), the translator ``dispatcher``, the slug charset
|
||||
helpers, and the database bootstrap hooks (demo seed, translator defaults).
|
||||
Importable by every other pagerite module without cycles.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import secrets
|
||||
from datetime import UTC, datetime
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
|
||||
import blake3
|
||||
from fastapi import HTTPException, Request
|
||||
from fastapi.responses import Response
|
||||
from kanta import Kanta
|
||||
from zstandard import ZstdCompressor
|
||||
|
||||
from pagerite import analytics, i18n, seed, translate, views
|
||||
from pagerite.__main__ import DEVMODE
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.config import load
|
||||
from pagerite.data import (
|
||||
Data,
|
||||
Node,
|
||||
append_order,
|
||||
find_slot,
|
||||
prettify,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
|
||||
config = load()
|
||||
|
||||
# Site identity: the hostname comes from the CLI (first positional argument,
|
||||
# passed in PAGERITE_CONFIG) and names the per-site data directory
|
||||
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
|
||||
HOSTNAME = config.hostname
|
||||
SITE_DIR = Path(HOSTNAME)
|
||||
#: Public origin of the site, used for absolute social/canonical/sitemap
|
||||
#: URLs. Localhost serves varying ports, so it falls back to the request's
|
||||
#: own base URL instead.
|
||||
SITE_URL = f"https://{HOSTNAME}" if HOSTNAME != "localhost" else ""
|
||||
|
||||
DB_PATH = os.getenv("PAGERITE_DB", str(SITE_DIR / "content.kantadb"))
|
||||
|
||||
# Visit analytics go to their own JSON file, not the kanta database.
|
||||
ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json")))
|
||||
analytics_store = analytics.Store(ANALYTICS_PATH)
|
||||
|
||||
# Content-addressed file store (uploads, seed assets, fetched favicons):
|
||||
# files on disk under hash-prefixed names, cached in RAM, served at /_f/.
|
||||
FILES_DIR = Path(os.getenv("PAGERITE_FILES", str(SITE_DIR / "files")))
|
||||
|
||||
# Uploaded images are thumbnailed to this size and recompressed to AVIF
|
||||
# (primary), with WebP and JPEG fallbacks re-encoded from the AVIF at
|
||||
# somewhat lower quality (similar or smaller file size); the untouched
|
||||
# original is kept alongside as ``<hash>.orig<ext>`` (never served).
|
||||
IMAGE_MAXSIZE = 1920
|
||||
IMAGE_QUALITY = 60
|
||||
IMAGE_WEBP_QUALITY = 50
|
||||
IMAGE_JPG_QUALITY = 55
|
||||
|
||||
# Favicons get the same derivatives but thumbnailed much smaller — 192px
|
||||
# is plenty (browsers scale down for the 16x16 tab icon themselves).
|
||||
FAVICON_MAXSIZE = 192
|
||||
|
||||
# Our own data root; kanta edits it in place, reads are plain attribute access.
|
||||
data = Data()
|
||||
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
|
||||
|
||||
# Dynamic HTML is compressed per request at level 9 (static assets are
|
||||
# already pre-compressed by fastapi-vue's Frontend).
|
||||
_zstd = ZstdCompressor(9)
|
||||
|
||||
|
||||
def _render_html(
|
||||
kind: str,
|
||||
path: str,
|
||||
base_url: str,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> str:
|
||||
"""Render one of the generated pages (see _html_response)."""
|
||||
if kind == "page":
|
||||
# A selected language without an actual translation renders the
|
||||
# original (translation is None; see docs/localization.md).
|
||||
original = i18n.primary_lang(data.menu, path)
|
||||
translation = (
|
||||
i18n.get_translation(data, path, lang) if lang != original else None
|
||||
)
|
||||
return views.render_page(
|
||||
data.menu,
|
||||
data,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
base_url,
|
||||
transition=data.transition,
|
||||
lang=lang,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if kind == "category":
|
||||
# A category has no Markdown of its own; only the title map
|
||||
# localizes (heading, navigation, card text).
|
||||
original = i18n.primary_lang(data.menu, path)
|
||||
translation = (
|
||||
i18n.Translation(titles=i18n.title_map(data, lang))
|
||||
if lang != original
|
||||
else None
|
||||
)
|
||||
return views.render_category(
|
||||
data.menu,
|
||||
data,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
base_url,
|
||||
transition=data.transition,
|
||||
lang=lang,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if kind == "not-found":
|
||||
return views.render_not_found(
|
||||
data.menu,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
transition=data.transition,
|
||||
)
|
||||
return views.render_analytics(
|
||||
data.menu,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
transition=data.transition,
|
||||
)
|
||||
|
||||
|
||||
# Render generation: bumped (and the body cache cleared) by every
|
||||
# content/settings write, so page ETags and cached copies invalidate when
|
||||
# navigation-affecting changes happen. In-memory only — not database state.
|
||||
_render_gen = 0
|
||||
|
||||
|
||||
def _invalidate_pages() -> None:
|
||||
"""Drop cached page bodies and bump the render generation (ETags);
|
||||
any content change also re-runs translation dispatch."""
|
||||
global _render_gen
|
||||
_render_gen += 1
|
||||
_cached_body.cache_clear()
|
||||
dispatcher.schedule()
|
||||
|
||||
|
||||
@lru_cache(maxsize=128)
|
||||
def _cached_body(
|
||||
kind: str,
|
||||
path: str,
|
||||
base_url: str,
|
||||
zstd: bool,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> bytes:
|
||||
"""Rendered page body; cleared by _invalidate_pages on any
|
||||
content/settings change. base_url feeds the social meta URLs, zstd
|
||||
selects the stored encoding (both variants are cached rather than
|
||||
re-compressed) and lang the selected language (not the raw
|
||||
Accept-Language header, which would blow up the cache key space).
|
||||
link_lang is the ?lang= override replicated onto the navigation links:
|
||||
a query render and a header-selected render of the same language differ
|
||||
in their links, so they are cached separately.
|
||||
"""
|
||||
body = _render_html(kind, path, base_url, lang, link_lang).encode()
|
||||
return _zstd.compress(body) if zstd else body
|
||||
|
||||
|
||||
def _html_response(
|
||||
request: Request,
|
||||
kind: str,
|
||||
path: str,
|
||||
status_code: int = 200,
|
||||
headers: dict | None = None,
|
||||
etag: bool = False,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> Response:
|
||||
"""Response for a generated page, zstd-compressed when the client
|
||||
accepts it (no gzip fallback).
|
||||
|
||||
Done per handler rather than in middleware so that Frontend's
|
||||
already-compressed asset responses are never touched. The ETag stays
|
||||
identical across encodings (revalidation compares it before
|
||||
compression); ``vary: accept-encoding`` keeps caches from mixing the
|
||||
representations. In dev the cache is bypassed so theme/design edits on
|
||||
disk apply immediately.
|
||||
|
||||
``etag=True`` derives the validator from a blake3 hash of the
|
||||
(uncompressed) body — for pages like /_a that have no Node whose
|
||||
modified timestamp could serve as one — and answers matching
|
||||
if-none-match revalidations with a 304.
|
||||
"""
|
||||
zstd = "zstd" in request.headers.get("accept-encoding", "")
|
||||
# Absolute social/canonical URLs use the site's public origin; on
|
||||
# localhost (varying ports) fall back to the request's own base URL.
|
||||
base_url = SITE_URL or str(request.base_url).rstrip("/")
|
||||
if DEVMODE:
|
||||
identity = _render_html(kind, path, base_url, lang, link_lang).encode()
|
||||
body = _zstd.compress(identity) if zstd else identity
|
||||
else:
|
||||
identity = _cached_body(kind, path, base_url, False, lang, link_lang)
|
||||
body = (
|
||||
_cached_body(kind, path, base_url, True, lang, link_lang)
|
||||
if zstd
|
||||
else identity
|
||||
)
|
||||
h = dict(headers or {})
|
||||
# Content varies by language (Accept-Language selects a translation)
|
||||
# and by encoding; keep caches from mixing either representation.
|
||||
h["vary"] = "accept-language" + (", accept-encoding" if zstd else "")
|
||||
if etag:
|
||||
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
|
||||
h["etag"] = tag
|
||||
if request.headers.get("if-none-match") == tag:
|
||||
return Response(status_code=304, headers=h)
|
||||
if zstd:
|
||||
h["content-encoding"] = "zstd"
|
||||
return Response(body, status_code, h, media_type="text/html")
|
||||
|
||||
|
||||
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
|
||||
|
||||
|
||||
def _is_reserved(path: str) -> bool:
|
||||
"""Slug shape that content may never use: each segment must be lower-case
|
||||
ASCII letters, digits, hyphens and underscores (underscores may not be
|
||||
the first character), and dots are never allowed.
|
||||
"""
|
||||
if path == "":
|
||||
return False
|
||||
return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
|
||||
|
||||
|
||||
def _check_reserved(path: str) -> None:
|
||||
"""Reject paths that do not follow the slug charset."""
|
||||
if _is_reserved(path):
|
||||
raise HTTPException(
|
||||
400,
|
||||
'slugs may only use a-z, 0-9, "-" and "_" (not as the first character), and no dots',
|
||||
)
|
||||
|
||||
|
||||
def _ensure(menu: dict[str, Node], path: str) -> Node:
|
||||
"""Return the node at ``path``, creating it and any missing ancestors
|
||||
(content-less category labels) appended at the end of their level."""
|
||||
nodes = menu
|
||||
node = None
|
||||
for seg in path.split("/"):
|
||||
node = nodes.get(seg)
|
||||
if node is None:
|
||||
node = Node(title=prettify(seg), order=append_order(nodes))
|
||||
nodes[seg] = node
|
||||
nodes = node.children
|
||||
return node
|
||||
|
||||
|
||||
def _remove_page(menu: dict[str, Node], path: str) -> bool:
|
||||
"""Delete the node at ``path`` (inside a transaction).
|
||||
|
||||
A node with children becomes a content-less category label; a childless
|
||||
node is removed entirely. Returns False if the path does not exist.
|
||||
"""
|
||||
slot = find_slot(menu, path)
|
||||
node = slot[0].get(slot[1]) if slot else None
|
||||
if node is None:
|
||||
return False
|
||||
if node.children:
|
||||
node.chunks = None
|
||||
node.modified = datetime.now(UTC)
|
||||
else:
|
||||
del slot[0][slot[1]]
|
||||
return True
|
||||
|
||||
|
||||
def _store_seed_file(
|
||||
markdown: str, banner: str, orig: str, body: bytes
|
||||
) -> tuple[str, str]:
|
||||
"""Store a seed file content-addressed and point references at /_f/.
|
||||
|
||||
Images get the same AVIF/WebP/JPEG derivatives as uploads and are
|
||||
linked extension-less; other content is stored as-is with its
|
||||
extension."""
|
||||
from pagerite.files import _ext, store_image # lazy: files imports state
|
||||
|
||||
ext = _ext(orig)
|
||||
name = store_image(body, ext, derive=ext != ".gif")
|
||||
markdown = markdown.replace(f"]({orig}", f"](/_f/{name}")
|
||||
banner = banner.replace(f'src="/{orig}"', f'src="/_f/{name}"')
|
||||
banner = banner.replace(f'src="{orig}"', f'src="/_f/{name}"')
|
||||
return markdown, banner
|
||||
|
||||
|
||||
@kanta.bootstrap
|
||||
def _seed(data: Data) -> None:
|
||||
"""Write the demo pages on database creation (never on existing dbs)."""
|
||||
for path in seed.PAGES:
|
||||
title, markdown, files, banner, order, design = seed.PAGES[path]
|
||||
for orig, body in files.items():
|
||||
markdown, banner = _store_seed_file(markdown, banner, orig, body)
|
||||
node = _ensure(data.menu, path)
|
||||
node.title = title
|
||||
# Empty markdown means a pure category label (e.g. "showcase",
|
||||
# seeded only to carry a banner design): leave chunks as None so
|
||||
# the node renders the placeholder and nav points at its children.
|
||||
if markdown:
|
||||
node.chunks = store_chunks(data.chunks, markdown)
|
||||
node.banner = banner
|
||||
node.banner_design = design
|
||||
node.order = order
|
||||
|
||||
|
||||
#: Translator key format: 12 lowercase alphanumeric characters — not
|
||||
#: brute-forceable over a WebSocket handshake, still human-manageable.
|
||||
#: The editor's lang tab generates further keys in the same format.
|
||||
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
|
||||
|
||||
|
||||
@kanta.bootstrap
|
||||
def _translator_defaults(data: Data) -> None:
|
||||
"""Translator defaults on database creation: the first service key and
|
||||
the wanted target languages (Spanish and Chinese — English is the
|
||||
original language, never a translation target). Further keys are
|
||||
managed in the editor shell's lang tab."""
|
||||
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
|
||||
data.translate_keys[key] = "default"
|
||||
data.translate_langs = {"es": True, "zh": True}
|
||||
|
||||
|
||||
# The translator dispatcher — protocol, connected clients and the job
|
||||
# pipeline live in translate.py; its WebSocket route is in api.py.
|
||||
dispatcher = translate.Dispatcher(data, kanta, _invalidate_pages)
|
||||
@@ -111,12 +111,13 @@ article h3 {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent);
|
||||
border-inline-start-color: var(--accent);
|
||||
background: color-mix(var(--accent) 6%, transparent);
|
||||
padding: 0.4rem 0.9rem;
|
||||
/* Keep the quoted text on the paragraph edge: the tinted box extends
|
||||
past it by its own border/padding, like code blocks. */
|
||||
margin: 0 -0.9rem 1rem calc(-0.25rem - 0.9rem);
|
||||
margin: 0 0 1rem;
|
||||
margin-inline: calc(-0.25rem - 0.9rem) -0.9rem;
|
||||
border-radius: 6px;
|
||||
}
|
||||
|
||||
@@ -125,9 +126,9 @@ blockquote {
|
||||
bar stays in both. */
|
||||
pre {
|
||||
border: 1px solid transparent;
|
||||
border-left: 0.25rem solid var(--accent);
|
||||
border-inline-start: 0.25rem solid var(--accent);
|
||||
/* Text on the paragraph edge: the box extends by padding + border. */
|
||||
margin-left: calc(-0.8rem - 0.25rem);
|
||||
margin-inline-start: calc(-0.8rem - 0.25rem);
|
||||
border-radius: 6px;
|
||||
}
|
||||
|
||||
|
||||
@@ -144,9 +144,9 @@ article h1 {
|
||||
font-weight: 700;
|
||||
padding-bottom: 0.5rem;
|
||||
/* The hazard-stripe underline breaks out of the page box: the negative
|
||||
right margin extends the h1's box (and thus its background) all the
|
||||
way to the viewport's right edge. */
|
||||
margin-right: calc((100% - 100vw) / 2);
|
||||
end margin extends the h1's box (and thus its background) all the
|
||||
way to the viewport's edge on that side. */
|
||||
margin-inline-end: calc((100% - 100vw) / 2);
|
||||
background:
|
||||
linear-gradient(-55deg,
|
||||
transparent 0 0.2rem,
|
||||
@@ -186,20 +186,21 @@ article ul ul ul li::before {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
background: color-mix(var(--accent2) 6%, transparent);
|
||||
padding: 0.25rem 0.75rem;
|
||||
/* Keep the quoted text on the paragraph edge: the tinted box extends
|
||||
past it by its own border/padding, like code blocks. */
|
||||
margin: 0 -0.75rem 1rem -1rem;
|
||||
margin: 0 0 1rem;
|
||||
margin-inline: -1rem -0.75rem;
|
||||
}
|
||||
|
||||
/* Code follows the color scheme; the dark-scheme well joins the violet
|
||||
family (--code-bg above). The orange side bar stays in both. */
|
||||
pre {
|
||||
border-left: 0.25rem solid var(--accent);
|
||||
border-inline-start: 0.25rem solid var(--accent);
|
||||
/* Text on the paragraph edge: the box extends by padding + border. */
|
||||
margin-left: calc(-0.8rem - 0.25rem);
|
||||
margin-inline-start: calc(-0.8rem - 0.25rem);
|
||||
border-radius: 3px;
|
||||
}
|
||||
|
||||
|
||||
@@ -66,7 +66,7 @@ article ul li::before {
|
||||
content: "◆";
|
||||
color: var(--accent);
|
||||
font-size: 0.8em;
|
||||
margin-left: calc(-1 * var(--list-indent) / 0.8);
|
||||
margin-inline-start: calc(-1 * var(--list-indent) / 0.8);
|
||||
width: calc(var(--list-indent) / 0.8);
|
||||
}
|
||||
|
||||
@@ -80,7 +80,7 @@ article ul ul ul li::before {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
}
|
||||
|
||||
/* Code panels sit slightly lighter than the page; the token colors come
|
||||
|
||||
@@ -160,7 +160,7 @@ main::before {
|
||||
top edge and a grassy shadow. */
|
||||
#sidebar {
|
||||
background: linear-gradient(160deg, #f4faddd9, #d9eec5cf);
|
||||
border-right: 1px solid #ffffff80;
|
||||
border-inline-end: 1px solid #ffffff80;
|
||||
border-bottom: 1px solid var(--line);
|
||||
box-shadow: 0 0.3rem 1rem #4f913b1f;
|
||||
border-radius: 1rem;
|
||||
@@ -225,7 +225,7 @@ article ul ul ul li::before {
|
||||
/* Quotes get a grassy edge and a wash of sunlight. */
|
||||
blockquote {
|
||||
color: #4d6849;
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
background: linear-gradient(90deg, #fff0a238, transparent 70%);
|
||||
padding-top: 0.25rem;
|
||||
padding-bottom: 0.25rem;
|
||||
|
||||
@@ -0,0 +1,489 @@
|
||||
"""Visit analytics: collection sockets, geoip enrichment, favicon fetch.
|
||||
|
||||
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
|
||||
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
|
||||
enriched in background tasks with reverse DNS (cached PTR lookups) and the
|
||||
DB-IP city MMDB (``GeoIP``, decompressed into RAM and opened once at
|
||||
startup);
|
||||
external referrers get their favicon fetched and stored content-hashed.
|
||||
Snapshot broadcasts to connected admin sockets are debounced.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import gzip
|
||||
import io
|
||||
import ipaddress
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import socket
|
||||
from datetime import date
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
from urllib.parse import urlparse
|
||||
|
||||
import httpx
|
||||
import msgspec
|
||||
from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
|
||||
from fastapi.responses import Response
|
||||
from uarite import uaparse
|
||||
|
||||
from pagerite import analytics
|
||||
from pagerite.data import resolve
|
||||
from pagerite.files import _hash_name, file_store
|
||||
from pagerite.state import SITE_URL, _html_response, analytics_store, data
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# httpx logs every request at INFO (e.g. the favicon fetches below); our own
|
||||
# one-line summary in _schedule_favicon_fetch replaces that noise.
|
||||
logging.getLogger("httpx").setLevel(logging.WARNING)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
# Live WebSocket clients for the analytics stream.
|
||||
_analytics_ws_clients: set[WebSocket] = set()
|
||||
_analytics_broadcast_task: asyncio.Task | None = None
|
||||
|
||||
|
||||
# DB-IP databases persist in the working directory (one download serves all
|
||||
# sites run from it). Not the package directory: reinstalls/upgrades wipe it.
|
||||
_DBIP_DIR = Path.cwd()
|
||||
|
||||
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
|
||||
|
||||
|
||||
def _download_dbip() -> None:
|
||||
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
|
||||
today = date.today()
|
||||
months = [f"{today:%Y-%m}"]
|
||||
# The current month's file may not be published yet; fall back to last month.
|
||||
prev = (today.replace(day=1) - date.resolution).replace(day=1)
|
||||
months.append(f"{prev:%Y-%m}")
|
||||
|
||||
existing = sorted(
|
||||
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
|
||||
for p in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*")
|
||||
)
|
||||
if existing and existing[-1] >= months[0]:
|
||||
logger.info("DB-IP database is current (%s), skipping download", existing[-1])
|
||||
return
|
||||
|
||||
for month in months:
|
||||
url = DBIP_URL.format(month=month)
|
||||
target = _DBIP_DIR / f"dbip-city-lite-{month}.mmdb.gz"
|
||||
tmp = target.with_suffix(".mmdb.gz.tmp")
|
||||
logger.info("Downloading %s", url)
|
||||
try:
|
||||
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
|
||||
if r.status_code == 404:
|
||||
continue
|
||||
r.raise_for_status()
|
||||
with open(tmp, "wb") as f:
|
||||
for chunk in r.iter_bytes():
|
||||
f.write(chunk)
|
||||
except httpx.HTTPError as e:
|
||||
logger.warning("DB-IP download failed: %s", e)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
# Verify it is actually gzip data before installing it.
|
||||
try:
|
||||
with gzip.open(tmp, "rb") as f:
|
||||
f.read(1)
|
||||
except OSError:
|
||||
logger.warning("DB-IP download for %s was not valid gzip", month)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
os.replace(tmp, target)
|
||||
# Drop older databases so the app never picks up a stale one.
|
||||
for old in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*"):
|
||||
if old.name != target.name:
|
||||
old.unlink()
|
||||
logger.info("DB-IP database updated to %s", target.name)
|
||||
return
|
||||
logger.warning("Could not download a DB-IP database")
|
||||
|
||||
|
||||
def _geoip_db_path() -> Path | None:
|
||||
"""Find a DB-IP MMDB in the working directory: the ``.mmdb.gz`` download
|
||||
is canonical (decompressed into RAM at open); a plain ``.mmdb`` left over
|
||||
from older versions is still usable, and removed once the matching ``.gz``
|
||||
is present so it does not linger on disk. Returns None if none is present.
|
||||
"""
|
||||
gz = sorted(_DBIP_DIR.glob("dbip-*.mmdb.gz"))
|
||||
if gz:
|
||||
for stale in _DBIP_DIR.glob("dbip-*.mmdb"):
|
||||
stale.unlink()
|
||||
return gz[0]
|
||||
mmdb = sorted(_DBIP_DIR.glob("dbip-*.mmdb"))
|
||||
if mmdb:
|
||||
return mmdb[0]
|
||||
return None
|
||||
|
||||
|
||||
class GeoIP:
|
||||
"""Lazy DB-IP MMDB reader. Call ``_load()`` once at startup before
|
||||
concurrent requests arrive; ``country()`` is read-only and safe to call
|
||||
from ``asyncio.to_thread`` workers afterwards.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._reader: object | None = None
|
||||
|
||||
def _load(self) -> None:
|
||||
if self._reader is not None:
|
||||
return
|
||||
source = _geoip_db_path()
|
||||
if source is None:
|
||||
return
|
||||
try:
|
||||
import maxminddb
|
||||
|
||||
if source.suffix == ".gz":
|
||||
# Only the .gz is kept on disk; the database is decompressed
|
||||
# into RAM (MODE_FD makes the pure-Python Reader .read() the
|
||||
# buffer — never mmap — and bypasses the C extension).
|
||||
buf = io.BytesIO(gzip.decompress(source.read_bytes()))
|
||||
self._reader = maxminddb.open_database(buf, maxminddb.MODE_FD)
|
||||
else:
|
||||
self._reader = maxminddb.open_database(str(source))
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
def country(self, ip: str) -> str:
|
||||
"""Two-letter ISO country code for ``ip``, or "" when unavailable."""
|
||||
if not ip or self._reader is None:
|
||||
return ""
|
||||
try:
|
||||
rec = self._reader.get(ip)
|
||||
if rec:
|
||||
return (rec.get("country") or {}).get("iso_code", "")
|
||||
except Exception:
|
||||
pass
|
||||
return ""
|
||||
|
||||
def city(self, ip: str) -> str:
|
||||
"""City name for ``ip``, or "" when unavailable.
|
||||
|
||||
GeoIP sometimes appends district names in parentheses (e.g.
|
||||
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
|
||||
the value is stored.
|
||||
"""
|
||||
if not ip or self._reader is None:
|
||||
return ""
|
||||
try:
|
||||
rec = self._reader.get(ip)
|
||||
if rec:
|
||||
city = (rec.get("city") or {}).get("names", {}).get("en", "")
|
||||
if city:
|
||||
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
|
||||
return city
|
||||
except Exception:
|
||||
pass
|
||||
return ""
|
||||
|
||||
|
||||
_geoip = GeoIP()
|
||||
|
||||
|
||||
def _client_ip(request: Request | WebSocket) -> str:
|
||||
"""Client IP: first X-Forwarded-For hop (we sit behind a proxy), else
|
||||
the direct peer."""
|
||||
forwarded = request.headers.get("x-forwarded-for", "").split(",")[0].strip()
|
||||
return forwarded or (request.client.host if request.client else "")
|
||||
|
||||
|
||||
def _query_suffix(request: Request) -> str:
|
||||
"""The request's query string as a "?..." suffix, or "" when absent."""
|
||||
query = str(request.url.query)
|
||||
return f"?{query}" if query else ""
|
||||
|
||||
|
||||
@lru_cache(maxsize=4096)
|
||||
def _cached_ptr(ip: str) -> str:
|
||||
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
|
||||
if not ip:
|
||||
return ""
|
||||
try:
|
||||
addr = ipaddress.ip_address(ip)
|
||||
except ValueError:
|
||||
return ""
|
||||
if (
|
||||
addr.is_private
|
||||
or addr.is_loopback
|
||||
or addr.is_reserved
|
||||
or addr.is_multicast
|
||||
or addr.is_link_local
|
||||
):
|
||||
return ""
|
||||
try:
|
||||
host, _, _ = socket.gethostbyaddr(ip)
|
||||
except socket.herror:
|
||||
return ""
|
||||
return host
|
||||
|
||||
|
||||
async def _lookup_host(ip: str) -> str:
|
||||
"""Async wrapper around ``_cached_ptr``; runs the blocking lookup in a thread."""
|
||||
return await asyncio.to_thread(_cached_ptr, ip)
|
||||
|
||||
|
||||
async def _geoip_country(ip: str) -> str:
|
||||
"""Async wrapper around the DB-IP MMDB lookup."""
|
||||
return await asyncio.to_thread(_geoip.country, ip)
|
||||
|
||||
|
||||
async def _geoip_city(ip: str) -> str:
|
||||
"""Async wrapper around the DB-IP MMDB city lookup."""
|
||||
return await asyncio.to_thread(_geoip.city, ip)
|
||||
|
||||
|
||||
async def _enrich_client(client_hash: bytes) -> None:
|
||||
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
|
||||
client = analytics_store.data.clients.get(client_hash)
|
||||
if not client or not client.ip:
|
||||
return
|
||||
host = await _lookup_host(client.ip)
|
||||
country = await _geoip_country(client.ip)
|
||||
city = await _geoip_city(client.ip)
|
||||
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
|
||||
|
||||
|
||||
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
|
||||
"""Start background host/geoip enrichment for the given client hashes."""
|
||||
for client_hash in client_hashes:
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
|
||||
|
||||
#: Icon MIME -> file extension for the stored favicon name. The extension
|
||||
#: reflects the actual content, not the /favicon.ico request path.
|
||||
_FAVICON_EXT = {
|
||||
"image/x-icon": ".ico",
|
||||
"image/vnd.microsoft.icon": ".ico",
|
||||
"image/png": ".png",
|
||||
"image/gif": ".gif",
|
||||
"image/jpeg": ".jpg",
|
||||
"image/webp": ".webp",
|
||||
"image/avif": ".avif",
|
||||
"image/svg+xml": ".svg",
|
||||
}
|
||||
|
||||
_FAVICON_MAX_BYTES = 65536
|
||||
|
||||
#: Origins with a fetch task currently in flight.
|
||||
_favicon_in_flight: set[str] = set()
|
||||
|
||||
|
||||
async def _fetch_favicon(origin: str) -> None:
|
||||
"""Fetch ``{origin}/favicon.ico`` and store it content-hashed on disk.
|
||||
|
||||
The result (icon file name, or "" for a miss) is recorded in the
|
||||
analytics store; misses are retried after analytics._FAVICON_RETRY.
|
||||
Never raises: analytics must not break page serving.
|
||||
"""
|
||||
try:
|
||||
async with httpx.AsyncClient(follow_redirects=True, timeout=8) as client:
|
||||
r = await client.get(f"{origin}/favicon.ico")
|
||||
body = r.content
|
||||
if (
|
||||
not (200 <= r.status_code < 300)
|
||||
or not body
|
||||
or len(body) > _FAVICON_MAX_BYTES
|
||||
):
|
||||
analytics_store.record_favicon(origin)
|
||||
return
|
||||
mime = r.headers.get("content-type", "").split(";")[0].strip().lower()
|
||||
if not mime.startswith("image/"):
|
||||
# Served without an image type: sniff SVG, else assume ICO.
|
||||
if b"<svg" in body[:1024]:
|
||||
mime = "image/svg+xml"
|
||||
elif mime in ("", "application/octet-stream", "text/plain"):
|
||||
mime = "image/x-icon"
|
||||
else:
|
||||
analytics_store.record_favicon(origin)
|
||||
return
|
||||
ext = _FAVICON_EXT.get(mime, ".ico")
|
||||
name = _hash_name(body, f"favicon{ext}")
|
||||
file_store.put(name, body)
|
||||
analytics_store.record_favicon(origin, name)
|
||||
except httpx.HTTPError, OSError:
|
||||
analytics_store.record_favicon(origin)
|
||||
finally:
|
||||
_favicon_in_flight.discard(origin)
|
||||
|
||||
|
||||
def _schedule_favicon_fetch() -> None:
|
||||
"""Start background favicon fetches for origins that need one."""
|
||||
origins = [
|
||||
origin
|
||||
for origin in analytics_store.favicon_origins_needed()
|
||||
if origin not in _favicon_in_flight
|
||||
]
|
||||
if not origins:
|
||||
return
|
||||
logger.info(
|
||||
"Fetching favicons: %s",
|
||||
", ".join(o.removeprefix("https://") for o in origins),
|
||||
)
|
||||
for origin in origins:
|
||||
_favicon_in_flight.add(origin)
|
||||
asyncio.create_task(_fetch_favicon(origin))
|
||||
|
||||
|
||||
async def _broadcast_analytics() -> None:
|
||||
"""Send the current analytics snapshot to every connected WS client."""
|
||||
if not _analytics_ws_clients:
|
||||
return
|
||||
payload = analytics_store.display_json(_in_menu)
|
||||
closed = set()
|
||||
for ws in _analytics_ws_clients:
|
||||
try:
|
||||
await ws.send_text(payload)
|
||||
except Exception:
|
||||
closed.add(ws)
|
||||
for ws in closed:
|
||||
_analytics_ws_clients.discard(ws)
|
||||
|
||||
|
||||
async def _debounced_analytics_broadcast() -> None:
|
||||
"""Wait briefly, then broadcast the latest snapshot once."""
|
||||
await asyncio.sleep(0.2)
|
||||
await _broadcast_analytics()
|
||||
|
||||
|
||||
def _schedule_analytics_broadcast() -> None:
|
||||
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
|
||||
global _analytics_broadcast_task
|
||||
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
|
||||
return
|
||||
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
|
||||
_debounced_analytics_broadcast()
|
||||
)
|
||||
|
||||
|
||||
def _in_menu(path: str) -> bool:
|
||||
"""True when ``path`` ("/a/b" or "/") resolves to a real menu node.
|
||||
|
||||
Category placeholders return 404 but are real nodes: their GETs must not
|
||||
count as misses in the display-time abuse classification.
|
||||
"""
|
||||
return resolve(data.menu, path.strip("/")) is not None
|
||||
|
||||
|
||||
def _record_get(request: Request, *, status: int = 200) -> None:
|
||||
"""Record the document GET as one raw access-log line in analytics.
|
||||
|
||||
Nothing is classified here — the true HTTP status, the full request path
|
||||
(query included), an external referer origin and the preload flag are
|
||||
stored, and visitor/crawler/abuse classification happens at display time
|
||||
(see analytics.Store.display). Idle-time preloads from pagerite.js
|
||||
(``x-pagerite-preload`` header) are recorded with ``pre=True``: never
|
||||
counted, but a navigation later served from the in-memory page cache is
|
||||
attributed this GET's status.
|
||||
|
||||
The devserver's health probe (``GET /?from=devserver.py`` from
|
||||
``127.0.0.1``) is ignored: it is not real traffic. The root-path and
|
||||
localhost checks prevent remote visitors from forging the same query.
|
||||
"""
|
||||
if (
|
||||
request.url.path == "/"
|
||||
and str(request.url.query) == "from=devserver.py"
|
||||
and _client_ip(request) == "127.0.0.1"
|
||||
):
|
||||
return
|
||||
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
|
||||
referer = request.headers.get("referer", "")
|
||||
if analytics._origin(referer) in (None, own_origin):
|
||||
referer = ""
|
||||
client_hash = analytics_store.record_get(
|
||||
_client_ip(request),
|
||||
request.headers.get("user-agent", ""),
|
||||
f"{request.url.path}{_query_suffix(request)}",
|
||||
status=status,
|
||||
referer=referer,
|
||||
accept_language=request.headers.get("accept-language", ""),
|
||||
pre=bool(request.headers.get("x-pagerite-preload")),
|
||||
)
|
||||
if client_hash is not None:
|
||||
_schedule_client_enrichment([client_hash])
|
||||
|
||||
|
||||
@router.get("/_a", response_model=None)
|
||||
async def analytics_page(request: Request) -> Response:
|
||||
"""Render the analytics viewer as a normal site page at /_a.
|
||||
|
||||
The page itself is public, but the data stream (/_api/ws/analytics) stays
|
||||
admin-gated like the rest of /_api, so only authorized users see the
|
||||
statistics; others get the viewer with a "could not be loaded" message.
|
||||
"""
|
||||
return _html_response(
|
||||
request,
|
||||
"analytics",
|
||||
"",
|
||||
headers={"cache-control": "no-cache"},
|
||||
etag=True,
|
||||
)
|
||||
|
||||
|
||||
@router.websocket("/_ws")
|
||||
async def activity_ws(ws: WebSocket) -> None:
|
||||
"""Collect visitor activity: navigations and reading-time updates.
|
||||
|
||||
Public, like the pages themselves (only /_api is gated); one connection
|
||||
follows a browsing session. Messages are ``analytics.Ping`` structs as
|
||||
JSON text frames; ``to`` set is a navigation, ``read`` alone a
|
||||
reading-time update. Everything is recorded raw — known bot UAs and
|
||||
abusive IPs are filtered at display time, not here. The reverse-DNS and
|
||||
DB-IP geoip lookups happen in background tasks so message handling is
|
||||
never delayed by slow DNS or the first MMDB decompress.
|
||||
"""
|
||||
ip = _client_ip(ws)
|
||||
ua = ws.headers.get("user-agent", "")
|
||||
accept_language = ws.headers.get("accept-language", "")
|
||||
# Identify the visitor on the access-log open/close lines (the IP is
|
||||
# already printed there): compact UA plus the browser's language tag.
|
||||
lang, _country = analytics._parse_accept_language(accept_language)
|
||||
ws.scope.setdefault("state", {})["log_extra"] = " ".join(
|
||||
part for part in (uaparse(ua).pretty, lang) if part
|
||||
)
|
||||
await ws.accept()
|
||||
try:
|
||||
while True:
|
||||
text = await ws.receive_text()
|
||||
try:
|
||||
msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
|
||||
except msgspec.DecodeError:
|
||||
continue
|
||||
new_client = analytics_store.record_msg(
|
||||
msg.fr,
|
||||
msg.to or None,
|
||||
ip,
|
||||
ua,
|
||||
accept_language,
|
||||
hide=msg.hide,
|
||||
read=msg.read,
|
||||
)
|
||||
if new_client is not None:
|
||||
_schedule_client_enrichment([new_client])
|
||||
_schedule_favicon_fetch()
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
|
||||
|
||||
@router.websocket("/_api/ws/analytics")
|
||||
async def analytics_websocket(ws: WebSocket) -> None:
|
||||
"""Stream the analytics snapshot, then push updates as they happen.
|
||||
|
||||
Admin-only via the /_api forward-auth gate, like every management
|
||||
endpoint. Powers the analytics viewer rendered at /_a.
|
||||
"""
|
||||
await ws.accept()
|
||||
await ws.send_text(analytics_store.display_json(_in_menu))
|
||||
_analytics_ws_clients.add(ws)
|
||||
try:
|
||||
while True:
|
||||
await ws.receive_text()
|
||||
except Exception:
|
||||
pass
|
||||
finally:
|
||||
_analytics_ws_clients.discard(ws)
|
||||
@@ -0,0 +1,417 @@
|
||||
"""Translator service protocol, dispatcher and its transport-independent core.
|
||||
|
||||
The external machine-translation service connects over WebSocket
|
||||
(``/_translate/<key>``, the route itself is in api.py) and exchanges JSON
|
||||
frames decoded into the tagged msgspec structs below (``bytes`` fields ride
|
||||
as base64 — no manual encoding anywhere). This module holds everything
|
||||
else: the message structs, the connected-client dispatcher (``Dispatcher``
|
||||
— one job at a time per connection, wanted ∩ capable language matching,
|
||||
requeue on disconnect), which fragments are pending for a language
|
||||
(``pending_items``) and storing a result (``store_results``).
|
||||
|
||||
Fragments cross the wire as **prose segments**: the model only ever
|
||||
receives plain text runs (Job.texts) plus per-segment context surrounds
|
||||
(Job.contexts) and returns their translations (Result.texts, same order);
|
||||
markup never leaves the server — reassembly is offset splicing
|
||||
(``pagerite/segments.py``).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
|
||||
import msgspec
|
||||
from fastapi import WebSocket, WebSocketDisconnect
|
||||
from kanta import Kanta
|
||||
|
||||
from pagerite import i18n
|
||||
from pagerite.chunks import chunk_key, needs_translation
|
||||
from pagerite.data import Data, Node, sorted_nodes
|
||||
from pagerite.segments import Span, join, split
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class Hello(msgspec.Struct, tag="hello"):
|
||||
"""Client greeting on connect: the language codes its model CAN produce
|
||||
(capabilities). The server offers jobs only in the intersection with
|
||||
the wanted target languages (``Data.translate_langs``)."""
|
||||
|
||||
langs: list[str]
|
||||
|
||||
|
||||
class TransItem(msgspec.Struct):
|
||||
"""One fragment to translate: original Markdown (or a node title)."""
|
||||
|
||||
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
|
||||
text: str
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Title jobs only: the article's opening prose, so the model sees the
|
||||
#: title as a heading in context, not a lone sentence.
|
||||
context: str = ""
|
||||
|
||||
|
||||
class Job(msgspec.Struct, tag="job"):
|
||||
"""Server push: ONE fragment to translate.
|
||||
|
||||
Exactly one job is in flight per connection — the next is sent only
|
||||
after this one's Result. Clients wanting parallelism open multiple
|
||||
connections."""
|
||||
|
||||
lang: str
|
||||
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
|
||||
#: The fragment's prose segments (pagerite/segments.py): plain text
|
||||
#: runs only — no markup, URLs, code or placeholders ever cross the
|
||||
#: wire. Translate each element independently.
|
||||
texts: list[str]
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Per segment (parallel to texts; "" = none): the surround to
|
||||
#: translate it in — a carved-out segment (link text, partial run)
|
||||
#: carries its block's plain text, a title the article's opening.
|
||||
#: Reference client behavior (scripts/translator.py): translate
|
||||
#: segment+context together, keep the segment's part (its own line /
|
||||
#: paragraph); fall back to the segment alone when the output holds no
|
||||
#: separator. Contexts are not part of the result.
|
||||
contexts: list[str] = msgspec.field(default_factory=list)
|
||||
|
||||
|
||||
class TransResult(msgspec.Struct):
|
||||
"""One translated fragment (storage level, see store_results)."""
|
||||
|
||||
key: bytes
|
||||
text: str
|
||||
|
||||
|
||||
class Result(msgspec.Struct, tag="result"):
|
||||
"""Client reply: the translation of the connection's current Job
|
||||
(must match its lang and key exactly)."""
|
||||
|
||||
lang: str
|
||||
key: bytes
|
||||
#: The job's segments, translated, same order and count. Each must be
|
||||
#: pure prose — the server rejects the result otherwise.
|
||||
texts: list[str]
|
||||
|
||||
|
||||
#: Union of the client -> server frames (the "type" tag selects).
|
||||
ClientMsg = Hello | Result
|
||||
|
||||
|
||||
def pending_items(data: Data, lang: str) -> list[TransItem]:
|
||||
"""Fragments of the site still untranslated for ``lang``, deduped by key.
|
||||
|
||||
Every node (published or not, pages and pure category labels alike)
|
||||
contributes its title; pages also contribute each chunk
|
||||
that needs translation (``needs_translation``), is not editor-flagged
|
||||
no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang``
|
||||
yet. Content-addressed text (shared paragraphs, repeated titles) appears
|
||||
once, under the first page in menu order that has it.
|
||||
"""
|
||||
items: list[TransItem] = []
|
||||
seen: set[bytes] = set()
|
||||
|
||||
def emit(key: bytes, text: str, path: str, kind: str, context: str = "") -> None:
|
||||
if key in seen or lang in data.trans.get(key, {}):
|
||||
return
|
||||
seen.add(key)
|
||||
items.append(
|
||||
TransItem(key=key, text=text, path=path, kind=kind, context=context)
|
||||
)
|
||||
|
||||
def opening(node: Node) -> str:
|
||||
"""The article's opening prose (first segment, capped): the title
|
||||
job's context — a lone word like "About" reads as a heading on top
|
||||
of an article, not as a sentence. Empty when there's no prose."""
|
||||
for h in node.chunks or ():
|
||||
text = data.chunks.get(h)
|
||||
if text and (segs := split(text)[1]):
|
||||
return segs[0][:400]
|
||||
return ""
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
# An article whose primary language IS the target needs no
|
||||
# translation into it — skip its title and chunks entirely.
|
||||
# Category labels (chunks is None) contribute only their title:
|
||||
# it is their nav-menu label.
|
||||
node_lang = node.language or inherited
|
||||
if node_lang != lang:
|
||||
if node.title:
|
||||
emit(
|
||||
chunk_key(node.title),
|
||||
node.title,
|
||||
path,
|
||||
"title",
|
||||
context=opening(node),
|
||||
)
|
||||
for h in node.chunks or ():
|
||||
text = data.chunks.get(h)
|
||||
if (
|
||||
text is not None
|
||||
and h not in node.no_trans
|
||||
and needs_translation(text)
|
||||
):
|
||||
emit(h, text, path, "chunk")
|
||||
walk(node.children, path, node_lang)
|
||||
|
||||
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
|
||||
return items
|
||||
|
||||
|
||||
def store_results(data: Data, lang: str, items: list[TransResult]) -> list[str]:
|
||||
"""Store machine translations for ``lang``; return the paths of the
|
||||
articles that gained at least one entry.
|
||||
|
||||
Pure data operations: the caller wraps this in a kanta transaction and
|
||||
invalidates pages. Unknown keys are stored anyway (unreferenced hashes
|
||||
are never read, and the content may simply have moved on since the job
|
||||
was pushed); re-storing an existing key overwrites, last wins. Every
|
||||
article that gained an entry gets ``node.langs[lang]`` set (the
|
||||
availability index, docs/migrate.md) — because chunks are
|
||||
content-addressed, that includes pages merely sharing a fragment.
|
||||
"""
|
||||
stored = {item.key for item in items}
|
||||
for item in items:
|
||||
data.trans.setdefault(item.key, {})[lang] = item.text
|
||||
pages: list[str] = []
|
||||
|
||||
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
node_lang = node.language or inherited
|
||||
if node_lang != lang:
|
||||
keys = set(node.chunks or ())
|
||||
if node.title:
|
||||
keys.add(chunk_key(node.title))
|
||||
if keys & stored:
|
||||
node.langs[lang] = True
|
||||
pages.append(path)
|
||||
walk(node.children, path, node_lang)
|
||||
|
||||
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
|
||||
return pages
|
||||
|
||||
|
||||
class _Connection:
|
||||
"""One connected translator socket: the language codes it announced as
|
||||
capabilities (Hello) and the (lang, chunk-key) job currently in flight
|
||||
on it, with the segment spans to splice its Result into
|
||||
(pagerite/segments.py) — one at a time, the next is sent only after its
|
||||
Result.
|
||||
|
||||
Per-connection only: in-flight lives solely here, so on disconnect the
|
||||
item simply becomes pending again and is re-offered to any free capable
|
||||
connection."""
|
||||
|
||||
def __init__(self, capable: set[str]) -> None:
|
||||
self.capable = capable
|
||||
self.inflight: tuple[str, bytes] | None = None
|
||||
#: Source spans of the in-flight job's segments (splice offsets
|
||||
#: and link marks).
|
||||
self.spans: list[Span] = []
|
||||
self.original: str = "" # its full source text (for the splicing)
|
||||
self.kind: str = "" # "chunk" | "title" (for the transaction action)
|
||||
|
||||
|
||||
class Dispatcher:
|
||||
"""The translator dispatcher: connected client sockets and the job
|
||||
pipeline (docs/localization.md).
|
||||
|
||||
One single-item job at a time per connection, offered in the
|
||||
intersection of the wanted languages (``Data.translate_langs``) and the
|
||||
connection's announced capabilities. Pending work is derived from the
|
||||
``trans`` store (``pending_items``) minus the items in flight on any
|
||||
connection, so a dropped connection's in-flight item is simply
|
||||
re-offered. Results are matched to content by chunk key alone. A
|
||||
(lang, key) whose Result fails segment validation is skipped for the
|
||||
rest of the run — generation is near-deterministic, so an immediate
|
||||
retry would just re-fail.
|
||||
"""
|
||||
|
||||
def __init__(self, data: Data, db: Kanta, invalidate) -> None:
|
||||
self.data = data
|
||||
self.db = db
|
||||
#: Sync content-change hook (state._invalidate_pages), called inside
|
||||
#: transactions; schedules the next dispatch pass.
|
||||
self.invalidate = invalidate
|
||||
#: Connected translator sockets and their per-connection state.
|
||||
self.clients: dict[WebSocket, _Connection] = {}
|
||||
#: (lang, chunk key) of fragments whose result failed validation
|
||||
#: (segment count, empty or non-prose segments, segments.py) this run.
|
||||
self.validation_failures: set[tuple[str, bytes]] = set()
|
||||
|
||||
def reset_validation_failures(self) -> None:
|
||||
"""Clear the skip list of fragments rejected this run (segment
|
||||
validation): a translations refresh is precisely the "another
|
||||
chance" for them."""
|
||||
self.validation_failures.clear()
|
||||
|
||||
def schedule(self) -> None:
|
||||
"""Schedule a dispatch pass, if any translator is connected.
|
||||
|
||||
The invalidate hook is sync and called inside transactions: the
|
||||
task first runs once the current coroutine awaits again, i.e. after
|
||||
the transaction has committed. No-op without a running loop (CLI
|
||||
use)."""
|
||||
if not self.clients:
|
||||
return
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return
|
||||
asyncio.create_task(self._dispatch())
|
||||
|
||||
async def _dispatch(self) -> None:
|
||||
"""Offer one pending item to every free capable connection."""
|
||||
wanted = {
|
||||
tag for lang in self.data.translate_langs if (tag := i18n.base_tag(lang))
|
||||
}
|
||||
if not wanted:
|
||||
return
|
||||
for ws, state in list(self.clients.items()):
|
||||
if state.inflight is not None:
|
||||
continue
|
||||
langs = wanted & state.capable
|
||||
if not langs:
|
||||
continue
|
||||
inflight = {s.inflight for s in self.clients.values() if s.inflight}
|
||||
job = None
|
||||
spans: list[Span] = []
|
||||
original = ""
|
||||
# Titles before articles — across languages too, so every menu
|
||||
# is named before any article body is worked on (a page's name
|
||||
# is its most visible string). pending_items emits in menu
|
||||
# order, a page's title before its chunks; filtering by kind
|
||||
# keeps that stable order within each kind.
|
||||
pending = {lang: pending_items(self.data, lang) for lang in sorted(langs)}
|
||||
for kind in ("title", "chunk"):
|
||||
for lang in sorted(langs):
|
||||
for item in pending[lang]:
|
||||
if (
|
||||
item.kind != kind
|
||||
or (lang, item.key) in inflight
|
||||
or (lang, item.key) in self.validation_failures
|
||||
):
|
||||
continue
|
||||
spans, texts, contexts = split(item.text)
|
||||
if not texts:
|
||||
continue # prose that could not be located for splicing
|
||||
original = item.text
|
||||
if item.kind == "title" and item.context:
|
||||
# A title's surround is the article's opening prose
|
||||
# (TransItem.context), not its own one-word block.
|
||||
contexts = [item.context] * len(texts)
|
||||
job = Job(
|
||||
lang=lang,
|
||||
key=item.key,
|
||||
texts=texts,
|
||||
path=item.path,
|
||||
kind=item.kind,
|
||||
contexts=contexts,
|
||||
)
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is None:
|
||||
continue
|
||||
state.inflight = (job.lang, job.key) # before the await: no double-assign
|
||||
state.spans = spans
|
||||
state.original = original
|
||||
state.kind = job.kind
|
||||
try:
|
||||
await ws.send_text(msgspec.json.encode(job).decode())
|
||||
except Exception: # send failed: the receive loop cleans up
|
||||
self.clients.pop(ws, None)
|
||||
|
||||
async def handle_ws(self, ws: WebSocket, clientkey: str) -> None:
|
||||
"""The /_translate/<key> channel (docs/localization.md).
|
||||
|
||||
A wrong/empty key rejects the handshake (closing before accept
|
||||
makes Starlette answer HTTP 403). Protocol (JSON frames): the
|
||||
client opens with Hello(langs) announcing its CAPABILITIES — the
|
||||
language codes its model can produce (normalized to translation
|
||||
tags; "en"/empty dropped) — then answers each Job with its
|
||||
Result(lang, key, texts). A Result without an in-flight job or with
|
||||
a different (lang, key), a duplicate Hello, or any malformed frame
|
||||
closes the socket with a protocol error.
|
||||
"""
|
||||
if clientkey not in self.data.translate_keys:
|
||||
await ws.close(code=1008) # policy violation; pre-accept = HTTP 403
|
||||
return
|
||||
await ws.accept()
|
||||
state: _Connection | None = None
|
||||
try:
|
||||
while True:
|
||||
raw = await ws.receive_text()
|
||||
try:
|
||||
msg = msgspec.json.decode(raw.encode(), type=ClientMsg)
|
||||
except msgspec.DecodeError:
|
||||
await ws.close(code=1002) # protocol error
|
||||
return
|
||||
if isinstance(msg, Hello):
|
||||
if state is not None: # one Hello per connection
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
state = _Connection(
|
||||
{tag for lang in msg.langs if (tag := i18n.base_tag(lang))}
|
||||
)
|
||||
self.clients[ws] = state
|
||||
self.schedule()
|
||||
else: # Result
|
||||
lang = i18n.base_tag(msg.lang)
|
||||
if (
|
||||
state is None # results before Hello
|
||||
or state.inflight is None # no job in flight
|
||||
or (lang, msg.key) != state.inflight # wrong job
|
||||
):
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
texts, spans, original = msg.texts, state.spans, state.original
|
||||
kind, state.kind = state.kind, ""
|
||||
state.inflight = None
|
||||
state.spans = []
|
||||
state.original = ""
|
||||
text = (
|
||||
join(original, spans, texts)
|
||||
if len(texts) == len(spans)
|
||||
else None
|
||||
)
|
||||
if text is None:
|
||||
# The model broke the segment contract (count
|
||||
# mismatch, empty or non-prose segment): drop the
|
||||
# result and skip the fragment for this run (it
|
||||
# stays pending; a restart, a refresh or a model
|
||||
# change gets another chance).
|
||||
self.validation_failures.add((lang, msg.key))
|
||||
logger.warning(
|
||||
"[%s] result for chunk %s rejected: invalid segments",
|
||||
lang,
|
||||
msg.key.hex(),
|
||||
)
|
||||
self.schedule()
|
||||
continue
|
||||
with self.db.transaction(
|
||||
f"translate:{lang}{':title' if kind == 'title' else ''}",
|
||||
user=clientkey,
|
||||
):
|
||||
paths = store_results(
|
||||
self.data, lang, [TransResult(key=msg.key, text=text)]
|
||||
)
|
||||
self.invalidate() # schedules the next dispatch
|
||||
if paths:
|
||||
logger.info(
|
||||
"[%s] now available for %d page(s): %s",
|
||||
lang,
|
||||
len(paths),
|
||||
", ".join(sorted(paths)),
|
||||
)
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
finally:
|
||||
if self.clients.pop(ws, None) is not None:
|
||||
# The in-flight item (if any) is pending again; offer it around.
|
||||
self.schedule()
|
||||
+351
-73
@@ -24,7 +24,10 @@ import re
|
||||
from html5tagger import HTML, Document, E, Template
|
||||
from platformdirs import site_data_dir, user_data_path
|
||||
|
||||
from pagerite.data import Node, prettify, resolve, sorted_nodes
|
||||
from pagerite import i18n
|
||||
from pagerite.config import load as _load_config
|
||||
from pagerite.data import Data, Node, node_markdown, prettify, resolve, sorted_nodes
|
||||
from pagerite.i18n import Translation
|
||||
from pagerite.markdown import render
|
||||
|
||||
SITE_NAME = "Pagerite"
|
||||
@@ -37,10 +40,16 @@ def _data_roots() -> list[Path]:
|
||||
roots = [user_data_path("pagerite", appauthor=False)]
|
||||
# site_data_dir keeps the multipath (site_data_path collapses it to the
|
||||
# first entry, since a Path cannot hold several).
|
||||
roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(os.pathsep)
|
||||
roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(
|
||||
os.pathsep
|
||||
)
|
||||
return [Path(r) for r in roots]
|
||||
|
||||
|
||||
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
|
||||
config = _load_config()
|
||||
|
||||
|
||||
def _theme_dirs() -> list[Path]:
|
||||
"""Theme search roots, most specific first; first match wins per file.
|
||||
|
||||
@@ -53,7 +62,7 @@ def _theme_dirs() -> list[Path]:
|
||||
"""
|
||||
return [
|
||||
Path("themes"),
|
||||
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "themes",
|
||||
Path(config.hostname) / "themes",
|
||||
*(root / "themes" for root in _data_roots()),
|
||||
Path(__file__).parent / "themes",
|
||||
]
|
||||
@@ -69,7 +78,7 @@ THEME_DIRS = _theme_dirs()
|
||||
# built-in --font-* variables.
|
||||
FONT_DIRS = [
|
||||
Path("fonts"),
|
||||
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "fonts",
|
||||
Path(config.hostname) / "fonts",
|
||||
*(root / "fonts" for root in _data_roots()),
|
||||
]
|
||||
|
||||
@@ -102,13 +111,15 @@ def _theme_color_schemes(theme: str) -> set[str]:
|
||||
return set()
|
||||
try:
|
||||
css = path.read_text()
|
||||
except (OSError, ValueError):
|
||||
except OSError, ValueError:
|
||||
return set()
|
||||
css = re.sub(r"/\*.*?\*/", "", css, flags=re.DOTALL)
|
||||
m = re.search(r"color-scheme\s*:\s*([^;]+);", css, re.IGNORECASE)
|
||||
if not m:
|
||||
return set()
|
||||
return {tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}}
|
||||
return {
|
||||
tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}
|
||||
}
|
||||
|
||||
|
||||
def _theme_mode(theme: str) -> str:
|
||||
@@ -252,20 +263,31 @@ def _transition_css_url(transition: str) -> str | None:
|
||||
|
||||
|
||||
def _editor_css_url(vite_url: str | None) -> str | None:
|
||||
"""URL for the editor-specific stylesheet (Vue component styles).
|
||||
"""URLs (comma-joined) for the editor-specific stylesheets (Vue
|
||||
component styles).
|
||||
|
||||
This is linked by the public-page edit pen so the editor styles are
|
||||
loaded before the editor JS dynamic-import resolves.
|
||||
loaded before the editor JS dynamic-import resolves. Component styles
|
||||
can land on shared chunks rather than the entry's own stylesheet —
|
||||
LangSelect's ride on the shared store chunk, as it is also used by the
|
||||
on-demand public language selector — so collect the stylesheets of the
|
||||
entry and its imported chunks (the same traversal _langselect_assets
|
||||
does).
|
||||
"""
|
||||
if vite_url:
|
||||
return None
|
||||
manifest = _manifest()
|
||||
entry = manifest["src/main.js"]
|
||||
base = manifest.get(_BASE_CSS_KEY, {}).get("file")
|
||||
for css in entry.get("css", []):
|
||||
if css != base:
|
||||
return f"/{css}"
|
||||
return None
|
||||
stylesheets, seen = [], set()
|
||||
queue = ["src/main.js"]
|
||||
for key in queue: # grows with imported chunks
|
||||
if key in seen:
|
||||
continue
|
||||
seen.add(key)
|
||||
entry = manifest[key]
|
||||
stylesheets += [f"/{css}" for css in entry.get("css", []) if css != base]
|
||||
queue += entry.get("imports", [])
|
||||
return ",".join(stylesheets) or None
|
||||
|
||||
|
||||
def _inline_asset(url: str) -> str:
|
||||
@@ -309,6 +331,9 @@ def _layout(
|
||||
transition: str = "cube",
|
||||
favicon: str = "",
|
||||
social: dict[str, str] | None = None,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
canonical: str = "",
|
||||
alternates: list[tuple[str, str]] = (),
|
||||
) -> Template:
|
||||
"""Page layout template with standard assets and ES-module scripts.
|
||||
|
||||
@@ -333,21 +358,33 @@ def _layout(
|
||||
|
||||
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
|
||||
property attributes, everything else (description, twitter:*) as name.
|
||||
|
||||
``lang`` is the served language for <html lang>; an RTL language (ar,
|
||||
fa, ...) also puts dir="rtl" on <html> (the editor panel and the
|
||||
analytics dashboard carry their own lang="en" dir="ltr", so they are
|
||||
unaffected). ``canonical`` and
|
||||
``alternates`` ((hreflang, href) pairs) are the page's language URLs
|
||||
(see docs/localization.md), emitted right after the viewport and before
|
||||
the social tags: canonical first, then the hreflang alternates.
|
||||
"""
|
||||
doc = Document(E.Title, lang="en")
|
||||
doc = Document(
|
||||
E.Title, lang=lang, dir="rtl" if lang in i18n.RTL_LANGUAGES else "ltr"
|
||||
)
|
||||
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs
|
||||
# the real device width, not the default 980px layout viewport.
|
||||
doc.meta(name="viewport", content="width=device-width, initial-scale=1")
|
||||
if canonical:
|
||||
doc.link(rel="canonical", href=canonical)
|
||||
for hreflang, href in alternates:
|
||||
doc.link(rel="alternate", hreflang=hreflang, href=href)
|
||||
for key, value in (social or {}).items():
|
||||
if value:
|
||||
if key.startswith(("og:", "article:")):
|
||||
doc.meta(property=key, content=value)
|
||||
elif key == "canonical":
|
||||
doc.link(rel="canonical", href=value)
|
||||
else:
|
||||
doc.meta(name=key, content=value)
|
||||
# A custom favicon (from the site editor) is linked explicitly; without
|
||||
# one, browsers fall back to the build's /favicon.ico by convention.
|
||||
# A custom favicon (from the site editor) is linked explicitly;
|
||||
# /favicon.ico redirects to the same store file for non-HTML contexts.
|
||||
if favicon:
|
||||
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
|
||||
# Asset URLs for the on-demand bundles (editor, analytics) for
|
||||
@@ -359,10 +396,14 @@ def _layout(
|
||||
# carries the on-demand URLs in one JSON script instead.
|
||||
vite_url = os.environ.get("PAGERITE_VITE_URL")
|
||||
editor_scripts, editor_css = _editor_assets()
|
||||
langselect_scripts, langselect_css = _langselect_assets()
|
||||
config = {
|
||||
"pagerite:editor-src": editor_scripts[-1],
|
||||
"pagerite:analytics-src": _analytics_assets()[0][0],
|
||||
"pagerite:langselect-src": langselect_scripts[-1],
|
||||
}
|
||||
if langselect_css:
|
||||
config["pagerite:langselect-css"] = ",".join(langselect_css)
|
||||
if editor_css:
|
||||
config["pagerite:editor-css"] = editor_css
|
||||
if vite_url:
|
||||
@@ -412,8 +453,7 @@ def _layout(
|
||||
if custom_css.strip():
|
||||
doc.style(custom_css, id="pagerite-user")
|
||||
body = (
|
||||
doc
|
||||
.header(
|
||||
doc.header(
|
||||
E.div(E.Banner, id="page-banner"),
|
||||
E.Brand,
|
||||
E.nav(E.Nav, id="nav"),
|
||||
@@ -443,26 +483,52 @@ def _layout(
|
||||
return Template(body)
|
||||
|
||||
|
||||
def _brand_link(brand: str, brand_html: str = "") -> HTML:
|
||||
def _brand_link(brand: str, brand_html: str = "", link_lang: str = "") -> HTML:
|
||||
"""Header brand: custom HTML (in a #brand wrapper, rendered instead of
|
||||
the link) when configured, else the plain brand link; omitted entirely
|
||||
when neither is set."""
|
||||
if brand_html.strip():
|
||||
return HTML(str(E.div(HTML(brand_html), id="brand")))
|
||||
return HTML(str(E.a(brand, href="/", id="brand"))) if brand else HTML("")
|
||||
return (
|
||||
HTML(str(E.a(brand, href=_href("", link_lang), id="brand")))
|
||||
if brand
|
||||
else HTML("")
|
||||
)
|
||||
|
||||
|
||||
def _title(slug: str, node: Node) -> str:
|
||||
"""Menu label: the configured title, prettified slug, "Home" fallback."""
|
||||
def _title(
|
||||
slug: str, node: Node, translation: Translation | None = None, path: str = ""
|
||||
) -> str:
|
||||
"""Menu label: the configured title, prettified slug, "Home" fallback.
|
||||
|
||||
With a translation, its title map (keyed by node path) wins, falling
|
||||
back per node to the original English title.
|
||||
"""
|
||||
if translation and (t := translation.titles.get(path)):
|
||||
return t
|
||||
return node.title or prettify(slug) or "Home"
|
||||
|
||||
|
||||
def _href(path: str, link_lang: str = "") -> str:
|
||||
"""Site-chrome link to a page: when the page was requested with a
|
||||
?lang= override the query is replicated onto the navigation links it
|
||||
renders, so clicks and prefetches (which take the href as-is) stay in
|
||||
the chosen language — even without JS (docs/localization.md)."""
|
||||
return f"/{path}?lang={link_lang}" if link_lang else f"/{path}"
|
||||
|
||||
|
||||
def _nav_link(
|
||||
doc, menu: dict[str, Node], node: Node, path: str, current: str,
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
node: Node,
|
||||
path: str,
|
||||
current: str,
|
||||
ancestors_current: bool = True,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> None:
|
||||
"""Render one <li> linking the node. Category labels (no content of
|
||||
their own — None, or empty markdown as left by the site editor's
|
||||
their own — chunks None, or an empty page as left by the site editor's
|
||||
page creation) link straight to their first child page, so normal
|
||||
navigation bypasses the placeholder/empty page at their own URL."""
|
||||
# The navbar highlights a top-level item also when viewing any of its
|
||||
@@ -470,17 +536,22 @@ def _nav_link(
|
||||
is_current = current == path or (
|
||||
ancestors_current and path and current.startswith(f"{path}/")
|
||||
)
|
||||
href = f"/{path}"
|
||||
if not node.content and (leaf := first_leaf(menu, path)) is not None:
|
||||
href = f"/{leaf}"
|
||||
href = _href(path, link_lang)
|
||||
if not node.chunks and (leaf := first_leaf(menu, path)) is not None:
|
||||
href = _href(leaf, link_lang)
|
||||
doc.li.a(
|
||||
_title(path.rpartition("/")[2], node),
|
||||
_title(path.rpartition("/")[2], node, translation, path),
|
||||
href=href,
|
||||
**{"class": "current"} if is_current else {},
|
||||
)
|
||||
|
||||
|
||||
def nav_html(menu: dict[str, Node], current: str) -> HTML:
|
||||
def nav_html(
|
||||
menu: dict[str, Node],
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the contents of the #nav element for the current path.
|
||||
|
||||
Top-level items in menu order; the front page (slug "", href "/")
|
||||
@@ -491,11 +562,24 @@ def nav_html(menu: dict[str, Node], current: str) -> HTML:
|
||||
with nav:
|
||||
for slug, node in sorted_nodes(menu):
|
||||
if node.published:
|
||||
_nav_link(nav, menu, node, slug, current)
|
||||
_nav_link(
|
||||
nav,
|
||||
menu,
|
||||
node,
|
||||
slug,
|
||||
current,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
return HTML(str(nav))
|
||||
|
||||
|
||||
def sidebar_html(menu: dict[str, Node], current: str) -> HTML:
|
||||
def sidebar_html(
|
||||
menu: dict[str, Node],
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the #sidebar element for the current path (empty when none).
|
||||
|
||||
The sidebar is the current main level section's sub-navigation: the
|
||||
@@ -529,24 +613,45 @@ def sidebar_html(menu: dict[str, Node], current: str) -> HTML:
|
||||
nav = E.ul
|
||||
with nav:
|
||||
for slug, child in items:
|
||||
_sidebar_item(nav, menu, child, f"{section}/{slug}", current)
|
||||
_sidebar_item(
|
||||
nav, menu, child, f"{section}/{slug}", current, translation, link_lang
|
||||
)
|
||||
return HTML(str(E.aside(nav, id="sidebar")))
|
||||
|
||||
|
||||
def _sidebar_item(doc, menu: dict[str, Node], node: Node, path: str, current: str) -> None:
|
||||
def _sidebar_item(
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
node: Node,
|
||||
path: str,
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> None:
|
||||
"""One sidebar <li>: the node link, with its published children as a
|
||||
nested list (third level and deeper, recursively)."""
|
||||
_nav_link(doc, menu, node, path, current, ancestors_current=False)
|
||||
_nav_link(
|
||||
doc,
|
||||
menu,
|
||||
node,
|
||||
path,
|
||||
current,
|
||||
ancestors_current=False,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
sub = [(s, c) for s, c in sorted_nodes(node.children) if c.published]
|
||||
if sub:
|
||||
# doc.li.a(...) above left the <li> open for nesting.
|
||||
with doc.ul:
|
||||
for slug, child in sub:
|
||||
_sidebar_item(doc, menu, child, f"{path}/{slug}", current)
|
||||
_sidebar_item(
|
||||
doc, menu, child, f"{path}/{slug}", current, translation, link_lang
|
||||
)
|
||||
|
||||
|
||||
def first_leaf(menu: dict[str, Node], path: str) -> str | None:
|
||||
"""First published descendant page (content set) in menu order.
|
||||
"""First published descendant page (chunks set) in menu order.
|
||||
|
||||
This is the nav-link target for content-less category labels.
|
||||
"""
|
||||
@@ -561,7 +666,7 @@ def _first_leaf(node: Node, path: str) -> str | None:
|
||||
if not child.published:
|
||||
continue
|
||||
cpath = f"{path}/{slug}" if path else slug
|
||||
if child.content:
|
||||
if child.chunks:
|
||||
return cpath
|
||||
if (leaf := _first_leaf(child, cpath)) is not None:
|
||||
return leaf
|
||||
@@ -674,16 +779,42 @@ def banner_source(menu: dict[str, Node], path: str) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
def page_content(menu: dict[str, Node], path: str) -> HTML:
|
||||
def page_content(
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the contents of the #main element for a page.
|
||||
|
||||
A page with published children (a category page) lists them as cards
|
||||
after the markdown content.
|
||||
after the markdown content. With a translation, its Markdown goes
|
||||
through the same render pipeline; missing pieces (markdown=None, absent
|
||||
title entries) fall back to the original. ``lang`` feeds the cards'
|
||||
per-target localization.
|
||||
"""
|
||||
node = resolve(menu, path)[-1]
|
||||
content = node_markdown(data, node) or ""
|
||||
title = node.title
|
||||
# The original text pins the section anchors: on a translated page the
|
||||
# heading slugs (and thus #hash URLs) stay in the original language.
|
||||
anchors_from = None
|
||||
if translation:
|
||||
anchors_from = (content, title)
|
||||
if translation.markdown is not None:
|
||||
content = translation.markdown
|
||||
title = (
|
||||
_title(path.rpartition("/")[2], node, translation, path)
|
||||
if node.title
|
||||
else title
|
||||
)
|
||||
# The title is injected into the markdown (as # title when it has no
|
||||
# h1 of its own), so title and content render as one article.
|
||||
rendered = render(node.content or "", path, node.created, node.modified, title=node.title)
|
||||
rendered = render(
|
||||
content, path, node.created, node.modified, title=title, anchors_from=anchors_from
|
||||
)
|
||||
# Long articles get .multicol: the article column cap lifts (see the
|
||||
# #content grid in pagerite.css) and the .cols segments lay out in at
|
||||
# most two columns. The html is already segmented by render() — the
|
||||
@@ -691,11 +822,20 @@ def page_content(menu: dict[str, Node], path: str) -> HTML:
|
||||
doc = E.article(class_="multicol") if rendered.multicol else E.article
|
||||
with doc:
|
||||
doc(HTML(rendered.html))
|
||||
_cards(doc, menu, node, path)
|
||||
_cards(doc, menu, data, node, path, translation, link_lang, lang)
|
||||
return HTML(str(doc))
|
||||
|
||||
|
||||
def _cards(doc, menu: dict[str, Node], node: Node, path: str) -> None:
|
||||
def _cards(
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
node: Node,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> None:
|
||||
"""Card stacks of the node's published children (nothing when childless).
|
||||
|
||||
One column per direct child, all in a single full-width row (the .wide
|
||||
@@ -721,35 +861,54 @@ def _cards(doc, menu: dict[str, Node], node: Node, path: str) -> None:
|
||||
continue
|
||||
with doc.div(class_="stack"):
|
||||
for epath, enode in entries:
|
||||
_card(doc, enode, epath)
|
||||
_card(doc, data, enode, epath, translation, link_lang, lang)
|
||||
|
||||
|
||||
def _walk(node: Node, path: str):
|
||||
"""Published content pages of a subtree, pre-order in menu order: the
|
||||
node itself first when it has content (the stack's landing card), then
|
||||
its descendants (content-less nodes contribute only their subtree)."""
|
||||
if node.content:
|
||||
if node.chunks:
|
||||
yield path, node
|
||||
for slug, child in sorted_nodes(node.children):
|
||||
if child.published:
|
||||
yield from _walk(child, f"{path}/{slug}")
|
||||
|
||||
|
||||
def _card(doc, node: Node, path: str) -> None:
|
||||
def _card(
|
||||
doc,
|
||||
data: Data,
|
||||
node: Node,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> None:
|
||||
"""One card in a stack: cover + title, plus the description when the
|
||||
page has no image (its card shows a gradient cover instead)."""
|
||||
page has no image (its card shows a gradient cover instead).
|
||||
|
||||
The card text localizes per target article where that page is
|
||||
available in the language: the title comes from the translation's
|
||||
title map and the cover/description heuristics run on the target's
|
||||
hybrid Markdown — with per-card fallback to the original otherwise.
|
||||
"""
|
||||
image = description = ""
|
||||
if node.content:
|
||||
html = render(node.content, path, node.created, node.modified).html
|
||||
if node.chunks:
|
||||
md = node_markdown(data, node) or ""
|
||||
if lang and lang in node.langs:
|
||||
md = i18n.hybrid_markdown(data, node, path, lang)
|
||||
html = render(md, path, node.created, node.modified).html
|
||||
image, _ = _media(html)
|
||||
if not image:
|
||||
description = _description(html, 150)
|
||||
with doc.a(href=f"/{path}", class_="card"):
|
||||
with doc.a(href=_href(path, link_lang), class_="card"):
|
||||
if image:
|
||||
doc.span(class_="cover", style=f'background-image: url("{image}")')
|
||||
else:
|
||||
doc.span(class_="cover")
|
||||
doc.span(_title(path.rpartition("/")[2], node), class_="title")
|
||||
doc.span(
|
||||
_title(path.rpartition("/")[2], node, translation, path), class_="title"
|
||||
)
|
||||
if description:
|
||||
doc.span(description, class_="desc")
|
||||
|
||||
@@ -776,7 +935,9 @@ def _description(html: str, limit: int = 200) -> str:
|
||||
return text
|
||||
# Prefer a clean cut: the last sentence ending within the limit, as
|
||||
# long as it does not reduce the description to a tiny fragment.
|
||||
if (end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)) > limit // 2:
|
||||
if (
|
||||
end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)
|
||||
) > limit // 2:
|
||||
return text[:end]
|
||||
return text[:limit].rsplit(" ", 1)[0] + "…"
|
||||
|
||||
@@ -832,7 +993,12 @@ def _share_media(html: str, base_url: str) -> tuple[str, str]:
|
||||
|
||||
|
||||
def _social_meta(
|
||||
node: Node, path: str, title: str, html: str, brand: str, base_url: str,
|
||||
node: Node,
|
||||
path: str,
|
||||
title: str,
|
||||
html: str,
|
||||
brand: str,
|
||||
base_url: str,
|
||||
) -> dict[str, str]:
|
||||
"""Open Graph/Twitter/SEO meta tags for a content page.
|
||||
|
||||
@@ -849,12 +1015,9 @@ def _social_meta(
|
||||
url = f"{base_url}/{path}" if base_url else ""
|
||||
text = _description(html)
|
||||
image, video = _share_media(html, base_url)
|
||||
twitter_image = (
|
||||
re.sub(r"(/_f/[0-9a-f]{12})$", r"\1.webp", image) if image else ""
|
||||
)
|
||||
twitter_image = re.sub(r"(/_f/[0-9a-f]{12})$", r"\1.webp", image) if image else ""
|
||||
return {
|
||||
"description": text,
|
||||
"canonical": url,
|
||||
"og:type": "article",
|
||||
"og:title": title,
|
||||
"og:description": text,
|
||||
@@ -869,8 +1032,45 @@ def _social_meta(
|
||||
}
|
||||
|
||||
|
||||
def _language_urls(
|
||||
data: Data,
|
||||
path: str,
|
||||
node: Node,
|
||||
lang: str,
|
||||
original: str,
|
||||
base_url: str,
|
||||
) -> tuple[str, list[tuple[str, str]]]:
|
||||
"""(canonical, hreflang alternates) for a page (docs/localization.md).
|
||||
|
||||
The canonical names the actually served language — the plain URL for
|
||||
the original (for SEO the non-query URL means the article's language),
|
||||
?lang= for a translation — regardless of how the language was arrived
|
||||
at (query or header). The alternates list the languages the page is
|
||||
actually available in (``node.langs``; a category label's title counts
|
||||
as its content): x-default first (the plain, autodetecting URL), then
|
||||
every available language — the original again by its plain URL,
|
||||
translations by ?lang=. The public language selector keys off these.
|
||||
("", []) without a base_url.
|
||||
"""
|
||||
if not base_url:
|
||||
return "", []
|
||||
url = f"{base_url}/{path}"
|
||||
canonical = url if lang == original else f"{url}?lang={lang}"
|
||||
alternates = []
|
||||
if data.translate_langs:
|
||||
# Only languages the page actually has AND that are still enabled
|
||||
# site-wide (a disabled target stops being advertised).
|
||||
enabled = {original, *data.translate_langs}
|
||||
alternates = [("x-default", url)] + [
|
||||
(tag, url if tag == original else f"{url}?lang={tag}")
|
||||
for tag in sorted({original, *node.langs} & enabled)
|
||||
]
|
||||
return canonical, alternates
|
||||
|
||||
|
||||
def render_page(
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
path: str,
|
||||
brand: str = SITE_NAME,
|
||||
custom_css: str = "",
|
||||
@@ -879,21 +1079,42 @@ def render_page(
|
||||
brand_html: str = "",
|
||||
base_url: str = "",
|
||||
transition: str = "cube",
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> str:
|
||||
"""Render a full HTML page for the slug path."""
|
||||
"""Render a full HTML page for the slug path.
|
||||
|
||||
``lang``/``translation`` serve a translated version (see
|
||||
docs/localization.md): None translation = the English original.
|
||||
``link_lang`` is the ?lang= override the page was requested with,
|
||||
replicated onto the navigation links so the language sticks.
|
||||
"""
|
||||
node = resolve(menu, path)[-1]
|
||||
title = _title(path.rpartition("/")[2], node)
|
||||
main = page_content(menu, path)
|
||||
original = i18n.primary_lang(menu, path)
|
||||
if translation is None:
|
||||
lang = original
|
||||
title = _title(path.rpartition("/")[2], node, translation, path)
|
||||
main = page_content(menu, data, path, translation, link_lang, lang)
|
||||
social = _social_meta(node, path, title, str(main), brand, base_url)
|
||||
canonical, alternates = _language_urls(data, path, node, lang, original, base_url)
|
||||
return str(
|
||||
_layout(
|
||||
*_page_assets(), custom_css, theme, banner_design(menu, path, theme),
|
||||
transition, favicon, social,
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
social,
|
||||
lang,
|
||||
canonical,
|
||||
alternates,
|
||||
)(
|
||||
Title=f"{title} – {brand}" if brand else title,
|
||||
Brand=_brand_link(brand, brand_html),
|
||||
Nav=nav_html(menu, path),
|
||||
Sidebar=sidebar_html(menu, path),
|
||||
Brand=_brand_link(brand, brand_html, link_lang),
|
||||
Nav=nav_html(menu, path, translation, link_lang),
|
||||
Sidebar=sidebar_html(menu, path, translation, link_lang),
|
||||
Banner=banner_html(menu, path, theme),
|
||||
Main=main,
|
||||
),
|
||||
@@ -902,13 +1123,18 @@ def render_page(
|
||||
|
||||
def render_category(
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
path: str,
|
||||
brand: str = SITE_NAME,
|
||||
custom_css: str = "",
|
||||
theme: str = "",
|
||||
favicon: str = "",
|
||||
brand_html: str = "",
|
||||
base_url: str = "",
|
||||
transition: str = "cube",
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> str:
|
||||
"""Render the listing for a content-less category label (404).
|
||||
|
||||
@@ -916,22 +1142,42 @@ def render_category(
|
||||
children are listed as cards, like on a category page with content.
|
||||
Nav links point straight at the first child, so this is mainly seen
|
||||
in the site editor, where the pen creates the landing page.
|
||||
|
||||
With a translation (titles only — the category has no Markdown) the
|
||||
heading, navigation and card text localize per target article
|
||||
(docs/localization.md); ``link_lang`` replicates the ?lang= override
|
||||
onto the navigation links as on content pages. The hreflang alternates
|
||||
are computed as on content pages — a translated title makes the
|
||||
language available here too.
|
||||
"""
|
||||
node = resolve(menu, path)[-1]
|
||||
title = _title(path.rpartition("/")[2], node)
|
||||
original = i18n.primary_lang(menu, path)
|
||||
if translation is None:
|
||||
lang = original
|
||||
title = _title(path.rpartition("/")[2], node, translation, path)
|
||||
_, alternates = _language_urls(data, path, node, lang, original, base_url)
|
||||
doc = E.article
|
||||
with doc:
|
||||
doc.h1(title)
|
||||
if any(c.published for c in node.children.values()):
|
||||
_cards(doc, menu, node, path)
|
||||
_cards(doc, menu, data, node, path, translation, link_lang, lang)
|
||||
else:
|
||||
doc.p("This section has no page of its own yet.")
|
||||
return str(
|
||||
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon)(
|
||||
_layout(
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
lang=lang,
|
||||
alternates=alternates,
|
||||
)(
|
||||
Title=f"{title} – {brand}" if brand else title,
|
||||
Brand=_brand_link(brand, brand_html),
|
||||
Nav=nav_html(menu, path),
|
||||
Sidebar=sidebar_html(menu, path),
|
||||
Brand=_brand_link(brand, brand_html, link_lang),
|
||||
Nav=nav_html(menu, path, translation, link_lang),
|
||||
Sidebar=sidebar_html(menu, path, translation, link_lang),
|
||||
Banner=banner_html(menu, path, theme),
|
||||
Main=HTML(str(doc)),
|
||||
),
|
||||
@@ -954,7 +1200,14 @@ def render_not_found(
|
||||
doc.h1("Not Found")
|
||||
doc.p(f"No article at /{path}. If there was before, it may have been deleted.")
|
||||
return str(
|
||||
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon)(
|
||||
_layout(
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
)(
|
||||
Title=f"Not Found – {brand}" if brand else "Not Found",
|
||||
Brand=_brand_link(brand, brand_html),
|
||||
Nav=nav_html(menu, path),
|
||||
@@ -1015,6 +1268,31 @@ def _analytics_assets() -> tuple[list[str], list[str]]:
|
||||
return _asset_cache["analytics"]
|
||||
|
||||
|
||||
def _langselect_assets() -> tuple[list[str], list[str]]:
|
||||
"""Script and stylesheet URLs for the on-demand public language selector."""
|
||||
vite_url = os.environ.get("PAGERITE_VITE_URL")
|
||||
if vite_url:
|
||||
return [f"{vite_url}/src/langselect-main.js"], []
|
||||
if "langselect" not in _asset_cache:
|
||||
manifest = _manifest()
|
||||
# import() loads no CSS automatically: collect the stylesheets of
|
||||
# the entry and its imported chunks (LangSelect's ride on the
|
||||
# shared langs chunk).
|
||||
scripts, stylesheets, seen = [], [], set()
|
||||
queue = ["src/langselect-main.js"]
|
||||
for key in queue: # grows with imported chunks
|
||||
if key in seen:
|
||||
continue
|
||||
seen.add(key)
|
||||
entry = manifest[key]
|
||||
if entry.get("isEntry"):
|
||||
scripts.append(f"/{entry['file']}")
|
||||
stylesheets += [f"/{css}" for css in entry.get("css", [])]
|
||||
queue += entry.get("imports", [])
|
||||
_asset_cache["langselect"] = scripts, stylesheets
|
||||
return _asset_cache["langselect"]
|
||||
|
||||
|
||||
def render_analytics(
|
||||
menu: dict[str, Node],
|
||||
brand: str = SITE_NAME,
|
||||
|
||||
+5
-3
@@ -17,11 +17,11 @@ readme = "README.md"
|
||||
requires-python = ">=3.14"
|
||||
dependencies = [
|
||||
"blake3>=1.0.9",
|
||||
"fastapi-vue>=1.3.1",
|
||||
"fastapi-vue~=1.4.2",
|
||||
"fastapi[standard]>=0.141.1",
|
||||
"html5tagger>=2.0.0",
|
||||
"httpx>=0.28.1",
|
||||
"kanta>=0.8.1",
|
||||
"kanta>=0.9.0",
|
||||
"markdown-it-py>=4.2.0",
|
||||
"maxminddb>=3.1.1",
|
||||
"mdit-py-plugins>=0.6.1",
|
||||
@@ -29,7 +29,7 @@ dependencies = [
|
||||
"platformdirs>=4.11.5",
|
||||
"pygments>=2.20.0",
|
||||
"python-slugify>=8.0.4",
|
||||
"ua-parser>=1.0.2",
|
||||
"uarite>=0.1.2",
|
||||
"zstandard>=0.25.0",
|
||||
]
|
||||
|
||||
@@ -38,6 +38,7 @@ pagerite = "pagerite.__main__:main"
|
||||
|
||||
[project.urls]
|
||||
Repository = "https://git.zi.fi/LeoVasanko/pagerite"
|
||||
Issues = "https://github.com/LeoVasanko/pagerite"
|
||||
|
||||
[dependency-groups]
|
||||
dev = []
|
||||
@@ -54,6 +55,7 @@ packages = ["pagerite"]
|
||||
# the frontend build).
|
||||
artifacts = ["pagerite/frontend-build", "pagerite/themes", "pagerite/seed-assets"]
|
||||
only-packages = true
|
||||
packages = ["pagerite"]
|
||||
|
||||
[tool.hatch.build.targets.sdist.hooks.custom]
|
||||
path = "scripts/fastapi-vue/buildhook.py"
|
||||
|
||||
@@ -8,6 +8,8 @@ import sys
|
||||
from contextlib import suppress
|
||||
from pathlib import Path
|
||||
|
||||
import tracerite
|
||||
|
||||
# Import util.py from scripts/fastapi-vue (not a package, so we adjust sys.path)
|
||||
sys.path.insert(0, str(Path(__file__).with_name("fastapi-vue")))
|
||||
from devutil import (
|
||||
@@ -39,7 +41,7 @@ async def run_devserver(
|
||||
viteurl, npm_install, vite = setup_vite(listen, DEFAULT_VITE_PORT)
|
||||
backurl, pagerite = setup_cli("pagerite", backend, DEFAULT_DEV_PORT)
|
||||
|
||||
# Tell the everyone by environment (vite proxy and backend devmode use these)
|
||||
# Tell everyone via environment (vite proxy and backend devmode use these)
|
||||
os.environ["PAGERITE_VITE_URL"] = viteurl
|
||||
os.environ["PAGERITE_BACKEND_URL"] = backurl
|
||||
os.environ["PAGERITE_DEV"] = "1"
|
||||
@@ -54,6 +56,7 @@ async def run_devserver(
|
||||
|
||||
def main() -> None:
|
||||
"""Parse CLI arguments and run the devserver."""
|
||||
tracerite.load()
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Run Vite and FastAPI development servers",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
|
||||
+27
-19
@@ -91,22 +91,22 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
|
||||
CrawlerProfile(
|
||||
"googlebot",
|
||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
|
||||
"66.249.64.66", # US, Google
|
||||
"66.249.64.66", # US, Google
|
||||
),
|
||||
CrawlerProfile(
|
||||
"bingbot",
|
||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
|
||||
"40.77.167.0", # US, Microsoft
|
||||
"40.77.167.0", # US, Microsoft
|
||||
),
|
||||
CrawlerProfile(
|
||||
"duckduckbot",
|
||||
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
|
||||
"95.217.0.1", # Germany, Hetzner VPS
|
||||
"95.217.0.1", # Germany, Hetzner VPS
|
||||
),
|
||||
CrawlerProfile(
|
||||
"curl",
|
||||
"curl/8.5.0",
|
||||
"139.162.0.1", # Singapore, Linode VPS
|
||||
"139.162.0.1", # Singapore, Linode VPS
|
||||
),
|
||||
]
|
||||
|
||||
@@ -115,14 +115,14 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
|
||||
# host part; the host may rotate once mid-session.
|
||||
RESIDENTIAL_SOURCE_IPS: list[str] = [
|
||||
# Residential IPv4
|
||||
"91.154.140.209", # Finland, Elisa
|
||||
"84.143.145.207", # Germany, Deutsche Telekom
|
||||
"220.165.255.254", # China, Chinanet / China Telecom
|
||||
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
|
||||
"91.154.140.209", # Finland, Elisa
|
||||
"84.143.145.207", # Germany, Deutsche Telekom
|
||||
"220.165.255.254", # China, Chinanet / China Telecom
|
||||
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
|
||||
# Residential IPv6 /64 prefixes
|
||||
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
|
||||
"240e:45d:1e60:5b0::/64", # China, China Telecom
|
||||
"2409:8904:6720:4123::/64", # China, China Unicom
|
||||
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
|
||||
"240e:45d:1e60:5b0::/64", # China, China Telecom
|
||||
"2409:8904:6720:4123::/64", # China, China Unicom
|
||||
]
|
||||
|
||||
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
|
||||
@@ -130,9 +130,9 @@ RESIDENTIAL_SOURCE_IPS: list[str] = [
|
||||
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
|
||||
# and index 2 uses a fixed crawler UA.
|
||||
ABUSE_SOURCE_IPS: list[str] = [
|
||||
"45.63.0.12", # US, Vultr VPS
|
||||
"138.197.0.89", # US, DigitalOcean / Cloudways
|
||||
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
|
||||
"45.63.0.12", # US, Vultr VPS
|
||||
"138.197.0.89", # US, DigitalOcean / Cloudways
|
||||
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
|
||||
]
|
||||
|
||||
# Paths commonly probed by attackers looking for exposed config, admin panels,
|
||||
@@ -284,10 +284,16 @@ TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
|
||||
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
|
||||
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
|
||||
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
|
||||
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
|
||||
(
|
||||
"https://news.ycombinator.com/",
|
||||
{"utm_source": "hackernews", "utm_medium": "referral"},
|
||||
),
|
||||
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
|
||||
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
|
||||
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
|
||||
(
|
||||
"https://www.producthunt.com/",
|
||||
{"utm_source": "producthunt", "utm_medium": "referral"},
|
||||
),
|
||||
]
|
||||
|
||||
# Fraction of referered sessions that also carry UTM tags.
|
||||
@@ -437,7 +443,7 @@ def _random_ipv6_host(prefix: str) -> str:
|
||||
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
|
||||
if base.endswith("::"):
|
||||
base = base[:-2]
|
||||
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
|
||||
host = ":".join(f"{random.randint(0, 0xFFFF):04x}" for _ in range(4))
|
||||
return f"{base}:{host}"
|
||||
|
||||
|
||||
@@ -760,9 +766,11 @@ def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
|
||||
if ua_mode == 0:
|
||||
get_ua = _abuse_ua
|
||||
elif ua_mode == 1:
|
||||
|
||||
def get_ua() -> str:
|
||||
return BROWSER_PROFILES[0].user_agent
|
||||
else:
|
||||
|
||||
def get_ua() -> str:
|
||||
return CRAWLER_PROFILES[0].user_agent
|
||||
|
||||
@@ -804,8 +812,8 @@ def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
|
||||
nargs="?",
|
||||
default="http://localhost:8200",
|
||||
help="Base URL of the Pagerite site (default: http://localhost:8200). "
|
||||
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
|
||||
"missing scheme defaults to http://.",
|
||||
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
|
||||
"missing scheme defaults to http://.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-t",
|
||||
|
||||
@@ -7,8 +7,8 @@ import sys
|
||||
from contextlib import suppress
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any, Self
|
||||
from urllib.parse import urlsplit
|
||||
|
||||
import httpx
|
||||
from buildutil import find_dev_tool, find_install_tool, logger
|
||||
from fastapi_vue.hostutil import parse_endpoint
|
||||
|
||||
@@ -105,18 +105,47 @@ class ProcessGroup:
|
||||
await p.wait()
|
||||
|
||||
|
||||
async def http_get_server(url: str, timeout: float) -> str | None: # noqa: ASYNC109
|
||||
"""GET url with plain asyncio streams, return the response Server header.
|
||||
|
||||
Returns an empty string when the server responds without a Server header,
|
||||
and None when the server is unreachable or doesn't answer in time.
|
||||
"""
|
||||
parts = urlsplit(url)
|
||||
host = parts.hostname or "localhost"
|
||||
port = parts.port or (443 if parts.scheme == "https" else 80)
|
||||
path = parts.path or "/"
|
||||
if parts.query:
|
||||
path += f"?{parts.query}"
|
||||
try:
|
||||
async with asyncio.timeout(timeout):
|
||||
reader, writer = await asyncio.open_connection(host, port)
|
||||
try:
|
||||
writer.write(f"GET {path} HTTP/1.0\r\nHost: {host}\r\n\r\n".encode())
|
||||
await writer.drain()
|
||||
data = await reader.readuntil(b"\r\n\r\n")
|
||||
finally:
|
||||
writer.close()
|
||||
except OSError, EOFError, ValueError, TimeoutError:
|
||||
return None
|
||||
for line in data.decode("latin-1").split("\r\n"):
|
||||
if line.lower().startswith("server:"):
|
||||
return line.split(":", 1)[1].strip()
|
||||
return ""
|
||||
|
||||
|
||||
async def check_ports_free(*urls: str) -> None:
|
||||
"""Verify URLs are not responding (ports are free). Raise SystemExit if any respond."""
|
||||
|
||||
async def check(client: httpx.AsyncClient, url: str) -> None:
|
||||
with suppress(httpx.RequestError):
|
||||
res = await client.get(url, timeout=0.1)
|
||||
server = res.headers.get("server", "server")
|
||||
logger.warning("Conflicting %s already running at %s", server, url)
|
||||
async def check(url: str) -> None:
|
||||
server = await http_get_server(url, timeout=0.1)
|
||||
if server is not None:
|
||||
logger.warning(
|
||||
"Conflicting %s already running at %s", server or "server", url
|
||||
)
|
||||
raise SystemExit(1)
|
||||
|
||||
async with httpx.AsyncClient() as client:
|
||||
await asyncio.gather(*[check(client, url) for url in urls])
|
||||
await asyncio.gather(*[check(url) for url in urls])
|
||||
|
||||
|
||||
async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
|
||||
@@ -128,18 +157,14 @@ async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
|
||||
if not path:
|
||||
return
|
||||
|
||||
async with httpx.AsyncClient() as client:
|
||||
for attempt in range(max_attempts):
|
||||
try:
|
||||
await client.get(f"{url}{path}", timeout=1.0)
|
||||
except httpx.RequestError:
|
||||
if attempt == max_attempts - 1:
|
||||
logger.warning("Backend didn't start in time")
|
||||
raise SystemExit(1) from None
|
||||
await asyncio.sleep(0.1)
|
||||
else:
|
||||
logger.info("✓ Backend ready!")
|
||||
return
|
||||
for attempt in range(max_attempts):
|
||||
if await http_get_server(f"{url}{path}", timeout=1.0) is not None:
|
||||
logger.info("✓ Backend ready!")
|
||||
return
|
||||
if attempt == max_attempts - 1:
|
||||
logger.warning("Backend didn't start in time")
|
||||
raise SystemExit(1)
|
||||
await asyncio.sleep(0.1)
|
||||
|
||||
|
||||
def setup_vite(
|
||||
@@ -222,5 +247,7 @@ def setup_cli(
|
||||
host = endpoints[0]["host"]
|
||||
port = endpoints[0]["port"]
|
||||
|
||||
cmd = [cli, f"--listen={host}:{port}"]
|
||||
# Run the package as a module with the current interpreter, instead of
|
||||
# relying on a PATH-installed CLI entry point.
|
||||
cmd = [sys.executable, "-m", cli, f"--listen={host}:{port}"]
|
||||
return f"http://{host}:{port}", cmd
|
||||
|
||||
@@ -0,0 +1,378 @@
|
||||
#!/usr/bin/env python3
|
||||
# /// script
|
||||
# requires-python = ">=3.14"
|
||||
# dependencies = [
|
||||
# "accelerate>=1.14.0",
|
||||
# "msgspec>=0.19.0",
|
||||
# "torch>=2.13.0",
|
||||
# "tracerite>=2.6.5",
|
||||
# "transformers>=5.16.1",
|
||||
# "websockets>=15.0.1",
|
||||
# ]
|
||||
# ///
|
||||
"""Pagerite translator service: translate site content with Seed-X-PPO-7B.
|
||||
|
||||
Connects to a Pagerite server's translator WebSocket — the full URL
|
||||
including the access key (the admin finds the key in the site settings,
|
||||
GET /_api/settings -> ``translate_keys``) —
|
||||
and announces the languages the
|
||||
model CAN translate (capabilities). The server dispatches one single-item
|
||||
job at a time per connection, offered only in its configured target
|
||||
languages (``Data.translate_langs``) ∩ the announced capabilities; a
|
||||
dropped connection's in-flight item is simply re-offered
|
||||
(docs/localization.md). For parallelism, run multiple instances.
|
||||
|
||||
Seed-X-PPO-7B (bf16, ~15 GB) is the only supported model. The script stays
|
||||
running and connected full time; the model loads at startup (a backlog is
|
||||
likely after a downtime) and is unloaded after 60 s idle, re-loading on the
|
||||
next job — the GPU is held only while actually translating.
|
||||
|
||||
Usage:
|
||||
uv run scripts/translator.py ws://localhost:8410/_translate/KEY
|
||||
uv run scripts/translator.py wss://example.com/_translate/KEY
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import gc
|
||||
import sys
|
||||
import time
|
||||
|
||||
import msgspec
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
import tracerite
|
||||
import websockets
|
||||
|
||||
tracerite.load()
|
||||
|
||||
SEED_X = "ByteDance-Seed/Seed-X-PPO-7B"
|
||||
|
||||
# Seed-X language tags (appended to the prompt; required by its PPO training)
|
||||
SEED_X_TAGS = {
|
||||
"arabic": "ar",
|
||||
"chinese": "zh",
|
||||
"czech": "cs",
|
||||
"danish": "da",
|
||||
"dutch": "nl",
|
||||
"english": "en",
|
||||
"finnish": "fi",
|
||||
"french": "fr",
|
||||
"german": "de",
|
||||
"greek": "el",
|
||||
"hungarian": "hu",
|
||||
"indonesian": "id",
|
||||
"italian": "it",
|
||||
"japanese": "ja",
|
||||
"korean": "ko",
|
||||
"malay": "ms",
|
||||
"norwegian": "no",
|
||||
"persian": "fa",
|
||||
"polish": "pl",
|
||||
"portuguese": "pt",
|
||||
"romanian": "ro",
|
||||
"russian": "ru",
|
||||
"spanish": "es",
|
||||
"swedish": "sv",
|
||||
"thai": "th",
|
||||
"turkish": "tr",
|
||||
"ukrainian": "uk",
|
||||
"vietnamese": "vi",
|
||||
}
|
||||
SEED_X_NAMES = {v: k for k, v in SEED_X_TAGS.items()}
|
||||
|
||||
#: The fragments arrive as prose segments (pagerite/segments.py): plain
|
||||
#: text runs only — no markup, URLs, code or placeholders. The wire
|
||||
#: invariant is that segments are PURE PROSE, and one character marks the
|
||||
#: boundary both ways: "<" never appears in a segment. Sources containing
|
||||
#: it are never dispatched (pagerite/segments.py keeps them in the
|
||||
#: original language); the model's output is cut at the first "<" — one
|
||||
#: rule that covers the whole class of markup bleed (an echoed <lang> tag,
|
||||
#: a "<br>", ...) instead of a pattern per artifact. (Generation-level
|
||||
#: stop strings can't do this job: the model's <s> framing token would
|
||||
#: trip a "<" stop at the first token; skip_special_tokens strips the
|
||||
#: framing at decode.)
|
||||
#:
|
||||
#: Two kinds cross the wire (Job.kind), each with its own prompt template:
|
||||
#: titles get told they ARE titles (a lone word otherwise invites
|
||||
#: context-free readings — "About" as "approximately"). Any segment may
|
||||
#: carry its surround in Job.contexts (a title: the article's opening; a
|
||||
#: carved-out segment like a link text: its block's plain text) and is then
|
||||
#: translated together with that surround (seed_x_chunk). No punctuation
|
||||
#: clause, on purpose: Seed-X handles trailing-punctuation instructions by
|
||||
#: slipping into its [COT] reasoning mode (observed for Chinese:
|
||||
#: minutes-long generations, reasoning text in the output) —
|
||||
#: match_punctuation handles stray punctuation deterministically instead.
|
||||
PROMPTS = {
|
||||
"chunk": "Translate the following {source_lang} text into {target_lang}:\n{text} <{tag}>",
|
||||
"title": "Translate the following {source_lang} title into {target_lang}:\n{text} <{tag}>",
|
||||
# Title with the article's opening as context (Job.contexts): the model
|
||||
# translates both; generation stops at the blank line separating them,
|
||||
# and the segment's own part of the output is the translation. No
|
||||
# separator in the output (the model merged them) → seed_x_chunk falls
|
||||
# back to the plain kind template.
|
||||
"title+context": "Translate the following {source_lang} title and the beginning of its article "
|
||||
"into {target_lang}:\n{text}\n\n{context} <{tag}>",
|
||||
# A segment carved out of a larger block (link text, partial run) with
|
||||
# its sentence as context — same mechanics as title+context.
|
||||
"chunk+context": "Translate the following {source_lang} text into {target_lang}:\n"
|
||||
"{text}\n\n{context} <{tag}>",
|
||||
}
|
||||
TERMINAL_PUNCT = ".,!?:;…。,!?;:、"
|
||||
|
||||
|
||||
def match_punctuation(source: str, translated: str) -> str:
|
||||
"""Drop terminal punctuation the model added.
|
||||
|
||||
When the source segment ends without terminal punctuation, the
|
||||
translation must not gain any either. A leading Spanish ¡/¿ only pairs
|
||||
with a terminal !/?, so it goes with it.
|
||||
"""
|
||||
if not source or source[-1] in TERMINAL_PUNCT:
|
||||
return translated
|
||||
trimmed = translated.rstrip(TERMINAL_PUNCT)
|
||||
if trimmed and trimmed[0] in "¡¿":
|
||||
trimmed = trimmed[1:].lstrip()
|
||||
return trimmed
|
||||
|
||||
|
||||
# The wire structs below duplicate pagerite/translate.py: this script runs
|
||||
# in its own uv environment and cannot import the server package. The
|
||||
# "type" tag selects the frame; bytes fields ride as base64.
|
||||
class Hello(msgspec.Struct, tag="hello"):
|
||||
"""Client greeting on connect: the language codes its model CAN produce
|
||||
(capabilities). The server offers jobs only in the intersection with
|
||||
its wanted target languages."""
|
||||
|
||||
langs: list[str]
|
||||
|
||||
|
||||
class Job(msgspec.Struct, tag="job"):
|
||||
"""Server push: ONE fragment to translate. Exactly one job is in flight
|
||||
per connection — the next arrives only after this one's Result."""
|
||||
|
||||
lang: str
|
||||
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
|
||||
#: The fragment's prose segments: plain text runs only, no markup —
|
||||
#: translate each element independently (pagerite/segments.py).
|
||||
texts: list[str]
|
||||
path: str #: article it came from ("" = front page), no leading slash
|
||||
kind: str #: "chunk" | "title"
|
||||
#: Per segment (parallel to texts; "" = none): the surround to
|
||||
#: translate it in — a link text carries its sentence, a title the
|
||||
#: article's opening. See seed_x_chunk for how they are used.
|
||||
contexts: list[str] = msgspec.field(default_factory=list)
|
||||
|
||||
|
||||
class Result(msgspec.Struct, tag="result"):
|
||||
"""Client reply: the translation of the connection's current Job
|
||||
(must match its lang and key exactly)."""
|
||||
|
||||
lang: str
|
||||
key: bytes
|
||||
texts: list[str] #: the job's segments translated, same order and count
|
||||
|
||||
|
||||
#: Idle seconds after the last job before the model is unloaded (the GPU
|
||||
#: is released; the WebSocket connection and tokenizer stay).
|
||||
IDLE_UNLOAD_S = 60
|
||||
|
||||
|
||||
class SeedX:
|
||||
"""The Seed-X model, loaded at startup and re-loaded on demand.
|
||||
|
||||
Loading up front covers the likely backlog after a downtime (and any
|
||||
first-run model download) before the server starts dispatching. After
|
||||
IDLE_UNLOAD_S without a job the model is dropped and re-loaded on the
|
||||
next one — the script stays connected the whole time, holding the GPU
|
||||
only while translating. The tokenizer (small, CPU) loads once.
|
||||
"""
|
||||
|
||||
def __init__(self):
|
||||
self.tokenizer = AutoTokenizer.from_pretrained(SEED_X)
|
||||
self.model = None
|
||||
self._unload_task = None
|
||||
self._load()
|
||||
|
||||
def _load(self):
|
||||
t0 = time.monotonic()
|
||||
self.model = AutoModelForCausalLM.from_pretrained(
|
||||
SEED_X, dtype=torch.bfloat16, device_map="auto"
|
||||
)
|
||||
print(f"[seed-x loaded in {time.monotonic() - t0:.0f}s]", file=sys.stderr)
|
||||
|
||||
def get(self):
|
||||
"""The (tokenizer, model) pair, re-loading the model if it was
|
||||
idle-unloaded, and cancelling any pending idle unload."""
|
||||
if self._unload_task:
|
||||
self._unload_task.cancel()
|
||||
self._unload_task = None
|
||||
if self.model is None:
|
||||
self._load()
|
||||
return self.tokenizer, self.model
|
||||
|
||||
def idle(self):
|
||||
"""Re-arm the idle unload after a job completes (arming it at job
|
||||
START could unload under a >IDLE_UNLOAD_S generation)."""
|
||||
self._unload_task = asyncio.create_task(self._unload_later())
|
||||
|
||||
async def _unload_later(self):
|
||||
try:
|
||||
await asyncio.sleep(IDLE_UNLOAD_S)
|
||||
except asyncio.CancelledError:
|
||||
return
|
||||
if self.model is not None:
|
||||
self.model = None
|
||||
gc.collect()
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.empty_cache()
|
||||
print(f"[seed-x unloaded after {IDLE_UNLOAD_S}s idle]", file=sys.stderr)
|
||||
|
||||
|
||||
def seed_x_chunk(
|
||||
tokenizer,
|
||||
model,
|
||||
text: str,
|
||||
target_lang: str,
|
||||
tag: str,
|
||||
kind: str = "chunk",
|
||||
context: str = "",
|
||||
source_lang: str = "English",
|
||||
):
|
||||
"""Translate one segment; returns (translation, output_tokens, generation_seconds).
|
||||
|
||||
With context, the segment is translated together with its surround (a
|
||||
link text with its sentence, a title with the article's opening), and
|
||||
the segment's own part of the output is the translation: its own line
|
||||
for a single-line source (a single-line segment's translation never
|
||||
contains a line break — generation stops at the blank line separating
|
||||
the two), its own paragraph for a multi-line one (softbreak-merged
|
||||
lines keep single newlines, the separator is the blank line). If the
|
||||
model merged them — no separator, or an empty first part — fall back to
|
||||
translating the segment alone; the wasted tokens are counted either
|
||||
way.
|
||||
"""
|
||||
# No chat template on this model; the trailing language tag is required (trans/ style prompt).
|
||||
template = PROMPTS.get(f"{kind}+context" if context else kind, PROMPTS["chunk"])
|
||||
prompt = template.format(
|
||||
source_lang=source_lang,
|
||||
target_lang=target_lang,
|
||||
text=text,
|
||||
tag=tag,
|
||||
context=context,
|
||||
)
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||
t0 = time.monotonic()
|
||||
# The only stop string is the context separator. "<" must NOT be one:
|
||||
# stopping works on the raw output, which always starts with the
|
||||
# model's <s> framing token. skip_special_tokens strips <s>/</s> at
|
||||
# decode; the post-decode cut at the first "<" then enforces the wire
|
||||
# invariant (prose only) against markup bleed.
|
||||
kwargs = {"stop_strings": ["\n\n"], "tokenizer": tokenizer} if context else {}
|
||||
out = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]),
|
||||
do_sample=False,
|
||||
**kwargs,
|
||||
)
|
||||
dt = time.monotonic() - t0
|
||||
n = out.shape[1] - inputs.input_ids.shape[1]
|
||||
decoded = tokenizer.decode(
|
||||
out[0][inputs.input_ids.shape[1] :], skip_special_tokens=True
|
||||
)
|
||||
translated = decoded.partition("<")[0]
|
||||
if not context:
|
||||
return translated.strip(), n, dt
|
||||
if "\n" in text:
|
||||
# Multi-line segment: its translation keeps single newlines; the
|
||||
# blank line is the separator from the context translation.
|
||||
sep = "\n\n" in translated
|
||||
out = translated.split("\n\n", 1)[0] if sep else ""
|
||||
else:
|
||||
out, sep, _ = translated.partition("\n")
|
||||
if not sep:
|
||||
out = ""
|
||||
out = out.strip()
|
||||
if out:
|
||||
return out, n, dt
|
||||
# The model merged segment and context (no separator, or an empty first
|
||||
# part): retry without the context.
|
||||
again, n2, dt2 = seed_x_chunk(
|
||||
tokenizer, model, text, target_lang, tag, kind=kind, source_lang=source_lang
|
||||
)
|
||||
return again, n + n2, dt + dt2
|
||||
|
||||
|
||||
async def do_job(ws, job: Job, seed_x: SeedX) -> None:
|
||||
"""Translate the job's segments (one model call each) and send them back."""
|
||||
lang_name = SEED_X_NAMES[job.lang].capitalize()
|
||||
tokenizer, model = seed_x.get()
|
||||
# Deliberately blocking: nothing else needs the loop while the job is
|
||||
# being answered, and the reconnect loop recovers a dropped connection
|
||||
# (the in-flight item is simply re-offered).
|
||||
texts = []
|
||||
tokens = dt = 0
|
||||
for i, text in enumerate(job.texts):
|
||||
ctx = job.contexts[i] if i < len(job.contexts) else ""
|
||||
translated, n, t = seed_x_chunk(
|
||||
tokenizer, model, text, lang_name, job.lang, kind=job.kind, context=ctx
|
||||
)
|
||||
texts.append(match_punctuation(text, translated))
|
||||
tokens += n
|
||||
dt += t
|
||||
print(
|
||||
f"[{job.lang} {job.kind} {job.path or '/'}: {len(texts)} segments, "
|
||||
f"{tokens} tokens in {dt:.1f}s = {tokens / dt:.1f} tok/s]",
|
||||
file=sys.stderr,
|
||||
)
|
||||
await ws.send(
|
||||
msgspec.json.encode(Result(lang=job.lang, key=job.key, texts=texts)).decode()
|
||||
)
|
||||
seed_x.idle()
|
||||
|
||||
|
||||
async def serve(url: str, seed_x: SeedX) -> None:
|
||||
"""Connect, announce capabilities, answer jobs; reconnect with backoff."""
|
||||
seed_x.idle() # the startup load also unloads when no work arrives
|
||||
backoff = 1
|
||||
while True:
|
||||
try:
|
||||
async with websockets.connect(url) as ws:
|
||||
backoff = 1
|
||||
await ws.send(
|
||||
msgspec.json.encode(Hello(langs=sorted(SEED_X_NAMES))).decode()
|
||||
)
|
||||
print(
|
||||
f"[connected; announced {len(SEED_X_NAMES)} language capabilities]",
|
||||
file=sys.stderr,
|
||||
)
|
||||
async for raw in ws:
|
||||
await do_job(ws, msgspec.json.decode(raw, type=Job), seed_x)
|
||||
except websockets.exceptions.InvalidHandshake:
|
||||
sys.exit("handshake rejected; check the URL (including the key)")
|
||||
except (OSError, websockets.exceptions.ConnectionClosed) as e:
|
||||
print(
|
||||
f"[connection lost ({e}); reconnecting in {backoff}s]", file=sys.stderr
|
||||
)
|
||||
await asyncio.sleep(backoff)
|
||||
backoff = min(backoff * 2, 60)
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(
|
||||
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
|
||||
)
|
||||
p.add_argument(
|
||||
"url",
|
||||
help="full translator WebSocket URL including the key, "
|
||||
"e.g. ws://localhost:8410/_translate/KEY",
|
||||
)
|
||||
args = p.parse_args()
|
||||
if not args.url.startswith(("ws://", "wss://")):
|
||||
p.error("url must start with ws:// or wss://")
|
||||
|
||||
asyncio.run(serve(args.url, SeedX()))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user