Compare commits

...
31 Commits
Author SHA1 Message Date
LeoVasanko cbd50cfece Clip chart plot curves to the chart area with a per-chart SVG clipPath.
Past-week overlays can run far above the autoscaled y range, and the svg
is overflow: visible for the axis labels — wrap the data paths (bars,
skyline, area/line curves) in a clipped group so they cannot paint outside
the plot rect.
2026-09-03 15:46:42 +00:00
LeoVasanko 85ef968cd6 Reclassify sub-5s visits as crawlers, show crawler referers and abuser articles.
Visits with under 5 s of total reported reading time are JS-running bots:
display() converts them to crawler hits (one per trail page, with referer
and UTM query) and excludes them from every aggregate.  Crawler referers
join the favicon fetch origins and render with their icon in the crawler
table.  Abuse hits now record the real response status, so the abuse table
splits 404 probes from the articles the abuser actually read (200 GETs),
shown as trail links like the visitor/crawler tables.
2026-09-03 15:43:43 +00:00
LeoVasanko 7631c9f0a3 Prioritize titles in the translator dispatch queue.
All pending titles are now offered before any article chunks (stable
sort, menu order kept within each kind) — a page's name in the menu is
its most visible string.
2026-09-03 15:26:36 +00:00
LeoVasanko d7d03754d1 Record the acting user and use noun-based transaction actions.
Every request-driven kanta transaction now passes user= from the
Remote-User header the SSO/forward-auth proxy sets (the editor socket
reads it from the WebSocket headers; the translator worker keeps its
client key). Action labels are short identifiers naming the object, not
sentences: page / page:{lang} / page:title / page:{lang}:title /
page:language / page:slug / page:delete / structure:reorder / settings /
translate:reset for admin actions, translate:{lang}[ :title ] for worker
submissions (title results identified via the job kind, now tracked in
the connection state).
2026-09-03 15:24:32 +00:00
LeoVasanko 65fec6c519 Drop the startup translator-URL log line; the key and URL are shown in the site settings UI. 2026-09-03 14:42:02 +00:00
LeoVasanko dc55445ae0 Find link/formatting mark boundaries in translations by fuzzy word alignment.
Weight-ratio mapping alone was routinely off by a word and could glue a
mark to its neighbor (losing the space between). Now each mark's source
words are aligned to the translation's words by form similarity
(sequence ratio + shared prefix, case-folded, capitalization bonus) with
cheap skip penalties, so inflection, dropped articles/prepositions and
reordering don't break the match; slices are cut exactly at word
boundaries. Alignments without an anchor pair fall back to the weight
ratio (still the CJK path).
2026-09-03 14:38:59 +00:00
LeoVasanko b133ad6dd6 Fix image .margin positioning where caption and image got that layout instead of the figure wrapper getting it. 2026-09-03 14:12:35 +00:00
LeoVasanko 4413c7efdf Cleanup: break up the massive app.py into separate modules of manageable size. 2026-09-03 13:36:29 +00:00
LeoVasanko 2f533eaf09 Suppress mediapreview info logs. 2026-09-03 03:57:12 +00:00
LeoVasanko 1d46843c76 Add missing proxy path to vite config. 2026-09-03 03:55:09 +00:00
LeoVasanko d3e2196c83 Localization (#1)
Implement comprehensive content localization, admin panels for editing each language, AI translation interface with automatic updates when base language version is changed.
- SEO tags for all language URLs
- Uses accept-language by default, ?lang=en overrides temporarily
- User edits patched on top of translations
- RTL language supportReviewed-on: #1
2026-09-03 03:54:27 +00:00
LeoVasanko b6e6e46cfb Update URLs 2026-08-31 22:33:42 +00:00
LeoVasanko 7a4544731d Fastapi-vue-setup updated to 1.4.0: prettier logs. 2026-08-31 20:50:38 +00:00
LeoVasanko fd1987c9b3 Slight color change for inline code text to make it stand out of paragraph text. Not applied to list items and headings. 2026-08-30 03:40:03 +00:00
LeoVasanko 85a306296f Never break inline code spans (nowrap)
word-break: keep-all does not suppress breaks at hard hyphens (an
explicit UAX #14 break opportunity), so `--arg` could split after the
dashes. Inline code is short — nowrap is the safe fix.
2026-08-30 02:20:54 +00:00
LeoVasanko 2b9c7635d3 Figure-wrap lone images even with trailing inline attrs
Space-separated brace attrs on a lone image's own line (e.g.
{style="max-width: 40em"}) are consumed onto the paragraph but leave
an empty text token in the inline children, which defeated the
lone-image check — the image stayed a bare inline <img> without the
figure wrapper (and without lightbox/zoom treatment). Strip empty text
tokens before the check.
2026-08-30 02:20:54 +00:00
LeoVasanko 7e8e0a2ea7 Click-to-enlarge lightbox for article figures
Clicking a figure image opens a full-viewport overlay: the image as
large as fits (100vw, flex-shrunk to leave exactly the caption's
height) with the caption below. Click or any key closes it; wheel and
touch scrolling stop at the overlay (overscroll-behavior: contain).

Styling lives in the base theme: blurred dark backdrop (theme-tunable
via --lightbox-bg / --lightbox-text), soft shadow, reduced-motion-aware
open animation.
2026-08-30 02:08:06 +00:00
LeoVasanko 8fc5dd7b9b Keep margin boxes in the column flow, position them out of flow
::: aside / .margin blocks no longer split the column segments: they
stay inside the .colseg at their anchor point, so the surrounding text
is measured and laid out as a single columned layout (split-off sides
previously lost .cols when too short on their own).

The zone placements (multicol side zone, sidebar track, wide gutter)
now position the boxes absolutely off the article's left border —
unaffected by any column layout inside — with the vertical spot coming
from the unset top (where the box occurs in the text). Narrow widths
keep the in-column float fallback. Trade-off: out-of-flow boxes no
longer stack via clear, so boxes anchored close together may overlap.
2026-08-30 01:41:47 +00:00
LeoVasanko fbebddeaa9 Style inline code; drop explicit colorspace from color-mix
Inline code gets padding-inline, keep-all word-break, and a color nudged
30% toward --muted from the inherited color, so code inside
accent-colored text keeps its hue.

All color-mix calls now rely on the default interpolation space (oklab)
instead of declaring srgb/oklab explicitly.
2026-08-30 01:22:22 +00:00
LeoVasanko 31895065f8 Collect analytics over a /_ws WebSocket instead of POST /_a pings 2026-08-29 21:20:04 +00:00
LeoVasanko 447a565b05 Disallow /auth/ and /_api in robots.txt 2026-08-29 20:45:39 +00:00
LeoVasanko 399aa95d44 Encapsulate all migrations in kanta migrate_vN, drop persisted Data.version
migrate_v1 now also rebuilds the legacy flat pages store as the menu
tree (moved from the app lifespan, raw-dict level); migrate_v2 now also
backfills missing AVIF/WebP/JPEG derivatives on disk (moved from the
lifespan) and drops the obsolete version field. The version render
counter was cache-invalidation state, not database state: replaced by an
in-memory render generation that clears the page-body LRU and feeds page
ETags. The legacy Page struct and Data.pages/version fields are removed;
old databases lose the stale keys on re-serialization.
2026-08-29 20:16:13 +00:00
LeoVasanko f793d21c5e Serve images extension-less at /_f/{hash} with Accept-negotiated AVIF/WebP/JPEG
Uploaded images (SVGs rasterized, GIFs excepted) are stored as the
original (hash.orig.ext, internal only, never served) plus AVIF primary
and WebP/JPEG fallback derivatives re-encoded from it. Pages link the
bare hash; the server serves a format only when Accept lists it
explicitly (image/avif -> AVIF, image/webp -> WebP, else JPEG) with
vary: accept, while an explicit extension pins the format. Favicons go
through the same pipeline at 192px. migrate_v2 rewrites old
/_f/{hash}.avif article links, a startup backfill creates missing
derivatives, and twitter:image pins the .webp variant for X's scraper.
2026-08-29 19:46:47 +00:00
LeoVasanko fb3e6d1a04 Use runtime-only Vue build, raise chunk warning to 1.2 MB 2026-08-29 06:27:16 +00:00
LeoVasanko fd75a260b5 Insert images and tables as block-level fresh lines
uploadImage and insertTable no longer inject at the cursor: on a
non-empty line (e.g. inside an existing image tag) the block goes on a
fresh blank-separated line after it, never into it.
2026-08-29 06:09:51 +00:00
LeoVasanko 6058853341 Recompress uploaded images to thumbnailed AVIF via mediapreview
PUT /_api/files now runs raster uploads through mediapreview.dispatch
(temp file for format routing: pyvips, ffmpeg for HEIC/HEIF/AVIF),
storing the untouched original as <hash>.orig<ext> and serving the
AVIF derivative <hash>.avif in links. SVG/GIF and failed conversions
fall back to plain <hash><ext> storage. FileStore.delete removes the
whole hash pair. Adds mediapreview[standard] dependency.
2026-08-29 05:59:37 +00:00
LeoVasanko 9adc48479f Scroll page on cursor move only when cursor leaves viewport
Cursor-driven editor→page scroll sync pinned the cursor's page position
at a fixed window height, so every cursor move dragged the page along.
Now the page scrolls only when the cursor's mapped position crosses a
viewport edge margin, and just enough to bring it back inside.
2026-08-29 05:08:52 +00:00
LeoVasanko 864492b897 Article editor toolbar: toggling fences/links, class pickers, sizes, Tab indent
- Image insert always adds an empty "" caption with the cursor inside.
- Fenced blocks (``` code, ::: aside) share one toggle: clicked inside
  one it is removed and the content selected; otherwise the selection
  (expanded to whole lines) is wrapped, cursor left on the opener line.
- Code button: inline wrap toggles (selection preserved, backtick runs),
  line-spanning selections make fenced blocks.
- Link button toggles: clicked inside [label](url) it unwraps.
- Block classes via two pickers (placement, AA size) with current-class
  indication and a "normal" reset; placement replaces ::: container
  names, fences get their own attribute line.
- New .small/.large/.huge text-size classes (0.7/1.5/3em); base body is
  exactly 1rem so the scale is uniform.
- .left/.right floats generalized from figures to any block.
- Toolbar polish: non-emoji glyphs, scaled-up symbols, borderless
  hover/active states, reordered (B/i and size last).
- Editor panel no longer closes on Escape.
- CodeMirror editors capture Tab/Shift-Tab for indent/dedent.
2026-08-29 05:04:01 +00:00
LeoVasanko 515c6e1435 Block attrs space-separated at end of a text line
_block_attrs now applies a trailing {...} to the block when it ends the
last text line after whitespace (some text {.small}), not only on a line
of its own; a space is what keeps the braces off an image/link ending the
line. Glued-to-text braces stay literal.
2026-08-29 05:03:30 +00:00
LeoVasanko daa1670653 Themeable selection color via --selection-bg, shared by page and CodeMirror
Nitro's orange accent selection fill clashed with accent-colored text.
Introduce --selection-bg (base: accent 30% mix, as before) used by both
::selection and the editors' selection layer; nitro overrides it with a
neutral grey. CodeMirror's selection needed a baseTheme with its exact
&light/&dark selectors to win, and is now hidden when unfocused like a
normal input.
2026-08-29 03:39:24 +00:00
LeoVasanko e3fbce8ece Fix slug input: free typing with live space→hyphen and lowercasing
Live-filtering through slugify ate hyphens and spaces (trailing hyphens
are stripped) and jumped the cursor to the end on mid-text edits. Now
only spaces→hyphens and lowercasing happen oninput (length-preserving,
cursor stays put); the full slugify runs at commit before the server
call, which keeps its own validation.
2026-08-29 03:14:57 +00:00
56 changed files with 7779 additions and 1983 deletions
+15 -3
View File
@@ -10,9 +10,18 @@ Please instead ask the user to see from dev tools what you need, e.g. to look up
Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes: Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes:
- `pagerite/` — Python backend package (hatchling build target). - `pagerite/` — Python backend package (hatchling build target).
- `app.py` — FastAPI app and route registration. - `app.py` thin FastAPI assembly: lifespan, `FastAPI(...)`, router includes (route ordering: api/tracking/files routers, then `frontend.route(app, "/")`, then the pages catch-all last).
- `state.py` — shared core, no routes: env-derived site constants, `data`/`kanta`, `analytics_store`, the fastapi-vue `frontend`, the render cache (`_html_response`, `_invalidate_pages`), the translator `dispatcher`, slug helpers, `@kanta.bootstrap` hooks.
- `files.py``FileStore` (content-addressed, RAM-cached), image derivative helpers (`store_image`), file routes (`/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, favicon settings).
- `api.py` — editor REST + WS: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, `/_translate/{key}`.
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetch, `/_ws`, `/_api/ws/analytics`, the `/_a` page (docs/analytics.md).
- `pages.py` — public content pages: `/`, `/sitemap.xml`, `/robots.txt`, the `/{path:path}` catch-all.
- `data.py` — msgspec Structs for the kanta database. - `data.py` — msgspec Structs for the kanta database.
- `migrations.py` — kanta schema migrations (`migrate_vN`), e.g. v1 moves legacy in-db file blobs to the on-disk store. - `chunks.py` — block-level Markdown chunking and content-hash keys for the chunk stores (docs/migrate.md).
- `i18n.py` — language selection, translation assembly (chunks + patches) and translated-edit recording (user patches, per-language title overrides, refresh).
- `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); api.py only registers the route.
- `segments.py` — the translation round trip: fragments split into pure-prose wire segments (via markdown.make_md's verbatim parser; link- and formatting-carrying blocks stay whole, link/formatted texts inline, Markdown stripped) and translations spliced back by source offset, link/formatting markdown re-inserted at weight-mapped positions (docs/localization.md).
- `migrations.py` — kanta migrations (`migrate_vN`); ALL schema/storage upgrades live here (raw state dict before struct decoding), never in the app lifespan: v1 moves legacy in-db file blobs to the on-disk store and rebuilds the legacy flat `pages` as the menu tree, v2 rewrites `/_f/{hash}.ext` image links to the extension-less form, backfills AVIF/WebP/JPEG derivatives on disk and drops the obsolete `version` field.
- `markdown.py` — markdown-it-py renderer. - `markdown.py` — markdown-it-py renderer.
- `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`). - `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`).
- `seed.py` — demo content, written only on first database creation. - `seed.py` — demo content, written only on first database creation.
@@ -21,8 +30,11 @@ Pagerite is a CMS. See `docs` for the full design and implementation details. Ke
- `main.js` — Vue editor app entry. - `main.js` — Vue editor app entry.
- `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`). - `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`).
- `pagerite.js` — public page entry. - `pagerite.js` — public page entry.
- `editorLang.js` + `LangSelect.vue` — the editor shell's shared language selection and its selector component (page + structure tabs; drives the page preview while the panel is open, via `swapdoc.setLangOverride`).
- `reconnect.js` — shared WebSocket pacing for all sockets (staggered connect slots, stuck-CONNECTING watchdog, exponential backoff): bursts and rapid retries trip the browser's WebSocket throttling.
- `assets/` — base CSS, Pygments styles, fonts. - `assets/` — base CSS, Pygments styles, fonts.
- `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test). - `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test).
- `scripts/translator.py` — Seed-X translator service client for the `/_translate/{key}` socket (reference client, runs in its own uv env via PEP 723); stays connected full time, unloads the model after 60 s idle and reloads on the next job.
Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed). Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed).
@@ -53,5 +65,5 @@ Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed).
- Keep dependencies minimal; add via `uv add` and mention it. - Keep dependencies minimal; add via `uv add` and mention it.
- The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content. - The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
- No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. - No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. The one keyed exception is `/_translate/{key}` (translator service; `Data.translate_keys`, see docs/localization.md).
- Update the relevant MarkDown files when architecture, tooling, or conventions change. - Update the relevant MarkDown files when architecture, tooling, or conventions change.
+79 -43
View File
@@ -8,10 +8,11 @@ directory, e.g. `localhost/analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`, - `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
`CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map, `CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map,
atomic JSON persistence). atomic JSON persistence).
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`), - `pagerite/pages.py` — entry-referer stashing in `show_page` (`_track_entry`,
the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics` in `pagerite/tracking.py`), 404 recording.
(admin-gated like every `/_api` endpoint). - `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
- `frontend/src/pagerite.js` — client navigation pings and the 📊 pen. `WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the - `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
normal site layout on the `/_a` analytics page. normal site layout on the `/_a` analytics page.
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView` - `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
@@ -19,46 +20,64 @@ directory, e.g. `localhost/analytics.json`).
## What is collected ## What is collected
The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
`fr`, `to`, `hide` and `read` as query parameters (`fr` = source path; whole browsing session and sends activity messages over it — JSON text
falsy values are omitted): frames matching the server's `Ping` msgspec struct with the fields `fr`
(source path), `to` (navigation target), `read` (active seconds on `fr`
since the last report) and `hide`; falsy fields are omitted. One channel
follows the session, so the activity of a visit stays tied together, and
while the user is active the accumulated reading time is flushed every few
seconds: the trail times are cumulative, so a disconnection simply leaves
the last reported time in place (no close beacon). After 5 minutes without
any activity the client closes the socket itself — a sleeping browser tab
would lose it anyway — and the next activity reconnects as a fresh session;
reconnects are attempted only on user activity, with an exponential backoff
between attempts so a failing endpoint is never hammered. Idle-time link preloads
stay plain `fetch()` calls so the browser may cache the responses; the
WebSocket reports actual navigations and active time spent on a page.
- **Initial page load**: only `to` — the loaded path — is sent, never `fr` - **Initial page load**: only `to` — the loaded path — is sent, never `fr`
(an `fr` equal to `to` would log a bogus self-transition when a session (an `fr` equal to `to` would log a bogus self-transition when a session
already exists, e.g. a second tab). This ping is what starts already exists, e.g. a second tab). This message is what starts
the visit and counts the entry page view — the document GET alone records the visit and counts the entry page view — the document GET alone records
nothing, so bots never register (admin browsing does register, but nothing, so bots never register (admin browsing does register, but
flagged `hide`; see **Admins** below). JS-running crawlers flagged `hide`; see **Admins** below). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent (Googlebot, GoogleOther, Applebot, ...) do connect and report, but their
gives them away: pings whose UA matches `_is_bot_ua` (anything calling User-Agent gives them away: messages whose UA matches `_is_bot_ua`
(anything calling
itself a "bot", plus known exceptions such as GoogleOther) are ignored itself a "bot", plus known exceptions such as GoogleOther) are ignored
server-side, and their document GETs land in the crawler list instead. server-side, and their document GETs land in the crawler list instead.
Real-browser bots whose UA does not match still register a visit, but
their reported reading time stays under 5 seconds, so they are
reclassified as crawler hits at display time (see **Crawler hits** below).
No source-IP verification is done: a spoofed bot UA merely lands in the No source-IP verification is done: a spoofed bot UA merely lands in the
crawler stats, and scanners that probe telltale paths are caught by the crawler stats, and scanners that probe telltale paths are caught by the
abuse rules regardless. Reloads are not abuse rules regardless. Reloads are not
visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
refresh neither counts a second view nor logs a self-transition. The GET refresh neither counts a second view nor logs a self-transition. The GET
handler stashes a cross-origin https `Referer` (origin part only — handler stashes a cross-origin https `Referer` (origin part only —
unavailable to JS once the page has loaded) and any unavailable to JS once the page has loaded) and any
`utm_*` query parameters in in-memory IP tables, consumed by the ping that `utm_*` query parameters in in-memory IP tables, consumed by the first
message that
starts the visit; internal or absent referers never touch the referer table. starts the visit; internal or absent referers never touch the referer table.
- **Internal fetch-navigations**: `to` is the target path, sent only after - **Internal fetch-navigations**: `to` is the target path, sent only after
the swap actually happened (a failed swap falls back to a full load, the swap actually happened (a failed swap falls back to a full load,
whose initial ping counts the view instead — no gap, no double count). whose initial message counts the view instead — no gap, no double count).
- **External links** (`https` only): `to` is the link's full URL. This is the - **External links** (`https` only): `to` is the link's full URL. This is the
exit-link record; the user may continue navigating afterwards (new tab, exit-link record; the user may continue navigating afterwards (new tab,
back), so the exit URL is not necessarily the last trail entry. Outbound back), so the exit URL is not necessarily the last trail entry. Outbound
links are stored by full URL so several links to the same domain remain links are stored by full URL so several links to the same domain remain
distinct. distinct.
- **Excluded**: back/forward (popstate) navigations, navigating *to* the - **Excluded**: back/forward (popstate) navigations, navigating *to* the
analytics page (`/_a` — its GET is untracked, and the server rejects it analytics page (`/_a` — its GET is untracked, and the server cannot
as a ping target anyway), and everything while the user has the editor record it as a navigation target anyway), and everything while the user has
the editor
open (`body.editing`). Admin noise, not visits. Navigating *away* from open (`body.editing`). Admin noise, not visits. Navigating *away* from
`/_a` does ping: the fetch-navigation already GET-ed the target page `/_a` does report: the fetch-navigation already GET-ed the target page
without the preload header, and without the ping that GET would flush to without the preload header, and without the message that GET would flush to
the crawler list. the crawler list.
- **Admins**: when SSO is in use and the session is known to be an admin, - **Admins**: when SSO is in use and the session is known to be an admin,
the client still pings but adds `hide=1`. The activity is recorded as the client still reports but adds `hide`. The activity is recorded as
usual (navigations and all), but the `hide` flag is set on the **client usual (navigations and all), but the `hide` flag is set on the **client
record** — so it covers everything that client ever did: visits and record** — so it covers everything that client ever did: visits and
crawler hits from before the login included. Hidden clients never appear crawler hits from before the login included. Hidden clients never appear
@@ -74,14 +93,14 @@ falsy values are omitted):
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the ("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
https origin and accepted only when the client sent exactly that. https origin and accepted only when the client sent exactly that.
- **External-site favicons**: for every external https origin seen as a visit - **External-site favicons**: for every external https origin seen as a visit
referer or an exit link, the server fetches `{origin}/favicon.ico` in a referer, a crawler-hit referer or an exit link, the server fetches `{origin}/favicon.ico` in a
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only — background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
SVG is sniffed from the body when served without an image type) and stores SVG is sniffed from the body when served without an image type) and stores
the icon content-hashed on disk in the FileStore (served at `/_f/{name}`, the icon content-hashed on disk in the FileStore (served at `/_f/{name}`,
extension matching the actual MIME). The origin → file name mapping is extension matching the actual MIME). The origin → file name mapping is
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
recorded too and retried only after 7 days. Fetches are scheduled after recorded too and retried only after 7 days. Fetches are scheduled after
each ping and once at startup, which backfills icons for already-recorded each activity message and once at startup, which backfills icons for already-recorded
data. The viewer payload carries `favicons` (origin → `/_f/...` path), data. The viewer payload carries `favicons` (origin → `/_f/...` path),
and the viewer shows the icon wherever an external site is mentioned: and the viewer shows the icon wherever an external site is mentioned:
referer/exit trail links in the visit table and the source/exit pills of referer/exit trail links in the visit table and the source/exit pills of
@@ -100,8 +119,8 @@ falsy values are omitted):
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB `host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
root, it is loaded at startup and used to look up `country`/`city`. These root, it is loaded at startup and used to look up `country`/`city`. These
lookups run in background tasks after the event is stored, so the `/_a` lookups run in background tasks after the event is stored, so WebSocket
response is never delayed. The decompressed `dbip-*.mmdb` file is kept in message handling is never delayed. The decompressed `dbip-*.mmdb` file is kept in
the repository root and ignored by git. The CLI flag `--dbip` the repository root and ignored by git. The CLI flag `--dbip`
(`uv run pagerite --dbip`) downloads the latest (`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts, `dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
@@ -110,18 +129,27 @@ falsy values are omitted):
file is used. file is used.
- **Crawler hits**: every document GET is queued in RAM as a pending crawler - **Crawler hits**: every document GET is queued in RAM as a pending crawler
hit — except idle-time link preloads from pagerite.js, which carry an hit — except idle-time link preloads from pagerite.js, which carry an
`x-pagerite-preload` header and are not tracked at all (the ping sent when `x-pagerite-preload` header and are not tracked at all (the navigation
the user actually navigates to a preloaded page does the counting; forging message sent when the user actually navigates to a preloaded page does
the counting; forging
the header only hides a GET from the crawler stats, the path-based abuse the header only hides a GET from the crawler stats, the path-based abuse
classification is unaffected). If a ping classification is unaffected). If a message
from the same client arrives within 10 seconds the hit is discarded; from the same client arrives within 10 seconds the hit is discarded;
otherwise it is written to `crawlers` — unless the client is hidden otherwise it is written to `crawlers` — unless the client is hidden
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as (admin), in which case the hit is discarded on expiry too. Crawlers do not count as
visits or views. The `Accept-Language` header is stored on the shared visits or views. Bots running real browsers can still slip past the UA
check: a visit whose total reported reading time stays under 5 seconds
(`_MIN_VISIT_READ`; durations are client-provided and trusted — such bots
report 02 s) is reclassified as crawler hits at display time, one hit
per internal trail page, and counts in no visit aggregate. The
`Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip `Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited; the crawler table lists a trail of internal pages that crawler visited, preceded by its referer
when there is one — spiders often advertise their own site as the
referer, and it is rendered with its favicon like visit referers (crawler
referers are included in the favicon fetch origins). The crawler table lists
the most recent crawler first, with the most active as a tie-breaker. the most recent crawler first, with the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment - **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php` starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
@@ -131,13 +159,16 @@ falsy values are omitted):
random-UA scanner no longer pollutes the crawler stats of the legitimate random-UA scanner no longer pollutes the crawler stats of the legitimate
bot it impersonates. Once classified, every document GET and 404 from the bot it impersonates. Once classified, every document GET and 404 from the
IP is recorded as an abuse hit with the full request path (query string IP is recorded as an abuse hit with the full request path (query string
included), and its pings are ignored. The classified IP set (`abuse_ips`) included), and its activity messages are ignored. The classified IP set (`abuse_ips`)
is persisted in the JSON file; the plain-404 counters are RAM-only. In the is persisted in the JSON file; the plain-404 counters are RAM-only. In the
viewer, abuse hits are grouped by IP (never by client/UA — scanners viewer, abuse hits are grouped by IP (never by client/UA — scanners
randomize theirs) in a separate "Abuse" table. Identical paths are randomize theirs) in a separate "Abuse" table. Identical paths are
collapsed into one entry with their hit count; flagged paths that collapsed into one entry with their hit count. The 404 probes ("paths
triggered classification are lifted to the top, followed by other 404s and abused": flagged paths that triggered classification first, then other
then document GETs from the abuser. Raw User-Agent strings are shown one 404s) are kept in a separate column from the real articles the abuser
actually read ("articles read": document GETs that returned 200, not the
404 fallback rendering — rendered as trail links like the visitor and
crawler tables, with the query string stripped). Raw User-Agent strings are shown one
per line with their occurrence counts, and the full lists are click-to-copy. per line with their occurrence counts, and the full lists are click-to-copy.
## Visits and sessions ## Visits and sessions
@@ -145,9 +176,9 @@ falsy values are omitted):
There are no cookies. A visit is tied together by a client hash — the first There are no cookies. A visit is tied together by a client hash — the first
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6 6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
/64 network), the raw `User-Agent` string and the extracted /64 network), the raw `User-Agent` string and the extracted
`Accept-Language` tag. The first ping from a client hash starts a new `Accept-Language` tag. The first message from a client hash starts a new
visit; subsequent pings extend it. Pings arriving with no known session visit; subsequent messages extend it. Messages arriving with no known session
(server restart) start a fresh visit from the first ping — treated as (server restart) start a fresh visit from the first message — treated as
missing data rather than dropped. The client-hash → visit map and the IP → missing data rather than dropped. The client-hash → visit map and the IP →
entry-referer/UTM tables are in-memory only; client metadata is stored in entry-referer/UTM tables are in-memory only; client metadata is stored in
`Analytics.clients` keyed by the client hash. `Analytics.clients` keyed by the client hash.
@@ -164,7 +195,7 @@ Each `Client` record:
- `ua` — raw `User-Agent` string, - `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when - `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string, parsable, otherwise the raw string,
- `hide` — true for admin clients (`hide=1` ping): all their visits, - `hide` — true for admin clients (`hide` message field): all their visits,
crawler hits and abuse hits are recorded but excluded from every crawler hits and abuse hits are recorded but excluded from every
statistic and from the viewer payload. statistic and from the viewer payload.
@@ -180,7 +211,7 @@ Each `Visit` record:
reading time in seconds (`read`) and the most recent HTTP status seen reading time in seconds (`read`) and the most recent HTTP status seen
for the target (`status`). Re-visiting an already seen target updates for the target (`status`). Re-visiting an already seen target updates
its item instead of appending. its item instead of appending.
- `navs` — every navigation ping (`fr`, `to`), keyed by its timestamp, - `navs` — every navigation message (`fr`, `to`), keyed by its timestamp,
repeats included. The aggregates are computed from this log at display repeats included. The aggregates are computed from this log at display
time. time.
- `utm``utm_*` query parameters from the landing URL, as a dict. - `utm``utm_*` query parameters from the landing URL, as a dict.
@@ -202,14 +233,16 @@ Each `AbuseHit` record:
- `client` — 6-byte blake3 hash referencing `Analytics.clients`, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the path that triggered abuse classification (telltale - `flag` — true for the path that triggered abuse classification (telltale
path or the 404 that crossed the threshold), path or the 404 that crossed the threshold),
- `is_404` — true for 404 responses, false for document GETs from the - `is_404` — true for 404 responses (probed paths and 404-fallback document
abuser. GETs), false for real (200) document GETs — articles the abuser read.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical paths are collapsed with their counts; flagged paths Abuse table identical paths are collapsed with their counts, split into the
that triggered classification are lifted to the top, followed by other 404s 404 probes (flagged paths that triggered classification first, then other
and then document GETs from the abuser. Within each category paths are 404s, shown verbatim) and the 200 document GETs shown as trail links in the
separate articles column.
Within each list paths are
sorted by count descending, then by their earliest hit. sorted by count descending, then by their earliest hit.
In the visitor and crawler tables, internal paths that returned a 404 status In the visitor and crawler tables, internal paths that returned a 404 status
@@ -220,14 +253,17 @@ to tell misses from real pages at a glance.
Aggregates are **not stored**; they are computed at display time by Aggregates are **not stored**; they are computed at display time by
`Store.display()` from the visit records (entry + `navs` log), skipping `Store.display()` from the visit records (entry + `navs` log), skipping
hidden clients' visits. This is what allows a client to become hidden after hidden clients' visits and short visits reclassified as crawler hits
(under `_MIN_VISIT_READ` seconds of total reported reading time). This is
what allows a client to become hidden after
navigations were already logged: no counts need reversing. The computed navigations were already logged: no counts need reversing. The computed
shapes, part of the WebSocket payload (`Display` struct alongside `visits`, shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
`crawlers`, `abuse` and `clients`): `crawlers`, `abuse` and `clients`):
- `transitions`: time series of page transitions, sparse nested dict - `transitions`: time series of page transitions, sparse nested dict
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the `from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
referer origin or `"(direct)"` for initial loads, a page path for pings. referer origin or `"(direct)"` for initial loads, a page path for
navigations.
- `views`: time series of page loads, `path -> bucket -> count`, sparse: only - `views`: time series of page loads, `path -> bucket -> count`, sparse: only
non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp). non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp).
Every load counts, including repeats within a visit; external exit origins Every load counts, including repeats within a visit; external exit origins
+14 -6
View File
@@ -4,13 +4,21 @@ The Python backend lives in `pagerite/`.
## `app.py` ## `app.py`
The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything defined earlier wins. The one exception is the content catch-all `/{path:path}`, registered AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all. Thin FastAPI assembly: lifespan (open the kanta database, load the file store, the frontend build and GeoIP), the `FastAPI(...)` instance with built-in API docs disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content, the `server` header middleware, and router includes. The routes themselves live in specialized modules:
- `state.py` — shared core, no routes: the environment-derived site constants (`HOSTNAME`, `SITE_URL`, `DB_PATH`, `FILES_DIR`, image/favicon tunables), the `data` root and its `kanta` handle (`Kanta(..., migrations="pagerite.migrations")`), the `analytics_store`, the fastapi-vue `frontend`, the page render cache and `_html_response`, the translator `dispatcher`, the slug charset helpers, and the `@kanta.bootstrap` hooks (demo seed, translator defaults).
- `files.py` — the `FileStore` and image derivative helpers, and the file routes: `/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, the favicon settings endpoints.
- `api.py` — the editor REST API and WebSockets: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, and the translator channel `/_translate/{clientkey}`.
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetching, debounced broadcasts, the `/_ws` activity socket, the admin stream `/_api/ws/analytics`, and the `/_a` viewer page.
- `pages.py` — the public content pages: `/`, `/sitemap.xml`, `/robots.txt` and the `/{path:path}` catch-all.
Route ordering is load-bearing and lives in `app.py`: the api/tracking/files routers are included BEFORE `frontend.route(app, "/")` is called — fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything registered earlier wins. The content catch-all `/{path:path}` is included AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours. The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding, and `data.version`, which bumps on every content/settings change and so transparently invalidates the whole cache. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and `data.version`; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304. Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `state.py`'s `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore`: content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Legacy databases that still carry blobs in a `files` kanta field are migrated to disk by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which pops the field from the raw state before struct decoding. Uploaded files, seed assets and fetched external-site favicons live in the `FileStore` (in `files.py`): content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
## `data.py` ## `data.py`
@@ -20,13 +28,13 @@ msgspec Structs for the kanta database. See `docs/content-model.md` for the full
markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists, admon, gfm_autolink, sub/superscript plugins; typographer + breaks on). In bodies with at least three top-level h1/h2 headings (nested ones, e.g. inside `::: aside`, never participate), each gets a slug id (`python-slugify`, mirroring the editor's `slugify.js` — unicode folds to ASCII, separators become single hyphens) unless the author set `{#id}`, and their text is wrapped in a self-link (`a.anchor`) so section links are copyable; anchored headings also carry `data-line` with their markdown source line (the page editor's section pens and piecewise scroll sync key off it); the first in-body h1 is the article title — when the markdown has no h1, `render(title=...)` injects it as `# {title}` so implicit and explicit titles take the same path — it gets no id and doesn't count toward the three, its self-link is `href=""` (scroll to top); shorter articles stay anchor-free, h3+ is never navigable, and duplicates get `-2`/`-3` suffixes. Custom image rule: relative srcs resolve against the page path; an image standing alone in its paragraph becomes a figure (captioned when titled), while inline-with-text images and raw `<img>` HTML stay plain. A `{dates}` line expands to the article's published/updated dateline (`p.dateline`, from `Node.created`/`modified`; left literal in previews of unsaved pages). Code fences take pandoc-style brace attributes on the info line (` ```{.python .wide #id key=val} ` — the first class is the language when no bare language word precedes the braces) as well as a trailing `{...}` line; both land on the `<pre>`, the `<code>` keeps only the language class. markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists, admon, gfm_autolink, sub/superscript plugins; typographer + breaks on). In bodies with at least three top-level h1/h2 headings (nested ones, e.g. inside `::: aside`, never participate), each gets a slug id (`python-slugify`, mirroring the editor's `slugify.js` — unicode folds to ASCII, separators become single hyphens) unless the author set `{#id}`, and their text is wrapped in a self-link (`a.anchor`) so section links are copyable; anchored headings also carry `data-line` with their markdown source line (the page editor's section pens and piecewise scroll sync key off it); the first in-body h1 is the article title — when the markdown has no h1, `render(title=...)` injects it as `# {title}` so implicit and explicit titles take the same path — it gets no id and doesn't count toward the three, its self-link is `href=""` (scroll to top); shorter articles stay anchor-free, h3+ is never navigable, and duplicates get `-2`/`-3` suffixes. Custom image rule: relative srcs resolve against the page path; an image standing alone in its paragraph becomes a figure (captioned when titled), while inline-with-text images and raw `<img>` HTML stay plain. A `{dates}` line expands to the article's published/updated dateline (`p.dateline`, from `Node.created`/`modified`; left literal in previews of unsaved pages). Code fences take pandoc-style brace attributes on the info line (` ```{.python .wide #id key=val} ` — the first class is the language when no bare language word precedes the braces) as well as a trailing `{...}` line; both land on the `<pre>`, the `<code>` keeps only the language class.
`render()` returns a `Rendered(html, multicol)`: the article content segmented for the column layout (there is no wrapper div — segments and bare blocks are direct `<article>` children) — h1/h2 headings, `.wide` blocks and margin-breakout blocks (`.margin`, `::: aside`) stand bare, the runs between them become `<div class="colseg">` (plus `.cols` on segments with enough text in at least two paragraphs or one long enough to split across columns, `::: nocols` opting out; in column segments, paragraphs past `BREAKABLE_TEXT` visible characters are marked `.breakable` so they may split across columns), and `multicol` flags bodies long enough to columnize (visible-text thresholds, code excluded). `views.py` puts the class on the article; pagerite.css takes it from there (at most two columns, the left-margin breakout, all viewport adaptation). `render()` returns a `Rendered(html, multicol)`: the article content segmented for the column layout (there is no wrapper div — segments and bare blocks are direct `<article>` children) — h1/h2 headings and `.wide` blocks stand bare, the runs between them become `<div class="colseg">` (margin-breakout boxes — `.margin`, `::: aside` stay inside the segment at their anchor point; the CSS positions them out of flow into the side zone) (plus `.cols` on segments with enough text in at least two paragraphs or one long enough to split across columns, `::: nocols` opting out; in column segments, paragraphs past `BREAKABLE_TEXT` visible characters are marked `.breakable` so they may split across columns), and `multicol` flags bodies long enough to columnize (visible-text thresholds, code excluded). `views.py` puts the class on the article; pagerite.css takes it from there (at most two columns, the left-margin breakout, all viewport adaptation).
## `views.py` ## `views.py`
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering. The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`SITE_URL``https://<hostname>` from the CLI hostname argument; on localhost the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. The page title is injected as `# {title}` when the markdown has no h1 of its own, so it never appears twice (it always supplies `<title>` and nav labels). Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`SITE_URL``https://<hostname>` from the CLI hostname argument; on localhost the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. Additionally `twitter:image` pins extension-less `/_f/{hash}` share images to the `.webp` variant — X only honors WebP via twitter:image (not og:image) and its scraper cannot be trusted to negotiate via Accept. The page title is injected as `# {title}` when the markdown has no h1 of its own, so it never appears twice (it always supplies `<title>` and nav labels).
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), rendered only from the second level down — main-level pages list their children as cards after the content instead. Below that, the sidebar renders when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, main-level pages, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap). The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), rendered only from the second level down — main-level pages list their children as cards after the content instead. Below that, the sidebar renders when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, main-level pages, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
@@ -34,4 +42,4 @@ Any page with published children — a category page — lists them as a card gr
## `seed.py` ## `seed.py`
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `app.py`. Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `state.py`.
+3 -3
View File
@@ -8,13 +8,13 @@ The site structure is stored in the kanta database managed by `pagerite/data.py`
`Node.content` is the Markdown page, or None for a pure category label whose URL renders a 404 listing its children as cards (while nav links to it point at its first child); every label's title and slug are editable. A page with published children — a category page — lists them as cards after its markdown content; the sidebar sub-navigation renders only from the second level down, never on main-level pages. `Node.content` is the Markdown page, or None for a pure category label whose URL renders a 404 listing its children as cards (while nav links to it point at its first child); every label's title and slug are editable. A page with published children — a category page — lists them as cards after its markdown content; the sidebar sub-navigation renders only from the second level down, never on main-level pages.
Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `Data.pages` (pre-tree databases) migrates into `menu` on startup. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`. Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `pages` (pre-tree databases) migrates into `menu` via `migrate_v1`. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`.
`Data.version` is bumped on every write and embedded in page ETags so nav-affecting changes invalidate caches. Every content/settings write calls `_invalidate_pages()` in state.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
## Files ## Files
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching; the `FileStore` in app.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`). Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in files.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
## Banners ## Banners
+4 -4
View File
@@ -19,8 +19,8 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
- Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.). - Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
- **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above. - **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they float in the side zone at the article's left (the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation. - Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's left (the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/{hash}.ext`. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/….avif "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, floating in the side zone left of the text on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section. - **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/… "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone left of the text on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
## Page structure and navigation ## Page structure and navigation
@@ -34,7 +34,7 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
## Reading experience ## Reading experience
- The article column is sized by the **viewport, never by content**: a symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps the layout stable across navigation. The sidebar occupies the left gutter, the right gutter balances it. Long articles (flagged `.multicol` by the backend render) lift the cap and become a bounded **composition**, centered in the available space with the surplus left vacant: a fluid text lane (up to 42rem) plus a 16rem **side zone at the article's left** — the region the nav sidebar overlays — which hosts margin boxes (`.margin`, `::: aside`, margin figures) at all but phone widths, without the text ever moving. On pages with a sidebar, the sidebar gets its own track at every width — flexible, 12rem when space is tight and growing up to 150% (18rem) once the viewport has room beyond the article, the sidebar keeping its left side on the viewport's edge — and the track is the left lane instead: no in-article zone, the text lane runs fluid up to 86rem leaning on the viewport's right edge (surplus extends the left lane), and the boxes hang into the lane off the article's left border (growing leftward with it, up to 18rem), sliding under the translucent sticky nav. Once two lanes fit beside the zone (≥96rem available in `main`), the text flows in two fluid lanes (36rem minimum, capped at 102rem total — technical content wants the wider lanes, and wider windows just add vacant space). The stages step by the space actually available in `main` (container queries + `cqw` units, so the docked editor's inset is automatic). `.wide` figures on multicol pages bleed to the viewport edges measured from `main` (`cqw`), sliding under the sidebar. The backend splits the body into `.colseg` segments at h1/h2 headings, `.wide` elements and margin blocks (full-width separators or margin boxes, never inside columns), tagging segments that hold enough text in at least two paragraphs (or one long enough to split) with `.cols` — code blocks are excluded from that measure, a `::: nocols` container opts its whole section out, and column-filling paragraphs are marked `.breakable` so they may split across the column gap (shorter paragraphs stay whole). On wide single-column pages (≥104rem), margin boxes lean into the vacant left gutter as well, growing with it up to 18rem. - The article column is sized by the **viewport, never by content**: a symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps the layout stable across navigation. The sidebar occupies the left gutter, the right gutter balances it. Long articles (flagged `.multicol` by the backend render) lift the cap and become a bounded **composition**, centered in the available space with the surplus left vacant: a fluid text lane (up to 42rem) plus a 16rem **side zone at the article's left** — the region the nav sidebar overlays — which hosts margin boxes (`.margin`, `::: aside`, margin figures) at all but phone widths, without the text ever moving. On pages with a sidebar, the sidebar gets its own track at every width — flexible, 12rem when space is tight and growing up to 150% (18rem) once the viewport has room beyond the article, the sidebar keeping its left side on the viewport's edge — and the track is the left lane instead: no in-article zone, the text lane runs fluid up to 86rem leaning on the viewport's right edge (surplus extends the left lane), and the boxes hang into the lane off the article's left border (growing leftward with it, up to 18rem), sliding under the translucent sticky nav. Once two lanes fit beside the zone (≥96rem available in `main`), the text flows in two fluid lanes (36rem minimum, capped at 102rem total — technical content wants the wider lanes, and wider windows just add vacant space). The stages step by the space actually available in `main` (container queries + `cqw` units, so the docked editor's inset is automatic). `.wide` figures on multicol pages bleed to the viewport edges measured from `main` (`cqw`), sliding under the sidebar. The backend splits the body into `.colseg` segments at h1/h2 headings and `.wide` elements (full-width separators, never inside columns); margin boxes stay inside the segment at their anchor point and the CSS takes them out of flow — absolutely positioned off the article's left border into the zone, the columns flowing through unaffected — tagging segments that hold enough text in at least two paragraphs (or one long enough to split) with `.cols` — code blocks are excluded from that measure, a `::: nocols` container opts its whole section out, and column-filling paragraphs are marked `.breakable` so they may split across the column gap (shorter paragraphs stay whole). On wide single-column pages (≥104rem), margin boxes lean into the vacant left gutter as well, growing with it up to 18rem.
- A gentle **scroll-reveal** of headings, figures and block-level elements (IntersectionObserver). It is layout-level: articles need no support for it, and `prefers-reduced-motion` disables all motion. - A gentle **scroll-reveal** of headings, figures and block-level elements (IntersectionObserver). It is layout-level: articles need no support for it, and `prefers-reduced-motion` disables all motion.
## Styling ## Styling
@@ -48,7 +48,7 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
- **Page mode** — the 🖊️ next to a page's heading (including 404s, which is how new pages start) opens a CodeMirror Markdown editor docked to the left of the article: the panel is fixed to the viewport's left edge (its top tracks the banner's bottom until the banner scrolls away), the content shifts right and the sidebar hides while editing. Preview renders server-side per keystroke (no debouncing) and swaps the whole visible article content in one go (the edit pen and category cards survive the swap). - **Page mode** — the 🖊️ next to a page's heading (including 404s, which is how new pages start) opens a CodeMirror Markdown editor docked to the left of the article: the panel is fixed to the viewport's left edge (its top tracks the banner's bottom until the banner scrolls away), the content shifts right and the sidebar hides while editing. Preview renders server-side per keystroke (no debouncing) and swaps the whole visible article content in one go (the edit pen and category cards survive the swap).
- **Site mode** — the ⚙️ at the top right (after the 📊 analytics link, before login) opens a panel with the site **brand** (applied to the header live), a **theme** selector (swapping the theme stylesheet in place), a **page transition** selector (`cube`/`crossfade`, swapping `#pagerite-transition` in place), **font** picks (heading/body/brand — stored as plain `:root` rows inside the custom CSS, referencing the base stylesheet's per-family font variables), a **site-wide custom CSS** field (injected into `<style id="pagerite-user">` in the live page head and swapped during fetch-navigation), the page's **banner design** selector (inherit / none / any design found on disk, inherited by children), the page's **banner HTML** field (supplementing the design, previewed into the real banner region, so you see exactly which banner you're editing) and the **structure tree**. Everything saves immediately as you edit — no save button, no edit mode. - **Site mode** — the ⚙️ at the top right (after the 📊 analytics link, before login) opens a panel with the site **brand** (applied to the header live), a **theme** selector (swapping the theme stylesheet in place), a **page transition** selector (`cube`/`crossfade`, swapping `#pagerite-transition` in place), **font** picks (heading/body/brand — stored as plain `:root` rows inside the custom CSS, referencing the base stylesheet's per-family font variables), a **site-wide custom CSS** field (injected into `<style id="pagerite-user">` in the live page head and swapped during fetch-navigation), the page's **banner design** selector (inherit / none / any design found on disk, inherited by children), the page's **banner HTML** field (supplementing the design, previewed into the real banner region, so you see exactly which banner you're editing) and the **structure tree**. Everything saves immediately as you edit — no save button, no edit mode.
- Clicking a pen again closes the editor (without saving; a dirty preview reloads the page). The pens are `<button>`s wired up by `pagerite.js` — editing is an action, not a navigation. The editor's WebSocket **reconnects automatically** with local text and pending saves preserved. (All users are trusted authors for now; access control later with SSO.) - Clicking a pen again closes the editor (without saving; a dirty preview reloads the page). The pens are `<button>`s wired up by `pagerite.js` — editing is an action, not a navigation. The editor's WebSocket **reconnects automatically** with local text and pending saves preserved. (All users are trusted authors for now; access control later with SSO.)
- **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published controls. Images can be pasted straight into the editor or chosen via a file input: they upload to the content store (`PUT /_api/files/...`) and insert `![alt](/_f/hash.ext)` at the cursor. - **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published controls. Images can be pasted straight into the editor or chosen via a file input: they upload to the content store (`PUT /_api/files/...`) and insert `![alt](/_f/hash)` at the cursor.
- The **structure panel** (vue-draggable tree of the whole site, in site mode) covers page management: reorder any menu level, drag across sections, add, delete (two clicks: the button arms, then deletes — no dialogs). Every node is a real label — content-less category rows offer a to give them a landing page. Deleting a category removes only its landing page (the label and its subpages stay). Every non-empty list ends with a row that starts a new page as a local-only tree row at that level; the row can be dragged into place before its title and slug are filled in and is persisted only on commit. While dragging, these rows double as "end of this list" drop targets; dropping ON the lower part of a row makes the page that row's first child (even a leaf's, creating a sublist), while a row's exposed top edge inserts a sibling before it. A dragged row's indentation previews the target list's depth. Rows are always editable: titles save while typing, slug edits commit on blur/Enter since they rename the path (moving the whole subtree). The front page is the root row with an empty slug — renaming it away leaves no front page ("/" redirects to the first nav item), and giving another top-level row the empty slug makes it the front page. - The **structure panel** (vue-draggable tree of the whole site, in site mode) covers page management: reorder any menu level, drag across sections, add, delete (two clicks: the button arms, then deletes — no dialogs). Every node is a real label — content-less category rows offer a to give them a landing page. Deleting a category removes only its landing page (the label and its subpages stay). Every non-empty list ends with a row that starts a new page as a local-only tree row at that level; the row can be dragged into place before its title and slug are filled in and is persisted only on commit. While dragging, these rows double as "end of this list" drop targets; dropping ON the lower part of a row makes the page that row's first child (even a leaf's, creating a sublist), while a row's exposed top edge inserts a sibling before it. A dragged row's indentation previews the target list's depth. Rows are always editable: titles save while typing, slug edits commit on blur/Enter since they rename the path (moving the whole subtree). The front page is the root row with an empty slug — renaming it away leaves no front page ("/" redirects to the first nav item), and giving another top-level row the empty slug makes it the front page.
- Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a stateless JSON protocol (`open`/`render`/`save`; on save all fields are optional and absent ones keep their old values, `move_from` renames), avoiding REST polling and races. Rendering always stays server-side. - Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a stateless JSON protocol (`open`/`render`/`save`; on save all fields are optional and absent ones keep their old values, `move_from` renames), avoiding REST polling and races. Rendering always stays server-side.
- A REST API also exists for scripting, all under `/_api/`: `GET pages` (the full tree), `PUT/DELETE pages/{path}`, `GET/PUT settings` (site brand, theme and custom CSS), `POST structure` (reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`. - A REST API also exists for scripting, all under `/_api/`: `GET pages` (the full tree), `PUT/DELETE pages/{path}`, `GET/PUT settings` (site brand, theme and custom CSS), `POST structure` (reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`.
+10 -5
View File
@@ -4,17 +4,22 @@ The Vue editor is a single tabbed `EditorShell.vue` mounted in a host div create
## Tabs ## Tabs
The shell hosts four kept-alive tabs (ordered site-wide first — site, structure — then, after a visual break, the per-page tabs — article, banner): The shell hosts five kept-alive tabs (ordered site-wide first — site, structure, localization — then, after a visual break, the per-page tabs — article, banner):
- `PageEditor.vue` — CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`, previewing into the visible article; editor and article scrolls are linked piecewise-linearly, keyed on the section anchors' `data-line` (markdown source line the backend stamps on top-level anchored h1/h2s): the page follows the cursor (fractional, wrap-aware, anchored at a fixed window height), the editor follows page scroll with a progress-based viewport anchor, applied instantly (the window keeps scrolling normally while any editor is open — the panel is fixed to the viewport's left edge, its top tracking the banner's bottom edge until the banner scrolls away — and the panel scrolls internally); anchored h2s carry their own edit pens that open the editor scrolled to that section; a format bar offers Markdown helpers — bold/italic/code/link/table/image upload, with Ctrl/Cmd-B/I/S bindings — for the hard-to-remember syntax. Edits content and title only, never the path. - `PageEditor.vue` — CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`, previewing into the visible article; editor and article scrolls are linked piecewise-linearly, keyed on the section anchors' `data-line` (markdown source line the backend stamps on top-level anchored h1/h2s): the page follows the cursor (fractional, wrap-aware, scrolling only when the cursor's page position leaves the viewport, with an edge margin), the editor follows page scroll with a progress-based viewport anchor, applied instantly (the window keeps scrolling normally while any editor is open — the panel is fixed to the viewport's left edge, its top tracking the banner's bottom edge until the banner scrolls away — and the panel scrolls internally); anchored h2s carry their own edit pens that open the editor scrolled to that section; a format bar offers Markdown helpers — bold/italic/code/link/table/image upload (always block-level on a fresh blank-separated line of its own — a cursor on a non-empty line, e.g. inside an existing image tag, inserts after that line, never into it; always with an empty `""` caption, cursor inside the quotes), toggling fences (` ``` ` code blocks and `::: aside` containers share the same machinery: clicked inside one they remove it and select the content, otherwise they wrap the selection or the cursor's line, keeping it selected), and `.left`/`.right`/`.wide`/`.margin` placement toggles plus `.small`/`.large`/`.huge` text-size toggles (brace attributes on the block at the cursor, mutually exclusive within each group; on `:::` containers a placement class replaces the container name instead — `::: aside``::: margin`), with Ctrl/Cmd-B/I/S bindings — for the hard-to-remember syntax. Edits content and title only, never the path.
- `BannerEditor.vue` — per-page banner HTML + banner design selector, previewed into `#page-banner`. - `BannerEditor.vue` — per-page banner HTML + banner design selector, previewed into `#page-banner`.
- `SiteEditor.vue` — site brand + optional custom brand HTML with image/video upload + theme selector + page-transition selector + font picker + favicon upload — clicking the preview tile picks a new one — + site-wide custom CSS, CSS injected into `<head id="pagerite-user">`. - `SiteEditor.vue` — site brand + optional custom brand HTML with image/video upload + theme selector + page-transition selector + font picker + favicon upload — clicking the preview tile picks a new one — + site-wide custom CSS, CSS injected into `<head id="pagerite-user">`.
- `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row. - `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row, plus a per-row flag dropdown setting the page's primary language (`Node.language`, inherited by the subtree).
- `LocalizationEditor.vue` — the site-wide translation settings: target languages as a flag grid (toggles, grouped in geographic rows; see docs/localization.md), the refresh-all-translations button, and the translator service WebSocket URL(s) to connect `scripts/translator.py` to.
Media uploads everywhere use the image icon buttons (pasting into the editor works too). The article, banner and site-settings pens are shorthands that open the shell on the matching tab; once open, clicking a pen switches tabs (and retargets the editors to the current page) instead of closing/remounting. The close button in the tab bar closes the shell (Escape too); tabs have no close buttons of their own. Closing only HIDES the shell — the Vue app stays mounted, so page-editor state (unsaved text included) survives until a real page reload; the editor always follows the URL, so fetch-navigating with the shell open (or before re-opening it) retargets it to the new page — unsaved text is stashed per path for the session and restored when returning, cleared on save. Saving there is explicit (Ctrl+S) and refreshes the page regions in place. Admin panels never reload the page. Media uploads everywhere use the image icon buttons (pasting into the editor works too). The article, banner and site-settings pens are shorthands that open the shell on the matching tab; once open, clicking a pen switches tabs (and retargets the editors to the current page) instead of closing/remounting. The close button in the tab bar closes the shell (deliberately NOT Escape — it fired too easily by accident); tabs have no close buttons of their own. Closing only HIDES the shell — the Vue app stays mounted, so page-editor state (unsaved text included) survives until a real page reload; the editor always follows the URL, so fetch-navigating with the shell open (or before re-opening it) retargets it to the new page — unsaved text is stashed per path for the session and restored when returning, cleared on save. Saving there is explicit (Ctrl+S) and refreshes the page regions in place. Admin panels never reload the page.
In-place page re-rendering shared by the banner/site/structure tabs lives in `swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic regions, replaceState). It also exports `dropPageCache`, which the editor tabs call after any save that can alter the rendered HTML of other pages (theme, headings, structure, banners, site brand/CSS, favicon). Dropping the cache while editing avoids re-fetching every page immediately; the public runtime re-preloads visible links once the editor panel closes. In-place page re-rendering shared by the banner/site/structure tabs lives in `swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic regions, replaceState). It also exports `dropPageCache`, which the editor tabs call after any save that can alter the rendered HTML of other pages (theme, headings, structure, banners, site brand/CSS, favicon). Dropping the cache while editing avoids re-fetching every page immediately; the public runtime re-preloads visible links once the editor panel closes.
The page and structure tabs share one language selector: `LangSelect.vue` (small flag + dropdown) v-modeled on the shell-wide selection in `editorLang.js` (`''` = primary). While the panel is open that selection overrides the page's normal language preferences: EditorShell calls `swapdoc.setLangOverride`, which pins every `loadPlain` fetch (`?lang=`, the primary by its own code) and pagerite.js's own fetches/prefetches (`pagerite:session-lang`), until the panel closes and the override clears.
All WebSockets (page/banner editors, analytics view, the pagerite.js activity channel) pace their connections through `reconnect.js`: new sockets are created a staggered slot apart (a page load opens Vite's HMR socket plus several of ours at the same moment, and such bursts — like rapid retries — trip the browser's WebSocket throttling, leaving every socket to the host "pending" for minutes), a watchdog closes sockets stuck CONNECTING so they reschedule instead of hanging forever, and retries follow an exponential backoff with jitter that only a healthy connection resets. While a socket is connecting or waiting to reconnect the panel says so (`ConnNote.vue`), and the CodeMirror editors stay locked until their document arrives (typing before the doc accept would be clobbered by it).
## Saving behavior ## Saving behavior
Everything saves immediately as you edit (brand/title/CSS debounced, slug on commit since it renames the path), theme change swaps the stylesheet in place, tree rows navigate in place without transitions when focused, and the front page is a root-only row whose empty slug is editable like any other. Saves that can affect other pages drop the prefetch cache; the cache is rebuilt when the editor panel closes so navigation stays instant. Everything saves immediately as you edit (brand/title/CSS debounced, slug on commit since it renames the path), theme change swaps the stylesheet in place, tree rows navigate in place without transitions when focused, and the front page is a root-only row whose empty slug is editable like any other. Saves that can affect other pages drop the prefetch cache; the cache is rebuilt when the editor panel closes so navigation stays instant.
@@ -25,6 +30,6 @@ Dropping ON the lower part of a row moves the page under that row (the child lis
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`. The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme, banner-design and transition CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-transition`/`#pagerite-user` elements to restore the canonical order (base < theme < design < transition < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod. `vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js`, `src/pagerite.js` and `src/analytics-main.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme, banner-design and transition CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-transition`/`#pagerite-user` elements to restore the canonical order (base < theme < design < transition < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin. `vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
+1 -1
View File
@@ -8,7 +8,7 @@ Vue editor app entry, mounts the tabbed `EditorShell`. See `docs/editing.md` for
## `pagerite.js` ## `pagerite.js`
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire). Editors can drop the entire cache with the `pagerite:drop-page-cache` event when site-wide or page changes (theme, headings, structure, banners, etc.) invalidate the cached HTML of other pages; `main.js` triggers a fresh `pagerite:preload-pages` pass when the editor panel closes so navigation is fast again. Navigation that starts while the editor is open bypasses the cache and fetches the target page on demand. Also runs scroll-reveal, a scroll-driven section hash (the location hash tracks the h1/h2 above the viewport middle via replaceState — removed above the first tagged heading and at the very top, never set on unscrollable pages), OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, and the auth check. Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire). Editors can drop the entire cache with the `pagerite:drop-page-cache` event when site-wide or page changes (theme, headings, structure, banners, etc.) invalidate the cached HTML of other pages; `main.js` triggers a fresh `pagerite:preload-pages` pass when the editor panel closes so navigation is fast again. Navigation that starts while the editor is open bypasses the cache and fetches the target page on demand. Also runs scroll-reveal, a scroll-driven section hash (the location hash tracks the h1/h2 above the viewport middle via replaceState — removed above the first tagged heading and at the very top, never set on unscrollable pages), OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, click-to-enlarge on article figure images (a full-viewport lightbox with the caption, closed by click or Esc), and the auth check.
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`. It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`.
+483
View File
@@ -0,0 +1,483 @@
# Localization
Pages are served in the visitor's language based on a `?lang=` query
parameter or the `Accept-Language` header.
- **Phase 1 (implemented):** negotiation, URL scheme, caching, rendering
plumbing. Translations are consumed through a stub interface; the database
still holds only the original language.
- **Phase 2 (implemented):** gettext-style fragment storage in the
database — machine-translated chunks plus user override patches, assembled
at render time. Storage details in `docs/migrate.md`.
## Phase 1: negotiation and URLs
### The primary language
Each article has a primary (original) language: `Node.language`, inherited
down the tree like `banner` — "" = the nearest ancestor's, the front page
last (it doubles as the site default), with `en` as the final fallback
(`ORIGINAL_LANGUAGE`, `primary_lang()` in `pagerite/i18n.py`). It is
configured per row in the structure editor. Everything per-article keys
off the resolved value: language selection, `<html lang>`, canonical URLs,
what counts as a translation, and the translation targets (a node's own
primary is never one — so the target set may include the site default, and
a page in another language can be translated into it).
### Language selection
Deliberately simple — **q-values are ignored**:
- All known `Accept-Language` implementations send the header **in order of
preference**, so we parse it as an ordered list and never reorder.
- Selection rule (`select_language` in `pagerite/i18n.py`):
1. If `?lang=<tag>` is present, use it (if a translation exists; otherwise
fall through to header logic).
2. If the article's original language appears anywhere in the header list,
use the **original**. Rationale: an AI translation is strictly worse
than the original for anyone who has that language configured at all
(e.g. `fi-FI, fi, en-US, en` gets English, not machine-translated
Finnish).
3. Otherwise walk the header list in order and use the first language for
which a translation exists.
4. Fall back to the original.
Region tags normalize to their base subtag (`fi-FI``fi`).
### URLs: pretty for users, indexable for search engines
- Canonical URLs stay pretty (`/some-page`). Each language version is
addressable as `/some-page?lang=fi` so search engines can index them.
- `<link rel="canonical">` names the **actually served language**: the plain
URL when serving the original (for SEO the non-query URL means the
article's own language), `?lang=xx` when serving a translation — however
the language was arrived at (query or header).
- `<link rel="alternate" hreflang="…">` entries follow the canonical
directly (before the social meta tags) and are the same set on every
page — the site-wide configured languages (`translate_langs`, which the
translator works to fill in): `x-default` first, pointing at the plain
autodetecting URL, then every language explicitly with `?lang=`, the
page's own primary language included.
- The override sticks for the session of clicks: a page requested with
`?lang=` replicates the query onto the navigation links it renders (nav,
sidebar, cards, brand — in-article links are content and stay as
authored), so plain clicks and no-JS navigation keep the language.
pagerite.js additionally strips the query from the address bar via
`history.replaceState` (pretty, shareable URLs), remembers the language,
and adds it to every internal fetch that lacks one (preloads,
fetch-navigations, history traversals); history entries stay query-less.
- A full page refresh or a shared link resets to automatic selection (header
only). This gives a clean one-time override without cookies.
### Response correctness
- Content responses carry `Vary: accept-language` (added to the existing
`accept-encoding` vary).
- `_cached_body` and the page ETag include the **selected language** (not the
raw header, which would blow up the cache key space) and the **replicated
link language**: a `?lang=fi` render and a header-selected Finnish render
of the same page differ in their navigation links, so they are cached as
separate variants.
- `<html lang="…">` reflects the served language, and an RTL language
(`i18n.RTL_LANGUAGES` — ar, fa, he, ur) also sets `dir="rtl"` on `<html>`
(the editor panel carries its own `lang="en" dir="ltr"` so it stays LTR).
Client-side page swaps (fetch navigation in pagerite.js, editor re-renders
in swapdoc.js) copy both attributes from the fetched document, so a hot
switch into or out of an RTL page flips the layout without a reload.
### Rendering
- The translated Markdown goes through the same `markdown.render` pipeline.
- Navigation/sidebar titles come from the translation's title map, with
per-node fallback to the original title (a partially translated tree must
still render).
- Category placeholder pages (the 404s for content-less labels) select a
language like content pages, but over the **subtree's** combined
availability (`subtree_languages`) — they have no chunks of their own;
the heading, navigation and card text localize from the title map and
the target articles' translations.
- Card descriptions and cover picks run on the target article's hybrid
Markdown where that page is available in the served language, with
per-card fallback to the original.
- Fixed UI strings ("Not Found" etc.) and the editor UI stay English for now.
- The markdown typographer (SmartyPants) is English-centric; per-language
typographer options are a possible follow-up, not blocking.
## Phase 2: fragment-based translation storage (implemented)
Phase 1 assumed whole-page translated Markdown delivered from outside. The
refined model is gettext-style: an article has **one primary version** (its
`content`, in its own language) plus, per target language, **machine
fragments** (translated chunks of Markdown) and **user patches** (minimal
editor overrides). Both are stored in the database and assembled into the
served Markdown at render time.
### The scenario this must handle
1. Article written in English.
2. Machine-translated into Spanish → fragments stored.
3. Editor fixes one Spanish paragraph and changes a link elsewhere to point
at a Spanish resource → user patch hunks stored.
4. English article edited → the edited chunk's key changes; its Spanish
fragment no longer matches.
5. Page requested before the machine translation refreshes → served as a
**hybrid**: old fragments for unchanged chunks, plain English for the
edited chunk. User patches are attempted against this hybrid, best effort,
each hunk independently: the text fix is stale (its search text no longer
exists) and silently skipped; the link change still applies even though
the link sits in the now-English paragraph.
6. Machine translation refreshes → full Spanish again, with both patch hunks
applying.
### Chunks
`chunk_markdown(markdown)` splits the source into block-level chunks —
blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
list blocks, tables, HTML blocks. A chunk's identity is its **source text**,
gettext-msgid style:
```python
chunk_key = blake3(normalize(chunk_text)).digest(9) # bytes; base64 at the JSON level
```
(`normalize`: strip trailing whitespace per line, collapse surrounding blank
lines — so whitespace-only source edits don't invalidate translations.)
Consequences:
- Editing the English source invalidates exactly the edited chunks; all
other fragments keep applying. Stale fragments are simply never referenced
again and can be garbage-collected lazily (or left; they are tiny).
- No explicit "source version" bookkeeping is needed — staleness falls out
of the keys.
### User patches
Editors always edit **full Markdown** in the existing editor UX — never
fragments. When editing a translated view (`?lang=es`), the editor is loaded
with the *current hybrid Markdown*; on save, the server computes a minimal
diff against that hybrid and stores it as a patch:
```python
class Patch(msgspec.Struct, omit_defaults=True):
"""One editing session's overrides, applied independently per hunk."""
hunks: list[tuple[str, str]] = [] # (search, replace) on hybrid Markdown
```
Hunks are produced from `difflib.SequenceMatcher` on the hybrid vs. the
edited text at block granularity: each `replace`/`delete`/`insert` opcode
becomes one `(search, replace)` pair, with the preceding block's tail as
left context for `insert` (pure inserts have empty search context otherwise).
Application is dead simple:
```python
def apply_patch(hybrid: str, patch: Patch) -> str:
for search, replace in patch.hunks:
if search and search in hybrid:
hybrid = hybrid.replace(search, replace, 1)
# missing search text = stale hunk -> silently skipped
return hybrid
```
Per-hunk independence is the robustness property from the scenario: a stale
text fix does not block a still-valid link change. Patches are stored as an
ordered list and applied in order.
### Storage
Full storage design and the `migrate_v3` restructuring live in
`docs/migrate.md`. The short version, as it concerns this document:
- Originals **and** translations are content-addressed text chunks in flat
stores: `Data.chunks: dict[bytes, str]` and
`Data.trans: dict[bytes, dict[str, str]]` (chunk hash → lang → text) —
path-independent, so repeated paragraphs and menu titles are translated
once and article moves touch nothing. `Node.chunks: list[bytes]` gives
each article its order.
- `Node` gains **`language: str = ""`**, inherited down the tree like
`banner` (empty = nearest ancestor, front page last, site default `en`
final). `select_language` and `<html lang>` use the resolved value instead
of the global `ORIGINAL_LANGUAGE` constant.
- **Known weakness:** changing a page's (or subtree's) `language` after
translations exist mis-keys everything — translations are keyed by
*source* chunks, so old entries silently stop matching and user patches
(searching for old-hybrid text) mostly go stale. That is acceptable:
the orphaned data is harmless and translations regenerate. We do not
migrate translations across a language change.
- Article paths are stored and keyed **without leading slashes**
(`"docs/setup"`, front page `""`); slashes are added only in hrefs.
### Render pipeline (the phase-1 `get_translation` stub, now real)
```python
def get_translation(data, path, lang) -> Translation | None:
if lang not in node.langs:
return None
hybrid = "\n\n".join(
chunks[h] if h in node.no_trans else trans.get(h, {}).get(lang, chunks[h])
for h in node.chunks
)
for patch in data.patches.get(f"{path}:{lang}", []):
hybrid = apply_patch(hybrid, patch)
return Translation(markdown=hybrid, titles=title_map(data, lang))
```
- Availability is an article-level index: `node.langs: dict[lang, True]`,
maintained by the translation writers (translator job, patch saves) in the
same transaction as their data writes — rendering and language selection
never probe the `trans` store chunk by chunk. A stale key is benign (the
"translation" just renders as the original).
- `titles` for nav/sidebar/cards: each node's translated title is
`trans.get(hash(node.title), {}).get(lang)` with per-node fallback — one
dict lookup per nav item at render time.
- Cache invalidation: writes to `chunks` / `trans` / `patches` (translator,
editor saves) call `_invalidate_pages()`, same as content writes.
### Editor flow
The page and structure editors share one language selector (`LangSelect.vue`:
a small flag button opening a dropdown; the same country-flag-icons set as
the analytics visitor cells), v-modeled on one shell-wide selection
(`editorLang.js`, `''` = the primary language). The page editor lists the
page's own primary language (`Node.language`, resolved through the
hierarchy and echoed in the WS doc as `primary_lang`) plus the union of
the page's translations (`node.langs`) and
the site-wide `translate_langs`; it always opens in the primary language,
even when the page itself was served in a translation. A note under the
toolbar states the blast radius:
edits to the primary language re-chunk the original (invalidating the
affected translation fragments everywhere); edits to a translation stay
local to that language.
While the editor panel is open, its language selection **overrides the
normal language preferences** for the page preview: EditorShell pins every
in-place re-render and pagerite.js fetch/prefetch to it (`?lang=` — a
primary selection pins by the current page's own resolved primary, which
`select_language` honors), and closing the panel restores the normal
preferences.
- WS `open` with a `lang` returns the effective **hybrid** Markdown and
title for that language (ungated by `node.langs` — a language without
any fragments yet starts from the original text), plus the language
metadata (`lang`, `primary_lang`, `langs`, `translate_langs`).
- The editor keeps a **shadow copy** of the Markdown it opened. WS `save`
with `lang` sends it as `base`; the server diffs `base` → submitted text
(`make_patch`) and appends a `Patch`. Diffing against the shadow (rather
than the current hybrid) keeps hunks correct when the original or the
machine translation moved under an open editor; application against the
then-current hybrid stays best-effort per hunk, as designed.
- A changed **title** on a translated save becomes a fragment in
`Data.trans` keyed by the original title's chunk hash — the same storage
as machine title translations. An untouched title field (holding the
served translation) is not sent, so saving never freezes a stale machine
title into an override.
- Saving never deletes; a translation additionally cannot be emptied (that
would render as a blank page in that language).
- The live preview renders the version being edited, whichever language
the page itself was loaded in (the render is just the edited Markdown +
title). A translated save keeps that preview in place — re-fetching the
page would come back in the header-selected language.
- Saving the primary-language version re-chunks the submitted Markdown and
updates `Data.chunks` / `node.chunks` — only genuinely new text lands in
the kanta change diff (see docs/migrate.md).
The **structure editor** selects from the same languages with the same
`LangSelect` (the selection is shared — switching in either tab switches
both, and the preview). It is also where a page's **primary language** is
configured: each row carries a small flag dropdown (the resolved flag,
dimmed while inherited) that sets `Node.language` via a structure op —
'' = inherit, so setting it on a section covers the whole subtree. The
tree it
lists (`GET /_api/pages?lang=`) comes back with per-language titles where a
translation exists (`translated` marks those rows; untranslated rows show
the original title, dimmed). Retitling in a non-primary language posts the
structure op with a `lang` and writes a per-language title fragment in
`Data.trans` (keyed by the original title's chunk hash, exactly like a
machine title translation — a user edit simply overwrites it); sending the
original's text drops the override. The structure itself — slugs,
hierarchy, order — is language-independent, so pending rows, slug edits,
drag-and-drop and deletes work identically in every language.
### Translator service API
An external machine-translation service connects over WebSocket at
`/_translate/{key}` — deliberately **not** under `/_api`: the SSO
forward-auth does not cover that route, and the key in the path is the
access control. Keys live in `Data.translate_keys` (key -> display name) —
12 lowercase alphanumeric characters each, the first one generated at
database bootstrap and multiple keys reserved for future management (e.g.
a web UI). The full WS URL(s) are printed in the startup log
(`ws://localhost:{port}/_translate/{key}` locally,
`wss://{hostname}/_translate/{key}` on a public hostname) and the keys are
surfaced to the admin in `GET /_api/settings` as `translate_keys`. An
unknown or empty key rejects the handshake (close-before-accept → HTTP
403). Transactions storing results record the connecting key as the kanta
transaction `user`.
Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
`bytes` fields ride as base64):
- `{"type": "hello", "langs": [...]}` — client greeting announcing its
**capabilities**: the language codes its model can produce (normalized
to base subtags; `en`/empty dropped).
- `{"type": "job", "lang", "key", "texts", "path", "kind", "contexts"}`
server push: ONE fragment to translate (an article title or a chunk), as
a list of **prose segments** (see Segmentation below). `contexts` is
parallel to `texts` ("" = none): the surround to translate the segment
in — for clients that translate better with context (see below).
Contexts are not part of the result.
- `{"type": "result", "lang", "key", "texts"}` — client reply: the
segments translated, same order and count, matching its job by (lang, key).
Which languages get translated is **server-configured**:
`Data.translate_langs` (presence-key dict, bootstrapped to Spanish and
Chinese — edited in the editor shell's localization tab, whose flag grid
lists every language including English, or set via `/_api/settings` as
`translate_langs`). A target equal to an article's own primary language is
skipped per article (its original already is that language), so the set
may freely contain the site default. The dispatcher offers a
connection jobs only in `wanted ∩ capable`; a connection without overlap
simply stays idle.
`DELETE /_api/translations` (the localization tab's "refresh all
translations" button) drops every machine translation (`Data.trans`) and
rebuilds the availability index (`node.langs`) from the surviving user
patches, so the dispatcher re-translates everything from scratch; the
run's validation skip-list is cleared with it, giving rejected fragments
another chance.
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; api.py only
registers the route):
- **One job at a time per connection** — the next job is sent only after
the current one's result. Clients wanting parallelism open multiple
connections (e.g. several `scripts/translator.py` instances).
- Pending work is derived from the `trans` store
(`translate.pending_items`) minus the items in flight on any connection,
so a **disconnect requeues** that connection's in-flight item and it is
offered to any free capable connection.
- Dispatch re-runs on every relevant event: Hello, result, disconnect and
content change (`_invalidate_pages()` schedules it, so the pass runs
after the writing transaction commits).
- A result with no job in flight, a mismatched (lang, key), a duplicate
hello, or any malformed frame closes the socket with a protocol error.
Results are stored into `trans` in one transaction and set
`node.langs[lang]` on every article they touch (shared chunks make several
pages gain a language from one fragment). Unknown keys are stored anyway
and re-storing overwrites — results are idempotent.
#### Segmentation
Fragments cross the wire as **prose segments** (`pagerite/segments.py`): the
fragment is parsed with the project's own markdown-it setup
(`markdown.make_md(verbatim=True)` — all extensions, but no typographer or
tasklist label wrapping, so token text stays byte-identical to the source)
and split into the runs a model may touch: paragraph/heading/table-cell text
(merged across soft line breaks), image alt texts and captions, footnote
bodies. A block of plain text, inline **links and paired text formatting**
(strong/em/s) **stays whole** — link and formatted texts cross inline, in
sentence context, with the Markdown stripped (see below). Everything else
never leaves the server: code spans and
fences, URLs and autolinks, link/image *destinations*, `{...}` spans
(placeholders like `{dates}` as well as attrs), reference and footnote
labels, container fences, GFM alert markers (`[!NOTE]`), raw HTML — and the
remaining markup punctuation (`|`, `:::`), which is a run boundary.
Chunks with no segments (a lone `{dates}`, container fences, pure
code/HTML) are never dispatched at all (`needs_translation`); every
language renders them from the original chunk. Each segment is accompanied
by a context string (a segment carved out of a larger block carries the
block's plain text; a whole-block segment carries "") — context is a
prompt aid only, never spliced into the result.
Reassembly is offset splicing, not text the model produced: each segment's
source span was located at dispatch (sequential search; a run that is not a
verbatim source substring — entity-decoded text, backslash escapes — is
skipped and stays in the original language), and the returned translations
are swapped in by offset. Markup corruption is therefore impossible by
construction; the failure modes that remain are a wrong segment count, an
empty segment, or markup injected INTO a segment (a `<br>` in a title
translation would splice live HTML) — each returned segment must parse as
pure prose, or the whole result is dropped and logged, and the (lang, key)
pair is skipped for the rest of the server run (generation is
near-deterministic, so an immediate retry would re-fail; the fragment stays
pending and gets another chance on restart or `DELETE /_api/translations`).
`Data.trans` therefore only ever holds clean translated Markdown.
Link- and formatting-carrying blocks are the one place a segment is not
spliced verbatim: a label translated apart from its sentence comes back
grammatically incompatible with it (case government, particles, word
order), and shown the Markdown the model mangles it (Seed-X dropped the
`**` and the glued-on colon in `**Pagerite**: …`), so the block crosses
whole — all Markdown stripped — and the server re-inserts the link and
formatting syntax into the translated block. The boundaries are found by
**text processing alone**
markers on the wire are hopeless (an earlier sentinel-masking design let
the model see and mangle exactly that punctuation: Seed-X renumbered the
tokens and turned `![` into `¡¡…!!`). Each mark's source words are aligned
to the translation's words by **form similarity** (`_find_mark` in
segments.py): a word-level alignment (sequence ratio plus shared prefix,
case-folded — inflection moves word endings, `banana``banaanilla`, and
articles or prepositions drop out; a mid-sentence capital on BOTH sides
earns a bonus, naming conventions being the likeliest shared cause) with
small penalties for skipped words, so reordering and dropped function
words don't break the match. An alignment is accepted only with an anchor
(one pair of similarity ≥ 0.7) and a decent average, and the slice is cut
exactly at word boundaries, so the whitespace between the mark and its
neighbors stays in the plain text. A mark with no convincing alignment
falls back to its weight ratio in the source block (word units before its
text boundaries over the block total) applied to the translation's units
— CJK ideographs count as one unit each, kana runs as one; for CJK targets
the fallback IS the path, cross-script form similarity being nil. Placement is
approximate and several reordered marks in one block can still cluster —
the accepted trade: better a coherent sentence with a slightly shifted link
than separately translated snippets that don't fit together. A boundary
that maps to an empty slice degrades to the source text rather than
emitting a broken `[](url)` or `**`. Blocks mixing in any other inline
markup (code spans, images, raw HTML) don't qualify and still split into
runs at those boundaries.
Punctuation is the translator's own job: Seed-X tends to "finish" short
labels (titles, nav items) with a comma or period the source never had.
Prompt wording is NOT the fix — a punctuation-instruction clause made
Seed-X slip into its `[COT]` reasoning mode (minutes-long generations with
reasoning text in the output, observed for Chinese). The reference client
enforces punctuation deterministically instead (`match_punctuation` in
scripts/translator.py): a translation of a segment without terminal
punctuation gets any added trailing marks (and a newly opened Spanish ¡/¿)
stripped before the result goes back.
The same client-side enforcement covers markup bleed as a CLASS, not per
artifact: `<` is the prose/markup boundary on the wire and never appears in
a segment in either direction. Source pieces containing `<` are never
dispatched (they stay in the original language — segments.py), and the
reference client cuts the model's output at the first `<`
(scripts/translator.py) — echoed language tags, stray `<br>`s and any
future variant are one handled case. (The cut is post-decode, not a
generation stop string: Seed-X opens every generation with its `<s>`
framing token, which would trip a `<` stop immediately.)
Short fragments get more than a bare prompt: each segment may carry its
surround in `Job.contexts` — a title carries the article's opening prose
(its own block is just the title word), a segment carved out of a larger
block (a partial run; a link text whose block didn't qualify for the
whole-block treatment) carries the block's plain text, and a
whole-block segment (a plain paragraph) is self-contextualizing and carries
"". The reference client translates segment and surround together, stops
generation at the blank line separating them, and keeps the segment's own
part of the output (its line resp. paragraph; a hard-break `␣␣\n` separator
works too). If the model merged them (no separator, or an empty first
part), it falls back to translating the segment alone. The surround fixes
context-free readings ("About" as "approximately" — with the opening it
becomes "Tietoa"/"Acerca de"; "here" as "就在这里" → the idiomatic
"点击这里") and, as a side effect, most stray trailing punctuation.
### Explicitly out of scope for phase 2
- The machine translation itself: the API above moves fragments in and out;
the translating is external. `scripts/translator.py` is the reference
client (Seed-X-PPO-7B only — its 28 languages are the ceiling).
- Garbage collection of orphaned chunks/translations (see docs/migrate.md).
- sitemap.xml per-language entries; translated UI chrome; per-language
typographer options; multi-locale date/number formatting.
+178
View File
@@ -0,0 +1,178 @@
# migrate_v3: content-addressed chunk storage
Status: **implemented**. `migrate_v3` restructures how article text and
translations are stored, motivated by the localization model in
`docs/localization.md` (phase 2). Since it is a full migration, it is free to
break the current `Node.content: str | None` layout.
## Goals
- **Minimal change diffs.** kanta persists change diffs; editing one
paragraph of a long article must not rewrite the whole article string, and
a translation refresh must touch only the re-translated chunks.
- **Fast, simple lookup.** Everything heavy lives in flat
`dict[hash, content]` stores; ordering lives in `list[hash]`. No large
nested structures, no deep paths.
- **Path-independent text.** Chunks and their translations are keyed by
content hash, not by article path — the same paragraph (or menu title)
appearing in several articles is stored and translated once. Moving or
renaming an article touches nothing.
## Design (chosen: global content-addressed stores)
Original articles are *also* stored as chunks; everything — originals and
translations — lives in flat hash-keyed dicts. Costs accepted: rendering does
one dict lookup per chunk (trivial), orphaned hashes need occasional garbage
collection, and the editor save path re-chunks server-side (it already
diffs). The rejected alternatives: per-article nested `LangVersion`
structures (churn, duplication, whole-string originals) and a hybrid with
whole originals plus global translations (keeps the worst change-diff
property).
## Target layout
```python
class Node(msgspec.Struct, omit_defaults=True):
...
#: Replaces `content: str | None`. None = pure category label;
#: a list (possibly empty) = a page, as ordered chunk hashes.
chunks: list[bytes] | None = None
#: Primary language of the article (BCP-47 base tag). "" = inherit
#: (nearest ancestor, front page last, site default "en" final).
language: str = ""
#: Chunk hashes the editor marked "do not translate" (always served
#: from the original). Presence-keys, value always True.
no_trans: dict[bytes, True] = {}
#: Languages this article is available in (besides its primary
#: language). Presence-keys, value always True — rendering, language
#: selection and hreflang alternates read this set instead of probing
#: the trans store chunk by chunk. Maintained by the writers (see
#: "Language index maintenance" below).
langs: dict[str, True] = {}
class Data(msgspec.Struct):
...
#: API keys gating the translator service WebSocket (/_translate/{key}):
#: key -> display name; the first is generated at bootstrap (state.py).
translate_keys: dict[str, str] = {}
#: Wanted target languages for the translator service (presence-keys);
#: jobs are offered only in these ∩ a connection's capabilities.
translate_langs: dict[str, True] = {}
#: All original-language text, content-addressed: blake3(normalized)
#: digest[:9] -> Markdown chunk. Shared by every article. Keys are
#: bytes; kanta/msgspec base64-encode them at the JSON level.
chunks: dict[bytes, str] = {}
#: Machine translations: chunk hash -> lang -> translated Markdown
#: (nested, not tuple keys: msgspec's JSON serializer rejects them).
#: Also used for node titles (hash of the title text).
trans: dict[bytes, dict[str, str]] = {}
#: User override patches per article and language:
#: f"{path}:{lang}" -> ordered patches (see localization.md).
patches: dict[str, list[Patch]] = {}
```
Notes:
- **Article paths never carry a leading slash** in the DB or in lookup keys
(`"docs/setup"`, front page `""`); the leading slash is added only when
building hrefs. `migrate_v3` audits existing stored paths (translation
keys, analytics references, any path-valued fields) and normalizes them.
- **Titles are chunks too**, by hash only: the nav renderer looks up
`trans.get(hash(node.title), {}).get(lang)`. No separate title storage;
editing a title invalidates its translations automatically.
- **Per-hunk options** live in two places: *inherent* options are derived at
chunking time (code fences, HTML blocks and prose-free chunks are
no-translate without storing anything — `needs_translation`, see
docs/localization.md "Masking"); *editor-set* flags are `node.no_trans`
(keyed by chunk
hash, so a heavy edit silently drops the flag — acceptable and
self-healing).
- **Patch payloads stay inline** in `Patch.hunks` — patches are small by
construction (minimal server-computed diffs). If a pathological case shows
up, hunks can be hash-stored later without schema pain.
## Language index maintenance (`node.langs`)
`node.langs` is a denormalized index over the `trans`/`patches` stores so
that article rendering, `select_language`'s availability check, and hreflang
alternate links never enumerate chunks. It is written by whoever writes
translation data, in the same transaction:
- **Translator service:** the WebSocket API at `/_translate/{key}` (see
docs/localization.md) offers pending fragments (titles + translatable
chunks lacking an entry for the language) as single-item jobs — one at
a time per connection, in `Data.translate_langs` ∩ the connection's
announced capabilities — and receives the matching result; storing it
writes the `trans[h][lang]` entry, sets `node.langs[lang] = True` on
every article that gained one and invalidates the page cache — all in
one transaction.
- **Translated-view save:** appending the first patch for `f"{path}:{lang}"`
sets `node.langs[lang] = True` (patches alone make the version exist).
- **Removals:** deleting a patch or GC'ing translations re-derives the key:
keep `lang` if any `trans` entry for the article's current chunks/title or
any patch remains, otherwise drop it. Stale `langs` keys are benign (an
advertised language that renders as the original), so removal can lag.
## Render / save pipeline (summary)
- **Render:** `text = "\n\n".join(chunks[h] for h in node.chunks)` for the
original; for language `L` (only ever attempted when `L in node.langs`),
per chunk `trans.get(h, {}).get(L)` unless missing or `h in node.no_trans`,
falling back to `chunks[h]`; then apply `patches.get(f"{path}:{L}", [])`
in order (per-hunk, best effort); then `markdown.render` as today. All of
this assembles the `Translation` the phase-1 plumbing already consumes.
- **Availability:** `node.langs` is the availability index; `?lang=`
handling uses exactly this set. (hreflang alternates are site-wide from
`translate_langs` instead — see docs/localization.md.)
- **Save (primary language):** server re-chunks the submitted Markdown,
inserts new hashes into `Data.chunks`, replaces `node.chunks`. Unchanged
chunks keep their hashes — only genuinely new text lands in the diff.
- **Save (translated view):** diff against the served hybrid, append a
`Patch` under `patches[f"{path}:{lang}"]`; `node.chunks` untouched.
- **Invalidate:** any write to `chunks` / `trans` / `patches` calls
`_invalidate_pages()`.
## migrate_v3 steps
1. Walk `menu`; for every node with a string `content`:
`chunks = chunk_markdown(content)`; write each into the new `chunks`
store; replace the field with the hash list (`None` stays `None`).
2. Initialize empty `chunks` / `trans` / `patches` stores.
3. Normalize stored paths: strip leading slashes anywhere paths are keys or
values.
4. `language`, `no_trans` and `langs` need nothing — struct defaults cover
them (`langs` starts empty; the translator job fills it as translations
land).
Chunking must be deterministic and shared with render/save, so
`chunk_markdown` + `chunk_key` live in `pagerite/i18n.py` (or a small
`pagerite/chunks.py`) and are imported by both `migrations.py` and
`views.py`/`state.py`.
## Implementation notes (deviations from the plan above)
- Chunking lives in `pagerite/chunks.py`; hashing uses the `blake3` package
(already a dependency), truncated to a 9-byte `bytes` digest (kanta's
JSON persistence base64-encodes bytes keys to 12-char strings).
- `trans` is keyed `hash -> lang -> text` (nested dict), not by
`f"{hash}:{lang}"` tuples: msgspec's JSON serializer only supports
str-like/number-like dict keys, and kanta persists as JSON lines.
- `Translation.titles` stayed keyed by node path (phase-1 shape, views
untouched): `get_translation` builds it by walking the menu with the same
per-title `trans.get(chunk_key(node.title), {}).get(lang)` lookups.
- Insert hunks anchor on the whole preceding block (not just its tail) —
a stronger, simpler search context.
- `make_patch` diffs with `SequenceMatcher(autojunk=False)` so patches are
deterministic (popular lines like blank separators never become junk).
- Step 3's path normalization is a no-op in practice: the only path-keyed
store (`patches`) starts empty at v3; analytics paths live outside the
kantadb. The code still strips leading slashes defensively.
## Garbage collection (later, manual or idle-time)
Orphaned entries accumulate: chunks no longer referenced by any
`node.chunks`/`node.title`, translations whose chunk hash is orphaned, patch
hunks that never match. All are harmless (never read). A GC pass is a single
tree walk collecting live hashes, then deleting the rest from `chunks` and
`trans`; patches whose every hunk is stale get pruned. Not part of
migrate_v3.
+41 -4
View File
@@ -27,6 +27,8 @@ import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue' import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue' import VisitorCharts from './VisitorCharts.vue'
import { VIEW_W } from './analytics/chart.js' import { VIEW_W } from './analytics/chart.js'
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
import ConnNote from './ConnNote.vue'
// Same centering margin as the charts, so the totals row's left edge // Same centering margin as the charts, so the totals row's left edge
// aligns with the chart svg above the natural width. // aligns with the chart svg above the natural width.
@@ -40,8 +42,20 @@ const error = ref('')
const now = ref(Date.now()) const now = ref(Date.now())
let ws = null let ws = null
let reconnectTimeout = null let reconnectTimeout = null
let connectWatchdog = null
const reconnects = reconnectPolicy()
let timeInterval = null let timeInterval = null
// The panel is live data over its socket: while it is connecting or waiting
// to reconnect, say so (ConnNote) instead of showing a silent stale view.
const conn = ref('connecting') // connecting | open | waiting
const retryIn = ref(0)
const connNote = computed(() =>
conn.value === 'connecting' ? 'connecting to the server…'
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
: '',
)
// The initial range comes from the URL hash (shareable links); without one, // The initial range comes from the URL hash (shareable links); without one,
// it is derived from the first analytics snapshot: day when the recorded // it is derived from the first analytics snapshot: day when the recorded
// history is shorter than 24 h, week otherwise. // history is shorter than 24 h, week otherwise.
@@ -51,9 +65,16 @@ let rangePinned = Boolean(RANGES[hashRange])
function connectAnalytics() { function connectAnalytics() {
if (ws) return if (ws) return
conn.value = 'connecting'
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:' const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`) ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
ws.onopen = () => { error.value = '' } clearTimeout(connectWatchdog)
connectWatchdog = watchConnecting(ws, 'analytics')
ws.onopen = () => {
conn.value = 'open'
reconnects.opened()
error.value = ''
}
ws.onmessage = (event) => { ws.onmessage = (event) => {
try { try {
data.value = JSON.parse(event.data) data.value = JSON.parse(event.data)
@@ -75,12 +96,19 @@ function connectAnalytics() {
} }
ws.onclose = () => { ws.onclose = () => {
ws = null ws = null
reconnectTimeout = setTimeout(connectAnalytics, 2000) // The policy paces the retry: doubling backoff with jitter, reset only
// by a healthy connection — a fixed rapid loop trips the browser's
// WebSocket throttling (all sockets then sit "pending" for minutes).
const wait = reconnects.closed()
retryIn.value = Math.max(1, Math.round(wait / 1000))
conn.value = 'waiting'
reconnectTimeout = setTimeout(connectAnalytics, wait)
} }
} }
onMounted(async () => { onMounted(async () => {
connectAnalytics() // The first connection takes a staggered slot (see ./reconnect).
reconnectTimeout = setTimeout(connectAnalytics, socketSlot())
now.value = Date.now() now.value = Date.now()
timeInterval = setInterval(() => { now.value = Date.now() }, 1000) timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
// The site tree for the transition map (all pages in menu order). Not // The site tree for the transition map (all pages in menu order). Not
@@ -93,6 +121,7 @@ onMounted(async () => {
onUnmounted(() => { onUnmounted(() => {
if (reconnectTimeout) clearTimeout(reconnectTimeout) if (reconnectTimeout) clearTimeout(reconnectTimeout)
if (connectWatchdog) clearTimeout(connectWatchdog)
if (timeInterval) clearInterval(timeInterval) if (timeInterval) clearInterval(timeInterval)
if (ws) { if (ws) {
ws.onclose = null ws.onclose = null
@@ -134,7 +163,7 @@ const favicons = computed(() => data.value?.favicons || {})
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value)) const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
const crawlers = computed(() => rangeData.value?.crawlers || []) const crawlers = computed(() => rangeData.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value)) const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, now.value)) const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, pageTree.value, now.value))
</script> </script>
@@ -151,6 +180,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
</nav> </nav>
<a href="/" class="close" title="home"></a> <a href="/" class="close" title="home"></a>
</header> </header>
<ConnNote :text="connNote" />
<p v-if="error" class="error"> {{ error }}</p> <p v-if="error" class="error"> {{ error }}</p>
<p v-else-if="!data" class="loading">loading</p> <p v-else-if="!data" class="loading">loading</p>
<template v-else> <template v-else>
@@ -214,6 +244,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
<tbody> <tbody>
<tr v-for="(c, i) in crawlerRows" :key="i"> <tr v-for="(c, i) in crawlerRows" :key="i">
<td class="trail"> <td class="trail">
<TrailLink v-if="c.refererStep" :step="c.refererStep" :favicons="favicons" @close="$emit('close')" />
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" /> <TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
</td> </td>
<VisitorCell <VisitorCell
@@ -240,6 +271,7 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
<thead> <thead>
<tr> <tr>
<th>paths abused</th> <th>paths abused</th>
<th>articles read</th>
<th>visitor</th> <th>visitor</th>
<th class="last-seen">last seen</th> <th class="last-seen">last seen</th>
</tr> </tr>
@@ -256,6 +288,11 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small> <small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
</div> </div>
</td> </td>
<td class="trail clickable-list"
@click="copyList(a.allArticles, $event)">
<TrailLink v-for="(s, si) in a.articles" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
<small v-if="!a.articles.length" class="muted"></small>
</td>
<VisitorCell <VisitorCell
:ip="a.ip" :ip="a.ip"
:ip-display="a.ipDisplay" :ip-display="a.ipDisplay"
+51 -10
View File
@@ -3,9 +3,13 @@
// the real #page-banner region. Close and tab switching live in EditorShell. // the real #page-banner region. Close and tab switching live in EditorShell.
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue' import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
import { EditorView, basicSetup } from 'codemirror' import { EditorView, basicSetup } from 'codemirror'
import { EditorState } from '@codemirror/state' import { Compartment, EditorState } from '@codemirror/state'
import { keymap } from '@codemirror/view'
import { indentWithTab } from '@codemirror/commands'
import { html } from '@codemirror/lang-html' import { html } from '@codemirror/lang-html'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
import ConnNote from './ConnNote.vue'
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
import { dropPageCache, loadPlain, runScripts } from './swapdoc' import { dropPageCache, loadPlain, runScripts } from './swapdoc'
const props = defineProps({ const props = defineProps({
@@ -23,9 +27,22 @@ const bannerEl = ref(null)
let ws = null let ws = null
let pendingSave = null let pendingSave = null
let reconnectTimer = null let reconnectTimer = null
let reconnectDelay = 2000 let connectWatchdog = null
const MAX_RECONNECT_DELAY = 16000 // Reconnection pacing lives in ./reconnect (shared with the other sockets).
const reconnects = reconnectPolicy()
let everConnected = false let everConnected = false
// Connection state drives the note at the top (ConnNote), and locks input
// until the banner's document has arrived (typing before it would be
// clobbered by the doc accept).
const conn = ref('connecting') // connecting | open | waiting
const retryIn = ref(0)
const docReady = ref(false)
const editable = new Compartment()
const connNote = computed(() =>
conn.value === 'connecting' ? 'connecting to the server…'
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
: docReady.value ? '' : 'loading the banner…',
)
let view = null // CodeMirror for the banner HTML let view = null // CodeMirror for the banner HTML
let syncing = false // set while replacing the document programmatically let syncing = false // set while replacing the document programmatically
@@ -88,6 +105,9 @@ function save() {
function openPath(p) { function openPath(p) {
path.value = p path.value = p
// Lock input until the doc arrives (typing would be clobbered by it).
docReady.value = false
view?.dispatch({ effects: editable.reconfigure(EditorView.editable.of(false)) })
send({ type: 'open', path: p }) send({ type: 'open', path: p })
} }
watch(() => props.pagePath, (p) => { openPath(normPath(p)) }) watch(() => props.pagePath, (p) => { openPath(normPath(p)) })
@@ -207,6 +227,8 @@ function onMessage(ev) {
const msg = JSON.parse(ev.data) const msg = JSON.parse(ev.data)
if (msg.type === 'doc' && msg.path === path.value) { if (msg.type === 'doc' && msg.path === path.value) {
setDocument(msg.banner ?? '') setDocument(msg.banner ?? '')
docReady.value = true
view.dispatch({ effects: editable.reconfigure(EditorView.editable.of(true)) })
bannerDesign.value = msg.banner_design ?? null bannerDesign.value = msg.banner_design ?? null
bannerDesignFrom.value = msg.banner_design_from ?? null bannerDesignFrom.value = msg.banner_design_from ?? null
bannerDesignInherited.value = msg.banner_design_inherited ?? '' bannerDesignInherited.value = msg.banner_design_inherited ?? ''
@@ -232,12 +254,22 @@ function onKeydown(ev) {
} }
function connect() { function connect() {
clearTimeout(reconnectTimer)
conn.value = 'connecting'
if (ws) {
// Replacing a stale socket: detach its handlers so its close is silent.
ws.onopen = ws.onmessage = ws.onclose = ws.onerror = null
if (ws.readyState !== WebSocket.CLOSED) ws.close()
}
ws = new WebSocket( ws = new WebSocket(
`${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`, `${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`,
) )
ws.onmessage = onMessage ws.onmessage = onMessage
clearTimeout(connectWatchdog)
connectWatchdog = watchConnecting(ws, 'banner')
ws.onopen = () => { ws.onopen = () => {
reconnectDelay = 2000 conn.value = 'open'
reconnects.opened()
if (everConnected) { if (everConnected) {
if (pendingSave) send(pendingSave) if (pendingSave) send(pendingSave)
} else { } else {
@@ -246,25 +278,32 @@ function connect() {
everConnected = true everConnected = true
} }
ws.onclose = () => { ws.onclose = () => {
clearTimeout(reconnectTimer) // The wait is the policy's: doubling backoff with jitter (./reconnect),
reconnectTimer = setTimeout(() => { // reset only by a healthy connection — rapid retries trip the browser's
connect() // WebSocket throttling (sockets stuck "pending" for minutes).
reconnectDelay = Math.min(reconnectDelay * 2, MAX_RECONNECT_DELAY) const wait = reconnects.closed()
}, reconnectDelay) retryIn.value = Math.max(1, Math.round(wait / 1000))
conn.value = 'waiting'
reconnectTimer = setTimeout(connect, wait)
} }
} }
onMounted(async () => { onMounted(async () => {
connect() // The first connection takes a staggered slot (see ./reconnect).
reconnectTimer = setTimeout(connect, socketSlot())
view = new EditorView({ view = new EditorView({
state: EditorState.create({ state: EditorState.create({
doc: '', doc: '',
extensions: [ extensions: [
basicSetup, basicSetup,
// Tab/Shift-Tab indent and dedent instead of moving focus.
keymap.of([indentWithTab]),
html(), html(),
cmTheme, cmTheme,
cmHighlight, cmHighlight,
EditorView.lineWrapping, EditorView.lineWrapping,
// Locked until the banner's document arrives (docReady/ConnNote).
editable.of(EditorView.editable.of(false)),
EditorView.updateListener.of((u) => { EditorView.updateListener.of((u) => {
if (u.docChanged && !syncing) { if (u.docChanged && !syncing) {
banner.value = view.state.doc.toString() banner.value = view.state.doc.toString()
@@ -282,6 +321,7 @@ onMounted(async () => {
onUnmounted(() => { onUnmounted(() => {
clearTimeout(reconnectTimer) clearTimeout(reconnectTimer)
clearTimeout(connectWatchdog)
for (const t of Object.values(timers)) clearTimeout(t) for (const t of Object.values(timers)) clearTimeout(t)
if (ws) { if (ws) {
ws.onclose = null // intentional close, no reconnect ws.onclose = null // intentional close, no reconnect
@@ -296,6 +336,7 @@ onUnmounted(() => {
<template> <template>
<div class="banner-editor"> <div class="banner-editor">
<div v-if="saveError">{{ saveError }}</div> <div v-if="saveError">{{ saveError }}</div>
<ConnNote :text="connNote" />
<section class="block" @paste="onBannerPaste"> <section class="block" @paste="onBannerPaste">
<div class="block-head"> <div class="block-head">
+21
View File
@@ -0,0 +1,21 @@
<script setup>
// Connection-state note for the WebSocket-backed panels (page/banner
// editors, analytics view): while the socket is connecting or waiting to
// reconnect the panel cannot load or save, and this says so. An empty
// text hides the note.
defineProps({ text: { type: String, default: '' } })
</script>
<template>
<div v-if="text" class="conn-note" role="status">{{ text }}</div>
</template>
<style scoped>
.conn-note {
padding: 0.2rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
color: var(--muted);
font-size: 0.8rem;
}
</style>
+47 -10
View File
@@ -1,5 +1,5 @@
<script setup> <script setup>
// Tabbed shell for the four admin editors. The individual pens are shorthands // Tabbed shell for the five admin editors. The individual pens are shorthands
// that open the shell on a given tab; once open, tabs switch instantly without // that open the shell on a given tab; once open, tabs switch instantly without
// closing the panel. Tabs are kept alive so switching preserves state. // closing the panel. Tabs are kept alive so switching preserves state.
import { onMounted, onUnmounted, provide, ref, watch } from 'vue' import { onMounted, onUnmounted, provide, ref, watch } from 'vue'
@@ -7,6 +7,9 @@ import PageEditor from './PageEditor.vue'
import BannerEditor from './BannerEditor.vue' import BannerEditor from './BannerEditor.vue'
import SiteEditor from './SiteEditor.vue' import SiteEditor from './SiteEditor.vue'
import StructureEditor from './StructureEditor.vue' import StructureEditor from './StructureEditor.vue'
import LocalizationEditor from './LocalizationEditor.vue'
import { editorLang, pagePrimary } from './editorLang'
import { loadPlain, setLangOverride } from './swapdoc'
const props = defineProps({ const props = defineProps({
pagePath: { type: String, default: '' }, pagePath: { type: String, default: '' },
@@ -17,11 +20,37 @@ const emit = defineEmits(['close'])
const currentPath = ref(props.pagePath) const currentPath = ref(props.pagePath)
const activeMode = ref(props.initialMode) const activeMode = ref(props.initialMode)
// Tab order: site-wide settings first (site, structure), then — after a // The shared language selection (./editorLang, v-modeled by the tabs'
// visual break — the per-page editors (article, banner). // LangSelects) also drives the page preview: while the shell is open it
// overrides the normal language preferences (?lang= / Accept-Language),
// so the page renders in the language being edited; closing restores.
// The primary selection pins by the CURRENT PAGE's own primary language
// (pages may differ — Node.language is inherited down the tree).
let pinned = false
function pinPreviewLang() {
pinned = true
// '' pagePrimary = not yet learned: pin 'en', the server's final fallback
// (i18n.ORIGINAL_LANGUAGE).
setLangOverride(editorLang.value || pagePrimary.value || 'en')
loadPlain(currentPath.value)
}
function unpinPreviewLang() {
if (!pinned) return
pinned = false
setLangOverride(null)
loadPlain(currentPath.value)
}
watch(editorLang, () => { if (pinned) pinPreviewLang() })
// The page's primary may be (re)learned while pinned on it (doc accept,
// tree refresh, a language change on the row) — re-pin with the new code.
watch(pagePrimary, () => { if (pinned && !editorLang.value) pinPreviewLang() })
// Tab order: site-wide settings first (site, structure, localization), then
// — after a visual break — the per-page editors (article, banner).
const MODES = [ const MODES = [
{ key: 'site', label: 'site', component: SiteEditor }, { key: 'site', label: 'site', component: SiteEditor },
{ key: 'structure', label: 'structure', component: StructureEditor }, { key: 'structure', label: 'structure', component: StructureEditor },
{ key: 'localization', label: 'lang', component: LocalizationEditor },
{ key: 'page', label: 'article', component: PageEditor, breakBefore: true }, { key: 'page', label: 'article', component: PageEditor, breakBefore: true },
{ key: 'banner', label: 'banner', component: BannerEditor }, { key: 'banner', label: 'banner', component: BannerEditor },
] ]
@@ -54,25 +83,33 @@ function onSwitchEvent(ev) {
// Closing the shell hides it but keeps it mounted (main.js); the tabs stay // Closing the shell hides it but keeps it mounted (main.js); the tabs stay
// cached in KeepAlive the whole time, so no state is ever lost until a real // cached in KeepAlive the whole time, so no state is ever lost until a real
// page reload. On re-show each active tab re-applies its window title and // page reload. On re-show each active tab re-applies its window title and
// preview via its own pagerite:editor-shown listener. // preview via its own pagerite:editor-shown listener. No Escape-to-close:
function onKeydown(ev) { // it fired too easily by accident (e.g. dismissing an editor popup).
if (ev.key === 'Escape' && document.body.classList.contains('editing')) close()
}
onMounted(() => { onMounted(() => {
document.body.dataset.editorMode = activeMode.value document.body.dataset.editorMode = activeMode.value
addEventListener('pagerite:switch-editor', onSwitchEvent) addEventListener('pagerite:switch-editor', onSwitchEvent)
addEventListener('keydown', onKeydown) addEventListener('pagerite:editor-shown', pinPreviewLang)
addEventListener('pagerite:editor-hidden', unpinPreviewLang)
// The shell mounts visible (openEditor), so pin immediately. The site
// default primary language comes from the settings — it only fills the
// unknown; the page/structure tabs refine pagePrimary per page as they
// learn it (their knowledge is strictly better).
pinPreviewLang()
fetch('/_api/settings').then((r) => r.json()).then((s) => {
if (!pagePrimary.value) pagePrimary.value = s.primary_lang || 'en'
}).catch(() => { /* keep the fallback */ })
}) })
onUnmounted(() => { onUnmounted(() => {
removeEventListener('pagerite:switch-editor', onSwitchEvent) removeEventListener('pagerite:switch-editor', onSwitchEvent)
removeEventListener('keydown', onKeydown) removeEventListener('pagerite:editor-shown', pinPreviewLang)
removeEventListener('pagerite:editor-hidden', unpinPreviewLang)
}) })
</script> </script>
<template> <template>
<div class="editor-root overlay"> <div class="editor-root overlay" lang="en" dir="ltr">
<header class="editor-tabs"> <header class="editor-tabs">
<template v-for="m in MODES" :key="m.key"> <template v-for="m in MODES" :key="m.key">
<span v-if="m.breakBefore" class="tab-break" /> <span v-if="m.breakBefore" class="tab-break" />
+147
View File
@@ -0,0 +1,147 @@
<script setup>
// The editor shell's one language selector (page + structure tabs): a small
// flag button opening a clean dropdown, v-modeled on the shared editorLang
// ('' = the primary language). The lang tab's flag grid is a different
// control (toggles, not a select) and stays as it is.
import { computed, ref } from 'vue'
const props = defineProps({
modelValue: { type: String, default: '' },
options: { type: Array, required: true }, // [{tag, code, name, flag, primary}]
title: { type: String, default: '' }, // toggle-button tooltip override
})
const emit = defineEmits(['update:modelValue'])
const open = ref(false)
const toggleBtn = ref(null)
const popStyle = ref({})
const current = computed(
() => props.options.find((o) => o.tag === props.modelValue) ?? props.options[0],
)
function toggle() {
open.value = !open.value
if (open.value) {
// Position: fixed so the popup overflows the scrolling editor panel
// onto the page area instead of being clipped by it.
const r = toggleBtn.value.getBoundingClientRect()
popStyle.value = { top: `${r.bottom + 2}px`, left: `${r.left}px` }
}
}
function select(tag) {
emit('update:modelValue', tag)
open.value = false
}
</script>
<template>
<span v-if="options.length > 1" class="lang-select">
<button
ref="toggleBtn"
type="button"
class="lang-current"
:class="{ open }"
:title="title || (current
? `language: ${current.name}${current.primary ? ' (primary)' : ''}`
: '')"
@click="toggle"
><span v-if="current?.flag" class="flag" v-html="current.flag" /></button>
<span v-if="open" class="lang-pop" :style="popStyle" @mouseleave="open = false">
<button
v-for="o in options"
:key="o.code"
type="button"
:class="{ active: o.tag === modelValue }"
:title="o.primary ? `${o.name} — the primary language` : `${o.name} — translation`"
@click="select(o.tag)"
><span v-if="o.flag" class="flag" v-html="o.flag" /> {{ o.name }}<small v-if="o.primary"> (primary)</small></button>
</span>
</span>
</template>
<style scoped>
.lang-select {
position: relative;
display: flex;
}
/* The closed state is just the small flag — no button chrome until hovered. */
.lang-current {
display: flex;
align-items: center;
padding: 2px;
background: none;
border: 1px solid transparent;
border-radius: 4px;
cursor: pointer;
}
.lang-current:hover,
.lang-current.open {
border-color: var(--line);
}
/* The dropdown matches the page's existing popups (.picker-pop look).
Fixed-positioned (anchored to the toggle's viewport rect on open) so it
is not clipped by the editor panel's scrolling overflow. */
.lang-pop {
position: fixed;
z-index: 20;
display: flex;
flex-direction: column;
align-items: stretch;
gap: 0.15rem;
padding: 0.3rem;
background: var(--bg);
border: 1px solid var(--line);
border-radius: 6px;
box-shadow: 0 4px 16px #0004;
white-space: nowrap;
}
.lang-pop button {
display: flex;
align-items: center;
gap: 0.4rem;
padding: 0.15rem 0.4rem;
font: inherit;
font-size: 0.9rem;
text-align: left;
color: var(--text);
background: none;
border: none;
border-radius: 4px;
cursor: pointer;
}
.lang-pop button:hover {
background: var(--surface);
}
.lang-pop button.active {
color: var(--accent);
}
.lang-pop small {
color: var(--muted);
}
/* Flags render like in the analytics visitor cells. */
.flag {
display: inline-flex;
width: 18px;
height: 12px;
flex: 0 0 auto;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+281
View File
@@ -0,0 +1,281 @@
<script setup>
// Lang tab: the site-wide translation target languages (translate_langs)
// and the translator service WebSocket URL(s) (translate_keys). ALL
// languages are listed, English included — a page whose primary language
// (Node.language, configured per row in the structure tab, inherited down
// the hierarchy) differs can be translated INTO any other. Flag clicks
// toggle and save immediately; the settings round-trip re-reads the
// payload, so this tab only ever changes translate_langs. The settings
// write's invalidation hook kicks the translation dispatcher. The refresh
// button drops all machine translations (user patches are kept), making
// the dispatcher re-translate everything.
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs'
import { dropPageCache } from './swapdoc'
defineProps({ pagePath: { type: String, default: '' } })
// close/path-change are wired by EditorShell; this tab never emits them.
defineEmits(['close', 'pathChange'])
const saveError = ref('')
const selected = ref(new Set())
const keyUrls = ref([])
// The toggleable targets: every translatable language, laid out in
// geographic/cultural groups (one row each) rather than alphabetized —
// related languages sit together (a node's own primary is excluded per
// article, server-side). Any code missing from LANG_GROUPS trails as an
// extra row.
const groups = computed(() => {
const tile = (code) => ({ code, name: langName(code), flag: flagFor(code) })
const rows = LANG_GROUPS.map((g) => g.filter((c) => c in TRANSLATABLE).map(tile))
const covered = new Set(LANG_GROUPS.flat())
const rest = Object.keys(TRANSLATABLE).filter((c) => !covered.has(c)).map(tile)
if (rest.length) rows.push(rest)
return rows.filter((r) => r.length)
})
function updateWindowTitle() {
document.title = 'lang 🖊️'
}
onActivated(updateWindowTitle)
// The shell stays mounted while hidden: when it is re-shown with this tab
// active, restore the window title.
function onEditorShown() {
if (document.body.dataset.editorMode === 'localization') updateWindowTitle()
}
onMounted(async () => {
addEventListener('pagerite:editor-shown', onEditorShown)
try {
const s = await (await fetch('/_api/settings')).json()
selected.value = new Set(s.translate_langs || [])
const wsBase = location.origin.replace(/^http/, 'ws')
keyUrls.value = Object.entries(s.translate_keys || {})
.map(([key, name]) => ({ name, url: `${wsBase}/_translate/${key}` }))
} catch { /* keep defaults */ }
})
onUnmounted(() => removeEventListener('pagerite:editor-shown', onEditorShown))
async function toggle(code) {
const next = new Set(selected.value)
if (next.has(code)) next.delete(code)
else next.add(code)
selected.value = next
try {
const s = await (await fetch('/_api/settings')).json()
const res = await fetch('/_api/settings', {
method: 'PUT',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ ...s, translate_langs: [...next] }),
})
if (res.ok) {
saveError.value = ''
dropPageCache()
} else {
saveError.value = '⚠️ changes could not be saved'
}
} catch {
saveError.value = '⚠️ changes could not be saved'
}
}
// Delete all machine translations server-side; the dispatcher re-fills
// them (a connected translator starts getting jobs right away). User
// patches survive — they are edits, not machine output.
const refreshing = ref(false)
async function refresh() {
if (refreshing.value) return
refreshing.value = true
try {
const res = await fetch('/_api/translations', { method: 'DELETE' })
saveError.value = res.ok ? '' : '⚠️ translations could not be refreshed'
if (res.ok) dropPageCache()
} catch {
saveError.value = '⚠️ translations could not be refreshed'
} finally {
refreshing.value = false
}
}
</script>
<template>
<div class="localization-editor">
<div v-if="saveError">{{ saveError }}</div>
<section class="block">
<div class="block-head">
<span class="field-label">languages</span>
</div>
<div class="flags">
<div v-for="(row, ri) in groups" :key="ri" class="flag-row">
<button
v-for="o in row"
:key="o.code"
type="button"
class="flag-tile"
:class="{ selected: selected.has(o.code) }"
:title="`${o.name} (${o.code})`"
@click="toggle(o.code)"
>
<span class="flag" v-html="o.flag" />
</button>
</div>
</div>
</section>
<section class="block">
<div class="block-head">
<span class="field-label">translations</span>
<small class="muted">deleting re-translates everything; user edits are kept</small>
</div>
<button
type="button"
class="refresh-btn"
:disabled="refreshing"
title="delete all machine translations and let the translator re-fill them"
@click="refresh"
>
{{ refreshing ? 'refreshing…' : 'refresh all translations' }}
</button>
</section>
<section v-if="keyUrls.length" class="block">
<div class="block-head">
<span class="field-label">translator service</span>
<small class="muted">connect scripts/translator.py to</small>
</div>
<div v-for="k in keyUrls" :key="k.url" class="key-row">
<code>{{ k.url }}</code>
<small class="muted">{{ k.name }}</small>
</div>
</section>
</div>
</template>
<style scoped>
.localization-editor {
overflow-y: auto;
background: var(--surface);
}
.block {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.5rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
}
.block-head {
display: flex;
align-items: center;
gap: 0.6rem;
}
.field-label {
color: var(--muted);
font-size: 0.85rem;
}
.muted {
color: var(--muted);
}
/* Flag grid: one geographic group per row. Deselected flags sit dimmed and
grayed; a click brings one to full color (selected = a translation
target) — the shading alone carries the state, no outline. */
.flags {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.2rem 0;
}
.flag-row {
display: flex;
flex-wrap: wrap;
gap: 0.5rem;
}
.flag-tile {
padding: 3px;
background: none;
border: 2px solid transparent;
border-radius: 5px;
cursor: pointer;
opacity: 0.4;
filter: grayscale(0.8);
transition: opacity 0.15s, filter 0.15s, border-color 0.15s;
}
.flag-tile:hover {
opacity: 0.8;
filter: none;
}
.flag-tile.selected {
opacity: 1;
filter: none;
}
/* Same flag chips as the PageEditor language picker / analytics cells. */
.flag {
display: inline-flex;
width: 18px;
height: 12px;
flex: 0 0 auto;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
}
.flag-tile .flag {
width: 36px;
height: 24px;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
.key-row {
display: flex;
align-items: baseline;
gap: 0.6rem;
}
.key-row code {
user-select: all;
}
.refresh-btn {
align-self: flex-start;
margin-bottom: 0.2rem;
padding: 0.3rem 0.8rem;
font: inherit;
font-size: 0.85rem;
color: var(--muted);
background: none;
border: 1px solid var(--line);
border-radius: 5px;
cursor: pointer;
}
.refresh-btn:hover:not(:disabled) {
color: var(--text);
border-color: var(--muted);
}
.refresh-btn:disabled {
opacity: 0.5;
cursor: default;
}
</style>
File diff suppressed because it is too large Load Diff
+6
View File
@@ -5,6 +5,8 @@
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue' import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
import { EditorView, basicSetup } from 'codemirror' import { EditorView, basicSetup } from 'codemirror'
import { EditorState } from '@codemirror/state' import { EditorState } from '@codemirror/state'
import { keymap } from '@codemirror/view'
import { indentWithTab } from '@codemirror/commands'
import { css } from '@codemirror/lang-css' import { css } from '@codemirror/lang-css'
import { html } from '@codemirror/lang-html' import { html } from '@codemirror/lang-html'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
@@ -487,6 +489,8 @@ onMounted(async () => {
doc: '', doc: '',
extensions: [ extensions: [
basicSetup, basicSetup,
// Tab/Shift-Tab indent and dedent instead of moving focus.
keymap.of([indentWithTab]),
css(), css(),
cmTheme, cmTheme,
cmHighlight, cmHighlight,
@@ -506,6 +510,8 @@ onMounted(async () => {
doc: '', doc: '',
extensions: [ extensions: [
basicSetup, basicSetup,
// Tab/Shift-Tab indent and dedent instead of moving focus.
keymap.of([indentWithTab]),
html(), html(),
cmTheme, cmTheme,
cmHighlight, cmHighlight,
+104 -10
View File
@@ -7,9 +7,20 @@
// real — a label with a title and slug, with content (landing page) or // real — a label with a title and slug, with content (landing page) or
// without (category whose URL renders a placeholder page). The front page // without (category whose URL renders a placeholder page). The front page
// is a top-level row with an empty slug, not the parent of the others. // is a top-level row with an empty slug, not the parent of the others.
import { inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue' //
// Languages: the LangSelect switches which language the TITLES are shown
// and edited in (rows without a translation show the original, dimmed) —
// the selection is shared shell-wide (./editorLang) with the page editor
// and the page preview. Translated title edits write a per-language
// fragment (POST /_api/structure with lang); the structure itself —
// slugs, order, hierarchy — is language-independent and always edits the
// same tree.
import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
import StructureTree from './StructureTree.vue' import StructureTree from './StructureTree.vue'
import LangSelect from './LangSelect.vue'
import { slugify } from './slugify' import { slugify } from './slugify'
import { flagFor, langName } from './langs'
import { editorLang, pagePrimary } from './editorLang'
import { dropPageCache, loadPlain } from './swapdoc' import { dropPageCache, loadPlain } from './swapdoc'
const props = defineProps({ const props = defineProps({
@@ -23,6 +34,50 @@ const path = ref('')
const saveError = ref('') const saveError = ref('')
const tree = ref([]) const tree = ref([])
// The language the tree's titles are shown and edited in: "" = primary.
const lang = editorLang
const primaryLang = ref('en')
const siteLangs = ref([])
// The strip's options: the primary language first, then the configured
// translation targets (the lang tab manages that set).
const langOptions = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
.map((code) => ({
tag: code === primaryLang.value ? '' : code,
code,
name: langName(code),
flag: flagFor(code),
primary: code === primaryLang.value,
})),
)
const currentLang = computed(
() => langOptions.value.find((o) => o.tag === lang.value) ?? langOptions.value[0],
)
// The selection is shared (./editorLang): a change re-fetches the tree's
// titles in it (and EditorShell swaps the page preview into it).
watch(lang, () => refreshPages())
// Per-row primary language (Node.language, '' = inherit): the row's
// dropdown lists "inherit" first (naming what it resolves to), then every
// site language. Setting it on a section covers its whole subtree.
const rowLangChoices = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
.map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })),
)
function rowLangOptions(el) {
const resolved = el.primary || primaryLang.value
return [
{ tag: '', code: '_inherit', name: `inherit (${langName(resolved)})`, flag: flagFor(resolved), primary: false },
...rowLangChoices.value,
]
}
async function setLanguage(node, tag) {
await postStructure({ path: node.path, language: tag })
}
function normPath(p) { function normPath(p) {
return p.trim().replace(/^\/+|\/+$/g, '') return p.trim().replace(/^\/+|\/+$/g, '')
} }
@@ -106,7 +161,8 @@ function discardPending() {
async function commitPending() { async function commitPending() {
const node = pending.value const node = pending.value
if (!node) return if (!node) return
// Empty slug: derive one from the title (transliterated to ASCII). // The typed slug is slugified at commit; empty derives one from the
// title (transliterated to ASCII).
const slug = slugify(node.slug.trim()) || slugify(node.title) const slug = slugify(node.slug.trim()) || slugify(node.title)
if (!slug) { if (!slug) {
return return
@@ -151,9 +207,22 @@ async function commitPending() {
} }
// --- Site structure tree (drag-and-drop ordering/moving) ---------------- // --- Site structure tree (drag-and-drop ordering/moving) ----------------
function findNode(nodes, p) {
for (const n of nodes) {
if (n.path === p) return n
const found = findNode(n.children, p)
if (found) return found
}
return null
}
async function refreshPages() { async function refreshPages() {
try { try {
tree.value = await (await fetch('/_api/pages')).json() const q = lang.value ? `?lang=${lang.value}` : ''
tree.value = await (await fetch(`/_api/pages${q}`)).json()
// The tree carries each node's resolved primary language: publish the
// current page's (the shell pins the preview by it on '' selection).
pagePrimary.value = findNode(tree.value, path.value)?.primary || 'en'
} catch { /* list stays stale; not fatal */ } } catch { /* list stays stale; not fatal */ }
} }
@@ -206,20 +275,25 @@ async function onReorder(parentPath, list, evt) {
} }
// Inline title/slug editing: rows are always editable. Title saves while // Inline title/slug editing: rows are always editable. Title saves while
// typing (debounced); the slug commits on blur/Enter, since it renames // typing (debounced) — in the selected language (a translation writes a
// the path (moving the whole subtree with it). // title fragment, the primary language the original); the slug commits on
// blur/Enter, since it renames the path (moving the whole subtree with it).
// Slugs are language-independent.
function onTitleInput(node, ev) { function onTitleInput(node, ev) {
const title = ev.target.value.trim() const title = ev.target.value.trim()
if (!title || title === node.title) return if (!title || title === node.title) return
debounce(`title:${node.path}`, async () => { debounce(`title:${node.path}`, async () => {
await postStructure({ path: node.path, title }) await postStructure({ path: node.path, title, lang: lang.value })
}) })
} }
// The slug inputs are filtered as you type (StructureTree onSlugInput, // Slug inputs are typed freely (spaces become hyphens live, see
// see slugify.js); the server re-validates and its reason is shown. // StructureTree onSlugInput); the value is slugified here at commit
// (blur/Enter) before talking to the server, which re-validates (e.g.
// reserved names) and its reason is shown.
async function commitSlug(node, ev) { async function commitSlug(node, ev) {
const slug = ev.target.value.trim() const slug = slugify(ev.target.value.trim())
ev.target.value = slug
if (slug === node.slug) return if (slug === node.slug) return
const parent = node.path.split('/').slice(0, -1).join('/') const parent = node.path.split('/').slice(0, -1).join('/')
// Empty slug at top level = the front page (path ""). // Empty slug at top level = the front page (path "").
@@ -262,12 +336,19 @@ provide('structureHandlers', {
commitPending, commitPending,
discardPending, discardPending,
newPage, newPage,
langOptions: rowLangOptions,
setLanguage,
}) })
onMounted(() => { onMounted(() => {
path.value = normPath(props.pagePath) path.value = normPath(props.pagePath)
refreshPages() refreshPages()
addEventListener('pagerite:editor-shown', onEditorShown) addEventListener('pagerite:editor-shown', onEditorShown)
// The language strip: site primary + configured targets.
fetch('/_api/settings').then((r) => r.json()).then((s) => {
primaryLang.value = s.primary_lang || 'en'
siteLangs.value = s.translate_langs || []
}).catch(() => { /* no strip */ })
}) })
onUnmounted(() => { onUnmounted(() => {
@@ -279,8 +360,15 @@ onUnmounted(() => {
<template> <template>
<div class="structure-editor"> <div class="structure-editor">
<div v-if="saveError">{{ saveError }}</div> <div v-if="saveError">{{ saveError }}</div>
<div v-if="langOptions.length > 1" class="block lang-block">
<div><LangSelect v-model="lang" :options="langOptions" /></div>
<small v-if="lang" class="muted">
viewing {{ currentLang.name }} titles dimmed rows are untranslated
(shown in the primary language); slugs never translate
</small>
</div>
<section class="block structure"> <section class="block structure">
<StructureTree :nodes="tree" /> <StructureTree :nodes="tree" :lang="lang" />
</section> </section>
</div> </div>
</template> </template>
@@ -305,4 +393,10 @@ onUnmounted(() => {
overflow-y: auto; overflow-y: auto;
min-height: 0; min-height: 0;
} }
/* The language selector is LangSelect.vue — its styles live there. */
.muted {
color: var(--muted);
}
</style> </style>
+63 -15
View File
@@ -1,7 +1,13 @@
<script setup> <script setup>
// Recursive site-structure tree with drag-and-drop ordering (vue-draggable). // Recursive site-structure tree with drag-and-drop ordering (vue-draggable).
// Nodes come from the server (GET /_api/pages via StructureEditor.vue) as // Nodes come from the server (GET /_api/pages via StructureEditor.vue) as
// {slug, path, title, order, published, has_content, children}. // {slug, path, title, translated, order, published, has_content, language,
// primary, children}. The row's flag (LangSelect) sets the node's primary
// language (language; '' = inherit — dimmed, showing the resolved flag);
// the setting covers the whole subtree.
// With a `lang` prop (StructureEditor's language strip) the titles shown
// are that language's; `translated` marks rows with an actual translation
// (untranslated rows show the original title, dimmed).
// Every node is real: a label whose title and slug are always editable // Every node is real: a label whose title and slug are always editable
// inline — the title saves while typing (and focusing it opens the page), // inline — the title saves while typing (and focusing it opens the page),
// the slug commits on blur/Enter since it renames the path, moving the // the slug commits on blur/Enter since it renames the path, moving the
@@ -24,26 +30,32 @@
import { inject } from 'vue' import { inject } from 'vue'
import draggable from 'vuedraggable' import draggable from 'vuedraggable'
import { slugify } from './slugify' import { slugify } from './slugify'
import LangSelect from './LangSelect.vue'
defineOptions({ name: 'StructureTree' }) defineOptions({ name: 'StructureTree' })
const props = defineProps({ const props = defineProps({
nodes: { type: Array, required: true }, nodes: { type: Array, required: true },
parentPath: { type: String, default: '' }, parentPath: { type: String, default: '' },
depth: { type: Number, default: 0 }, depth: { type: Number, default: 0 },
// StructureEditor's selected language ('' = original). Only used for the
// untranslated-title styling here; the fetch and title edits live in the
// parent (handlers.titleInput posts the lang with the op).
lang: { type: String, default: '' },
}) })
const handlers = inject('structureHandlers') const handlers = inject('structureHandlers')
// Live-filter the slug inputs as they are typed (oninput): invalid // Slug inputs accept free typing; the only live rewrites are turning
// characters are simply not accepted, spaces become hyphens and unicode // spaces into hyphens and lowercasing (both keep the length for ASCII,
// folds to ASCII (see slugify.js). Existing rows commit on change, the // so the cursor stays put). Anything else (unicode folding, stripping,
// pending row is v-modeled. // collapsing) is left for commit time, where the value is run through
function onSlugInput(ev) { // slugify before talking to the server (StructureEditor). `element` is
ev.target.value = slugify(ev.target.value) // the pending row (v-modeled), null for existing rows (plain :value
} // binding, read back on commit).
function onSlugInput(element, ev) {
function onPendingSlugInput(element, ev) { const v = ev.target.value.replace(/\s/g, '-').toLowerCase()
element.slug = slugify(ev.target.value) ev.target.value = v
if (element) element.slug = v
} }
// Focus the title input of a fresh pending row. // Focus the title input of a fresh pending row.
@@ -113,7 +125,7 @@ function onEnd() {
class="edit slug-edit" class="edit slug-edit"
:placeholder="slugify(element.title)" :placeholder="slugify(element.title)"
title="Slug (last path segment) — empty: derived from the title" title="Slug (last path segment) — empty: derived from the title"
@input="onPendingSlugInput(element, $event)" @input="onSlugInput(element, $event)"
@keyup.enter="handlers.commitPending()" @keyup.enter="handlers.commitPending()"
@keyup.esc="handlers.discardPending()" @keyup.esc="handlers.discardPending()"
/> />
@@ -125,8 +137,11 @@ function onEnd() {
<template v-else> <template v-else>
<input <input
class="edit title-edit" class="edit title-edit"
:class="{ untranslated: lang && !element.translated }"
:value="element.title" :value="element.title"
title="Label in the navigation — saves while typing; click opens the page" :title="lang && !element.translated
? 'No translation yet showing the original; typing creates the translated title'
: 'Label in the navigation saves while typing; click opens the page'"
@input="handlers.titleInput(element, $event)" @input="handlers.titleInput(element, $event)"
@focus="handlers.open(element.path)" @focus="handlers.open(element.path)"
/> />
@@ -135,10 +150,21 @@ function onEnd() {
:value="element.slug" :value="element.slug"
placeholder="front page" placeholder="front page"
title="Slug (last path segment) — renames move the whole subtree. Empty at top level = front page" title="Slug (last path segment) — renames move the whole subtree. Empty at top level = front page"
@input="onSlugInput" @input="onSlugInput(null, $event)"
@change="handlers.commitSlug(element, $event)" @change="handlers.commitSlug(element, $event)"
/> />
<span class="acts"> <span class="acts">
<span
class="row-lang"
:class="{ inherited: !element.language }"
><LangSelect
:model-value="element.language"
:options="handlers.langOptions(element)"
:title="element.language
? `primary language: set on this page (subtree inherits)`
: `primary language: inherited — set it here (subtree inherits)`"
@update:model-value="handlers.setLanguage(element, $event)"
/></span>
<span v-if="!element.published" class="draft">draft</span> <span v-if="!element.published" class="draft">draft</span>
<button <button
v-if="element.has_content || !element.children.length" v-if="element.has_content || !element.children.length"
@@ -157,6 +183,7 @@ function onEnd() {
:nodes="element.children" :nodes="element.children"
:parent-path="element.path" :parent-path="element.path"
:depth="depth + 1" :depth="depth + 1"
:lang="lang"
/> />
</div> </div>
</template> </template>
@@ -222,7 +249,7 @@ body.tree-dragging .treelist {
level, not across levels). */ level, not across levels). */
.row { .row {
display: grid; display: grid;
grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem 5rem; grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem auto;
align-items: baseline; align-items: baseline;
gap: 0.35rem; gap: 0.35rem;
/* Vertical spacing widens the drop zones: the exposed top strip is the /* Vertical spacing widens the drop zones: the exposed top strip is the
@@ -273,6 +300,13 @@ body.tree-dragging .treelist {
cursor: text; cursor: text;
} }
/* With a language selected (StructureEditor's strip), rows without an
actual translation show the original title dimmed and italic. */
.title-edit.untranslated {
color: var(--muted);
font-style: italic;
}
.slug-edit { .slug-edit {
font-family: var(--font-code); font-family: var(--font-code);
} }
@@ -284,6 +318,20 @@ body.tree-dragging .treelist {
justify-content: end; justify-content: end;
} }
/* Row language selector (LangSelect): the effective primary language's
flag; dimmed while the setting is inherited rather than set on the row. */
.row-lang {
display: inline-flex;
}
.row-lang.inherited :deep(.lang-current) {
opacity: 0.45;
}
.row-lang.inherited:hover :deep(.lang-current) {
opacity: 0.85;
}
.draft { .draft {
color: var(--muted); color: var(--muted);
font-size: 0.75rem; font-size: 0.75rem;
+20 -12
View File
@@ -75,25 +75,33 @@ const viewChart = computed(() => buildChart(viewSeries.value, now.value))
<svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`" <svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`"
:style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }" :style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }"
role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)"> role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)">
<!-- Clip the plot curves to the chart area: past-week overlays can
run far above the autoscaled y range, and the svg itself is
overflow: visible for the axis labels. -->
<clipPath :id="`plot-${c.ylabel}`">
<rect x="0" y="0" :width="CHART_W" :height="CHART_H" />
</clipPath>
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value" <line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" /> :x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
<template v-for="t in c.chart.xticks" :key="'t' + t.x"> <template v-for="t in c.chart.xticks" :key="'t' + t.x">
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H" <line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
class="minor vertical" /> class="minor vertical" />
</template> </template>
<template v-if="c.chart.bars"> <g :clip-path="`url(#plot-${c.ylabel})`">
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i" <template v-if="c.chart.bars">
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" /> <rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
<path :d="c.chart.skyline" class="line" /> :x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
</template> <path :d="c.chart.skyline" class="line" />
<template v-else>
<!-- Oldest overlay weeks first so the current week paints on top. -->
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
<path v-if="s.area" :d="s.area" class="area" />
<path :d="s.line" class="line" :class="{ past: s.past }"
:style="{ opacity: s.opacity }" />
</template> </template>
</template> <template v-else>
<!-- Oldest overlay weeks first so the current week paints on top. -->
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
<path v-if="s.area" :d="s.area" class="area" />
<path :d="s.line" class="line" :class="{ past: s.past }"
:style="{ opacity: s.opacity }" />
</template>
</template>
</g>
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5" <line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
class="axis" /> class="axis" />
<text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y" <text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y"
+31 -17
View File
@@ -370,7 +370,9 @@ export function mainDomain(host, limit = 24) {
/** /**
* Group raw crawler hits by client hash and format each group as a row showing * Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by most recent hit * every internal page that crawler visited. Rows are sorted by most recent hit
* first, with total hits as a tie-breaker. * first, with total hits as a tie-breaker. The group's ``refererStep`` is the
* latest external referer seen for the crawler — spiders often advertise
* their own site there — rendered with its favicon like visit referers.
* ``clients`` maps client hashes to client records. * ``clients`` maps client hashes to client records.
*/ */
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) { export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
@@ -382,10 +384,12 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
clientHash: c.client, clientHash: c.client,
client, client,
lastStart: 0, lastStart: 0,
referer: '',
pages: new Map(), pages: new Map(),
} }
const start = new Date(c.start).getTime() const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start if (start > g.lastStart) g.lastStart = start
if (c.referer) g.referer = c.referer
if (c.entry?.startsWith('/')) { if (c.entry?.startsWith('/')) {
const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 } const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
existing.count += 1 existing.count += 1
@@ -410,6 +414,7 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
lastSeen: formatWhen(g.lastStart, now), lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart), lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart), lastSeenLocal: formatWhenLocal(g.lastStart),
refererStep: stepOf(g.referer, titles),
pages: [...g.pages.entries()] pages: [...g.pages.entries()]
.sort((a, b) => b[1].count - a[1].count) .sort((a, b) => b[1].count - a[1].count)
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })), .map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
@@ -430,16 +435,19 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
/** /**
* Group abuse hits by IP and format each group as a row with the full paths * Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical paths are collapsed into one entry with their hit count. * probed. Identical paths are collapsed into one entry with their hit count.
* Flagged paths (the ones that triggered abuse classification) are lifted to * The paths split into two lists: ``paths`` holds the 404 probes (flagged
* the top, followed by other 404s, then document GETs from the abuser. Within * paths — the ones that triggered abuse classification — first, then other
* each category paths are sorted by count descending, then earliest first. * 404s) shown verbatim, query string included, and ``articles`` holds the
* real (200) document GETs as trail steps resolved against the page tree
* (query string stripped), rendered like the visitor/crawler trails. Within
* each list paths are sorted by count descending, then earliest first.
* Rows are sorted by most recent hit first. Visitor metadata comes from the * Rows are sorted by most recent hit first. Visitor metadata comes from the
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell * latest client hash seen for the IP; ``clientCount`` tells the visitor cell
* how many distinct client variations the IP produced. Paths are shown * how many distinct client variations the IP produced.
* verbatim (query string included), not resolved against the page tree.
* ``clients`` maps client hashes to client records. * ``clients`` maps client hashes to client records.
*/ */
export function formatAbuseRows(abuse, clients, now = Date.now()) { export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
const groups = new Map() const groups = new Map()
for (const a of abuse || []) { for (const a of abuse || []) {
const client = (clients || {})[a.client] || {} const client = (clients || {})[a.client] || {}
@@ -481,13 +489,14 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
.sort((a, b) => b.lastStart - a.lastStart) .sort((a, b) => b.lastStart - a.lastStart)
.slice(0, 10) .slice(0, 10)
.map((g) => { .map((g) => {
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2) const all = [...g.pathCounts.values()]
const paths = [...g.pathCounts.values()].sort( const byCount = (a, b) => b.count - a.count || a.firstStart - b.firstStart
(a, b) => const paths = all
pathCategory(a) - pathCategory(b) || .filter((p) => p.flag || p.is_404)
b.count - a.count || .sort((a, b) => (a.flag ? 0 : 1) - (b.flag ? 0 : 1) || byCount(a, b))
a.firstStart - b.firstStart, const articles = all.filter((p) => !p.flag && !p.is_404).sort(byCount)
) const pathList = (list) =>
list.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path)).join('\n')
const client = (clients || {})[g.lastClient] || {} const client = (clients || {})[g.lastClient] || {}
const host = client.host || '' const host = client.host || ''
const isHost = !!host const isHost = !!host
@@ -501,9 +510,14 @@ export function formatAbuseRows(abuse, clients, now = Date.now()) {
flag: p.flag, flag: p.flag,
is_404: p.is_404, is_404: p.is_404,
})), })),
allPaths: paths allPaths: pathList(paths),
.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path)) articles: articles
.join('\n'), .map((p) => {
const step = stepOf(p.path.split('?')[0], titles)
return step ? { ...step, count: p.count } : null
})
.filter(Boolean),
allArticles: pathList(articles),
clientCount: g.clientHashes.size, clientCount: g.clientHashes.size,
ip: client.ip || g.ip, ip: client.ip || g.ip,
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
+215 -61
View File
@@ -17,7 +17,15 @@
--line: #0000001a; --line: #0000001a;
/* Links stay quiet — mostly text color with a hint of the accent — /* Links stay quiet — mostly text color with a hint of the accent —
and light up to the full accent on hover. */ and light up to the full accent on hover. */
--link: color-mix(in oklab, var(--text) 50%, var(--accent)); --link: color-mix(var(--text) 50%, var(--accent));
/* Inline code tint in body paragraphs: halfway between text and muted.
Both endpoints are theme constants, so the mix is a definite color
per theme — themes may also pin it outright. Outside paragraphs code
inherits the context's color (accent headings stay accent). */
--code-inline: color-mix(var(--text), var(--muted));
/* Selection fill for page text and the CodeMirror editors; themes
override when the accent tint clashes with accent-colored text. */
--selection-bg: color-mix(var(--accent) 30%, transparent);
/* Code highlighting palette, consumed by pygments.css: complete light and /* Code highlighting palette, consumed by pygments.css: complete light and
dark sets (background included), resolved by light-dark() from the dark sets (background included), resolved by light-dark() from the
used color-scheme. A theme picks a set simply by declaring used color-scheme. A theme picks a set simply by declaring
@@ -71,10 +79,25 @@
} }
::selection { ::selection {
background: color-mix(var(--accent) 30%, transparent); background: var(--selection-bg);
color: inherit; color: inherit;
} }
/* Text size classes for any block ({.small} {.large} {.huge}, set via the
format bar or by hand): em units on the single 1rem base scale, so they
compose with the theme's typography. */
.small {
font-size: 0.7em;
}
.large {
font-size: 1.5em;
}
.huge {
font-size: 3em;
}
/* Links never underline — including SVG link text, which the UA stylesheet /* Links never underline — including SVG link text, which the UA stylesheet
underlines by default. */ underlines by default. */
a { a {
@@ -93,13 +116,13 @@ html {
auto-hidden scrollbars styled by the --os-* variables below, so they auto-hidden scrollbars styled by the --os-* variables below, so they
never reserve layout space or shift the page when appearing. */ never reserve layout space or shift the page when appearing. */
scrollbar-width: thin; scrollbar-width: thin;
scrollbar-color: color-mix(in srgb, var(--muted) 45%, transparent) transparent; scrollbar-color: color-mix(var(--muted) 45%, transparent) transparent;
} }
.os-scrollbar { .os-scrollbar {
--os-size: 0.5rem; --os-size: 0.5rem;
--os-thumb-bg: color-mix(in srgb, var(--muted) 45%, transparent); --os-thumb-bg: color-mix(var(--muted) 45%, transparent);
--os-thumb-hover-bg: color-mix(in srgb, var(--muted) 65%, transparent); --os-thumb-hover-bg: color-mix(var(--muted) 65%, transparent);
--os-thumb-active-bg: var(--muted); --os-thumb-active-bg: var(--muted);
--os-track-bg: transparent; --os-track-bg: transparent;
--os-thumb-border-radius: 0.25rem; --os-thumb-border-radius: 0.25rem;
@@ -107,7 +130,10 @@ html {
body { body {
font-family: var(--font-body); font-family: var(--font-body);
font-size: 1.05rem; /* 1rem exactly: the text scale (headings, .small/.large/.huge) keys off
one base size — no per-context tweaking, or consistent sizing becomes
impossible. */
font-size: 1rem;
line-height: 1.65; line-height: 1.65;
/* Tabular numerals wherever the active font supports them; avoids /* Tabular numerals wherever the active font supports them; avoids
numbers jumping in width as counters/values change. */ numbers jumping in width as counters/values change. */
@@ -385,7 +411,7 @@ body.editing #sidebar {
overflow-y: auto; overflow-y: auto;
padding: 1rem 1rem 1rem 1.25rem; padding: 1rem 1rem 1rem 1.25rem;
border-radius: 0 0 0.5rem 0; border-radius: 0 0 0.5rem 0;
background: color-mix(in srgb, var(--bg) 75%, transparent); background: color-mix(var(--bg) 75%, transparent);
backdrop-filter: blur(0.5rem); backdrop-filter: blur(0.5rem);
} }
@@ -797,7 +823,7 @@ article dd {
two fluid lanes (36rem minimum, never more than two) once they fit two fluid lanes (36rem minimum, never more than two) once they fit
beside the zone, capped at 102rem total. The zone — the region the nav beside the zone, capped at 102rem total. The zone — the region the nav
sidebar overlays — is a margin indent on the lane content; margin sidebar overlays — is a margin indent on the lane content; margin
boxes float into it, and the text never moves. */ boxes are placed into it out of flow, and the text never moves. */
article.multicol { article.multicol {
margin-inline: auto; margin-inline: auto;
max-width: 58rem; /* 42rem lane + 16rem zone */ max-width: 58rem; /* 42rem lane + 16rem zone */
@@ -806,32 +832,27 @@ article.multicol {
@container (min-width: 45rem) { @container (min-width: 45rem) {
/* The side zone (not on phones): lane content indents 16rem; margin /* The side zone (not on phones): lane content indents 16rem; margin
boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures) boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures)
float at the article's left edge — the same region the nav sidebar are taken out of flow and placed against the article's left edge —
overlays. Scoped to direct article children (the backend render keeps the same region the nav sidebar overlays. The boxes stay in the
margin blocks out of the column segments); nested ones keep the column segment at their anchor point (the backend render no longer
in-column float fallback. */ splits segments around them); absolute positioning off the article —
always position: relative — pins the horizontal side to the zone
regardless of any column layout inside, while the unset top keeps
the box at the vertical position where it occurs in the text. */
article.multicol>.colseg, article.multicol>.colseg,
article.multicol>h1, article.multicol>h1,
article.multicol>h2 { article.multicol>h2 {
margin-left: 16rem; margin-left: 16rem;
} }
article.multicol>.margin, article.multicol .margin,
article.multicol>.aside, article.multicol .aside,
article.multicol>figure:has(.margin) { article.multicol figure.margin {
float: left; position: absolute;
clear: left; left: 0;
width: 14rem; width: 14rem;
max-width: none; max-width: none;
margin: 0.3rem 2rem 1rem 0; margin: 0.3rem 0 0;
}
/* Wide separators start below any margin box — their bleed must not
wrap around it. */
article.multicol>figure:has(.wide),
article.multicol>div.wide,
article.multicol>pre.wide {
clear: left;
} }
} }
@@ -875,11 +896,10 @@ article.multicol {
margin-left: 0; margin-left: 0;
} }
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.margin, body:has(#sidebar):has(.multicol):not(.editing) article.multicol .margin,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.aside, body:has(#sidebar):has(.multicol):not(.editing) article.multicol .aside,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>figure:has(.margin) { body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure.margin {
float: left; position: absolute;
clear: left;
/* Attached to the article's left border (1.25rem gap), hanging into /* Attached to the article's left border (1.25rem gap), hanging into
the left lane and growing leftward with it: 12rem when the lane is the left lane and growing leftward with it: 12rem when the lane is
tight, up to 150% (18rem) when the track or the surplus has room tight, up to 150% (18rem) when the track or the surplus has room
@@ -888,14 +908,15 @@ article.multicol {
--box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem); --box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem);
width: var(--box-w); width: var(--box-w);
max-width: none; max-width: none;
margin: 0.3rem 0 1rem calc(-1.25rem - var(--box-w)); left: calc(-1.25rem - var(--box-w));
margin: 0.3rem 0 0;
} }
} }
/* A shrink-wrapped figure (explicit image width) centers in the plain /* A shrink-wrapped figure (explicit image width) centers in the plain
layout; inside a column the centering looks adrift — left-align. layout; inside a column the centering looks adrift — left-align.
Floated figures keep their own margins (the text gap). */ Floated figures keep their own margins (the text gap). */
.multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), :has(.margin)) { .multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), .margin) {
margin-inline: 0; margin-inline: 0;
} }
@@ -999,7 +1020,7 @@ blockquote p + p {
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
border-left: 0.25rem solid var(--admonition-color, var(--accent)); border-left: 0.25rem solid var(--admonition-color, var(--accent));
border-radius: 0 0.3rem 0.3rem 0; border-radius: 0 0.3rem 0.3rem 0;
background: color-mix(in srgb, var(--admonition-color, var(--accent)) 7%, transparent); background: color-mix(var(--admonition-color, var(--accent)) 7%, transparent);
} }
.admonition> :last-child, .admonition> :last-child,
@@ -1072,14 +1093,18 @@ blockquote p + p {
--admonition-color: var(--accent3); --admonition-color: var(--accent3);
} }
/* Side boxes: ::: aside is a muted floated box (consecutive asides stack /* Side boxes: ::: aside is a muted floated box; {.margin} / ::: margin
via clear: left); {.margin} / ::: margin is a plainer margin note, and is a plainer margin note, and figures take the class directly on the
figures take {.margin} like {.left}. On multicol pages they float in <figure> (the renderer moves it off the img), floating like {.left}.
the composition's left side zone — or in the sidebar's track when the Where the layout has room for a side zone — multicol pages, the
layout reserves one (see the article section); on wide single-column sidebar's track, the wide single-column gutter (see the article
pages they lean into the left gutter (with the figure rules below); section and the figure rules below) — the boxes are taken out of flow
otherwise they stay in-column left floats. Headings already clear and absolutely positioned into it, off the article's left border, each
floats, so boxes never bleed into the next section. */ at the vertical spot where it occurs in the text (boxes occurring
closer together than their heights may overlap — keep them apart);
otherwise they stay in-column left floats (consecutive floats stack
via clear: left). Headings already clear floats, so in-column boxes
never bleed into the next section. */
.aside { .aside {
float: left; float: left;
clear: left; clear: left;
@@ -1089,7 +1114,7 @@ blockquote p + p {
padding: 0.6rem 0.9rem; padding: 0.6rem 0.9rem;
font-size: 0.9rem; font-size: 0.9rem;
color: var(--muted); color: var(--muted);
background: color-mix(in srgb, var(--accent) 6%, transparent); background: color-mix(var(--accent) 6%, transparent);
border-radius: 0.3rem; border-radius: 0.3rem;
} }
@@ -1148,6 +1173,32 @@ code {
font-size-adjust: ex-height var(--code-x-height); font-size-adjust: ex-height var(--code-x-height);
} }
/* Inline code: a slight nudge toward the muted tone rather than a fixed
color, so code inside accent-colored headings keeps the heading's hue
and the distinction stays subtle everywhere. color-mix resolves against
the inherited color (currentColor in a color declaration refers to the
inherited value), shifting it 30% toward --muted. */
code:not(pre code) {
padding-inline: 0.2em;
/* Never break inside a code span. word-break: keep-all would not
suffice: a hard hyphen is an explicit break opportunity (UAX #14),
which keep-all does not suppress — `--arg` could still split after
the dashes. Inline code is short, so nowrap is safe here. */
white-space: nowrap;
}
/* In body paragraphs the text color is a known constant (--text), so
code can take a fixed tint (--code-inline) instead of an unreliable
mix off the inherited color. A paragraph never wraps pre, so no guard
is needed here. */
p code {
color: var(--code-inline);
}
code:not(pre code):first-child {
padding-left: 0;
}
/* Click-to-copy button (added by pagerite.js) */ /* Click-to-copy button (added by pagerite.js) */
.copy { .copy {
position: absolute; position: absolute;
@@ -1192,18 +1243,18 @@ td {
th { th {
text-align: left; text-align: left;
background: linear-gradient(180deg, background: linear-gradient(180deg,
var(--table-head-a, color-mix(in srgb, var(--accent) 10%, var(--surface))), var(--table-head-a, color-mix(var(--accent) 10%, var(--surface))),
var(--table-head-b, color-mix(in srgb, var(--accent) 18%, var(--surface)))); var(--table-head-b, color-mix(var(--accent) 18%, var(--surface))));
} }
td { td {
background: linear-gradient(160deg, background: linear-gradient(160deg,
color-mix(in srgb, var(--table-tint, var(--accent)) 5%, transparent), color-mix(var(--table-tint, var(--accent)) 5%, transparent),
transparent 75%); transparent 75%);
} }
tbody tr+tr td { tbody tr+tr td {
border-top: 1px solid color-mix(in srgb, var(--table-tint, var(--accent)) 12%, transparent); border-top: 1px solid color-mix(var(--table-tint, var(--accent)) 12%, transparent);
} }
/* Definition lists are laid out as a lightweight two-column grid — terms /* Definition lists are laid out as a lightweight two-column grid — terms
@@ -1237,7 +1288,8 @@ dd {
Markdown images standing alone in a paragraph render as a block Markdown images standing alone in a paragraph render as a block
<figure> (with <figcaption> when the image has a title); the <figure> (with <figcaption> when the image has a title); the
brace-attribute positioning class ({.left}, {.right}, {.wide}) lives brace-attribute positioning class ({.left}, {.right}, {.wide}) lives
on the img inside, but only the figure is ever positioned, so the on the img inside — except {.margin}, which the renderer moves onto
the figure itself — but only the figure is ever positioned, so the
caption stays below the image. Raw <img> HTML written by the author caption stays below the image. Raw <img> HTML written by the author
stays inline and unstyled beyond these defaults. */ stays inline and unstyled beyond these defaults. */
img { img {
@@ -1277,6 +1329,24 @@ figure:has(.left) {
margin: 0.3rem 1em 1rem 0; margin: 0.3rem 1em 1rem 0;
} }
/* The same floats for other blocks: ::: left / ::: right containers
(rendered div.left/right), paragraphs, code fences, blockquotes and
tables all take the class directly ({.right} at the end of a
paragraph's last line, a trailing {.left} line after a fence, ...). */
:is(div, p, pre, blockquote, table).right {
float: right;
width: 30%;
max-width: 50%;
margin: 0.3rem 0 1rem 1em;
}
:is(div, p, pre, blockquote, table).left {
float: left;
width: 30%;
max-width: 50%;
margin: 0.3rem 1em 1rem 0;
}
/* An image with an explicit width attribute shrink-wraps instead: the /* An image with an explicit width attribute shrink-wraps instead: the
figure fits the image and, per the auto inline margins above, centers figure fits the image and, per the auto inline margins above, centers
in the column. Placed after the percentage widths above so it in the column. Placed after the percentage widths above so it
@@ -1285,30 +1355,106 @@ figure:has(img[width]) {
width: fit-content; width: fit-content;
} }
/* {.margin} figures float left like {.left} ones — until they fall into /* {.margin} figures (the class moves onto the figure wrapper at render —
see markdown.py) float left like {.left} ones — until they fall into
the side zone (see the composition rules up in the article section). */ the side zone (see the composition rules up in the article section). */
figure:has(.margin) { figure.margin {
float: left; float: left;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 1em 1rem 0; margin: 0.3rem 1em 1rem 0;
} }
/* Click-to-enlarge (pagerite.js): article figure images open in a
full-viewport lightbox — the image as large as fits with the caption
below; click or Esc closes. The backdrop keeps a hint of the page
behind it (blur + dark tint) rather than going solid black; themes
retune it via --lightbox-bg / --lightbox-text. */
article figure img {
cursor: zoom-in;
}
#lightbox {
--lightbox-bg: rgb(0 0 0 / 0.68);
--lightbox-text: #e8e8e8;
position: fixed;
inset: 0;
z-index: 100;
display: flex;
flex-direction: column;
align-items: center;
justify-content: center;
background: var(--lightbox-bg);
backdrop-filter: blur(1rem) saturate(0.85);
/* Wheel/touch scrolling stops at the overlay instead of chaining to
the page behind it. overscroll-behavior needs a scroll container,
hence overflow: auto — the content is clamped to fit, so it never
actually scrolls. */
overflow: auto;
overscroll-behavior: contain;
cursor: zoom-out;
}
#lightbox img {
max-width: 100vw;
/* Flex-shrink does the vertical fit: the image yields exactly the
space the caption needs, no fixed reservation. */
max-height: 100%;
flex: 0 1 auto;
min-height: 0;
object-fit: contain;
border-radius: 3px;
box-shadow: 0 1rem 4rem rgb(0 0 0 / 0.55);
}
#lightbox .caption {
max-width: 65ch;
padding: 0.8rem 1rem;
color: var(--lightbox-text);
font-size: 0.95rem;
text-align: center;
opacity: 0.85;
}
@media (prefers-reduced-motion: no-preference) {
#lightbox {
animation: lightbox-in 0.18s ease-out;
}
#lightbox img {
animation: lightbox-img 0.22s ease-out;
}
@keyframes lightbox-in {
from {
opacity: 0;
}
}
@keyframes lightbox-img {
from {
opacity: 0;
transform: scale(0.96);
}
}
}
/* Wide single-column pages: margin boxes lean into the vacant left /* Wide single-column pages: margin boxes lean into the vacant left
gutter instead (below 104rem the gutter cannot hold the box, and while gutter instead (below 104rem the gutter cannot hold the box, and while
editing the docked panel reshapes the gutters — in both they stay editing the docked panel reshapes the gutters — in both they stay
plain floats). The box grows with the gutter up to 150% (18rem), its plain floats). Out of flow like on multicol pages: the box hangs off
right side 1.25rem off the article's left border. */ the article's left border, growing with the gutter up to 150% (18rem),
its right side 1.25rem off the border. */
@media (min-width: 104rem) { @media (min-width: 104rem) {
body:not(.editing):not(:has(.multicol)) article>.margin, body:not(.editing):not(:has(.multicol)) article .margin,
body:not(.editing):not(:has(.multicol)) article>.aside, body:not(.editing):not(:has(.multicol)) article .aside,
body:not(.editing):not(:has(.multicol)) article>figure:has(.margin) { body:not(.editing):not(:has(.multicol)) article figure.margin {
float: left; position: absolute;
clear: left;
--box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem); --box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem);
width: var(--box-w); width: var(--box-w);
max-width: none; max-width: none;
margin: 0.3rem 0 1rem calc(-1.25rem - var(--box-w)); left: calc(-1.25rem - var(--box-w));
margin: 0.3rem 0 0;
} }
} }
@@ -1457,8 +1603,8 @@ article h2 {
/* Phones and other narrow viewports: single-column layout with the /* Phones and other narrow viewports: single-column layout with the
sidebar lifted above the article as a wrapping link strip, and no sidebar lifted above the article as a wrapping link strip, and no
floats at all — .left/.right/.margin figures fall back to plain floats at all — .left/.right/.margin figures and blocks fall back to
centered figures (explicit img widths still shrink-wrap), margin boxes plain full-width (explicit img widths still shrink-wrap), margin boxes
go full width, while .wide keeps its full viewport bleed. */ go full width, while .wide keeps its full viewport bleed. */
@media (max-width: 48rem) { @media (max-width: 48rem) {
@@ -1525,13 +1671,21 @@ article h2 {
figure:has(.right), figure:has(.right),
figure:has(.left), figure:has(.left),
figure:has(.margin) { figure.margin {
float: none; float: none;
width: 100%; width: 100%;
max-width: none; max-width: none;
margin: 0 auto 1.5rem; margin: 0 auto 1.5rem;
} }
/* Non-figure floated blocks flatten to plain full-width blocks too. */
:is(div, p, pre, blockquote, table):is(.right, .left) {
float: none;
width: auto;
max-width: none;
margin: 0 0 1rem;
}
/* Margin boxes go full width too — no room for side floats on a /* Margin boxes go full width too — no room for side floats on a
phone. */ phone. */
.aside, .aside,
+2 -2
View File
@@ -25,12 +25,12 @@ pre code .cp { color: var(--code-comment); font-weight: bold; font-style: italic
pre code .cpf { color: var(--code-comment); font-style: italic } /* Comment.PreprocFile */ pre code .cpf { color: var(--code-comment); font-style: italic } /* Comment.PreprocFile */
pre code .c1 { color: var(--code-comment); font-style: italic } /* Comment.Single */ pre code .c1 { color: var(--code-comment); font-style: italic } /* Comment.Single */
pre code .cs { color: var(--code-comment); font-weight: bold; font-style: italic } /* Comment.Special */ pre code .cs { color: var(--code-comment); font-weight: bold; font-style: italic } /* Comment.Special */
pre code .gd { color: var(--code-error); background-color: color-mix(in oklab, var(--code-error) 25%, var(--code-bg)) } /* Generic.Deleted */ pre code .gd { color: var(--code-error); background-color: color-mix(var(--code-error) 25%, var(--code-bg)) } /* Generic.Deleted */
pre code .ge { color: var(--code-text); font-style: italic } /* Generic.Emph */ pre code .ge { color: var(--code-text); font-style: italic } /* Generic.Emph */
pre code .ges { color: var(--code-text); font-weight: bold; font-style: italic } /* Generic.EmphStrong */ pre code .ges { color: var(--code-text); font-weight: bold; font-style: italic } /* Generic.EmphStrong */
pre code .gr { color: var(--code-error) } /* Generic.Error */ pre code .gr { color: var(--code-error) } /* Generic.Error */
pre code .gh { color: var(--code-builtin); font-weight: bold } /* Generic.Heading */ pre code .gh { color: var(--code-builtin); font-weight: bold } /* Generic.Heading */
pre code .gi { color: var(--code-added); background-color: color-mix(in oklab, var(--code-added) 25%, var(--code-bg)) } /* Generic.Inserted */ pre code .gi { color: var(--code-added); background-color: color-mix(var(--code-added) 25%, var(--code-bg)) } /* Generic.Inserted */
pre code .go { color: var(--code-muted) } /* Generic.Output */ pre code .go { color: var(--code-muted) } /* Generic.Output */
pre code .gp { color: var(--code-muted) } /* Generic.Prompt */ pre code .gp { color: var(--code-muted) } /* Generic.Prompt */
pre code .gs { color: var(--code-text); font-weight: bold } /* Generic.Strong */ pre code .gs { color: var(--code-text); font-weight: bold } /* Generic.Strong */
+17 -3
View File
@@ -8,7 +8,7 @@ import { tags } from '@lezer/highlight'
// The base theme sets monospace on .cm-scroller, so the font must be set // The base theme sets monospace on .cm-scroller, so the font must be set
// there, not on "&". // there, not on "&".
export const cmTheme = EditorView.theme({ const cmEditorTheme = EditorView.theme({
"&": { "&": {
backgroundColor: "var(--bg)", backgroundColor: "var(--bg)",
color: "var(--text)", color: "var(--text)",
@@ -33,11 +33,25 @@ export const cmTheme = EditorView.theme({
".cm-cursor": { borderLeftColor: "var(--text)" }, ".cm-cursor": { borderLeftColor: "var(--text)" },
// basicSetup's active-line highlight assumes a dark theme. // basicSetup's active-line highlight assumes a dark theme.
".cm-activeLine": { backgroundColor: "transparent" }, ".cm-activeLine": { backgroundColor: "transparent" },
"&.cm-focused .cm-selectionBackground, .cm-selectionBackground":
{ backgroundColor: "var(--line)" },
"&.cm-focused": { outline: "none" }, "&.cm-focused": { outline: "none" },
}) })
// Selection color needs a baseTheme: only base themes support the
// &light/&dark selectors, and @codemirror/view's own selection rules use
// them — we must match its selectors exactly (equal specificity) and rely
// on mounting later to win. Focused: the page's --selection-bg (the base
// accents tint; themes may override it). Unfocused: hidden, like a normal
// input (CodeMirror greys it by default).
const cmSelection = EditorView.baseTheme({
"&light .cm-selectionBackground, &dark .cm-selectionBackground":
{ backgroundColor: "transparent" },
"&light.cm-focused > .cm-scroller > .cm-selectionLayer .cm-selectionBackground, &dark.cm-focused > .cm-scroller > .cm-selectionLayer .cm-selectionBackground":
{ backgroundColor: "var(--selection-bg)" },
})
// Exported as one extension so the editors just list `cmTheme`.
export const cmTheme = [cmEditorTheme, cmSelection]
export const cmHighlight = syntaxHighlighting(HighlightStyle.define([ export const cmHighlight = syntaxHighlighting(HighlightStyle.define([
{ tag: tags.heading, fontWeight: "600", color: "var(--accent)" }, { tag: tags.heading, fontWeight: "600", color: "var(--accent)" },
{ tag: tags.strong, fontWeight: "700" }, { tag: tags.strong, fontWeight: "700" },
+15
View File
@@ -0,0 +1,15 @@
// The editor shell's shared language selection ('' = the primary language):
// one state, v-modeled by the LangSelect of every tab that has one (page,
// structure). While the panel is open it also drives the page preview —
// EditorShell applies it as the fetch-time language override (swapdoc).
import { ref } from 'vue'
export const editorLang = ref('')
// The CURRENT PAGE's primary language ('' = not yet learned): the shell's
// settings fetch fills it with the site default; the page/structure tabs
// then refine it per page (doc accept / tree rows — strictly better
// sources, so they overwrite freely while the settings fetch only fills
// the unknown). EditorShell pins the preview by it when the selection is
// '' (the primary).
export const pagePrimary = ref('')
+57
View File
@@ -0,0 +1,57 @@
// Language helpers shared by the editors (the PageEditor language picker,
// the localization settings tab). Flags come from the country-flag-icons
// set, same as the analytics visitor cells.
import * as flagSvgs from 'country-flag-icons/string/3x2'
// The Seed-X reference translator's languages (scripts/translator.py) — the
// translation-target ceiling — each mapped to the language's home country
// flag (England for English, Portugal for Portuguese — not the most
// populous variant). Internal tags are the bare 2-letter base subtags; the
// translator decides the variant. A variant tag (en-US, pt-BR) is still a
// valid explicit selection for a future translator that distinguishes them
// — flagFor shows its own region then.
export const TRANSLATABLE = {
ar: 'EG', cs: 'CZ', da: 'DK', de: 'DE', el: 'GR', en: 'GB', es: 'ES',
fa: 'IR', fi: 'FI', fr: 'FR', hu: 'HU', id: 'ID', it: 'IT', ja: 'JP',
ko: 'KR', ms: 'MY', nl: 'NL', no: 'NO', pl: 'PL', pt: 'PT', ro: 'RO',
ru: 'RU', sv: 'SE', th: 'TH', tr: 'TR', uk: 'UA', vi: 'VN', zh: 'CN',
}
// The languages in geographic/cultural groups (the lang tab's flag grid
// lays them out one group per row, in this order): English with the
// Nordics, then Western/Central and Eastern Europe, Southern Europe with
// the Middle East, and Asia.
export const LANG_GROUPS = [
['en', 'nl', 'da', 'no', 'sv', 'fi', 'ru'],
['fr', 'de', 'pl', 'cs', 'hu', 'ro', 'uk'],
['es', 'pt', 'it', 'el', 'tr', 'ar', 'fa'],
['zh', 'ja', 'ko', 'vi', 'th', 'id', 'ms'],
]
const displayNames = new Intl.DisplayNames(['en'], { type: 'language' })
// English display name for a language tag ("fi" -> "Finnish").
export function langName(tag) {
try {
return displayNames.of(tag) || tag
} catch {
return tag
}
}
// Flag SVG string for a language tag: an explicit region variant (en-US)
// gets its own region's flag; a bare base tag maps to the language's home
// country (en → GB, pt → PT); languages outside the list fall back to the
// tag's most likely region.
export function flagFor(tag) {
tag = tag || ''
if (!tag.includes('-')) {
const country = TRANSLATABLE[tag.split('-')[0].toLowerCase()]
if (country) return flagSvgs[country] || ''
}
try {
return flagSvgs[new Intl.Locale(tag).maximize().region] || ''
} catch {
return ''
}
}
+230 -68
View File
@@ -6,6 +6,7 @@
// support from the article itself and are re-applied after each swap. // support from the article itself and are re-applied after each swap.
import { OverlayScrollbars } from "overlayscrollbars"; import { OverlayScrollbars } from "overlayscrollbars";
import "overlayscrollbars/overlayscrollbars.css"; import "overlayscrollbars/overlayscrollbars.css";
import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
(() => { (() => {
// Overlay scrollbars: the native kind reserves a strip of layout (or // Overlay scrollbars: the native kind reserves a strip of layout (or
@@ -34,6 +35,54 @@ import "overlayscrollbars/overlayscrollbars.css";
}); });
} }
// --- Language override (?lang=) ---------------------------------------
// /page?lang=fi serves a translated, indexable version (each language is
// its own canonical). The chosen language sticks for the session of
// clicks: the server replicates ?lang= onto the navigation links it
// renders (nav, sidebar, cards — in-article links are content and stay
// as authored), and pageUrl adds it to internal fetches that lack one.
// The address bar keeps the pretty URL: the query is stripped on load
// and never pushed into history. A full refresh or a shared link resets
// to automatic selection (the browser's own Accept-Language — every
// plain fetch carries it by default). See docs/localization.md.
const langParam = new URL(location.href).searchParams.get("lang");
if (langParam) {
const url = new URL(location.href);
url.searchParams.delete("lang");
history.replaceState(history.state, "", url);
}
// The session language. While the editor panel is open, its language
// selection overrides the normal preference (swapdoc.setLangOverride):
// internal fetches and prefetches follow it until the panel closes and
// the override clears (null restores the initial ?lang=, if any).
let sessionLang = langParam;
addEventListener("pagerite:session-lang", (ev) => {
sessionLang = ev.detail?.lang || langParam;
});
// An internal URL as fetched: carries the session's ?lang= unless the
// link already pins a language of its own. With no ?lang= on the initial
// load nothing is ever added.
const pageUrl = (url) => {
const u = new URL(url, location.href);
if (sessionLang && u.origin === location.origin && !u.searchParams.has("lang")) {
u.searchParams.set("lang", sessionLang);
}
return u;
};
// The in-memory page cache is keyed by path + query: the same pathname
// holds different HTML for each language version.
const rawKey = (url) => {
const u = new URL(url, location.href);
return u.pathname + u.search;
};
const cacheKey = (url) => rawKey(pageUrl(url));
// What goes into the address bar and history: the pretty URL, no ?lang=.
const prettyUrl = (url) => {
const u = pageUrl(url);
u.searchParams.delete("lang");
return u;
};
// Regions every page has. #sidebar is NOT among them: it is omitted // Regions every page has. #sidebar is NOT among them: it is omitted
// entirely when the section has no sub-navigation, and handled below. // entirely when the section has no sub-navigation, and handled below.
const REGIONS = ["page-banner", "nav", "main"]; const REGIONS = ["page-banner", "nav", "main"];
@@ -268,6 +317,41 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
} }
// --- Figure lightbox (click to enlarge) --------------------------------
// Clicking an article figure's image opens it in a full-viewport box:
// the image as large as fits with its caption below; click or Esc
// closes. The zoomed <img> reuses the source URL — /_f/ is immutable
// and content-negotiated, so the large view comes from the browser
// cache at no extra cost.
let lightbox = null;
function closeLightbox() {
lightbox?.remove();
lightbox = null;
}
function openLightbox(figure) {
const img = figure.querySelector("img");
if (!img) return;
closeLightbox();
lightbox = document.createElement("div");
lightbox.id = "lightbox";
const big = document.createElement("img");
big.src = img.currentSrc || img.src;
big.alt = img.alt;
lightbox.append(big);
const cap = figure.querySelector("figcaption");
if (cap) {
const c = document.createElement("div");
c.className = "caption";
c.textContent = cap.textContent;
lightbox.append(c);
}
lightbox.addEventListener("click", closeLightbox);
document.body.append(lightbox);
}
// Any key dismisses the lightbox — Esc included, and the rest would
// only scroll the page behind it anyway.
addEventListener("keydown", () => closeLightbox());
// Tuck the article edit pen at the end of the first h1 (which may come // Tuck the article edit pen at the end of the first h1 (which may come
// from the markdown itself). Re-runs when the editor replaces the // from the markdown itself). Re-runs when the editor replaces the
// previewed article, since that wipes elements inside it. // previewed article, since that wipes elements inside it.
@@ -317,9 +401,12 @@ import "overlayscrollbars/overlayscrollbars.css";
// received it as the document (re-fetching would be redundant, and // received it as the document (re-fetching would be redundant, and
// browser heuristics may send it without if-none-match, defeating the // browser heuristics may send it without if-none-match, defeating the
// conditional request); it enters the cache when navigated to. // conditional request); it enters the cache when navigated to.
const pageCache = new Map(); // pathname -> HTML text const pageCache = new Map(); // rawKey/cacheKey(url) -> HTML text
addEventListener("pagerite:page-fetched", (ev) => { addEventListener("pagerite:page-fetched", (ev) => {
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html); // Key by the URL as announced, exactly as the editor fetched it: a
// copy pinned to a language (?lang=) caches under its own key, where
// navigation with the same session language finds it.
pageCache.set(rawKey(ev.detail.url), ev.detail.html);
}); });
// Editors mutate site-wide state (theme, structure, headings, banners), // Editors mutate site-wide state (theme, structure, headings, banners),
@@ -338,21 +425,22 @@ import "overlayscrollbars/overlayscrollbars.css";
}); });
function preload() { function preload() {
const urls = new Set(); const urls = new Map(); // cache key -> URL, deduped (hashes collapse)
for (const a of document.querySelectorAll( for (const a of document.querySelectorAll(
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]', '#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
)) { )) {
urls.add(a.pathname); const u = pageUrl(a.href);
urls.set(rawKey(u), u);
} }
for (const url of urls) { for (const [key, u] of urls) {
if (pageCache.has(url)) continue; if (pageCache.has(key)) continue;
// x-pagerite-preload: idle cache warm-up, not a page view — the // x-pagerite-preload: idle cache warm-up, not a page view — the
// server excludes these GETs from analytics (the ping sent on actual // server excludes these GETs from analytics (the navigation message
// navigation does the counting). // sent on actual navigation does the counting).
fetch(url, { headers: { "x-pagerite-preload": "1" } }) fetch(u, { headers: { "x-pagerite-preload": "1" } })
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html") .then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
? r.text() : "")) ? r.text() : ""))
.then((html) => { if (html) pageCache.set(url, html); }) .then((html) => { if (html) pageCache.set(key, html); })
.catch(() => {}); .catch(() => {});
} }
} }
@@ -432,61 +520,116 @@ import "overlayscrollbars/overlayscrollbars.css";
}); });
}, { passive: true }); }, { passive: true });
// --- Analytics pings --------------------------------------------------- // --- Analytics over WebSocket ------------------------------------------
// Fire-and-forget POSTs to /_a with the fields as query parameters (a // One /_ws connection follows the whole browsing session: the initial page
// beacon can carry no body, and query args show in server logs next to // load (starts the visit — the server counts nothing from the document GET
// the document GET they refer to): on the initial page load (starts the // alone), internal fetch-navigations, external https exits, and frequent
// visit — the server counts nothing from the document GET alone), for // active reading-time updates. Messages are JSON text frames matching the
// internal fetch-navigations, for external https exits, and on window // server's msgspec Ping struct: {fr?, to?, read?, hide?} — falsy fields
// close. ``read`` is the active time (ms) spent on ``fr``. // are omitted. ``read`` is the active time (ms) accumulated on ``fr`` since
// Reading time pauses after 1 minute of inactivity and resumes on the // the last report; reading time pauses after 1 minute of inactivity and
// next mouse/touch/scroll/keyboard event. // resumes on the next mouse/touch/scroll/keyboard event. While the user is
// Excluded: back/forward (popstate never pings), everything while the // active, accumulated reading time is flushed every few seconds, so a
// disconnect simply leaves the last reported time on the server — no close
// beacon is needed. After 5 minutes without any activity the client closes
// the channel itself (a sleeping tab would lose it anyway); the next
// activity reconnects and the server sees a new session.
// Excluded: back/forward (popstate never reports), everything while the
// editor is open (body.editing — admin noise, not visits), and // editor is open (body.editing — admin noise, not visits), and
// navigations TO the analytics page (/_a — admin machinery, and the // navigations TO the analytics page (/_a — admin machinery). Navigations
// server rejects it as a ping target anyway). Navigations AWAY from /_a // AWAY from /_a must report: load() already fetched the target page
// must ping: load() already fetched the target page without the preload // without the preload header, and without the message that GET would flush
// header, and without the ping that GET would flush to the crawler list. // to the crawler list.
// Admins (when SSO is actually in use — with no auth proxy "admin" is // Admins (when SSO is actually in use — with no auth proxy "admin" is
// everyone's state) ping normally but with hide=1: the server then // everyone's state) report normally but with hide: the server then flags
// records nothing and scrubs any session the same browser accumulated // the client record, scrubbing everything it ever did from the statistics,
// before logging in, so admins never show up as visits or crawlers. // so admins never show up as visits or crawlers.
// See docs/analytics.md. // See docs/analytics.md.
// fetch wrapper: every key of ``params`` becomes a query arg on /_a // The activity WebSocket. Messages sent before the connection opens are
// (falsy values are omitted). Admins get hide=1. ``beacon`` uses // queued (the queue keeps the interim activity). Reconnects are driven by
// sendBeacon when available, for unload-time pings. // user activity only — never by timers while the page sits idle — with an
function pingFetch(params, { beacon = false } = {}) { // exponential falloff between attempts so a failing endpoint cannot make
const query = new URLSearchParams(); // us hammer the server (or trip its security limits). After a longer
if (ssoAvailable && isAdmin) params = { ...params, hide: 1 }; // stretch without any activity we close the socket proactively: the user
for (const [key, value] of Object.entries(params)) { // has moved on and left the tab open (a sleeping browser tab would lose
if (value) query.set(key, value); // the connection anyway), so the next activity reconnects and registers
} // as a fresh session. Analytics must never break navigation: every send
const url = `/_a?${query}`; // is wrapped, and a server without the endpoint just leaves the socket
// failing in the background.
let ws = null;
const wsQueue = [];
const wsPolicy = reconnectPolicy({ min: 1000 });
// The first attempt is staggered too: page load opens several sockets at
// once (Vite's HMR socket, the editors), and the burst trips the browser's
// WebSocket throttling (sockets then sit "pending" for minutes).
let wsNotBefore = Date.now() + socketSlot();
let wsWatchdog = null;
function activityWs() {
if (ws || Date.now() < wsNotBefore) return;
const url = new URL("/_ws", location.href);
url.protocol = url.protocol === "https:" ? "wss:" : "ws:";
try { try {
if (beacon && navigator.sendBeacon) { ws = new WebSocket(url);
navigator.sendBeacon(url); } catch {
} else { return;
fetch(url, { method: "POST", keepalive: true }); }
} clearTimeout(wsWatchdog);
} catch { /* analytics must never break navigation */ } wsWatchdog = watchConnecting(ws, "activity");
ws.onopen = () => {
wsPolicy.opened();
for (const msg of wsQueue.splice(0)) ws.send(JSON.stringify(msg));
};
ws.onclose = () => {
ws = null;
// No timer here: the next user activity retries, after the backoff.
wsNotBefore = Date.now() + wsPolicy.closed();
};
ws.onerror = () => ws.close();
} }
function ping({ to, fr = currentPath, read = 0, beacon = false } = {}) { function report(msg) {
if (document.body.classList.contains("editing")) return; if (document.body.classList.contains("editing")) return;
if (to === "/_a") return; if (msg.to === "/_a") return;
pingFetch({ fr, to, read: Math.round(read / 1000) }, { beacon }); if (ssoAvailable && isAdmin) msg.hide = true;
activityWs();
if (ws?.readyState === WebSocket.OPEN) {
try {
ws.send(JSON.stringify(msg));
return;
} catch { /* fall through to queueing */ }
}
wsQueue.push(msg);
}
function ping({ to, fr = currentPath, read = 0 } = {}) {
// Reading-time updates from the analytics page itself are not tracked
// (/_a is admin machinery; the server would reject the path anyway).
if (!to && currentPath === "/_a") return;
const msg = {};
if (fr) msg.fr = fr;
if (to) msg.to = to;
const secs = Math.round(read / 1000);
if (secs > 0) msg.read = secs;
if (!msg.to && !msg.read) return;
report(msg);
} }
// Active reading time for the current page. The clock stops after 1 minute // Active reading time for the current page. The clock stops after 1 minute
// without activity and restarts on the next mouse/touch/scroll/keyboard // without activity and restarts on the next mouse/touch/scroll/keyboard
// event. // event. Every READ_FLUSH_MS of accumulated activity is reported. After
// IDLE_MS with no activity at all, the remaining read time is flushed and
// the WebSocket is closed: the user has moved on, and the next activity
// reconnects as a new session.
const INACTIVE_MS = 60_000; const INACTIVE_MS = 60_000;
const READ_FLUSH_MS = 5_000;
const IDLE_MS = 5 * 60_000;
let readStart = performance.now(); let readStart = performance.now();
let readElapsed = 0; let readElapsed = 0;
let reading = true; let reading = true;
let readInactivityTimer = null; let readInactivityTimer = null;
let closePingedFor = null; let idleTimer = null;
function markReadActivity() { function markReadActivity() {
if (!reading) { if (!reading) {
@@ -500,6 +643,24 @@ import "overlayscrollbars/overlayscrollbars.css";
reading = false; reading = false;
} }
}, INACTIVE_MS); }, INACTIVE_MS);
// Any activity is a sign of life: (re)connect the channel if it was
// dropped or idle-closed (not while editing — admin noise), and push
// the idle disconnect forward.
if (!document.body.classList.contains("editing")) activityWs();
clearTimeout(idleTimer);
idleTimer = setTimeout(() => {
if (ws) {
// Flush what is left unsent, then hang up. report() would try to
// reconnect a dead socket, which is exactly what we avoid here.
const left = takeReadTime();
if (Math.round(left / 1000) > 0) ping({ read: left });
ws.close();
}
}, IDLE_MS);
// Frequently update the article read time on the server.
if (readElapsed + performance.now() - readStart >= READ_FLUSH_MS) {
ping({ read: takeReadTime() });
}
} }
function takeReadTime() { function takeReadTime() {
@@ -519,28 +680,19 @@ import "overlayscrollbars/overlayscrollbars.css";
clearTimeout(readInactivityTimer); clearTimeout(readInactivityTimer);
} }
function sendClosePing() {
if (closePingedFor === currentPath) return;
closePingedFor = currentPath;
const read = takeReadTime();
if (Math.round(read / 1000) <= 0) return;
ping({ read, beacon: true });
}
for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) { for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) {
addEventListener(ev, markReadActivity, { passive: true }); addEventListener(ev, markReadActivity, { passive: true });
} }
addEventListener("pagehide", sendClosePing);
// The initial page load pings too — it is what starts the visit and // The initial page load reports too — it is what starts the visit and
// counts the entry page view (the document GET alone records nothing). // counts the entry page view (the document GET alone records nothing).
// It carries only ``to``: the server attributes the entry to the referer // It carries only ``to``: the server attributes the entry to the referer
// it saw on the document GET (unavailable to JS once loaded), and an // it saw on the document GET (unavailable to JS once loaded), and an
// ``fr`` equal to ``to`` would log a bogus self-transition when a // ``fr`` equal to ``to`` would log a bogus self-transition when a
// session already exists (e.g. a second tab). // session already exists (e.g. a second tab).
// Sent once per load, after the auth probes so the admin gate applies; // Sent once per load, after the auth probes so the admin gate applies;
// the pageshow re-probe must not ping again. Reloads are not visits: // the pageshow re-probe must not report again. Reloads are not visits:
// pinging them would double-count the view and log a self-transition. // reporting them would double-count the view and log a self-transition.
let entryPinged = false; let entryPinged = false;
function pingEntryOnce() { function pingEntryOnce() {
if (entryPinged) return; if (entryPinged) return;
@@ -603,12 +755,12 @@ import "overlayscrollbars/overlayscrollbars.css";
teardownAnalytics(); teardownAnalytics();
let doc; let doc;
let finalUrl = url; let finalUrl = url;
const cached = !editing && pageCache.get(new URL(url, location.href).pathname); const cached = !editing && pageCache.get(cacheKey(url));
if (cached) { if (cached) {
doc = new DOMParser().parseFromString(cached, "text/html"); doc = new DOMParser().parseFromString(cached, "text/html");
} else { } else {
try { try {
const res = await fetch(url); const res = await fetch(pageUrl(url));
const type = res.headers.get("content-type") || ""; const type = res.headers.get("content-type") || "";
if (!res.ok || !type.includes("text/html")) throw new Error("not a page"); if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
// Reflect any redirect the server issued. // Reflect any redirect the server issued.
@@ -616,15 +768,15 @@ import "overlayscrollbars/overlayscrollbars.css";
const html = await res.text(); const html = await res.text();
// Populate the cache too, so returning here (back/forward, or a // Populate the cache too, so returning here (back/forward, or a
// self-link in the nav) is served from memory. // self-link in the nav) is served from memory.
pageCache.set(new URL(finalUrl, location.href).pathname, html); pageCache.set(cacheKey(finalUrl), html);
doc = new DOMParser().parseFromString(html, "text/html"); doc = new DOMParser().parseFromString(html, "text/html");
} catch { } catch {
location.href = url; // fall back to a normal navigation location.href = pageUrl(url); // fall back to a normal navigation
return false; return false;
} }
} }
if (REGIONS.some((id) => !doc.getElementById(id))) { if (REGIONS.some((id) => !doc.getElementById(id))) {
location.href = url; location.href = pageUrl(url);
return false; return false;
} }
const doit = () => { const doit = () => {
@@ -680,6 +832,11 @@ import "overlayscrollbars/overlayscrollbars.css";
// stylesheet after the server-rendered tag. // stylesheet after the server-rendered tag.
const userStyle = document.getElementById("pagerite-user"); const userStyle = document.getElementById("pagerite-user");
if (userStyle) document.head.appendChild(userStyle); if (userStyle) document.head.appendChild(userStyle);
// The served language rides on <html> (lang + dir, rtl for e.g.
// Arabic) — follow the swapped page (the editor panel carries its
// own lang="en" dir="ltr", so it is unaffected).
document.documentElement.lang = doc.documentElement.lang;
document.documentElement.dir = doc.documentElement.dir;
document.title = doc.title; document.title = doc.title;
// Banners may contain scripts (canvas etc.), content pages may too. // Banners may contain scripts (canvas etc.), content pages may too.
runScripts(document.getElementById("page-banner")); runScripts(document.getElementById("page-banner"));
@@ -704,7 +861,7 @@ import "overlayscrollbars/overlayscrollbars.css";
doit(); doit();
} }
currentPath = new URL(finalUrl, location.href).pathname; currentPath = new URL(finalUrl, location.href).pathname;
if (push) history.pushState({ idx: ++historyIdx }, "", finalUrl); if (push) history.pushState({ idx: ++historyIdx }, "", prettyUrl(finalUrl));
// The open editor follows the URL: retarget the per-page tabs to the // The open editor follows the URL: retarget the per-page tabs to the
// navigated-to page (unsaved text of the previous page is discarded — // navigated-to page (unsaved text of the previous page is discarded —
// the article it previewed into is gone). // the article it previewed into is gone).
@@ -763,6 +920,13 @@ import "overlayscrollbars/overlayscrollbars.css";
.catch((e) => console.error("editor load failed:", e)); .catch((e) => console.error("editor load failed:", e));
return; return;
} }
// Article figure images enlarge into the lightbox.
const fig = ev.target.closest("#main article figure");
if (fig && ev.target.closest("img")) {
ev.preventDefault();
openLightbox(fig);
return;
}
const a = ev.target.closest("a[href]"); const a = ev.target.closest("a[href]");
if (!a || a.target || a.hasAttribute("download")) return; if (!a || a.target || a.hasAttribute("download")) return;
const url = new URL(a.href, location.href); const url = new URL(a.href, location.href);
@@ -770,7 +934,6 @@ import "overlayscrollbars/overlayscrollbars.css";
// External link: the browser navigates; record the full https URL so // External link: the browser navigates; record the full https URL so
// different links to the same domain stay distinct in analytics. // different links to the same domain stay distinct in analytics.
if (url.protocol === "https:") { if (url.protocol === "https:") {
closePingedFor = currentPath;
ping({ to: url.href, read: takeReadTime() }); ping({ to: url.href, read: takeReadTime() });
} }
return; return;
@@ -794,7 +957,6 @@ import "overlayscrollbars/overlayscrollbars.css";
const from = currentPath; const from = currentPath;
load(url).then((ok) => { load(url).then((ok) => {
if (!ok) return; if (!ok) return;
closePingedFor = null;
ping({ to: url.pathname, fr: from, read: takeReadTime() }); ping({ to: url.pathname, fr: from, read: takeReadTime() });
resetReadTime(); resetReadTime();
}); });
+56
View File
@@ -0,0 +1,56 @@
// Shared reconnect policy for the WebSockets (page/banner editors,
// analytics view, the activity channel). Two things trip a browser's
// WebSocket throttling, after which every socket to the host sits
// "pending" (never opens, never closes) for minutes:
//
// 1. A burst of simultaneous attempts — page load opens Vite's HMR
// socket plus several of ours at the same moment, and every refresh
// repeats the burst. socketSlot() spaces new sockets out.
// 2. Too-frequent retries — so failed attempts back off exponentially
// (a few seconds, doubling to half a minute), reset only after a
// connection stayed open long enough to count as healthy. A socket
// that closes right after opening must NOT reset the backoff.
export function reconnectPolicy({ min = 2000, max = 30000, healthyAfter = 30000 } = {}) {
let delay = min
let openedAt = 0
return {
// Stamp a socket that just opened.
opened() {
openedAt = Date.now()
},
// The socket closed: the wait before the next attempt (up to 50%
// jitter; the base doubles per failure). A healthy streak resets it.
closed() {
if (openedAt && Date.now() - openedAt >= healthyAfter) delay = min
openedAt = 0
const wait = Math.round(delay * (1 + Math.random() * 0.5))
delay = Math.min(delay * 2, max)
return wait
},
}
}
// Sockets created at the same moment (page load: Vite's HMR socket plus
// ours) read as one burst to the browser's throttling. Space new sockets
// out: each call reserves a slot a beat after the previous one.
let nextSlot = 0
export function socketSlot() {
const now = Date.now()
const wait = Math.max(0, nextSlot - now)
nextSlot = Math.max(now, nextSlot) + 300
return wait
}
// A socket still CONNECTING after this long counts as a failed attempt:
// browser throttling leaves sockets "pending" (no open, no close) for
// minutes, and without a watchdog the app would wait on one forever (the
// recurring empty editor). Closing it fires onclose, which reschedules
// through the policy's backoff — it never reconnects aggressively itself.
export function watchConnecting(ws, label) {
return setTimeout(() => {
if (ws.readyState === WebSocket.CONNECTING) {
console.warn(`[pagerite] ${label} socket stuck connecting — closing it, retrying with backoff`)
ws.close()
}
}, 10_000)
}
+28 -3
View File
@@ -11,6 +11,19 @@ export function dropPageCache() {
dispatchEvent(new CustomEvent('pagerite:drop-page-cache')) dispatchEvent(new CustomEvent('pagerite:drop-page-cache'))
} }
// The editor's language override (set by EditorShell): while the panel is
// open, its language selection wins over the normal preferences (?lang= /
// Accept-Language) — every in-place re-render asks for that language
// explicitly, and pagerite.js applies it to its own fetches and prefetches
// (pagerite:session-lang). The primary selection pins by its code:
// ?lang=<primary> selects the original explicitly (i18n.select_language).
let overrideLang = null // the ?lang= value in force, null = normal prefs
export function setLangOverride(queryLang) {
overrideLang = queryLang || null
dispatchEvent(new CustomEvent('pagerite:session-lang', { detail: { lang: overrideLang } }))
}
export function runScripts(root) { export function runScripts(root) {
// Scripts injected via innerHTML do not execute; re-create them. // Scripts injected via innerHTML do not execute; re-create them.
if (!root) return if (!root) return
@@ -101,6 +114,11 @@ function swapRegions(doc) {
} }
anchor = imported anchor = imported
} }
// The served language rides on <html> (lang + dir, rtl for e.g. Arabic):
// follow the swapped page. The editor panel carries its own lang="en"
// dir="ltr", so it is unaffected.
document.documentElement.lang = doc.documentElement.lang
document.documentElement.dir = doc.documentElement.dir
// The editor keeps its own title while open; only inherit the server title // The editor keeps its own title while open; only inherit the server title
// when navigating outside the editor (e.g. fetch-navigation swaps). // when navigating outside the editor (e.g. fetch-navigation swaps).
if (!document.body.classList.contains('editing')) { if (!document.body.classList.contains('editing')) {
@@ -111,13 +129,14 @@ function swapRegions(doc) {
// Fetch /p, swap its regions into the live page and replaceState to it. // Fetch /p, swap its regions into the live page and replaceState to it.
// Returns the final URL (after redirects), or null when the fetch did not // Returns the final URL (after redirects), or null when the fetch did not
// yield a page. Category and missing URLs render a placeholder 404 page — // yield a page. Category and missing URLs render a placeholder 404 page —
// fine to swap in (new pages are created by editing them). // fine to swap in (new pages are created by editing them). While the
// editor's language override is set the fetch pins that language.
export async function loadPlain(p) { export async function loadPlain(p) {
let doc let doc
let finalUrl = `/${p}` let finalUrl = `/${p}`
let html let html
try { try {
const res = await fetch(finalUrl) const res = await fetch(overrideLang ? `${finalUrl}?lang=${overrideLang}` : finalUrl)
const type = res.headers.get('content-type') || '' const type = res.headers.get('content-type') || ''
if (!type.includes('text/html')) return null if (!type.includes('text/html')) return null
if (res.redirected) finalUrl = res.url if (res.redirected) finalUrl = res.url
@@ -126,10 +145,16 @@ export async function loadPlain(p) {
} catch { return null } } catch { return null }
if (!doc.getElementById('main')) return null if (!doc.getElementById('main')) return null
swapRegions(doc) swapRegions(doc)
history.replaceState(history.state, '', finalUrl) // The address bar keeps the pretty URL: a language query is a fetch
// detail, never shown (pagerite.js's initial ?lang= works the same).
const pretty = new URL(finalUrl, location.href)
pretty.searchParams.delete('lang')
history.replaceState(history.state, '', pretty)
runScripts(document.getElementById('page-banner')) runScripts(document.getElementById('page-banner'))
runScripts(document.getElementById('main')) runScripts(document.getElementById('main'))
// Keep pagerite.js's in-memory page cache in sync with the fresh copy. // Keep pagerite.js's in-memory page cache in sync with the fresh copy.
// The URL is announced as fetched: a language-pinned copy caches under
// its own ?lang= key, where navigation with the same pin finds it.
dispatchEvent(new CustomEvent('pagerite:page-fetched', { detail: { url: finalUrl, html } })) dispatchEvent(new CustomEvent('pagerite:page-fetched', { detail: { url: finalUrl, html } }))
dispatchEvent(new CustomEvent('pagerite:preview')) // re-inject + re-tuck the edit pens dispatchEvent(new CustomEvent('pagerite:preview')) // re-inject + re-tuck the edit pens
return finalUrl return finalUrl
+2
View File
@@ -5,6 +5,7 @@
* Configures Vite for FastAPI backend integration: * Configures Vite for FastAPI backend integration:
* - Proxies /api/* requests to the FastAPI backend * - Proxies /api/* requests to the FastAPI backend
* - Builds to the Python module's frontend-build directory * - Builds to the Python module's frontend-build directory
* - Disables Vite's screen clearing on startup
* *
* Options: * Options:
* paths - Array of paths to proxy (default: ["/api"]) * paths - Array of paths to proxy (default: ["/api"])
@@ -26,6 +27,7 @@ export default function fastapiVue({ paths = ["/api"] } = {}) {
return { return {
name: "vite-plugin-fastapi-pagerite", name: "vite-plugin-fastapi-pagerite",
config: () => ({ config: () => ({
clearScreen: false,
server: { proxy }, server: { proxy },
build: { build: {
outDir: "../pagerite/frontend-build", outDir: "../pagerite/frontend-build",
+10 -1
View File
@@ -16,7 +16,7 @@ const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
// https://vite.dev/config/ // https://vite.dev/config/
export default defineConfig({ export default defineConfig({
plugins: [ plugins: [
fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_fonts", "/_a"] }), fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_fonts", "/_a", "/_ws", "/_translate"] }),
vue(), vue(),
vueDevTools(), vueDevTools(),
], ],
@@ -26,7 +26,16 @@ export default defineConfig({
}, },
}, },
appType: 'mpa', // no SPA fallback; every HTML page is served by FastAPI appType: 'mpa', // no SPA fallback; every HTML page is served by FastAPI
resolve: {
alias: {
// All components are precompiled SFCs — drop the runtime template
// compiler (~60 kB min) from the bundle.
vue: 'vue/dist/vue.runtime.esm-bundler.js',
},
},
build: { build: {
// The main editor bundle (CodeMirror + Vue) is intentionally one chunk.
chunkSizeWarningLimit: 1200,
// Mirror the URL space in the build output: hashed files land under // Mirror the URL space in the build output: hashed files land under
// frontend-build/_assets/ and the Frontend serves the build directory // frontend-build/_assets/ and the Frontend serves the build directory
// at the site root (frontend/public/favicon.ico -> /favicon.ico). // at the site root (frontend/public/favicon.ico -> /favicon.ico).
+20 -7
View File
@@ -9,6 +9,7 @@ from pathlib import Path
import httpx import httpx
from fastapi_vue import server from fastapi_vue import server
from fastapi_vue.hostutil import parse_endpoints
DEFAULT_PORT = 8100 DEFAULT_PORT = 8100
DEVMODE = os.getenv("PAGERITE_DEV") == "1" DEVMODE = os.getenv("PAGERITE_DEV") == "1"
@@ -32,7 +33,9 @@ def _download_dbip() -> None:
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*") for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
) )
if existing and existing[-1] >= months[0]: if existing and existing[-1] >= months[0]:
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download") print(
f"pagerite: DB-IP database is current ({existing[-1]}), skipping download"
)
return return
for month in months: for month in months:
@@ -57,7 +60,10 @@ def _download_dbip() -> None:
with gzip.open(tmp, "rb") as f: with gzip.open(tmp, "rb") as f:
f.read(1) f.read(1)
except OSError: except OSError:
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr) print(
f"pagerite: DB-IP download for {month} was not valid gzip",
file=sys.stderr,
)
tmp.unlink(missing_ok=True) tmp.unlink(missing_ok=True)
continue continue
os.replace(tmp, target) os.replace(tmp, target)
@@ -77,9 +83,11 @@ def main() -> None:
"hostname", "hostname",
nargs="?", nargs="?",
default="localhost", default="localhost",
help=("Public hostname of the site; names the data directory " help=(
"<hostname>/{content.kantadb, analytics.json, files} under the " "Public hostname of the site; names the data directory "
"cwd (default: localhost)."), "<hostname>/{content.kantadb, analytics.json, files} under the "
"cwd (default: localhost)."
),
) )
parser.add_argument( parser.add_argument(
"-l", "-l",
@@ -96,15 +104,20 @@ def main() -> None:
# Export the hostname before pagerite.app is imported: it derives the # Export the hostname before pagerite.app is imported: it derives the
# data directory and public origin from it at import time. # data directory and public origin from it at import time.
os.environ["PAGERITE_HOSTNAME"] = args.hostname os.environ["PAGERITE_HOSTNAME"] = args.hostname
# And the listen port: the app prints the translator WS URL at startup,
# which for localhost includes the actual port.
for endpoint in parse_endpoints(args.listen, DEFAULT_PORT):
if "port" in endpoint:
os.environ["PAGERITE_PORT"] = str(endpoint["port"])
break
if args.dbip: if args.dbip:
_download_dbip() _download_dbip()
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
server.run( server.run(
"pagerite.app:app", "pagerite.app:app",
listen=args.listen, listen=args.listen,
default_port=DEFAULT_PORT, default_port=DEFAULT_PORT,
server_header=False, server_header=False,
**dev, reload=Path(__file__).parent if DEVMODE else False,
) )
+124 -43
View File
@@ -1,18 +1,27 @@
"""Server-side visit analytics (collection only; see docs/analytics.md). """Server-side visit analytics (collection only; see docs/analytics.md).
Events come from navigation pings POSTed to /_a by pagerite.js: the first Events come from pagerite.js over the /_ws WebSocket (``Ping`` messages as
ping on page load starts a visit, later pings extend it, and pings with no JSON text frames): the first navigation message on page load starts a
known session start a fresh one (missing data, not dropped). The document visit, later messages extend it, and messages with no known session start a
fresh one (missing data, not dropped). Active reading time is reported as
frequent ``read`` updates while the user is active; the times are
cumulative per trail item, so a disconnect simply leaves the last logged
time in place. The document
GET handler stashes the entry referer (external https origin) and any GET handler stashes the entry referer (external https origin) and any
utm_* query parameters in in-memory IP tables, consumed when the ping utm_* query parameters in in-memory IP tables, consumed when the first
starts the visit; nothing is counted without a ping (plain bots that only message starts the visit; nothing is counted without a message (plain bots
fetch documents end up in the crawler list). JS-running crawlers that only fetch documents end up in the crawler list). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their UA gives them (Googlebot, GoogleOther, Applebot, ...) do connect and send messages, but
away (``_is_bot_ua``) and their pings are ignored, so they land in the their UA gives them
crawler list too. Idle-time link preloads from pagerite.js carry an away (``_is_bot_ua``) and their messages are ignored, so they land in the
``x-pagerite-preload`` header and are not tracked at all — the ping sent crawler list too. Bots whose UA does not match are caught by engagement:
when the user actually navigates does the counting. a visit whose total reported reading time is under ``_MIN_VISIT_READ``
Admin clients ping with ``hide=1``: the client record is flagged ``hide``, seconds is reclassified as crawler hits at display time (durations are
client-provided and trusted — real-browser bots report 02 s), so it never
counts in the visit aggregates either. Idle-time link preloads from pagerite.js carry an
``x-pagerite-preload`` header and are not tracked at all — the navigation
message sent when the user actually navigates does the counting.
Admin clients send ``hide``: the client record is flagged ``hide``,
which covers everything that client ever did — visits and crawler hits which covers everything that client ever did — visits and crawler hits
from before the login included. Aggregates (site visits, page views, from before the login included. Aggregates (site visits, page views,
transitions) are not stored; they are computed at display time from the transitions) are not stored; they are computed at display time from the
@@ -38,7 +47,7 @@ from collections.abc import Callable
from contextlib import suppress from contextlib import suppress
from datetime import UTC, datetime, timedelta from datetime import UTC, datetime, timedelta
from pathlib import Path from pathlib import Path
from urllib.parse import parse_qs, urlparse from urllib.parse import parse_qs, urlencode, urlparse
import blake3 import blake3
import msgspec import msgspec
@@ -70,6 +79,27 @@ def _compact_user_agent(ua: str) -> str:
return " ".join(p for p in parts if p).strip() return " ".join(p for p in parts if p).strip()
class Ping(msgspec.Struct, omit_defaults=True):
"""One client message on the /_ws activity WebSocket.
Sent as a JSON text frame (msgspec-encoded, decoded to str for the
wire). ``to`` set: a navigation — internal page path or external https
exit URL. ``read`` alone (with ``fr``): an active reading-time update
for the page ``fr``; these arrive frequently while the user is active
and accumulate on the trail item. ``hide`` flags the client as an
admin: everything it ever did is excluded from the statistics.
"""
#: Path of the page the activity happened on ("" for the initial load).
fr: str = ""
#: Navigation target: internal path or external https exit URL.
to: str = ""
#: Active reading time (seconds) spent on ``fr`` since the last report.
read: int = 0
#: Admin client: record but hide everything from the statistics.
hide: bool = False
class Client(msgspec.Struct, omit_defaults=True): class Client(msgspec.Struct, omit_defaults=True):
"""Client metadata shared by visits, crawler hits and abuse hits. """Client metadata shared by visits, crawler hits and abuse hits.
@@ -151,7 +181,7 @@ class Visit(msgspec.Struct, omit_defaults=True):
class CrawlerHit(msgspec.Struct, omit_defaults=True): class CrawlerHit(msgspec.Struct, omit_defaults=True):
"""A document GET that was never followed by an analytics ping. """A document GET that was never followed by an activity message.
Client metadata is held in ``Analytics.clients`` keyed by ``client``. Client metadata is held in ``Analytics.clients`` keyed by ``client``.
""" """
@@ -174,8 +204,10 @@ class AbuseHit(msgspec.Struct, omit_defaults=True):
Unlike crawler hits the full request path (query string included) is Unlike crawler hits the full request path (query string included) is
kept: the interesting part is exactly which paths were probed. kept: the interesting part is exactly which paths were probed.
``flag`` marks the path that triggered classification; ``is_404`` ``flag`` marks the path that triggered classification; ``is_404``
distinguishes 404 responses from document GETs made by the abuser. distinguishes 404 responses (probed paths and 404-fallback document
Client metadata is held in ``Analytics.clients`` keyed by ``client``. GETs) from real 200 document GETs — the abuser actually reading
articles. Client metadata is held in ``Analytics.clients`` keyed by
``client``.
""" """
start: datetime start: datetime
@@ -186,7 +218,7 @@ class AbuseHit(msgspec.Struct, omit_defaults=True):
#: True when this path triggered abuse classification (telltale path #: True when this path triggered abuse classification (telltale path
#: or the 404 that crossed the threshold). #: or the 404 that crossed the threshold).
flag: bool = False flag: bool = False
#: True for 404 responses; false for document GETs from the abuser. #: True for 404 responses; false for real (200) document GETs.
is_404: bool = False is_404: bool = False
@@ -321,6 +353,11 @@ def _utm_tags(query: str) -> dict[str, str]:
_CRAWLER_TIMEOUT = timedelta(seconds=10) _CRAWLER_TIMEOUT = timedelta(seconds=10)
#: Minimum total reported reading time (seconds, summed over the trail) for
#: a session to count as a visit; shorter sessions are JS-running bots and
#: are shown as crawler hits instead.
_MIN_VISIT_READ = 5
#: How long a failed favicon fetch suppresses retries for the same origin. #: How long a failed favicon fetch suppresses retries for the same origin.
_FAVICON_RETRY = timedelta(days=7) _FAVICON_RETRY = timedelta(days=7)
@@ -336,6 +373,7 @@ def _is_bot_ua(ua: str) -> bool:
"""True when the UA claims a crawler identity (bot or spider).""" """True when the UA claims a crawler identity (bot or spider)."""
return bool(_BOT_UA.search(ua)) return bool(_BOT_UA.search(ua))
#: Plain-404 count per IP that classifies it as abuse even without a #: Plain-404 count per IP that classifies it as abuse even without a
#: telltale path hit. #: telltale path hit.
_ABUSE_404_THRESHOLD = 10 _ABUSE_404_THRESHOLD = 10
@@ -373,9 +411,7 @@ def _client_hash(ip: str, ua: str, lang: str) -> bytes:
The key is the prettified IP (IPv6 /64), the raw UA string and the The key is the prettified IP (IPv6 /64), the raw UA string and the
extracted language tag, separated by null bytes. extracted language tag, separated by null bytes.
""" """
return blake3.blake3( return blake3.blake3(f"{_network_ip(ip)}\0{ua}\0{lang}".encode()).digest()[:6]
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
).digest()[:6]
class Store: class Store:
@@ -392,19 +428,19 @@ class Store:
#: client hash -> index of the current visit in data.visits #: client hash -> index of the current visit in data.visits
self.sessions: dict[bytes, int] = {} self.sessions: dict[bytes, int] = {}
#: ip -> external https origin of the latest document GET carrying #: ip -> external https origin of the latest document GET carrying
#: one, stashed for the visit the client's initial ping starts. #: one, stashed for the visit the client's initial message starts.
#: Internal or absent referers never touch the table. #: Internal or absent referers never touch the table.
self.pending_referers: dict[str, str] = {} self.pending_referers: dict[str, str] = {}
#: ip -> utm_* query parameters from the latest document GET that #: ip -> utm_* query parameters from the latest document GET that
#: carried any, stashed for the visit the client's initial ping starts. #: carried any, stashed for the visit the client's initial message starts.
#: Only non-empty sets are stored, so a later parameter-less page #: Only non-empty sets are stored, so a later parameter-less page
#: does not overwrite an earlier tagged landing URL. #: does not overwrite an earlier tagged landing URL.
self.pending_utms: dict[str, dict[str, str]] = {} self.pending_utms: dict[str, dict[str, str]] = {}
#: Document GETs that have not yet been matched by a ping. Kept #: Document GETs that have not yet been matched by a message. Kept
#: in RAM only; expired entries are written to ``data.crawlers``. #: in RAM only; expired entries are written to ``data.crawlers``.
self.pending_crawlers: list[CrawlerHit] = [] self.pending_crawlers: list[CrawlerHit] = []
#: client hash -> {path: status} for recent document GETs, consumed #: client hash -> {path: status} for recent document GETs, consumed
#: by the matching ping to record the status of each visited path. #: by the matching message to record the status of each visited path.
self.pending_statuses: dict[bytes, dict[str, int]] = {} self.pending_statuses: dict[bytes, dict[str, int]] = {}
#: ip -> number of plain (non-telltale) 404s seen, in RAM only; #: ip -> number of plain (non-telltale) 404s seen, in RAM only;
#: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse. #: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse.
@@ -480,16 +516,43 @@ class Store:
def display(self) -> Display: def display(self) -> Display:
"""Build the viewer payload, excluding hidden clients. """Build the viewer payload, excluding hidden clients.
Visits whose total reported reading time is under
``_MIN_VISIT_READ`` seconds are JS-running bots, not readers: they
are converted to crawler hits (one per internal trail page) and left
out of the visit list and every aggregate.
The aggregates (site visits, page views, transitions) are computed The aggregates (site visits, page views, transitions) are computed
here from the visit records rather than stored, so a client that here from the visit records rather than stored, so a client that
becomes hidden after navigations were already logged disappears becomes hidden after navigations were already logged disappears
from every statistic. Internal-path navigations count as page from every statistic. Internal-path navigations count as page
views; external https targets are transitions only. views; external https targets are transitions only.
""" """
visits = [v for v in self.data.visits if not self._hidden(v.client)] visits: list[Visit] = []
crawlers = [h for h in self.data.crawlers if not self._hidden(h.client)]
for visit in self.data.visits:
if self._hidden(visit.client):
continue
if sum(item.read for item in visit.trail.values()) >= _MIN_VISIT_READ:
visits.append(visit)
continue
query = urlencode(visit.utm)
first = True
for t, item in visit.trail.items():
if not item.to.startswith("/"):
continue
crawlers.append(
CrawlerHit(
start=t,
entry=item.to,
client=visit.client,
referer=visit.referer if first else "",
query=query if first else "",
status=item.status,
)
)
first = False
display = Display( display = Display(
visits=visits, visits=visits,
crawlers=[h for h in self.data.crawlers if not self._hidden(h.client)], crawlers=crawlers,
abuse=[h for h in self.data.abuse if not self._hidden(h.client)], abuse=[h for h in self.data.abuse if not self._hidden(h.client)],
clients={h: c for h, c in self.data.clients.items() if not c.hide}, clients={h: c for h, c in self.data.clients.items() if not c.hide},
favicons={ favicons={
@@ -512,7 +575,9 @@ class Store:
if nav.to.startswith("/"): if nav.to.startswith("/"):
nav_views = display.views.setdefault(nav.to, {}) nav_views = display.views.setdefault(nav.to, {})
nav_views[nb] = nav_views.get(nb, 0) + 1 nav_views[nb] = nav_views.get(nb, 0) + 1
nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(nav.to, {}) nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(
nav.to, {}
)
nbuckets[nb] = nbuckets.get(nb, 0) + 1 nbuckets[nb] = nbuckets.get(nb, 0) + 1
return display return display
@@ -574,7 +639,9 @@ class Store:
def favicon_origins_needed(self) -> list[str]: def favicon_origins_needed(self) -> list[str]:
"""External https origins seen in visits whose favicon needs fetching. """External https origins seen in visits whose favicon needs fetching.
Covers visit referers and external exit targets (trail and navs). Covers visit referers and external exit targets (trail and navs),
plus crawler-hit referers — spiders often advertise their own site
as the referer, so the icon identifies them in the crawler table.
Origins with a stored icon, or a miss younger than Origins with a stored icon, or a miss younger than
``_FAVICON_RETRY``, are skipped. ``_FAVICON_RETRY``, are skipped.
""" """
@@ -586,6 +653,9 @@ class Store:
origin = _origin(target.to) origin = _origin(target.to)
if origin is not None: if origin is not None:
origins.add(origin) origins.add(origin)
for hit in self.data.crawlers:
if hit.referer:
origins.add(hit.referer)
now = datetime.now(UTC) now = datetime.now(UTC)
return [ return [
origin origin
@@ -638,21 +708,29 @@ class Store:
self.data.abuse_ips[ip] = True self.data.abuse_ips[ip] = True
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip] moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
if moved: if moved:
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip] self.data.crawlers = [
h for h in self.data.crawlers if self._client_ip(h.client) != ip
]
for h in moved: for h in moved:
self._abuse_hit( self._abuse_hit(
h.client, h.client,
h.entry + (f"?{h.query}" if h.query else ""), h.entry + (f"?{h.query}" if h.query else ""),
start=h.start, start=h.start,
is_404=h.status != 200,
) )
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip] pending = [
h for h in self.pending_crawlers if self._client_ip(h.client) == ip
]
if pending: if pending:
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip] self.pending_crawlers = [
h for h in self.pending_crawlers if self._client_ip(h.client) != ip
]
for h in pending: for h in pending:
self._abuse_hit( self._abuse_hit(
h.client, h.client,
h.entry + (f"?{h.query}" if h.query else ""), h.entry + (f"?{h.query}" if h.query else ""),
start=h.start, start=h.start,
is_404=h.status != 200,
) )
self._abuse_hit(client_hash, path, flag=flag, is_404=is_404) self._abuse_hit(client_hash, path, flag=flag, is_404=is_404)
self._save() self._save()
@@ -721,16 +799,16 @@ class Store:
) -> list[bytes]: ) -> list[bytes]:
"""Stash the entry referer/UTM tags and queue a pending crawler hit. """Stash the entry referer/UTM tags and queue a pending crawler hit.
Nothing is counted here — the client's initial /_a ping starts the Nothing is counted here — the client's first /_ws message starts the
visit (only non-admin clients ping). Only a cross-origin https visit (only non-admin clients report). Only a cross-origin https
referer updates the table; an internal or absent referer leaves any referer updates the table; an internal or absent referer leaves any
stashed origin untouched. UTM parameters are kept only when the stashed origin untouched. UTM parameters are kept only when the
landing URL actually carries them, so a subsequent parameter-less page landing URL actually carries them, so a subsequent parameter-less page
does not erase an earlier tagged landing. does not erase an earlier tagged landing.
Every document GET is also queued as a pending crawler hit. If a ping Every document GET is also queued as a pending crawler hit. If a
from the same client arrives within ``_CRAWLER_TIMEOUT``, the hit is message from the same client arrives within ``_CRAWLER_TIMEOUT``, the
discarded; otherwise it is flushed to ``data.crawlers``. The hit is discarded; otherwise it is flushed to ``data.crawlers``. The
Accept-Language header is stored on the client record immediately; Accept-Language header is stored on the client record immediately;
host/geoip are filled in later by async enrichment. host/geoip are filled in later by async enrichment.
@@ -746,7 +824,7 @@ class Store:
client_hash = self._ensure_client(ip, ua, lang, country=country) client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips: if ip in self.data.abuse_ips:
flushed = self._flush_crawlers() flushed = self._flush_crawlers()
self._abuse_hit(client_hash, full_path, is_404=False, flag=False) self._abuse_hit(client_hash, full_path, is_404=status != 200, flag=False)
self._save() self._save()
return flushed return flushed
now = datetime.now(UTC) now = datetime.now(UTC)
@@ -794,15 +872,16 @@ class Store:
hide: bool = False, hide: bool = False,
read: int = 0, read: int = 0,
) -> tuple[int | None, list[bytes]]: ) -> tuple[int | None, list[bytes]]:
"""Record a client navigation ping ({from, to, read} from pagerite.js). """Record a client activity message (``Ping`` from pagerite.js over /_ws).
``to`` is an internal path ("/...") or an https URL for exit links; a ``to`` is an internal path ("/...") or an https URL for exit links; a
missing/empty ``to`` means the page is being closed and only the missing/empty ``to`` means a pure reading-time update and only the
``read`` time should be recorded. The transition is always counted when ``read`` time should be recorded. The transition is always counted when
``to`` is present; the trail only grows on first sight of a page within ``to`` is present; the trail only grows on first sight of a page within
the visit. ``read`` is the active time (seconds) spent on ``from_``. the visit. ``read`` is the active time (seconds) spent on ``from_``
since the previous report.
A ping with no known session starts a fresh visit, consuming the A message with no known session starts a fresh visit, consuming the
referer and UTM tags stashed by the document GET if there are any. referer and UTM tags stashed by the document GET if there are any.
``hide`` is set by admin clients: the client record is flagged ``hide`` is set by admin clients: the client record is flagged
@@ -811,7 +890,7 @@ class Store:
normally. Hidden clients are excluded from every statistic and list normally. Hidden clients are excluded from every statistic and list
at display time, and their pending crawler hits are discarded. at display time, and their pending crawler hits are discarded.
Pings from IPs classified as abuse, and pings whose User-Agent Messages from IPs classified as abuse, and messages whose User-Agent
claims a JS-running crawler identity (``_is_bot_ua``), are ignored claims a JS-running crawler identity (``_is_bot_ua``), are ignored
entirely — the crawler's pending hits stay queued and flush to entirely — the crawler's pending hits stay queued and flush to
``data.crawlers`` normally. ``data.crawlers`` normally.
@@ -890,5 +969,7 @@ class Store:
else: else:
visit.trail[now] = TrailItem(to=target, status=target_status) visit.trail[now] = TrailItem(to=target, status=target_status)
self._save() self._save()
visit_index = index if index is not None and index < len(self.data.visits) else None visit_index = (
index if index is not None and index < len(self.data.visits) else None
)
return visit_index, flushed return visit_index, flushed
+653
View File
@@ -0,0 +1,653 @@
"""Editor REST API and WebSocket sessions.
The management endpoints behind the SSO forward-auth gate: the site tree
(``/_api/pages``), structure operations (``/_api/structure``), site-wide
settings (``/_api/settings``), task-list toggles (``/_api/toggle-task``),
the translations refresh (``/_api/translations``), the editor session
socket (``/_api/ws/editor``), and the translator service channel
(``/_translate/{clientkey}`` — deliberately NOT under ``/_api``: the
server-generated key in the path is the access control).
"""
import logging
from datetime import UTC, datetime
from fastapi import (
APIRouter,
HTTPException,
Request,
WebSocket,
WebSocketDisconnect,
)
from pydantic import BaseModel
from pagerite import i18n, views
from pagerite.chunks import store_chunks
from pagerite.data import (
Node,
append_order,
find_slot,
node_markdown,
resolve,
sorted_nodes,
)
from pagerite.markdown import render, toggle_task
from pagerite.state import (
_check_reserved,
_ensure,
_invalidate_pages,
_remove_page,
data,
dispatcher,
kanta,
)
logger = logging.getLogger(__name__)
router = APIRouter()
class PageIn(BaseModel):
"""Payload for creating or replacing a page."""
title: str
markdown: str
published: bool = True
banner: str | None = None # None keeps the existing banner
@router.get("/_api/pages")
async def list_pages(lang: str | None = None) -> list[dict]:
"""The site tree for the structure editor (all nodes, drafts included).
Nested by slug; each node carries its full path, menu order, flags and
language settings (``language`` is the node's own primary-language
setting, "" = inherit; ``primary`` is the resolved effective one).
With a ``?lang=`` translation, titles come out in that language where a
translation exists (``translated`` flags it — true trivially for rows
whose primary language IS the selected one; other rows fall back to
the original title, dimmed) — the structure itself (slugs, order,
hierarchy) is language-independent.
"""
tag = i18n.base_tag(lang or "")
titles = i18n.title_map(data, tag) if tag else {}
def dump(nodes: dict[str, Node], prefix: str, inherited: str) -> list[dict]:
out = []
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
primary = node.language or inherited
out.append(
{
"slug": slug,
"path": path,
"title": titles.get(path) or node.title,
"translated": path in titles or (bool(tag) and primary == tag),
"order": node.order,
"published": node.published,
"has_content": node.chunks is not None,
"language": node.language,
"primary": primary,
"children": dump(node.children, path, primary),
}
)
return out
return dump(data.menu, "", i18n.ORIGINAL_LANGUAGE)
@router.put("/_api/pages/{path:path}", status_code=204)
async def save_page(
path: str, page: PageIn, request: Request, lang: str | None = None
) -> None:
"""Create or replace the page at a slug path ("" or "/" = front page).
Missing ancestors are created as content-less category labels. Giving
a category markdown turns it into a landing page. Empty markdown (after
stripping) creates an empty page that renders with just its title —
saving never deletes; use DELETE to remove a page (the page editor
issues DELETE when you save empty text).
With a ``?lang=`` query (a translation, not the primary language) the
save is a translated-view edit (docs/localization.md): the markdown is
diffed against the currently served hybrid and the minimal diff is
appended as a Patch under ``patches[f"{path}:{lang}"]`` — node.chunks
and the original-language fields (title, published, banner) stay
untouched.
"""
path = path.strip("/")
_check_reserved(path)
lang = i18n.base_tag(lang or "")
if lang and lang != i18n.primary_lang(data.menu, path):
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None:
raise HTTPException(404, "no such page")
with kanta.transaction(
f"page:{lang}", user=request.headers.get("remote-user"), extra=path
):
# Patches alone make the translated version exist.
if i18n.add_patch(data, node, path, lang, page.markdown):
_invalidate_pages()
return
with kanta.transaction(
"page", user=request.headers.get("remote-user"), extra=path
):
node = _ensure(data.menu, path)
node.title = page.title
node.chunks = store_chunks(data.chunks, page.markdown)
node.published = page.published
if page.banner is not None:
node.banner = page.banner
node.modified = datetime.now(UTC)
_invalidate_pages()
@router.delete("/_api/pages/{path:path}", status_code=204)
async def delete_page(path: str, request: Request) -> None:
"""Delete a node by slug path.
A category (node with children) loses only its landing page and stays
as a content-less label; a childless node is removed entirely.
"""
path = path.strip("/")
_check_reserved(path)
with kanta.transaction(
"page:delete", user=request.headers.get("remote-user"), extra=path
):
if not _remove_page(data.menu, path):
raise HTTPException(404, "no such page")
_invalidate_pages()
class StructureOp(BaseModel):
"""Rearrange the site tree: reorder, move/rename or retitle a node.
`order` is a fresh fractional key computed client-side from the node's
new siblings (a value halfway between them); all other items keep
theirs. `move_to` is the full target path — the parent must exist and
the new slug be free. Moves carry the whole subtree. The front page is
just the top-level node with slug "": renaming it away leaves no front
page ("/" then redirects to the first nav item), and any childless
top-level node can take the empty slug to become the front page.
With `lang` (a translation, not the node's primary language) a `title`
edit writes a per-language title fragment instead of the original — the
same storage as machine title translations (docs/localization.md);
sending the original's text removes the override. Structural fields are
not combinable with a translated title edit.
`language` sets the node's primary language (a BCP-47 base tag; "" =
inherit from the nearest ancestor, the front page last, site default
"en" final — Node.language), inherited by the whole subtree.
"""
path: str
order: float | None = None
move_to: str | None = None
title: str | None = None
lang: str | None = None
language: str | None = None
@router.post("/_api/structure", status_code=204)
async def update_structure(op: StructureOp, request: Request) -> None:
"""Apply one structure operation (see StructureOp)."""
path = op.path.strip("/")
chain = resolve(data.menu, path)
if chain is None:
raise HTTPException(404, "no such page")
node = chain[-1]
lang = i18n.base_tag(op.lang or "")
if op.language is not None:
# Primary-language setting (inherited by the subtree): reselects
# what "the original" means for the node — its language is part of
# every render, so a change invalidates everywhere.
language = i18n.base_tag(op.language)
with kanta.transaction(
"page:language", user=request.headers.get("remote-user"), extra=path
):
if language != node.language:
node.language = language
_invalidate_pages()
return
if op.title is not None and lang and lang != i18n.primary_lang(data.menu, path):
# Translated title (i18n.set_title_translation): original title,
# slugs and hierarchy stay untouched.
with kanta.transaction(
f"page:{lang}:title", user=request.headers.get("remote-user"), extra=path
):
if i18n.set_title_translation(data, node, lang, op.title):
_invalidate_pages()
return
target = op.move_to.strip("/") if op.move_to is not None else None
if target is not None and target != path:
_check_reserved(target)
if path and target.startswith(f"{path}/"):
raise HTTPException(400, "cannot move a page under itself")
slot = find_slot(data.menu, target)
if slot is None:
raise HTTPException(404, "target parent does not exist")
tnodes, tslug = slot
if tslug in tnodes:
raise HTTPException(400, "target path exists")
if not tslug and node.children:
raise HTTPException(400, "the front page cannot have children")
# One structure call can combine a title set, a move/rename and a
# reorder; the action names the most significant of them.
action = (
"page:slug"
if target is not None and target != path
else "page:title"
if op.title is not None
else "structure:reorder"
)
with kanta.transaction(
action, user=request.headers.get("remote-user"), extra=path
):
if op.title is not None:
node.title = op.title
if target is not None and target != path:
snodes, sslug = find_slot(data.menu, path)
del snodes[sslug]
# A pure rename (same parent) keeps its position; only a move
# to another level appends at the end (unless an order came
# with the drop).
same_level = path.rpartition("/")[0] == target.rpartition("/")[0]
node.order = (
op.order
if op.order is not None
else node.order
if same_level
else append_order(tnodes)
)
tnodes[tslug] = node
elif op.order is not None:
node.order = op.order
node.modified = datetime.now(UTC)
_invalidate_pages()
@router.get("/_api/settings")
async def get_settings() -> dict:
"""Site-wide settings (brand, theme, custom CSS and favicon URL), plus
the themes, banner designs and user fonts available on disk for the
selectors, the translator service keys and the wanted translation
languages (for the /_translate socket)."""
return {
"brand": data.brand,
"brand_html": data.brand_html,
"theme": data.theme,
"custom_css": data.custom_css,
"favicon": f"/_f/{data.favicon}" if data.favicon else "",
"themes": views._theme_info(),
"banner_designs": views._banner_design_names(),
"fonts": views._user_fonts(),
"transition": data.transition,
"transitions": views._transition_names(),
"translate_keys": data.translate_keys,
# The site default primary language: the front page's resolved
# setting (every page may override it, inherited down the tree).
"primary_lang": i18n.primary_lang(data.menu, ""),
"translate_langs": sorted(data.translate_langs),
}
class SettingsIn(BaseModel):
"""Payload for updating site-wide settings."""
brand: str
theme: str
custom_css: str
brand_html: str = ""
transition: str = "cube"
translate_langs: list[str] | None = None # None keeps the current set
@router.put("/_api/settings", status_code=204)
async def put_settings(settings: SettingsIn, request: Request) -> None:
"""Update site-wide settings; invalidates cached pages and ETags."""
with kanta.transaction(
"settings", user=request.headers.get("remote-user")
):
data.brand = settings.brand
data.brand_html = settings.brand_html
data.theme = settings.theme
data.custom_css = settings.custom_css
data.transition = settings.transition
if settings.translate_langs is not None:
# Any language may be a target — including the site default
# (an article in another language can be translated INTO it);
# a node's own primary is excluded per article, not here.
data.translate_langs = {
tag: True
for lang in settings.translate_langs
if (tag := i18n.base_tag(lang))
}
_invalidate_pages()
@router.delete("/_api/translations", status_code=204)
async def delete_translations(request: Request) -> None:
"""Drop all machine translations (Data.trans) so the dispatcher
re-translates everything from scratch (a translate:reset action:
the invalidation hook re-offers every fragment to connected
translators). User patches are kept; the availability index
(node.langs) is rebuilt from them — patches alone still make a language
exist on a page."""
with kanta.transaction(
"translate:reset", user=request.headers.get("remote-user")
):
i18n.clear_translations(data)
_invalidate_pages()
# Fragments rejected this run (segment validation) stay skipped no
# longer: a refresh is precisely the "another chance" for them.
dispatcher.reset_validation_failures()
class ToggleTaskIn(BaseModel):
"""Payload for toggling one task-list checkbox."""
path: str
index: int
markdown: str | None = None
@router.post("/_api/toggle-task")
async def toggle_task_endpoint(body: ToggleTaskIn, request: Request) -> dict[str, str]:
"""Toggle the Nth task-list checkbox in a page's Markdown source.
If ``markdown`` is provided the source is left untouched and the toggled
Markdown is returned (used while the page editor is open, so the live
CodeMirror document can be updated). Otherwise the stored page at
``path`` is read, toggled, and saved.
"""
path = body.path.strip("/")
_check_reserved(path)
if body.markdown is not None:
new_markdown = toggle_task(body.markdown, body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
return {"markdown": new_markdown}
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None:
raise HTTPException(404, "no such page")
new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
with kanta.transaction(
"page", user=request.headers.get("remote-user"), extra=path
):
# Re-chunk like any save: only the chunk containing the toggled
# checkbox gets a new hash, the rest keep theirs.
node.chunks = store_chunks(data.chunks, new_markdown)
node.modified = datetime.now(UTC)
_invalidate_pages()
return {"markdown": new_markdown}
# WebSocket API for external translation services (not under /_api: it is keyed
# with Data.translate_keys instead of the SSO forward-auth). The dispatcher —
# protocol, connected clients and the job pipeline — lives in translate.py.
@router.websocket("/_translate/{clientkey}")
async def translate_ws(ws: WebSocket, clientkey: str) -> None:
"""Translator service channel (docs/localization.md).
Deliberately NOT under /_api/: the external forward-auth is skipped;
the server-generated client key in the path is the access control
(``Data.translate_keys``: key -> display name; the first is generated
at bootstrap, all are shown in the admin's /_api/settings).
"""
await dispatcher.handle_ws(ws, clientkey)
@router.websocket("/_api/ws/editor")
async def editor_ws(ws: WebSocket) -> None:
"""Editor session: open pages, render previews, save — over one socket.
Stateless protocol (each message carries the path):
<- {"type": "open", "path", "lang"?}
-> {"type": "doc", "path", "exists", "title", "markdown", "published",
"banner", "banner_design", "lang", "primary_lang", "langs",
"translate_langs"}
<- {"type": "render", "path", "markdown"}
-> {"type": "html", "path", "html"}
<- {"type": "save", "path", "title"?, "markdown"?, "published"?,
"banner"?, "banner_design"?, "move_from"?, "lang"?, "base"?}
(absent fields keep their old values; move_from: rename/move a
page, subtree included)
-> {"type": "saved", "path"} | {"type": "error", "detail"}
With "lang" (a translation, not the primary language), open returns the
effective hybrid Markdown and title for that language plus the language
metadata the picker's UI needs; save diffs the submitted Markdown
against "base" (the editor's shadow copy of the hybrid it started from
— absent: the current hybrid) and stores it as a user Patch, and a
changed title becomes a fragment in Data.trans — node.chunks and the
other fields stay untouched (docs/localization.md).
"""
await ws.accept()
try:
while True:
msg = await ws.receive_json()
path = msg.get("path", "").strip("/")
try:
_check_reserved(path)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
match msg.get("type"):
case "open":
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
# The article's primary language: its own setting,
# inherited down the tree ("en" final fallback).
node_lang = i18n.primary_lang(data.menu, path)
lang = i18n.base_tag(str(msg.get("lang") or ""))
if lang == node_lang:
lang = ""
markdown = ""
title = node.title if node else ""
if node is not None:
markdown = node_markdown(data, node) or ""
if lang and node.chunks is not None:
# Translation view: the effective (hybrid)
# Markdown and title for that language —
# machine fragments + user patches over the
# original (docs/localization.md editor flow).
markdown = i18n.hybrid_markdown(data, node, path, lang)
title = i18n.title_map(data, lang).get(path) or title
await ws.send_json(
{
"type": "doc",
"path": path,
"exists": node is not None,
"title": title,
"markdown": markdown,
"published": node.published if node else True,
"banner": node.banner if node else "",
# Own banner design setting: null = inherit,
# "" = none, otherwise a design name.
"banner_design": node.banner_design if node else None,
# Which node's banner applies here ("" = front page,
# null = default artwork); the site editor shows it
# as the banner field's placeholder.
"banner_from": views.banner_source(data.menu, path),
# Which node's banner-design setting would apply on
# inherit ("" = front page, null = the active
# theme's default) and what design that resolves to.
"banner_design_from": (
src := views.banner_design_source(
data.menu, path, data.theme
)
),
"banner_design_inherited": (
views.banner_design(data.menu, src, data.theme)
if src is not None
else views.theme_banner_design(data.theme)
),
# Language context for the editor's picker: the
# language this Markdown represents ("" = primary),
# the page's own primary language, the translations
# this page already has, and the site-wide
# configured target languages.
"lang": lang,
"primary_lang": node_lang,
"langs": sorted(node.langs) if node else [],
"translate_langs": sorted(data.translate_langs),
}
)
case "render":
markdown = msg.get("markdown", "")
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
rendered = render(
markdown,
path,
node.created if node else None,
node.modified if node else None,
# The title is injected as h1 when the markdown has
# none; the editor's title field edits live-preview.
title=msg.get("title") or (node.title if node else ""),
)
await ws.send_json(
{
"type": "html",
"path": path,
"html": rendered.html,
# Column-layout flag: the preview toggles the
# article's .multicol class and swaps in the
# segmented (.colseg/.cols) article html.
"multicol": rendered.multicol,
}
)
case "save":
move_from = (msg.get("move_from") or path).strip("/")
lang = i18n.base_tag(str(msg.get("lang") or ""))
translated = bool(
lang and lang != i18n.primary_lang(data.menu, move_from)
)
try:
_check_reserved(move_from)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
old_chain = resolve(data.menu, move_from)
old = old_chain[-1] if old_chain else None
if old is None and move_from != path:
move_from = path # nothing to carry over; plain save
if move_from != path:
# Rename/move: detach the node (subtree included)
# and attach it at the new path. The target slug
# must be free and the front page childless.
if move_from and path.startswith(f"{move_from}/"):
await ws.send_json(
{
"type": "error",
"detail": "cannot move a page under itself",
}
)
continue
tslug = path.rpartition("/")[2]
if not tslug and old.children:
await ws.send_json(
{
"type": "error",
"detail": "the front page cannot have children",
}
)
continue
tchain = resolve(data.menu, path)
if tchain is not None:
await ws.send_json(
{
"type": "error",
"detail": "target path exists",
}
)
continue
if translated and (
move_from != path or old is None or old.chunks is None
):
# A translated-view save patches an existing
# original; it cannot create or move pages.
await ws.send_json({"type": "error", "detail": "no such page"})
continue
if translated and "markdown" in msg and not msg["markdown"].strip():
# Saving never deletes; an emptied translation would
# render as a blank page in that language.
await ws.send_json(
{
"type": "error",
"detail": "a translation cannot be emptied",
}
)
continue
with kanta.transaction(
f"page:{lang}" if translated else "page",
user=ws.headers.get("remote-user"),
extra=path,
):
if move_from != path:
same_menu = (
move_from.rpartition("/")[0] == path.rpartition("/")[0]
)
snodes, sslug = find_slot(data.menu, move_from)
node = snodes.pop(sslug)
parent = path.rpartition("/")[0]
if parent:
_ensure(data.menu, parent)
tnodes, tslug = find_slot(data.menu, path)
node.order = (
node.order if same_menu else append_order(tnodes)
)
tnodes[tslug] = node
else:
node = old if old is not None else _ensure(data.menu, path)
if translated:
# node.chunks and the original-language fields
# stay untouched: the markdown diff (against the
# editor's shadow "base" — the hybrid it started
# from; absent: the current hybrid) is appended
# as a Patch, a changed title becomes a
# per-language title override (i18n).
changed = False
if "markdown" in msg:
base = msg.get("base")
changed = i18n.add_patch(
data,
node,
path,
lang,
msg["markdown"],
base=base if isinstance(base, str) else None,
)
if "title" in msg and node.title:
changed = (
i18n.set_title_translation(
data, node, lang, msg["title"]
)
or changed
)
if changed:
_invalidate_pages()
else:
if "markdown" in msg:
# Saving never deletes; empty markdown is an
# empty page. Deletion is an explicit choice
# by the page editor (REST DELETE).
node.chunks = store_chunks(data.chunks, msg["markdown"])
if "title" in msg:
node.title = msg["title"]
if "published" in msg:
node.published = bool(msg["published"])
if "banner" in msg:
node.banner = msg["banner"]
if "banner_design" in msg:
node.banner_design = msg["banner_design"]
node.modified = datetime.now(UTC)
_invalidate_pages()
await ws.send_json({"type": "saved", "path": path})
except WebSocketDisconnect:
pass
+53 -1348
View File
File diff suppressed because it is too large Load Diff
+164
View File
@@ -0,0 +1,164 @@
"""Block-level Markdown chunking for content-addressed storage.
A page's Markdown is split into deterministic block-level chunks, each
stored once under its content hash in ``Data.chunks`` (docs/migrate.md).
Shared by the render/save pipeline (app.py, views.py, i18n.py) and the
schema migration (migrations.py), so a chunk's key is stable no matter
where the split happens.
"""
import re
import blake3
from pagerite.segments import has_prose
#: Fenced code block opener/closer: up to 3 spaces indent, then 3+
#: backticks or tildes (CommonMark).
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
#: HTML block openers that may span blank lines (CommonMark types 1-5:
#: script/pre/style/textarea, comments, processing instructions,
#: declarations, CDATA) with their closing condition. Other HTML blocks
#: end at the first blank line, which the generic blank-line split
#: already does.
_HTML_ATOMIC = (
(
re.compile(r"^ {0,3}<(?:script|pre|style|textarea)(?:\s|>|$)", re.I),
re.compile(r"</(?:script|pre|style|textarea)\s*>", re.I),
),
(re.compile(r"^ {0,3}<!--"), re.compile(r"-->")),
(re.compile(r"^ {0,3}<\?"), re.compile(r"\?>")),
(re.compile(r"^ {0,3}<!\[CDATA\["), re.compile(r"\]\]>")),
(re.compile(r"^ {0,3}<![A-Za-z]"), re.compile(r">")),
)
#: First line of a generic HTML block (a block-level tag).
_HTML_TAG = re.compile(r"^ {0,3}</?[A-Za-z][^>]*>")
def _fence_close(line: str, opener: str) -> bool:
"""True when ``line`` closes a code fence opened by ``opener``: the
same marker char, at least as many, and nothing else on the line."""
stripped = line.strip()
return (
len(stripped) >= len(opener)
and stripped[0] == opener[0]
and set(stripped) == {opener[0]}
)
def chunk_markdown(markdown: str) -> list[str]:
"""Split Markdown into block-level chunks, deterministically.
Blocks are separated by blank lines; fenced code blocks and the
multi-line HTML blocks (comments, script/pre/style, CDATA...) are
kept atomic, even across blank lines, and end at their closing
condition. Chunks carry no surrounding blank lines and no trailing
newline; rejoining with ``join_chunks`` reproduces the source modulo
blank-line normalization.
"""
chunks: list[str] = []
buf: list[str] = []
fence = "" # opener marker of the code fence we are in ("" = outside)
html_end: re.Pattern | None = None # closes the atomic HTML block we are in
def flush() -> None:
text = "\n".join(buf).strip("\n")
if text.strip():
chunks.append(text)
buf.clear()
for line in markdown.split("\n"):
if fence:
buf.append(line)
if _fence_close(line, fence):
fence = ""
flush()
continue
if html_end is not None:
buf.append(line)
if html_end.search(line):
html_end = None
flush()
continue
if not line.strip():
flush()
continue
if m := _FENCE_OPEN.match(line):
# Fences interrupt paragraphs (CommonMark): start a new block.
flush()
fence = m.group(1)
buf.append(line)
continue
if not buf:
for open_re, close_re in _HTML_ATOMIC:
if open_re.match(line):
buf.append(line)
if close_re.search(line): # opens and closes on one line
flush()
else:
html_end = close_re
break
else:
buf.append(line)
continue
buf.append(line)
flush() # an unterminated fence/HTML block runs to EOF, kept as code/HTML
return chunks
def _normalize(text: str) -> str:
"""Whitespace-insensitive chunk identity: strip trailing whitespace
per line and collapse surrounding blank lines, so whitespace-only
source edits don't invalidate translations."""
return "\n".join(line.rstrip() for line in text.split("\n")).strip("\n")
def chunk_key(text: str) -> bytes:
"""Content key of a chunk: the first 9 bytes of the blake3 digest of
the normalized text (72 bits — a site's chunk count stays far below
the birthday bound), using the same hasher as app.py's file store.
Keys are bytes: kanta/msgspec base64-encode them at the JSON
persistence level, so the raw database dicts carry 12-char strings.
"""
return blake3.blake3(_normalize(text).encode()).digest(9)
def needs_translation(chunk: str) -> bool:
"""False for chunks without prose: pure code fences, HTML blocks, and
anything that yields no translatable segments (pagerite/segments.py) —
container fences, lone {placeholders}, reference definitions.
These are inherently no-translate (docs/migrate.md): derived from the
chunk text itself, nothing is stored. Every language renders them from
the original chunk via the hybrid fallback.
"""
if _FENCE_OPEN.match(chunk):
return False
first = chunk.split("\n", 1)[0]
if any(open_re.match(first) for open_re, _ in _HTML_ATOMIC):
return False
if _HTML_TAG.match(first):
return False
return has_prose(chunk)
def join_chunks(chunks: list[str]) -> str:
"""The stored page form of chunks: blocks joined by a blank line,
with a trailing newline ("" for no chunks)."""
return "\n\n".join(chunks) + "\n" if chunks else ""
def store_chunks(store: dict[bytes, str], markdown: str) -> list[bytes]:
"""Chunk ``markdown`` into ``store`` (hash -> text); return the ordered
hashes. Unchanged chunks keep their hashes, so only genuinely new text
lands in the kanta change diff. First writer wins: variants sharing a
key differ only in insignificant whitespace (see chunk_key)."""
hashes = []
for chunk in chunk_markdown(markdown):
key = chunk_key(chunk)
store.setdefault(key, chunk)
hashes.append(key)
return hashes
+65 -31
View File
@@ -2,8 +2,9 @@
The site structure is a tree of Nodes. Every node is a menu label with a The site structure is a tree of Nodes. Every node is a menu label with a
configurable title and slug (its key in the parent's ``children``); the configurable title and slug (its key in the parent's ``children``); the
URL path is the chain of slugs from the top level. ``content`` is the URL path is the chain of slugs from the top level. ``chunks`` is the
node's Markdown page, or None for a pure category label, whose URL renders node's Markdown page as ordered content-hash keys into ``Data.chunks``
(docs/migrate.md), or None for a pure category label, whose URL renders
a placeholder page while nav links point at its first child. a placeholder page while nav links point at its first child.
""" """
@@ -11,6 +12,16 @@ from datetime import UTC, datetime
import msgspec import msgspec
from pagerite.chunks import join_chunks
class Patch(msgspec.Struct, omit_defaults=True):
"""One editing session's overrides on a translated view, applied
independently per hunk (docs/localization.md)."""
#: (search, replace) pairs on the served hybrid Markdown.
hunks: list[tuple[str, str]] = []
class Node(msgspec.Struct, omit_defaults=True): class Node(msgspec.Struct, omit_defaults=True):
"""One item of the site hierarchy. """One item of the site hierarchy.
@@ -28,9 +39,21 @@ class Node(msgspec.Struct, omit_defaults=True):
title: str = "" title: str = ""
order: float = 0 order: float = 0
#: Markdown source of the node's page; None = pure category label #: Ordered chunk hashes (9-byte keys into ``Data.chunks``); None =
#: (its URL renders a placeholder page). #: pure category label (its URL renders a placeholder page), a list
content: str | None = None #: (possibly empty) = a page.
chunks: list[bytes] | None = None
#: Primary language of the article (BCP-47 base tag). "" = inherit
#: (nearest ancestor, front page last, site default "en" final).
language: str = ""
#: Chunk hashes the editor marked "do not translate" (always served
#: from the original). Presence-keys, value always True.
no_trans: dict[bytes, bool] = {}
#: Languages this article is available in (besides its primary
#: language). Presence-keys, value always True — the availability
#: index for rendering and language selection; maintained by whoever
#: writes translation data (docs/migrate.md).
langs: dict[str, bool] = {}
#: Raw HTML for the header banner (img, styled div, canvas+script...), #: Raw HTML for the header banner (img, styled div, canvas+script...),
#: rendered after the banner design's artwork so author code always #: rendered after the banner design's artwork so author code always
#: wins over the design's own styles. #: wins over the design's own styles.
@@ -50,34 +73,11 @@ class Node(msgspec.Struct, omit_defaults=True):
) )
class Page(msgspec.Struct, omit_defaults=True):
"""Legacy flat page record, from before the tree model.
Kept only so old databases still decode; app.py migrates any entries
into ``Data.menu`` on startup and clears this.
"""
title: str
markdown: str
published: bool = True
order: float = 0
banner: str = ""
created: datetime = msgspec.field(
default_factory=lambda: datetime.now(UTC),
)
modified: datetime = msgspec.field(
default_factory=lambda: datetime.now(UTC),
)
class Data(msgspec.Struct): class Data(msgspec.Struct):
"""Root object of the kanta database. Owned and edited in place by us.""" """Root object of the kanta database. Owned and edited in place by us."""
#: Top-level menu items by slug; "" is the front page. #: Top-level menu items by slug; "" is the front page.
menu: dict[str, Node] = {} menu: dict[str, Node] = {}
#: Bumped on every structure/content write, so page ETags (which embed
#: it) invalidate cached copies when navigation-affecting changes happen.
version: int = 0
#: Site name shown in the header and <title> suffix; editable in the #: Site name shown in the header and <title> suffix; editable in the
#: site editor. Empty = no brand link in the header, no title suffix. #: site editor. Empty = no brand link in the header, no title suffix.
brand: str = "Pagerite" brand: str = "Pagerite"
@@ -101,9 +101,43 @@ class Data(msgspec.Struct):
#: linked as <link rel="icon"> on every page. Empty = the build's #: linked as <link rel="icon"> on every page. Empty = the build's
#: /favicon.ico. #: /favicon.ico.
favicon: str = "" favicon: str = ""
#: Legacy flat page store (pre-tree databases); migrated into `menu` #: API keys gating the translator service WebSocket (/_translate/{key};
#: on startup, then cleared. Never written otherwise. #: the external forward-auth does not cover that route): key -> display
pages: dict[str, Page] = {} #: name. Keys are 12 lowercase alphanumeric characters; the first is
#: generated at database bootstrap (see app.py), multiple keys are a
#: future reservation (e.g. managed via a web interface).
translate_keys: dict[str, str] = {}
#: Wanted target languages for the translator service (presence-keys,
#: value always True). The dispatcher offers jobs only in the
#: intersection of these and a connection's announced capabilities.
#: Bootstrapped to es+zh; edited in the editor shell's localization
#: tab (or via /_api/settings).
translate_langs: dict[str, bool] = {}
#: All original-language page text, content-addressed:
#: chunk_key (9 bytes; base64 at the JSON level) -> Markdown chunk.
#: Shared by every article.
chunks: dict[bytes, str] = {}
#: Machine translations: chunk hash -> lang -> translated Markdown
#: (a nested dict rather than tuple keys, which msgspec's JSON
#: serializer does not support). Also used for node titles (hash of
#: the title text).
trans: dict[bytes, dict[str, str]] = {}
#: User override patches per article and language:
#: f"{path}:{lang}" -> ordered patches (paths without leading slash).
patches: dict[str, list[Patch]] = {}
def node_markdown(data: Data, node: Node) -> str | None:
"""The node's original Markdown assembled from the chunk store.
None for category labels (chunks is None); an empty page gives "".
Hashes missing from the store (shouldn't happen) are skipped.
"""
if node.chunks is None:
return None
return join_chunks(
[t for h in node.chunks if (t := data.chunks.get(h)) is not None]
)
def prettify(slug: str) -> str: def prettify(slug: str) -> str:
+375
View File
@@ -0,0 +1,375 @@
"""Content-addressed file store, image derivatives, and file routes.
``FileStore`` keeps uploads, seed assets and fetched favicons on disk under
hash-prefixed names, fully cached in RAM (uncompressed plus a zstd copy
when compression shrinks the body), served immutable at ``/_f/``. Raster
images and SVGs are recompressed into AVIF/WebP/JPEG derivatives
(``store_image`` and helpers); the untouched original is kept alongside as
``<hash>.orig<ext>`` (never served). Routes: upload/delete under
``/_api/files``, the favicon settings endpoints, the ``/_f/`` server with
Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``).
"""
import asyncio
import logging
import mimetypes
import tempfile
from contextlib import suppress
from pathlib import Path
import blake3
from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import Response
from mediapreview import dispatch
from pagerite import views
from pagerite.state import (
FAVICON_MAXSIZE,
FILES_DIR,
IMAGE_JPG_QUALITY,
IMAGE_MAXSIZE,
IMAGE_QUALITY,
IMAGE_WEBP_QUALITY,
_invalidate_pages,
_zstd,
data,
kanta,
)
logger = logging.getLogger(__name__)
# mediapreview logs pyvips noise ("VipsForeignSaveJpegTarget argument strip is
# deprecated", "threadpool completed with N workers") at INFO; keep warnings.
logging.getLogger("mediapreview").setLevel(logging.WARNING)
router = APIRouter()
class FileStore:
"""Content-addressed files on disk, fully cached in RAM.
Every file is kept in RAM uncompressed and zstd-compressed (the
compressed copy only when it actually shrinks the body), so ``/_f``
serves both encodings without touching disk or re-compressing.
"""
def __init__(self, path: Path) -> None:
self.path = path
#: name -> (uncompressed body, zstd body or None)
self._cache: dict[str, tuple[bytes, bytes | None]] = {}
@staticmethod
def _entry(body: bytes) -> tuple[bytes, bytes | None]:
compressed = _zstd.compress(body)
return body, compressed if len(compressed) < len(body) else None
def load(self) -> None:
"""Read every stored file into the RAM cache (startup)."""
try:
entries = sorted(self.path.iterdir())
except FileNotFoundError:
return
for f in entries:
if f.is_file() and not f.name.startswith("."):
self._cache.setdefault(f.name, self._entry(f.read_bytes()))
def get(self, name: str) -> tuple[bytes, bytes | None] | None:
return self._cache.get(name)
def put(self, name: str, body: bytes) -> None:
"""Store ``body`` under ``name`` on disk and in the RAM cache."""
if name in self._cache:
return
self.path.mkdir(parents=True, exist_ok=True)
(self.path / name).write_bytes(body)
self._cache[name] = self._entry(body)
def delete(self, name: str) -> None:
"""Delete a file plus its derivatives/original counterparts, if any.
An image upload is stored as a group sharing the hash prefix
(``<hash>.orig.<ext>`` + ``<hash>.avif/.webp/.jpg``); deleting any
of the names removes them all.
"""
stem = name.partition(".")[0]
for key in [k for k in self._cache if k.partition(".")[0] == stem]:
self._cache.pop(key, None)
with suppress(FileNotFoundError):
(self.path / key).unlink()
def __contains__(self, name: str) -> bool:
return name in self._cache
file_store = FileStore(FILES_DIR)
def _ext(orig: str) -> str:
"""Sanitized lowercase extension (with dot) of an original file name."""
return "".join(c for c in Path(orig).suffix.lower() if c.isalnum() or c == ".")
def _hash_name(body: bytes, orig: str) -> str:
"""Content-addressed file name: blake3 hash prefix + original extension."""
return blake3.blake3(body).hexdigest()[:12] + _ext(orig)
def _to_avif(body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE) -> bytes | None:
"""Recompress an image body to a thumbnailed AVIF via mediapreview's
dispatch (pyvips for common formats, ffmpeg for HEIC/HEIF/AVIF), or
None if the body is not a decodable image (stored as-is by the caller).
Dispatch needs a real file for format routing, so the body goes
through a temp file.
"""
with tempfile.NamedTemporaryFile(suffix=ext) as tmp:
tmp.write(body)
tmp.flush()
try:
avif, _resp = dispatch(
Path(tmp.name),
quality=IMAGE_QUALITY,
maxsize=maxsize,
maxzoom=1,
)
except Exception:
return None
return avif
def _svg_to_png(body: bytes, maxsize: int) -> bytes | None:
"""Rasterize an SVG to PNG via pyvips, scaled so the long side is
``maxsize`` SVGs often carry no meaningful intrinsic resolution, so
we rasterize at full image size rather than the tiny nominal one."""
import pyvips
try:
img = pyvips.Image.new_from_buffer(body, "")
scale = (
maxsize / max(img.width, img.height)
if img.width and img.height
else maxsize
)
if scale != 1:
img = pyvips.Image.new_from_buffer(body, "", scale=scale)
return img.write_to_buffer(".png")
except pyvips.Error:
return None
def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes:
"""Re-encode the AVIF derivative into a fallback format (WebP/JPEG)
via pyvips. JPEG has no alpha, so it is flattened onto white;
``strip`` keeps metadata (EXIF) out of the fallbacks."""
import pyvips
img = pyvips.Image.new_from_buffer(avif, "")
if suffix == ".jpg" and img.hasalpha():
img = img.flatten(background=[255, 255, 255])
return img.write_to_buffer(suffix, Q=quality, strip=True)
def _image_derivatives(
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE
) -> dict[str, bytes] | None:
"""The served variants of an uploaded image: ``avif`` (primary,
thumbnailed to ``maxsize``) plus ``webp`` and ``jpg`` fallbacks
re-encoded from it. SVGs are rasterized first (they are vector, so
the raster replaces nothing the .svg itself stays servable).
Returns None for non-decodable content (stored as-is by the caller).
"""
if ext == ".svg":
png = _svg_to_png(body, maxsize)
if png is None:
return None
body, ext = png, ".png"
avif = _to_avif(body, ext, maxsize)
if avif is None:
return None
return {
"avif": avif,
"webp": _avif_to_format(avif, ".webp", IMAGE_WEBP_QUALITY),
"jpg": _avif_to_format(avif, ".jpg", IMAGE_JPG_QUALITY),
}
def store_image(
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE, *, derive: bool = True
) -> str:
"""Store an image body content-addressed and return its file name.
Decodable images get AVIF/WebP/JPEG derivatives thumbnailed to
``maxsize``; the original is kept as ``<hash>.orig<ext>`` (SVG
originals as ``<hash>.svg``, still servable) and the bare ``<hash>``
name is returned (the server negotiates the format by Accept header).
Anything else undecodable content, or ``derive=False`` (GIFs, whose
animation recompression would lose) is stored as-is and returned with
its extension. Blocking (pyvips/ffmpeg); call via ``asyncio.to_thread``
from async code.
"""
digest = blake3.blake3(body).hexdigest()[:12]
derivatives = _image_derivatives(body, ext, maxsize) if derive else None
if derivatives is None: # store the body as-is
file_store.put(digest + ext, body)
return digest + ext
file_store.put(f"{digest}.svg" if ext == ".svg" else f"{digest}.orig{ext}", body)
for fmt, variant in derivatives.items():
file_store.put(f"{digest}.{fmt}", variant)
return digest
@router.put("/_api/files/{name}")
async def upload_file(name: str, request: Request) -> dict[str, str]:
"""Store an upload (image, video...) in the content-addressed store.
The stored name is a blake3 hash prefix + the original extension,
served immutable at "/_f/{name}"; returns {"path": "/_f/..."}.
Raster images and SVGs are recompressed (SVGs rasterized) into AVIF
(primary) plus WebP and JPEG fallbacks: the original goes to
``<hash>.orig<ext>`` (kept for reprocessing, never served it may
carry EXIF data; SVG originals stay servable as ``<hash>.svg`` since
vector carries no EXIF) and pages link the bare ``/_f/<hash>``, the
server picking the format from the request's Accept header. GIFs are
stored as-is (animation would be lost), as is other non-decodable
content.
"""
if "/" in name or name in {".", ".."}:
raise HTTPException(400, "bad file name")
body = await request.body()
if not body:
raise HTTPException(400, "empty file")
ext = _ext(name)
stored = await asyncio.to_thread(store_image, body, ext, derive=ext != ".gif")
return {"path": f"/_f/{stored}"}
@router.delete("/_api/files/{name}", status_code=204)
async def delete_file(name: str) -> None:
"""Remove a file from the content-addressed store (no refcounting:
other pages referencing the same content will 404)."""
if name not in file_store:
raise HTTPException(404, "no such file")
file_store.delete(name)
@router.put("/_api/settings/favicon")
async def put_favicon(request: Request) -> dict[str, str]:
"""Upload a favicon into the content-addressed store and activate it.
Raw image body (ico/png/svg...). Decodable images are thumbnailed to
FAVICON_MAXSIZE (192px browsers scale down from there themselves)
and stored as AVIF/WebP/JPEG derivatives linked extension-less; SVG
originals also stay servable under their ``.svg`` name. Undecodable
bodies are stored as-is. Pages link it as <link rel="icon">. Returns
{"path": "/_f/..."}.
"""
body = await request.body()
if not body:
raise HTTPException(400, "empty file")
ext = _ext(request.headers.get("x-filename", "favicon.ico"))
stored = await asyncio.to_thread(store_image, body, ext, FAVICON_MAXSIZE)
with kanta.transaction("settings", user=request.headers.get("remote-user")):
data.favicon = stored
_invalidate_pages()
return {"path": f"/_f/{stored}"}
@router.delete("/_api/settings/favicon", status_code=204)
async def delete_favicon(request: Request) -> None:
"""Clear the custom favicon (back to the build's /favicon.ico).
The blob stays in the content-addressed store; only the reference goes.
"""
with kanta.transaction("settings", user=request.headers.get("remote-user")):
data.favicon = ""
_invalidate_pages()
async def _serve_user_file(path: Path | None, request: Request) -> Response:
"""Serve a user-asset file resolved on disk, with mtime etag.
Read from disk on every request (etag by mtime+size): user assets are
never built or content-hashed, so edits on disk show on the next page
load, in prod as well as dev.
"""
if path is None:
raise HTTPException(404)
stat = path.stat()
etag = f'"{stat.st_mtime_ns:x}-{stat.st_size:x}"'
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
return Response(
path.read_bytes(),
media_type=mime,
headers={"etag": etag, "cache-control": "no-cache"},
)
@router.get("/_themes/{name}/{filename}")
async def theme_file(name: str, filename: str, request: Request) -> Response:
"""Serve a theme/banner-design file, resolved across views.THEME_DIRS.
Stylesheets plus any extra assets the CSS references (like summer's
grass.svg).
"""
return await _serve_user_file(views.theme_file(name, filename), request)
@router.get("/_fonts/{name}/{filename}")
async def user_font_file(name: str, filename: str, request: Request) -> Response:
"""Serve a user font file, resolved across views.FONT_DIRS.
The folder's font.css (@font-face rules + --font-{name} stack variable)
is linked on every page; the woff2 files it references come from here.
"""
return await _serve_user_file(views.font_file(name, filename), request)
@router.get("/_f/{name}")
async def stored_file(name: str, request: Request) -> Response:
"""Serve a file from the content-addressed store (immutable: the name
is its own hash, so cache forever). Bodies are served from the RAM
cache, zstd-compressed when the client accepts it and compression
actually shrank the file.
A bare ``/_f/{hash}`` (no extension, how pages link uploaded images)
content-negotiates between the stored derivatives: a format is served
only when the Accept header lists it explicitly ``image/avif``
AVIF, ``image/webp`` WebP, anything else (including ``image/*`` and
``*/*``) JPEG. An explicit extension pins the format. ``.orig.``
originals are internal (they may carry EXIF data) and never served."""
if ".orig." in name:
raise HTTPException(404)
etag = name
vary = ""
entry = file_store.get(name)
if entry is None and "." not in name:
# Extension-less image link: negotiate avif/webp/jpg by Accept.
vary = "accept"
accept = request.headers.get("accept", "")
if "image/avif" in accept:
order = ("avif", "webp", "jpg")
elif "image/webp" in accept:
order = ("webp", "jpg", "avif")
else:
order = ("jpg", "webp", "avif")
for ext in order:
etag = f"{name}.{ext}"
entry = file_store.get(etag)
if entry is not None:
break
if entry is None:
raise HTTPException(404)
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
body, compressed = entry
headers = {"etag": etag, "cache-control": "public, max-age=31536000, immutable"}
if compressed is not None and "zstd" in request.headers.get("accept-encoding", ""):
headers["content-encoding"] = "zstd"
vary = f"{vary}, accept-encoding".lstrip(", ")
body = compressed
if vary:
headers["vary"] = vary
mime = mimetypes.guess_type(etag)[0] or "application/octet-stream"
return Response(body, media_type=mime, headers=headers)
+282
View File
@@ -0,0 +1,282 @@
"""Localization: language selection, translation storage and assembly.
See docs/localization.md and docs/migrate.md. Each article's primary
language is ``Node.language``, inherited down the hierarchy (front page =
site default, ORIGINAL_LANGUAGE as the final fallback). The database holds
the original language as content-addressed chunks (``Data.chunks``); per
target language there are machine-translated fragments (``Data.trans``)
and user override patches (``Data.patches``), assembled into the served
Markdown at render time, with per-node fallback to the original titles.
"""
from collections.abc import Callable
from difflib import SequenceMatcher
import msgspec
from pagerite.chunks import chunk_key, chunk_markdown, join_chunks
from pagerite.data import Data, Node, Patch, resolve
#: Final fallback for a page's primary language when neither it nor any
#: ancestor (up to the front page) sets one (Node.language, "" = inherit).
ORIGINAL_LANGUAGE = "en"
#: Languages written right-to-left; pages served in one get dir="rtl" on
#: <html> (views._layout).
RTL_LANGUAGES = frozenset({"ar", "fa", "he", "ur"})
def primary_lang(menu: dict[str, Node], path: str) -> str:
"""The primary language of the article at ``path``: its own
``language`` setting, else the nearest ancestor's (the front page
last it doubles as the site default), falling back to
ORIGINAL_LANGUAGE. Missing tail segments (a page being created)
resolve to the nearest existing ancestor."""
p = path.strip("/")
while True:
chain = resolve(menu, p)
if chain:
for node in reversed(chain):
if node.language:
return node.language
if not p:
return ORIGINAL_LANGUAGE
p = p.rpartition("/")[0]
class Translation(msgspec.Struct, omit_defaults=True):
"""Translated content for one page and language.
``markdown`` is the translated page source in the same format as the
original (None = keep the original Markdown); ``titles`` maps node paths
(top-level slug, then slash-joined) to translated navigation titles, so a
partially translated tree still renders with per-node English fallback.
"""
markdown: str | None = None
titles: dict[str, str] = {}
def base_tag(tag: str) -> str:
"""The lowercase base subtag of a language tag (fi-FI -> fi)."""
return tag.strip().lower().partition("-")[0]
def parse_accept_language(header: str) -> list[str]:
"""Accept-Language header as an ordered, deduped list of base subtags.
q-values are deliberately ignored: all known implementations send the
header in order of preference. Region tags normalize to their base
subtag (fi-FI -> fi); "*" and empties are dropped.
"""
langs = []
for part in header.split(","):
tag = base_tag(part.split(";", 1)[0])
if tag and tag != "*" and tag not in langs:
langs.append(tag)
return langs
def select_language(
query_lang: str | None,
accept_language: str | None,
is_available: Callable[[str], bool],
original: str = ORIGINAL_LANGUAGE,
) -> str:
"""The language to serve (see docs/localization.md).
1. ``?lang=`` wins when a translation exists for it (otherwise falls
through to the header logic).
2. The original language anywhere in the header list wins an AI
translation is strictly worse than the original for anyone who has
English configured at all.
3. Otherwise the first header language with an available translation.
4. Fall back to the original.
"""
if query_lang:
tag = base_tag(query_lang)
if tag == original or (tag and is_available(tag)):
return tag
langs = parse_accept_language(accept_language or "")
if original in langs:
return original
for lang in langs:
if lang != original and is_available(lang):
return lang
return original
def apply_patch(hybrid: str, patch: Patch) -> str:
"""Apply one patch to the hybrid Markdown, best effort, each hunk
independently: a hunk whose search text no longer exists is stale and
silently skipped (docs/localization.md)."""
for search, replace in patch.hunks:
if search and search in hybrid:
hybrid = hybrid.replace(search, replace, 1)
return hybrid
def make_patch(base: str, edited: str) -> Patch:
"""The minimal diff of ``edited`` against the served ``base`` hybrid as
(search, replace) hunks at block granularity (docs/localization.md).
Blocks are the chunk_markdown split, so hunks align with translation
units and code fences never straddle a hunk boundary. Pure inserts
anchor on the preceding block (an empty search would never match);
inserts at the very top anchor on the first block. autojunk is off:
the diff must be deterministic, and pages are small.
"""
a, b = chunk_markdown(base), chunk_markdown(edited)
hunks: list[tuple[str, str]] = []
for tag, i1, i2, j1, j2 in SequenceMatcher(
None, a, b, autojunk=False
).get_opcodes():
if tag == "equal":
continue
search = "\n\n".join(a[i1:i2])
replace = "\n\n".join(b[j1:j2])
if tag == "insert":
if i1:
search = a[i1 - 1]
replace = f"{a[i1 - 1]}\n\n{replace}"
elif a:
search = a[0]
replace = f"{replace}\n\n{a[0]}"
# else: base is empty — the hunk is inert (empty search is
# skipped by apply_patch); saving a translation of an empty
# page records nothing applicable.
hunks.append((search, replace))
return Patch(hunks=hunks)
def hybrid_markdown(data: Data, node: Node, path: str, lang: str) -> str:
"""The served Markdown for ``lang``: per chunk the translation from
``Data.trans``, unless missing or marked no-translate (fallback to the
original chunk), then the language's user patches applied in order.
Not gated on ``node.langs`` (get_translation is the gated view): the
editor save path diffs against this even for a language's first patch.
"""
hybrid = join_chunks(
[
data.chunks.get(h, "")
if h in node.no_trans
else data.trans.get(h, {}).get(lang) or data.chunks.get(h, "")
for h in node.chunks or []
]
)
for patch in data.patches.get(f"{path}:{lang}", []):
hybrid = apply_patch(hybrid, patch)
return hybrid
def add_patch(
data: Data, node: Node, path: str, lang: str, edited: str, base: str | None = None
) -> bool:
"""Record a translated-view edit as a user Patch: the minimal diff of
``edited`` against ``base`` (default: the currently served hybrid),
appended to the language's patch list. Patches alone make the
translated version exist, so ``node.langs`` is set. Returns True when
a patch was stored. Pure data ops the caller wraps in a transaction
and invalidates."""
patch = make_patch(
base if base is not None else hybrid_markdown(data, node, path, lang), edited
)
if not patch.hunks:
return False
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
node.langs[lang] = True
return True
def set_title_translation(data: Data, node: Node, lang: str, title: str) -> bool:
"""Record (or drop) a per-language title override: a fragment in
``Data.trans`` keyed by the ORIGINAL title's chunk hash — the same
storage machine title translations use, overriding them. Sending the
original's text drops the override. Returns True when anything changed.
Pure data ops the caller wraps in a transaction and invalidates."""
key = chunk_key(node.title)
current = data.trans.get(key, {}).get(lang)
if title == node.title:
if current is None:
return False
del data.trans[key][lang]
return True
if current == title:
return False
data.trans.setdefault(key, {})[lang] = title
node.langs[lang] = True
return True
def clear_translations(data: Data) -> None:
"""Drop all machine translations (``Data.trans``) and rebuild the
availability index (``node.langs``) from the surviving user patches
patches alone make a language exist on a page. Pure data ops the
caller wraps in a transaction and invalidates."""
data.trans.clear()
patch_langs: dict[str, set[str]] = {}
for key in data.patches:
path, _, lang = key.rpartition(":")
patch_langs.setdefault(path, set()).add(lang)
def walk(nodes: dict[str, Node], prefix: str) -> None:
for slug, node in nodes.items():
path = f"{prefix}/{slug}" if prefix else slug
node.langs = {lang: True for lang in patch_langs.get(path, ())}
walk(node.children, path)
walk(data.menu, "")
def title_map(data: Data, lang: str) -> dict[str, str]:
"""path -> translated title for every node that has one.
Titles are chunks too (docs/migrate.md): keyed by the hash of the
title text, so editing a title invalidates its translations. Nodes
without an entry fall back to their original title in views as do
nodes whose primary language IS ``lang`` (their original title already
is in that language).
"""
titles = {}
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in nodes.items():
path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited
if node.title and node_lang != lang:
t = data.trans.get(chunk_key(node.title), {}).get(lang)
if t:
titles[path] = t
walk(node.children, path, node_lang)
walk(data.menu, "", ORIGINAL_LANGUAGE)
return titles
def subtree_languages(node: Node) -> set[str]:
"""Languages available anywhere in the node's subtree (the union of the
``langs`` indexes). Category placeholder pages select their language
from this: they have no chunks of their own, but their title,
navigation and card text localize wherever a translation exists."""
langs = set(node.langs)
for child in node.children.values():
langs |= subtree_languages(child)
return langs
def get_translation(data: Data, path: str, lang: str) -> Translation | None:
"""The translation of the page at ``path`` for ``lang``, or None.
None when the page does not exist or is not available in ``lang``:
``node.langs`` is the availability index (a stale key is benign the
"translation" then just renders as the original).
"""
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None or lang not in node.langs:
return None
return Translation(
markdown=hybrid_markdown(data, node, path, lang),
titles=title_map(data, lang),
)
+114 -68
View File
@@ -25,14 +25,18 @@ schemes stay visible; manually labelled links are untouched), and
``H~2~O`` / ``x^2^`` give sub/superscripts. ``H~2~O`` / ``x^2^`` give sub/superscripts.
render() also builds the layout structure: the top-level blocks are render() also builds the layout structure: the top-level blocks are
segmented for the column layout h1/h2 headings, ``.wide`` blocks and segmented for the column layout h1/h2 headings and ``.wide`` blocks
margin-breakout blocks (``.margin``, ``::: aside``) stand on their own, stand on their own, the runs between them are wrapped in
the runs between them are wrapped in ``<div class="colseg">`` (tagged ``<div class="colseg">`` (tagged
``.cols`` when the segment holds enough text COLS_TEXT in at least ``.cols`` when the segment holds enough text COLS_TEXT in at least
COLS_PARAS paragraphs or one paragraph long enough to turn .breakable, COLS_PARAS paragraphs or one paragraph long enough to turn .breakable,
unless a ``::: nocols`` container opts it out; unless a ``::: nocols`` container opts it out;
in column segments, paragraphs past BREAKABLE_TEXT are marked in column segments, paragraphs past BREAKABLE_TEXT are marked
``.breakable`` so they may split across columns). The result carries ``.breakable`` so they may split across columns). Margin-breakout boxes
(``.margin``, ``::: aside``) stay inside the segment at their anchor
point; pagerite.css takes them out of flow (absolute, off the article's
left border, into the side zone), so the columns flow through as if the
box wasn't there. The result carries
``multicol`` when the whole body justifies columns (views.py puts the ``multicol`` when the whole body justifies columns (views.py puts the
class on the article); how many columns (never more than two), whether class on the article); how many columns (never more than two), whether
the margin breakout applies and every other viewport adaptation is then the margin breakout applies and every other viewport adaptation is then
@@ -152,14 +156,27 @@ def _image_rule(
page = env.get("page_path", "") page = env.get("page_path", "")
token.attrs["src"] = f"/{page}/{src}" if page else f"/{src}" token.attrs["src"] = f"/{page}/{src}" if page else f"/{src}"
token.attrs["alt"] = self.renderInlineAsText(token.children, options, env) token.attrs["alt"] = self.renderInlineAsText(token.children, options, env)
img = self.renderToken(tokens, idx, options, env)
if len(tokens) == 1: if len(tokens) == 1:
# The only inline content of its paragraph: render as a block # The only inline content of its paragraph: render as a block
# figure, captioned when titled. (The <p> wrapper is dropped by # figure, captioned when titled. (The <p> wrapper is dropped by
# _unwrap_lone_figures below.) # _unwrap_lone_figures below.) {.margin} positions the whole
# figure, so it moves from the img onto the figure wrapper — left
# on the img, the margin-breakout CSS would pull the image out of
# the figure (and mostly off-screen), leaving the caption behind.
classes = (token.attrs.get("class") or "").split()
figure_class = ""
if "margin" in classes:
classes.remove("margin")
if classes:
token.attrs["class"] = " ".join(classes)
else:
del token.attrs["class"]
figure_class = ' class="margin"'
img = self.renderToken(tokens, idx, options, env)
title = token.attrs.get("title") title = token.attrs.get("title")
caption = f"<figcaption>{escapeHtml(title)}</figcaption>" if title else "" caption = f"<figcaption>{escapeHtml(title)}</figcaption>" if title else ""
return f"<figure>{img}{caption}</figure>" return f"<figure{figure_class}>{img}{caption}</figure>"
img = self.renderToken(tokens, idx, options, env)
# Inline with other content: a plain inline image. # Inline with other content: a plain inline image.
return img return img
@@ -176,7 +193,13 @@ def _unwrap_lone_figures(state) -> None:
for i, token in enumerate(tokens): for i, token in enumerate(tokens):
if token.type != "inline" or not token.children: if token.type != "inline" or not token.children:
continue continue
[child] = token.children if len(token.children) == 1 else [None] # Attrs consumed out of the text (e.g. {style=...} space-separated
# on the image's own line) leave empty text tokens behind — strip
# them so the lone-image check is not thrown off by user styling.
children = [c for c in token.children if c.type != "text" or c.content]
if children:
token.children = children
[child] = children if len(children) == 1 else [None]
if child and child.type == "image": if child and child.type == "image":
if ( if (
tokens[i - 1].type == "paragraph_open" tokens[i - 1].type == "paragraph_open"
@@ -268,8 +291,8 @@ def _container_attrs(state) -> None:
The container plugin's default render is a plain renderToken, so the The container plugin's default render is a plain renderToken, so the
name and brace attributes must live on the token itself and being a name and brace attributes must live on the token itself and being a
core rule (rather than a render rule) lets the segmentation in core rule (rather than a render rule) lets the segmentation in
render() see the classes (::: aside's margin breakout, the ::: nocols render() see the classes (the ::: nocols opt-out, {.wide}
opt-out, {.wide} containers). containers).
""" """
for token in state.tokens: for token in state.tokens:
if token.type != "container_block_open": if token.type != "container_block_open":
@@ -289,9 +312,11 @@ def _block_attrs(state) -> None:
"""Apply `{.class key=value}` on a block's last line to the block. """Apply `{.class key=value}` on a block's last line to the block.
The inline attrs plugin only covers attributes right after an image, The inline attrs plugin only covers attributes right after an image,
code span or link; this extends the same brace syntax to whole blocks, code span or link; this extends the same brace syntax to whole blocks.
e.g. a paragraph ending with a `{.wide}` line (no blank line between) A paragraph takes them at the end of its last line, either directly
gets the `wide` class and thereby breaks out of the column layout. (a trailing `{.wide}` line, no blank line between) or space-separated
at the end of the text (`some text {.small}`) a space means the
braces belong to the block, not to an image or link before them.
A lone `{...}` paragraph applies to the previous block instead (this A lone `{...}` paragraph applies to the previous block instead (this
is how headings take attributes, since a heading's next line always is how headings take attributes, since a heading's next line always
starts a new paragraph). Runs before the typographer so quotes inside starts a new paragraph). Runs before the typographer so quotes inside
@@ -302,18 +327,20 @@ def _block_attrs(state) -> None:
if token.type != "inline" or not token.children: if token.type != "inline" or not token.children:
continue continue
text = token.children[-1] text = token.children[-1]
if ( if text.type != "text":
text.type != "text"
or not text.content.startswith("{")
or not text.content.endswith("}")
):
continue continue
m = re.search(r"(\{[^{}]*\})\s*$", text.content)
if not m:
continue
start = m.start(1)
if start and not text.content[start - 1].isspace():
continue # glued to the text — literal, or inline attrs
try: try:
_, attrs = parse_attrs(text.content.strip()) _, attrs = parse_attrs(m.group(1))
except ParseError: except ParseError:
continue continue
standalone = len(token.children) == 1 standalone = len(token.children) == 1
if not standalone and token.children[-2].type != "softbreak": if not standalone and start == 0 and token.children[-2].type != "softbreak":
continue continue
# The target: the enclosing block for a trailing attrs line, or the # The target: the enclosing block for a trailing attrs line, or the
# previous same-level block for a standalone attrs paragraph — # previous same-level block for a standalone attrs paragraph —
@@ -347,8 +374,15 @@ def _block_attrs(state) -> None:
tokens[own].hidden = True tokens[own].hidden = True
token.children = [] token.children = []
tokens[i + 1].hidden = True tokens[i + 1].hidden = True
else: elif start == 0:
del token.children[-2:] del token.children[-2:]
else:
# Braces space-separated at the end of a text line: strip them
# (a whitespace-only remainder means they were on a line of
# their own after all — drop the softbreak too).
text.content = text.content[:start].rstrip()
if not text.content and token.children[-2].type == "softbreak":
del token.children[-2:]
#: Minimum number of in-body h1/h2 headings for section anchors to be #: Minimum number of in-body h1/h2 headings for section anchors to be
@@ -401,7 +435,10 @@ def _heading_ids(state) -> None:
heads = [ heads = [
(i, token) (i, token)
for i, token in enumerate(tokens) for i, token in enumerate(tokens)
if token.type == "heading_open" and token.tag in ("h1", "h2") and token.level == 0 and i != first_h1 if token.type == "heading_open"
and token.tag in ("h1", "h2")
and token.level == 0
and i != first_h1
] ]
if len(heads) < ANCHOR_MIN_HEADINGS: if len(heads) < ANCHOR_MIN_HEADINGS:
return return
@@ -426,39 +463,53 @@ def _heading_ids(state) -> None:
wrap(i, token, f"#{hid}") wrap(i, token, f"#{hid}")
md = ( def make_md(*, verbatim: bool = False) -> MarkdownIt:
MarkdownIt( """A fully configured parser. The module-level ``md`` (below) is the
"default", render instance; ``verbatim=True`` builds the segmentation instance for
{ segments.py, where token text must stay byte-identical to the source so
"html": True, prose spans can be spliced back by offset: no typographer (quotes and
"highlight": _highlight, dashes stay straight), no tasklist label wrapping (the item text stays
"typographer": True, a plain text token), and soft line breaks (wrapped prose merges into
"breaks": True, one segment instead of splitting at hardbreaks)."""
}, parser = (
MarkdownIt(
"default",
{
"html": True,
"highlight": _highlight,
"typographer": not verbatim,
"breaks": not verbatim,
},
)
.use(attrs_plugin)
.use(admon_plugin)
.use(container_plugin, "block", validate=_container_validate)
.use(footnote_plugin)
.use(deflist_plugin)
# label wrapping (render) puts the item text inside the checkbox
# <label> html_inline; without it the text stays a plain token.
.use(
tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim
)
.use(gfm_autolink_plugin)
.use(sub_plugin)
.use(superscript_plugin)
) )
.use(attrs_plugin) parser.add_render_rule("image", _image_rule)
.use(admon_plugin) parser.add_render_rule("fence", _fence_rule)
.use(container_plugin, "block", validate=_container_validate) # GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
.use(footnote_plugin) parser.options["alerts"] = True
.use(deflist_plugin) # Block attrs must be stripped before the typographer curlifies their quotes.
# label_after: the item text is wrapped in <label for> after the parser.core.ruler.before("replacements", "block_attrs", _block_attrs)
# checkbox, so clicking the text toggles it. parser.core.ruler.push("container_attrs", _container_attrs)
.use(tasklists_plugin, enabled=True, label=True, label_after=True) parser.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
.use(gfm_autolink_plugin) parser.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
.use(sub_plugin) parser.core.ruler.push("shorten_autolinks", _shorten_autolinks)
.use(superscript_plugin) parser.core.ruler.push("heading_ids", _heading_ids)
) return parser
md.add_render_rule("image", _image_rule)
md.add_render_rule("fence", _fence_rule)
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule. md = make_md()
md.options["alerts"] = True
# Block attrs must be stripped before the typographer curlifies their quotes.
md.core.ruler.before("replacements", "block_attrs", _block_attrs)
md.core.ruler.push("container_attrs", _container_attrs)
md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
md.core.ruler.push("shorten_autolinks", _shorten_autolinks)
md.core.ruler.push("heading_ids", _heading_ids)
# Text-length thresholds (visible characters, code blocks excluded) for the # Text-length thresholds (visible characters, code blocks excluded) for the
@@ -482,11 +533,12 @@ _PARA_OPEN_RE = re.compile(r"<p[\s>]")
_PARA_RE = re.compile(r"<p((?:\s[^>]*)?)>(.*?)</p>", re.S) _PARA_RE = re.compile(r"<p((?:\s[^>]*)?)>(.*?)</p>", re.S)
# Classes that take their block out of the column flow: .wide is a # Classes that take their block out of the column flow: .wide is a
# full-width separator, .margin/.aside float in the side zone at the # full-width separator that splits the column segments. Margin-breakout
# article's left (they must be direct article children for that — the zone # boxes (.margin/.aside) are NOT boundaries: they stay inside the segment
# rules key off it — never inside a column). # at their anchor point, and CSS positions them absolutely out of the
# article's left border (the zone rules anchor off the article), so the
# column flow is unaffected.
_WIDE = "wide" _WIDE = "wide"
_BREAKOUT = ("margin", "aside")
class Rendered(NamedTuple): class Rendered(NamedTuple):
@@ -546,14 +598,10 @@ def _top_level_blocks(tokens: list) -> list[list]:
def _is_boundary(block: list) -> bool: def _is_boundary(block: list) -> bool:
"""True for blocks that never go inside a column segment (see the """True for blocks that never go inside a column segment (see the
_WIDE/_BREAKOUT comment above): h1/h2 headings, anything carrying _WIDE comment above): h1/h2 headings and anything carrying .wide."""
.wide, and blocks whose own element carries .margin/.aside for a
lone-image paragraph (which renders as a <figure>) the image's classes
count as the block's own."""
first = block[0] first = block[0]
if first.type == "heading_open" and first.tag in ("h1", "h2"): if first.type == "heading_open" and first.tag in ("h1", "h2"):
return True return True
own = _classes(first)
for token in block: for token in block:
if _WIDE in _classes(token): if _WIDE in _classes(token):
return True return True
@@ -561,9 +609,7 @@ def _is_boundary(block: list) -> bool:
children = token.children or [] children = token.children or []
if any(_WIDE in _classes(c) for c in children): if any(_WIDE in _classes(c) for c in children):
return True return True
if len(children) == 1 and children[0].type == "image": return False
own |= _classes(children[0])
return bool(own & set(_BREAKOUT))
def render( def render(
@@ -580,8 +626,8 @@ def render(
same pipeline as an explicit one (first-h1 anchor treatment included). same pipeline as an explicit one (first-h1 anchor treatment included).
The top-level blocks are grouped into column segments: boundary blocks The top-level blocks are grouped into column segments: boundary blocks
(h1/h2 headings, .wide, margin-breakout blocks see _is_boundary) are (h1/h2 headings, .wide see _is_boundary) are rendered bare, the runs
rendered bare, the runs between them wrapped in <div class="colseg">. between them wrapped in <div class="colseg">.
A segment is tagged .cols when it holds enough text (COLS_TEXT) in at A segment is tagged .cols when it holds enough text (COLS_TEXT) in at
least two paragraphs (COLS_PARAS) or one breakable-length paragraph, least two paragraphs (COLS_PARAS) or one breakable-length paragraph,
and no ::: nocols container; its long paragraphs are marked .breakable; and no ::: nocols container; its long paragraphs are marked .breakable;
+166 -11
View File
@@ -1,23 +1,178 @@
"""Kanta schema migrations, discovered by name (``migrate_vN``). """Kanta schema migrations, discovered by name (``migrate_vN``).
Each function receives the raw state dict (JSON-level: bytes are base64 Each function receives the raw state dict (JSON-level: bytes are base64
strings) before it is decoded into ``Data`` structs, and runs exactly once strings, datetimes RFC 3339 strings, struct fields with default values
omitted) before it is decoded into ``Data`` structs, and runs exactly once
per database based on its recorded version. per database based on its recorded version.
All storage/schema upgrades live here including on-disk file work, which
runs through files.py's file store (imported lazily: files.py owns the store
and state.py passes this module to Kanta; at migration time, during lifespan
``kanta.open()``, both modules are fully loaded).
""" """
import base64 import base64
import re
from pathlib import Path
from pagerite.chunks import chunk_key, chunk_markdown
from pagerite.data import prettify
def _append_order(nodes: dict) -> float:
"""Raw-dict equivalent of data.append_order (order keys may be absent)."""
return max((n.get("order", 0) for n in nodes.values()), default=0) + 1
def _ensure(menu: dict, path: str) -> dict:
"""Raw-dict equivalent of state._ensure: the node dict at ``path``,
creating it and any missing ancestors (content-less category labels)
appended at the end of their level."""
nodes = menu
node = None
for seg in path.split("/"):
node = nodes.get(seg)
if node is None:
node = {"title": prettify(seg), "order": _append_order(nodes)}
nodes[seg] = node
nodes = node.setdefault("children", {})
return node
def migrate_v1(d: dict) -> None: def migrate_v1(d: dict) -> None:
"""Move in-database file blobs to the on-disk content-addressed store.""" """Move in-database file blobs to the on-disk content-addressed store,
and rebuild the legacy flat page store (``pages``) as the menu tree."""
files = d.pop("files", None) files = d.pop("files", None)
if not files: if files:
return from pagerite.files import file_store
# Deferred import: app.py owns the file store and passes this module to
# Kanta; at migration time (lifespan open) the module is fully loaded.
from pagerite.app import file_store
for name, body in files.items(): for name, body in files.items():
if isinstance(body, str): # JSON-level bytes are base64 strings if isinstance(body, str): # JSON-level bytes are base64 strings
body = base64.b64decode(body) body = base64.b64decode(body)
file_store.put(name, body) file_store.put(name, body)
pages = d.pop("pages", None)
if not pages:
return
menu = d.setdefault("menu", {})
for path, page in pages.items():
node = _ensure(menu, path)
node["title"] = page["title"]
node["content"] = page["markdown"]
for key in ("banner", "published", "order", "created", "modified"):
if key in page:
node[key] = page[key]
#: Extension-less file links: uploaded images are linked as /_f/<hash>
#: and the server negotiates avif/webp/jpg from the Accept header.
_DERIVATIVE_LINK = re.compile(r"(/_f/[0-9a-f]{12})\.(?:avif|webp)\b")
def _backfill_derivatives() -> None:
"""Create missing AVIF/WebP/JPEG derivatives for files stored before
they were introduced (older uploads may have only the original plus
AVIF, and SVGs no raster variants at all). WebP/JPEG are re-encoded
from an existing AVIF when available, everything else from the
original (SVGs rasterized first)."""
from pagerite.files import (
IMAGE_MAXSIZE,
IMAGE_WEBP_QUALITY,
IMAGE_JPG_QUALITY,
_avif_to_format,
_svg_to_png,
_to_avif,
file_store,
)
try:
paths = [f for f in file_store.path.iterdir() if f.is_file()]
except FileNotFoundError:
return
groups: dict[str, list[Path]] = {}
for p in paths:
groups.setdefault(p.name.partition(".")[0], []).append(p)
for digest, files in groups.items():
names = {p.name for p in files}
source = next(
(p for p in files if ".orig." in p.name or p.suffix == ".svg"), None
)
if source is None:
continue # plain as-is file, no derivatives to make
avif = file_store.get(f"{digest}.avif")
if avif is None:
ext = source.suffix
body = source.read_bytes()
if ext == ".svg":
png = _svg_to_png(body, IMAGE_MAXSIZE)
if png is None:
continue
body, ext = png, ".png"
converted = _to_avif(body, ext)
if converted is None:
continue
file_store.put(f"{digest}.avif", converted)
avif = file_store.get(f"{digest}.avif")
for fmt, quality in (
("webp", IMAGE_WEBP_QUALITY),
("jpg", IMAGE_JPG_QUALITY),
):
if f"{digest}.{fmt}" not in names:
file_store.put(
f"{digest}.{fmt}", _avif_to_format(avif[0], f".{fmt}", quality)
)
def migrate_v2(d: dict) -> None:
"""Extension-less image links: strip .avif/.webp extensions from /_f/
links in page content and banners (the server now negotiates the format
by Accept header), backfill missing AVIF/WebP/JPEG derivatives on disk,
and drop the obsolete render-counter field ``version`` (invalidation is
an in-memory concern now, not database state)."""
def walk(nodes: dict) -> None:
for node in nodes.values():
for field in ("content", "banner"):
if isinstance(node.get(field), str):
node[field] = _DERIVATIVE_LINK.sub(r"\1", node[field])
walk(node.get("children") or {})
walk(d.get("menu") or {})
d.pop("version", None)
_backfill_derivatives()
def migrate_v3(d: dict) -> None:
"""Content-addressed chunk storage (docs/migrate.md): split every
node's string ``content`` into block chunks stored once per content
hash in the new ``chunks`` store; the node keeps the ordered hash
list as ``chunks`` (an absent content stays absent, i.e. None = a
pure category label; "" chunks to an empty list = an empty page).
Chunk keys are 9-byte blake3 digests; at this raw JSON level they are
base64 strings (decoding into the structs restores ``bytes`` keys).
``trans``/``patches`` start empty; the translator job fills them and
maintains the ``langs`` index as translations land. ``language``,
``no_trans`` and ``langs`` need nothing struct defaults cover them.
"""
store = d.setdefault("chunks", {})
d.setdefault("trans", {})
patches = d.setdefault("patches", {})
def walk(nodes: dict) -> None:
for node in nodes.values():
content = node.pop("content", None)
if isinstance(content, str):
hashes = []
for chunk in chunk_markdown(content):
key = base64.b64encode(chunk_key(chunk)).decode()
store.setdefault(key, chunk)
hashes.append(key)
node["chunks"] = hashes
walk(node.get("children") or {})
walk(d.get("menu") or {})
# Article paths never carry a leading slash in keys (docs/migrate.md).
# The only path-keyed store starts empty here, so this is defensive
# for databases that went through a downgrade/upgrade cycle.
for key in [k for k in patches if k.startswith("/")]:
patches[key.lstrip("/")] = patches.pop(key)
+253
View File
@@ -0,0 +1,253 @@
"""Public content pages: front page, sitemap, robots, and the catch-all.
``GET /{path:path}`` resolves a slug path against the menu tree and renders
the page (or a category placeholder, or 404); it must be registered AFTER
the fastapi-vue asset routes so built frontend files win over content slugs
(see app.py). Requests are recorded in analytics (crawler hits and 404s
here, visits via the /_ws socket in tracking.py).
"""
import asyncio
import logging
from datetime import UTC, datetime
from email.utils import format_datetime
from xml.sax.saxutils import escape as xml_escape
from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import RedirectResponse, Response
from pagerite import i18n, state
from pagerite.data import Node, resolve, sorted_nodes
from pagerite.state import (
SITE_URL,
_html_response,
_is_reserved,
analytics_store,
data,
)
from pagerite.tracking import (
_client_ip,
_enrich_client,
_query_suffix,
_schedule_client_enrichment,
_track_entry,
)
logger = logging.getLogger(__name__)
router = APIRouter()
def _http_date(dt: datetime) -> str:
"""RFC 7231 date for the Last-Modified header."""
return format_datetime(dt.astimezone(UTC), usegmt=True)
def _is_trackable_path(path: str) -> bool:
"""Content URLs only: skip auth endpoints and reserved/machinery paths."""
if not path:
return True
if path == "auth" or path.startswith("auth/"):
return False
return not _is_reserved(path)
@router.get("/")
async def front_page(request: Request) -> Response:
"""Render the front page (slug path "")."""
return await show_page(request, "")
@router.get("/sitemap.xml")
async def sitemap(request: Request) -> Response:
"""Dynamically generate a sitemap of all published article pages."""
base = SITE_URL or str(request.base_url).rstrip("/")
entries: list[tuple[str, datetime, int]] = []
def walk(
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
) -> None:
first_content_slug = next(
(
slug
for slug, node in sorted_nodes(nodes)
if node.published and node.chunks is not None
),
None,
)
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
depth = path.count("/") if path else 0
if (
not parent_has_content
and slug == first_content_slug
and node.published
and node.chunks is not None
and depth > 0
):
depth -= 1
if node.published and node.chunks is not None:
entries.append((path, node.modified, depth))
if node.children:
walk(node.children, path, node.chunks is not None)
walk(data.menu, "")
def priority(depth: int) -> float:
return max(0.1, 1.0 - depth * 0.2)
lines = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
]
for path, modified, depth in entries:
loc = xml_escape(f"{base}/{path}" if path else base)
lastmod = (
modified.astimezone(UTC)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
lines.append(
f" <url>"
f"<loc>{loc}</loc>"
f"<lastmod>{lastmod}</lastmod>"
f"<priority>{priority(depth):.1f}</priority>"
f"</url>"
)
lines.append("</urlset>")
return Response(
"\n".join(lines),
media_type="application/xml",
headers={"cache-control": "no-cache"},
)
@router.get("/robots.txt")
async def robots_txt(request: Request) -> Response:
"""Allow content crawling, keep the SSO login (/auth/) and the
admin-gated API (/_api) out of search results, and point crawlers at
the sitemap."""
base = SITE_URL or str(request.base_url).rstrip("/")
body = f"User-agent: *\nAllow: /\nDisallow: /auth/\nDisallow: /_api\nSitemap: {base}/sitemap.xml\n"
return Response(
body,
media_type="text/plain",
headers={"cache-control": "no-cache"},
)
@router.get("/{path:path}", response_model=None)
async def show_page(request: Request, path: str) -> Response:
"""Render the content page at a slug path, or 404.
A node without content is a category label: its URL renders a
placeholder page (nav links point straight at its first child).
"""
path = path.strip("/")
ua = request.headers.get("user-agent", "")
accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page.
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
# as abuse in analytics.
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
raise HTTPException(404)
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is not None and node.published and node.chunks is not None:
# Language selection (docs/localization.md): ?lang= wins when a
# translation exists, else header logic. Analytics keep the raw
# Accept-Language header regardless of the selection.
query_lang = request.query_params.get("lang")
lang = i18n.select_language(
query_lang,
accept_language,
lambda tag: tag in node.langs,
original=i18n.primary_lang(data.menu, path),
)
# A ?lang= override is replicated onto the page's navigation links
# (link_lang), so clicks and prefetches stay in the chosen language.
# Query and header-selected renders of the same language differ in
# their links, so link_lang is part of the ETag and body cache key.
link_lang = i18n.base_tag(query_lang or "")
# no-cache forbids serving a stored page without revalidation
# (browsers would otherwise cache heuristically and serve stale
# pages, e.g. after a theme change). In-session speed instead comes
# from pagerite.js's in-memory page cache (preload everything, never
# fetch on navigation); the ETag just makes those one-time preload
# fetches and any revalidation cheap.
etag = f'"{path}@{node.modified.timestamp()}g{state._render_gen}l{lang}q{link_lang}"'
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
if _is_trackable_path(path):
flushed = _track_entry(path, request)
_schedule_client_enrichment(flushed)
return _html_response(
request,
"page",
path,
headers={
"etag": etag,
"last-modified": _http_date(node.modified),
"cache-control": "no-cache",
},
lang=lang,
link_lang=link_lang,
)
if node is not None and node.published and node.chunks is None:
# Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real).
# Language selection as on content pages, but over the whole
# subtree's availability: the category has no chunks of its own —
# its heading, the navigation and the cards' text localize from
# the title map and the target articles' translations.
query_lang = request.query_params.get("lang")
subtree_langs = i18n.subtree_languages(node)
lang = i18n.select_language(
query_lang,
accept_language,
lambda tag: tag in subtree_langs,
original=i18n.primary_lang(data.menu, path),
)
link_lang = i18n.base_tag(query_lang or "")
if _is_trackable_path(path):
flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(
request,
"category",
path,
404,
headers={
"last-modified": _http_date(node.modified),
"cache-control": "no-cache",
},
lang=lang,
link_lang=link_lang,
)
if node is None and not path:
# No front page (no top-level node with slug ""): "/" opens the
# first item of the navigation instead.
for slug, item in sorted_nodes(data.menu):
if item.published:
return RedirectResponse(f"/{slug}")
if _is_trackable_path(path):
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(request, "not-found", path, 404)
+4 -3
View File
@@ -21,6 +21,7 @@ ASSETS = Path(__file__).with_name("seed-assets")
def _asset(name: str) -> bytes: def _asset(name: str) -> bytes:
return (ASSETS / name).read_bytes() return (ASSETS / name).read_bytes()
WELCOME = """\ WELCOME = """\
Welcome to your new **Pagerite** site. Everything you see is a page written in Markdown, served from a pretty URL, and editable right here in the browser. Welcome to your new **Pagerite** site. Everything you see is a page written in Markdown, served from a pretty URL, and editable right here in the browser.
@@ -51,7 +52,7 @@ The 🖊️ pens open a tabbed editor over the page you are viewing:
The URL is the structure: a page at `docs/markdown` lives under `docs`, and the menus are derived from that. Slugs are lowercase ASCII (`a-z 0-9 - _`). A node without content is a category label it renders a placeholder and its menu link points at its first child page. This site's own `docs` label demonstrates that, and the sidebar on this page shows the two submenu levels below it. The URL is the structure: a page at `docs/markdown` lives under `docs`, and the menus are derived from that. Slugs are lowercase ASCII (`a-z 0-9 - _`). A node without content is a category label it renders a placeholder and its menu link points at its first child page. This site's own `docs` label demonstrates that, and the sidebar on this page shows the two submenu levels below it.
Images and files uploaded anywhere land in a content-addressed store served from `/_f/{hash}.ext`, so links survive page moves. The article editor's format bar and copy-paste both upload images for you. Images and files uploaded anywhere land in a content-addressed store served from `/_f/{hash}`, so links survive page moves. The server picks AVIF, WebP or JPEG from your browser's Accept header. The article editor's format bar and copy-paste both upload images for you.
{dates} {dates}
""" """
@@ -114,7 +115,7 @@ Headings from `##` down organize the article. On pages with at least three of th
> and a blank `>` line starts a new paragraph. > and a blank `>` line starts a new paragraph.
> [!NOTE] > [!NOTE]
> GitHub-style alerts NOTE, TIP, IMPORTANT, WARNING, CAUTION > GitHub-style alerts `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION`
> render as callout boxes. > render as callout boxes.
``` ```
@@ -122,7 +123,7 @@ Headings from `##` down organize the article. On pages with at least three of th
> and a blank `>` line starts a new paragraph. > and a blank `>` line starts a new paragraph.
> [!NOTE] > [!NOTE]
> GitHub-style alerts NOTE, TIP, IMPORTANT, WARNING, CAUTION > GitHub-style alerts `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION`
> render as callout boxes. > render as callout boxes.
## Code ## Code
+622
View File
@@ -0,0 +1,622 @@
"""Segmented translation round trip: prose out, translations back in.
A translator model mangles anything that is not plain prose sentinels get
renumbered, ``![`` becomes sentence punctuation, stray ``<br>`` tags appear.
So the model is never shown any of it: a fragment (a Markdown chunk or a
node title) is parsed with the project's own markdown-it setup
(``markdown.make_md(verbatim=True)`` extensions included, so container,
attrs, footnote and tasklist syntax never leaks into text tokens) and split
into **prose segments**: the merged text runs, plus image alt texts and
link/image titles. Only those cross the wire, as a plain list of strings
(Job.texts / Result.texts in translate.py) accompanied, per segment, by
a CONTEXT (Job.contexts): a segment carved out of a larger block (a link
text, a partial run) carries the block's plain text, so the model sees the
sentence it lives in; whole-block segments are self-contextualizing and
carry "". Title fragments carry the article's opening instead (assigned by
the dispatcher from TransItem.context).
Reassembly is server-side offset splicing, not text the model produced:
each segment's source span was located at dispatch (``split``), and
``join`` swaps in the translations. Markup therefore cannot break it
never left the server. A returned segment must still be pure prose itself
(the model could inject markup INTO a segment); anything else count
mismatch, empty segment, markup tokens rejects the whole result and the
fragment stays pending.
A block of plain text, prose links and paired text formatting
(strong/em/s) crosses as ONE segment link texts and formatted text
inline, in sentence context, with the Markdown stripped (the model
mangles it: sentinels get renumbered, ``**`` gets dropped or moved)
because a label translated apart from its sentence comes back
grammatically incompatible with it (case government, particles, word
order). ``join`` re-inserts the link/formatting markdown into the
translated block at fuzzily matched positions (``_place_marks``): no
markers on the wire, the boundaries are found by aligning the mark's
source words to the translation's words by form similarity (``_find_mark``
inflection, dropped articles and reordering tolerated), with the
source/translation weight ratio as fallback (the CJK path, where
cross-script form similarity is nil). Placement is approximate: better a
coherent sentence with a slightly shifted link than separately translated
snippets that don't fit together. Blocks with any other inline markup
(code, images, HTML) still split into runs at those boundaries.
Locating is best effort: a run that is not a verbatim source substring
(entity-decoded text, backslash escapes) is skipped it simply stays in
the original language. So is any piece containing "<": "<" is the
prose/markup boundary on the wire translators cut their output there,
so such pieces could not survive the round trip.
"""
import bisect
import difflib
import re
from typing import NamedTuple
from pagerite.markdown import make_md
#: The segmentation parser: the project's own markdown-it, verbatim flavor
#: (see make_md). Never used for rendering.
_MD = make_md(verbatim=True)
#: Any Unicode letter (digits and underscore are not prose).
_LETTER = re.compile(r"[^\W\d_]")
#: A GFM alert marker ([!NOTE] etc.) at the start of a blockquote's first
#: paragraph: syntax, not prose — stripped from the first segment.
_ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
#: Any {...} span: {placeholders} and attrs that ended up inside prose
#: (inline attrs are consumed by the parser; a lone {dates} is not).
_BRACES = re.compile(r"\{[^{}\n]*\}")
#: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]",
#: "[]" or a bare "]" (shortcut reference); the destination may nest one
#: level of parens. Best effort — a mis-scan fails the span-reconstruction
#: check in _linked_block and the block falls back to per-run segments.
_LINK_TAIL = re.compile(r"\](?:\((?:\\.|[^()\\]|\([^()]*\))*\)|\[(?:\\.|[^\]])*\])?")
#: Weight units for mapping link boundaries from source to translation:
#: a word counts 1 and so does every single CJK ideograph (kana runs count
#: as one) — CJK has no spaces to count words by. Punctuation and
#: whitespace count nothing, so mapped boundaries always land on unit
#: starts.
_UNIT = re.compile(
r"[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]" # CJK ideographs: one unit each
r"|[\u3040-\u309f\u30a0-\u30ff]+" # kana runs: one unit each
r"|\w+" # anything else word-like (Latin, Cyrillic, Hangul, digits)
)
class Mark(NamedTuple):
"""One inline link or paired formatting (strong/em/s) inside a
whole-block segment: the source weight (unit count, see _UNIT) at the
inner text's start and end (fallback for mapping the boundaries into
the translation when fuzzy word alignment finds nothing, _find_mark),
the exact source syntax around the text ("[" / "](url)", "**" / "**",
...) and the source text itself the words fuzzy alignment looks for,
and the fallback when the mapped slice comes out empty (better an
untranslated label than a broken "[](url)")."""
w_start: int
w_end: int
pre: str
post: str
inner: str
class Span(NamedTuple):
"""A segment's source span in the fragment: offsets for splicing the
translation back, the segment's source weight and the links to
re-insert into its translation (empty = a plain prose segment)."""
start: int
end: int
weight: int
marks: list[Mark]
def _weight(text: str) -> int:
"""The text's weight in translation-mapping units (see _UNIT)."""
return len(_UNIT.findall(text))
def _runs(children: list) -> list[str]:
"""Prose runs of an inline token's children, in order.
Text tokens merge across soft breaks into one run; every markup token
(emphasis, links, code, images, HTML, footnote refs, hard breaks) is a
run boundary. Link and image *text* is prose; autolink text (the URL
itself) is not. Image tokens contribute their alt-text children and
their title attribute.
"""
runs: list[str] = []
cur: list[str] = []
def flush() -> None:
if cur:
s = "".join(cur)
cur.clear()
if _LETTER.search(s):
runs.append(s)
skip = 0 # inside an autolink (its text is the URL — not prose)
for t in children:
if skip:
if t.type == "link_close":
skip -= 1
continue
if t.type == "text":
cur.append(t.content)
elif t.type == "softbreak":
cur.append("\n")
elif t.type == "link_open" and t.markup == "autolink":
flush()
skip = 1
elif t.type == "image":
flush()
if t.children:
runs.extend(_runs(t.children))
title = t.attrGet("title")
if title and _LETTER.search(title):
runs.append(title)
else:
flush()
if t.children:
runs.extend(_runs(t.children))
flush()
return runs
def _block_text(children: list) -> str:
"""The block's text as a reader sees it: text runs and link texts
merged (softbreaks as newlines); image alts, autolink URLs, code and
other markup content excluded. Used as the translation CONTEXT for
segments carved out of the block (link texts, partial runs): a lone
word translates differently than the same word inside its sentence."""
parts: list[str] = []
skip = 0 # inside an autolink (its text is the URL)
for t in children:
if skip:
if t.type == "link_close":
skip -= 1
continue
if t.type == "text":
parts.append(t.content)
elif t.type == "softbreak":
parts.append("\n")
elif t.type == "link_open" and t.markup == "autolink":
skip = 1
elif t.type == "image":
continue
elif t.children:
parts.append(_block_text(t.children))
return "".join(parts)
def _locate(source: str, needle: str, cursor: int) -> int:
"""The needle's offset in source at/after cursor, -1 when absent.
An occurrence preceded by a backslash is an escaped character, not the
token's source: keep looking (failing that, the run is skipped — it
stays in the original language).
"""
pos = source.find(needle, cursor)
while pos > 0 and source[pos - 1] == "\\":
pos = source.find(needle, pos + 1)
return pos
def _linked_block(
source: str, kids: list, cursor: int, strip_alert: bool
) -> tuple[Span, str] | None:
"""A whole-block segment for an inline of plain text, prose links and
paired text formatting (strong/em/s): (Span, wire text) with the links
and formatting as marks, or None when the block has any other shape
the caller then falls back to per-run segments.
The block crosses the wire as one prose piece, link texts and formatted
text inline (the model is never shown any Markdown it mangles it),
so a translation that inflects or reorders around them stays coherent;
join re-inserts the link/formatting syntax at weight-mapped positions.
The source span is located piece by piece and verified by
reconstruction; anything not byte-exact (entities, escapes, an odd
link tail) bails to the fallback.
"""
pieces: list[
tuple[str, str]
] = [] # (text, mark): "" plain, "link", else the delimiter
buf: list[str] = [] # current plain piece
link: list[str] | None = None # current mark's text parts
mark_kind = "" # the current mark's opener ("link" or the delimiter)
for tok in kids:
if tok.type in ("link_open", "strong_open", "em_open", "s_open"):
if link is not None or tok.markup == "autolink":
return None
if buf:
pieces.append(("".join(buf), ""))
buf = []
link = []
mark_kind = "link" if tok.type == "link_open" else tok.markup
elif tok.type in ("link_close", "strong_close", "em_close", "s_close"):
if (
link is None
or ("link" if tok.type == "link_close" else tok.markup) != mark_kind
):
return None
inner = "".join(link)
if not _LETTER.search(inner):
return None
pieces.append((inner, mark_kind))
link = None
elif tok.type in ("text", "softbreak"):
(link if link is not None else buf).append(
"\n" if tok.type == "softbreak" else tok.content
)
else: # code, images, HTML, footnote refs: run boundaries
return None
if link is not None:
return None # unbalanced (the parser should not do this)
if buf:
pieces.append(("".join(buf), ""))
if not any(mark for _, mark in pieces):
return None
if strip_alert and pieces and not pieces[0][1]:
# A GFM alert marker leading the blockquote's first paragraph is
# syntax; strip it from the wire text (it stays out of the span).
first = _ALERT.sub("", pieces[0][0], count=1)
if first.strip():
pieces[0] = (first, "")
else:
pieces.pop(0)
if not pieces:
return None
raw = "".join(text for text, _ in pieces)
lead = len(raw) - len(raw.lstrip())
wire = raw.strip()
if not _LETTER.search(wire) or "<" in wire or _BRACES.search(wire):
return None
# Locate each piece verbatim, in order; the source slices between the
# located pieces are then the link syntax, exact by construction.
located: list[tuple[int, int]] = []
pos = cursor
for text_, _ in pieces:
at = _locate(source, text_, pos)
if at == -1:
return None
located.append((at, at + len(text_)))
pos = at + len(text_)
span_start, span_end = located[0][0], located[-1][1]
marks: list[Mark] = []
offset = 0 # raw (pre-strip) plain-text offset of the current piece
for i, ((text_, kind), (s, e)) in enumerate(zip(pieces, located)):
if not kind:
offset += len(text_)
continue
# The syntax around the text: the gap between pieces goes to the
# mark on its left as post (so between two marks the whole "](u)["
# or "**" is the first's post); a block-leading mark takes its
# opener in front of its text ("[" or the delimiter), a
# block-trailing one the scanned link tail or the close delimiter.
if i == 0:
opener = "[" if kind == "link" else kind
if s < len(opener) or source[s - len(opener) : s] != opener:
return None
pre, span_start = opener, s - len(opener)
elif pieces[i - 1][1]:
pre = "" # the previous mark's post covers the whole gap
else:
pre = source[located[i - 1][1] : s]
if i + 1 < len(pieces):
post = source[e : located[i + 1][0]]
elif kind == "link":
m = _LINK_TAIL.match(source, e)
if m is None:
return None
post, span_end = m.group(), m.end()
else:
if source[e : e + len(kind)] != kind:
return None
post, span_end = kind, e + len(kind)
ps = min(max(offset - lead, 0), len(wire))
pe = min(max(offset + len(text_) - lead, 0), len(wire))
if pe <= ps:
return None
marks.append(
Mark(_weight(wire[:ps]), _weight(wire[:pe]), pre, post, wire[ps:pe])
)
offset += len(text_)
# Verify: the marks must reconstruct the source span exactly (the only
# real risk is the guessed tail of a trailing link).
rec: list[str] = []
mi = 0
for text_, kind in pieces:
if kind:
mark = marks[mi]
mi += 1
rec += [mark.pre, text_, mark.post]
else:
rec.append(text_)
if source[span_start:span_end] != "".join(rec):
return None
return Span(span_start, span_end, _weight(wire), marks), wire
def split(text: str) -> tuple[list[Span], list[str], list[str]]:
"""Split a fragment into (spans, segments, contexts): prose segments to
translate, their source spans in ``text`` for splicing the translations
back, and per-segment translation context.
A block of plain text, prose links and paired formatting (strong/em/s)
becomes ONE segment (link/formatted text inline, in context, Markdown
stripped), the links and formatting recorded as marks on its Span for
weight-mapped re-insertion in join. Other blocks split into text runs
at markup boundaries; runs containing {...} spans are carved further
the braces stay out of the wire text. A run that cannot be located
verbatim in the source contributes no segment. A segment's context is
its block's plain text when the segment was carved OUT of a larger
block (a partial run); a segment that IS the whole block (a plain
paragraph, a heading, a linked block) is self-contextualizing and gets
"".
"""
spans: list[Span] = []
segments: list[str] = []
contexts: list[str] = []
cursor = 0
blockquote_fresh = 0 # blockquote depth whose first inline is upcoming
def emit(run: str, at: int, ctx: str) -> None:
"""Carve {...} spans out of the located run; emit the prose pieces,
stripped padding whitespace stays in the template, off the wire.
Pieces containing "<" are never emitted: translators cut output at
the first "<" (the prose/markup boundary, scripts/translator.py),
so such a piece could not survive the round trip it stays in the
original language instead."""
pieces = []
pos = 0
for m in _BRACES.finditer(run):
pieces.append((pos, m.start()))
pos = m.end()
pieces.append((pos, len(run)))
for p0, p1 in pieces:
raw = run[p0:p1]
piece = raw.strip()
if _LETTER.search(piece) and "<" not in piece:
start = at + p0 + (len(raw) - len(raw.lstrip()))
spans.append(Span(start, start + len(piece), 0, []))
segments.append(piece)
contexts.append(ctx)
tokens = _MD.parse(text)
for t in tokens:
if t.type == "blockquote_open":
blockquote_fresh += 1
elif t.type == "blockquote_close":
blockquote_fresh -= 1
elif t.type == "inline":
kids = t.children or []
# An alert marker ([!NOTE]) leading a blockquote's first
# paragraph is syntax; both paths strip it. (Only the first
# inline of the blockquote can carry it — the flag clears on
# the first inline seen.)
alert = bool(blockquote_fresh)
blockquote_fresh = 0
linked = _linked_block(text, kids, cursor, strip_alert=alert)
if linked is not None:
span, wire = linked
spans.append(span)
segments.append(wire)
contexts.append("")
cursor = span.end
continue
runs = _runs(kids)
block = _block_text(kids).strip()
if alert and runs:
run = _ALERT.sub("", runs[0], count=1)
if _LETTER.search(run):
runs[0] = run
else:
runs.pop(0)
for run in runs:
ctx = block if block and run.strip() != block else ""
pos = _locate(text, run, cursor)
if pos != -1:
emit(run, pos, ctx)
cursor = pos + len(run)
elif "\n" in run:
# Indented continuation lines etc. break the verbatim
# match: locate each line separately instead.
for part in run.split("\n"):
if not _LETTER.search(part):
continue
pos = _locate(text, part, cursor)
if pos != -1:
emit(part, pos, ctx)
cursor = pos + len(part)
return spans, segments, contexts
def pure_prose(text: str) -> bool:
"""True when the text parses as nothing but prose (text and softbreak
tokens) the acceptance test for a translated segment: the model may
not return markup of its own (a `<br>` here would splice live HTML into
the fragment)."""
children = _MD.parseInline(text)[0].children or []
return all(t.type in ("text", "softbreak") for t in children)
def _word_sim(a: str, b: str) -> float:
"""How likely two words are the same term across a translation, 0..1.
A case-folded exact match is 1; otherwise the better of the sequence
ratio and the shared-prefix ratio inflection and derivational change
mostly move the ending ("banana" -> "banaanilla") or drop an article or
preposition around it. Case-folded so capitalization differences across
languages don't hide a term, with a small bonus when BOTH sides are
capitalized: a mid-sentence capital on both sides is likely the same
name (capitalization conventions differ per language, so its absence
proves nothing).
"""
bonus = 0.1 if a[:1].isupper() and b[:1].isupper() else 0.0
a, b = a.casefold(), b.casefold()
if a == b:
return 1.0
prefix = 0
for ca, cb in zip(a, b):
if ca != cb:
break
prefix += 1
sim = max(
difflib.SequenceMatcher(None, a, b).ratio(),
prefix / max(len(a), len(b)),
)
return min(1.0, sim + bonus)
#: Alignment costs for _find_mark: skipping a translation word (an article
#: or preposition the target language added) is cheap, skipping a source
#: word (one the translation dropped) costs more — a mark whose words
#: mostly vanished is no match at all. Every matched pair pays _MATCH, so
#: aligning a word to a lookalike-nothing (similarity below _MATCH) is
#: worse than skipping it.
_GAP_T = 0.25
_GAP_S = 0.6
_MATCH = 0.3
def _find_mark(src: list[str], units: list[re.Match], start: int) -> tuple[int, int] | None:
"""Locate a mark's source words in the translation's units (from unit
index ``start`` on), as the (start, end) unit-index span of the best
fuzzy alignment; None when no alignment is convincing (the caller falls
back to the weight ratio).
Word-for-word alignment with skips (_word_sim per pair, _GAP_T/_GAP_S
per skipped word): reordering is handled by the search itself, an added
or dropped article/preposition by the skip penalties. Accepted only
with an anchor one pair of similarity >= 0.7 and a decent average,
so a fully reworded label doesn't snap onto chance lookalikes.
"""
tgt = [u.group() for u in units[start:]]
n, m = len(src), len(tgt)
if not n or not m:
return None
# dp[i][j]: best score aligning src[:i] to tgt[:j]; a free tail (the
# answer is the best dp[n][j] over j) keeps trailing words costless.
dp = [[0.0] * (m + 1) for _ in range(n + 1)]
back: list[list[tuple[int, int]]] = [[(0, 0)] * (m + 1) for _ in range(n + 1)]
for i in range(1, n + 1):
dp[i][0] = dp[i - 1][0] - _GAP_S
back[i][0] = (i - 1, 0)
for j in range(1, m + 1):
options = [
(dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH, (i - 1, j - 1)),
(dp[i][j - 1] - _GAP_T, (i, j - 1)),
(dp[i - 1][j] - _GAP_S, (i - 1, j)),
]
dp[i][j], back[i][j] = max(options, key=lambda o: o[0])
j_end = max(range(m + 1), key=lambda j: dp[n][j])
pairs: list[tuple[int, int]] = [] # matched (source, target) indices
i, j = n, j_end
while i > 0:
pi, pj = back[i][j]
if (pi, pj) == (i - 1, j - 1):
pairs.append((i - 1, j - 1))
i, j = pi, pj
if not pairs:
return None
pairs.reverse() # backtracking collected them last-first
sims = [_word_sim(src[a], tgt[t]) for a, t in pairs]
# Weak pairs at the span's ends are not part of the label (a declined
# neighbor the DP matched for a pittance) — trim them off.
while len(sims) > 1 and sims[0] < 0.5:
pairs.pop(0)
sims.pop(0)
while len(sims) > 1 and sims[-1] < 0.5:
pairs.pop()
sims.pop()
if max(sims) < 0.7 or sum(sims) / len(sims) < 0.45:
return None
return start + pairs[0][1], start + pairs[-1][1] + 1
def _place_marks(translation: str, weight: int, marks: list[Mark]) -> str | None:
"""Re-insert a whole-block segment's links into its translation.
Each mark's boundaries are found by fuzzy word-form alignment
(_find_mark): the mark's source words are matched against the
translation's units by form similarity — no markers on the wire
(sentinels never survived the model), no assumption that word order or
count survived either. Slicing exactly at unit boundaries keeps the
whitespace between the mark and its neighbors in the plain text, where
it belongs. A mark with no convincing alignment falls back to its
source weight ratio (units before the boundary / total applied to the
translation's units) — the pre-fuzz heuristic, still the CJK path,
where form similarity across scripts is nil. A boundary landing empty
degrades to the source link text: better an untranslated label than a
broken "[](url)". None when the translation has no units to map onto
(the caller rejects the result).
"""
units = list(_UNIT.finditer(translation))
total = len(units)
if not total or not weight:
return None
starts = [u.start() for u in units]
bounds = starts + [len(translation)]
out: list[str] = []
cur = 0 # char cursor: never before the previous mark's end
ucur = 0 # unit cursor, the same monotonicity in unit indices
for mark in marks:
found = _find_mark(_UNIT.findall(mark.inner), units, ucur)
if found is not None:
u1, u2 = found
x1, x2 = units[u1].start(), units[u2 - 1].end()
else:
x1 = bounds[min(round(mark.w_start / weight * total), total)]
x2 = bounds[min(round(mark.w_end / weight * total), total)]
x1 = max(x1, cur)
x2 = max(x2, x1)
# The slice ends at the next unit's start, so the whitespace
# and punctuation before that unit is inside it — but it
# belongs BETWEEN the mark and the following word, not in the
# inner text: end the inner text at its last unit and leave
# the rest for the following slice.
raw = translation[x1:x2]
inner_units = list(_UNIT.finditer(raw))
x2 = x1 + inner_units[-1].end() if inner_units else x1
inner = translation[x1:x2].strip() or mark.inner
out += [translation[cur:x1], mark.pre, inner, mark.post]
cur = x2
ucur = bisect.bisect_left(starts, x2)
out.append(translation[cur:])
return "".join(out)
def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
"""Splice translated segments back into the original fragment; None on
any validation failure (count mismatch, empty or non-prose segment)
the caller drops the result and the fragment stays pending. Segments
with marks (a block that crossed as one piece) get their links
re-inserted at weight-mapped positions after the prose check."""
if len(texts) != len(spans):
return None
out: list[str] = []
cursor = 0
for span, translation in zip(spans, texts):
if not translation.strip() or not pure_prose(translation):
return None
if span.marks:
translation = _place_marks(translation, span.weight, span.marks)
if translation is None:
return None
out.append(original[cursor : span.start])
out.append(translation)
cursor = span.end
out.append(original[cursor:])
return "".join(out)
def has_prose(text: str) -> bool:
"""True when the fragment yields at least one translatable segment.
Chunks that are all markup, code, placeholders or reference definitions
have no business reaching the model: every language renders them from
the original chunk."""
return bool(split(text)[1])
+366
View File
@@ -0,0 +1,366 @@
"""Shared core: site constants, the kanta database, and the render cache.
Everything the route modules (files, api, tracking, pages) need that is not
a route itself: environment-derived paths and tunables, the ``Data`` root
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
store, the fastapi-vue ``Frontend``, the page render cache
(``_render_html``/``_cached_body``/``_html_response`` plus the
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
content/settings write), the translator ``dispatcher``, the slug charset
helpers, and the database bootstrap hooks (demo seed, translator defaults).
Importable by every other pagerite module without cycles.
"""
import logging
import os
import re
import secrets
from datetime import UTC, datetime
from functools import lru_cache
from pathlib import Path
import blake3
from fastapi import HTTPException, Request
from fastapi.responses import Response
from fastapi_vue import Frontend
from kanta import Kanta
from zstandard import ZstdCompressor
from pagerite import analytics, i18n, seed, translate, views
from pagerite.__main__ import DEVMODE
from pagerite.chunks import store_chunks
from pagerite.data import (
Data,
Node,
append_order,
find_slot,
prettify,
)
logger = logging.getLogger(__name__)
# Site identity: the hostname comes from the CLI (first positional argument,
# exported as PAGERITE_HOSTNAME) and names the per-site data directory
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
HOSTNAME = os.getenv("PAGERITE_HOSTNAME", "localhost")
SITE_DIR = Path(HOSTNAME)
#: Public origin of the site, used for absolute social/canonical/sitemap
#: URLs. Localhost serves varying ports, so it falls back to the request's
#: own base URL instead.
SITE_URL = f"https://{HOSTNAME}" if HOSTNAME != "localhost" else ""
DB_PATH = os.getenv("PAGERITE_DB", str(SITE_DIR / "content.kantadb"))
# Visit analytics go to their own JSON file, not the kanta database.
ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json")))
analytics_store = analytics.Store(ANALYTICS_PATH)
# Content-addressed file store (uploads, seed assets, fetched favicons):
# files on disk under hash-prefixed names, cached in RAM, served at /_f/.
FILES_DIR = Path(os.getenv("PAGERITE_FILES", str(SITE_DIR / "files")))
# Uploaded images are thumbnailed to this size and recompressed to AVIF
# (primary), with WebP and JPEG fallbacks re-encoded from the AVIF at
# somewhat lower quality (similar or smaller file size); the untouched
# original is kept alongside as ``<hash>.orig<ext>`` (never served).
IMAGE_MAXSIZE = 1920
IMAGE_QUALITY = 60
IMAGE_WEBP_QUALITY = 50
IMAGE_JPG_QUALITY = 55
# Favicons get the same derivatives but thumbnailed much smaller — 192px
# is plenty (browsers scale down for the 16x16 tab icon themselves).
FAVICON_MAXSIZE = 192
# Our own data root; kanta edits it in place, reads are plain attribute access.
data = Data()
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
# Vue build served at the site root, no SPA catch-all (assets only). The
# build mirrors the URL space: hashed, immutable files live under
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico.
BUILD_DIR = Path(__file__).with_name("frontend-build")
frontend = Frontend(BUILD_DIR, spa=False, cached="/_assets/")
# Dynamic HTML is compressed per request at level 9 (static assets are
# already pre-compressed by fastapi-vue's Frontend).
_zstd = ZstdCompressor(9)
def _render_html(
kind: str,
path: str,
base_url: str,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> str:
"""Render one of the generated pages (see _html_response)."""
if kind == "page":
# A selected language without an actual translation renders the
# original (translation is None; see docs/localization.md).
original = i18n.primary_lang(data.menu, path)
translation = (
i18n.get_translation(data, path, lang) if lang != original else None
)
return views.render_page(
data.menu,
data,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
base_url,
transition=data.transition,
lang=lang,
translation=translation,
link_lang=link_lang,
)
if kind == "category":
# A category has no Markdown of its own; only the title map
# localizes (heading, navigation, card text).
original = i18n.primary_lang(data.menu, path)
translation = (
i18n.Translation(titles=i18n.title_map(data, lang))
if lang != original
else None
)
return views.render_category(
data.menu,
data,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
transition=data.transition,
lang=lang,
translation=translation,
link_lang=link_lang,
)
if kind == "not-found":
return views.render_not_found(
data.menu,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
transition=data.transition,
)
return views.render_analytics(
data.menu,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
transition=data.transition,
)
# Render generation: bumped (and the body cache cleared) by every
# content/settings write, so page ETags and cached copies invalidate when
# navigation-affecting changes happen. In-memory only — not database state.
_render_gen = 0
def _invalidate_pages() -> None:
"""Drop cached page bodies and bump the render generation (ETags);
any content change also re-runs translation dispatch."""
global _render_gen
_render_gen += 1
_cached_body.cache_clear()
dispatcher.schedule()
@lru_cache(maxsize=128)
def _cached_body(
kind: str,
path: str,
base_url: str,
zstd: bool,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> bytes:
"""Rendered page body; cleared by _invalidate_pages on any
content/settings change. base_url feeds the social meta URLs, zstd
selects the stored encoding (both variants are cached rather than
re-compressed) and lang the selected language (not the raw
Accept-Language header, which would blow up the cache key space).
link_lang is the ?lang= override replicated onto the navigation links:
a query render and a header-selected render of the same language differ
in their links, so they are cached separately.
"""
body = _render_html(kind, path, base_url, lang, link_lang).encode()
return _zstd.compress(body) if zstd else body
def _html_response(
request: Request,
kind: str,
path: str,
status_code: int = 200,
headers: dict | None = None,
etag: bool = False,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> Response:
"""Response for a generated page, zstd-compressed when the client
accepts it (no gzip fallback).
Done per handler rather than in middleware so that Frontend's
already-compressed asset responses are never touched. The ETag stays
identical across encodings (revalidation compares it before
compression); ``vary: accept-encoding`` keeps caches from mixing the
representations. In dev the cache is bypassed so theme/design edits on
disk apply immediately.
``etag=True`` derives the validator from a blake3 hash of the
(uncompressed) body for pages like /_a that have no Node whose
modified timestamp could serve as one and answers matching
if-none-match revalidations with a 304.
"""
zstd = "zstd" in request.headers.get("accept-encoding", "")
# Absolute social/canonical URLs use the site's public origin; on
# localhost (varying ports) fall back to the request's own base URL.
base_url = SITE_URL or str(request.base_url).rstrip("/")
if DEVMODE:
identity = _render_html(kind, path, base_url, lang, link_lang).encode()
body = _zstd.compress(identity) if zstd else identity
else:
identity = _cached_body(kind, path, base_url, False, lang, link_lang)
body = (
_cached_body(kind, path, base_url, True, lang, link_lang)
if zstd
else identity
)
h = dict(headers or {})
# Content varies by language (Accept-Language selects a translation)
# and by encoding; keep caches from mixing either representation.
h["vary"] = "accept-language" + (", accept-encoding" if zstd else "")
if etag:
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
h["etag"] = tag
if request.headers.get("if-none-match") == tag:
return Response(status_code=304, headers=h)
if zstd:
h["content-encoding"] = "zstd"
return Response(body, status_code, h, media_type="text/html")
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
def _is_reserved(path: str) -> bool:
"""Slug shape that content may never use: each segment must be lower-case
ASCII letters, digits, hyphens and underscores (underscores may not be
the first character), and dots are never allowed.
"""
if path == "":
return False
return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
def _check_reserved(path: str) -> None:
"""Reject paths that do not follow the slug charset."""
if _is_reserved(path):
raise HTTPException(
400,
'slugs may only use a-z, 0-9, "-" and "_" (not as the first character), and no dots',
)
def _ensure(menu: dict[str, Node], path: str) -> Node:
"""Return the node at ``path``, creating it and any missing ancestors
(content-less category labels) appended at the end of their level."""
nodes = menu
node = None
for seg in path.split("/"):
node = nodes.get(seg)
if node is None:
node = Node(title=prettify(seg), order=append_order(nodes))
nodes[seg] = node
nodes = node.children
return node
def _remove_page(menu: dict[str, Node], path: str) -> bool:
"""Delete the node at ``path`` (inside a transaction).
A node with children becomes a content-less category label; a childless
node is removed entirely. Returns False if the path does not exist.
"""
slot = find_slot(menu, path)
node = slot[0].get(slot[1]) if slot else None
if node is None:
return False
if node.children:
node.chunks = None
node.modified = datetime.now(UTC)
else:
del slot[0][slot[1]]
return True
def _store_seed_file(
markdown: str, banner: str, orig: str, body: bytes
) -> tuple[str, str]:
"""Store a seed file content-addressed and point references at /_f/.
Images get the same AVIF/WebP/JPEG derivatives as uploads and are
linked extension-less; other content is stored as-is with its
extension."""
from pagerite.files import _ext, store_image # lazy: files imports state
ext = _ext(orig)
name = store_image(body, ext, derive=ext != ".gif")
markdown = markdown.replace(f"]({orig}", f"](/_f/{name}")
banner = banner.replace(f'src="/{orig}"', f'src="/_f/{name}"')
banner = banner.replace(f'src="{orig}"', f'src="/_f/{name}"')
return markdown, banner
@kanta.bootstrap
def _seed(data: Data) -> None:
"""Write the demo pages on database creation (never on existing dbs)."""
for path in seed.PAGES:
title, markdown, files, banner, order, design = seed.PAGES[path]
for orig, body in files.items():
markdown, banner = _store_seed_file(markdown, banner, orig, body)
node = _ensure(data.menu, path)
node.title = title
# Empty markdown means a pure category label (e.g. "showcase",
# seeded only to carry a banner design): leave chunks as None so
# the node renders the placeholder and nav points at its children.
if markdown:
node.chunks = store_chunks(data.chunks, markdown)
node.banner = banner
node.banner_design = design
node.order = order
#: Translator key format: 12 lowercase alphanumeric characters — not
#: brute-forceable over a WebSocket handshake, still human-manageable.
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
@kanta.bootstrap
def _translator_defaults(data: Data) -> None:
"""Translator defaults on database creation: the first service key and
the wanted target languages (Spanish and Chinese English is the
original language, never a translation target).
Keys are a dict (key -> display name) with the future reservation that
multiple keys could be managed (e.g. via a web interface)."""
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
data.translate_keys[key] = "default"
data.translate_langs = {"es": True, "zh": True}
# The translator dispatcher — protocol, connected clients and the job
# pipeline live in translate.py; its WebSocket route is in api.py.
dispatcher = translate.Dispatcher(data, kanta, _invalidate_pages)
+1 -1
View File
@@ -112,7 +112,7 @@ article h3 {
blockquote { blockquote {
border-left-color: var(--accent); border-left-color: var(--accent);
background: color-mix(in oklab, var(--accent) 6%, transparent); background: color-mix(var(--accent) 6%, transparent);
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
+4 -5
View File
@@ -42,6 +42,9 @@
--font-body: var(--font-montserrat); --font-body: var(--font-montserrat);
--font-heading: var(--font-literata); --font-heading: var(--font-literata);
--code-x-height: 0.517; /* Montserrat's x-height ratio */ --code-x-height: 0.517; /* Montserrat's x-height ratio */
/* Neutral grey selection instead of the accent tint: accent-colored
text (h2, links, markers) stays readable on it in both schemes. */
--selection-bg: #6664;
} }
/* Dark scheme: same identity, but the page goes deep violet (never muddy /* Dark scheme: same identity, but the page goes deep violet (never muddy
@@ -62,10 +65,6 @@
} }
} }
::selection {
background: var(--accent);
}
/* Any banner used is separated from page by a thick orange line */ /* Any banner used is separated from page by a thick orange line */
#banner { #banner {
border-bottom: 4px solid var(--accent); border-bottom: 4px solid var(--accent);
@@ -188,7 +187,7 @@ article ul ul ul li::before {
blockquote { blockquote {
border-left-color: var(--accent2); border-left-color: var(--accent2);
background: color-mix(in oklab, var(--accent2) 6%, transparent); background: color-mix(var(--accent2) 6%, transparent);
padding: 0.25rem 0.75rem; padding: 0.25rem 0.75rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
+1 -1
View File
@@ -24,7 +24,7 @@
--sun: #ffe071; --sun: #ffe071;
--line: #4f913b29; --line: #4f913b29;
--link: color-mix(in oklab, var(--text) 38%, var(--accent)); --link: color-mix(var(--text) 38%, var(--accent));
--font-body: var(--font-cause); --font-body: var(--font-cause);
--font-heading: var(--font-new-rocker); --font-heading: var(--font-new-rocker);
+410
View File
@@ -0,0 +1,410 @@
"""Visit analytics: collection sockets, geoip enrichment, favicon fetch.
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
enriched in background tasks with reverse DNS (cached PTR lookups) and the
DB-IP city MMDB (``GeoIP``, decompressed and opened once at startup);
external referrers get their favicon fetched and stored content-hashed.
Snapshot broadcasts to connected admin sockets are debounced.
"""
import asyncio
import gzip
import ipaddress
import logging
import os
import re
import shutil
import socket
from functools import lru_cache
from pathlib import Path
from urllib.parse import urlparse
import httpx
import msgspec
from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
from fastapi.responses import Response
from pagerite import analytics
from pagerite.files import _hash_name, file_store
from pagerite.state import SITE_URL, _html_response, analytics_store
logger = logging.getLogger(__name__)
router = APIRouter()
# Live WebSocket clients for the analytics stream.
_analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None
# Repository root from this file's location (pagerite/tracking.py -> ..).
_REPO_ROOT = Path(__file__).resolve().parent.parent
def _geoip_db_path() -> Path | None:
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present.
"""
mmdb = sorted(_REPO_ROOT.glob("dbip-*.mmdb"))
if mmdb:
return mmdb[0]
gz = sorted(_REPO_ROOT.glob("dbip-*.mmdb.gz"))
if gz:
return gz[0]
return None
class GeoIP:
"""Lazy DB-IP MMDB reader. Call ``_load()`` once at startup before
concurrent requests arrive; ``country()`` is read-only and safe to call
from ``asyncio.to_thread`` workers afterwards.
"""
def __init__(self) -> None:
self._reader: object | None = None
def _decompress(self, source: Path, target: Path) -> None:
if target.exists():
return
tmp = target.with_suffix(target.suffix + ".tmp")
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
shutil.copyfileobj(src, dst)
os.replace(tmp, target)
def _load(self) -> None:
if self._reader is not None:
return
source = _geoip_db_path()
if source is None:
return
if source.suffix == ".gz":
target = source.with_suffix("")
self._decompress(source, target)
source = target
try:
import maxminddb
self._reader = maxminddb.open_database(str(source))
except Exception:
pass
def country(self, ip: str) -> str:
"""Two-letter ISO country code for ``ip``, or "" when unavailable."""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
return (rec.get("country") or {}).get("iso_code", "")
except Exception:
pass
return ""
def city(self, ip: str) -> str:
"""City name for ``ip``, or "" when unavailable.
GeoIP sometimes appends district names in parentheses (e.g.
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
the value is stored.
"""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
city = (rec.get("city") or {}).get("names", {}).get("en", "")
if city:
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
return city
except Exception:
pass
return ""
_geoip = GeoIP()
def _client_ip(request: Request | WebSocket) -> str:
"""Client IP: first X-Forwarded-For hop (we sit behind a proxy), else
the direct peer."""
forwarded = request.headers.get("x-forwarded-for", "").split(",")[0].strip()
return forwarded or (request.client.host if request.client else "")
def _query_suffix(request: Request) -> str:
"""The request's query string as a "?..." suffix, or "" when absent."""
query = str(request.url.query)
return f"?{query}" if query else ""
@lru_cache(maxsize=4096)
def _cached_ptr(ip: str) -> str:
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
if not ip:
return ""
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ""
if (
addr.is_private
or addr.is_loopback
or addr.is_reserved
or addr.is_multicast
or addr.is_link_local
):
return ""
try:
host, _, _ = socket.gethostbyaddr(ip)
except socket.herror:
return ""
return host
async def _lookup_host(ip: str) -> str:
"""Async wrapper around ``_cached_ptr``; runs the blocking lookup in a thread."""
return await asyncio.to_thread(_cached_ptr, ip)
async def _geoip_country(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB lookup."""
return await asyncio.to_thread(_geoip.country, ip)
async def _geoip_city(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB city lookup."""
return await asyncio.to_thread(_geoip.city, ip)
async def _enrich_client(client_hash: bytes) -> None:
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
client = analytics_store.data.clients.get(client_hash)
if not client or not client.ip:
return
host = await _lookup_host(client.ip)
country = await _geoip_country(client.ip)
city = await _geoip_city(client.ip)
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
"""Start background host/geoip enrichment for the given client hashes."""
for client_hash in client_hashes:
asyncio.create_task(_enrich_client(client_hash))
#: Icon MIME -> file extension for the stored favicon name. The extension
#: reflects the actual content, not the /favicon.ico request path.
_FAVICON_EXT = {
"image/x-icon": ".ico",
"image/vnd.microsoft.icon": ".ico",
"image/png": ".png",
"image/gif": ".gif",
"image/jpeg": ".jpg",
"image/webp": ".webp",
"image/avif": ".avif",
"image/svg+xml": ".svg",
}
_FAVICON_MAX_BYTES = 65536
#: Origins with a fetch task currently in flight.
_favicon_in_flight: set[str] = set()
async def _fetch_favicon(origin: str) -> None:
"""Fetch ``{origin}/favicon.ico`` and store it content-hashed on disk.
The result (icon file name, or "" for a miss) is recorded in the
analytics store; misses are retried after analytics._FAVICON_RETRY.
Never raises: analytics must not break page serving.
"""
try:
async with httpx.AsyncClient(follow_redirects=True, timeout=8) as client:
r = await client.get(f"{origin}/favicon.ico")
body = r.content
if (
not (200 <= r.status_code < 300)
or not body
or len(body) > _FAVICON_MAX_BYTES
):
analytics_store.record_favicon(origin)
return
mime = r.headers.get("content-type", "").split(";")[0].strip().lower()
if not mime.startswith("image/"):
# Served without an image type: sniff SVG, else assume ICO.
if b"<svg" in body[:1024]:
mime = "image/svg+xml"
elif mime in ("", "application/octet-stream", "text/plain"):
mime = "image/x-icon"
else:
analytics_store.record_favicon(origin)
return
ext = _FAVICON_EXT.get(mime, ".ico")
name = _hash_name(body, f"favicon{ext}")
file_store.put(name, body)
analytics_store.record_favicon(origin, name)
except httpx.HTTPError, OSError:
analytics_store.record_favicon(origin)
finally:
_favicon_in_flight.discard(origin)
def _schedule_favicon_fetch() -> None:
"""Start background favicon fetches for origins that need one."""
for origin in analytics_store.favicon_origins_needed():
if origin in _favicon_in_flight:
continue
_favicon_in_flight.add(origin)
asyncio.create_task(_fetch_favicon(origin))
async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients:
return
payload = analytics_store.display_json()
closed = set()
for ws in _analytics_ws_clients:
try:
await ws.send_text(payload)
except Exception:
closed.add(ws)
for ws in closed:
_analytics_ws_clients.discard(ws)
async def _debounced_analytics_broadcast() -> None:
"""Wait briefly, then broadcast the latest snapshot once."""
await asyncio.sleep(0.2)
await _broadcast_analytics()
def _schedule_analytics_broadcast() -> None:
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
global _analytics_broadcast_task
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
return
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
_debounced_analytics_broadcast()
)
def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]:
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
Nothing is counted on the GET itself the client's first /_ws message
starts the visit, so bots never register as visits (JS-running crawlers
connect too, but the WebSocket handler ignores known bot UAs). (Admin
clients report too, but with hide, which flags their visit hidden: it is
recorded but excluded from all statistics and from the crawler list.)
The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
logged as a crawler hit. The root-path and localhost checks prevent
remote visitors from hiding traffic with the same query string.
Returns the client hashes of any pending crawler hits flushed to persistent
storage, so callers can schedule async geoip and reverse-DNS enrichment.
"""
if request.headers.get("x-pagerite-preload"):
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
# activity message sent when the user actually navigates does the
# counting.
# (Forging the header only hides a GET from the crawler stats; the
# path-based abuse classification is unaffected.)
return []
if (
path == ""
and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1"
):
return []
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
full_path = f"{request.url.path}{_query_suffix(request)}"
return analytics_store.track_entry(
request.headers.get("referer", ""),
own_origin,
_client_ip(request),
request.headers.get("user-agent", ""),
full_path,
request.headers.get("accept-language", ""),
status=status,
)
@router.get("/_a", response_model=None)
async def analytics_page(request: Request) -> Response:
"""Render the analytics viewer as a normal site page at /_a.
The page itself is public, but the data stream (/_api/ws/analytics) stays
admin-gated like the rest of /_api, so only authorized users see the
statistics; others get the viewer with a "could not be loaded" message.
"""
return _html_response(
request,
"analytics",
"",
headers={"cache-control": "no-cache"},
etag=True,
)
@router.websocket("/_ws")
async def activity_ws(ws: WebSocket) -> None:
"""Collect visitor activity: navigations and reading-time updates.
Public, like the pages themselves (only /_api is gated); one connection
follows a browsing session. Messages are ``analytics.Ping`` structs as
JSON text frames; ``to`` set is a navigation, ``read`` alone a
reading-time update. The reverse-DNS and DB-IP geoip lookups happen in
background tasks so message handling is never delayed by slow DNS or
the first MMDB decompress.
"""
await ws.accept()
ip = _client_ip(ws)
ua = ws.headers.get("user-agent", "")
accept_language = ws.headers.get("accept-language", "")
try:
while True:
text = await ws.receive_text()
try:
msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
except msgspec.DecodeError:
continue
visit_index, flushed_clients = analytics_store.ping(
msg.fr,
msg.to or None,
ip,
ua,
accept_language,
hide=msg.hide,
read=msg.read,
)
if visit_index is not None:
visit = analytics_store.data.visits[visit_index]
asyncio.create_task(_enrich_client(visit.client))
_schedule_client_enrichment(flushed_clients)
_schedule_favicon_fetch()
except WebSocketDisconnect:
pass
@router.websocket("/_api/ws/analytics")
async def analytics_websocket(ws: WebSocket) -> None:
"""Stream the analytics snapshot, then push updates as they happen.
Admin-only via the /_api forward-auth gate, like every management
endpoint. Powers the analytics viewer rendered at /_a.
"""
await ws.accept()
await ws.send_text(analytics_store.display_json())
_analytics_ws_clients.add(ws)
try:
while True:
await ws.receive_text()
except Exception:
pass
finally:
_analytics_ws_clients.discard(ws)
+408
View File
@@ -0,0 +1,408 @@
"""Translator service protocol, dispatcher and its transport-independent core.
The external machine-translation service connects over WebSocket
(``/_translate/<key>``, the route itself is in api.py) and exchanges JSON
frames decoded into the tagged msgspec structs below (``bytes`` fields ride
as base64 no manual encoding anywhere). This module holds everything
else: the message structs, the connected-client dispatcher (``Dispatcher``
one job at a time per connection, wanted capable language matching,
requeue on disconnect), which fragments are pending for a language
(``pending_items``) and storing a result (``store_results``).
Fragments cross the wire as **prose segments**: the model only ever
receives plain text runs (Job.texts) plus per-segment context surrounds
(Job.contexts) and returns their translations (Result.texts, same order);
markup never leaves the server reassembly is offset splicing
(``pagerite/segments.py``).
"""
import asyncio
import logging
import msgspec
from fastapi import WebSocket, WebSocketDisconnect
from kanta import Kanta
from pagerite import i18n
from pagerite.chunks import chunk_key, needs_translation
from pagerite.data import Data, Node, sorted_nodes
from pagerite.segments import Span, join, split
logger = logging.getLogger(__name__)
class Hello(msgspec.Struct, tag="hello"):
"""Client greeting on connect: the language codes its model CAN produce
(capabilities). The server offers jobs only in the intersection with
the wanted target languages (``Data.translate_langs``)."""
langs: list[str]
class TransItem(msgspec.Struct):
"""One fragment to translate: original Markdown (or a node title)."""
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
text: str
path: str #: article it came from ("" = front page), no leading slash
kind: str #: "chunk" | "title"
#: Title jobs only: the article's opening prose, so the model sees the
#: title as a heading in context, not a lone sentence.
context: str = ""
class Job(msgspec.Struct, tag="job"):
"""Server push: ONE fragment to translate.
Exactly one job is in flight per connection the next is sent only
after this one's Result. Clients wanting parallelism open multiple
connections."""
lang: str
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
#: The fragment's prose segments (pagerite/segments.py): plain text
#: runs only — no markup, URLs, code or placeholders ever cross the
#: wire. Translate each element independently.
texts: list[str]
path: str #: article it came from ("" = front page), no leading slash
kind: str #: "chunk" | "title"
#: Per segment (parallel to texts; "" = none): the surround to
#: translate it in — a carved-out segment (link text, partial run)
#: carries its block's plain text, a title the article's opening.
#: Reference client behavior (scripts/translator.py): translate
#: segment+context together, keep the segment's part (its own line /
#: paragraph); fall back to the segment alone when the output holds no
#: separator. Contexts are not part of the result.
contexts: list[str] = msgspec.field(default_factory=list)
class TransResult(msgspec.Struct):
"""One translated fragment (storage level, see store_results)."""
key: bytes
text: str
class Result(msgspec.Struct, tag="result"):
"""Client reply: the translation of the connection's current Job
(must match its lang and key exactly)."""
lang: str
key: bytes
#: The job's segments, translated, same order and count. Each must be
#: pure prose — the server rejects the result otherwise.
texts: list[str]
#: Union of the client -> server frames (the "type" tag selects).
ClientMsg = Hello | Result
def pending_items(data: Data, lang: str) -> list[TransItem]:
"""Fragments of the site still untranslated for ``lang``, deduped by key.
Every page node (published or not) contributes its title and each chunk
that needs translation (``needs_translation``), is not editor-flagged
no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang``
yet. Content-addressed text (shared paragraphs, repeated titles) appears
once, under the first page in menu order that has it.
"""
items: list[TransItem] = []
seen: set[bytes] = set()
def emit(key: bytes, text: str, path: str, kind: str, context: str = "") -> None:
if key in seen or lang in data.trans.get(key, {}):
return
seen.add(key)
items.append(
TransItem(key=key, text=text, path=path, kind=kind, context=context)
)
def opening(node: Node) -> str:
"""The article's opening prose (first segment, capped): the title
job's context — a lone word like "About" reads as a heading on top
of an article, not as a sentence. Empty when there's no prose."""
for h in node.chunks or ():
text = data.chunks.get(h)
if text and (segs := split(text)[1]):
return segs[0][:400]
return ""
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
# An article whose primary language IS the target needs no
# translation into it — skip its title and chunks entirely.
node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang:
if node.title:
emit(
chunk_key(node.title),
node.title,
path,
"title",
context=opening(node),
)
for h in node.chunks:
text = data.chunks.get(h)
if (
text is not None
and h not in node.no_trans
and needs_translation(text)
):
emit(h, text, path, "chunk")
walk(node.children, path, node_lang)
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
return items
def store_results(data: Data, lang: str, items: list[TransResult]) -> list[str]:
"""Store machine translations for ``lang``; return the paths of the
articles that gained at least one entry.
Pure data operations: the caller wraps this in a kanta transaction and
invalidates pages. Unknown keys are stored anyway (unreferenced hashes
are never read, and the content may simply have moved on since the job
was pushed); re-storing an existing key overwrites, last wins. Every
article that gained an entry gets ``node.langs[lang]`` set (the
availability index, docs/migrate.md) because chunks are
content-addressed, that includes pages merely sharing a fragment.
"""
stored = {item.key for item in items}
for item in items:
data.trans.setdefault(item.key, {})[lang] = item.text
pages: list[str] = []
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang:
keys = set(node.chunks)
if node.title:
keys.add(chunk_key(node.title))
if keys & stored:
node.langs[lang] = True
pages.append(path)
walk(node.children, path, node_lang)
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
return pages
class _Connection:
"""One connected translator socket: the language codes it announced as
capabilities (Hello) and the (lang, chunk-key) job currently in flight
on it, with the segment spans to splice its Result into
(pagerite/segments.py) one at a time, the next is sent only after its
Result.
Per-connection only: in-flight lives solely here, so on disconnect the
item simply becomes pending again and is re-offered to any free capable
connection."""
def __init__(self, capable: set[str]) -> None:
self.capable = capable
self.inflight: tuple[str, bytes] | None = None
#: Source spans of the in-flight job's segments (splice offsets
#: and link marks).
self.spans: list[Span] = []
self.original: str = "" # its full source text (for the splicing)
self.kind: str = "" # "chunk" | "title" (for the transaction action)
class Dispatcher:
"""The translator dispatcher: connected client sockets and the job
pipeline (docs/localization.md).
One single-item job at a time per connection, offered in the
intersection of the wanted languages (``Data.translate_langs``) and the
connection's announced capabilities. Pending work is derived from the
``trans`` store (``pending_items``) minus the items in flight on any
connection, so a dropped connection's in-flight item is simply
re-offered. Results are matched to content by chunk key alone. A
(lang, key) whose Result fails segment validation is skipped for the
rest of the run generation is near-deterministic, so an immediate
retry would just re-fail.
"""
def __init__(self, data: Data, db: Kanta, invalidate) -> None:
self.data = data
self.db = db
#: Sync content-change hook (state._invalidate_pages), called inside
#: transactions; schedules the next dispatch pass.
self.invalidate = invalidate
#: Connected translator sockets and their per-connection state.
self.clients: dict[WebSocket, _Connection] = {}
#: (lang, chunk key) of fragments whose result failed validation
#: (segment count, empty or non-prose segments, segments.py) this run.
self.validation_failures: set[tuple[str, bytes]] = set()
def reset_validation_failures(self) -> None:
"""Clear the skip list of fragments rejected this run (segment
validation): a translations refresh is precisely the "another
chance" for them."""
self.validation_failures.clear()
def schedule(self) -> None:
"""Schedule a dispatch pass, if any translator is connected.
The invalidate hook is sync and called inside transactions: the
task first runs once the current coroutine awaits again, i.e. after
the transaction has committed. No-op without a running loop (CLI
use)."""
if not self.clients:
return
try:
asyncio.get_running_loop()
except RuntimeError:
return
asyncio.create_task(self._dispatch())
async def _dispatch(self) -> None:
"""Offer one pending item to every free capable connection."""
wanted = {
tag for lang in self.data.translate_langs if (tag := i18n.base_tag(lang))
}
if not wanted:
return
for ws, state in list(self.clients.items()):
if state.inflight is not None:
continue
langs = wanted & state.capable
if not langs:
continue
inflight = {s.inflight for s in self.clients.values() if s.inflight}
job = None
spans: list[Span] = []
original = ""
for lang in sorted(langs):
# Titles first: a page's name in the menu is its most
# visible string (stable: menu order kept within each kind).
for item in sorted(
pending_items(self.data, lang), key=lambda it: it.kind != "title"
):
if (lang, item.key) in inflight or (
lang,
item.key,
) in self.validation_failures:
continue
spans, texts, contexts = split(item.text)
if not texts:
continue # prose that could not be located for splicing
original = item.text
if item.kind == "title" and item.context:
# A title's surround is the article's opening prose
# (TransItem.context), not its own one-word block.
contexts = [item.context] * len(texts)
job = Job(
lang=lang,
key=item.key,
texts=texts,
path=item.path,
kind=item.kind,
contexts=contexts,
)
break
if job is not None:
break
if job is None:
continue
state.inflight = (job.lang, job.key) # before the await: no double-assign
state.spans = spans
state.original = original
state.kind = job.kind
try:
await ws.send_text(msgspec.json.encode(job).decode())
except Exception: # send failed: the receive loop cleans up
self.clients.pop(ws, None)
async def handle_ws(self, ws: WebSocket, clientkey: str) -> None:
"""The /_translate/<key> channel (docs/localization.md).
A wrong/empty key rejects the handshake (closing before accept
makes Starlette answer HTTP 403). Protocol (JSON frames): the
client opens with Hello(langs) announcing its CAPABILITIES the
language codes its model can produce (normalized to translation
tags; "en"/empty dropped) then answers each Job with its
Result(lang, key, texts). A Result without an in-flight job or with
a different (lang, key), a duplicate Hello, or any malformed frame
closes the socket with a protocol error.
"""
if clientkey not in self.data.translate_keys:
await ws.close(code=1008) # policy violation; pre-accept = HTTP 403
return
await ws.accept()
state: _Connection | None = None
try:
while True:
raw = await ws.receive_text()
try:
msg = msgspec.json.decode(raw.encode(), type=ClientMsg)
except msgspec.DecodeError:
await ws.close(code=1002) # protocol error
return
if isinstance(msg, Hello):
if state is not None: # one Hello per connection
await ws.close(code=1002)
return
state = _Connection(
{tag for lang in msg.langs if (tag := i18n.base_tag(lang))}
)
self.clients[ws] = state
self.schedule()
else: # Result
lang = i18n.base_tag(msg.lang)
if (
state is None # results before Hello
or state.inflight is None # no job in flight
or (lang, msg.key) != state.inflight # wrong job
):
await ws.close(code=1002)
return
texts, spans, original = msg.texts, state.spans, state.original
kind, state.kind = state.kind, ""
state.inflight = None
state.spans = []
state.original = ""
text = (
join(original, spans, texts)
if len(texts) == len(spans)
else None
)
if text is None:
# The model broke the segment contract (count
# mismatch, empty or non-prose segment): drop the
# result and skip the fragment for this run (it
# stays pending; a restart, a refresh or a model
# change gets another chance).
self.validation_failures.add((lang, msg.key))
logger.warning(
"[%s] result for chunk %s rejected: invalid segments",
lang,
msg.key.hex(),
)
self.schedule()
continue
with self.db.transaction(
f"translate:{lang}{':title' if kind == 'title' else ''}",
user=clientkey,
):
paths = store_results(
self.data, lang, [TransResult(key=msg.key, text=text)]
)
self.invalidate() # schedules the next dispatch
if paths:
logger.info(
"[%s] now available for %d page(s): %s",
lang,
len(paths),
", ".join(sorted(paths)),
)
except WebSocketDisconnect:
pass
finally:
if self.clients.pop(ws, None) is not None:
# The in-flight item (if any) is pending again; offer it around.
self.schedule()
+272 -60
View File
@@ -24,7 +24,9 @@ import re
from html5tagger import HTML, Document, E, Template from html5tagger import HTML, Document, E, Template
from platformdirs import site_data_dir, user_data_path from platformdirs import site_data_dir, user_data_path
from pagerite.data import Node, prettify, resolve, sorted_nodes from pagerite import i18n
from pagerite.data import Data, Node, node_markdown, prettify, resolve, sorted_nodes
from pagerite.i18n import Translation
from pagerite.markdown import render from pagerite.markdown import render
SITE_NAME = "Pagerite" SITE_NAME = "Pagerite"
@@ -37,7 +39,9 @@ def _data_roots() -> list[Path]:
roots = [user_data_path("pagerite", appauthor=False)] roots = [user_data_path("pagerite", appauthor=False)]
# site_data_dir keeps the multipath (site_data_path collapses it to the # site_data_dir keeps the multipath (site_data_path collapses it to the
# first entry, since a Path cannot hold several). # first entry, since a Path cannot hold several).
roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(os.pathsep) roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(
os.pathsep
)
return [Path(r) for r in roots] return [Path(r) for r in roots]
@@ -102,13 +106,15 @@ def _theme_color_schemes(theme: str) -> set[str]:
return set() return set()
try: try:
css = path.read_text() css = path.read_text()
except (OSError, ValueError): except OSError, ValueError:
return set() return set()
css = re.sub(r"/\*.*?\*/", "", css, flags=re.DOTALL) css = re.sub(r"/\*.*?\*/", "", css, flags=re.DOTALL)
m = re.search(r"color-scheme\s*:\s*([^;]+);", css, re.IGNORECASE) m = re.search(r"color-scheme\s*:\s*([^;]+);", css, re.IGNORECASE)
if not m: if not m:
return set() return set()
return {tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}} return {
tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}
}
def _theme_mode(theme: str) -> str: def _theme_mode(theme: str) -> str:
@@ -309,6 +315,9 @@ def _layout(
transition: str = "cube", transition: str = "cube",
favicon: str = "", favicon: str = "",
social: dict[str, str] | None = None, social: dict[str, str] | None = None,
lang: str = i18n.ORIGINAL_LANGUAGE,
canonical: str = "",
alternates: list[tuple[str, str]] = (),
) -> Template: ) -> Template:
"""Page layout template with standard assets and ES-module scripts. """Page layout template with standard assets and ES-module scripts.
@@ -333,17 +342,28 @@ def _layout(
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as ``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
property attributes, everything else (description, twitter:*) as name. property attributes, everything else (description, twitter:*) as name.
``lang`` is the served language for <html lang>; an RTL language (ar,
fa, ...) also puts dir="rtl" on <html> (the editor panel carries its own
lang="en" dir="ltr", so it is unaffected). ``canonical`` and
``alternates`` ((hreflang, href) pairs) are the page's language URLs
(see docs/localization.md), emitted right after the viewport and before
the social tags: canonical first, then the hreflang alternates.
""" """
doc = Document(E.Title, lang="en") doc = Document(
E.Title, lang=lang, dir="rtl" if lang in i18n.RTL_LANGUAGES else "ltr"
)
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs # Responsive layout (see the 48rem breakpoint in pagerite.css) needs
# the real device width, not the default 980px layout viewport. # the real device width, not the default 980px layout viewport.
doc.meta(name="viewport", content="width=device-width, initial-scale=1") doc.meta(name="viewport", content="width=device-width, initial-scale=1")
if canonical:
doc.link(rel="canonical", href=canonical)
for hreflang, href in alternates:
doc.link(rel="alternate", hreflang=hreflang, href=href)
for key, value in (social or {}).items(): for key, value in (social or {}).items():
if value: if value:
if key.startswith(("og:", "article:")): if key.startswith(("og:", "article:")):
doc.meta(property=key, content=value) doc.meta(property=key, content=value)
elif key == "canonical":
doc.link(rel="canonical", href=value)
else: else:
doc.meta(name=key, content=value) doc.meta(name=key, content=value)
# A custom favicon (from the site editor) is linked explicitly; without # A custom favicon (from the site editor) is linked explicitly; without
@@ -412,8 +432,7 @@ def _layout(
if custom_css.strip(): if custom_css.strip():
doc.style(custom_css, id="pagerite-user") doc.style(custom_css, id="pagerite-user")
body = ( body = (
doc doc.header(
.header(
E.div(E.Banner, id="page-banner"), E.div(E.Banner, id="page-banner"),
E.Brand, E.Brand,
E.nav(E.Nav, id="nav"), E.nav(E.Nav, id="nav"),
@@ -443,26 +462,52 @@ def _layout(
return Template(body) return Template(body)
def _brand_link(brand: str, brand_html: str = "") -> HTML: def _brand_link(brand: str, brand_html: str = "", link_lang: str = "") -> HTML:
"""Header brand: custom HTML (in a #brand wrapper, rendered instead of """Header brand: custom HTML (in a #brand wrapper, rendered instead of
the link) when configured, else the plain brand link; omitted entirely the link) when configured, else the plain brand link; omitted entirely
when neither is set.""" when neither is set."""
if brand_html.strip(): if brand_html.strip():
return HTML(str(E.div(HTML(brand_html), id="brand"))) return HTML(str(E.div(HTML(brand_html), id="brand")))
return HTML(str(E.a(brand, href="/", id="brand"))) if brand else HTML("") return (
HTML(str(E.a(brand, href=_href("", link_lang), id="brand")))
if brand
else HTML("")
)
def _title(slug: str, node: Node) -> str: def _title(
"""Menu label: the configured title, prettified slug, "Home" fallback.""" slug: str, node: Node, translation: Translation | None = None, path: str = ""
) -> str:
"""Menu label: the configured title, prettified slug, "Home" fallback.
With a translation, its title map (keyed by node path) wins, falling
back per node to the original English title.
"""
if translation and (t := translation.titles.get(path)):
return t
return node.title or prettify(slug) or "Home" return node.title or prettify(slug) or "Home"
def _href(path: str, link_lang: str = "") -> str:
"""Site-chrome link to a page: when the page was requested with a
?lang= override the query is replicated onto the navigation links it
renders, so clicks and prefetches (which take the href as-is) stay in
the chosen language even without JS (docs/localization.md)."""
return f"/{path}?lang={link_lang}" if link_lang else f"/{path}"
def _nav_link( def _nav_link(
doc, menu: dict[str, Node], node: Node, path: str, current: str, doc,
menu: dict[str, Node],
node: Node,
path: str,
current: str,
ancestors_current: bool = True, ancestors_current: bool = True,
translation: Translation | None = None,
link_lang: str = "",
) -> None: ) -> None:
"""Render one <li> linking the node. Category labels (no content of """Render one <li> linking the node. Category labels (no content of
their own None, or empty markdown as left by the site editor's their own chunks None, or an empty page as left by the site editor's
page creation) link straight to their first child page, so normal page creation) link straight to their first child page, so normal
navigation bypasses the placeholder/empty page at their own URL.""" navigation bypasses the placeholder/empty page at their own URL."""
# The navbar highlights a top-level item also when viewing any of its # The navbar highlights a top-level item also when viewing any of its
@@ -470,17 +515,22 @@ def _nav_link(
is_current = current == path or ( is_current = current == path or (
ancestors_current and path and current.startswith(f"{path}/") ancestors_current and path and current.startswith(f"{path}/")
) )
href = f"/{path}" href = _href(path, link_lang)
if not node.content and (leaf := first_leaf(menu, path)) is not None: if not node.chunks and (leaf := first_leaf(menu, path)) is not None:
href = f"/{leaf}" href = _href(leaf, link_lang)
doc.li.a( doc.li.a(
_title(path.rpartition("/")[2], node), _title(path.rpartition("/")[2], node, translation, path),
href=href, href=href,
**{"class": "current"} if is_current else {}, **{"class": "current"} if is_current else {},
) )
def nav_html(menu: dict[str, Node], current: str) -> HTML: def nav_html(
menu: dict[str, Node],
current: str,
translation: Translation | None = None,
link_lang: str = "",
) -> HTML:
"""Render the contents of the #nav element for the current path. """Render the contents of the #nav element for the current path.
Top-level items in menu order; the front page (slug "", href "/") Top-level items in menu order; the front page (slug "", href "/")
@@ -491,11 +541,24 @@ def nav_html(menu: dict[str, Node], current: str) -> HTML:
with nav: with nav:
for slug, node in sorted_nodes(menu): for slug, node in sorted_nodes(menu):
if node.published: if node.published:
_nav_link(nav, menu, node, slug, current) _nav_link(
nav,
menu,
node,
slug,
current,
translation=translation,
link_lang=link_lang,
)
return HTML(str(nav)) return HTML(str(nav))
def sidebar_html(menu: dict[str, Node], current: str) -> HTML: def sidebar_html(
menu: dict[str, Node],
current: str,
translation: Translation | None = None,
link_lang: str = "",
) -> HTML:
"""Render the #sidebar element for the current path (empty when none). """Render the #sidebar element for the current path (empty when none).
The sidebar is the current main level section's sub-navigation: the The sidebar is the current main level section's sub-navigation: the
@@ -529,24 +592,45 @@ def sidebar_html(menu: dict[str, Node], current: str) -> HTML:
nav = E.ul nav = E.ul
with nav: with nav:
for slug, child in items: for slug, child in items:
_sidebar_item(nav, menu, child, f"{section}/{slug}", current) _sidebar_item(
nav, menu, child, f"{section}/{slug}", current, translation, link_lang
)
return HTML(str(E.aside(nav, id="sidebar"))) return HTML(str(E.aside(nav, id="sidebar")))
def _sidebar_item(doc, menu: dict[str, Node], node: Node, path: str, current: str) -> None: def _sidebar_item(
doc,
menu: dict[str, Node],
node: Node,
path: str,
current: str,
translation: Translation | None = None,
link_lang: str = "",
) -> None:
"""One sidebar <li>: the node link, with its published children as a """One sidebar <li>: the node link, with its published children as a
nested list (third level and deeper, recursively).""" nested list (third level and deeper, recursively)."""
_nav_link(doc, menu, node, path, current, ancestors_current=False) _nav_link(
doc,
menu,
node,
path,
current,
ancestors_current=False,
translation=translation,
link_lang=link_lang,
)
sub = [(s, c) for s, c in sorted_nodes(node.children) if c.published] sub = [(s, c) for s, c in sorted_nodes(node.children) if c.published]
if sub: if sub:
# doc.li.a(...) above left the <li> open for nesting. # doc.li.a(...) above left the <li> open for nesting.
with doc.ul: with doc.ul:
for slug, child in sub: for slug, child in sub:
_sidebar_item(doc, menu, child, f"{path}/{slug}", current) _sidebar_item(
doc, menu, child, f"{path}/{slug}", current, translation, link_lang
)
def first_leaf(menu: dict[str, Node], path: str) -> str | None: def first_leaf(menu: dict[str, Node], path: str) -> str | None:
"""First published descendant page (content set) in menu order. """First published descendant page (chunks set) in menu order.
This is the nav-link target for content-less category labels. This is the nav-link target for content-less category labels.
""" """
@@ -561,7 +645,7 @@ def _first_leaf(node: Node, path: str) -> str | None:
if not child.published: if not child.published:
continue continue
cpath = f"{path}/{slug}" if path else slug cpath = f"{path}/{slug}" if path else slug
if child.content: if child.chunks:
return cpath return cpath
if (leaf := _first_leaf(child, cpath)) is not None: if (leaf := _first_leaf(child, cpath)) is not None:
return leaf return leaf
@@ -674,16 +758,36 @@ def banner_source(menu: dict[str, Node], path: str) -> str | None:
return None return None
def page_content(menu: dict[str, Node], path: str) -> HTML: def page_content(
menu: dict[str, Node],
data: Data,
path: str,
translation: Translation | None = None,
link_lang: str = "",
lang: str = "",
) -> HTML:
"""Render the contents of the #main element for a page. """Render the contents of the #main element for a page.
A page with published children (a category page) lists them as cards A page with published children (a category page) lists them as cards
after the markdown content. after the markdown content. With a translation, its Markdown goes
through the same render pipeline; missing pieces (markdown=None, absent
title entries) fall back to the original. ``lang`` feeds the cards'
per-target localization.
""" """
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
content = node_markdown(data, node) or ""
title = node.title
if translation:
if translation.markdown is not None:
content = translation.markdown
title = (
_title(path.rpartition("/")[2], node, translation, path)
if node.title
else title
)
# The title is injected into the markdown (as # title when it has no # The title is injected into the markdown (as # title when it has no
# h1 of its own), so title and content render as one article. # h1 of its own), so title and content render as one article.
rendered = render(node.content or "", path, node.created, node.modified, title=node.title) rendered = render(content, path, node.created, node.modified, title=title)
# Long articles get .multicol: the article column cap lifts (see the # Long articles get .multicol: the article column cap lifts (see the
# #content grid in pagerite.css) and the .cols segments lay out in at # #content grid in pagerite.css) and the .cols segments lay out in at
# most two columns. The html is already segmented by render() — the # most two columns. The html is already segmented by render() — the
@@ -691,11 +795,20 @@ def page_content(menu: dict[str, Node], path: str) -> HTML:
doc = E.article(class_="multicol") if rendered.multicol else E.article doc = E.article(class_="multicol") if rendered.multicol else E.article
with doc: with doc:
doc(HTML(rendered.html)) doc(HTML(rendered.html))
_cards(doc, menu, node, path) _cards(doc, menu, data, node, path, translation, link_lang, lang)
return HTML(str(doc)) return HTML(str(doc))
def _cards(doc, menu: dict[str, Node], node: Node, path: str) -> None: def _cards(
doc,
menu: dict[str, Node],
data: Data,
node: Node,
path: str,
translation: Translation | None = None,
link_lang: str = "",
lang: str = "",
) -> None:
"""Card stacks of the node's published children (nothing when childless). """Card stacks of the node's published children (nothing when childless).
One column per direct child, all in a single full-width row (the .wide One column per direct child, all in a single full-width row (the .wide
@@ -721,35 +834,54 @@ def _cards(doc, menu: dict[str, Node], node: Node, path: str) -> None:
continue continue
with doc.div(class_="stack"): with doc.div(class_="stack"):
for epath, enode in entries: for epath, enode in entries:
_card(doc, enode, epath) _card(doc, data, enode, epath, translation, link_lang, lang)
def _walk(node: Node, path: str): def _walk(node: Node, path: str):
"""Published content pages of a subtree, pre-order in menu order: the """Published content pages of a subtree, pre-order in menu order: the
node itself first when it has content (the stack's landing card), then node itself first when it has content (the stack's landing card), then
its descendants (content-less nodes contribute only their subtree).""" its descendants (content-less nodes contribute only their subtree)."""
if node.content: if node.chunks:
yield path, node yield path, node
for slug, child in sorted_nodes(node.children): for slug, child in sorted_nodes(node.children):
if child.published: if child.published:
yield from _walk(child, f"{path}/{slug}") yield from _walk(child, f"{path}/{slug}")
def _card(doc, node: Node, path: str) -> None: def _card(
doc,
data: Data,
node: Node,
path: str,
translation: Translation | None = None,
link_lang: str = "",
lang: str = "",
) -> None:
"""One card in a stack: cover + title, plus the description when the """One card in a stack: cover + title, plus the description when the
page has no image (its card shows a gradient cover instead).""" page has no image (its card shows a gradient cover instead).
The card text localizes per target article where that page is
available in the language: the title comes from the translation's
title map and the cover/description heuristics run on the target's
hybrid Markdown with per-card fallback to the original otherwise.
"""
image = description = "" image = description = ""
if node.content: if node.chunks:
html = render(node.content, path, node.created, node.modified).html md = node_markdown(data, node) or ""
if lang and lang in node.langs:
md = i18n.hybrid_markdown(data, node, path, lang)
html = render(md, path, node.created, node.modified).html
image, _ = _media(html) image, _ = _media(html)
if not image: if not image:
description = _description(html, 150) description = _description(html, 150)
with doc.a(href=f"/{path}", class_="card"): with doc.a(href=_href(path, link_lang), class_="card"):
if image: if image:
doc.span(class_="cover", style=f'background-image: url("{image}")') doc.span(class_="cover", style=f'background-image: url("{image}")')
else: else:
doc.span(class_="cover") doc.span(class_="cover")
doc.span(_title(path.rpartition("/")[2], node), class_="title") doc.span(
_title(path.rpartition("/")[2], node, translation, path), class_="title"
)
if description: if description:
doc.span(description, class_="desc") doc.span(description, class_="desc")
@@ -776,7 +908,9 @@ def _description(html: str, limit: int = 200) -> str:
return text return text
# Prefer a clean cut: the last sentence ending within the limit, as # Prefer a clean cut: the last sentence ending within the limit, as
# long as it does not reduce the description to a tiny fragment. # long as it does not reduce the description to a tiny fragment.
if (end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)) > limit // 2: if (
end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)
) > limit // 2:
return text[:end] return text[:end]
return text[:limit].rsplit(" ", 1)[0] + "" return text[:limit].rsplit(" ", 1)[0] + ""
@@ -815,7 +949,10 @@ def _share_media(html: str, base_url: str) -> tuple[str, str]:
"""(image, video) share URLs from the rendered article. """(image, video) share URLs from the rendered article.
The _media picks as absolute URLs built from the request base The _media picks as absolute URLs built from the request base
social scrapers cannot use relative ones. social scrapers cannot use relative ones. Extension-less store links
(``/_f/<hash>``) are used as-is: the server negotiates the format
from the scraper's Accept header (no explicit image/avif|webp → JPEG,
which every scraper supports).
""" """
if not base_url: if not base_url:
return "", "" return "", ""
@@ -829,7 +966,12 @@ def _share_media(html: str, base_url: str) -> tuple[str, str]:
def _social_meta( def _social_meta(
node: Node, path: str, title: str, html: str, brand: str, base_url: str, node: Node,
path: str,
title: str,
html: str,
brand: str,
base_url: str,
) -> dict[str, str]: ) -> dict[str, str]:
"""Open Graph/Twitter/SEO meta tags for a content page. """Open Graph/Twitter/SEO meta tags for a content page.
@@ -838,13 +980,17 @@ def _social_meta(
share image the article's first <img> — authors lead with their most share image the article's first <img> — authors lead with their most
representative figure. Absolute URLs are built from the request's base representative figure. Absolute URLs are built from the request's base
(social scrapers cannot use relative ones). (social scrapers cannot use relative ones).
``twitter:image`` pins extension-less store links to the ``.webp``
variant: X only honors WebP via twitter:image (not og:image) and its
scraper cannot be trusted to negotiate via Accept.
""" """
url = f"{base_url}/{path}" if base_url else "" url = f"{base_url}/{path}" if base_url else ""
text = _description(html) text = _description(html)
image, video = _share_media(html, base_url) image, video = _share_media(html, base_url)
twitter_image = re.sub(r"(/_f/[0-9a-f]{12})$", r"\1.webp", image) if image else ""
return { return {
"description": text, "description": text,
"canonical": url,
"og:type": "article", "og:type": "article",
"og:title": title, "og:title": title,
"og:description": text, "og:description": text,
@@ -855,11 +1001,13 @@ def _social_meta(
"article:published_time": node.created.isoformat(), "article:published_time": node.created.isoformat(),
"article:modified_time": node.modified.isoformat(), "article:modified_time": node.modified.isoformat(),
"twitter:card": "summary_large_image" if image else "summary", "twitter:card": "summary_large_image" if image else "summary",
"twitter:image": twitter_image,
} }
def render_page( def render_page(
menu: dict[str, Node], menu: dict[str, Node],
data: Data,
path: str, path: str,
brand: str = SITE_NAME, brand: str = SITE_NAME,
custom_css: str = "", custom_css: str = "",
@@ -868,21 +1016,59 @@ def render_page(
brand_html: str = "", brand_html: str = "",
base_url: str = "", base_url: str = "",
transition: str = "cube", transition: str = "cube",
lang: str = i18n.ORIGINAL_LANGUAGE,
translation: Translation | None = None,
link_lang: str = "",
) -> str: ) -> str:
"""Render a full HTML page for the slug path.""" """Render a full HTML page for the slug path.
``lang``/``translation`` serve a translated version (see
docs/localization.md): None translation = the English original.
``link_lang`` is the ?lang= override the page was requested with,
replicated onto the navigation links so the language sticks.
"""
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
title = _title(path.rpartition("/")[2], node) original = i18n.primary_lang(menu, path)
main = page_content(menu, path) if translation is None:
lang = original
title = _title(path.rpartition("/")[2], node, translation, path)
main = page_content(menu, data, path, translation, link_lang, lang)
social = _social_meta(node, path, title, str(main), brand, base_url) social = _social_meta(node, path, title, str(main), brand, base_url)
# Canonical/hreflang URLs (docs/localization.md): the canonical names
# the actually served language — the plain URL for the original (for
# SEO the non-query URL means the article's language), ?lang= for a
# translation — regardless of how the language was arrived at (query
# or header). The alternates are site-wide, the same set on every
# page: the configured translate_langs (the translator works to fill
# them all in), x-default first (the plain, autodetecting URL), then
# every language explicitly, the page's own primary included.
canonical = ""
alternates = []
if base_url:
url = f"{base_url}/{path}"
canonical = url if lang == original else f"{url}?lang={lang}"
if data.translate_langs:
alternates = [("x-default", url)] + [
(tag, f"{url}?lang={tag}")
for tag in sorted({original, *data.translate_langs})
]
return str( return str(
_layout( _layout(
*_page_assets(), custom_css, theme, banner_design(menu, path, theme), *_page_assets(),
transition, favicon, social, custom_css,
theme,
banner_design(menu, path, theme),
transition,
favicon,
social,
lang,
canonical,
alternates,
)( )(
Title=f"{title} {brand}" if brand else title, Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html, link_lang),
Nav=nav_html(menu, path), Nav=nav_html(menu, path, translation, link_lang),
Sidebar=sidebar_html(menu, path), Sidebar=sidebar_html(menu, path, translation, link_lang),
Banner=banner_html(menu, path, theme), Banner=banner_html(menu, path, theme),
Main=main, Main=main,
), ),
@@ -891,6 +1077,7 @@ def render_page(
def render_category( def render_category(
menu: dict[str, Node], menu: dict[str, Node],
data: Data,
path: str, path: str,
brand: str = SITE_NAME, brand: str = SITE_NAME,
custom_css: str = "", custom_css: str = "",
@@ -898,6 +1085,9 @@ def render_category(
favicon: str = "", favicon: str = "",
brand_html: str = "", brand_html: str = "",
transition: str = "cube", transition: str = "cube",
lang: str = i18n.ORIGINAL_LANGUAGE,
translation: Translation | None = None,
link_lang: str = "",
) -> str: ) -> str:
"""Render the listing for a content-less category label (404). """Render the listing for a content-less category label (404).
@@ -905,22 +1095,37 @@ def render_category(
children are listed as cards, like on a category page with content. children are listed as cards, like on a category page with content.
Nav links point straight at the first child, so this is mainly seen Nav links point straight at the first child, so this is mainly seen
in the site editor, where the pen creates the landing page. in the site editor, where the pen creates the landing page.
With a translation (titles only the category has no Markdown) the
heading, navigation and card text localize per target article
(docs/localization.md); ``link_lang`` replicates the ?lang= override
onto the navigation links as on content pages.
""" """
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
title = _title(path.rpartition("/")[2], node) if translation is None:
lang = i18n.primary_lang(menu, path)
title = _title(path.rpartition("/")[2], node, translation, path)
doc = E.article doc = E.article
with doc: with doc:
doc.h1(title) doc.h1(title)
if any(c.published for c in node.children.values()): if any(c.published for c in node.children.values()):
_cards(doc, menu, node, path) _cards(doc, menu, data, node, path, translation, link_lang, lang)
else: else:
doc.p("This section has no page of its own yet.") doc.p("This section has no page of its own yet.")
return str( return str(
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon)( _layout(
*_page_assets(),
custom_css,
theme,
banner_design(menu, path, theme),
transition,
favicon,
lang=lang,
)(
Title=f"{title} {brand}" if brand else title, Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html, link_lang),
Nav=nav_html(menu, path), Nav=nav_html(menu, path, translation, link_lang),
Sidebar=sidebar_html(menu, path), Sidebar=sidebar_html(menu, path, translation, link_lang),
Banner=banner_html(menu, path, theme), Banner=banner_html(menu, path, theme),
Main=HTML(str(doc)), Main=HTML(str(doc)),
), ),
@@ -943,7 +1148,14 @@ def render_not_found(
doc.h1("Not Found") doc.h1("Not Found")
doc.p(f"No article at /{path}. If there was before, it may have been deleted.") doc.p(f"No article at /{path}. If there was before, it may have been deleted.")
return str( return str(
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon)( _layout(
*_page_assets(),
custom_css,
theme,
banner_design(menu, path, theme),
transition,
favicon,
)(
Title=f"Not Found {brand}" if brand else "Not Found", Title=f"Not Found {brand}" if brand else "Not Found",
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html),
Nav=nav_html(menu, path), Nav=nav_html(menu, path),
+5 -2
View File
@@ -17,14 +17,15 @@ readme = "README.md"
requires-python = ">=3.14" requires-python = ">=3.14"
dependencies = [ dependencies = [
"blake3>=1.0.9", "blake3>=1.0.9",
"fastapi-vue>=1.3.1", "fastapi-vue~=1.4.0",
"fastapi[standard]>=0.141.1", "fastapi[standard]>=0.141.1",
"html5tagger>=2.0.0", "html5tagger>=2.0.0",
"httpx>=0.28.1", "httpx>=0.28.1",
"kanta>=0.8.1", "kanta>=0.9.0",
"markdown-it-py>=4.2.0", "markdown-it-py>=4.2.0",
"maxminddb>=3.1.1", "maxminddb>=3.1.1",
"mdit-py-plugins>=0.6.1", "mdit-py-plugins>=0.6.1",
"mediapreview[standard]>=0.2.3",
"platformdirs>=4.11.5", "platformdirs>=4.11.5",
"pygments>=2.20.0", "pygments>=2.20.0",
"python-slugify>=8.0.4", "python-slugify>=8.0.4",
@@ -37,6 +38,7 @@ pagerite = "pagerite.__main__:main"
[project.urls] [project.urls]
Repository = "https://git.zi.fi/LeoVasanko/pagerite" Repository = "https://git.zi.fi/LeoVasanko/pagerite"
Issues = "https://github.com/LeoVasanko/pagerite"
[dependency-groups] [dependency-groups]
dev = [] dev = []
@@ -53,6 +55,7 @@ packages = ["pagerite"]
# the frontend build). # the frontend build).
artifacts = ["pagerite/frontend-build", "pagerite/themes", "pagerite/seed-assets"] artifacts = ["pagerite/frontend-build", "pagerite/themes", "pagerite/seed-assets"]
only-packages = true only-packages = true
packages = ["pagerite"]
[tool.hatch.build.targets.sdist.hooks.custom] [tool.hatch.build.targets.sdist.hooks.custom]
path = "scripts/fastapi-vue/buildhook.py" path = "scripts/fastapi-vue/buildhook.py"
+4 -1
View File
@@ -8,6 +8,8 @@ import sys
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
import tracerite
# Import util.py from scripts/fastapi-vue (not a package, so we adjust sys.path) # Import util.py from scripts/fastapi-vue (not a package, so we adjust sys.path)
sys.path.insert(0, str(Path(__file__).with_name("fastapi-vue"))) sys.path.insert(0, str(Path(__file__).with_name("fastapi-vue")))
from devutil import ( from devutil import (
@@ -39,7 +41,7 @@ async def run_devserver(
viteurl, npm_install, vite = setup_vite(listen, DEFAULT_VITE_PORT) viteurl, npm_install, vite = setup_vite(listen, DEFAULT_VITE_PORT)
backurl, pagerite = setup_cli("pagerite", backend, DEFAULT_DEV_PORT) backurl, pagerite = setup_cli("pagerite", backend, DEFAULT_DEV_PORT)
# Tell the everyone by environment (vite proxy and backend devmode use these) # Tell everyone via environment (vite proxy and backend devmode use these)
os.environ["PAGERITE_VITE_URL"] = viteurl os.environ["PAGERITE_VITE_URL"] = viteurl
os.environ["PAGERITE_BACKEND_URL"] = backurl os.environ["PAGERITE_BACKEND_URL"] = backurl
os.environ["PAGERITE_DEV"] = "1" os.environ["PAGERITE_DEV"] = "1"
@@ -54,6 +56,7 @@ async def run_devserver(
def main() -> None: def main() -> None:
"""Parse CLI arguments and run the devserver.""" """Parse CLI arguments and run the devserver."""
tracerite.load()
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
description="Run Vite and FastAPI development servers", description="Run Vite and FastAPI development servers",
formatter_class=argparse.RawDescriptionHelpFormatter, formatter_class=argparse.RawDescriptionHelpFormatter,
+27 -19
View File
@@ -91,22 +91,22 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
CrawlerProfile( CrawlerProfile(
"googlebot", "googlebot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"66.249.64.66", # US, Google "66.249.64.66", # US, Google
), ),
CrawlerProfile( CrawlerProfile(
"bingbot", "bingbot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"40.77.167.0", # US, Microsoft "40.77.167.0", # US, Microsoft
), ),
CrawlerProfile( CrawlerProfile(
"duckduckbot", "duckduckbot",
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)", "DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
"95.217.0.1", # Germany, Hetzner VPS "95.217.0.1", # Germany, Hetzner VPS
), ),
CrawlerProfile( CrawlerProfile(
"curl", "curl",
"curl/8.5.0", "curl/8.5.0",
"139.162.0.1", # Singapore, Linode VPS "139.162.0.1", # Singapore, Linode VPS
), ),
] ]
@@ -115,14 +115,14 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
# host part; the host may rotate once mid-session. # host part; the host may rotate once mid-session.
RESIDENTIAL_SOURCE_IPS: list[str] = [ RESIDENTIAL_SOURCE_IPS: list[str] = [
# Residential IPv4 # Residential IPv4
"91.154.140.209", # Finland, Elisa "91.154.140.209", # Finland, Elisa
"84.143.145.207", # Germany, Deutsche Telekom "84.143.145.207", # Germany, Deutsche Telekom
"220.165.255.254", # China, Chinanet / China Telecom "220.165.255.254", # China, Chinanet / China Telecom
"84.235.83.162", # Saudi Arabia, SaudiNet / STC "84.235.83.162", # Saudi Arabia, SaudiNet / STC
# Residential IPv6 /64 prefixes # Residential IPv6 /64 prefixes
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom "2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
"240e:45d:1e60:5b0::/64", # China, China Telecom "240e:45d:1e60:5b0::/64", # China, China Telecom
"2409:8904:6720:4123::/64", # China, China Unicom "2409:8904:6720:4123::/64", # China, China Unicom
] ]
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for # Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
@@ -130,9 +130,9 @@ RESIDENTIAL_SOURCE_IPS: list[str] = [
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA, # Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
# and index 2 uses a fixed crawler UA. # and index 2 uses a fixed crawler UA.
ABUSE_SOURCE_IPS: list[str] = [ ABUSE_SOURCE_IPS: list[str] = [
"45.63.0.12", # US, Vultr VPS "45.63.0.12", # US, Vultr VPS
"138.197.0.89", # US, DigitalOcean / Cloudways "138.197.0.89", # US, DigitalOcean / Cloudways
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS "2a01:4f8:0:2::1234", # Germany, Hetzner VPS
] ]
# Paths commonly probed by attackers looking for exposed config, admin panels, # Paths commonly probed by attackers looking for exposed config, admin panels,
@@ -284,10 +284,16 @@ TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}), ("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}), ("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}), ("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}), (
"https://news.ycombinator.com/",
{"utm_source": "hackernews", "utm_medium": "referral"},
),
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}), ("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}), ("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}), (
"https://www.producthunt.com/",
{"utm_source": "producthunt", "utm_medium": "referral"},
),
] ]
# Fraction of referered sessions that also carry UTM tags. # Fraction of referered sessions that also carry UTM tags.
@@ -437,7 +443,7 @@ def _random_ipv6_host(prefix: str) -> str:
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}") raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
if base.endswith("::"): if base.endswith("::"):
base = base[:-2] base = base[:-2]
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4)) host = ":".join(f"{random.randint(0, 0xFFFF):04x}" for _ in range(4))
return f"{base}:{host}" return f"{base}:{host}"
@@ -760,9 +766,11 @@ def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
if ua_mode == 0: if ua_mode == 0:
get_ua = _abuse_ua get_ua = _abuse_ua
elif ua_mode == 1: elif ua_mode == 1:
def get_ua() -> str: def get_ua() -> str:
return BROWSER_PROFILES[0].user_agent return BROWSER_PROFILES[0].user_agent
else: else:
def get_ua() -> str: def get_ua() -> str:
return CRAWLER_PROFILES[0].user_agent return CRAWLER_PROFILES[0].user_agent
@@ -804,8 +812,8 @@ def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
nargs="?", nargs="?",
default="http://localhost:8200", default="http://localhost:8200",
help="Base URL of the Pagerite site (default: http://localhost:8200). " help="Base URL of the Pagerite site (default: http://localhost:8200). "
"A bare :PORT or PORT is treated as http://localhost:PORT; a " "A bare :PORT or PORT is treated as http://localhost:PORT; a "
"missing scheme defaults to http://.", "missing scheme defaults to http://.",
) )
parser.add_argument( parser.add_argument(
"-t", "-t",
+48 -21
View File
@@ -7,8 +7,8 @@ import sys
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, Any, Self from typing import TYPE_CHECKING, Any, Self
from urllib.parse import urlsplit
import httpx
from buildutil import find_dev_tool, find_install_tool, logger from buildutil import find_dev_tool, find_install_tool, logger
from fastapi_vue.hostutil import parse_endpoint from fastapi_vue.hostutil import parse_endpoint
@@ -105,18 +105,47 @@ class ProcessGroup:
await p.wait() await p.wait()
async def http_get_server(url: str, timeout: float) -> str | None: # noqa: ASYNC109
"""GET url with plain asyncio streams, return the response Server header.
Returns an empty string when the server responds without a Server header,
and None when the server is unreachable or doesn't answer in time.
"""
parts = urlsplit(url)
host = parts.hostname or "localhost"
port = parts.port or (443 if parts.scheme == "https" else 80)
path = parts.path or "/"
if parts.query:
path += f"?{parts.query}"
try:
async with asyncio.timeout(timeout):
reader, writer = await asyncio.open_connection(host, port)
try:
writer.write(f"GET {path} HTTP/1.0\r\nHost: {host}\r\n\r\n".encode())
await writer.drain()
data = await reader.readuntil(b"\r\n\r\n")
finally:
writer.close()
except OSError, EOFError, ValueError, TimeoutError:
return None
for line in data.decode("latin-1").split("\r\n"):
if line.lower().startswith("server:"):
return line.split(":", 1)[1].strip()
return ""
async def check_ports_free(*urls: str) -> None: async def check_ports_free(*urls: str) -> None:
"""Verify URLs are not responding (ports are free). Raise SystemExit if any respond.""" """Verify URLs are not responding (ports are free). Raise SystemExit if any respond."""
async def check(client: httpx.AsyncClient, url: str) -> None: async def check(url: str) -> None:
with suppress(httpx.RequestError): server = await http_get_server(url, timeout=0.1)
res = await client.get(url, timeout=0.1) if server is not None:
server = res.headers.get("server", "server") logger.warning(
logger.warning("Conflicting %s already running at %s", server, url) "Conflicting %s already running at %s", server or "server", url
)
raise SystemExit(1) raise SystemExit(1)
async with httpx.AsyncClient() as client: await asyncio.gather(*[check(url) for url in urls])
await asyncio.gather(*[check(client, url) for url in urls])
async def ready(url: str, path: str = "", max_attempts: int = 50) -> None: async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
@@ -128,18 +157,14 @@ async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
if not path: if not path:
return return
async with httpx.AsyncClient() as client: for attempt in range(max_attempts):
for attempt in range(max_attempts): if await http_get_server(f"{url}{path}", timeout=1.0) is not None:
try: logger.info("✓ Backend ready!")
await client.get(f"{url}{path}", timeout=1.0) return
except httpx.RequestError: if attempt == max_attempts - 1:
if attempt == max_attempts - 1: logger.warning("Backend didn't start in time")
logger.warning("Backend didn't start in time") raise SystemExit(1)
raise SystemExit(1) from None await asyncio.sleep(0.1)
await asyncio.sleep(0.1)
else:
logger.info("✓ Backend ready!")
return
def setup_vite( def setup_vite(
@@ -222,5 +247,7 @@ def setup_cli(
host = endpoints[0]["host"] host = endpoints[0]["host"]
port = endpoints[0]["port"] port = endpoints[0]["port"]
cmd = [cli, f"--listen={host}:{port}"] # Run the package as a module with the current interpreter, instead of
# relying on a PATH-installed CLI entry point.
cmd = [sys.executable, "-m", cli, f"--listen={host}:{port}"]
return f"http://{host}:{port}", cmd return f"http://{host}:{port}", cmd
+378
View File
@@ -0,0 +1,378 @@
#!/usr/bin/env python3
# /// script
# requires-python = ">=3.14"
# dependencies = [
# "accelerate>=1.14.0",
# "msgspec>=0.19.0",
# "torch>=2.13.0",
# "tracerite>=2.6.5",
# "transformers>=5.16.1",
# "websockets>=15.0.1",
# ]
# ///
"""Pagerite translator service: translate site content with Seed-X-PPO-7B.
Connects to a Pagerite server's translator WebSocket — the full URL
including the access key (the admin finds the key in the site settings,
GET /_api/settings -> ``translate_keys``)
and announces the languages the
model CAN translate (capabilities). The server dispatches one single-item
job at a time per connection, offered only in its configured target
languages (``Data.translate_langs``) the announced capabilities; a
dropped connection's in-flight item is simply re-offered
(docs/localization.md). For parallelism, run multiple instances.
Seed-X-PPO-7B (bf16, ~15 GB) is the only supported model. The script stays
running and connected full time; the model loads at startup (a backlog is
likely after a downtime) and is unloaded after 60 s idle, re-loading on the
next job the GPU is held only while actually translating.
Usage:
uv run scripts/translator.py ws://localhost:8410/_translate/KEY
uv run scripts/translator.py wss://example.com/_translate/KEY
"""
import argparse
import asyncio
import gc
import sys
import time
import msgspec
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import tracerite
import websockets
tracerite.load()
SEED_X = "ByteDance-Seed/Seed-X-PPO-7B"
# Seed-X language tags (appended to the prompt; required by its PPO training)
SEED_X_TAGS = {
"arabic": "ar",
"chinese": "zh",
"czech": "cs",
"danish": "da",
"dutch": "nl",
"english": "en",
"finnish": "fi",
"french": "fr",
"german": "de",
"greek": "el",
"hungarian": "hu",
"indonesian": "id",
"italian": "it",
"japanese": "ja",
"korean": "ko",
"malay": "ms",
"norwegian": "no",
"persian": "fa",
"polish": "pl",
"portuguese": "pt",
"romanian": "ro",
"russian": "ru",
"spanish": "es",
"swedish": "sv",
"thai": "th",
"turkish": "tr",
"ukrainian": "uk",
"vietnamese": "vi",
}
SEED_X_NAMES = {v: k for k, v in SEED_X_TAGS.items()}
#: The fragments arrive as prose segments (pagerite/segments.py): plain
#: text runs only — no markup, URLs, code or placeholders. The wire
#: invariant is that segments are PURE PROSE, and one character marks the
#: boundary both ways: "<" never appears in a segment. Sources containing
#: it are never dispatched (pagerite/segments.py keeps them in the
#: original language); the model's output is cut at the first "<" — one
#: rule that covers the whole class of markup bleed (an echoed <lang> tag,
#: a "<br>", ...) instead of a pattern per artifact. (Generation-level
#: stop strings can't do this job: the model's <s> framing token would
#: trip a "<" stop at the first token; skip_special_tokens strips the
#: framing at decode.)
#:
#: Two kinds cross the wire (Job.kind), each with its own prompt template:
#: titles get told they ARE titles (a lone word otherwise invites
#: context-free readings — "About" as "approximately"). Any segment may
#: carry its surround in Job.contexts (a title: the article's opening; a
#: carved-out segment like a link text: its block's plain text) and is then
#: translated together with that surround (seed_x_chunk). No punctuation
#: clause, on purpose: Seed-X handles trailing-punctuation instructions by
#: slipping into its [COT] reasoning mode (observed for Chinese:
#: minutes-long generations, reasoning text in the output) —
#: match_punctuation handles stray punctuation deterministically instead.
PROMPTS = {
"chunk": "Translate the following {source_lang} text into {target_lang}:\n{text} <{tag}>",
"title": "Translate the following {source_lang} title into {target_lang}:\n{text} <{tag}>",
# Title with the article's opening as context (Job.contexts): the model
# translates both; generation stops at the blank line separating them,
# and the segment's own part of the output is the translation. No
# separator in the output (the model merged them) → seed_x_chunk falls
# back to the plain kind template.
"title+context": "Translate the following {source_lang} title and the beginning of its article "
"into {target_lang}:\n{text}\n\n{context} <{tag}>",
# A segment carved out of a larger block (link text, partial run) with
# its sentence as context — same mechanics as title+context.
"chunk+context": "Translate the following {source_lang} text into {target_lang}:\n"
"{text}\n\n{context} <{tag}>",
}
TERMINAL_PUNCT = ".,!?:;…。,!?;:、"
def match_punctuation(source: str, translated: str) -> str:
"""Drop terminal punctuation the model added.
When the source segment ends without terminal punctuation, the
translation must not gain any either. A leading Spanish ¡/¿ only pairs
with a terminal !/?, so it goes with it.
"""
if not source or source[-1] in TERMINAL_PUNCT:
return translated
trimmed = translated.rstrip(TERMINAL_PUNCT)
if trimmed and trimmed[0] in "¡¿":
trimmed = trimmed[1:].lstrip()
return trimmed
# The wire structs below duplicate pagerite/translate.py: this script runs
# in its own uv environment and cannot import the server package. The
# "type" tag selects the frame; bytes fields ride as base64.
class Hello(msgspec.Struct, tag="hello"):
"""Client greeting on connect: the language codes its model CAN produce
(capabilities). The server offers jobs only in the intersection with
its wanted target languages."""
langs: list[str]
class Job(msgspec.Struct, tag="job"):
"""Server push: ONE fragment to translate. Exactly one job is in flight
per connection the next arrives only after this one's Result."""
lang: str
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
#: The fragment's prose segments: plain text runs only, no markup —
#: translate each element independently (pagerite/segments.py).
texts: list[str]
path: str #: article it came from ("" = front page), no leading slash
kind: str #: "chunk" | "title"
#: Per segment (parallel to texts; "" = none): the surround to
#: translate it in — a link text carries its sentence, a title the
#: article's opening. See seed_x_chunk for how they are used.
contexts: list[str] = msgspec.field(default_factory=list)
class Result(msgspec.Struct, tag="result"):
"""Client reply: the translation of the connection's current Job
(must match its lang and key exactly)."""
lang: str
key: bytes
texts: list[str] #: the job's segments translated, same order and count
#: Idle seconds after the last job before the model is unloaded (the GPU
#: is released; the WebSocket connection and tokenizer stay).
IDLE_UNLOAD_S = 60
class SeedX:
"""The Seed-X model, loaded at startup and re-loaded on demand.
Loading up front covers the likely backlog after a downtime (and any
first-run model download) before the server starts dispatching. After
IDLE_UNLOAD_S without a job the model is dropped and re-loaded on the
next one the script stays connected the whole time, holding the GPU
only while translating. The tokenizer (small, CPU) loads once.
"""
def __init__(self):
self.tokenizer = AutoTokenizer.from_pretrained(SEED_X)
self.model = None
self._unload_task = None
self._load()
def _load(self):
t0 = time.monotonic()
self.model = AutoModelForCausalLM.from_pretrained(
SEED_X, dtype=torch.bfloat16, device_map="auto"
)
print(f"[seed-x loaded in {time.monotonic() - t0:.0f}s]", file=sys.stderr)
def get(self):
"""The (tokenizer, model) pair, re-loading the model if it was
idle-unloaded, and cancelling any pending idle unload."""
if self._unload_task:
self._unload_task.cancel()
self._unload_task = None
if self.model is None:
self._load()
return self.tokenizer, self.model
def idle(self):
"""Re-arm the idle unload after a job completes (arming it at job
START could unload under a >IDLE_UNLOAD_S generation)."""
self._unload_task = asyncio.create_task(self._unload_later())
async def _unload_later(self):
try:
await asyncio.sleep(IDLE_UNLOAD_S)
except asyncio.CancelledError:
return
if self.model is not None:
self.model = None
gc.collect()
if torch.cuda.is_available():
torch.cuda.empty_cache()
print(f"[seed-x unloaded after {IDLE_UNLOAD_S}s idle]", file=sys.stderr)
def seed_x_chunk(
tokenizer,
model,
text: str,
target_lang: str,
tag: str,
kind: str = "chunk",
context: str = "",
source_lang: str = "English",
):
"""Translate one segment; returns (translation, output_tokens, generation_seconds).
With context, the segment is translated together with its surround (a
link text with its sentence, a title with the article's opening), and
the segment's own part of the output is the translation: its own line
for a single-line source (a single-line segment's translation never
contains a line break generation stops at the blank line separating
the two), its own paragraph for a multi-line one (softbreak-merged
lines keep single newlines, the separator is the blank line). If the
model merged them no separator, or an empty first part fall back to
translating the segment alone; the wasted tokens are counted either
way.
"""
# No chat template on this model; the trailing language tag is required (trans/ style prompt).
template = PROMPTS.get(f"{kind}+context" if context else kind, PROMPTS["chunk"])
prompt = template.format(
source_lang=source_lang,
target_lang=target_lang,
text=text,
tag=tag,
context=context,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
t0 = time.monotonic()
# The only stop string is the context separator. "<" must NOT be one:
# stopping works on the raw output, which always starts with the
# model's <s> framing token. skip_special_tokens strips <s>/</s> at
# decode; the post-decode cut at the first "<" then enforces the wire
# invariant (prose only) against markup bleed.
kwargs = {"stop_strings": ["\n\n"], "tokenizer": tokenizer} if context else {}
out = model.generate(
**inputs,
max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]),
do_sample=False,
**kwargs,
)
dt = time.monotonic() - t0
n = out.shape[1] - inputs.input_ids.shape[1]
decoded = tokenizer.decode(
out[0][inputs.input_ids.shape[1] :], skip_special_tokens=True
)
translated = decoded.partition("<")[0]
if not context:
return translated.strip(), n, dt
if "\n" in text:
# Multi-line segment: its translation keeps single newlines; the
# blank line is the separator from the context translation.
sep = "\n\n" in translated
out = translated.split("\n\n", 1)[0] if sep else ""
else:
out, sep, _ = translated.partition("\n")
if not sep:
out = ""
out = out.strip()
if out:
return out, n, dt
# The model merged segment and context (no separator, or an empty first
# part): retry without the context.
again, n2, dt2 = seed_x_chunk(
tokenizer, model, text, target_lang, tag, kind=kind, source_lang=source_lang
)
return again, n + n2, dt + dt2
async def do_job(ws, job: Job, seed_x: SeedX) -> None:
"""Translate the job's segments (one model call each) and send them back."""
lang_name = SEED_X_NAMES[job.lang].capitalize()
tokenizer, model = seed_x.get()
# Deliberately blocking: nothing else needs the loop while the job is
# being answered, and the reconnect loop recovers a dropped connection
# (the in-flight item is simply re-offered).
texts = []
tokens = dt = 0
for i, text in enumerate(job.texts):
ctx = job.contexts[i] if i < len(job.contexts) else ""
translated, n, t = seed_x_chunk(
tokenizer, model, text, lang_name, job.lang, kind=job.kind, context=ctx
)
texts.append(match_punctuation(text, translated))
tokens += n
dt += t
print(
f"[{job.lang} {job.kind} {job.path or '/'}: {len(texts)} segments, "
f"{tokens} tokens in {dt:.1f}s = {tokens / dt:.1f} tok/s]",
file=sys.stderr,
)
await ws.send(
msgspec.json.encode(Result(lang=job.lang, key=job.key, texts=texts)).decode()
)
seed_x.idle()
async def serve(url: str, seed_x: SeedX) -> None:
"""Connect, announce capabilities, answer jobs; reconnect with backoff."""
seed_x.idle() # the startup load also unloads when no work arrives
backoff = 1
while True:
try:
async with websockets.connect(url) as ws:
backoff = 1
await ws.send(
msgspec.json.encode(Hello(langs=sorted(SEED_X_NAMES))).decode()
)
print(
f"[connected; announced {len(SEED_X_NAMES)} language capabilities]",
file=sys.stderr,
)
async for raw in ws:
await do_job(ws, msgspec.json.decode(raw, type=Job), seed_x)
except websockets.exceptions.InvalidHandshake:
sys.exit("handshake rejected; check the URL (including the key)")
except (OSError, websockets.exceptions.ConnectionClosed) as e:
print(
f"[connection lost ({e}); reconnecting in {backoff}s]", file=sys.stderr
)
await asyncio.sleep(backoff)
backoff = min(backoff * 2, 60)
def main():
p = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
)
p.add_argument(
"url",
help="full translator WebSocket URL including the key, "
"e.g. ws://localhost:8410/_translate/KEY",
)
args = p.parse_args()
if not args.url.startswith(("ws://", "wss://")):
p.error("url must start with ws:// or wss://")
asyncio.run(serve(args.url, SeedX()))
if __name__ == "__main__":
main()