Compare commits

..
31 Commits
Author SHA1 Message Date
LeoVasanko 7724290921 analytics viewer: referer + UTM as one neutral-grey badge
Favicon flush to the badge's left edge, then the referer host and UTM
summary in one link, with a single newline-formatted tooltip (origin,
then each utm_* pair). Fixed --badge-bg/--badge-text/--badge-muted
palette (defined in pagerite.css, deliberately unthemed) so transparent
favicons stay legible in both light and dark themes; own-article trail
links remain plain, lifting referer info visually apart from them.
Replaces the outlined utm-tag pill.
2026-09-15 05:20:57 +00:00
LeoVasanko 3b1804b093 analytics: record and show the rendered content language of page views
GETs store the resolved page language, activity pings carry <html lang>
(including a new ping on in-place language switches, which also keeps the
switch's GET out of crawler classification), and trail items/crawler hits
surface it. The viewer shows flags discreetly: nothing on single-language
sites or primary-language-only rows, one leading flag for uniform visits,
transition flags on mixed trails, a flag row per crawler.
2026-09-15 05:20:57 +00:00
LeoVasanko 4cd8dbfc73 Update fastapi-vue-setup 1.6.1 2026-09-15 05:20:57 +00:00
LeoVasanko dc4bdf6efc Avoid deprecation warning from pyvips; remove unnecessary logging config. 2026-09-15 05:20:57 +00:00
LeoVasanko a458bded09 Create the per-hostname data directory on first run 2026-09-15 05:20:57 +00:00
LeoVasanko 3a4745389e analytics viewer: copy-to-clipboard UAs, crawler info links, shared icon-btn
- UA lines click to copy the raw UA; abuse rows with rotating clients
  copy every variation, one per line (uaRaws)
- "Copied!" popup is viewport-fixed at the click point; cell overflow
  clipped the old absolute one
- language tags show their Intl.DisplayNames full name on hover
- a 🔗 link after the pretty UA opens the crawler info page uarite
  provides
- the emoji-symbol dim-until-hover idiom (opacity 0.7 -> 1) is now one
  global .icon-btn class in pagerite.css, shared by the edit pens,
  auth links, editor icon buttons and the new UA link
2026-09-09 00:55:44 +00:00
LeoVasanko 4eea3ae3af analytics: parse UAs at display time, never store ua_pretty
The stored ua_pretty froze each record at the uarite version of its
record time (old records showed disguised Meta crawlers as
"Chrome/145 Windows").  Client now carries a display-time-only
"uarite" field holding the full uarite.UA dataclass (pretty/engine/
os/provider/kind/url), filled when the viewer payload is built, so old
data always follows the current uarite.  uarite 0.2.0 fixes the Meta
misdetection itself; _compact_user_agent is gone, tracking.py calls
uaparse directly.
2026-09-09 00:11:41 +00:00
LeoVasanko 69583fb9fe Replace ua-parser with uarite for UA formatting and bot detection 2026-09-05 06:15:43 +00:00
LeoVasanko b57b7060ec Keep container fence lines out of prose chunks
A closing ::: glued to a paragraph (no blank line before it) rode inside
the prose chunk and crossed to the translator as part of the text run;
when the model dropped it, validation passed and the splice lost the
fence — the rest of the page rendered inside the container (seen in the
Spanish translation). Container fence lines (::: openers and closers
alike) are now always their own prose-free chunk, never reaching the
translator. Affected pages re-chunk on next save and re-translate under
the new hashes, repairing themselves.
2026-09-04 19:04:24 +00:00
LeoVasanko 3f27a0a292 Fix unstyled editor language selector, order all language menus logically
LangSelect's scoped CSS landed on the shared store chunk, whose stylesheet
the editor never loaded (only the public selector path injected it), so the
editor's language selector rendered unstyled on untranslated pages. Collect
editor stylesheets from the entry's imported chunks too (same traversal as
the langselect assets).

Also: hreflang alternates now skip languages disabled site-wide, and all
selectors (page editor, structure tab, public selector) order languages the
same way — primary first, then the lang tab's geographic grouping.
2026-09-04 18:52:22 +00:00
LeoVasanko b30d909a23 Don't extract GeoIP .mmdb.gz on filesystem, only in RAM. 2026-09-04 18:30:55 +00:00
LeoVasanko 13fecd2118 Hreflang alternates on category placeholder pages too 2026-09-04 18:12:48 +00:00
LeoVasanko 11a138e19f Translate category-label titles, not just page titles 2026-09-04 18:09:40 +00:00
LeoVasanko ebd5911a38 Store DBIP database in folder where the program is ran, not where it is installed. 2026-09-04 18:04:30 +00:00
LeoVasanko db57125953 Public language selector, linked with the editor's language selection. Shared popup menu operation with edit toolbar with consistent closing logic. 2026-09-04 17:52:51 +00:00
LeoVasanko 986e28c220 Serve /favicon.ico as a redirect to the configured site icon 2026-09-04 14:59:38 +00:00
LeoVasanko 33a4a76364 Skip npm's security audit on install (registry endpoint stalls for minutes) 2026-09-04 02:03:10 +00:00
LeoVasanko 9fb4b5a681 Reject translations that would splice block-level Markdown 2026-09-04 01:43:29 +00:00
LeoVasanko 0c1349b037 Keep section anchors in the original language on translated pages 2026-09-03 23:59:10 +00:00
LeoVasanko 54f8c8e09b Carry literal < through translation as fullwidth < 2026-09-03 23:52:39 +00:00
LeoVasanko 78f4ddb2f0 Dispatch translations titles-first across languages 2026-09-03 23:03:47 +00:00
LeoVasanko e6a0c57446 Mirror content layout with text direction; admin stays LTR 2026-09-03 20:01:23 +00:00
LeoVasanko d6d07db2c5 Access-log extras: remote-user on /_api, visitor info on /_ws 2026-09-03 19:38:15 +00:00
LeoVasanko 38af57218a Translator API key management. 2026-09-03 19:23:56 +00:00
LeoVasanko eb2e8f8273 Raw access-log analytics storage with display-time classification
Replace the pre-classified store (visits/crawlers/abuse lists written at
collection time, plus abuse_ips and in-memory pending/session tables) with
a raw append-only log: one Get record per document GET (full path, true
HTTP status, referer origin, preload flag) and one Msg per /_ws activity
message.  Visitor/crawler/abuse classification, visit grouping (30-minute
inactivity gap), status/referer/UTM attribution and all aggregates are
derived in Store.display(), so future rule changes never invalidate stored
data.  The viewer payload keeps its exact shape.

Fixes structurally:

- Abuser 404s on slug-format paths showed up as "articles read": the
  not-found branch recorded the request twice through separate status
  plumbing.  Each request is now recorded once with its true status.
- 404 trail links never rendered red: cache-served navigations issue no
  GET and the only real GET (the idle preload) was discarded before status
  recording.  Preloads are now recorded with pre=True, never counted, and
  used for status attribution.
- formatAbuseRows merged a path's 404 probes and 200 reads into one entry;
  the collapse is now keyed by (path, status class).

Rule improvements enabled by the redesign:

- The plain-404 abuse threshold counts within a 1-hour sliding window, so
  long-time readers accumulating misses never classify (scanners spray).
- Hidden (admin) clients never trigger abuse classification: editing means
  visiting not-found pages.
- /.well-known/ probes (RFC 8615, e.g. Chrome devtools) are never abuse
  evidence; //foo-style empty path segments are an instant telltale.
- Visits can no longer open on an external exit URL; favicon fetches skip
  hidden clients' referers/exits.

Legacy analytics.json files are set aside as .bak-legacy on startup.
2026-09-03 18:52:16 +00:00
LeoVasanko 792b9e7aa9 fastapi-vue-setup 1.4.2 upgrade 2026-09-03 17:40:55 +00:00
LeoVasanko 4c6ab3dde6 Ruff format 2026-09-03 17:40:51 +00:00
LeoVasanko 9d70f17587 Pass CLI config to the app as JSON in PAGERITE_CONFIG 2026-09-03 16:14:59 +00:00
LeoVasanko 029bfe105e Show public https URL in startup box for non-localhost sites 2026-09-03 16:09:08 +00:00
LeoVasanko 62031fd5dd Move DB-IP download into app lifespan with proper logging 2026-09-03 16:08:15 +00:00
LeoVasanko 13cf716bb1 Silence httpx request logs, log favicon fetches in one line 2026-09-03 16:03:55 +00:00
55 changed files with 2671 additions and 1544 deletions
+4 -2
View File
@@ -25,10 +25,12 @@ Pagerite is a CMS. See `docs` for the full design and implementation details. Ke
- `markdown.py` — markdown-it-py renderer. - `markdown.py` — markdown-it-py renderer.
- `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`). - `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`).
- `seed.py` — demo content, written only on first database creation. - `seed.py` — demo content, written only on first database creation.
- `analytics.py` — visit analytics collection (see `docs/analytics.md`). - `analytics.py` — visit analytics collection (see `docs/analytics.md`). UA formatting/bot detection comes from the **uarite** package.
- `frontend/src/` — Vue editor and public-page JS entries. - `frontend/src/` — Vue editor and public-page JS entries.
- `main.js` — Vue editor app entry. - `main.js` — Vue editor app entry.
- `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`). - `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`).
- `langselect-main.js` + `LangSelector.vue` — public language selector, imported on demand by pagerite.js on pages with more than one hreflang alternate (the editors' `LangSelect` flag dropdown).
- `store.js` — the shared Pinia store (`useStore`, id `pagerite`) for cross-bundle UI state.
- `pagerite.js` — public page entry. - `pagerite.js` — public page entry.
- `editorLang.js` + `LangSelect.vue` — the editor shell's shared language selection and its selector component (page + structure tabs; drives the page preview while the panel is open, via `swapdoc.setLangOverride`). - `editorLang.js` + `LangSelect.vue` — the editor shell's shared language selection and its selector component (page + structure tabs; drives the page preview while the panel is open, via `swapdoc.setLangOverride`).
- `reconnect.js` — shared WebSocket pacing for all sockets (staggered connect slots, stuck-CONNECTING watchdog, exponential backoff): bursts and rapid retries trip the browser's WebSocket throttling. - `reconnect.js` — shared WebSocket pacing for all sockets (staggered connect slots, stuck-CONNECTING watchdog, exponential backoff): bursts and rapid retries trip the browser's WebSocket throttling.
@@ -64,6 +66,6 @@ Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed).
## Conventions ## Conventions
- Keep dependencies minimal; add via `uv add` and mention it. - Keep dependencies minimal; add via `uv add` and mention it.
- The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content. - The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` (backend redirect to the configured site icon). Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
- No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. The one keyed exception is `/_translate/{key}` (translator service; `Data.translate_keys`, see docs/localization.md). - No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. The one keyed exception is `/_translate/{key}` (translator service; `Data.translate_keys`, see docs/localization.md).
- Update the relevant MarkDown files when architecture, tooling, or conventions change. - Update the relevant MarkDown files when architecture, tooling, or conventions change.
+232 -185
View File
@@ -1,15 +1,15 @@
# Analytics # Analytics
Server-side visit analytics. Data lives in a plain JSON file — a msgspec Server-side visit analytics built on a **raw access-log-style event store**.
Struct dumped to disk — separate from the kanta content database, path from Data lives in a plain JSON file — a msgspec Struct dumped to disk — separate
`PAGERITE_ANALYTICS` (default: `analytics.json` in the per-site data from the kanta content database, path from `PAGERITE_ANALYTICS` (default:
directory, e.g. `localhost/analytics.json`). `analytics.json` in the per-site data directory, e.g. `localhost/analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`, - `pagerite/analytics.py` — data model (`Analytics`, `Get`, `Msg`, `Client`,
`CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map, `Favicon`), the `Store` (raw log + atomic JSON persistence) and
atomic JSON persistence). `Store.display()`, where **all** classification happens.
- `pagerite/pages.py`entry-referer stashing in `show_page` (`_track_entry`, - `pagerite/pages.py`records every served document as one raw GET line
in `pagerite/tracking.py`), 404 recording. (`_record_get`, in `pagerite/tracking.py`) with its true HTTP status.
- `pagerite/tracking.py` — the `/_ws` activity WebSocket, and - `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
`WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint). `WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen. - `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
@@ -18,48 +18,116 @@ directory, e.g. `localhost/analytics.json`).
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView` - `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
into `#analytics-app` inside `#main`. into `#analytics-app` inside `#main`.
## What is collected ## Raw records
The store is deliberately close to an access log: two append-only lists plus
shared metadata. **Nothing is classified when recorded** — whether a client
turns out to be a reader, a crawler or a scanner is decided by
`Store.display()` from the raw events, so the stored data survives any future
change to the classification rules.
Each `Get` record (one per served document):
- `t` — timestamp of the request,
- `path` — full request path, query string included (e.g. `/.env?x=1`),
- `status` — the true HTTP status of the response (200, or 404 for a category
placeholder or a missing page),
- `ref` — external https origin of the `Referer`, `""` for direct/internal
(same-origin referers are dropped by the recorder),
- `pre` — true for idle-time link preloads from pagerite.js
(`x-pagerite-preload` header): never counted as a view, crawler hit or
abuse — recorded only so a navigation later served from the in-memory page
cache (which issues no GET at all) can be attributed this GET's status,
- `lang` — rendered content language of the served document (the resolved
language of a localized page), `""` for non-localized responses (404
probes, reserved paths),
- `client` — 6-byte blake3 hash referencing `Analytics.clients`.
304 revalidation responses return before recording and are not logged.
Each `Msg` record (one per pagerite.js activity message over `/_ws`):
- `t` — timestamp,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `fr` — path of the page the activity happened on (`""` for the initial
load),
- `to` — navigation target (validated at record time: internal slug path or
external https URL; anything else is dropped — sanitation, not
classification),
- `read` — active seconds spent on `fr` since the previous report,
- `lang` — rendered language reported by the client for the page the
activity happened on (the page's `<html lang>`; `""` from old clients).
Each `Client` record (shared by every event, keyed by hash):
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `hide` — true for admin clients (`hide` message field): everything this
client ever did is recorded but excluded from every statistic and from the
viewer payload. This is the one flag set at record time — it is a client
property, not a classification.
The viewer payload adds one display-time field to each client, never
persisted (stored records keep the default and old data always follows the
current uarite version):
- `uarite` — the `uarite.UA` dataclass from parsing the raw UA
(`pretty`/`engine`/`os`/`provider`/`kind`/`url`): the crawler name for
bots,
with a category suffix only where a provider runs crawlers of more than
one kind (`GPTBot (AI)` vs `OAI-SearchBot (search)`, `Googlebot (search)`
vs `Google-Extended (AI)`; single-kind providers stay plain: `Facebook`,
`WhatsApp`), `Browser/major OS` on the desktop, the device where that is
the relevant information (iPhone reports its iOS version, Android phones
their model instead of the OS), otherwise the raw string; `url` is the
crawler's info page when uarite knows one (rendered as a 🔗 link after the
pretty UA in the viewer), `kind` drives the bot classification.
A reverse-DNS lookup is attempted for each new client and the result, when
available, is stored as `host`; local/reserved/multicast addresses are
skipped. If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present
in the working directory, it is loaded at startup and used to look up
`country`/`city`. These lookups run in background tasks after the event is
stored, so WebSocket message handling is never delayed. Only the downloaded
`.mmdb.gz` is kept on disk (in the working directory, ignored by git); it is
decompressed into RAM when opened. The
CLI flag `--dbip` (`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP at startup (in the app lifespan,
before the MMDB is opened), skipping the download when the local database is
already current and removing older versions after an update; without the flag
only an existing file is used.
## What the client sends
The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
whole browsing session and sends activity messages over it — JSON text whole browsing session and sends activity messages over it — JSON text
frames matching the server's `Ping` msgspec struct with the fields `fr` frames matching the server's `Ping` msgspec struct with the fields `fr`
(source path), `to` (navigation target), `read` (active seconds on `fr` (source path), `to` (navigation target), `read` (active seconds on `fr`
since the last report) and `hide`; falsy fields are omitted. One channel since the last report), `lang` (the rendered language of the page the
activity happened on — its `<html lang>`) and `hide`; falsy fields are
omitted. One channel
follows the session, so the activity of a visit stays tied together, and follows the session, so the activity of a visit stays tied together, and
while the user is active the accumulated reading time is flushed every few while the user is active the accumulated reading time is flushed every few
seconds: the trail times are cumulative, so a disconnection simply leaves seconds: the times are incremental, so a disconnection simply leaves the
the last reported time in place (no close beacon). After 5 minutes without last reported time in place (no close beacon). After 5 minutes without
any activity the client closes the socket itself — a sleeping browser tab any activity the client closes the socket itself — a sleeping browser tab
would lose it anyway — and the next activity reconnects as a fresh session; would lose it anyway — and the next activity reconnects; reconnects are
reconnects are attempted only on user activity, with an exponential backoff attempted only on user activity, with an exponential backoff between
between attempts so a failing endpoint is never hammered. Idle-time link preloads attempts so a failing endpoint is never hammered. Idle-time link preloads
stay plain `fetch()` calls so the browser may cache the responses; the stay plain `fetch()` calls so the browser may cache the responses; the
WebSocket reports actual navigations and active time spent on a page. WebSocket reports actual navigations and active time spent on a page.
- **Initial page load**: only `to` — the loaded path — is sent, never `fr` - **Initial page load**: only `to` — the loaded path — is sent, never `fr`
(an `fr` equal to `to` would log a bogus self-transition when a session (an `fr` equal to `to` would log a bogus self-transition when a session
already exists, e.g. a second tab). This message is what starts already exists, e.g. a second tab). Reloads are not
the visit and counts the entry page view — the document GET alone records
nothing, so bots never register (admin browsing does register, but
flagged `hide`; see **Admins** below). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do connect and report, but their
User-Agent gives them away: messages whose UA matches `_is_bot_ua`
(anything calling
itself a "bot", plus known exceptions such as GoogleOther) are ignored
server-side, and their document GETs land in the crawler list instead.
Real-browser bots whose UA does not match still register a visit, but
their reported reading time stays under 5 seconds, so they are
reclassified as crawler hits at display time (see **Crawler hits** below).
No source-IP verification is done: a spoofed bot UA merely lands in the
crawler stats, and scanners that probe telltale paths are caught by the
abuse rules regardless. Reloads are not
visits: the message is skipped (PerformanceNavigationTiming `reload`), so a visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
refresh neither counts a second view nor logs a self-transition. The GET refresh neither counts a second view nor logs a self-transition.
handler stashes a cross-origin https `Referer` (origin part only —
unavailable to JS once the page has loaded) and any
`utm_*` query parameters in in-memory IP tables, consumed by the first
message that
starts the visit; internal or absent referers never touch the referer table.
- **Internal fetch-navigations**: `to` is the target path, sent only after - **Internal fetch-navigations**: `to` is the target path, sent only after
the swap actually happened (a failed swap falls back to a full load, the swap actually happened (a failed swap falls back to a full load,
whose initial message counts the view instead — no gap, no double count). whose initial message counts the view instead — no gap, no double count).
@@ -71,152 +139,135 @@ WebSocket reports actual navigations and active time spent on a page.
- **Excluded**: back/forward (popstate) navigations, navigating *to* the - **Excluded**: back/forward (popstate) navigations, navigating *to* the
analytics page (`/_a` — its GET is untracked, and the server cannot analytics page (`/_a` — its GET is untracked, and the server cannot
record it as a navigation target anyway), and everything while the user has record it as a navigation target anyway), and everything while the user has
the editor the editor open (`body.editing`). Admin noise, not visits. Navigating
open (`body.editing`). Admin noise, not visits. Navigating *away* from *away* from `/_a` does report.
`/_a` does report: the fetch-navigation already GET-ed the target page
without the preload header, and without the message that GET would flush to
the crawler list.
- **Admins**: when SSO is in use and the session is known to be an admin, - **Admins**: when SSO is in use and the session is known to be an admin,
the client still reports but adds `hide`. The activity is recorded as the client still reports but adds `hide`. The activity is recorded as
usual (navigations and all), but the `hide` flag is set on the **client usual (navigations and all), but the `hide` flag is set on the **client
record** — so it covers everything that client ever did: visits and record** — so it covers everything that client ever did, including the
crawler hits from before the login included. Hidden clients never appear time before the login. Hidden clients never appear in the viewer payload:
in the viewer payload: `Store.display()` drops their visits, crawler `Store.display()` drops their events and metadata, and computes every
hits, abuse hits and metadata, and computes every aggregate (site visits, aggregate (site visits, page views, transitions) from the visible visits
page views, transitions) from the visible visits only, so nothing needs only, so nothing needs to be reversed or redacted. With no auth proxy
to be reversed or redacted. Pending crawler hits from a hidden client
are discarded when they expire, so admin browsing never lands in the
crawler list either. With no auth proxy
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything (dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
is recorded. is recorded.
- The server validates `to`: internal paths must be valid slug paths - **External-site favicons**: for every external https origin seen as a GET
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the referer or an exit link, the server fetches `{origin}/favicon.ico` in a
https origin and accepted only when the client sent exactly that.
- **External-site favicons**: for every external https origin seen as a visit
referer, a crawler-hit referer or an exit link, the server fetches `{origin}/favicon.ico` in a
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only — background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
SVG is sniffed from the body when served without an image type) and stores SVG is sniffed from the body when served without an image type) and stores
the icon content-hashed on disk in the FileStore (served at `/_f/{name}`, the icon content-hashed on disk in the FileStore (served at `/_f/{name}`,
extension matching the actual MIME). The origin → file name mapping is extension matching the actual MIME). The origin → file name mapping is
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
recorded too and retried only after 7 days. Fetches are scheduled after recorded too and retried only after 7 days. Fetches are scheduled after
each activity message and once at startup, which backfills icons for already-recorded each activity message and once at startup, which backfills icons for
data. The viewer payload carries `favicons` (origin → `/_f/...` path), already-recorded data. The viewer payload carries `favicons` (origin →
and the viewer shows the icon wherever an external site is mentioned: `/_f/...` path), and the viewer shows the icon wherever an external site
referer/exit trail links in the visit table and the source/exit pills of is mentioned: referer/exit trail links in the visit table and the
the transition map (UTM-attributed source nodes without an https origin source/exit pills of the transition map (UTM-attributed source nodes
stay text-only). without an https origin stay text-only).
- **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
`User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
the first 6 bytes identify a shared `Client` record. The `Client` stores
the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
`country` from the language-region subtag, and asynchronously-filled
`country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
crawler hits and abuse hits all reference this record by its hash, so
client metadata is stored once instead of repeated per event.
- The visitor IP is stored in the `Client`. A reverse-DNS lookup is
attempted for each new client and the result, when available, is stored as
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
root, it is loaded at startup and used to look up `country`/`city`. These
lookups run in background tasks after the event is stored, so WebSocket
message handling is never delayed. The decompressed `dbip-*.mmdb` file is kept in
the repository root and ignored by git. The CLI flag `--dbip`
(`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
skipping the download when the local database is already current and
removing older versions after an update; without the flag only an existing
file is used.
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
hit — except idle-time link preloads from pagerite.js, which carry an
`x-pagerite-preload` header and are not tracked at all (the navigation
message sent when the user actually navigates to a preloaded page does
the counting; forging
the header only hides a GET from the crawler stats, the path-based abuse
classification is unaffected). If a message
from the same client arrives within 10 seconds the hit is discarded;
otherwise it is written to `crawlers` — unless the client is hidden
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as
visits or views. Bots running real browsers can still slip past the UA
check: a visit whose total reported reading time stays under 5 seconds
(`_MIN_VISIT_READ`; durations are client-provided and trusted — such bots
report 02 s) is reclassified as crawler hits at display time, one hit
per internal trail page, and counts in no visit aggregate. The
`Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited, preceded by its referer
when there is one — spiders often advertise their own site as the
referer, and it is rendered with its favicon like visit referers (crawler
referers are included in the favicon fetch origins). The crawler table lists
the most recent crawler first, with the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
classifies the source IP as abuse immediately, and ten plain 404s from one
IP do too. Classification reclassifies history: all earlier crawler hits
from that IP (persisted and pending) move to the `abuse` list, so a
random-UA scanner no longer pollutes the crawler stats of the legitimate
bot it impersonates. Once classified, every document GET and 404 from the
IP is recorded as an abuse hit with the full request path (query string
included), and its activity messages are ignored. The classified IP set (`abuse_ips`)
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
viewer, abuse hits are grouped by IP (never by client/UA — scanners
randomize theirs) in a separate "Abuse" table. Identical paths are
collapsed into one entry with their hit count. The 404 probes ("paths
abused": flagged paths that triggered classification first, then other
404s) are kept in a separate column from the real articles the abuser
actually read ("articles read": document GETs that returned 200, not the
404 fallback rendering — rendered as trail links like the visitor and
crawler tables, with the query string stripped). Raw User-Agent strings are shown one
per line with their occurrence counts, and the full lists are click-to-copy.
## Visits and sessions ## Display-time classification
There are no cookies. A visit is tied together by a client hash — the first `Store.display(in_menu)` derives the viewer payload from the raw events on
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6 every (debounced) broadcast — O(n log n) over the log, cheap enough for a
/64 network), the raw `User-Agent` string and the extracted small CMS. `in_menu(path)` resolves a path against the current menu (passed
`Accept-Language` tag. The first message from a client hash starts a new in from `tracking.py`, which owns the content database import) so 404
visit; subsequent messages extend it. Messages arriving with no known session responses for real menu nodes — category placeholders — are not mistaken
(server restart) start a fresh visit from the first message — treated as for misses.
missing data rather than dropped. The client-hash → visit map and the IP →
entry-referer/UTM tables are in-memory only; client metadata is stored in
`Analytics.clients` keyed by the client hash.
Each `Client` record: - **Visits and sessions**: a client's messages are grouped into visits
chronologically; a new visit starts after 30 minutes of inactivity
(`_SESSION_GAP`). A fresh page load with an already-open visit (second
tab) extends it, logging a `(direct)` transition. The visit's trail holds
first-seen targets in order; `read` updates accumulate active seconds on
the trail item matching `fr` (preferring the item whose language matches
the report, so seconds after a language switch land on the new-language
step). Each trail item's HTTP status comes from
the client's latest GET for that path — preloads included, which is what
allows 404 pages to render red in the viewer even when the navigation
itself was served from the page cache. Each trail item also carries the
rendered language: the client's report, for the entry page falling back
to its GET's rendered language (old clients don't send one); a page
re-visited in a different language becomes a distinct trail step instead
of merging into the existing item. The entry page's referer and
`utm_*` tags come from the GET that loaded it (within 10 s before the
first message).
- **Crawler hits**: a document GET no activity message matched within
`_CRAWLER_TIMEOUT` (10 s) is a crawler hit — plain bots that only fetch
documents never register as visits. JS-running crawlers (Googlebot,
GoogleOther, Applebot, ...) do connect and send messages, but their UA
gives them away (`_is_bot_ua`, backed by `uarite.uaparse` — which
also knows the disguised ones: facebookexternalhit, Google-Extended,
WhatsApp, ...): their messages are ignored at display
time, so their GETs never match and land in the crawler list too. Real-
browser bots whose UA does not match are caught by engagement: a visit
whose total reported reading time is under 5 seconds (`_MIN_VISIT_READ`;
durations are client-provided and trusted — such bots report 02 s) is
reclassified as crawler hits, one per internal trail page, and counts in
no visit aggregate. No source-IP verification is done: a spoofed bot UA
merely lands in the crawler stats, and scanners that probe telltale paths
are caught by the abuse rules regardless. In the viewer, crawler hits are
grouped by client hash and shown as a trail of pages, preceded by the
referer when there is one (rendered with its favicon like visit
referers). The crawler table lists the most recent crawler first, with
the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 on a telltale path — an empty URL segment
(`//foo` — no real client generates those), any segment starting with a
dot (`/.env`, `/.git/config`) or ending in `.php` — classifies the source
IP as abuse, and ten plain 404s within one hour (`_ABUSE_404_WINDOW`) on
paths that don't resolve to a menu node do too. Two exemptions keep
legitimate traffic out: RFC 8615 well-known URIs (`/.well-known/…`
browsers and services probe them, e.g. Chrome's devtools fetch of
`appspecific/com.chrome.devtools.json`) are never telltale and never
count toward the threshold, and category placeholders return 404 but are
real menu nodes, so they never count either. The window keeps a
long-time reader's slowly accumulating misses from ever crossing the
threshold — scanners spray in bursts. Hidden (admin) clients never
trigger classification: editing means visiting not-found pages, since
that is where the create pen lives. Once an IP is classified, **all** its document GETs are shown in the abuse list —
including any that arrived before classification, since the raw log keeps
everything — and its activity messages are ignored. In the viewer, abuse
hits are grouped by IP (never by client/UA — scanners randomize theirs)
in a separate "Abuse" table, split by the recorded status: the 404 probes
("paths abused" — flagged paths that triggered classification first, then
other 404s, shown verbatim with query strings) versus the real articles
the abuser actually read ("articles read" — the 200 document GETs,
rendered as trail links like the visitor and crawler tables, query string
stripped). Raw User-Agent strings are shown one per line with their
occurrence counts, and the full lists are click-to-copy.
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer), In the visitor and crawler tables, internal paths that returned a 404 status
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`, are shown in red and the link title includes the status code, so it is easy
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`), to tell misses from real pages at a glance.
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string,
- `hide` — true for admin clients (`hide` message field): all their visits,
crawler hits and abuse hits are recorded but excluded from every
statistic and from the viewer payload.
Each `Visit` record: ## Derived shapes (the viewer payload)
- `start` — timestamp of the first event, The `Display` payload contains the derived `visits`, `crawlers` and `abuse`
rows (structs `Visit`/`Nav`/`TrailItem`, `CrawlerHit`, `AbuseHit` — display
DTOs only, never persisted), the visible `clients`, the fetched `favicons`,
the site language context (`multilingual` — translation languages are
configured, so the viewer can suppress language UI on single-language
sites — and `primary_lang` — the front page's primary language, so the
viewer can skip the primary-language default case),
and the aggregates below.
Each derived `Visit`:
- `start` — timestamp of the first activity,
- `entry` — first page (path) seen, - `entry` — first page (path) seen,
- `referer` — external https origin of the initial load, `""` for direct, - `referer` — external https origin of the entry GET, `""` for direct,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `trail` — the entry page and everything seen afterwards, keyed by the - `trail` — the entry page and everything seen afterwards, keyed by the
timestamp of first sight (insertion order = first-seen order). Each item timestamp of first sight (insertion order = first-seen order). Each item
holds `to` (page path or external exit URL), the accumulated active holds `to` (page path or external exit URL), the accumulated active
reading time in seconds (`read`) and the most recent HTTP status seen reading time in seconds (`read`), the most recent HTTP status seen
for the target (`status`). Re-visiting an already seen target updates for the target (`status`) and the rendered language (`lang`; a page
its item instead of appending. seen in two languages within one visit gets one item per language),
- `navs` — every navigation message (`fr`, `to`), keyed by its timestamp, - `navs` — every navigation (`fr`, `to`), keyed by its timestamp, repeats
repeats included. The aggregates are computed from this log at display included. The aggregates are computed from this log,
time.
- `utm``utm_*` query parameters from the landing URL, as a dict. - `utm``utm_*` query parameters from the landing URL, as a dict.
Each `CrawlerHit` record: Each derived `CrawlerHit`:
- `start` — timestamp of the document GET, - `start` — timestamp of the document GET,
- `entry` — page path requested, - `entry` — page path requested,
@@ -224,41 +275,34 @@ Each `CrawlerHit` record:
- `referer` — external https origin of the request, `""` for direct/none, - `referer` — external https origin of the request, `""` for direct/none,
- `query` — raw query string of the request, - `query` — raw query string of the request,
- `status` — HTTP status of the served response (200 for a real page, 404 - `status` — HTTP status of the served response (200 for a real page, 404
for a category placeholder or missing page). for a category placeholder or missing page),
- `lang` — rendered content language of the served document (from the GET).
Each `AbuseHit` record: Each derived `AbuseHit`:
- `start` — timestamp of the request, - `start` — timestamp of the request,
- `path` — full request path including the query string (e.g. `/.env?x=1`), - `path` — full request path including the query string,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the path that triggered abuse classification (telltale - `flag` — true for the paths that triggered abuse classification (telltale
path or the 404 that crossed the threshold), paths, or the 404 that crossed the threshold),
- `is_404` — true for 404 responses (probed paths and 404-fallback document - `is_404` — true for 404 responses, false for real (200) document GETs.
GETs), false for real (200) document GETs — articles the abuser read.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical paths are collapsed with their counts, split into the Abuse table identical requests (same path and status class) are collapsed
404 probes (flagged paths that triggered classification first, then other with their counts — a path's 404 probes and its later 200 reads never
404s, shown verbatim) and the 200 document GETs shown as trail links in the merge. Within each list paths are sorted by count descending, then by their
separate articles column. earliest hit.
Within each list paths are
sorted by count descending, then by their earliest hit.
In the visitor and crawler tables, internal paths that returned a 404 status
are shown in red and the link title includes the status code, so it is easy
to tell misses from real pages at a glance.
## Aggregates ## Aggregates
Aggregates are **not stored**; they are computed at display time by Aggregates are **not stored**; they are computed at display time by
`Store.display()` from the visit records (entry + `navs` log), skipping `Store.display()` from the derived visits (entry + `navs` log), skipping
hidden clients' visits and short visits reclassified as crawler hits hidden clients and short visits reclassified as crawler hits. This is
(under `_MIN_VISIT_READ` seconds of total reported reading time). This is what allows a client to become hidden after navigations were already
what allows a client to become hidden after logged: no counts need reversing. The computed shapes, part of the
navigations were already logged: no counts need reversing. The computed WebSocket payload (`Display` struct alongside `visits`, `crawlers`, `abuse`
shapes, part of the WebSocket payload (`Display` struct alongside `visits`, and `clients`):
`crawlers`, `abuse` and `clients`):
- `transitions`: time series of page transitions, sparse nested dict - `transitions`: time series of page transitions, sparse nested dict
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the `from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
@@ -272,13 +316,16 @@ shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
5-minute bucketing. 5-minute bucketing.
Sparseness keeps quiet sites small; dropping old data is a matter of deleting Sparseness keeps quiet sites small; dropping old data is a matter of deleting
list entries (`visits` is a plain append-only list). list entries (`gets`/`msgs` are plain append-only lists).
## Persistence ## Persistence
The whole `Analytics` struct is JSON-encoded and written atomically The whole `Analytics` struct is JSON-encoded and written atomically
(temp file + rename) on every recorded event. Traffic on a small CMS makes (temp file + rename) on every recorded event. Traffic on a small CMS makes
this cheap enough; batching can be added later without changing the format. this cheap enough; batching can be added later without changing the format.
A file written by the pre-redesign schema (stored `visits`/`crawlers`/`abuse`
lists) is not convertible; it is renamed to `analytics.json.bak-legacy` and
recording starts fresh.
## Viewing ## Viewing
+2 -2
View File
@@ -19,8 +19,8 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
- Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.). - Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
- **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above. - **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's left (the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation. - Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's start edge (left in LTR, right in RTL — the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/… "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone left of the text on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section. - **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/… "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column, to its end/start edge following the text direction (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone at the text's start edge on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
## Page structure and navigation ## Page structure and navigation
+2 -1
View File
@@ -26,6 +26,7 @@ Vite builds ES-module `.js` outputs; in dev the backend links them as `<script t
All site data lives under `<hostname>/` in the cwd — `content.kantadb`, All site data lives under `<hostname>/` in the cwd — `content.kantadb`,
`analytics.json` and `files/` — where `<hostname>` is the CLI's first `analytics.json` and `files/` — where `<hostname>` is the CLI's first
positional argument (default `localhost`, exported as `PAGERITE_HOSTNAME`; positional argument (default `localhost`, passed to the app as JSON in
`PAGERITE_CONFIG`, see `pagerite/config.py`;
`PAGERITE_DB`/`PAGERITE_ANALYTICS`/`PAGERITE_FILES` override individual `PAGERITE_DB`/`PAGERITE_ANALYTICS`/`PAGERITE_FILES` override individual
paths). gitignored. Do not delete it without asking. paths). gitignored. Do not delete it without asking.
+54 -20
View File
@@ -53,11 +53,13 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
article's own language), `?lang=xx` when serving a translation — however article's own language), `?lang=xx` when serving a translation — however
the language was arrived at (query or header). the language was arrived at (query or header).
- `<link rel="alternate" hreflang="…">` entries follow the canonical - `<link rel="alternate" hreflang="…">` entries follow the canonical
directly (before the social meta tags) and are the same set on every directly (before the social meta tags) and list the languages the page
page — the site-wide configured languages (`translate_langs`, which the is **actually available in**: `x-default` first, pointing at the plain
translator works to fill in): `x-default` first, pointing at the plain autodetecting URL, then every available language — the original again by
autodetecting URL, then every language explicitly with `?lang=`, the its plain URL, translations by `?lang=`. The public language selector
page's own primary language included. keys off these: pagerite.js mounts the editors' flag dropdown in the
top-right corner when the head advertises x-default plus more than one
language, loading its bundle (Vue + the flag SVG set) on demand.
- The override sticks for the session of clicks: a page requested with - The override sticks for the session of clicks: a page requested with
`?lang=` replicates the query onto the navigation links it renders (nav, `?lang=` replicates the query onto the navigation links it renders (nav,
sidebar, cards, brand — in-article links are content and stay as sidebar, cards, brand — in-article links are content and stay as
@@ -66,6 +68,11 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
`history.replaceState` (pretty, shareable URLs), remembers the language, `history.replaceState` (pretty, shareable URLs), remembers the language,
and adds it to every internal fetch that lacks one (preloads, and adds it to every internal fetch that lacks one (preloads,
fetch-navigations, history traversals); history entries stay query-less. fetch-navigations, history traversals); history entries stay query-less.
- The public selector's pick is the same override, pure JS state
(`pagerite:set-session-lang`): the session language changes and the page
swaps in place — no `?lang=` in the address bar, no reload. The choice is
linked with the editor panel's language dropdown both ways; closing the
panel keeps the chosen language instead of reverting.
- A full page refresh or a shared link resets to automatic selection (header - A full page refresh or a shared link resets to automatic selection (header
only). This gives a clean one-time override without cookies. only). This gives a clean one-time override without cookies.
@@ -88,6 +95,10 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
### Rendering ### Rendering
- The translated Markdown goes through the same `markdown.render` pipeline. - The translated Markdown goes through the same `markdown.render` pipeline.
- Section anchors (`#hash` ids on h1/h2 headings) stay in the original
language: render(anchors_from=...) pins the translated render's heading
ids to the original text's slugs, matched by heading position, so links
to sections don't break across languages.
- Navigation/sidebar titles come from the translation's title map, with - Navigation/sidebar titles come from the translation's title map, with
per-node fallback to the original title (a partially translated tree must per-node fallback to the original title (a partially translated tree must
still render). still render).
@@ -95,7 +106,9 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
language like content pages, but over the **subtree's** combined language like content pages, but over the **subtree's** combined
availability (`subtree_languages`) — they have no chunks of their own; availability (`subtree_languages`) — they have no chunks of their own;
the heading, navigation and card text localize from the title map and the heading, navigation and card text localize from the title map and
the target articles' translations. the target articles' translations. Their hreflang alternates are
computed exactly like a content page's (a translated title counts as
availability, so the language selector is offered there too).
- Card descriptions and cover picks run on the target article's hybrid - Card descriptions and cover picks run on the target article's hybrid
Markdown where that page is available in the served language, with Markdown where that page is available in the served language, with
per-card fallback to the original. per-card fallback to the original.
@@ -133,7 +146,11 @@ served Markdown at render time.
`chunk_markdown(markdown)` splits the source into block-level chunks — `chunk_markdown(markdown)` splits the source into block-level chunks —
blank-line-separated blocks: headings, paragraphs, code fences (kept whole), blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
list blocks, tables, HTML blocks. A chunk's identity is its **source text**, list blocks, tables, HTML blocks. Container fence lines (`::: name` openers
and `:::` closers) are always their own chunk, blank lines or not — folded
into a prose chunk the closer would cross to the translator as part of the
text, where the model can drop it (the rest of the page then renders inside
the container). A chunk's identity is its **source text**,
gettext-msgid style: gettext-msgid style:
```python ```python
@@ -305,14 +322,15 @@ An external machine-translation service connects over WebSocket at
`/_translate/{key}` — deliberately **not** under `/_api`: the SSO `/_translate/{key}` — deliberately **not** under `/_api`: the SSO
forward-auth does not cover that route, and the key in the path is the forward-auth does not cover that route, and the key in the path is the
access control. Keys live in `Data.translate_keys` (key -> display name) — access control. Keys live in `Data.translate_keys` (key -> display name) —
12 lowercase alphanumeric characters each, the first one generated at 12 lowercase alphanumeric characters each; the first is generated at
database bootstrap and multiple keys reserved for future management (e.g. database bootstrap, further ones are managed in the editor's lang tab
a web UI). The full WS URL(s) are printed in the startup log (add/rename/delete ride the `PUT /_api/settings` round-trip; the name is
(`ws://localhost:{port}/_translate/{key}` locally, an inline display label only). The full WS URL(s) are printed in the
`wss://{hostname}/_translate/{key}` on a public hostname) and the keys are startup log (`ws://localhost:{port}/_translate/{key}` locally,
surfaced to the admin in `GET /_api/settings` as `translate_keys`. An `wss://{hostname}/_translate/{key}` on a public hostname) and shown in the
unknown or empty key rejects the handshake (close-before-accept → HTTP lang tab as click-to-copy links; the keys are also surfaced in
403). Transactions storing results record the connecting key as the kanta `GET /_api/settings` as `translate_keys`. An unknown or empty key rejects
the handshake (close-before-accept → HTTP 403). Transactions storing results record the connecting key as the kanta
transaction `user`. transaction `user`.
Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`; Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
@@ -397,9 +415,14 @@ verbatim source substring — entity-decoded text, backslash escapes — is
skipped and stays in the original language), and the returned translations skipped and stays in the original language), and the returned translations
are swapped in by offset. Markup corruption is therefore impossible by are swapped in by offset. Markup corruption is therefore impossible by
construction; the failure modes that remain are a wrong segment count, an construction; the failure modes that remain are a wrong segment count, an
empty segment, or markup injected INTO a segment (a `<br>` in a title empty segment, markup injected INTO a segment (a `<br>` in a title
translation would splice live HTML) — each returned segment must parse as translation would splice live HTML), or a line that would start a new
pure prose, or the whole result is dropped and logged, and the (lang, key) block where the segment lands (a ``` or ::: fence line would eat the rest
of the block it splices into, closing fence included — segments are
inline prose, so `pure_prose` alone cannot see this) — each returned
segment must parse as
pure prose with no block-starting line or blank line, or the whole result
is dropped and logged, and the (lang, key)
pair is skipped for the rest of the server run (generation is pair is skipped for the rest of the server run (generation is
near-deterministic, so an immediate retry would re-fail; the fragment stays near-deterministic, so an immediate retry would re-fail; the fragment stays
pending and gets another chance on restart or `DELETE /_api/translations`). pending and gets another chance on restart or `DELETE /_api/translations`).
@@ -450,14 +473,25 @@ stripped before the result goes back.
The same client-side enforcement covers markup bleed as a CLASS, not per The same client-side enforcement covers markup bleed as a CLASS, not per
artifact: `<` is the prose/markup boundary on the wire and never appears in artifact: `<` is the prose/markup boundary on the wire and never appears in
a segment in either direction. Source pieces containing `<` are never a segment in either direction. A literal `<` in the source text (`<1MB` is
dispatched (they stay in the original language — segments.py), and the text, not markup — a tag needs a letter or `/!?`) crosses encoded as the
fullwidth `` and is decoded on return, before the result is validated and
spliced (segments.py) — the wire itself still never carries `<`, and the
reference client cuts the model's output at the first `<` reference client cuts the model's output at the first `<`
(scripts/translator.py) — echoed language tags, stray `<br>`s and any (scripts/translator.py) — echoed language tags, stray `<br>`s and any
future variant are one handled case. (The cut is post-decode, not a future variant are one handled case. (The cut is post-decode, not a
generation stop string: Seed-X opens every generation with its `<s>` generation stop string: Seed-X opens every generation with its `<s>`
framing token, which would trip a `<` stop immediately.) framing token, which would trip a `<` stop immediately.)
Server-side, a second layer covers what the inline parser cannot: ASCII
punctuation that is plain prose on the wire but Markdown syntax in the
splice context — quotes (a translated `"` would close the quoted image
title it lands in), brackets (alt texts, re-inserted link texts), `|` in
table rows, `\` escapes. Rather than rejecting such results, `join` swaps
them for Unicode look-alikes before splicing (`_NEUTRAL` in
segments.py — curly quotes, fullwidth brackets; the renderer's
typographer curls straight quotes anyway).
Short fragments get more than a bare prompt: each segment may carry its Short fragments get more than a bare prompt: each segment may carry its
surround in `Job.contexts` — a title carries the article's opening prose surround in `Job.contexts` — a title carries the article's opening prose
(its own block is just the title word), a segment carved out of a larger (its own block is just the title word), a segment carved out of a larger
+3
View File
@@ -37,3 +37,6 @@ __screenshots__/
# Playwright browser downloads (if ever installed locally) # Playwright browser downloads (if ever installed locally)
.pw-browsers/ .pw-browsers/
# npm project config (audit/fund off: the audit endpoint stalls installs)
!.npmrc
+2
View File
@@ -0,0 +1,2 @@
audit=false
fund=false
+1
View File
@@ -20,6 +20,7 @@
"codemirror": "^6.0.2", "codemirror": "^6.0.2",
"country-flag-icons": "^1.6.20", "country-flag-icons": "^1.6.20",
"overlayscrollbars": "^2.16.0", "overlayscrollbars": "^2.16.0",
"pinia": "^4.0.3",
"transliteration": "^2.6.1", "transliteration": "^2.6.1",
"vue": "^3.5.26", "vue": "^3.5.26",
"vuedraggable": "^4.1.0" "vuedraggable": "^4.1.0"
+42 -31
View File
@@ -23,6 +23,7 @@ import {
formatVisitRows, formatVisitRows,
} from './analytics/format.js' } from './analytics/format.js'
import TrailLink from './TrailLink.vue' import TrailLink from './TrailLink.vue'
import RefererBadge from './RefererBadge.vue'
import VisitorCell from './VisitorCell.vue' import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue' import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue' import VisitorCharts from './VisitorCharts.vue'
@@ -160,15 +161,22 @@ watch(range, (r) => {
const clients = computed(() => data.value?.clients || {}) const clients = computed(() => data.value?.clients || {})
const favicons = computed(() => data.value?.favicons || {}) const favicons = computed(() => data.value?.favicons || {})
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value)) // Site language context from the payload: drives the discreet rendered-
// language markers in the visit/crawler rows (multilingual sites only).
const site = computed(() => ({
multilingual: !!data.value?.multilingual,
primaryLang: data.value?.primary_lang || '',
}))
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value, site.value))
const crawlers = computed(() => rangeData.value?.crawlers || []) const crawlers = computed(() => rangeData.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value)) const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value, site.value))
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, pageTree.value, now.value)) const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, pageTree.value, now.value))
</script> </script>
<template> <template>
<div class="analytics-view"> <!-- Untranslated admin dashboard: always LTR, like the editor panel. -->
<div class="analytics-view" lang="en" dir="ltr">
<div class="analytics-panel"> <div class="analytics-panel">
<header> <header>
<h1>Analytics</h1> <h1>Analytics</h1>
@@ -208,15 +216,16 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
<tbody> <tbody>
<tr v-for="(v, i) in visitRows" :key="i"> <tr v-for="(v, i) in visitRows" :key="i">
<td class="trail"> <td class="trail">
<TrailLink v-if="v.refererStep" :step="v.refererStep" :favicons="favicons" @close="$emit('close')" /> <RefererBadge v-if="v.refererBadge" :badge="v.refererBadge" :favicons="favicons" />
<span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span> <span v-if="v.rowFlag" class="flag" v-html="v.rowFlag" :title="v.rowFlagTitle"></span>
<TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" :favicons="favicons" @close="$emit('close')" /> <TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" :favicons="favicons" :flags="s.langFlags" @close="$emit('close')" />
</td> </td>
<VisitorCell <VisitorCell
:ip="v.ip" :ip="v.ip"
:ip-display="v.ipDisplay" :ip-display="v.ipDisplay"
:ua="v.ua" :ua="v.ua"
:ua-raw="v.uaRaw" :ua-raw="v.uaRaw"
:ua-url="v.uaUrl"
:country="v.country" :country="v.country"
:city="v.city" :city="v.city"
:lang="v.lang" :lang="v.lang"
@@ -244,14 +253,16 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
<tbody> <tbody>
<tr v-for="(c, i) in crawlerRows" :key="i"> <tr v-for="(c, i) in crawlerRows" :key="i">
<td class="trail"> <td class="trail">
<TrailLink v-if="c.refererStep" :step="c.refererStep" :favicons="favicons" @close="$emit('close')" /> <RefererBadge v-if="c.refererBadge" :badge="c.refererBadge" :favicons="favicons" />
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" /> <TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
<span v-for="(f, fi) in c.readFlags" :key="fi" class="flag" v-html="f.flag" :title="f.name"></span>
</td> </td>
<VisitorCell <VisitorCell
:ip="c.ip" :ip="c.ip"
:ip-display="c.ipDisplay" :ip-display="c.ipDisplay"
:ua="c.ua" :ua="c.ua"
:ua-raw="c.uaRaw" :ua-raw="c.uaRaw"
:ua-url="c.uaUrl"
:country="c.country" :country="c.country"
:city="c.city" :city="c.city"
:lang="c.lang" :lang="c.lang"
@@ -298,6 +309,8 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
:ip-display="a.ipDisplay" :ip-display="a.ipDisplay"
:ua="a.ua" :ua="a.ua"
:ua-raw="a.uaRaw" :ua-raw="a.uaRaw"
:ua-url="a.uaUrl"
:ua-raws="a.uaRaws"
:country="a.country" :country="a.country"
:city="a.city" :city="a.city"
:lang="a.lang" :lang="a.lang"
@@ -467,16 +480,30 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
color: var(--error, #c00); color: var(--error, #c00);
} }
.visit-table .utm-tag { /* The referer badge outgrows the 8rem trail-link cap (it carries the UTM
display: inline-block; summary too); keep the inline-flex layout from the component. */
.visit-table .trail a.referer-badge {
display: inline-flex;
max-width: 100%; max-width: 100%;
padding: 0.05rem 0.4rem; }
border: 1px solid var(--line);
border-radius: 0.25rem; /* Same flag chip as the visitor cells (VisitorCell.vue); the flags here
white-space: nowrap; mark the language the page was read in. */
.visit-table .flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden; overflow: hidden;
text-overflow: ellipsis; border: 1px solid var(--line);
vertical-align: bottom; box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.visit-table .flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
} }
.visit-table .clickable-list { .visit-table .clickable-list {
@@ -505,22 +532,6 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
.visit-table .clickable-list, .visit-table .clickable-list,
.visit-table .last-seen { .visit-table .last-seen {
cursor: pointer; cursor: pointer;
position: relative;
}
.visit-table :deep(.copy-popup) {
position: absolute;
bottom: calc(100% + 0.25rem);
left: 50%;
transform: translateX(-50%);
padding: 0.15rem 0.4rem;
background: var(--text, CanvasText);
color: var(--bg, Canvas);
border-radius: 0.25rem;
font-size: 0.75rem;
white-space: nowrap;
pointer-events: none;
z-index: 10;
} }
.crawler-top-uas { .crawler-top-uas {
-8
View File
@@ -403,14 +403,6 @@ onUnmounted(() => {
margin-left: auto; margin-left: auto;
padding: 0 0.2rem; padding: 0 0.2rem;
font-size: 1rem; font-size: 1rem;
background: none;
border: none;
cursor: pointer;
opacity: 0.7;
}
.block-head .icon-btn:hover {
opacity: 1;
} }
/* The banner design selector stays compact; the upload button is pushed /* The banner design selector stays compact; the upload button is pushed
+18 -9
View File
@@ -21,11 +21,11 @@ const currentPath = ref(props.pagePath)
const activeMode = ref(props.initialMode) const activeMode = ref(props.initialMode)
// The shared language selection (./editorLang, v-modeled by the tabs' // The shared language selection (./editorLang, v-modeled by the tabs'
// LangSelects) also drives the page preview: while the shell is open it // LangSelects) is linked to the whole-page language: while the shell is
// overrides the normal language preferences (?lang= / Accept-Language), // open it drives the page preview (overrides ?lang= / Accept-Language),
// so the page renders in the language being edited; closing restores. // and closing keeps the pick as the session language. The primary
// The primary selection pins by the CURRENT PAGE's own primary language // selection pins by the CURRENT PAGE's own primary language (pages may
// (pages may differ — Node.language is inherited down the tree). // differ — Node.language is inherited down the tree).
let pinned = false let pinned = false
function pinPreviewLang() { function pinPreviewLang() {
pinned = true pinned = true
@@ -34,6 +34,15 @@ function pinPreviewLang() {
setLangOverride(editorLang.value || pagePrimary.value || 'en') setLangOverride(editorLang.value || pagePrimary.value || 'en')
loadPlain(currentPath.value) loadPlain(currentPath.value)
} }
// Opening the panel must not switch the page's language: adopt the
// session's chosen language (public selector / earlier pick) once, then
// pin. Runs only on (re)open — after that the selection is the user's.
function openShell() {
const session = window.__pageriteLang
if (!editorLang.value && session && session !== (pagePrimary.value || 'en'))
editorLang.value = session
pinPreviewLang()
}
function unpinPreviewLang() { function unpinPreviewLang() {
if (!pinned) return if (!pinned) return
pinned = false pinned = false
@@ -89,13 +98,13 @@ function onSwitchEvent(ev) {
onMounted(() => { onMounted(() => {
document.body.dataset.editorMode = activeMode.value document.body.dataset.editorMode = activeMode.value
addEventListener('pagerite:switch-editor', onSwitchEvent) addEventListener('pagerite:switch-editor', onSwitchEvent)
addEventListener('pagerite:editor-shown', pinPreviewLang) addEventListener('pagerite:editor-shown', openShell)
addEventListener('pagerite:editor-hidden', unpinPreviewLang) addEventListener('pagerite:editor-hidden', unpinPreviewLang)
// The shell mounts visible (openEditor), so pin immediately. The site // The shell mounts visible (openEditor), so open immediately. The site
// default primary language comes from the settings — it only fills the // default primary language comes from the settings — it only fills the
// unknown; the page/structure tabs refine pagePrimary per page as they // unknown; the page/structure tabs refine pagePrimary per page as they
// learn it (their knowledge is strictly better). // learn it (their knowledge is strictly better).
pinPreviewLang() openShell()
fetch('/_api/settings').then((r) => r.json()).then((s) => { fetch('/_api/settings').then((r) => r.json()).then((s) => {
if (!pagePrimary.value) pagePrimary.value = s.primary_lang || 'en' if (!pagePrimary.value) pagePrimary.value = s.primary_lang || 'en'
}).catch(() => { /* keep the fallback */ }) }).catch(() => { /* keep the fallback */ })
@@ -103,7 +112,7 @@ onMounted(() => {
onUnmounted(() => { onUnmounted(() => {
removeEventListener('pagerite:switch-editor', onSwitchEvent) removeEventListener('pagerite:switch-editor', onSwitchEvent)
removeEventListener('pagerite:editor-shown', pinPreviewLang) removeEventListener('pagerite:editor-shown', openShell)
removeEventListener('pagerite:editor-hidden', unpinPreviewLang) removeEventListener('pagerite:editor-hidden', unpinPreviewLang)
}) })
</script> </script>
+30 -9
View File
@@ -3,7 +3,8 @@
// flag button opening a clean dropdown, v-modeled on the shared editorLang // flag button opening a clean dropdown, v-modeled on the shared editorLang
// ('' = the primary language). The lang tab's flag grid is a different // ('' = the primary language). The lang tab's flag grid is a different
// control (toggles, not a select) and stays as it is. // control (toggles, not a select) and stays as it is.
import { computed, ref } from 'vue' import { computed, nextTick, ref } from 'vue'
import { usePopup } from './dropdown'
const props = defineProps({ const props = defineProps({
modelValue: { type: String, default: '' }, modelValue: { type: String, default: '' },
@@ -13,8 +14,12 @@ const props = defineProps({
const emit = defineEmits(['update:modelValue']) const emit = defineEmits(['update:modelValue'])
const open = ref(false) const open = ref(false)
const root = ref(null)
const toggleBtn = ref(null) const toggleBtn = ref(null)
const pop = ref(null)
const popStyle = ref({}) const popStyle = ref({})
// Closes on outside click / Escape (./dropdown), not on mouseleave.
usePopup(open, root)
const current = computed( const current = computed(
() => props.options.find((o) => o.tag === props.modelValue) ?? props.options[0], () => props.options.find((o) => o.tag === props.modelValue) ?? props.options[0],
) )
@@ -26,6 +31,17 @@ function toggle() {
// onto the page area instead of being clipped by it. // onto the page area instead of being clipped by it.
const r = toggleBtn.value.getBoundingClientRect() const r = toggleBtn.value.getBoundingClientRect()
popStyle.value = { top: `${r.bottom + 2}px`, left: `${r.left}px` } popStyle.value = { top: `${r.bottom + 2}px`, left: `${r.left}px` }
// A toggle mounted near the right window edge (the public page
// selector sits top-right) opens the popup flush against that edge.
nextTick(() => {
const p = pop.value?.getBoundingClientRect()
if (p && p.right > innerWidth - 4) {
popStyle.value = {
...popStyle.value,
left: `${Math.max(4, innerWidth - 4 - p.width)}px`,
}
}
})
} }
} }
@@ -36,7 +52,7 @@ function select(tag) {
</script> </script>
<template> <template>
<span v-if="options.length > 1" class="lang-select"> <span v-if="options.length > 1" ref="root" class="lang-select">
<button <button
ref="toggleBtn" ref="toggleBtn"
type="button" type="button"
@@ -47,7 +63,7 @@ function select(tag) {
: '')" : '')"
@click="toggle" @click="toggle"
><span v-if="current?.flag" class="flag" v-html="current.flag" /></button> ><span v-if="current?.flag" class="flag" v-html="current.flag" /></button>
<span v-if="open" class="lang-pop" :style="popStyle" @mouseleave="open = false"> <span v-if="open" ref="pop" class="lang-pop" :style="popStyle">
<button <button
v-for="o in options" v-for="o in options"
:key="o.code" :key="o.code"
@@ -66,20 +82,23 @@ function select(tag) {
display: flex; display: flex;
} }
/* The closed state is just the small flag — no button chrome until hovered. */ /* The closed state is just the small flag — no button chrome at all, on
hover either (it sits among borderless emoji-icon buttons); like them it
rests dimmed and brightens on hover. */
.lang-current { .lang-current {
display: flex; display: flex;
align-items: center; align-items: center;
padding: 2px; padding: 2px;
background: none; background: none;
border: 1px solid transparent; border: none;
border-radius: 4px; border-radius: 4px;
cursor: pointer; cursor: pointer;
opacity: 0.7;
} }
.lang-current:hover, .lang-current:hover,
.lang-current.open { .lang-current.open {
border-color: var(--line); opacity: 1;
} }
/* The dropdown matches the page's existing popups (.picker-pop look). /* The dropdown matches the page's existing popups (.picker-pop look).
@@ -127,11 +146,13 @@ function select(tag) {
color: var(--muted); color: var(--muted);
} }
/* Flags render like in the analytics visitor cells. */ /* em-sized so the chip matches the surrounding text/icon size in each
context; the hairline border delineates white-flagged countries (not
button chrome). */
.flag { .flag {
display: inline-flex; display: inline-flex;
width: 18px; width: 1.5em;
height: 12px; height: 1em;
flex: 0 0 auto; flex: 0 0 auto;
border-radius: 2px; border-radius: 2px;
overflow: hidden; overflow: hidden;
+46
View File
@@ -0,0 +1,46 @@
<script setup>
// The public page's language selector: the editors' flag dropdown
// (LangSelect) as the first item of the banner's corner container, fed
// from the shared store (pagerite.js sets the page's hreflang alternates
// and served language per navigation). It binds the same store.lang the
// editor's dropdown binds, so both always show the same selection. A pick
// also dispatches pagerite:set-session-lang — pagerite.js swaps the page
// in place when the editor is closed (open, the editor reacts to the
// store and re-renders it).
import { computed } from 'vue'
import LangSelect from './LangSelect.vue'
import { flagFor, langName, langSort } from './langs'
import { useStore } from './store'
const store = useStore()
// The "(primary)" marker is admin-panel information; the public selector
// lists plain languages. Order: the primary language first, then the rest
// in the lang tab's geographic grouping (./langs langSort) — the head's
// hreflang order is just alphabetical.
const primaryTag = computed(() => store.langAlternates.find((a) => a.primary)?.tag ?? '')
const options = computed(() => {
const rest = langSort(
store.langAlternates.map((a) => a.tag).filter((t) => t !== primaryTag.value),
)
return [primaryTag.value, ...rest].filter(Boolean).map((tag) => ({
tag,
code: tag,
name: langName(tag),
flag: flagFor(tag),
primary: false,
}))
})
// The explicit pick, else the served language (header-autodetected pages
// may have neither), else the primary.
const model = computed(() => store.lang || store.servedLang || primaryTag.value)
function go(tag) {
store.lang = tag === primaryTag.value ? '' : tag
dispatchEvent(new CustomEvent('pagerite:set-session-lang', { detail: { lang: tag } }))
}
</script>
<template>
<LangSelect :model-value="model" :options="options" @update:model-value="go" />
</template>
+161 -32
View File
@@ -1,16 +1,21 @@
<script setup> <script setup>
// Lang tab: the site-wide translation target languages (translate_langs) // Lang tab: the site-wide translation target languages (translate_langs)
// and the translator service WebSocket URL(s) (translate_keys). ALL // and the translator service keys (translate_keys) with their WebSocket
// languages are listed, English included — a page whose primary language // URLs. ALL languages are listed, English included — a page whose primary
// (Node.language, configured per row in the structure tab, inherited down // language (Node.language, configured per row in the structure tab,
// the hierarchy) differs can be translated INTO any other. Flag clicks // inherited down the hierarchy) differs can be translated INTO any other.
// toggle and save immediately; the settings round-trip re-reads the // Flag clicks toggle and save immediately; the settings round-trip
// payload, so this tab only ever changes translate_langs. The settings // re-reads the payload, so this tab only ever changes translate_langs. The
// write's invalidation hook kicks the translation dispatcher. The refresh // settings write's invalidation hook kicks the translation dispatcher. The
// button drops all machine translations (user patches are kept), making // refresh button drops all machine translations (user patches are kept),
// the dispatcher re-translate everything. // making the dispatcher re-translate everything. Translator keys are
// managed inline ( add, name edit, ✕ delete); new keys are generated
// here in the server's format and everything rides the settings
// round-trip. Clicking a key copies its full URL (following ws:// would
// fail).
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue' import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs' import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs'
import { copyList } from './analytics/format.js'
import { dropPageCache } from './swapdoc' import { dropPageCache } from './swapdoc'
defineProps({ pagePath: { type: String, default: '' } }) defineProps({ pagePath: { type: String, default: '' } })
@@ -21,6 +26,16 @@ const saveError = ref('')
const selected = ref(new Set()) const selected = ref(new Set())
const keyUrls = ref([]) const keyUrls = ref([])
// Full WebSocket URL for a key. New keys are generated right here: 12
// lowercase alphanumerics, the server-side format (state._KEY_ALPHABET).
const wsUrl = (key) =>
`${location.origin.replace(/^http/, 'ws')}/_translate/${key}`
const KEY_ALPHABET = 'abcdefghijklmnopqrstuvwxyz0123456789'
const newKey = () =>
[...crypto.getRandomValues(new Uint8Array(12))]
.map((b) => KEY_ALPHABET[b % KEY_ALPHABET.length])
.join('')
// The toggleable targets: every translatable language, laid out in // The toggleable targets: every translatable language, laid out in
// geographic/cultural groups (one row each) rather than alphabetized — // geographic/cultural groups (one row each) rather than alphabetized —
// related languages sit together (a node's own primary is excluded per // related languages sit together (a node's own primary is excluded per
@@ -52,9 +67,8 @@ onMounted(async () => {
try { try {
const s = await (await fetch('/_api/settings')).json() const s = await (await fetch('/_api/settings')).json()
selected.value = new Set(s.translate_langs || []) selected.value = new Set(s.translate_langs || [])
const wsBase = location.origin.replace(/^http/, 'ws')
keyUrls.value = Object.entries(s.translate_keys || {}) keyUrls.value = Object.entries(s.translate_keys || {})
.map(([key, name]) => ({ name, url: `${wsBase}/_translate/${key}` })) .map(([key, name]) => ({ key, name, url: wsUrl(key) }))
} catch { /* keep defaults */ } } catch { /* keep defaults */ }
}) })
@@ -100,6 +114,39 @@ async function refresh() {
refreshing.value = false refreshing.value = false
} }
} }
// Key management rides the settings round-trip, like toggle() above:
// mutate keyUrls, then PUT the whole settings payload with the new
// translate_keys. adds a fresh unnamed key, names save on every
// keystroke (@input — spamming the server is fine), ✕ deletes without
// confirmation.
async function saveKeys() {
try {
const s = await (await fetch('/_api/settings')).json()
const res = await fetch('/_api/settings', {
method: 'PUT',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
...s,
translate_keys: Object.fromEntries(keyUrls.value.map((k) => [k.key, k.name])),
}),
})
saveError.value = res.ok ? '' : '⚠️ changes could not be saved'
} catch {
saveError.value = '⚠️ changes could not be saved'
}
}
function addKey() {
const key = newKey()
keyUrls.value.push({ key, name: '', url: wsUrl(key) })
saveKeys()
}
function removeKey(k) {
keyUrls.value = keyUrls.value.filter((x) => x.key !== k.key)
saveKeys()
}
</script> </script>
<template> <template>
@@ -129,28 +176,37 @@ async function refresh() {
<section class="block"> <section class="block">
<div class="block-head"> <div class="block-head">
<span class="field-label">translations</span> <span class="field-label">Translator API</span>
<small class="muted">deleting re-translates everything; user edits are kept</small>
</div> </div>
<button <div v-for="k in keyUrls" :key="k.key" class="key-row">
type="button" <a
class="refresh-btn" :href="k.url"
:disabled="refreshing" class="key-link"
title="delete all machine translations and let the translator re-fill them" title="click to copy the URL"
@click="refresh" @click.prevent="copyList(k.url, $event)"
> >{{ k.key }}</a>
{{ refreshing ? 'refreshing…' : 'refresh all translations' }} <input
</button> v-model="k.name"
</section> type="text"
class="edit key-name"
<section v-if="keyUrls.length" class="block"> title="display name"
<div class="block-head"> @input="saveKeys()"
<span class="field-label">translator service</span> >
<small class="muted">connect scripts/translator.py to</small> <button type="button" class="act del" title="delete key" @click="removeKey(k)"></button>
</div> </div>
<div v-for="k in keyUrls" :key="k.url" class="key-row"> <div class="add-row">
<code>{{ k.url }}</code> <button type="button" class="add" title="new translator key" @click="addKey()"> API key</button>
<small class="muted">{{ k.name }}</small> </div>
<p><small class="muted">AI translator agents can connect with the API keys to do machine translations to your selected languages. Click the button below to delete all translations and start over. User edits are kept.</small></p>
<div class="refresh-row">
<button
type="button"
class="refresh-btn"
:disabled="refreshing"
@click="refresh"
>
{{ refreshing ? 'Reseting…' : 'Reset' }}
</button>
</div> </div>
</section> </section>
</div> </div>
@@ -252,10 +308,83 @@ async function refresh() {
gap: 0.6rem; gap: 0.6rem;
} }
.key-row code { /* Real links (handy for right-click/drag) showing just the key, but the
click copies the full URL instead of following — ws:// would fail to
navigate. Normal text color, not link-styled; position: relative
anchors the "Copied!" popup (analytics/format.js). */
.key-link {
position: relative;
color: var(--text);
font-family: var(--font-code);
user-select: all; user-select: all;
} }
.refresh-row {
display: flex;
align-items: baseline;
gap: 0.6rem;
}
/* Name input / ✕ / follow the structure tab's conventions: inputs stay
borderless until interacted with, glyph buttons redden / solidify on
hover. */
.key-name {
flex: 0 0 9rem;
}
.edit {
font: inherit;
font-size: 0.85rem;
padding: 0.1rem 0.4rem;
background: transparent;
color: var(--text);
border: 1px solid transparent;
border-radius: 4px;
min-width: 0;
}
.edit:hover {
border-color: var(--line);
}
.edit:focus {
background: var(--bg);
border-color: var(--accent);
outline: none;
}
.act {
padding: 0 0.25rem;
background: none;
border: none;
color: var(--muted);
font-size: 0.8rem;
cursor: pointer;
white-space: nowrap;
}
.del:hover {
color: #e06c75;
}
.add-row {
display: flex;
align-items: center;
}
.add {
padding: 0 0.3rem;
background: none;
border: none;
font-size: 0.9rem;
cursor: pointer;
opacity: 0.5;
}
.add:hover {
opacity: 1;
}
.refresh-btn { .refresh-btn {
align-self: flex-start; align-self: flex-start;
margin-bottom: 0.2rem; margin-bottom: 0.2rem;
+40 -14
View File
@@ -26,13 +26,14 @@
// renders the version being edited, whichever language the page itself // renders the version being edited, whichever language the page itself
// was loaded in. // was loaded in.
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue' import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
import { usePopup } from './dropdown'
import { EditorView, basicSetup } from 'codemirror' import { EditorView, basicSetup } from 'codemirror'
import { Compartment, EditorState } from '@codemirror/state' import { Compartment, EditorState } from '@codemirror/state'
import { keymap } from '@codemirror/view' import { keymap } from '@codemirror/view'
import { indentWithTab } from '@codemirror/commands' import { indentWithTab } from '@codemirror/commands'
import { markdown } from '@codemirror/lang-markdown' import { markdown } from '@codemirror/lang-markdown'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
import { flagFor, langName } from './langs' import { flagFor, langName, langSort } from './langs'
import { editorLang, pagePrimary } from './editorLang' import { editorLang, pagePrimary } from './editorLang'
import LangSelect from './LangSelect.vue' import LangSelect from './LangSelect.vue'
import ConnNote from './ConnNote.vue' import ConnNote from './ConnNote.vue'
@@ -120,11 +121,13 @@ function normPath(p) {
// localization settings tab). // localization settings tab).
// The picker's options: the primary language first, then the union of the // The picker's options: the primary language first, then the union of the
// page's translations and the site-wide configured targets, sorted. // page's translations and the site-wide configured targets in the lang
// tab's geographic grouping (./langs langSort).
const langOptions = computed(() => { const langOptions = computed(() => {
const others = [...new Set([...siteLangs.value, ...pageLangs.value])] const others = langSort(
.filter((l) => l && l !== primaryLang.value) [...new Set([...siteLangs.value, ...pageLangs.value])]
.sort() .filter((l) => l && l !== primaryLang.value),
)
return [primaryLang.value, ...others].map((code) => ({ return [primaryLang.value, ...others].map((code) => ({
tag: code === primaryLang.value ? '' : code, tag: code === primaryLang.value ? '' : code,
code, code,
@@ -621,8 +624,15 @@ const TABLE_MAX_ROWS = 6
// Class pickers: popup listing the block class toggles (placement ↔︎, // Class pickers: popup listing the block class toggles (placement ↔︎,
// text size AA), closed after applying. The block's current class of the // text size AA), closed after applying. The block's current class of the
// group is marked; choosing "normal" (or the current class) removes it. // group is marked; choosing "normal" (or the current class) removes it.
// All popups share the close behavior of ./dropdown (outside click /
// Escape; never mouseleave).
const classPicker = ref(null) // 'place' | 'size' | null const classPicker = ref(null) // 'place' | 'size' | null
const activeClasses = ref(new Set()) const activeClasses = ref(new Set())
const placeRoot = ref(null)
const sizeRoot = ref(null)
const tableRoot = ref(null)
usePopup(classPicker, computed(() => (classPicker.value === 'place' ? placeRoot : sizeRoot).value))
usePopup(tablePicker, tableRoot)
function openClassPicker(which) { function openClassPicker(which) {
classPicker.value = classPicker.value === which ? null : which classPicker.value = classPicker.value === which ? null : which
@@ -1136,15 +1146,31 @@ onUnmounted(() => {
<div class="format-bar"> <div class="format-bar">
<button type="button" class="code-btn" title="code — inline wrap, or a fenced block for line-spanning selections; click again to unwrap" @click="insertCode"><code>&lt;/&gt;</code></button> <button type="button" class="code-btn" title="code — inline wrap, or a fenced block for line-spanning selections; click again to unwrap" @click="insertCode"><code>&lt;/&gt;</code></button>
<button type="button" title="link (toggle: click inside a link to unwrap it)" @click="insertLink">🔗</button> <button type="button" title="link (toggle: click inside a link to unwrap it)" @click="insertLink">🔗</button>
<button <span class="picker" ref="tableRoot">
type="button" <button
title="table" type="button"
:class="{ active: tablePicker }" title="table"
@click="tablePicker = !tablePicker" :class="{ active: tablePicker }"
></button> @click="tablePicker = !tablePicker"
></button>
<div v-if="tablePicker" class="table-picker" @mouseleave="tableSize = { cols: 0, rows: 0 }">
<div class="tp-grid" :style="{ gridTemplateColumns: `repeat(${TABLE_MAX_COLS}, 1fr)` }">
<button
v-for="n in TABLE_MAX_COLS * TABLE_MAX_ROWS"
:key="n"
type="button"
class="tp-cell"
:class="{ on: tableSize.cols >= (n - 1) % TABLE_MAX_COLS + 1 && tableSize.rows >= Math.floor((n - 1) / TABLE_MAX_COLS) + 1 }"
@mouseenter="tableSize = { cols: (n - 1) % TABLE_MAX_COLS + 1, rows: Math.floor((n - 1) / TABLE_MAX_COLS) + 1 }"
@click="insertTable(tableSize.cols, tableSize.rows)"
/>
</div>
<div class="tp-size">{{ tableSize.cols || '' }} × {{ tableSize.rows || '' }}</div>
</div>
</span>
<button type="button" title="insert image (upload) — pasting works too" @click="fileInput.click()">🖼</button> <button type="button" title="insert image (upload) — pasting works too" @click="fileInput.click()">🖼</button>
<button type="button" title="aside box (::: aside) — wraps the selection or the cursor's line; clicked inside one, removes it" @click="insertAside"></button> <button type="button" title="aside box (::: aside) — wraps the selection or the cursor's line; clicked inside one, removes it" @click="insertAside"></button>
<span class="picker"> <span class="picker" ref="placeRoot">
<button <button
type="button" type="button"
title="block placement class" title="block placement class"
@@ -1166,7 +1192,7 @@ onUnmounted(() => {
</span> </span>
<button type="button" title="bold" @click="wrapInline('**')"><b>B</b></button> <button type="button" title="bold" @click="wrapInline('**')"><b>B</b></button>
<button type="button" title="italic" @click="wrapInline('*')"><i>i</i></button> <button type="button" title="italic" @click="wrapInline('*')"><i>i</i></button>
<span class="picker"> <span class="picker" ref="sizeRoot">
<button <button
type="button" type="button"
title="text size class" title="text size class"
@@ -1368,7 +1394,7 @@ onUnmounted(() => {
.table-picker { .table-picker {
position: absolute; position: absolute;
top: 100%; top: 100%;
left: 6.5rem; left: 0;
z-index: 20; z-index: 20;
padding: 0.5rem; padding: 0.5rem;
background: var(--bg); background: var(--bg);
+68
View File
@@ -0,0 +1,68 @@
<script setup>
// Referer + UTM as one badge in the analytics visit/crawler tables: the
// referer's favicon flush on the left, its host as the link text, then the
// UTM summary smaller/muted inside the same badge. The whole badge links to
// the referer origin (external, new tab) and carries a single one-fact-
// per-line tooltip (badge.title: origin, then each utm pair) — no titles on
// the inner elements. Referers are external, so there is no close event.
import { computed } from 'vue'
const props = defineProps({
badge: { type: Object, required: true },
favicons: { type: Object, default: null },
})
const favicon = computed(() => (props.badge.origin ? props.favicons?.[props.badge.origin] : null))
</script>
<template>
<a class="referer-badge"
:href="badge.href || undefined"
:title="badge.title"
:target="badge.href ? '_blank' : undefined"
:rel="badge.href ? 'noopener' : undefined">
<img v-if="favicon" class="badge-favicon" :src="favicon" alt="" />
<span v-if="badge.label">{{ badge.label }}</span>
<small v-if="badge.utm" class="small">{{ badge.utm }}</small>
</a>
</template>
<style scoped>
/* Browser-chrome chip on a fixed neutral palette (--badge-* in
pagerite.css, deliberately unthemed): black-on-transparent and
white-on-transparent favicons both stay legible on it, and the text is
always dark regardless of the theme's link/text colors. Colors go on the
inner elements, so the theme's a / a:hover color rules (which target the
anchor) cannot cascade in. */
.referer-badge {
display: inline-flex;
align-items: center;
gap: 0.35em;
padding-right: 0.45em;
border-radius: 0.25rem;
background: var(--badge-bg);
color: var(--badge-text);
white-space: nowrap;
overflow: hidden;
}
/* Flush with the badge's top/left/bottom edges: a full-height square (the
badge has no padding on those sides), corners clipped by the badge's
overflow: hidden border-radius. */
.badge-favicon {
width: 1.5em;
height: 1.5em;
object-fit: cover;
flex: none;
}
.referer-badge > span,
.referer-badge > small {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
}
.referer-badge > span { color: var(--badge-text); }
.referer-badge > small { color: var(--badge-muted); }
</style>
-8
View File
@@ -782,14 +782,6 @@ onUnmounted(() => {
margin-left: auto; margin-left: auto;
padding: 0 0.2rem; padding: 0 0.2rem;
font-size: 1rem; font-size: 1rem;
background: none;
border: none;
cursor: pointer;
opacity: 0.7;
}
.block-head .icon-btn:hover {
opacity: 1;
} }
.text-input { .text-input {
+5 -4
View File
@@ -19,7 +19,7 @@ import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, wa
import StructureTree from './StructureTree.vue' import StructureTree from './StructureTree.vue'
import LangSelect from './LangSelect.vue' import LangSelect from './LangSelect.vue'
import { slugify } from './slugify' import { slugify } from './slugify'
import { flagFor, langName } from './langs' import { flagFor, langName, langSort } from './langs'
import { editorLang, pagePrimary } from './editorLang' import { editorLang, pagePrimary } from './editorLang'
import { dropPageCache, loadPlain } from './swapdoc' import { dropPageCache, loadPlain } from './swapdoc'
@@ -40,9 +40,10 @@ const primaryLang = ref('en')
const siteLangs = ref([]) const siteLangs = ref([])
// The strip's options: the primary language first, then the configured // The strip's options: the primary language first, then the configured
// translation targets (the lang tab manages that set). // translation targets (the lang tab manages that set) in the lang tab's
// geographic grouping (./langs langSort).
const langOptions = computed(() => const langOptions = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)] [primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
.map((code) => ({ .map((code) => ({
tag: code === primaryLang.value ? '' : code, tag: code === primaryLang.value ? '' : code,
code, code,
@@ -63,7 +64,7 @@ watch(lang, () => refreshPages())
// dropdown lists "inherit" first (naming what it resolves to), then every // dropdown lists "inherit" first (naming what it resolves to), then every
// site language. Setting it on a section covers its whole subtree. // site language. Setting it on a section covers its whole subtree.
const rowLangChoices = computed(() => const rowLangChoices = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)] [primaryLang.value, ...langSort(siteLangs.value.filter((l) => l !== primaryLang.value))]
.map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })), .map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })),
) )
function rowLangOptions(el) { function rowLangOptions(el) {
+21
View File
@@ -6,6 +6,7 @@ const props = defineProps({
step: { type: Object, required: true }, step: { type: Object, required: true },
count: { type: Number, default: 0 }, count: { type: Number, default: 0 },
favicons: { type: Object, default: null }, favicons: { type: Object, default: null },
flags: { type: Array, default: () => [] },
}) })
defineEmits(['close']) defineEmits(['close'])
@@ -38,6 +39,7 @@ const title = computed(() => {
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small> <small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
<img v-if="favicon" class="favicon" :src="favicon" alt="" /> <img v-if="favicon" class="favicon" :src="favicon" alt="" />
<span>{{ step.slug }}</span> <span>{{ step.slug }}</span>
<span v-for="(f, fi) in flags" :key="fi" class="flag" v-html="f"></span>
</a> </a>
</template> </template>
@@ -48,4 +50,23 @@ const title = computed(() => {
margin-right: 0.25em; margin-right: 0.25em;
vertical-align: -0.1em; vertical-align: -0.1em;
} }
/* Same flag chip as the visitor cells (VisitorCell.vue). */
.flag {
display: inline-flex;
width: 18px;
height: 12px;
margin-left: 0.25em;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style> </style>
+23 -4
View File
@@ -4,15 +4,20 @@
// Clicking the IP copies the full address to the clipboard. // Clicking the IP copies the full address to the clipboard.
// ``variantCount`` overrides the UA line to warn when multiple client // ``variantCount`` overrides the UA line to warn when multiple client
// fingerprints share the same IP (e.g. a scanner rotating UAs). // fingerprints share the same IP (e.g. a scanner rotating UAs).
// Clicking the UA line copies the raw UA(s) to the clipboard, one per line
// (``uaRaws`` carries every variation for multi-client IPs).
import { computed } from 'vue' import { computed } from 'vue'
import * as flagSvgs from 'country-flag-icons/string/3x2' import * as flagSvgs from 'country-flag-icons/string/3x2'
import { copyIp, formatLang } from './analytics/format.js' import { copyIp, copyList, formatLang } from './analytics/format.js'
import { langName } from './langs.js'
const props = defineProps({ const props = defineProps({
ip: { type: String, default: '' }, ip: { type: String, default: '' },
ipDisplay: { type: String, default: '—' }, ipDisplay: { type: String, default: '—' },
ua: { type: String, default: '' }, ua: { type: String, default: '' },
uaRaw: { type: String, default: '' }, uaRaw: { type: String, default: '' },
uaRaws: { type: String, default: '' },
uaUrl: { type: String, default: '' },
country: { type: String, default: '' }, country: { type: String, default: '' },
city: { type: String, default: '' }, city: { type: String, default: '' },
lang: { type: String, default: '' }, lang: { type: String, default: '' },
@@ -26,6 +31,7 @@ const hasCity = computed(() => !!(props.city && props.city !== '—'))
const hasLocale = computed(() => hasCountry.value || hasCity.value) const hasLocale = computed(() => hasCountry.value || hasCity.value)
const langValue = computed(() => props.langDisplay || formatLang(props.lang)) const langValue = computed(() => props.langDisplay || formatLang(props.lang))
const showLang = computed(() => langValue.value && langValue.value !== '—') const showLang = computed(() => langValue.value && langValue.value !== '—')
const uaCopy = computed(() => props.uaRaws || props.uaRaw)
function flagSvg(code) { function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || '' return flagSvgs[code?.toUpperCase()] || ''
@@ -58,10 +64,16 @@ function countryName(code) {
</div> </div>
<div class="visitor-row"> <div class="visitor-row">
<div class="ua-line"> <div class="ua-line">
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small> <small v-if="variantCount > 1" class="muted variant-hint clickable-ip"
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small> :title="uaCopy"
@click="copyList(uaCopy, $event)">{{ variantCount }} client variations</small>
<small v-else class="muted clickable-ip" :title="uaRaw"
@click="copyList(uaCopy, $event)">{{ ua || '' }}</small><a v-if="uaUrl && variantCount <= 1"
class="ua-link icon-btn" :href="uaUrl"
target="_blank" rel="noopener noreferrer"
@click.stop>🔗</a>
</div> </div>
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div> <div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted" :title="langName(lang)">{{ langValue }}</small></div>
</div> </div>
</div> </div>
</td> </td>
@@ -121,6 +133,13 @@ function countryName(code) {
text-align: left; text-align: left;
} }
.ua-link {
text-decoration: none;
font-size: 0.75em;
margin-left: 0.2em;
vertical-align: middle;
}
.locale-lang { .locale-lang {
flex: 0 0 auto; flex: 0 0 auto;
overflow: hidden; overflow: hidden;
+132 -45
View File
@@ -2,6 +2,7 @@
* Formatters and aggregators for summary sections: totals and the recent * Formatters and aggregators for summary sections: totals and the recent
* visit trail. * visit trail.
*/ */
import { flagFor, langName } from '../langs.js'
/** /**
* IPv4 unchanged, IPv6 returns the /64 network prefix in compact form. * IPv4 unchanged, IPv6 returns the /64 network prefix in compact form.
@@ -25,32 +26,31 @@ export const hostIP = (ip) => {
} }
} }
function showCopiedFeedback(el) { function showCopiedFeedback(el, event) {
if (!el || typeof document === 'undefined') return if (typeof document === 'undefined') return
const popup = document.createElement('span') const popup = document.createElement('span')
popup.textContent = 'Copied!' popup.textContent = 'Copied!'
popup.className = 'copy-popup' popup.className = 'copy-popup'
// Fixed to the viewport at the click point: table cells clip absolute
// popups with their overflow: hidden ellipsis styling.
const x = event?.clientX ?? 0
const y = event?.clientY ?? 0
popup.style.cssText = popup.style.cssText =
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' + `position:fixed;left:${x}px;top:${y}px;` +
'transform:translateX(-50%);padding:0.15rem 0.4rem;' + 'transform:translate(-50%, calc(-100% - 0.5rem));padding:0.15rem 0.4rem;' +
'background:var(--text, CanvasText);color:var(--bg, Canvas);' + 'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' + 'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
'pointer-events:none;z-index:10;' 'pointer-events:none;z-index:100;'
el.classList.add('has-copy-popup') document.body.appendChild(popup)
el.appendChild(popup) setTimeout(() => popup.remove(), 1200)
setTimeout(() => {
popup.remove()
el.classList.remove('has-copy-popup')
}, 1200)
} }
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */ /** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
export async function copyIp(ip, event) { export async function copyIp(ip, event) {
if (!ip) return if (!ip) return
const el = event?.currentTarget
try { try {
await navigator.clipboard.writeText(ip) await navigator.clipboard.writeText(ip)
showCopiedFeedback(el) showCopiedFeedback(event?.currentTarget, event)
} catch { } catch {
/* ignore */ /* ignore */
} }
@@ -59,10 +59,9 @@ export async function copyIp(ip, event) {
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */ /** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
export async function copyList(text, event) { export async function copyList(text, event) {
if (!text) return if (!text) return
const el = event?.currentTarget
try { try {
await navigator.clipboard.writeText(text) await navigator.clipboard.writeText(text)
showCopiedFeedback(el) showCopiedFeedback(event?.currentTarget, event)
} catch { } catch {
/* ignore */ /* ignore */
} }
@@ -178,6 +177,32 @@ function stepOf(path, titles) {
return null return null
} }
/**
* Badge data combining a visit's/crawler's external referer origin with the
* visit's UTM tags: the origin as the badge link/label (the favicon is
* looked up by origin in the component), the known UTM values as a short
* inline summary, and a one-fact-per-line tooltip — the full origin URL on
* the first line, then every ``utm_*=value`` pair. Null when there is no
* external referer and no UTM tag (a plain direct visit).
*/
function refererBadgeOf(referer, titles, utmTags = {}) {
const step = stepOf(referer, titles)
const external = step?.external ? step : null
const known = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
const utm = known.map((k) => utmTags[k]).filter(Boolean).join(' · ')
if (!external && !utm) return null
return {
href: external?.origin || '',
label: external?.slug || '',
origin: external?.origin || '',
utm,
title: [
...(external ? [external.origin] : []),
...Object.entries(utmTags).map(([k, value]) => `${k}=${value}`),
].join('\n'),
}
}
/** /**
* Human-readable relative timestamp. Adapted from cista-storage: uses * Human-readable relative timestamp. Adapted from cista-storage: uses
* ``Intl.RelativeTimeFormat`` for short intervals and a compact date for * ``Intl.RelativeTimeFormat`` for short intervals and a compact date for
@@ -342,7 +367,7 @@ export function countCrawlerUas(crawlers, clients) {
const counts = {} const counts = {}
for (const c of crawlers || []) { for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {} const client = (clients || {})[c.client] || {}
const value = client.ua_pretty || client.ua || '(no UA)' const value = client.uarite?.pretty || client.ua || '(no UA)'
counts[value] = (counts[value] || 0) + 1 counts[value] = (counts[value] || 0) + 1
} }
return Object.entries(counts).sort((a, b) => b[1] - a[1]) return Object.entries(counts).sort((a, b) => b[1] - a[1])
@@ -370,12 +395,13 @@ export function mainDomain(host, limit = 24) {
/** /**
* Group raw crawler hits by client hash and format each group as a row showing * Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by most recent hit * every internal page that crawler visited. Rows are sorted by most recent hit
* first, with total hits as a tie-breaker. The group's ``refererStep`` is the * first, with total hits as a tie-breaker. The group's ``refererBadge`` is
* latest external referer seen for the crawler — spiders often advertise * the latest external referer seen for the crawler — spiders often advertise
* their own site there — rendered with its favicon like visit referers. * their own site there — rendered as a badge with its favicon like visit
* referers.
* ``clients`` maps client hashes to client records. * ``clients`` maps client hashes to client records.
*/ */
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) { export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now(), site = { multilingual: false, primaryLang: '' }) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
const groups = new Map() const groups = new Map()
for (const c of crawlers || []) { for (const c of crawlers || []) {
@@ -386,10 +412,12 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
lastStart: 0, lastStart: 0,
referer: '', referer: '',
pages: new Map(), pages: new Map(),
langs: new Set(),
} }
const start = new Date(c.start).getTime() const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start if (start > g.lastStart) g.lastStart = start
if (c.referer) g.referer = c.referer if (c.referer) g.referer = c.referer
if (c.lang) g.langs.add(c.lang)
if (c.entry?.startsWith('/')) { if (c.entry?.startsWith('/')) {
const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 } const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
existing.count += 1 existing.count += 1
@@ -410,19 +438,28 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
const client = g.client || {} const client = g.client || {}
const host = client.host || '' const host = client.host || ''
const isHost = !!host const isHost = !!host
// Rendered languages read, shown only when they say something the
// primary language alone would not (multilingual sites only).
const langs = [...g.langs].sort()
const showLangs =
site.multilingual && (langs.length > 1 || (langs[0] && langs[0] !== site.primaryLang))
return { return {
lastSeen: formatWhen(g.lastStart, now), lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart), lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart), lastSeenLocal: formatWhenLocal(g.lastStart),
refererStep: stepOf(g.referer, titles), refererBadge: refererBadgeOf(g.referer, titles),
pages: [...g.pages.entries()] pages: [...g.pages.entries()]
.sort((a, b) => b[1].count - a[1].count) .sort((a, b) => b[1].count - a[1].count)
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })), .map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
readFlags: showLangs
? langs.map((l) => ({ flag: flagFor(l), name: langName(l) })).filter((f) => f.flag)
: [],
ip: client.ip || '', ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost, isHost,
ua: client.ua_pretty || client.ua || '—', ua: client.uarite?.pretty || client.ua || '—',
uaRaw: client.ua || '', uaRaw: client.ua || '',
uaUrl: client.uarite?.url || '',
lang: client.lang || '—', lang: client.lang || '—',
langDisplay: formatLang(client.lang), langDisplay: formatLang(client.lang),
country: client.country || '—', country: client.country || '—',
@@ -434,7 +471,9 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
/** /**
* Group abuse hits by IP and format each group as a row with the full paths * Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical paths are collapsed into one entry with their hit count. * probed. Identical requests (same path and status class) are collapsed
* into one entry with their hit count; a path's 404 probes and its real
* (200) reads never merge.
* The paths split into two lists: ``paths`` holds the 404 probes (flagged * The paths split into two lists: ``paths`` holds the 404 probes (flagged
* paths — the ones that triggered abuse classification — first, then other * paths — the ones that triggered abuse classification — first, then other
* 404s) shown verbatim, query string included, and ``articles`` holds the * 404s) shown verbatim, query string included, and ``articles`` holds the
@@ -465,7 +504,11 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
g.lastClient = a.client g.lastClient = a.client
} }
const path = a.path || '' const path = a.path || ''
const existing = g.pathCounts.get(path) || { // Collapse identical requests, but never merge a path's 404 probes with
// its real (200) reads — a page probed while missing and later created
// must show up in both columns, not flip to "articles read".
const key = `${a.is_404 ? '4' : '2'}${path}`
const existing = g.pathCounts.get(key) || {
path, path,
count: 0, count: 0,
firstStart: start, firstStart: start,
@@ -475,8 +518,7 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
existing.count += 1 existing.count += 1
if (start < existing.firstStart) existing.firstStart = start if (start < existing.firstStart) existing.firstStart = start
if (a.flag) existing.flag = true if (a.flag) existing.flag = true
if (!a.is_404) existing.is_404 = false g.pathCounts.set(key, existing)
g.pathCounts.set(path, existing)
g.clientHashes.add(a.client) g.clientHashes.add(a.client)
groups.set(ip, g) groups.set(ip, g)
} }
@@ -500,6 +542,13 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
const client = (clients || {})[g.lastClient] || {} const client = (clients || {})[g.lastClient] || {}
const host = client.host || '' const host = client.host || ''
const isHost = !!host const isHost = !!host
const uaRaws = [
...new Set(
[...g.clientHashes]
.map((h) => (clients || {})[h]?.ua)
.filter(Boolean),
),
].join('\n')
return { return {
lastSeen: formatWhen(g.lastStart, now), lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart), lastSeenIso: formatWhenIso(g.lastStart),
@@ -522,8 +571,10 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
ip: client.ip || g.ip, ip: client.ip || g.ip,
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
isHost, isHost,
ua: client.ua_pretty || client.ua || '—', ua: client.uarite?.pretty || client.ua || '—',
uaRaw: client.ua || '', uaRaw: client.ua || '',
uaUrl: client.uarite?.url || '',
uaRaws,
lang: client.lang || '—', lang: client.lang || '—',
langDisplay: formatLang(client.lang), langDisplay: formatLang(client.lang),
country: client.country || '—', country: client.country || '—',
@@ -535,52 +586,88 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
/** /**
* Format raw visit records as rows for a technical table. Returns objects * Format raw visit records as rows for a technical table. Returns objects
* with display strings; missing values become "—". ``trail`` starts with the * with display strings; missing values become "—". The external referer
* external referer (when present), then the entry page and any further internal * (when present) and the UTM tags ride along as ``refererBadge``; ``trail``
* pages or external exit origins. Only the 20 most recent visits are shown. * holds the entry page and any further internal pages or external exit
* ``clients`` maps client hashes to client records. * origins; consecutive views of the same page (e.g. a
* language switch re-view) merge into one step that keeps the
* consecutive-distinct rendered languages, summed read time, and the latest
* status. On multilingual sites the rendered languages surface as flag
* icons: a visit read entirely in one non-primary language gets ``rowFlag``,
* and a visit spanning languages gets per-step ``langFlags`` markers where
* the language begins or changes. Only the 20 most recent visits are shown.
* ``clients`` maps client hashes to client records; ``site`` carries the
* payload's multilingual/primary-language context.
*/ */
export function formatVisitRows(visits, clients, pageTree, now = Date.now()) { export function formatVisitRows(visits, clients, pageTree, now = Date.now(), site = { multilingual: false, primaryLang: '' }) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().slice(0, 20).map((v) => { return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const client = (clients || {})[v.client] || {} const client = (clients || {})[v.client] || {}
const trail = Object.values(v.trail || {}) const steps = Object.values(v.trail || {})
.map((item) => { .map((item) => {
const step = stepOf(item.to, titles) const step = stepOf(item.to, titles)
if (step) { if (step) {
if (item.read) step.readSeconds = item.read if (item.read) step.readSeconds = item.read
if (item.status) step.status = item.status if (item.status) step.status = item.status
if (item.lang) step.lang = item.lang
} }
return step return step
}) })
.filter(Boolean) .filter(Boolean)
const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content'] const trail = []
const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean) for (const step of steps) {
const utm = utmValues.length ? utmValues.join(' · ') : '' const prev = trail[trail.length - 1]
const utmTitle = Object.entries(v.utm || {}) if (prev && prev.path === step.path) {
.map(([k, value]) => `${k}=${value}`) if (step.lang && step.lang !== prev.langs[prev.langs.length - 1]) prev.langs.push(step.lang)
.join(', ') if (step.readSeconds) prev.readSeconds = (prev.readSeconds || 0) + step.readSeconds
if (step.status) prev.status = step.status
} else {
step.langs = step.lang ? [step.lang] : []
trail.push(step)
}
}
const distinctLangs = new Set(trail.flatMap((s) => s.langs))
const dash = (s) => (s || '—') const dash = (s) => (s || '—')
const host = client.host || '' const host = client.host || ''
const isHost = !!host const isHost = !!host
return { const row = {
lastSeen: formatWhen(v.start, now), lastSeen: formatWhen(v.start, now),
lastSeenIso: formatWhenIso(v.start), lastSeenIso: formatWhenIso(v.start),
lastSeenLocal: formatWhenLocal(v.start), lastSeenLocal: formatWhenLocal(v.start),
langDisplay: formatLang(client.lang), langDisplay: formatLang(client.lang),
trail, trail,
refererStep: stepOf(v.referer, titles), refererBadge: refererBadgeOf(v.referer, titles, v.utm),
referer: dash(v.referer),
ip: client.ip || '', ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost, isHost,
lang: dash(client.lang), lang: dash(client.lang),
country: dash(client.country), country: dash(client.country),
city: dash(client.city), city: dash(client.city),
ua: client.ua_pretty || client.ua || '—', ua: client.uarite?.pretty || client.ua || '—',
uaRaw: client.ua || '', uaRaw: client.ua || '',
utm: utm || '', uaUrl: client.uarite?.url || '',
utmTitle,
} }
if (site.multilingual && distinctLangs.size) {
if (distinctLangs.size === 1) {
const [tag] = distinctLangs
const flag = flagFor(tag)
if (flag && tag !== site.primaryLang) {
row.rowFlag = flag
row.rowFlagTitle = langName(tag)
}
} else {
// Flag the steps where the rendered language begins or changes;
// lang-less steps keep the comparison chain going, they never flag.
let lastLang = null
for (const step of trail) {
if (!step.langs.length) continue
if (!lastLang || step.langs[step.langs.length - 1] !== lastLang) {
step.langFlags = step.langs.map(flagFor).filter(Boolean)
}
lastLang = step.langs[step.langs.length - 1]
}
}
}
return row
}) })
} }
+140 -115
View File
@@ -26,6 +26,14 @@
/* Selection fill for page text and the CodeMirror editors; themes /* Selection fill for page text and the CodeMirror editors; themes
override when the accent tint clashes with accent-colored text. */ override when the accent tint clashes with accent-colored text. */
--selection-bg: color-mix(var(--accent) 30%, transparent); --selection-bg: color-mix(var(--accent) 30%, transparent);
/* Referer-badge chip in the analytics viewer: a fixed neutral palette,
deliberately NOT themed — the chip sits behind transparent favicons, so
black-on-transparent and white-on-transparent glyphs must both stay
legible on every theme (dark greys kill black glyphs, pure white kills
white ones); its text likewise stays dark on any theme. */
--badge-bg: #c9d1d9;
--badge-text: #1f2328;
--badge-muted: #59636e;
/* Code highlighting palette, consumed by pygments.css: complete light and /* Code highlighting palette, consumed by pygments.css: complete light and
dark sets (background included), resolved by light-dark() from the dark sets (background included), resolved by light-dark() from the
used color-scheme. A theme picks a set simply by declaring used color-scheme. A theme picks a set simply by declaring
@@ -288,8 +296,10 @@ body {
/* Sidebar + main row. A symmetric grid: the article column is sized by the /* Sidebar + main row. A symmetric grid: the article column is sized by the
viewport alone (never by content), with equally sized flexible gutters viewport alone (never by content), with equally sized flexible gutters
on both sides. The sidebar sits in the left gutter, so it appearing or on both sides. The sidebar sits in the start gutter (grid columns are
disappearing never shifts the article; the right gutter balances it. flow-relative: on RTL pages the whole composition mirrors), so it
appearing or disappearing never shifts the article; the end gutter
balances it.
The outer tracks are minmax(0, 1fr) — a plain 1fr has an `auto` minimum, The outer tracks are minmax(0, 1fr) — a plain 1fr has an `auto` minimum,
which let the 12rem sidebar expand its track at narrow widths and push which let the 12rem sidebar expand its track at narrow widths and push
the article off-center; now the sidebar overlays the gutter edge instead the article off-center; now the sidebar overlays the gutter edge instead
@@ -305,16 +315,16 @@ body {
content length — code excluded) lift the 78rem cap: main takes the full content length — code excluded) lift the 78rem cap: main takes the full
width and the article composes itself inside it — fluid, bounded text width and the article composes itself inside it — fluid, bounded text
lanes with the surplus left vacant (see the article layout rules lanes with the surplus left vacant (see the article layout rules
below). The left track — main always sits in column 2 — collapses to below). The start track — main always sits in column 2 — collapses to
zero when the page has no sidebar; the sidebar then simply overlays zero when the page has no sidebar; the sidebar then simply overlays
the vacant zone, as it does on single-column pages. */ the vacant zone, as it does on single-column pages. */
body:has(.multicol) #content { body:has(.multicol) #content {
grid-template-columns: 0 minmax(0, 1fr); grid-template-columns: 0 minmax(0, 1fr);
} }
/* With a sidebar the left lane gets its own track at every width, so the /* With a sidebar the start lane gets its own track at every width, so the
sidebar never overlaps the article and the article leans on the sidebar never overlaps the article and the article leans on the
viewport's right edge (surplus extends the lane). The lane is flexible: viewport's end edge (surplus extends the lane). The lane is flexible:
12rem when space is tight, growing up to 150% (18rem) once the viewport 12rem when space is tight, growing up to 150% (18rem) once the viewport
exceeds the article's 88.5rem (86rem + main's side padding). The --lane exceeds the article's 88.5rem (86rem + main's side padding). The --lane
variable doubles as the measure for the margin boxes and the .wide variable doubles as the measure for the margin boxes and the .wide
@@ -398,7 +408,7 @@ body.editing #sidebar {
#sidebar { #sidebar {
grid-column: 1; grid-column: 1;
/* Pinned to the page's left edge (not the article's) and kept in view /* Pinned to the page's start edge (not the article's) and kept in view
while scrolling. Translucent + blurred rather than an opaque box, so while scrolling. Translucent + blurred rather than an opaque box, so
full-bleed .wide images can pass underneath without a hard edge. */ full-bleed .wide images can pass underneath without a hard edge. */
justify-self: start; justify-self: start;
@@ -409,8 +419,11 @@ body.editing #sidebar {
width: 12rem; width: 12rem;
max-height: 100vh; max-height: 100vh;
overflow-y: auto; overflow-y: auto;
padding: 1rem 1rem 1rem 1.25rem; /* The extra inline-start padding (the viewport-edge side, mirroring with
border-radius: 0 0 0.5rem 0; the direction) matches main's side padding. */
padding: 1rem;
padding-inline: 1.25rem 1rem;
border-end-start-radius: 0.5rem;
background: color-mix(var(--bg) 75%, transparent); background: color-mix(var(--bg) 75%, transparent);
backdrop-filter: blur(0.5rem); backdrop-filter: blur(0.5rem);
} }
@@ -437,7 +450,7 @@ body.editing #sidebar {
#sidebar ul ul li::before { #sidebar ul ul li::before {
content: "🔹"; content: "🔹";
display: inline-block; display: inline-block;
margin-left: -1.3em; margin-inline-start: -1.3em;
width: 1.3em; width: 1.3em;
} }
@@ -578,7 +591,7 @@ main {
background: color-mix(var(--bg) 30%, transparent); background: color-mix(var(--bg) 30%, transparent);
color: var(--muted); color: var(--muted);
font-size: 0.95rem; font-size: 0.95rem;
text-align: left; text-align: start;
hyphens: none; hyphens: none;
} }
@@ -679,19 +692,19 @@ article * + h6 {
so the gap becomes the figure's margin plus the full --list-indent. */ so the gap becomes the figure's margin plus the full --list-indent. */
article :is(ul, ol:not([type])) { article :is(ul, ol:not([type])) {
list-style: none; list-style: none;
padding-left: 0; padding-inline-start: 0;
--list-indent: 2em; --list-indent: 2em;
display: flow-root; display: flow-root;
} }
article :is(ul, ol:not([type])) > li { article :is(ul, ol:not([type])) > li {
padding-left: var(--list-indent); padding-inline-start: var(--list-indent);
} }
article ul li::before { article ul li::before {
content: "🔹"; content: "🔹";
display: inline-block; display: inline-block;
margin-left: calc(-1 * var(--list-indent)); margin-inline-start: calc(-1 * var(--list-indent));
width: var(--list-indent); width: var(--list-indent);
text-align: center; text-align: center;
} }
@@ -713,13 +726,13 @@ article ol:not([type]) > li::before {
content: counter(item) "."; content: counter(item) ".";
color: var(--muted); color: var(--muted);
display: inline-block; display: inline-block;
margin-left: calc(-1 * var(--list-indent)); margin-inline-start: calc(-1 * var(--list-indent));
width: var(--list-indent); width: var(--list-indent);
text-align: left; text-align: start;
} }
/* Task lists: real clickable checkboxes. The checkbox stands in for the /* Task lists: real clickable checkboxes. The checkbox stands in for the
list marker — taken out of flow, left-aligned in the indent box and list marker — taken out of flow, start-aligned in the indent box and
centered on the first line's middle, so the item text starts at the centered on the first line's middle, so the item text starts at the
same edge as every other list item's. */ same edge as every other list item's. */
article .task-list-item { article .task-list-item {
@@ -732,7 +745,7 @@ article .task-list-item::before {
article .task-list-item-checkbox { article .task-list-item-checkbox {
position: absolute; position: absolute;
left: 0; inset-inline-start: 0;
top: calc(0.5lh - .15ex); top: calc(0.5lh - .15ex);
translate: 0 -50%; translate: 0 -50%;
font-size: inherit; font-size: inherit;
@@ -751,6 +764,20 @@ article {
position: relative; position: relative;
} }
/* Emoji/symbol icon buttons and links: dim until hovered. */
.icon-btn {
padding: 0;
font: inherit;
background: none;
border: none;
cursor: pointer;
opacity: 0.7;
}
.icon-btn:hover {
opacity: 1;
}
.edit-link { .edit-link {
position: absolute; position: absolute;
top: 0.2rem; top: 0.2rem;
@@ -758,12 +785,6 @@ article {
left: -2.2rem; left: -2.2rem;
z-index: 2; z-index: 2;
/* stay above full-bleed .wide images */ /* stay above full-bleed .wide images */
font: inherit;
background: none;
border: none;
padding: 0;
cursor: pointer;
opacity: 0.7;
text-shadow: 0 0 0.1em black; text-shadow: 0 0 0.1em black;
} }
@@ -772,7 +793,7 @@ article h1 .edit-link {
position: static; position: static;
font-size: 1.1rem; font-size: 1.1rem;
vertical-align: 0.3em; vertical-align: 0.3em;
margin-left: 0.4rem; margin-inline-start: 0.4rem;
} }
/* Section pens sit at the end of anchored h2s, dimmer than the page pen /* Section pens sit at the end of anchored h2s, dimmer than the page pen
@@ -781,28 +802,18 @@ article h2 .edit-section {
position: static; position: static;
font-size: 0.85rem; font-size: 0.85rem;
vertical-align: 0.35em; vertical-align: 0.35em;
margin-left: 0.4rem; margin-inline-start: 0.4rem;
opacity: 0.35; opacity: 0.35;
} }
.edit-link:hover {
opacity: 1;
}
/* Login/profile links injected by pagerite.js when Paskia SSO is in use. /* Login/profile links injected by pagerite.js when Paskia SSO is in use.
They live inside the .editor-pens flex container in the banner's top-right They live inside the .editor-pens flex container in the banner's top-right
corner and inherit its reset; keep only their opacity/text-shadow tweaks. */ corner and inherit its reset; keep only their text-shadow tweak. */
.editor-pens a.login-link, .editor-pens a.login-link,
.editor-pens a.profile-link { .editor-pens a.profile-link {
opacity: 0.7;
text-shadow: 0 0 0.1em black; text-shadow: 0 0 0.1em black;
} }
.editor-pens a.login-link:hover,
.editor-pens a.profile-link:hover {
opacity: 1;
}
article p, article p,
article li, article li,
article dd { article dd {
@@ -811,10 +822,10 @@ article dd {
} }
/* The long-article composition (.multicol): a fluid but bounded text /* The long-article composition (.multicol): a fluid but bounded text
lane with a 16rem side zone at the article's left, centered in main — lane with a 16rem side zone at the article's start side, centered in
surplus width becomes vacant space, never endless text (with a sidebar main — surplus width becomes vacant space, never endless text (with a
the article leans right instead and the sidebar's track is the left sidebar the article leans to the end edge instead and the sidebar's
lane; see below). (The backend render splits the body into .colseg track is the start lane; see below). (The backend render splits the body into .colseg
segments separated by full-width h2s and .wide elements, tags segments separated by full-width h2s and .wide elements, tags
text-heavy segments of several paragraphs .cols — a ::: nocols text-heavy segments of several paragraphs .cols — a ::: nocols
container opts its section out, and column-filling paragraphs are container opts its section out, and column-filling paragraphs are
@@ -832,7 +843,7 @@ article.multicol {
@container (min-width: 45rem) { @container (min-width: 45rem) {
/* The side zone (not on phones): lane content indents 16rem; margin /* The side zone (not on phones): lane content indents 16rem; margin
boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures) boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures)
are taken out of flow and placed against the article's left edge — are taken out of flow and placed against the article's start edge —
the same region the nav sidebar overlays. The boxes stay in the the same region the nav sidebar overlays. The boxes stay in the
column segment at their anchor point (the backend render no longer column segment at their anchor point (the backend render no longer
splits segments around them); absolute positioning off the article — splits segments around them); absolute positioning off the article —
@@ -842,14 +853,14 @@ article.multicol {
article.multicol>.colseg, article.multicol>.colseg,
article.multicol>h1, article.multicol>h1,
article.multicol>h2 { article.multicol>h2 {
margin-left: 16rem; margin-inline-start: 16rem;
} }
article.multicol .margin, article.multicol .margin,
article.multicol .aside, article.multicol .aside,
article.multicol figure.margin { article.multicol figure.margin {
position: absolute; position: absolute;
left: 0; inset-inline-start: 0;
width: 14rem; width: 14rem;
max-width: none; max-width: none;
margin: 0.3rem 0 0; margin: 0.3rem 0 0;
@@ -870,11 +881,11 @@ article.multicol {
} }
} }
/* With a sidebar, the sidebar's 12rem track IS the left lane at every /* With a sidebar, the sidebar's 12rem track IS the start lane at every
width (see #content): no in-article zone, the text lane runs fluid (up width (see #content): no in-article zone, the text lane runs fluid (up
to 86rem) and leans on main's right edge — surplus width extends the to 86rem) and leans on main's end edge — surplus width extends the
left lane instead of balancing out on the right — and margin boxes start lane instead of balancing out at the end — and margin boxes
hang into the lane off the article's left border, sliding under the hang into the lane off the article's start border, sliding under the
translucent sticky nav, which only ever occupies its top. (Not below translucent sticky nav, which only ever occupies its top. (Not below
48rem: there the sidebar becomes a link strip above the article and 48rem: there the sidebar becomes a link strip above the article and
there is no lane to fall into.) */ there is no lane to fall into.) */
@@ -884,8 +895,8 @@ article.multicol {
margin-inline: auto 0; margin-inline: auto 0;
} }
/* The sidebar fills the flexible lane (its left side stays on the /* The sidebar fills the flexible lane (its start side stays on the
viewport's left edge, growing rightward). */ viewport's start edge, growing toward the article). */
body:has(#sidebar):has(.multicol):not(.editing) #sidebar { body:has(#sidebar):has(.multicol):not(.editing) #sidebar {
width: 100%; width: 100%;
} }
@@ -893,29 +904,29 @@ article.multicol {
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.colseg, body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.colseg,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h1, body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h1,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h2 { body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h2 {
margin-left: 0; margin-inline-start: 0;
} }
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .margin, body:has(#sidebar):has(.multicol):not(.editing) article.multicol .margin,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .aside, body:has(#sidebar):has(.multicol):not(.editing) article.multicol .aside,
body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure.margin { body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure.margin {
position: absolute; position: absolute;
/* Attached to the article's left border (1.25rem gap), hanging into /* Attached to the article's start border (1.25rem gap), hanging into
the left lane and growing leftward with it: 12rem when the lane is the start lane and growing with it: 12rem when the lane is
tight, up to 150% (18rem) when the track or the surplus has room tight, up to 150% (18rem) when the track or the surplus has room
(100cqw - 100% is the surplus left of the right-leaning article). (100cqw - 100% is the surplus beside the end-leaning article).
The lane (track + main's padding) always guarantees the room. */ The lane (track + main's padding) always guarantees the room. */
--box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem); --box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem);
width: var(--box-w); width: var(--box-w);
max-width: none; max-width: none;
left: calc(-1.25rem - var(--box-w)); inset-inline-start: calc(-1.25rem - var(--box-w));
margin: 0.3rem 0 0; margin: 0.3rem 0 0;
} }
} }
/* A shrink-wrapped figure (explicit image width) centers in the plain /* A shrink-wrapped figure (explicit image width) centers in the plain
layout; inside a column the centering looks adrift — left-align. layout; inside a column the centering looks adrift — align to the start
Floated figures keep their own margins (the text gap). */ edge. Floated figures keep their own margins (the text gap). */
.multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), .margin) { .multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), .margin) {
margin-inline: 0; margin-inline: 0;
} }
@@ -988,13 +999,15 @@ article a:hover {
/* Blockquotes: spacing comes from the blockquote itself (bottom-only like /* Blockquotes: spacing comes from the blockquote itself (bottom-only like
everything else in articles); inner paragraphs keep only the gap between everything else in articles); inner paragraphs keep only the gap between
them. The negative left margin pushes the bar out past the text edge, so them. The negative start margin pushes the bar out past the text edge, so
quoted text aligns with the surrounding paragraphs — same trick as code quoted text aligns with the surrounding paragraphs — same trick as code
blocks. */ blocks. */
blockquote { blockquote {
margin: 0 0 1rem -0.5rem; margin: 0 0 1rem;
padding: 0 0 0 0.25rem; margin-inline-start: -0.5rem;
border-left: 0.25rem solid var(--accent2); padding: 0;
padding-inline-start: 0.25rem;
border-inline-start: 0.25rem solid var(--accent2);
color: var(--muted); color: var(--muted);
} }
@@ -1009,17 +1022,19 @@ blockquote p + p {
/* Admonitions (markdown !!! note/warning/...) and GitHub-style alerts /* Admonitions (markdown !!! note/warning/...) and GitHub-style alerts
(> [!NOTE] ...): a lightweight callout in the blockquote idiom — accent (> [!NOTE] ...): a lightweight callout in the blockquote idiom — accent
bar and a faint wash, recolored per type, with a type emoji on the bar and a faint wash, recolored per type, with a type emoji on the
title. The negative left margin pushes bar and wash out past the text title. The negative start margin pushes bar and wash out past the text
edge so the inner text aligns with surrounding paragraphs — same trick edge so the inner text aligns with surrounding paragraphs — same trick
as blockquotes and code blocks (margin-left = border + padding-left). as blockquotes and code blocks (margin-inline-start = border +
Bottom-only margins like everything else in articles; inner paragraphs padding-inline-start). Bottom-only margins like everything else in
carry no margins of their own. */ articles; inner paragraphs carry no margins of their own. */
.admonition, .admonition,
.markdown-alert { .markdown-alert {
margin: 0 0 1rem -1.15rem; margin: 0 0 1rem;
margin-inline-start: -1.15rem;
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
border-left: 0.25rem solid var(--admonition-color, var(--accent)); border-inline-start: 0.25rem solid var(--admonition-color, var(--accent));
border-radius: 0 0.3rem 0.3rem 0; border-start-end-radius: 0.3rem;
border-end-end-radius: 0.3rem;
background: color-mix(var(--admonition-color, var(--accent)) 7%, transparent); background: color-mix(var(--admonition-color, var(--accent)) 7%, transparent);
} }
@@ -1037,7 +1052,7 @@ blockquote p + p {
.admonition-title::before, .admonition-title::before,
.markdown-alert-title::before { .markdown-alert-title::before {
padding-right: 0.35em; padding-inline-end: 0.35em;
} }
.admonition.note .admonition-title::before, .admonition.note .admonition-title::before,
@@ -1099,18 +1114,19 @@ blockquote p + p {
Where the layout has room for a side zone — multicol pages, the Where the layout has room for a side zone — multicol pages, the
sidebar's track, the wide single-column gutter (see the article sidebar's track, the wide single-column gutter (see the article
section and the figure rules below) — the boxes are taken out of flow section and the figure rules below) — the boxes are taken out of flow
and absolutely positioned into it, off the article's left border, each and absolutely positioned into it, off the article's start border, each
at the vertical spot where it occurs in the text (boxes occurring at the vertical spot where it occurs in the text (boxes occurring
closer together than their heights may overlap — keep them apart); closer together than their heights may overlap — keep them apart);
otherwise they stay in-column left floats (consecutive floats stack otherwise they stay in-column start floats (consecutive floats stack
via clear: left). Headings already clear floats, so in-column boxes via clear: inline-start). Headings already clear floats, so in-column
never bleed into the next section. */ boxes never bleed into the next section. */
.aside { .aside {
float: left; float: inline-start;
clear: left; clear: inline-start;
width: 30%; width: 30%;
max-width: 20rem; max-width: 20rem;
margin: 0.3rem 1.2rem 1rem 0; margin: 0.3rem 0 1rem;
margin-inline-end: 1.2rem;
padding: 0.6rem 0.9rem; padding: 0.6rem 0.9rem;
font-size: 0.9rem; font-size: 0.9rem;
color: var(--muted); color: var(--muted);
@@ -1140,11 +1156,12 @@ blockquote p + p {
} }
.margin { .margin {
float: left; float: inline-start;
clear: left; clear: inline-start;
width: 30%; width: 30%;
max-width: 20rem; max-width: 20rem;
margin: 0.3rem 1.2rem 1rem 0; margin: 0.3rem 0 1rem;
margin-inline-end: 1.2rem;
font-size: 0.9rem; font-size: 0.9rem;
color: var(--muted); color: var(--muted);
} }
@@ -1153,10 +1170,10 @@ pre {
overflow-x: auto; overflow-x: auto;
padding: 0.5rem 0.8rem; padding: 0.5rem 0.8rem;
/* Code text aligns with the surrounding paragraphs: the box extends /* Code text aligns with the surrounding paragraphs: the box extends
past them by its own padding. Themes that add a left border must past them by its own padding. Themes that add a leading border must
extend margin-left by the border width to keep this alignment. */ extend margin-inline-start by the border width to keep this
margin-left: -0.8rem; alignment. */
margin-right: -0.8rem; margin-inline: -0.8rem;
background: var(--code-bg); background: var(--code-bg);
border-radius: 4px; border-radius: 4px;
position: relative; position: relative;
@@ -1196,7 +1213,7 @@ p code {
} }
code:not(pre code):first-child { code:not(pre code):first-child {
padding-left: 0; padding-inline-start: 0;
} }
/* Click-to-copy button (added by pagerite.js) */ /* Click-to-copy button (added by pagerite.js) */
@@ -1241,7 +1258,7 @@ td {
} }
th { th {
text-align: left; text-align: start;
background: linear-gradient(180deg, background: linear-gradient(180deg,
var(--table-head-a, color-mix(var(--accent) 10%, var(--surface))), var(--table-head-a, color-mix(var(--accent) 10%, var(--surface))),
var(--table-head-b, color-mix(var(--accent) 18%, var(--surface)))); var(--table-head-b, color-mix(var(--accent) 18%, var(--surface))));
@@ -1313,20 +1330,24 @@ figure img:not([width]) {
width: 100%; width: 100%;
} }
/* Floated figures: {.right} / {.left}, defaulting to 30% of the column /* Floated figures: {.right} / {.left} float to the text column's end/start
and capped at half of it. */ edge — the class names are author-facing and fixed, but the sides follow
the text direction (in RTL, .left floats right) — defaulting to 30% of
the column and capped at half of it. */
figure:has(.right) { figure:has(.right) {
float: right; float: inline-end;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 0 1rem 1em; margin: 0.3rem 0 1rem;
margin-inline-start: 1em;
} }
figure:has(.left) { figure:has(.left) {
float: left; float: inline-start;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 1em 1rem 0; margin: 0.3rem 0 1rem;
margin-inline-end: 1em;
} }
/* The same floats for other blocks: ::: left / ::: right containers /* The same floats for other blocks: ::: left / ::: right containers
@@ -1334,17 +1355,19 @@ figure:has(.left) {
tables all take the class directly ({.right} at the end of a tables all take the class directly ({.right} at the end of a
paragraph's last line, a trailing {.left} line after a fence, ...). */ paragraph's last line, a trailing {.left} line after a fence, ...). */
:is(div, p, pre, blockquote, table).right { :is(div, p, pre, blockquote, table).right {
float: right; float: inline-end;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 0 1rem 1em; margin: 0.3rem 0 1rem;
margin-inline-start: 1em;
} }
:is(div, p, pre, blockquote, table).left { :is(div, p, pre, blockquote, table).left {
float: left; float: inline-start;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 1em 1rem 0; margin: 0.3rem 0 1rem;
margin-inline-end: 1em;
} }
/* An image with an explicit width attribute shrink-wraps instead: the /* An image with an explicit width attribute shrink-wraps instead: the
@@ -1356,13 +1379,15 @@ figure:has(img[width]) {
} }
/* {.margin} figures (the class moves onto the figure wrapper at render — /* {.margin} figures (the class moves onto the figure wrapper at render —
see markdown.py) float left like {.left} ones — until they fall into see markdown.py) float to the start edge like {.left} ones — until
the side zone (see the composition rules up in the article section). */ they fall into the side zone (see the composition rules up in the
article section). */
figure.margin { figure.margin {
float: left; float: inline-start;
width: 30%; width: 30%;
max-width: 50%; max-width: 50%;
margin: 0.3rem 1em 1rem 0; margin: 0.3rem 0 1rem;
margin-inline-end: 1em;
} }
/* Click-to-enlarge (pagerite.js): article figure images open in a /* Click-to-enlarge (pagerite.js): article figure images open in a
@@ -1439,12 +1464,12 @@ article figure img {
} }
} }
/* Wide single-column pages: margin boxes lean into the vacant left /* Wide single-column pages: margin boxes lean into the vacant start-side
gutter instead (below 104rem the gutter cannot hold the box, and while gutter instead (below 104rem the gutter cannot hold the box, and while
editing the docked panel reshapes the gutters — in both they stay editing the docked panel reshapes the gutters — in both they stay
plain floats). Out of flow like on multicol pages: the box hangs off plain floats). Out of flow like on multicol pages: the box hangs off
the article's left border, growing with the gutter up to 150% (18rem), the article's start border, growing with the gutter up to 150% (18rem),
its right side 1.25rem off the border. */ its end side 1.25rem off the border. */
@media (min-width: 104rem) { @media (min-width: 104rem) {
body:not(.editing):not(:has(.multicol)) article .margin, body:not(.editing):not(:has(.multicol)) article .margin,
body:not(.editing):not(:has(.multicol)) article .aside, body:not(.editing):not(:has(.multicol)) article .aside,
@@ -1453,14 +1478,14 @@ article figure img {
--box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem); --box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem);
width: var(--box-w); width: var(--box-w);
max-width: none; max-width: none;
left: calc(-1.25rem - var(--box-w)); inset-inline-start: calc(-1.25rem - var(--box-w));
margin: 0.3rem 0 0; margin: 0.3rem 0 0;
} }
} }
/* In the wide symmetric gutters (where the sidebar overlays the flexible /* In the wide symmetric gutters (where the sidebar overlays the flexible
left gutter rather than a reserved track) the sidebar flexes with the start gutter rather than a reserved track) the sidebar flexes with the
gutter up to 150% — its left side stays on the viewport's edge. */ gutter up to 150% — its start side stays on the viewport's edge. */
@media (min-width: 102rem) { @media (min-width: 102rem) {
body:not(.editing) #sidebar { body:not(.editing) #sidebar {
width: min(18rem, 100%); width: min(18rem, 100%);
@@ -1509,7 +1534,7 @@ body.editing pre.wide {
/* Narrow single-column pages with a sidebar: below 102rem the symmetric /* Narrow single-column pages with a sidebar: below 102rem the symmetric
gutters can no longer both hold the sidebar, so #content reserves it gutters can no longer both hold the sidebar, so #content reserves it
with a flexible left track (see the matching media query below) and the with a flexible start track (see the matching media query below) and the
article always starts at the lane's width (+ main's 1.25rem padding) — article always starts at the lane's width (+ main's 1.25rem padding) —
the bleed margin measures off --lane. Scoped by :has(#sidebar) since the bleed margin measures off --lane. Scoped by :has(#sidebar) since
the sidebar element is omitted entirely on pages without the sidebar element is omitted entirely on pages without
@@ -1536,9 +1561,9 @@ body:has(.multicol) pre.wide {
} }
/* Multicol with a sidebar track (≥48rem, see #content): main starts at /* Multicol with a sidebar track (≥48rem, see #content): main starts at
the flexible lane's width and the article leans right, so the bleed the flexible lane's width and the article leans to the end edge, so the
extends left past the surplus and the lane to the true viewport edge — bleed extends past the surplus and the lane to the true viewport start
sliding under the translucent sidebar — and right past main's edge — sliding under the translucent sidebar — and past main's end
padding. */ padding. */
@media (min-width: 48rem) { @media (min-width: 48rem) {
body:has(#sidebar):has(.multicol):not(.editing) figure:has(.wide), body:has(#sidebar):has(.multicol):not(.editing) figure:has(.wide),
@@ -1556,7 +1581,7 @@ figcaption {
hyphens: auto; hyphens: auto;
-webkit-hyphens: auto; -webkit-hyphens: auto;
text-wrap: pretty; text-wrap: pretty;
text-align: left; text-align: start;
/* Never let a long caption stretch a shrink-to-fit figure wider than the /* Never let a long caption stretch a shrink-to-fit figure wider than the
image; the caption wraps at the figure's width instead. */ image; the caption wraps at the figure's width instead. */
width: 0; width: 0;
@@ -1582,7 +1607,7 @@ article h2 {
/* Narrow windows with a sidebar: below 102rem the symmetric gutters can no /* Narrow windows with a sidebar: below 102rem the symmetric gutters can no
longer both hold the 12rem sidebar, so reserve its space with a flexible longer both hold the 12rem sidebar, so reserve its space with a flexible
left track instead of letting it overlap the article (multicol pages use start track instead of letting it overlap the article (multicol pages use
the same track at every width — see the #content rules above; their the same track at every width — see the #content rules above; their
higher-specificity rule wins there). The lane is 12rem when space is higher-specificity rule wins there). The lane is 12rem when space is
tight, growing up to 150% (18rem) once the viewport exceeds the tight, growing up to 150% (18rem) once the viewport exceeds the
@@ -1702,10 +1727,10 @@ article h2 {
width: fit-content; width: fit-content;
} }
/* The single-column sidebar .wide margins assume a left sidebar column; /* The single-column sidebar .wide margins assume a start-side sidebar
with the sidebar on top the article is viewport-wide and the plain column; with the sidebar on top the article is viewport-wide and the
centered bleed applies again. (Multicol pages need no override: their plain centered bleed applies again. (Multicol pages need no override:
cqw bleed is exact at any width.) */ their cqw bleed is exact at any width.) */
body:has(#sidebar):not(.editing):not(:has(.multicol)) figure:has(.wide) { body:has(#sidebar):not(.editing):not(:has(.multicol)) figure:has(.wide) {
margin-inline: calc(50% - 50vw); margin-inline: calc(50% - 50vw);
} }
@@ -1746,7 +1771,7 @@ article h2 {
.dateline { .dateline {
color: var(--muted); color: var(--muted);
font-size: 0.85rem; font-size: 0.85rem;
text-align: left; text-align: start;
} }
.footnotes { .footnotes {
+24
View File
@@ -0,0 +1,24 @@
// Shared popup open-state behavior: while `open` (a ref, truthy = open)
// is set, a pointerdown outside `root` (a template ref covering both the
// toggle button and the popup) or Escape resets it to null. One logic for
// every dropdown (LangSelect, the page editor's class/table pickers), so
// they can't drift apart.
import { onBeforeUnmount, watch } from 'vue'
export function usePopup(open, root) {
let off = null
const stop = watch(open, (v) => {
off?.()
off = null
if (!v) return
const down = (ev) => { if (!root.value?.contains(ev.target)) open.value = null }
const key = (ev) => { if (ev.key === 'Escape') open.value = null }
addEventListener('pointerdown', down, true)
addEventListener('keydown', key)
off = () => {
removeEventListener('pointerdown', down, true)
removeEventListener('keydown', key)
}
})
onBeforeUnmount(() => { off?.(); stop() })
}
+10 -5
View File
@@ -1,10 +1,15 @@
// The editor shell's shared language selection ('' = the primary language): // The editor shell's shared language selection ('' = the primary language):
// one state, v-modeled by the LangSelect of every tab that has one (page, // backed by the app-wide store (./store), so the editor tabs' LangSelects
// structure). While the panel is open it also drives the page preview — // and the public corner selector bind the same value. Linked to the
// EditorShell applies it as the fetch-time language override (swapdoc). // whole-page language: while the panel is open it drives the page preview
import { ref } from 'vue' // (EditorShell applies it as the fetch-time language override, swapdoc).
import { computed, ref } from 'vue'
import { pinia, useStore } from './store'
export const editorLang = ref('') export const editorLang = computed({
get: () => useStore(pinia).lang,
set: (v) => { useStore(pinia).lang = v },
})
// The CURRENT PAGE's primary language ('' = not yet learned): the shell's // The CURRENT PAGE's primary language ('' = not yet learned): the shell's
// settings fetch fills it with the site default; the page/structure tabs // settings fetch fills it with the site default; the page/structure tabs
+14
View File
@@ -30,6 +30,20 @@ export const LANG_GROUPS = [
const displayNames = new Intl.DisplayNames(['en'], { type: 'language' }) const displayNames = new Intl.DisplayNames(['en'], { type: 'language' })
// Consistent menu ordering for language selectors: the geographic/cultural
// grouping above (similar languages sit together, and it does not vary with
// the display language the way alphabetical-by-name would). Tags outside
// the groups trail, ordered by tag. The primary language is not special
// here — callers put it first themselves.
const groupOrder = new Map(LANG_GROUPS.flat().map((c, i) => [c, i]))
export function langSort(codes) {
return [...codes].sort(
(a, b) =>
(groupOrder.get(a) ?? groupOrder.size) - (groupOrder.get(b) ?? groupOrder.size)
|| a.localeCompare(b),
)
}
// English display name for a language tag ("fi" -> "Finnish"). // English display name for a language tag ("fi" -> "Finnish").
export function langName(tag) { export function langName(tag) {
try { try {
+50
View File
@@ -0,0 +1,50 @@
// Public language-selector entry: imported on demand by pagerite.js on
// pages advertising more than one language in their hreflang alternates.
// Vue, Pinia and the flag SVG set live in this chunk only — untranslated
// pages never pay for them. The selector's state lives in the shared
// store (./store), not the DOM: the corner container is rebuilt freely
// and ensureMounted re-mounts from the store.
import { createApp } from 'vue'
import LangSelector from './LangSelector.vue'
import { pinia, useStore } from './store'
let app = null
function store() {
return useStore(pinia)
}
// The current page's languages (called on every navigation).
export function setLanguages(alternates, current) {
Object.assign(store(), {
langAlternates: alternates,
servedLang: current,
langSelectorActive: true,
})
}
// The current page is single-language: the selector goes away.
export function hide() {
store().langSelectorActive = false
app?.unmount()
app = null
}
// Mount the selector as the container's first item; re-mount when its
// element went away with a container rebuild (a live app updates from the
// store reactively).
export function ensureMounted(host) {
if (!store().langSelectorActive || !host) {
app?.unmount()
app = null
return
}
if (app && host.contains(app._container)) return
app?.unmount()
const el = document.createElement('div')
el.id = 'lang-selector'
host.prepend(el)
app = createApp(LangSelector)
app.use(pinia)
app.mount(el)
}
+116 -33
View File
@@ -51,13 +51,18 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
url.searchParams.delete("lang"); url.searchParams.delete("lang");
history.replaceState(history.state, "", url); history.replaceState(history.state, "", url);
} }
// The session language. While the editor panel is open, its language // The session language: the user's explicit pick (initial ?lang=, public
// selection overrides the normal preference (swapdoc.setLangOverride): // selector, editor dropdown) is kept in chosenLang; while the editor is
// internal fetches and prefetches follow it until the panel closes and // open its selection overrides it (swapdoc.setLangOverride), and closing
// the override clears (null restores the initial ?lang=, if any). // falls back to chosenLang. JS state only — pretty URLs, no reloads.
// window.__pageriteLang is the pin for swapdoc.loadPlain's fetches.
let chosenLang = langParam;
let sessionLang = langParam; let sessionLang = langParam;
window.__pageriteLang = sessionLang;
addEventListener("pagerite:session-lang", (ev) => { addEventListener("pagerite:session-lang", (ev) => {
sessionLang = ev.detail?.lang || langParam; if (ev.detail?.lang) chosenLang = ev.detail.lang;
sessionLang = ev.detail?.lang || chosenLang;
window.__pageriteLang = sessionLang;
}); });
// An internal URL as fetched: carries the session's ?lang= unless the // An internal URL as fetched: carries the session's ?lang= unless the
// link already pins a language of its own. With no ?lang= on the initial // link already pins a language of its own. With no ?lang= on the initial
@@ -127,16 +132,16 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
if (line != null) { if (line != null) {
// Section pen on an anchored h2: opens the page editor at the // Section pen on an anchored h2: opens the page editor at the
// section's markdown source line (data-line, from the backend). // section's markdown source line (data-line, from the backend).
btn.className = "edit-link edit-section"; btn.className = "edit-link edit-section icon-btn";
btn.title = "edit section"; btn.title = "edit section";
btn.textContent = "🖊️"; btn.textContent = "🖊️";
btn.dataset.editorLine = line; btn.dataset.editorLine = line;
} else if (mode === "page") { } else if (mode === "page") {
btn.className = "edit-link edit-page"; btn.className = "edit-link edit-page icon-btn";
btn.title = "edit page"; btn.title = "edit page";
btn.textContent = "🖊️"; btn.textContent = "🖊️";
} else { } else {
btn.className = "edit-link site-edit-link"; btn.className = "edit-link site-edit-link icon-btn";
btn.title = "site settings"; btn.title = "site settings";
btn.textContent = "⚙️"; btn.textContent = "⚙️";
} }
@@ -160,13 +165,29 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
function makeAuthLink(admin) { function makeAuthLink(admin) {
const a = document.createElement("a"); const a = document.createElement("a");
a.className = admin ? "profile-link" : "login-link"; a.className = (admin ? "profile-link" : "login-link") + " icon-btn";
a.href = "/auth/"; a.href = "/auth/";
a.title = admin ? "profile" : "log in"; a.title = admin ? "profile" : "log in";
a.textContent = admin ? "\u{1F510}" : "\u{1F511}"; a.textContent = admin ? "\u{1F510}" : "\u{1F511}";
return a; return a;
} }
// The banner top-right corner container: the language selector (first
// item) plus the admin pens and auth links. renderAuthUi rebuilds it from
// scratch; the selector's state lives in the shared store, not the DOM,
// so the langselect bundle re-mounts it into the fresh container.
function pensContainer() {
let pens = document.querySelector(".editor-pens");
if (!pens) {
const banner = document.getElementById("page-banner");
if (!banner) return null;
pens = document.createElement("div");
pens.className = "editor-pens";
banner.after(pens);
}
return pens;
}
function removePens() { function removePens() {
document.querySelectorAll(".editor-pens, #main article button.edit-link") document.querySelectorAll(".editor-pens, #main article button.edit-link")
.forEach((el) => el.remove()); .forEach((el) => el.remove());
@@ -178,32 +199,33 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
// pens that may have been injected while the browser cache made us look // pens that may have been injected while the browser cache made us look
// authenticated. // authenticated.
removePens(); removePens();
if (!authReady) return; if (authReady) {
// Editing is open for admins and, as a dev/no-proxy fallback, when no
// Editing is open for admins and, as a dev/no-proxy fallback, when no // Paskia SSO is detected at all.
// Paskia SSO is detected at all. const canEdit = isAdmin || !ssoAvailable;
const canEdit = isAdmin || !ssoAvailable; // The analytics page is a read-only dashboard: editing pens and the side
// The analytics page is a read-only dashboard: editing pens and the side // panel do not apply there. Login/logout links are still useful.
// panel do not apply there. Login/logout links are still useful. const onAnalytics = currentPath === "/_a";
const onAnalytics = currentPath === "/_a"; if (document.getElementById("page-banner")) {
const banner = document.getElementById("page-banner"); const pens = pensContainer();
if (banner) { if (canEdit && !onAnalytics) {
const pens = document.createElement("div"); // Analytics viewer is now a normal page at /_a.
pens.className = "editor-pens"; const a = document.createElement("a");
if (canEdit && !onAnalytics) { a.className = "edit-link analytics-link icon-btn";
// Analytics viewer is now a normal page at /_a. a.href = "/_a";
const a = document.createElement("a"); a.title = "analytics";
a.className = "edit-link analytics-link"; a.textContent = "📊";
a.href = "/_a"; pens.append(a);
a.title = "analytics"; pens.append(makePen("site"));
a.textContent = "📊"; }
pens.append(a); if (ssoAvailable) pens.append(makeAuthLink(isAdmin));
pens.append(makePen("site")); if (!pens.firstElementChild) pens.remove();
} }
if (ssoAvailable) pens.append(makeAuthLink(isAdmin)); if (canEdit && !onAnalytics) injectPagePen();
banner.after(pens);
} }
if (canEdit && !onAnalytics) injectPagePen(); // Re-mount the selector into the fresh container (no-op until the
// bundle has been loaded once).
langselectMod?.ensureMounted(document.querySelector(".editor-pens"));
} }
async function setupAuth() { async function setupAuth() {
@@ -407,6 +429,9 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
// copy pinned to a language (?lang=) caches under its own key, where // copy pinned to a language (?lang=) caches under its own key, where
// navigation with the same session language finds it. // navigation with the same session language finds it.
pageCache.set(rawKey(ev.detail.url), ev.detail.html); pageCache.set(rawKey(ev.detail.url), ev.detail.html);
// Editor-driven swaps don't go through load(): re-evaluate the
// language selector from the fresh copy too.
mountLangselect(new DOMParser().parseFromString(ev.detail.html, "text/html"));
}); });
// Editors mutate site-wide state (theme, structure, headings, banners), // Editors mutate site-wide state (theme, structure, headings, banners),
@@ -612,6 +637,8 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
if (to) msg.to = to; if (to) msg.to = to;
const secs = Math.round(read / 1000); const secs = Math.round(read / 1000);
if (secs > 0) msg.read = secs; if (secs > 0) msg.read = secs;
const lang = document.documentElement.lang;
if (lang) msg.lang = lang;
if (!msg.to && !msg.read) return; if (!msg.to && !msg.read) return;
report(msg); report(msg);
} }
@@ -742,6 +769,60 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
} }
} }
// --- Public language selector ------------------------------------------
// Pages translated into more than one language advertise it via hreflang
// alternates (x-default + one link per language). Those pages get the
// editors' flag dropdown as the first item of the corner container; its
// bundle (Vue + the flag SVG set) loads on demand. Re-evaluated from the
// fresh document on every swap (the head's own alternates stay stale).
let langselectMod = null;
async function mountLangselect(doc) {
const links = [...doc.head.querySelectorAll('link[rel="alternate"][hreflang]')];
const dflt = links.find((l) => l.hreflang === "x-default");
const langs = links.filter((l) => l.hreflang && l.hreflang !== "x-default");
if (!dflt || langs.length <= 1) return langselectMod?.hide();
try {
langselectMod ??= await import(/* @vite-ignore */ assets["pagerite:langselect-src"]);
for (const css of (assets["pagerite:langselect-css"] || "").split(",")) {
if (css && !document.querySelector(`link[href="${css}"]`)) {
const link = document.createElement("link");
link.rel = "stylesheet";
link.href = css;
link.dataset.pagerite = "langselect-css";
document.head.append(link);
}
}
langselectMod.setLanguages(
// The original's alternate is the plain URL — x-default's href —
// which also marks it as the primary option.
langs.map((l) => ({ tag: l.hreflang, href: l.href, primary: l.href === dflt.href })),
doc.documentElement.lang,
);
langselectMod.ensureMounted(pensContainer());
} catch (e) {
console.error("language selector mount failed:", e);
}
}
// The selector's pick (LangSelector dispatches this): make it the
// session language and swap the page in place. With the editor open the
// pick already landed in the shared store — the editor's watch re-renders
// the page itself, so there is nothing to do here.
addEventListener("pagerite:set-session-lang", async (ev) => {
const tag = ev.detail?.lang;
if (!tag || tag === sessionLang) return;
if (document.body.classList.contains("editing")) return;
chosenLang = sessionLang = tag;
window.__pageriteLang = tag;
const y = scrollY; // a language switch is not a navigation: keep scroll
await load(currentPath, false);
scrollTo(0, y);
// Log the switch as a trail event in the new language (load() updated
// <html lang>): the ping matches the switch's GET server-side, so it is
// not misclassified as a crawler hit.
ping({ to: currentPath });
});
// --- Fetch navigation ------------------------------------------------ // --- Fetch navigation ------------------------------------------------
async function load(url, push = true, back = false) { async function load(url, push = true, back = false) {
// Navigating with the editor open closes it; unsaved edits are lost // Navigating with the editor open closes it; unsaved edits are lost
@@ -843,6 +924,7 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
runScripts(document.getElementById("main")); runScripts(document.getElementById("main"));
applyEffects(); applyEffects();
mountAnalytics(doc); mountAnalytics(doc);
mountLangselect(doc);
}; };
// Rotating cube page transition (styles injected as #pagerite-transition // Rotating cube page transition (styles injected as #pagerite-transition
// from the selected design's transition.css, e.g. themes/cube/); // from the selected design's transition.css, e.g. themes/cube/);
@@ -1076,4 +1158,5 @@ import { reconnectPolicy, socketSlot, watchConnecting } from "./reconnect";
setupAuth(); setupAuth();
applyEffects(); applyEffects();
mountAnalytics(document); mountAnalytics(document);
mountLangselect(document);
})(); })();
+26
View File
@@ -0,0 +1,26 @@
// The app's shared Pinia store — cross-bundle UI state lives here. Every
// entry chunk imports its own copy of this module, so the Pinia instance
// is parked on window (Vue itself is a shared chunk, so reactivity works
// across the copies). Pass `pinia` explicitly when calling useStore
// outside a component (module code, no active instance).
import { createPinia, defineStore } from 'pinia'
export const pinia = (window.__pageritePinia ??= createPinia())
export const useStore = defineStore('pagerite', {
state: () => ({
// The ONE language selection, v-modeled by both dropdowns (editor
// tabs, public corner selector): '' = no explicit pick (the page's
// primary / autodetect), else a concrete tag. A pick from either
// dropdown is visible to everyone immediately.
lang: '',
// The language the current page was actually served in (set by
// pagerite.js per navigation) — the selector's highlight fallback
// when there is no explicit pick.
servedLang: '',
// The public selector's page data: hreflang alternates
// ([{tag, href, primary}]) and whether to show at all.
langAlternates: [],
langSelectorActive: false,
}),
})
+12 -9
View File
@@ -12,12 +12,13 @@ export function dropPageCache() {
} }
// The editor's language override (set by EditorShell): while the panel is // The editor's language override (set by EditorShell): while the panel is
// open, its language selection wins over the normal preferences (?lang= / // open, its language selection wins over the normal preferences — every
// Accept-Language) — every in-place re-render asks for that language // in-place re-render asks for that language explicitly, and pagerite.js
// explicitly, and pagerite.js applies it to its own fetches and prefetches // applies it to its own fetches and prefetches (pagerite:session-lang).
// (pagerite:session-lang). The primary selection pins by its code: // The primary selection pins by its code: ?lang=<primary> selects the
// ?lang=<primary> selects the original explicitly (i18n.select_language). // original explicitly (i18n.select_language). Panel closed, the session's
let overrideLang = null // the ?lang= value in force, null = normal prefs // chosen language (window.__pageriteLang) takes over — the pick stays.
let overrideLang = null // the ?lang= value in force, null = the session's
export function setLangOverride(queryLang) { export function setLangOverride(queryLang) {
overrideLang = queryLang || null overrideLang = queryLang || null
@@ -129,14 +130,16 @@ function swapRegions(doc) {
// Fetch /p, swap its regions into the live page and replaceState to it. // Fetch /p, swap its regions into the live page and replaceState to it.
// Returns the final URL (after redirects), or null when the fetch did not // Returns the final URL (after redirects), or null when the fetch did not
// yield a page. Category and missing URLs render a placeholder 404 page — // yield a page. Category and missing URLs render a placeholder 404 page —
// fine to swap in (new pages are created by editing them). While the // fine to swap in (new pages are created by editing them). The fetch pins
// editor's language override is set the fetch pins that language. // the editor's language override, or — panel closed — the session's chosen
// language (window.__pageriteLang).
export async function loadPlain(p) { export async function loadPlain(p) {
let doc let doc
let finalUrl = `/${p}` let finalUrl = `/${p}`
let html let html
try { try {
const res = await fetch(overrideLang ? `${finalUrl}?lang=${overrideLang}` : finalUrl) const pin = overrideLang || window.__pageriteLang
const res = await fetch(pin ? `${finalUrl}?lang=${pin}` : finalUrl)
const type = res.headers.get('content-type') || '' const type = res.headers.get('content-type') || ''
if (!type.includes('text/html')) return null if (!type.includes('text/html')) return null
if (res.redirected) finalUrl = res.url if (res.redirected) finalUrl = res.url
+5 -5
View File
@@ -8,11 +8,11 @@
* - Disables Vite's screen clearing on startup * - Disables Vite's screen clearing on startup
* *
* Options: * Options:
* paths - Array of paths to proxy (default: ["/api"]) * paths - Array of paths to proxy (default: ['/api'])
*/ */
export default function fastapiVue({ paths = ["/api"] } = {}) { export default function fastapiVue({ paths = ['/api'] } = {}) {
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:8210" const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:8210'
// Build proxy configuration for each path // Build proxy configuration for each path
const proxy = {} const proxy = {}
@@ -25,12 +25,12 @@ export default function fastapiVue({ paths = ["/api"] } = {}) {
} }
return { return {
name: "vite-plugin-fastapi-pagerite", name: 'vite-plugin-fastapi-pagerite',
config: () => ({ config: () => ({
clearScreen: false, clearScreen: false,
server: { proxy }, server: { proxy },
build: { build: {
outDir: "../pagerite/frontend-build", outDir: '../pagerite/frontend-build',
emptyOutDir: true, emptyOutDir: true,
}, },
}), }),
+3 -2
View File
@@ -5,7 +5,7 @@ import { defineConfig } from 'vite'
import vue from '@vitejs/plugin-vue' import vue from '@vitejs/plugin-vue'
import vueDevTools from 'vite-plugin-vue-devtools' import vueDevTools from 'vite-plugin-vue-devtools'
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200' const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:8210'
// Proxy everything except Vite's own dev-time paths and the backend machinery // Proxy everything except Vite's own dev-time paths and the backend machinery
// to the FastAPI backend in dev. /_api, /_f, /_themes, /_fonts and /_a are // to the FastAPI backend in dev. /_api, /_f, /_themes, /_fonts and /_a are
@@ -38,7 +38,7 @@ export default defineConfig({
chunkSizeWarningLimit: 1200, chunkSizeWarningLimit: 1200,
// Mirror the URL space in the build output: hashed files land under // Mirror the URL space in the build output: hashed files land under
// frontend-build/_assets/ and the Frontend serves the build directory // frontend-build/_assets/ and the Frontend serves the build directory
// at the site root (frontend/public/favicon.ico -> /favicon.ico). // at the site root.
manifest: true, manifest: true,
assetsDir: '_assets', assetsDir: '_assets',
rollupOptions: { rollupOptions: {
@@ -50,6 +50,7 @@ export default defineConfig({
main: fileURLToPath(new URL('./src/main.js', import.meta.url)), main: fileURLToPath(new URL('./src/main.js', import.meta.url)),
pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)), pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)),
analytics: fileURLToPath(new URL('./src/analytics-main.js', import.meta.url)), analytics: fileURLToPath(new URL('./src/analytics-main.js', import.meta.url)),
langselect: fileURLToPath(new URL('./src/langselect-main.js', import.meta.url)),
// Only the base CSS is built; theme/banner-design stylesheets live // Only the base CSS is built; theme/banner-design stylesheets live
// in pagerite/themes/{name}/ and are served by the backend as-is. // in pagerite/themes/{name}/ and are served by the backend as-is.
pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)), pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)),
+18 -80
View File
@@ -1,79 +1,16 @@
"""Command-line entry point for running the backend server.""" """Command-line entry point for running the backend server."""
import argparse import argparse
import gzip
import os import os
import sys
from datetime import date
from pathlib import Path from pathlib import Path
import httpx import msgspec
from fastapi_vue import server from fastapi_vue import env, server
from fastapi_vue.hostutil import parse_endpoints
from pagerite.config import Config
DEFAULT_PORT = 8100 DEFAULT_PORT = 8100
DEVMODE = os.getenv("PAGERITE_DEV") == "1" os.environ["FASTAPI_VUE"] = "PAGERITE"
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
_REPO_ROOT = Path(__file__).resolve().parent.parent
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
def _download_dbip() -> None:
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
today = date.today()
months = [f"{today:%Y-%m}"]
# The current month's file may not be published yet; fall back to last month.
prev = (today.replace(day=1) - date.resolution).replace(day=1)
months.append(f"{prev:%Y-%m}")
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
print(
f"pagerite: DB-IP database is current ({existing[-1]}), skipping download"
)
return
for month in months:
url = DBIP_URL.format(month=month)
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
print(f"pagerite: downloading {url}")
try:
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
if r.status_code == 404:
continue
r.raise_for_status()
with open(tmp, "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk)
except httpx.HTTPError as e:
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
tmp.unlink(missing_ok=True)
continue
# Verify it is actually gzip data before installing it.
try:
with gzip.open(tmp, "rb") as f:
f.read(1)
except OSError:
print(
f"pagerite: DB-IP download for {month} was not valid gzip",
file=sys.stderr,
)
tmp.unlink(missing_ok=True)
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
print(f"pagerite: DB-IP database updated to {target.name}")
return
print("pagerite: could not download a DB-IP database", file=sys.stderr)
def main() -> None: def main() -> None:
@@ -101,23 +38,24 @@ def main() -> None:
help="Download/update the DB-IP city lite database before starting.", help="Download/update the DB-IP city lite database before starting.",
) )
args = parser.parse_args() args = parser.parse_args()
# Export the hostname before pagerite.app is imported: it derives the # Hand configuration to the app as JSON in PAGERITE_CONFIG; it must be
# data directory and public origin from it at import time. # set before pagerite.app is imported, as state.py reads it at import
os.environ["PAGERITE_HOSTNAME"] = args.hostname # time (data directory, public origin).
# And the listen port: the app prints the translator WS URL at startup, os.environ["PAGERITE_CONFIG"] = msgspec.json.encode(
# which for localhost includes the actual port. Config(hostname=args.hostname, dbip=args.dbip)
for endpoint in parse_endpoints(args.listen, DEFAULT_PORT): ).decode()
if "port" in endpoint: run_args: dict = {}
os.environ["PAGERITE_PORT"] = str(endpoint["port"]) if args.hostname != "localhost":
break # A public site sits behind TLS on its hostname; show that URL in the
if args.dbip: # startup box instead of the local listen address.
_download_dbip() run_args["startup_box"] = f"{{Name}} {{version}}\nhttps://{args.hostname}"
server.run( server.run(
"pagerite.app:app", "pagerite.app:app",
listen=args.listen, listen=args.listen,
default_port=DEFAULT_PORT, default_port=DEFAULT_PORT,
server_header=False, server_header=False,
reload=Path(__file__).parent if DEVMODE else False, reload=Path(__file__).parent if env.dev else False,
**run_args,
) )
+619 -540
View File
File diff suppressed because it is too large Load Diff
+16 -15
View File
@@ -130,9 +130,7 @@ async def save_page(
if i18n.add_patch(data, node, path, lang, page.markdown): if i18n.add_patch(data, node, path, lang, page.markdown):
_invalidate_pages() _invalidate_pages()
return return
with kanta.transaction( with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
"page", user=request.headers.get("remote-user"), extra=path
):
node = _ensure(data.menu, path) node = _ensure(data.menu, path)
node.title = page.title node.title = page.title
node.chunks = store_chunks(data.chunks, page.markdown) node.chunks = store_chunks(data.chunks, page.markdown)
@@ -242,9 +240,7 @@ async def update_structure(op: StructureOp, request: Request) -> None:
if op.title is not None if op.title is not None
else "structure:reorder" else "structure:reorder"
) )
with kanta.transaction( with kanta.transaction(action, user=request.headers.get("remote-user"), extra=path):
action, user=request.headers.get("remote-user"), extra=path
):
if op.title is not None: if op.title is not None:
node.title = op.title node.title = op.title
if target is not None and target != path: if target is not None and target != path:
@@ -302,14 +298,13 @@ class SettingsIn(BaseModel):
brand_html: str = "" brand_html: str = ""
transition: str = "cube" transition: str = "cube"
translate_langs: list[str] | None = None # None keeps the current set translate_langs: list[str] | None = None # None keeps the current set
translate_keys: dict[str, str] | None = None # None keeps the current keys
@router.put("/_api/settings", status_code=204) @router.put("/_api/settings", status_code=204)
async def put_settings(settings: SettingsIn, request: Request) -> None: async def put_settings(settings: SettingsIn, request: Request) -> None:
"""Update site-wide settings; invalidates cached pages and ETags.""" """Update site-wide settings; invalidates cached pages and ETags."""
with kanta.transaction( with kanta.transaction("settings", user=request.headers.get("remote-user")):
"settings", user=request.headers.get("remote-user")
):
data.brand = settings.brand data.brand = settings.brand
data.brand_html = settings.brand_html data.brand_html = settings.brand_html
data.theme = settings.theme data.theme = settings.theme
@@ -324,6 +319,8 @@ async def put_settings(settings: SettingsIn, request: Request) -> None:
for lang in settings.translate_langs for lang in settings.translate_langs
if (tag := i18n.base_tag(lang)) if (tag := i18n.base_tag(lang))
} }
if settings.translate_keys is not None:
data.translate_keys = settings.translate_keys
_invalidate_pages() _invalidate_pages()
@@ -335,9 +332,7 @@ async def delete_translations(request: Request) -> None:
translators). User patches are kept; the availability index translators). User patches are kept; the availability index
(node.langs) is rebuilt from them — patches alone still make a language (node.langs) is rebuilt from them — patches alone still make a language
exist on a page.""" exist on a page."""
with kanta.transaction( with kanta.transaction("translate:reset", user=request.headers.get("remote-user")):
"translate:reset", user=request.headers.get("remote-user")
):
i18n.clear_translations(data) i18n.clear_translations(data)
_invalidate_pages() _invalidate_pages()
# Fragments rejected this run (segment validation) stay skipped no # Fragments rejected this run (segment validation) stay skipped no
@@ -376,9 +371,7 @@ async def toggle_task_endpoint(body: ToggleTaskIn, request: Request) -> dict[str
new_markdown = toggle_task(node_markdown(data, node) or "", body.index) new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
if new_markdown is None: if new_markdown is None:
raise HTTPException(400, "invalid task index") raise HTTPException(400, "invalid task index")
with kanta.transaction( with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
"page", user=request.headers.get("remote-user"), extra=path
):
# Re-chunk like any save: only the chunk containing the toggled # Re-chunk like any save: only the chunk containing the toggled
# checkbox gets a new hash, the rest keep theirs. # checkbox gets a new hash, the rest keep theirs.
node.chunks = store_chunks(data.chunks, new_markdown) node.chunks = store_chunks(data.chunks, new_markdown)
@@ -510,6 +503,14 @@ async def editor_ws(ws: WebSocket) -> None:
# The title is injected as h1 when the markdown has # The title is injected as h1 when the markdown has
# none; the editor's title field edits live-preview. # none; the editor's title field edits live-preview.
title=msg.get("title") or (node.title if node else ""), title=msg.get("title") or (node.title if node else ""),
# Pin section anchors to the original language so the
# preview of a translation matches the served page
# (no-op when the previewed markdown is the original).
anchors_from=(
(node_markdown(data, node) or "", node.title)
if node
else None
),
) )
await ws.send_json( await ws.send_json(
{ {
+44 -7
View File
@@ -30,18 +30,50 @@ walking the tree (``resolve``), moves are slot detach/attach
import asyncio import asyncio
import logging import logging
from collections.abc import AsyncGenerator
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from pathlib import Path
from fastapi import FastAPI, Request from fastapi import FastAPI, Request
from fastapi.responses import Response from fastapi.responses import Response
from fastapi_vue import Frontend, env
from starlette.types import ASGIApp, Receive, Scope, Send
from pagerite import api, files, pages, tracking from pagerite import api, files, pages, tracking
from pagerite.__main__ import DEVMODE
from pagerite.files import file_store from pagerite.files import file_store
from pagerite.state import analytics_store, frontend, kanta from pagerite.state import analytics_store, config, kanta
from collections.abc import AsyncGenerator
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Vue build served at the site root, no SPA catch-all (assets only). The
# build mirrors the URL space: hashed, immutable files live under
# /_assets/ (assetsDir: '_/assets').
frontend = Frontend(
Path(__file__).with_name("frontend-build"), spa=False, cached="/_assets/"
)
class _AccessLogExtraMiddleware:
"""Fill the ``log_extra`` slot of fastapi_vue's access log.
Everything under ``/_api`` is gated by the SSO forward-auth, which names
the authenticated user in the ``remote-user`` header; put that user on
the access-log line, for plain requests and WebSocket open/close alike.
The scope dict is shared with the outer AccessLogMiddleware, which reads
the slot back at response/accept/close time.
"""
def __init__(self, app: ASGIApp) -> None:
self.app = app
async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
if scope["type"] in ("http", "websocket") and scope["path"].startswith("/_api"):
headers = dict(scope["headers"])
user = headers.get(b"remote-user", b"").decode("latin-1")
if user:
scope.setdefault("state", {})["log_extra"] = user
await self.app(scope, receive, send)
@asynccontextmanager @asynccontextmanager
async def lifespan(_app: FastAPI) -> AsyncGenerator: async def lifespan(_app: FastAPI) -> AsyncGenerator:
@@ -49,8 +81,11 @@ async def lifespan(_app: FastAPI) -> AsyncGenerator:
async with kanta: async with kanta:
await asyncio.to_thread(file_store.load) await asyncio.to_thread(file_store.load)
await frontend.load() await frontend.load()
# Decompress/open the DB-IP MMDB once at startup. Lookups are then # --dbip: update the DB-IP database first, then decompress/open the
# read-only and safe to run in background ``to_thread`` workers. # MMDB once. Lookups are then read-only and safe to run in
# background ``to_thread`` workers.
if config.dbip:
await asyncio.to_thread(tracking._download_dbip)
await asyncio.to_thread(tracking._geoip._load) await asyncio.to_thread(tracking._geoip._load)
analytics_store.subscribe(tracking._schedule_analytics_broadcast) analytics_store.subscribe(tracking._schedule_analytics_broadcast)
# Backfill favicons for external sites already in the recorded data. # Backfill favicons for external sites already in the recorded data.
@@ -63,13 +98,15 @@ async def lifespan(_app: FastAPI) -> AsyncGenerator:
# is not meant to be browsable by the public anyway. # is not meant to be browsable by the public anyway.
app = FastAPI( app = FastAPI(
title="Pagerite", title="Pagerite",
debug=DEVMODE, debug=env.dev,
lifespan=lifespan, lifespan=lifespan,
docs_url=None, docs_url=None,
redoc_url=None, redoc_url=None,
openapi_url=None, openapi_url=None,
) )
app.add_middleware(_AccessLogExtraMiddleware)
@app.middleware("http") @app.middleware("http")
async def _headers(request: Request, call_next) -> Response: async def _headers(request: Request, call_next) -> Response:
@@ -86,7 +123,7 @@ app.include_router(tracking.router)
app.include_router(files.router) app.include_router(files.router)
# Vue build asset routes are inserted at this position during load(): the # Vue build asset routes are inserted at this position during load(): the
# build mirrors the URL space (/_assets/*, /favicon.ico at the root). # build mirrors the URL space (/_assets/*).
frontend.route(app, "/") frontend.route(app, "/")
# The content catch-all goes last: built assets win over content slugs, # The content catch-all goes last: built assets win over content slugs,
+19 -3
View File
@@ -17,6 +17,13 @@ from pagerite.segments import has_prose
#: backticks or tildes (CommonMark). #: backticks or tildes (CommonMark).
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})") _FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
#: A container fence line (mdit-py-plugins container): the "::: aside"
#: opener and the ":::" closer alike. Always its own block, even with no
#: blank line around it: folded into a prose paragraph it would cross to
#: the translator as part of the text run, where the model can drop it —
#: the rest of the page then renders inside the container.
_CONTAINER = re.compile(r"^ {0,3}:{3,}(?:[ \t]|$)")
#: HTML block openers that may span blank lines (CommonMark types 1-5: #: HTML block openers that may span blank lines (CommonMark types 1-5:
#: script/pre/style/textarea, comments, processing instructions, #: script/pre/style/textarea, comments, processing instructions,
#: declarations, CDATA) with their closing condition. Other HTML blocks #: declarations, CDATA) with their closing condition. Other HTML blocks
@@ -54,9 +61,11 @@ def chunk_markdown(markdown: str) -> list[str]:
Blocks are separated by blank lines; fenced code blocks and the Blocks are separated by blank lines; fenced code blocks and the
multi-line HTML blocks (comments, script/pre/style, CDATA...) are multi-line HTML blocks (comments, script/pre/style, CDATA...) are
kept atomic, even across blank lines, and end at their closing kept atomic, even across blank lines, and end at their closing
condition. Chunks carry no surrounding blank lines and no trailing condition. Container fence lines (:::, open and close alike) are
newline; rejoining with ``join_chunks`` reproduces the source modulo always their own block, blank lines or not (see _CONTAINER). Chunks
blank-line normalization. carry no surrounding blank lines and no trailing newline; rejoining
with ``join_chunks`` reproduces the source modulo blank-line
normalization.
""" """
chunks: list[str] = [] chunks: list[str] = []
buf: list[str] = [] buf: list[str] = []
@@ -91,6 +100,13 @@ def chunk_markdown(markdown: str) -> list[str]:
fence = m.group(1) fence = m.group(1)
buf.append(line) buf.append(line)
continue continue
if _CONTAINER.match(line):
# Container fence lines (open and close alike) are their own
# block — never part of a prose chunk (see _CONTAINER).
flush()
buf.append(line)
flush()
continue
if not buf: if not buf:
for open_re, close_re in _HTML_ATOMIC: for open_re, close_re in _HTML_ATOMIC:
if open_re.match(line): if open_re.match(line):
+28
View File
@@ -0,0 +1,28 @@
"""CLI → app configuration, passed as JSON in the ``PAGERITE_CONFIG`` env var.
Kept dependency-free (msgspec only) so ``__main__`` can build and serialize
the config before any app module is imported, and the app side parses the
same struct back. Import-time safe: nothing here reads the environment
until ``load()`` is called.
"""
import os
import msgspec
class Config(msgspec.Struct):
"""Configuration passed from the CLI entry point to the app."""
#: Public hostname of the site; names the per-site data directory
#: ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
hostname: str = "localhost"
#: Download/update the DB-IP city lite database at startup (--dbip).
dbip: bool = False
def load() -> Config:
"""Parse ``PAGERITE_CONFIG``, or the defaults when unset."""
if raw := os.getenv("PAGERITE_CONFIG"):
return msgspec.json.decode(raw.encode(), type=Config)
return Config()
+4 -4
View File
@@ -98,14 +98,14 @@ class Data(msgspec.Struct):
#: Trusted author content; not sanitized. #: Trusted author content; not sanitized.
custom_css: str = "" custom_css: str = ""
#: Favicon: content-addressed file name (served at "/_f/{name}"), #: Favicon: content-addressed file name (served at "/_f/{name}"),
#: linked as <link rel="icon"> on every page. Empty = the build's #: linked as <link rel="icon"> on every page; /favicon.ico redirects
#: /favicon.ico. #: to it. Empty = no icon (and /favicon.ico 404s).
favicon: str = "" favicon: str = ""
#: API keys gating the translator service WebSocket (/_translate/{key}; #: API keys gating the translator service WebSocket (/_translate/{key};
#: the external forward-auth does not cover that route): key -> display #: the external forward-auth does not cover that route): key -> display
#: name. Keys are 12 lowercase alphanumeric characters; the first is #: name. Keys are 12 lowercase alphanumeric characters; the first is
#: generated at database bootstrap (see app.py), multiple keys are a #: generated at database bootstrap, more are managed in the editor
#: future reservation (e.g. managed via a web interface). #: shell's lang tab (via /_api/settings).
translate_keys: dict[str, str] = {} translate_keys: dict[str, str] = {}
#: Wanted target languages for the translator service (presence-keys, #: Wanted target languages for the translator service (presence-keys,
#: value always True). The dispatcher offers jobs only in the #: value always True). The dispatcher offers jobs only in the
+21 -10
View File
@@ -6,7 +6,8 @@ when compression shrinks the body), served immutable at ``/_f/``. Raster
images and SVGs are recompressed into AVIF/WebP/JPEG derivatives images and SVGs are recompressed into AVIF/WebP/JPEG derivatives
(``store_image`` and helpers); the untouched original is kept alongside as (``store_image`` and helpers); the untouched original is kept alongside as
``<hash>.orig<ext>`` (never served). Routes: upload/delete under ``<hash>.orig<ext>`` (never served). Routes: upload/delete under
``/_api/files``, the favicon settings endpoints, the ``/_f/`` server with ``/_api/files``, the favicon settings endpoints, the /favicon.ico
redirect to the configured icon, the ``/_f/`` server with
Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``). Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``).
""" """
@@ -19,7 +20,7 @@ from pathlib import Path
import blake3 import blake3
from fastapi import APIRouter, HTTPException, Request from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import Response from fastapi.responses import RedirectResponse, Response
from mediapreview import dispatch from mediapreview import dispatch
from pagerite import views from pagerite import views
@@ -38,10 +39,6 @@ from pagerite.state import (
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# mediapreview logs pyvips noise ("VipsForeignSaveJpegTarget argument strip is
# deprecated", "threadpool completed with N workers") at INFO; keep warnings.
logging.getLogger("mediapreview").setLevel(logging.WARNING)
router = APIRouter() router = APIRouter()
@@ -158,14 +155,13 @@ def _svg_to_png(body: bytes, maxsize: int) -> bytes | None:
def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes: def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes:
"""Re-encode the AVIF derivative into a fallback format (WebP/JPEG) """Re-encode the AVIF derivative into a fallback format (WebP/JPEG)
via pyvips. JPEG has no alpha, so it is flattened onto white; via pyvips. JPEG has no alpha, so it is flattened onto white."""
``strip`` keeps metadata (EXIF) out of the fallbacks."""
import pyvips import pyvips
img = pyvips.Image.new_from_buffer(avif, "") img = pyvips.Image.new_from_buffer(avif, "")
if suffix == ".jpg" and img.hasalpha(): if suffix == ".jpg" and img.hasalpha():
img = img.flatten(background=[255, 255, 255]) img = img.flatten(background=[255, 255, 255])
return img.write_to_buffer(suffix, Q=quality, strip=True) return img.write_to_buffer(suffix, Q=quality, keep="none")
def _image_derivatives( def _image_derivatives(
@@ -252,6 +248,20 @@ async def delete_file(name: str) -> None:
file_store.delete(name) file_store.delete(name)
@router.get("/favicon.ico", include_in_schema=False)
async def favicon_ico() -> Response:
"""The conventional /favicon.ico: redirect to the configured site icon.
Browsers request this path on their own (tabs, bookmarks, feeds and
other non-HTML contexts) regardless of the <link rel="icon"> pages
carry. Redirect to the icon's store URL, which negotiates the format
and caches immutably; 404 when no custom icon is configured.
"""
if not data.favicon:
raise HTTPException(404)
return RedirectResponse(f"/_f/{data.favicon}")
@router.put("/_api/settings/favicon") @router.put("/_api/settings/favicon")
async def put_favicon(request: Request) -> dict[str, str]: async def put_favicon(request: Request) -> dict[str, str]:
"""Upload a favicon into the content-addressed store and activate it. """Upload a favicon into the content-addressed store and activate it.
@@ -276,7 +286,8 @@ async def put_favicon(request: Request) -> dict[str, str]:
@router.delete("/_api/settings/favicon", status_code=204) @router.delete("/_api/settings/favicon", status_code=204)
async def delete_favicon(request: Request) -> None: async def delete_favicon(request: Request) -> None:
"""Clear the custom favicon (back to the build's /favicon.ico). """Clear the custom favicon (/favicon.ico goes back to 404, pages drop
the <link rel="icon">).
The blob stays in the content-addressed store; only the reference goes. The blob stays in the content-addressed store; only the reference goes.
""" """
+55 -7
View File
@@ -402,7 +402,9 @@ def _heading_ids(state) -> None:
its self-link is ``href=""`` (back to the top of the page). An its self-link is ``href=""`` (back to the top of the page). An
author-set `{#id}` always wins; auto ids slugify the heading text author-set `{#id}` always wins; auto ids slugify the heading text
(python-slugify, mirroring the editor's slugify.js) and dedupe with (python-slugify, mirroring the editor's slugify.js) and dedupe with
-2/-3 suffixes per render. Headings that already contain a link are -2/-3 suffixes per render — unless env["anchor_ids"] presets them, as
render(anchors_from=...) does for translated pages so section URLs
stay in the original language. Headings that already contain a link are
``data-line`` records the heading's markdown source line (0-based, after ``data-line`` records the heading's markdown source line (0-based, after
undoing the render(title=...) injection offset via ``env``) — the page undoing the render(title=...) injection offset via ``env``) — the page
editor uses it for section pens and piecewise-linear scroll sync. editor uses it for section pens and piecewise-linear scroll sync.
@@ -443,15 +445,25 @@ def _heading_ids(state) -> None:
if len(heads) < ANCHOR_MIN_HEADINGS: if len(heads) < ANCHOR_MIN_HEADINGS:
return return
seen: set[str] = set() seen: set[str] = set()
for i, token in heads: preset = state.env.get("anchor_ids")
for k, (i, token) in enumerate(heads):
inline = tokens[i + 1] inline = tokens[i + 1]
hid = token.attrGet("id") hid = token.attrGet("id")
if not isinstance(hid, str) or not hid: if not isinstance(hid, str) or not hid:
# Slug the visible text, not the raw markdown (`## [a](url)`). if preset is not None and k < len(preset):
text = "".join( # Translated render: the original language's slug, matched
c.content for c in inline.children if c.type in ("text", "code_inline") # by heading position (a translation never adds, removes or
) # reorders headings; a patched one that does falls back to
base = slugify(text) or "section" # slugging its own text past the end of the list).
base = preset[k]
else:
# Slug the visible text, not the raw markdown (`## [a](url)`).
text = "".join(
c.content
for c in inline.children
if c.type in ("text", "code_inline")
)
base = slugify(text) or "section"
hid, n = base, 2 hid, n = base, 2
while hid in seen: while hid in seen:
hid = f"{base}-{n}" hid = f"{base}-{n}"
@@ -463,6 +475,36 @@ def _heading_ids(state) -> None:
wrap(i, token, f"#{hid}") wrap(i, token, f"#{hid}")
def anchor_ids(text: str, title: str | None = None) -> list[str]:
"""The section anchor ids of text, in heading order.
render(anchors_from=...) feeds these to _heading_ids via
env["anchor_ids"], pinning a translated render's anchors to the
original language's slugs. The selection mirrors _heading_ids exactly
(the same md instance assigns the ids during this parse, author-set
{#id} included as-is); the in-body title h1 is excluded.
"""
if title and not has_h1(text):
text = f"# {title}\n\n{text}"
tokens = md.parse(text, {"page_path": ""})
first_h1 = next(
(
i
for i, t in enumerate(tokens)
if t.type == "heading_open" and t.tag == "h1" and t.level == 0
),
None,
)
return [
t.attrGet("id")
for i, t in enumerate(tokens)
if t.type == "heading_open"
and t.tag in ("h1", "h2")
and t.level == 0
and i != first_h1
]
def make_md(*, verbatim: bool = False) -> MarkdownIt: def make_md(*, verbatim: bool = False) -> MarkdownIt:
"""A fully configured parser. The module-level ``md`` (below) is the """A fully configured parser. The module-level ``md`` (below) is the
render instance; ``verbatim=True`` builds the segmentation instance for render instance; ``verbatim=True`` builds the segmentation instance for
@@ -618,12 +660,16 @@ def render(
created: datetime | None = None, created: datetime | None = None,
modified: datetime | None = None, modified: datetime | None = None,
title: str | None = None, title: str | None = None,
anchors_from: tuple[str, str] | None = None,
) -> Rendered: ) -> Rendered:
"""Render Markdown text to the article body's HTML and layout flags. """Render Markdown text to the article body's HTML and layout flags.
``title`` injects a ``# {title}`` line at the top when the markdown has ``title`` injects a ``# {title}`` line at the top when the markdown has
no h1 of its own, so the implicit page title goes through the exact no h1 of its own, so the implicit page title goes through the exact
same pipeline as an explicit one (first-h1 anchor treatment included). same pipeline as an explicit one (first-h1 anchor treatment included).
``anchors_from`` is the (markdown, title) of the ORIGINAL language when
rendering a translation: section anchors are pinned to its slugs so
localized pages keep the original #hash URLs.
The top-level blocks are grouped into column segments: boundary blocks The top-level blocks are grouped into column segments: boundary blocks
(h1/h2 headings, .wide — see _is_boundary) are rendered bare, the runs (h1/h2 headings, .wide — see _is_boundary) are rendered bare, the runs
@@ -641,6 +687,8 @@ def render(
right after the article's h1. right after the article's h1.
""" """
env = {"page_path": page_path, "line_offset": 0} env = {"page_path": page_path, "line_offset": 0}
if anchors_from is not None:
env["anchor_ids"] = anchor_ids(*anchors_from)
if title and not has_h1(text): if title and not has_h1(text):
text = f"# {title}\n\n{text}" text = f"# {title}\n\n{text}"
# The injected title shifts source lines by two; _heading_ids # The injected title shifts source lines by two; _heading_ids
+12 -35
View File
@@ -3,11 +3,11 @@
``GET /{path:path}`` resolves a slug path against the menu tree and renders ``GET /{path:path}`` resolves a slug path against the menu tree and renders
the page (or a category placeholder, or 404); it must be registered AFTER the page (or a category placeholder, or 404); it must be registered AFTER
the fastapi-vue asset routes so built frontend files win over content slugs the fastapi-vue asset routes so built frontend files win over content slugs
(see app.py). Requests are recorded in analytics (crawler hits and 404s (see app.py). Every served document is recorded raw in analytics (one
here, visits via the /_ws socket in tracking.py). access-log line with its true HTTP status; classification happens at
display time — see pagerite/analytics.py).
""" """
import asyncio
import logging import logging
from datetime import UTC, datetime from datetime import UTC, datetime
from email.utils import format_datetime from email.utils import format_datetime
@@ -22,16 +22,9 @@ from pagerite.state import (
SITE_URL, SITE_URL,
_html_response, _html_response,
_is_reserved, _is_reserved,
analytics_store,
data, data,
) )
from pagerite.tracking import ( from pagerite.tracking import _record_get
_client_ip,
_enrich_client,
_query_suffix,
_schedule_client_enrichment,
_track_entry,
)
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -146,27 +139,21 @@ async def show_page(request: Request, path: str) -> Response:
placeholder page (nav links point straight at its first child). placeholder page (nav links point straight at its first child).
""" """
path = path.strip("/") path = path.strip("/")
ua = request.headers.get("user-agent", "")
accept_language = request.headers.get("accept-language", "") accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path): if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its # Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page. # built-in 404 instead of rendering an editable article page.
# Scanner telltales (dotpaths like /.env, *.php) classify the IP # Recorded like any other GET: telltale scanner paths (dotpaths
# as abuse in analytics. # like /.env, *.php) classify the IP as abuse at display time.
client_hash = analytics_store.track_404( _record_get(request, status=404)
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
raise HTTPException(404) raise HTTPException(404)
chain = resolve(data.menu, path) chain = resolve(data.menu, path)
node = chain[-1] if chain else None node = chain[-1] if chain else None
if node is not None and node.published and node.chunks is not None: if node is not None and node.published and node.chunks is not None:
# Language selection (docs/localization.md): ?lang= wins when a # Language selection (docs/localization.md): ?lang= wins when a
# translation exists, else header logic. Analytics keep the raw # translation exists, else header logic. Analytics keep the raw
# Accept-Language header regardless of the selection. # Accept-Language header regardless of the selection, and record
# the resolved language as the GET's rendered language.
query_lang = request.query_params.get("lang") query_lang = request.query_params.get("lang")
lang = i18n.select_language( lang = i18n.select_language(
query_lang, query_lang,
@@ -189,8 +176,7 @@ async def show_page(request: Request, path: str) -> Response:
if request.headers.get("if-none-match") == etag: if request.headers.get("if-none-match") == etag:
return Response(status_code=304) return Response(status_code=304)
if _is_trackable_path(path): if _is_trackable_path(path):
flushed = _track_entry(path, request) _record_get(request, lang=lang)
_schedule_client_enrichment(flushed)
return _html_response( return _html_response(
request, request,
"page", "page",
@@ -220,8 +206,7 @@ async def show_page(request: Request, path: str) -> Response:
) )
link_lang = i18n.base_tag(query_lang or "") link_lang = i18n.base_tag(query_lang or "")
if _is_trackable_path(path): if _is_trackable_path(path):
flushed = _track_entry(path, request, status=404) _record_get(request, status=404, lang=lang)
_schedule_client_enrichment(flushed)
return _html_response( return _html_response(
request, request,
"category", "category",
@@ -241,13 +226,5 @@ async def show_page(request: Request, path: str) -> Response:
if item.published: if item.published:
return RedirectResponse(f"/{slug}") return RedirectResponse(f"/{slug}")
if _is_trackable_path(path): if _is_trackable_path(path):
client_hash = analytics_store.track_404( _record_get(request, status=404)
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(request, "not-found", path, 404) return _html_response(request, "not-found", path, 404)
+100 -22
View File
@@ -20,8 +20,13 @@ each segment's source span was located at dispatch (``split``), and
``join`` swaps in the translations. Markup therefore cannot break — it ``join`` swaps in the translations. Markup therefore cannot break — it
never left the server. A returned segment must still be pure prose itself never left the server. A returned segment must still be pure prose itself
(the model could inject markup INTO a segment); anything else — count (the model could inject markup INTO a segment); anything else — count
mismatch, empty segment, markup tokens — rejects the whole result and the mismatch, empty segment, markup tokens, a line that would start a new
fragment stays pending. block (a ``` or ::: fence would eat the rest of the block it lands in) —
rejects the whole result and the
fragment stays pending. Punctuation that is prose on the wire but syntax
in the splice context (quotes in a title attribute, brackets in an alt
text, "|" in a table row) is not worth a rejection either: it is swapped
for Unicode look-alikes (``_NEUTRAL``) before splicing.
A block of plain text, prose links and paired text formatting A block of plain text, prose links and paired text formatting
(strong/em/s) crosses as ONE segment — link texts and formatted text (strong/em/s) crosses as ONE segment — link texts and formatted text
@@ -42,9 +47,11 @@ snippets that don't fit together. Blocks with any other inline markup
Locating is best effort: a run that is not a verbatim source substring Locating is best effort: a run that is not a verbatim source substring
(entity-decoded text, backslash escapes) is skipped — it simply stays in (entity-decoded text, backslash escapes) is skipped — it simply stays in
the original language. So is any piece containing "<": "<" is the the original language. A literal "<" in prose ("<1MB") is text, not
prose/markup boundary on the wire — translators cut their output there, markup, but cannot cross as-is — "<" is the prose/markup boundary on the
so such pieces could not survive the round trip. wire, translators cut their output there — so it crosses encoded as the
fullwidth "" (``_encode``) and ``join`` decodes it back before
validating and splicing.
""" """
import bisect import bisect
@@ -69,6 +76,38 @@ _ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
#: (inline attrs are consumed by the parser; a lone {dates} is not). #: (inline attrs are consumed by the parser; a lone {dates} is not).
_BRACES = re.compile(r"\{[^{}\n]*\}") _BRACES = re.compile(r"\{[^{}\n]*\}")
def _encode(text: str) -> str:
"""Wire form of a segment or context: a literal "<" as fullwidth "".
A "<" in prose is text, not markup ("<1MB" — a tag needs a letter or
/!?), but "<" is the prose/markup boundary on the wire (translators
cut output at the first "<", scripts/translator.py), so it cannot
cross as-is. join decodes it back before the pure_prose check and
splicing — anything tag-like the model may have formed around it is
still rejected there.
"""
return text.replace("<", "")
#: ASCII punctuation that is plain prose to the inline parser (so
#: pure_prose cannot catch it) but Markdown SYNTAX in a splice context:
#: quotes close a quoted image/link title, brackets the [...] of alt and
#: re-inserted link texts, "|" splits a table row, and "\" escapes the
#: character after it (a trailing one eats a title's closing quote).
#: Neutralized to Unicode look-alikes (join), which Markdown treats as
#: plain text everywhere — the quotes are curled the way typographer=True
#: renders them anyway.
_NEUTRAL = str.maketrans(
{
'"': "",
"'": "",
"[": "",
"]": "",
"\\": "",
"|": "",
}
)
#: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]", #: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]",
#: "[]" or a bare "]" (shortcut reference); the destination may nest one #: "[]" or a bare "]" (shortcut reference); the destination may nest one
#: level of parens. Best effort — a mis-scan fails the span-reconstruction #: level of parens. Best effort — a mis-scan fails the span-reconstruction
@@ -273,7 +312,7 @@ def _linked_block(
raw = "".join(text for text, _ in pieces) raw = "".join(text for text, _ in pieces)
lead = len(raw) - len(raw.lstrip()) lead = len(raw) - len(raw.lstrip())
wire = raw.strip() wire = raw.strip()
if not _LETTER.search(wire) or "<" in wire or _BRACES.search(wire): if not _LETTER.search(wire) or _BRACES.search(wire):
return None return None
# Locate each piece verbatim, in order; the source slices between the # Locate each piece verbatim, in order; the source slices between the
# located pieces are then the link syntax, exact by construction. # located pieces are then the link syntax, exact by construction.
@@ -338,7 +377,7 @@ def _linked_block(
rec.append(text_) rec.append(text_)
if source[span_start:span_end] != "".join(rec): if source[span_start:span_end] != "".join(rec):
return None return None
return Span(span_start, span_end, _weight(wire), marks), wire return Span(span_start, span_end, _weight(wire), marks), _encode(wire)
def split(text: str) -> tuple[list[Span], list[str], list[str]]: def split(text: str) -> tuple[list[Span], list[str], list[str]]:
@@ -367,10 +406,8 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
def emit(run: str, at: int, ctx: str) -> None: def emit(run: str, at: int, ctx: str) -> None:
"""Carve {...} spans out of the located run; emit the prose pieces, """Carve {...} spans out of the located run; emit the prose pieces,
stripped — padding whitespace stays in the template, off the wire. stripped — padding whitespace stays in the template, off the wire.
Pieces containing "<" are never emitted: translators cut output at A literal "<" crosses encoded (``_encode``): it is text, not
the first "<" (the prose/markup boundary, scripts/translator.py), markup, but the wire keeps "<" as the prose/markup boundary."""
so such a piece could not survive the round trip — it stays in the
original language instead."""
pieces = [] pieces = []
pos = 0 pos = 0
for m in _BRACES.finditer(run): for m in _BRACES.finditer(run):
@@ -380,10 +417,10 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
for p0, p1 in pieces: for p0, p1 in pieces:
raw = run[p0:p1] raw = run[p0:p1]
piece = raw.strip() piece = raw.strip()
if _LETTER.search(piece) and "<" not in piece: if _LETTER.search(piece):
start = at + p0 + (len(raw) - len(raw.lstrip())) start = at + p0 + (len(raw) - len(raw.lstrip()))
spans.append(Span(start, start + len(piece), 0, [])) spans.append(Span(start, start + len(piece), 0, []))
segments.append(piece) segments.append(_encode(piece))
contexts.append(ctx) contexts.append(ctx)
tokens = _MD.parse(text) tokens = _MD.parse(text)
@@ -409,7 +446,7 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
cursor = span.end cursor = span.end
continue continue
runs = _runs(kids) runs = _runs(kids)
block = _block_text(kids).strip() block = _encode(_block_text(kids).strip())
if alert and runs: if alert and runs:
run = _ALERT.sub("", runs[0], count=1) run = _ALERT.sub("", runs[0], count=1)
if _LETTER.search(run): if _LETTER.search(run):
@@ -417,7 +454,7 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
else: else:
runs.pop(0) runs.pop(0)
for run in runs: for run in runs:
ctx = block if block and run.strip() != block else "" ctx = block if block and _encode(run.strip()) != block else ""
pos = _locate(text, run, cursor) pos = _locate(text, run, cursor)
if pos != -1: if pos != -1:
emit(run, pos, ctx) emit(run, pos, ctx)
@@ -435,6 +472,22 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
return spans, segments, contexts return spans, segments, contexts
#: Block-level Markdown a translation must not introduce: a segment is
#: spliced INSIDE a block of the fragment, so a line starting a heading,
#: quote, list, code/container fence or a setext/thematic-break underline
#: would break the fragment's block structure — a ``` or ::: line eats the
#: rest of the fence it lands in, closing fence included. pure_prose only
#: parses inline and lets such lines through as softbreak prose, so join
#: rejects them here. Blank lines split the host block and are rejected
#: too (a faithful translation of a single block has none).
_BLOCK = re.compile(
r"^[ \t]*(?:#{1,6}(?:[ \t]|$)|>[ \t]?|(?:[-+*]|\d{1,9}[.)])[ \t]|`{3,}|~{3,}|:{3,}(?:[ \t]|$)"
r"|-(?:[ \t]*-){2,}[ \t]*$|=[ =]*$|_(?:[ \t]*_){2,}[ \t]*$)",
re.M,
)
_BLANK = re.compile(r"\n[ \t]*\n")
def pure_prose(text: str) -> bool: def pure_prose(text: str) -> bool:
"""True when the text parses as nothing but prose (text and softbreak """True when the text parses as nothing but prose (text and softbreak
tokens) — the acceptance test for a translated segment: the model may tokens) — the acceptance test for a translated segment: the model may
@@ -483,7 +536,9 @@ _GAP_S = 0.6
_MATCH = 0.3 _MATCH = 0.3
def _find_mark(src: list[str], units: list[re.Match], start: int) -> tuple[int, int] | None: def _find_mark(
src: list[str], units: list[re.Match], start: int
) -> tuple[int, int] | None:
"""Locate a mark's source words in the translation's units (from unit """Locate a mark's source words in the translation's units (from unit
index ``start`` on), as the (start, end) unit-index span of the best index ``start`` on), as the (start, end) unit-index span of the best
fuzzy alignment; None when no alignment is convincing (the caller falls fuzzy alignment; None when no alignment is convincing (the caller falls
@@ -508,7 +563,10 @@ def _find_mark(src: list[str], units: list[re.Match], start: int) -> tuple[int,
back[i][0] = (i - 1, 0) back[i][0] = (i - 1, 0)
for j in range(1, m + 1): for j in range(1, m + 1):
options = [ options = [
(dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH, (i - 1, j - 1)), (
dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH,
(i - 1, j - 1),
),
(dp[i][j - 1] - _GAP_T, (i, j - 1)), (dp[i][j - 1] - _GAP_T, (i, j - 1)),
(dp[i - 1][j] - _GAP_S, (i - 1, j)), (dp[i - 1][j] - _GAP_S, (i - 1, j)),
] ]
@@ -592,17 +650,37 @@ def _place_marks(translation: str, weight: int, marks: list[Mark]) -> str | None
def join(original: str, spans: list[Span], texts: list[str]) -> str | None: def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
"""Splice translated segments back into the original fragment; None on """Splice translated segments back into the original fragment; None on
any validation failure (count mismatch, empty or non-prose segment) — any validation failure (count mismatch, empty, non-prose or
the caller drops the result and the fragment stays pending. Segments block-structure segment) — the caller drops the result and the fragment
with marks (a block that crossed as one piece) get their links stays pending. Segments with marks (a block that crossed as one piece)
re-inserted at weight-mapped positions after the prose check.""" get their links re-inserted at weight-mapped positions after the prose
check.
Markdown-significant ASCII punctuation that pure_prose cannot see
(plain text inline, syntax in the splice context — quoted titles, alt
and link texts, table rows) is neutralized to Unicode look-alikes
(``_NEUTRAL``) before splicing and mark placement (the swap is
char-for-char, so unit alignment is unaffected); lines that would
start a new block (a heading, a ``` or ::: fence — they would eat the
rest of the block/fence they land in) reject the result outright
(``_BLOCK``, ``_BLANK``)."""
if len(texts) != len(spans): if len(texts) != len(spans):
return None return None
out: list[str] = [] out: list[str] = []
cursor = 0 cursor = 0
for span, translation in zip(spans, texts): for span, translation in zip(spans, texts):
if not translation.strip() or not pure_prose(translation): # Decode the wire form ("" back to "<") first: pure_prose then
# validates exactly what gets spliced — a "" the model formed
# into anything tag-like is markup and rejects the result.
translation = translation.replace("", "<")
if (
not translation.strip()
or not pure_prose(translation)
or _BLOCK.search(translation)
or _BLANK.search(translation.strip())
):
return None return None
translation = translation.translate(_NEUTRAL)
if span.marks: if span.marks:
translation = _place_marks(translation, span.weight, span.marks) translation = _place_marks(translation, span.weight, span.marks)
if translation is None: if translation is None:
+17 -16
View File
@@ -3,7 +3,7 @@
Everything the route modules (files, api, tracking, pages) need that is not Everything the route modules (files, api, tracking, pages) need that is not
a route itself: environment-derived paths and tunables, the ``Data`` root a route itself: environment-derived paths and tunables, the ``Data`` root
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
store, the fastapi-vue ``Frontend``, the page render cache store, the page render cache
(``_render_html``/``_cached_body``/``_html_response`` plus the (``_render_html``/``_cached_body``/``_html_response`` plus the
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every ``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
content/settings write), the translator ``dispatcher``, the slug charset content/settings write), the translator ``dispatcher``, the slug charset
@@ -22,13 +22,13 @@ from pathlib import Path
import blake3 import blake3
from fastapi import HTTPException, Request from fastapi import HTTPException, Request
from fastapi.responses import Response from fastapi.responses import Response
from fastapi_vue import Frontend from fastapi_vue import env
from kanta import Kanta from kanta import Kanta
from zstandard import ZstdCompressor from zstandard import ZstdCompressor
from pagerite import analytics, i18n, seed, translate, views from pagerite import analytics, i18n, seed, translate, views
from pagerite.__main__ import DEVMODE
from pagerite.chunks import store_chunks from pagerite.chunks import store_chunks
from pagerite.config import load
from pagerite.data import ( from pagerite.data import (
Data, Data,
Node, Node,
@@ -39,10 +39,13 @@ from pagerite.data import (
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
config = load()
# Site identity: the hostname comes from the CLI (first positional argument, # Site identity: the hostname comes from the CLI (first positional argument,
# exported as PAGERITE_HOSTNAME) and names the per-site data directory # passed in PAGERITE_CONFIG) and names the per-site data directory
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd. # ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
HOSTNAME = os.getenv("PAGERITE_HOSTNAME", "localhost") HOSTNAME = config.hostname
SITE_DIR = Path(HOSTNAME) SITE_DIR = Path(HOSTNAME)
#: Public origin of the site, used for absolute social/canonical/sitemap #: Public origin of the site, used for absolute social/canonical/sitemap
#: URLs. Localhost serves varying ports, so it falls back to the request's #: URLs. Localhost serves varying ports, so it falls back to the request's
@@ -53,6 +56,10 @@ DB_PATH = os.getenv("PAGERITE_DB", str(SITE_DIR / "content.kantadb"))
# Visit analytics go to their own JSON file, not the kanta database. # Visit analytics go to their own JSON file, not the kanta database.
ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json"))) ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json")))
# The per-hostname data directory may not exist yet on first run; kanta
# creates the database file but not its parent directory.
Path(DB_PATH).parent.mkdir(parents=True, exist_ok=True)
ANALYTICS_PATH.parent.mkdir(parents=True, exist_ok=True)
analytics_store = analytics.Store(ANALYTICS_PATH) analytics_store = analytics.Store(ANALYTICS_PATH)
# Content-addressed file store (uploads, seed assets, fetched favicons): # Content-addressed file store (uploads, seed assets, fetched favicons):
@@ -76,12 +83,6 @@ FAVICON_MAXSIZE = 192
data = Data() data = Data()
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations") kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
# Vue build served at the site root, no SPA catch-all (assets only). The
# build mirrors the URL space: hashed, immutable files live under
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico.
BUILD_DIR = Path(__file__).with_name("frontend-build")
frontend = Frontend(BUILD_DIR, spa=False, cached="/_assets/")
# Dynamic HTML is compressed per request at level 9 (static assets are # Dynamic HTML is compressed per request at level 9 (static assets are
# already pre-compressed by fastapi-vue's Frontend). # already pre-compressed by fastapi-vue's Frontend).
_zstd = ZstdCompressor(9) _zstd = ZstdCompressor(9)
@@ -135,6 +136,7 @@ def _render_html(
data.theme, data.theme,
data.favicon, data.favicon,
data.brand_html, data.brand_html,
base_url,
transition=data.transition, transition=data.transition,
lang=lang, lang=lang,
translation=translation, translation=translation,
@@ -228,7 +230,7 @@ def _html_response(
# Absolute social/canonical URLs use the site's public origin; on # Absolute social/canonical URLs use the site's public origin; on
# localhost (varying ports) fall back to the request's own base URL. # localhost (varying ports) fall back to the request's own base URL.
base_url = SITE_URL or str(request.base_url).rstrip("/") base_url = SITE_URL or str(request.base_url).rstrip("/")
if DEVMODE: if env.dev:
identity = _render_html(kind, path, base_url, lang, link_lang).encode() identity = _render_html(kind, path, base_url, lang, link_lang).encode()
body = _zstd.compress(identity) if zstd else identity body = _zstd.compress(identity) if zstd else identity
else: else:
@@ -345,6 +347,7 @@ def _seed(data: Data) -> None:
#: Translator key format: 12 lowercase alphanumeric characters — not #: Translator key format: 12 lowercase alphanumeric characters — not
#: brute-forceable over a WebSocket handshake, still human-manageable. #: brute-forceable over a WebSocket handshake, still human-manageable.
#: The editor's lang tab generates further keys in the same format.
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789" _KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
@@ -352,10 +355,8 @@ _KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
def _translator_defaults(data: Data) -> None: def _translator_defaults(data: Data) -> None:
"""Translator defaults on database creation: the first service key and """Translator defaults on database creation: the first service key and
the wanted target languages (Spanish and Chinese — English is the the wanted target languages (Spanish and Chinese — English is the
original language, never a translation target). original language, never a translation target). Further keys are
managed in the editor shell's lang tab."""
Keys are a dict (key -> display name) with the future reservation that
multiple keys could be managed (e.g. via a web interface)."""
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12)) key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
data.translate_keys[key] = "default" data.translate_keys[key] = "default"
data.translate_langs = {"es": True, "zh": True} data.translate_langs = {"es": True, "zh": True}
+5 -4
View File
@@ -111,12 +111,13 @@ article h3 {
} }
blockquote { blockquote {
border-left-color: var(--accent); border-inline-start-color: var(--accent);
background: color-mix(var(--accent) 6%, transparent); background: color-mix(var(--accent) 6%, transparent);
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
margin: 0 -0.9rem 1rem calc(-0.25rem - 0.9rem); margin: 0 0 1rem;
margin-inline: calc(-0.25rem - 0.9rem) -0.9rem;
border-radius: 6px; border-radius: 6px;
} }
@@ -125,9 +126,9 @@ blockquote {
bar stays in both. */ bar stays in both. */
pre { pre {
border: 1px solid transparent; border: 1px solid transparent;
border-left: 0.25rem solid var(--accent); border-inline-start: 0.25rem solid var(--accent);
/* Text on the paragraph edge: the box extends by padding + border. */ /* Text on the paragraph edge: the box extends by padding + border. */
margin-left: calc(-0.8rem - 0.25rem); margin-inline-start: calc(-0.8rem - 0.25rem);
border-radius: 6px; border-radius: 6px;
} }
+8 -7
View File
@@ -144,9 +144,9 @@ article h1 {
font-weight: 700; font-weight: 700;
padding-bottom: 0.5rem; padding-bottom: 0.5rem;
/* The hazard-stripe underline breaks out of the page box: the negative /* The hazard-stripe underline breaks out of the page box: the negative
right margin extends the h1's box (and thus its background) all the end margin extends the h1's box (and thus its background) all the
way to the viewport's right edge. */ way to the viewport's edge on that side. */
margin-right: calc((100% - 100vw) / 2); margin-inline-end: calc((100% - 100vw) / 2);
background: background:
linear-gradient(-55deg, linear-gradient(-55deg,
transparent 0 0.2rem, transparent 0 0.2rem,
@@ -186,20 +186,21 @@ article ul ul ul li::before {
} }
blockquote { blockquote {
border-left-color: var(--accent2); border-inline-start-color: var(--accent2);
background: color-mix(var(--accent2) 6%, transparent); background: color-mix(var(--accent2) 6%, transparent);
padding: 0.25rem 0.75rem; padding: 0.25rem 0.75rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
margin: 0 -0.75rem 1rem -1rem; margin: 0 0 1rem;
margin-inline: -1rem -0.75rem;
} }
/* Code follows the color scheme; the dark-scheme well joins the violet /* Code follows the color scheme; the dark-scheme well joins the violet
family (--code-bg above). The orange side bar stays in both. */ family (--code-bg above). The orange side bar stays in both. */
pre { pre {
border-left: 0.25rem solid var(--accent); border-inline-start: 0.25rem solid var(--accent);
/* Text on the paragraph edge: the box extends by padding + border. */ /* Text on the paragraph edge: the box extends by padding + border. */
margin-left: calc(-0.8rem - 0.25rem); margin-inline-start: calc(-0.8rem - 0.25rem);
border-radius: 3px; border-radius: 3px;
} }
+2 -2
View File
@@ -66,7 +66,7 @@ article ul li::before {
content: "◆"; content: "◆";
color: var(--accent); color: var(--accent);
font-size: 0.8em; font-size: 0.8em;
margin-left: calc(-1 * var(--list-indent) / 0.8); margin-inline-start: calc(-1 * var(--list-indent) / 0.8);
width: calc(var(--list-indent) / 0.8); width: calc(var(--list-indent) / 0.8);
} }
@@ -80,7 +80,7 @@ article ul ul ul li::before {
} }
blockquote { blockquote {
border-left-color: var(--accent2); border-inline-start-color: var(--accent2);
} }
/* Code panels sit slightly lighter than the page; the token colors come /* Code panels sit slightly lighter than the page; the token colors come
+2 -2
View File
@@ -160,7 +160,7 @@ main::before {
top edge and a grassy shadow. */ top edge and a grassy shadow. */
#sidebar { #sidebar {
background: linear-gradient(160deg, #f4faddd9, #d9eec5cf); background: linear-gradient(160deg, #f4faddd9, #d9eec5cf);
border-right: 1px solid #ffffff80; border-inline-end: 1px solid #ffffff80;
border-bottom: 1px solid var(--line); border-bottom: 1px solid var(--line);
box-shadow: 0 0.3rem 1rem #4f913b1f; box-shadow: 0 0.3rem 1rem #4f913b1f;
border-radius: 1rem; border-radius: 1rem;
@@ -225,7 +225,7 @@ article ul ul ul li::before {
/* Quotes get a grassy edge and a wash of sunlight. */ /* Quotes get a grassy edge and a wash of sunlight. */
blockquote { blockquote {
color: #4d6849; color: #4d6849;
border-left-color: var(--accent2); border-inline-start-color: var(--accent2);
background: linear-gradient(90deg, #fff0a238, transparent 70%); background: linear-gradient(90deg, #fff0a238, transparent 70%);
padding-top: 0.25rem; padding-top: 0.25rem;
padding-bottom: 0.25rem; padding-bottom: 0.25rem;
+165 -67
View File
@@ -3,19 +3,21 @@
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
enriched in background tasks with reverse DNS (cached PTR lookups) and the enriched in background tasks with reverse DNS (cached PTR lookups) and the
DB-IP city MMDB (``GeoIP``, decompressed and opened once at startup); DB-IP city MMDB (``GeoIP``, decompressed into RAM and opened once at
startup);
external referrers get their favicon fetched and stored content-hashed. external referrers get their favicon fetched and stored content-hashed.
Snapshot broadcasts to connected admin sockets are debounced. Snapshot broadcasts to connected admin sockets are debounced.
""" """
import asyncio import asyncio
import gzip import gzip
import io
import ipaddress import ipaddress
import logging import logging
import os import os
import re import re
import shutil
import socket import socket
from datetime import date
from functools import lru_cache from functools import lru_cache
from pathlib import Path from pathlib import Path
from urllib.parse import urlparse from urllib.parse import urlparse
@@ -24,13 +26,19 @@ import httpx
import msgspec import msgspec
from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
from fastapi.responses import Response from fastapi.responses import Response
from uarite import uaparse
from pagerite import analytics from pagerite import analytics, i18n
from pagerite.data import resolve
from pagerite.files import _hash_name, file_store from pagerite.files import _hash_name, file_store
from pagerite.state import SITE_URL, _html_response, analytics_store from pagerite.state import SITE_URL, _html_response, analytics_store, data
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# httpx logs every request at INFO (e.g. the favicon fetches below); our own
# one-line summary in _schedule_favicon_fetch replaces that noise.
logging.getLogger("httpx").setLevel(logging.WARNING)
router = APIRouter() router = APIRouter()
# Live WebSocket clients for the analytics stream. # Live WebSocket clients for the analytics stream.
@@ -38,20 +46,78 @@ _analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None _analytics_broadcast_task: asyncio.Task | None = None
# Repository root from this file's location (pagerite/tracking.py -> ..). # DB-IP databases persist in the working directory (one download serves all
_REPO_ROOT = Path(__file__).resolve().parent.parent # sites run from it). Not the package directory: reinstalls/upgrades wipe it.
_DBIP_DIR = Path.cwd()
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
def _download_dbip() -> None:
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
today = date.today()
months = [f"{today:%Y-%m}"]
# The current month's file may not be published yet; fall back to last month.
prev = (today.replace(day=1) - date.resolution).replace(day=1)
months.append(f"{prev:%Y-%m}")
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
logger.info("DB-IP database is current (%s), skipping download", existing[-1])
return
for month in months:
url = DBIP_URL.format(month=month)
target = _DBIP_DIR / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
logger.info("Downloading %s", url)
try:
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
if r.status_code == 404:
continue
r.raise_for_status()
with open(tmp, "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk)
except httpx.HTTPError as e:
logger.warning("DB-IP download failed: %s", e)
tmp.unlink(missing_ok=True)
continue
# Verify it is actually gzip data before installing it.
try:
with gzip.open(tmp, "rb") as f:
f.read(1)
except OSError:
logger.warning("DB-IP download for %s was not valid gzip", month)
tmp.unlink(missing_ok=True)
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
logger.info("DB-IP database updated to %s", target.name)
return
logger.warning("Could not download a DB-IP database")
def _geoip_db_path() -> Path | None: def _geoip_db_path() -> Path | None:
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed """Find a DB-IP MMDB in the working directory: the ``.mmdb.gz`` download
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present. is canonical (decompressed into RAM at open); a plain ``.mmdb`` left over
from older versions is still usable, and removed once the matching ``.gz``
is present so it does not linger on disk. Returns None if none is present.
""" """
mmdb = sorted(_REPO_ROOT.glob("dbip-*.mmdb")) gz = sorted(_DBIP_DIR.glob("dbip-*.mmdb.gz"))
if gz:
for stale in _DBIP_DIR.glob("dbip-*.mmdb"):
stale.unlink()
return gz[0]
mmdb = sorted(_DBIP_DIR.glob("dbip-*.mmdb"))
if mmdb: if mmdb:
return mmdb[0] return mmdb[0]
gz = sorted(_REPO_ROOT.glob("dbip-*.mmdb.gz"))
if gz:
return gz[0]
return None return None
@@ -64,28 +130,23 @@ class GeoIP:
def __init__(self) -> None: def __init__(self) -> None:
self._reader: object | None = None self._reader: object | None = None
def _decompress(self, source: Path, target: Path) -> None:
if target.exists():
return
tmp = target.with_suffix(target.suffix + ".tmp")
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
shutil.copyfileobj(src, dst)
os.replace(tmp, target)
def _load(self) -> None: def _load(self) -> None:
if self._reader is not None: if self._reader is not None:
return return
source = _geoip_db_path() source = _geoip_db_path()
if source is None: if source is None:
return return
if source.suffix == ".gz":
target = source.with_suffix("")
self._decompress(source, target)
source = target
try: try:
import maxminddb import maxminddb
self._reader = maxminddb.open_database(str(source)) if source.suffix == ".gz":
# Only the .gz is kept on disk; the database is decompressed
# into RAM (MODE_FD makes the pure-Python Reader .read() the
# buffer — never mmap — and bypasses the C extension).
buf = io.BytesIO(gzip.decompress(source.read_bytes()))
self._reader = maxminddb.open_database(buf, maxminddb.MODE_FD)
else:
self._reader = maxminddb.open_database(str(source))
except Exception: except Exception:
pass pass
@@ -253,9 +314,18 @@ async def _fetch_favicon(origin: str) -> None:
def _schedule_favicon_fetch() -> None: def _schedule_favicon_fetch() -> None:
"""Start background favicon fetches for origins that need one.""" """Start background favicon fetches for origins that need one."""
for origin in analytics_store.favicon_origins_needed(): origins = [
if origin in _favicon_in_flight: origin
continue for origin in analytics_store.favicon_origins_needed()
if origin not in _favicon_in_flight
]
if not origins:
return
logger.info(
"Fetching favicons: %s",
", ".join(o.removeprefix("https://") for o in origins),
)
for origin in origins:
_favicon_in_flight.add(origin) _favicon_in_flight.add(origin)
asyncio.create_task(_fetch_favicon(origin)) asyncio.create_task(_fetch_favicon(origin))
@@ -264,7 +334,7 @@ async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client.""" """Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients: if not _analytics_ws_clients:
return return
payload = analytics_store.display_json() payload = _display_json()
closed = set() closed = set()
for ws in _analytics_ws_clients: for ws in _analytics_ws_clients:
try: try:
@@ -291,47 +361,69 @@ def _schedule_analytics_broadcast() -> None:
) )
def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]: def _in_menu(path: str) -> bool:
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET. """True when ``path`` ("/a/b" or "/") resolves to a real menu node.
Nothing is counted on the GET itself — the client's first /_ws message Category placeholders return 404 but are real nodes: their GETs must not
starts the visit, so bots never register as visits (JS-running crawlers count as misses in the display-time abuse classification.
connect too, but the WebSocket handler ignores known bot UAs). (Admin """
clients report too, but with hide, which flags their visit hidden: it is return resolve(data.menu, path.strip("/")) is not None
recorded but excluded from all statistics and from the crawler list.)
def _display_json() -> str:
"""The current analytics snapshot as JSON for the admin stream.
Adds the site's language context: ``multilingual`` (translation
languages configured) lets the viewer suppress language UI on
single-language sites, ``primary_lang`` (the front page's) lets it skip
the primary-language default case.
"""
return analytics_store.display_json(
_in_menu,
multilingual=bool(data.translate_langs),
primary_lang=i18n.primary_lang(data.menu, ""),
)
def _record_get(request: Request, *, status: int = 200, lang: str = "") -> None:
"""Record the document GET as one raw access-log line in analytics.
Nothing is classified here — the true HTTP status, the full request path
(query included), an external referer origin, the preload flag and the
rendered content language (``lang``, "" for non-localized responses such
as 404 probes and reserved paths) are stored, and
visitor/crawler/abuse classification happens at display time
(see analytics.Store.display). Idle-time preloads from pagerite.js
(``x-pagerite-preload`` header) are recorded with ``pre=True``: never
counted, but a navigation later served from the in-memory page cache is
attributed this GET's status.
The devserver's health probe (``GET /?from=devserver.py`` from The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be ``127.0.0.1``) is ignored: it is not real traffic. The root-path and
logged as a crawler hit. The root-path and localhost checks prevent localhost checks prevent remote visitors from forging the same query.
remote visitors from hiding traffic with the same query string.
Returns the client hashes of any pending crawler hits flushed to persistent
storage, so callers can schedule async geoip and reverse-DNS enrichment.
""" """
if request.headers.get("x-pagerite-preload"):
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
# activity message sent when the user actually navigates does the
# counting.
# (Forging the header only hides a GET from the crawler stats; the
# path-based abuse classification is unaffected.)
return []
if ( if (
path == "" request.url.path == "/"
and str(request.url.query) == "from=devserver.py" and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1" and _client_ip(request) == "127.0.0.1"
): ):
return [] return
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}" own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
full_path = f"{request.url.path}{_query_suffix(request)}" referer = request.headers.get("referer", "")
return analytics_store.track_entry( if analytics._origin(referer) in (None, own_origin):
request.headers.get("referer", ""), referer = ""
own_origin, client_hash = analytics_store.record_get(
_client_ip(request), _client_ip(request),
request.headers.get("user-agent", ""), request.headers.get("user-agent", ""),
full_path, f"{request.url.path}{_query_suffix(request)}",
request.headers.get("accept-language", ""),
status=status, status=status,
referer=referer,
accept_language=request.headers.get("accept-language", ""),
pre=bool(request.headers.get("x-pagerite-preload")),
lang=lang,
) )
if client_hash is not None:
_schedule_client_enrichment([client_hash])
@router.get("/_a", response_model=None) @router.get("/_a", response_model=None)
@@ -358,14 +450,21 @@ async def activity_ws(ws: WebSocket) -> None:
Public, like the pages themselves (only /_api is gated); one connection Public, like the pages themselves (only /_api is gated); one connection
follows a browsing session. Messages are ``analytics.Ping`` structs as follows a browsing session. Messages are ``analytics.Ping`` structs as
JSON text frames; ``to`` set is a navigation, ``read`` alone a JSON text frames; ``to`` set is a navigation, ``read`` alone a
reading-time update. The reverse-DNS and DB-IP geoip lookups happen in reading-time update. Everything is recorded raw — known bot UAs and
background tasks so message handling is never delayed by slow DNS or abusive IPs are filtered at display time, not here. The reverse-DNS and
the first MMDB decompress. DB-IP geoip lookups happen in background tasks so message handling is
never delayed by slow DNS or the first MMDB decompress.
""" """
await ws.accept()
ip = _client_ip(ws) ip = _client_ip(ws)
ua = ws.headers.get("user-agent", "") ua = ws.headers.get("user-agent", "")
accept_language = ws.headers.get("accept-language", "") accept_language = ws.headers.get("accept-language", "")
# Identify the visitor on the access-log open/close lines (the IP is
# already printed there): compact UA plus the browser's language tag.
lang, _country = analytics._parse_accept_language(accept_language)
ws.scope.setdefault("state", {})["log_extra"] = " ".join(
part for part in (uaparse(ua).pretty, lang) if part
)
await ws.accept()
try: try:
while True: while True:
text = await ws.receive_text() text = await ws.receive_text()
@@ -373,7 +472,7 @@ async def activity_ws(ws: WebSocket) -> None:
msg = msgspec.json.decode(text.encode(), type=analytics.Ping) msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
except msgspec.DecodeError: except msgspec.DecodeError:
continue continue
visit_index, flushed_clients = analytics_store.ping( new_client = analytics_store.record_msg(
msg.fr, msg.fr,
msg.to or None, msg.to or None,
ip, ip,
@@ -381,11 +480,10 @@ async def activity_ws(ws: WebSocket) -> None:
accept_language, accept_language,
hide=msg.hide, hide=msg.hide,
read=msg.read, read=msg.read,
lang=msg.lang,
) )
if visit_index is not None: if new_client is not None:
visit = analytics_store.data.visits[visit_index] _schedule_client_enrichment([new_client])
asyncio.create_task(_enrich_client(visit.client))
_schedule_client_enrichment(flushed_clients)
_schedule_favicon_fetch() _schedule_favicon_fetch()
except WebSocketDisconnect: except WebSocketDisconnect:
pass pass
@@ -399,7 +497,7 @@ async def analytics_websocket(ws: WebSocket) -> None:
endpoint. Powers the analytics viewer rendered at /_a. endpoint. Powers the analytics viewer rendered at /_a.
""" """
await ws.accept() await ws.accept()
await ws.send_text(analytics_store.display_json()) await ws.send_text(_display_json())
_analytics_ws_clients.add(ws) _analytics_ws_clients.add(ws)
try: try:
while True: while True:
+42 -33
View File
@@ -101,7 +101,8 @@ ClientMsg = Hello | Result
def pending_items(data: Data, lang: str) -> list[TransItem]: def pending_items(data: Data, lang: str) -> list[TransItem]:
"""Fragments of the site still untranslated for ``lang``, deduped by key. """Fragments of the site still untranslated for ``lang``, deduped by key.
Every page node (published or not) contributes its title and each chunk Every node (published or not, pages and pure category labels alike)
contributes its title; pages also contribute each chunk
that needs translation (``needs_translation``), is not editor-flagged that needs translation (``needs_translation``), is not editor-flagged
no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang`` no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang``
yet. Content-addressed text (shared paragraphs, repeated titles) appears yet. Content-addressed text (shared paragraphs, repeated titles) appears
@@ -133,8 +134,10 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
path = f"{prefix}/{slug}" if prefix else slug path = f"{prefix}/{slug}" if prefix else slug
# An article whose primary language IS the target needs no # An article whose primary language IS the target needs no
# translation into it — skip its title and chunks entirely. # translation into it — skip its title and chunks entirely.
# Category labels (chunks is None) contribute only their title:
# it is their nav-menu label.
node_lang = node.language or inherited node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang: if node_lang != lang:
if node.title: if node.title:
emit( emit(
chunk_key(node.title), chunk_key(node.title),
@@ -143,7 +146,7 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
"title", "title",
context=opening(node), context=opening(node),
) )
for h in node.chunks: for h in node.chunks or ():
text = data.chunks.get(h) text = data.chunks.get(h)
if ( if (
text is not None text is not None
@@ -178,8 +181,8 @@ def store_results(data: Data, lang: str, items: list[TransResult]) -> list[str]:
for slug, node in sorted_nodes(nodes): for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited node_lang = node.language or inherited
if node.chunks is not None and node_lang != lang: if node_lang != lang:
keys = set(node.chunks) keys = set(node.chunks or ())
if node.title: if node.title:
keys.add(chunk_key(node.title)) keys.add(chunk_key(node.title))
if keys & stored: if keys & stored:
@@ -277,34 +280,40 @@ class Dispatcher:
job = None job = None
spans: list[Span] = [] spans: list[Span] = []
original = "" original = ""
for lang in sorted(langs): # Titles before articles — across languages too, so every menu
# Titles first: a page's name in the menu is its most # is named before any article body is worked on (a page's name
# visible string (stable: menu order kept within each kind). # is its most visible string). pending_items emits in menu
for item in sorted( # order, a page's title before its chunks; filtering by kind
pending_items(self.data, lang), key=lambda it: it.kind != "title" # keeps that stable order within each kind.
): pending = {lang: pending_items(self.data, lang) for lang in sorted(langs)}
if (lang, item.key) in inflight or ( for kind in ("title", "chunk"):
lang, for lang in sorted(langs):
item.key, for item in pending[lang]:
) in self.validation_failures: if (
continue item.kind != kind
spans, texts, contexts = split(item.text) or (lang, item.key) in inflight
if not texts: or (lang, item.key) in self.validation_failures
continue # prose that could not be located for splicing ):
original = item.text continue
if item.kind == "title" and item.context: spans, texts, contexts = split(item.text)
# A title's surround is the article's opening prose if not texts:
# (TransItem.context), not its own one-word block. continue # prose that could not be located for splicing
contexts = [item.context] * len(texts) original = item.text
job = Job( if item.kind == "title" and item.context:
lang=lang, # A title's surround is the article's opening prose
key=item.key, # (TransItem.context), not its own one-word block.
texts=texts, contexts = [item.context] * len(texts)
path=item.path, job = Job(
kind=item.kind, lang=lang,
contexts=contexts, key=item.key,
) texts=texts,
break path=item.path,
kind=item.kind,
contexts=contexts,
)
break
if job is not None:
break
if job is not None: if job is not None:
break break
if job is None: if job is None:
+117 -39
View File
@@ -21,10 +21,12 @@ import json
import os import os
import re import re
from fastapi_vue import env
from html5tagger import HTML, Document, E, Template from html5tagger import HTML, Document, E, Template
from platformdirs import site_data_dir, user_data_path from platformdirs import site_data_dir, user_data_path
from pagerite import i18n from pagerite import i18n
from pagerite.config import load as _load_config
from pagerite.data import Data, Node, node_markdown, prettify, resolve, sorted_nodes from pagerite.data import Data, Node, node_markdown, prettify, resolve, sorted_nodes
from pagerite.i18n import Translation from pagerite.i18n import Translation
from pagerite.markdown import render from pagerite.markdown import render
@@ -45,6 +47,10 @@ def _data_roots() -> list[Path]:
return [Path(r) for r in roots] return [Path(r) for r in roots]
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
config = _load_config()
def _theme_dirs() -> list[Path]: def _theme_dirs() -> list[Path]:
"""Theme search roots, most specific first; first match wins per file. """Theme search roots, most specific first; first match wins per file.
@@ -57,7 +63,7 @@ def _theme_dirs() -> list[Path]:
""" """
return [ return [
Path("themes"), Path("themes"),
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "themes", Path(config.hostname) / "themes",
*(root / "themes" for root in _data_roots()), *(root / "themes" for root in _data_roots()),
Path(__file__).parent / "themes", Path(__file__).parent / "themes",
] ]
@@ -73,7 +79,7 @@ THEME_DIRS = _theme_dirs()
# built-in --font-* variables. # built-in --font-* variables.
FONT_DIRS = [ FONT_DIRS = [
Path("fonts"), Path("fonts"),
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "fonts", Path(config.hostname) / "fonts",
*(root / "fonts" for root in _data_roots()), *(root / "fonts" for root in _data_roots()),
] ]
@@ -258,20 +264,31 @@ def _transition_css_url(transition: str) -> str | None:
def _editor_css_url(vite_url: str | None) -> str | None: def _editor_css_url(vite_url: str | None) -> str | None:
"""URL for the editor-specific stylesheet (Vue component styles). """URLs (comma-joined) for the editor-specific stylesheets (Vue
component styles).
This is linked by the public-page edit pen so the editor styles are This is linked by the public-page edit pen so the editor styles are
loaded before the editor JS dynamic-import resolves. loaded before the editor JS dynamic-import resolves. Component styles
can land on shared chunks rather than the entry's own stylesheet —
LangSelect's ride on the shared store chunk, as it is also used by the
on-demand public language selector — so collect the stylesheets of the
entry and its imported chunks (the same traversal _langselect_assets
does).
""" """
if vite_url: if vite_url:
return None return None
manifest = _manifest() manifest = _manifest()
entry = manifest["src/main.js"]
base = manifest.get(_BASE_CSS_KEY, {}).get("file") base = manifest.get(_BASE_CSS_KEY, {}).get("file")
for css in entry.get("css", []): stylesheets, seen = [], set()
if css != base: queue = ["src/main.js"]
return f"/{css}" for key in queue: # grows with imported chunks
return None if key in seen:
continue
seen.add(key)
entry = manifest[key]
stylesheets += [f"/{css}" for css in entry.get("css", []) if css != base]
queue += entry.get("imports", [])
return ",".join(stylesheets) or None
def _inline_asset(url: str) -> str: def _inline_asset(url: str) -> str:
@@ -321,7 +338,7 @@ def _layout(
) -> Template: ) -> Template:
"""Page layout template with standard assets and ES-module scripts. """Page layout template with standard assets and ES-module scripts.
In dev (PAGERITE_VITE_URL set) assets are linked from the Vite dev In dev (Vite dev-server URL set) assets are linked from the Vite dev
server and stylesheets use ``blocking="render"`` so the browser waits server and stylesheets use ``blocking="render"`` so the browser waits
for them before showing the page, avoiding a flash of unstyled content. for them before showing the page, avoiding a flash of unstyled content.
In production all page assets are inlined into the document: stylesheets In production all page assets are inlined into the document: stylesheets
@@ -344,8 +361,9 @@ def _layout(
property attributes, everything else (description, twitter:*) as name. property attributes, everything else (description, twitter:*) as name.
``lang`` is the served language for <html lang>; an RTL language (ar, ``lang`` is the served language for <html lang>; an RTL language (ar,
fa, ...) also puts dir="rtl" on <html> (the editor panel carries its own fa, ...) also puts dir="rtl" on <html> (the editor panel and the
lang="en" dir="ltr", so it is unaffected). ``canonical`` and analytics dashboard carry their own lang="en" dir="ltr", so they are
unaffected). ``canonical`` and
``alternates`` ((hreflang, href) pairs) are the page's language URLs ``alternates`` ((hreflang, href) pairs) are the page's language URLs
(see docs/localization.md), emitted right after the viewport and before (see docs/localization.md), emitted right after the viewport and before
the social tags: canonical first, then the hreflang alternates. the social tags: canonical first, then the hreflang alternates.
@@ -366,8 +384,8 @@ def _layout(
doc.meta(property=key, content=value) doc.meta(property=key, content=value)
else: else:
doc.meta(name=key, content=value) doc.meta(name=key, content=value)
# A custom favicon (from the site editor) is linked explicitly; without # A custom favicon (from the site editor) is linked explicitly;
# one, browsers fall back to the build's /favicon.ico by convention. # /favicon.ico redirects to the same store file for non-HTML contexts.
if favicon: if favicon:
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon") doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
# Asset URLs for the on-demand bundles (editor, analytics) for # Asset URLs for the on-demand bundles (editor, analytics) for
@@ -377,12 +395,16 @@ def _layout(
# dev-server URLs as meta tags (Vite serves the modules and injects # dev-server URLs as meta tags (Vite serves the modules and injects
# their CSS for hot reloads); production inlines all page assets and # their CSS for hot reloads); production inlines all page assets and
# carries the on-demand URLs in one JSON script instead. # carries the on-demand URLs in one JSON script instead.
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = env.vite_url
editor_scripts, editor_css = _editor_assets() editor_scripts, editor_css = _editor_assets()
langselect_scripts, langselect_css = _langselect_assets()
config = { config = {
"pagerite:editor-src": editor_scripts[-1], "pagerite:editor-src": editor_scripts[-1],
"pagerite:analytics-src": _analytics_assets()[0][0], "pagerite:analytics-src": _analytics_assets()[0][0],
"pagerite:langselect-src": langselect_scripts[-1],
} }
if langselect_css:
config["pagerite:langselect-css"] = ",".join(langselect_css)
if editor_css: if editor_css:
config["pagerite:editor-css"] = editor_css config["pagerite:editor-css"] = editor_css
if vite_url: if vite_url:
@@ -777,7 +799,11 @@ def page_content(
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
content = node_markdown(data, node) or "" content = node_markdown(data, node) or ""
title = node.title title = node.title
# The original text pins the section anchors: on a translated page the
# heading slugs (and thus #hash URLs) stay in the original language.
anchors_from = None
if translation: if translation:
anchors_from = (content, title)
if translation.markdown is not None: if translation.markdown is not None:
content = translation.markdown content = translation.markdown
title = ( title = (
@@ -787,7 +813,9 @@ def page_content(
) )
# The title is injected into the markdown (as # title when it has no # The title is injected into the markdown (as # title when it has no
# h1 of its own), so title and content render as one article. # h1 of its own), so title and content render as one article.
rendered = render(content, path, node.created, node.modified, title=title) rendered = render(
content, path, node.created, node.modified, title=title, anchors_from=anchors_from
)
# Long articles get .multicol: the article column cap lifts (see the # Long articles get .multicol: the article column cap lifts (see the
# #content grid in pagerite.css) and the .cols segments lay out in at # #content grid in pagerite.css) and the .cols segments lay out in at
# most two columns. The html is already segmented by render() — the # most two columns. The html is already segmented by render() — the
@@ -1005,6 +1033,42 @@ def _social_meta(
} }
def _language_urls(
data: Data,
path: str,
node: Node,
lang: str,
original: str,
base_url: str,
) -> tuple[str, list[tuple[str, str]]]:
"""(canonical, hreflang alternates) for a page (docs/localization.md).
The canonical names the actually served language — the plain URL for
the original (for SEO the non-query URL means the article's language),
?lang= for a translation — regardless of how the language was arrived
at (query or header). The alternates list the languages the page is
actually available in (``node.langs``; a category label's title counts
as its content): x-default first (the plain, autodetecting URL), then
every available language — the original again by its plain URL,
translations by ?lang=. The public language selector keys off these.
("", []) without a base_url.
"""
if not base_url:
return "", []
url = f"{base_url}/{path}"
canonical = url if lang == original else f"{url}?lang={lang}"
alternates = []
if data.translate_langs:
# Only languages the page actually has AND that are still enabled
# site-wide (a disabled target stops being advertised).
enabled = {original, *data.translate_langs}
alternates = [("x-default", url)] + [
(tag, url if tag == original else f"{url}?lang={tag}")
for tag in sorted({original, *node.langs} & enabled)
]
return canonical, alternates
def render_page( def render_page(
menu: dict[str, Node], menu: dict[str, Node],
data: Data, data: Data,
@@ -1034,24 +1098,7 @@ def render_page(
title = _title(path.rpartition("/")[2], node, translation, path) title = _title(path.rpartition("/")[2], node, translation, path)
main = page_content(menu, data, path, translation, link_lang, lang) main = page_content(menu, data, path, translation, link_lang, lang)
social = _social_meta(node, path, title, str(main), brand, base_url) social = _social_meta(node, path, title, str(main), brand, base_url)
# Canonical/hreflang URLs (docs/localization.md): the canonical names canonical, alternates = _language_urls(data, path, node, lang, original, base_url)
# the actually served language — the plain URL for the original (for
# SEO the non-query URL means the article's language), ?lang= for a
# translation — regardless of how the language was arrived at (query
# or header). The alternates are site-wide, the same set on every
# page: the configured translate_langs (the translator works to fill
# them all in), x-default first (the plain, autodetecting URL), then
# every language explicitly, the page's own primary included.
canonical = ""
alternates = []
if base_url:
url = f"{base_url}/{path}"
canonical = url if lang == original else f"{url}?lang={lang}"
if data.translate_langs:
alternates = [("x-default", url)] + [
(tag, f"{url}?lang={tag}")
for tag in sorted({original, *data.translate_langs})
]
return str( return str(
_layout( _layout(
*_page_assets(), *_page_assets(),
@@ -1084,6 +1131,7 @@ def render_category(
theme: str = "", theme: str = "",
favicon: str = "", favicon: str = "",
brand_html: str = "", brand_html: str = "",
base_url: str = "",
transition: str = "cube", transition: str = "cube",
lang: str = i18n.ORIGINAL_LANGUAGE, lang: str = i18n.ORIGINAL_LANGUAGE,
translation: Translation | None = None, translation: Translation | None = None,
@@ -1099,12 +1147,16 @@ def render_category(
With a translation (titles only — the category has no Markdown) the With a translation (titles only — the category has no Markdown) the
heading, navigation and card text localize per target article heading, navigation and card text localize per target article
(docs/localization.md); ``link_lang`` replicates the ?lang= override (docs/localization.md); ``link_lang`` replicates the ?lang= override
onto the navigation links as on content pages. onto the navigation links as on content pages. The hreflang alternates
are computed as on content pages — a translated title makes the
language available here too.
""" """
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
original = i18n.primary_lang(menu, path)
if translation is None: if translation is None:
lang = i18n.primary_lang(menu, path) lang = original
title = _title(path.rpartition("/")[2], node, translation, path) title = _title(path.rpartition("/")[2], node, translation, path)
_, alternates = _language_urls(data, path, node, lang, original, base_url)
doc = E.article doc = E.article
with doc: with doc:
doc.h1(title) doc.h1(title)
@@ -1121,6 +1173,7 @@ def render_category(
transition, transition,
favicon, favicon,
lang=lang, lang=lang,
alternates=alternates,
)( )(
Title=f"{title} {brand}" if brand else title, Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html, link_lang), Brand=_brand_link(brand, brand_html, link_lang),
@@ -1174,7 +1227,7 @@ def _page_assets() -> tuple[list[str], list[str]]:
by the entry (e.g. overlayscrollbars.css) is extracted by Vite and must by the entry (e.g. overlayscrollbars.css) is extracted by Vite and must
be linked separately. be linked separately.
""" """
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = env.vite_url
if vite_url: if vite_url:
return [f"{vite_url}/src/pagerite.js"], [] return [f"{vite_url}/src/pagerite.js"], []
if "page" not in _asset_cache: if "page" not in _asset_cache:
@@ -1192,7 +1245,7 @@ def _editor_assets() -> tuple[list[str], str | None]:
The shared CSS is already linked on the page, so the pen only needs the The shared CSS is already linked on the page, so the pen only needs the
editor-specific stylesheet. editor-specific stylesheet.
""" """
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = env.vite_url
if vite_url: if vite_url:
return [f"{vite_url}/@vite/client", f"{vite_url}/src/main.js"], None return [f"{vite_url}/@vite/client", f"{vite_url}/src/main.js"], None
if "editor" not in _asset_cache: if "editor" not in _asset_cache:
@@ -1204,7 +1257,7 @@ def _editor_assets() -> tuple[list[str], str | None]:
def _analytics_assets() -> tuple[list[str], list[str]]: def _analytics_assets() -> tuple[list[str], list[str]]:
"""Script and stylesheet URLs for the analytics page entry.""" """Script and stylesheet URLs for the analytics page entry."""
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = env.vite_url
if vite_url: if vite_url:
return [f"{vite_url}/src/analytics-main.js"], [] return [f"{vite_url}/src/analytics-main.js"], []
if "analytics" not in _asset_cache: if "analytics" not in _asset_cache:
@@ -1216,6 +1269,31 @@ def _analytics_assets() -> tuple[list[str], list[str]]:
return _asset_cache["analytics"] return _asset_cache["analytics"]
def _langselect_assets() -> tuple[list[str], list[str]]:
"""Script and stylesheet URLs for the on-demand public language selector."""
vite_url = env.vite_url
if vite_url:
return [f"{vite_url}/src/langselect-main.js"], []
if "langselect" not in _asset_cache:
manifest = _manifest()
# import() loads no CSS automatically: collect the stylesheets of
# the entry and its imported chunks (LangSelect's ride on the
# shared langs chunk).
scripts, stylesheets, seen = [], [], set()
queue = ["src/langselect-main.js"]
for key in queue: # grows with imported chunks
if key in seen:
continue
seen.add(key)
entry = manifest[key]
if entry.get("isEntry"):
scripts.append(f"/{entry['file']}")
stylesheets += [f"/{css}" for css in entry.get("css", [])]
queue += entry.get("imports", [])
_asset_cache["langselect"] = scripts, stylesheets
return _asset_cache["langselect"]
def render_analytics( def render_analytics(
menu: dict[str, Node], menu: dict[str, Node],
brand: str = SITE_NAME, brand: str = SITE_NAME,
+2 -2
View File
@@ -17,7 +17,7 @@ readme = "README.md"
requires-python = ">=3.14" requires-python = ">=3.14"
dependencies = [ dependencies = [
"blake3>=1.0.9", "blake3>=1.0.9",
"fastapi-vue~=1.4.0", "fastapi-vue~=1.6.1",
"fastapi[standard]>=0.141.1", "fastapi[standard]>=0.141.1",
"html5tagger>=2.0.0", "html5tagger>=2.0.0",
"httpx>=0.28.1", "httpx>=0.28.1",
@@ -29,7 +29,7 @@ dependencies = [
"platformdirs>=4.11.5", "platformdirs>=4.11.5",
"pygments>=2.20.0", "pygments>=2.20.0",
"python-slugify>=8.0.4", "python-slugify>=8.0.4",
"ua-parser>=1.0.2", "uarite>=0.1.2",
"zstandard>=0.25.0", "zstandard>=0.25.0",
] ]
+10 -5
View File
@@ -1,11 +1,12 @@
#!/usr/bin/env -S uv run #!/usr/bin/env -S uv run
# auto-upgrade@fastapi-vue-setup - remove this if you modify this file
"""Run Vite development server for Vue app and FastAPI backend with auto-reload.""" """Run Vite development server for Vue app and FastAPI backend with auto-reload."""
import argparse import argparse
import asyncio import asyncio
import os import os
import subprocess
import sys import sys
from contextlib import suppress
from pathlib import Path from pathlib import Path
import tracerite import tracerite
@@ -47,11 +48,11 @@ async def run_devserver(
os.environ["PAGERITE_DEV"] = "1" os.environ["PAGERITE_DEV"] = "1"
async with ProcessGroup() as pg: async with ProcessGroup() as pg:
pg.create_task(check_ports_free(viteurl, backurl))
npm_i = await pg.spawn(*npm_install, cwd=front) npm_i = await pg.spawn(*npm_install, cwd=front)
await check_ports_free(viteurl, backurl) await pg.spawn(*pagerite, *(extra_args or []), vital=True)
await pg.spawn(*pagerite, *(extra_args or []))
await pg.wait(npm_i, ready(backurl, path=HEALTH)) await pg.wait(npm_i, ready(backurl, path=HEALTH))
await pg.spawn(*vite, cwd=front) await pg.spawn(*vite, cwd=front, vital=True)
def main() -> None: def main() -> None:
@@ -74,8 +75,12 @@ def main() -> None:
help=f"FastAPI (default: localhost:{DEFAULT_DEV_PORT})", help=f"FastAPI (default: localhost:{DEFAULT_DEV_PORT})",
) )
args, extra_args = parser.parse_known_args() args, extra_args = parser.parse_known_args()
with suppress(KeyboardInterrupt): try:
asyncio.run(run_devserver(args.listen, args.backend, extra_args)) asyncio.run(run_devserver(args.listen, args.backend, extra_args))
except* KeyboardInterrupt:
pass # user stopped the devserver: normal exit
except* subprocess.SubprocessError, RuntimeError:
raise SystemExit(1) from None # logged in devutil already; exit 1
HELP_EPILOG = """ HELP_EPILOG = """
-1
View File
@@ -1,4 +1,3 @@
# ruff: noqa: INP001
"""Hatch build hook for building Vue frontend during package build.""" """Hatch build hook for building Vue frontend during package build."""
import sys import sys
+2 -3
View File
@@ -1,4 +1,3 @@
# ruff: noqa: INP001
"""Utilities used at build time and in devserver script. No dependencies.""" """Utilities used at build time and in devserver script. No dependencies."""
import logging import logging
@@ -33,7 +32,7 @@ def _check_node_version(node_path: str) -> None:
Raises RuntimeError if version is too old or cannot be determined. Raises RuntimeError if version is too old or cannot be determined.
""" """
try: try:
result = subprocess.run( # noqa: S603 result = subprocess.run(
[node_path, "--version"], [node_path, "--version"],
capture_output=True, capture_output=True,
text=True, text=True,
@@ -221,7 +220,7 @@ def build(folder: str = "frontend") -> None:
def run(cmd: list[str]) -> None: def run(cmd: list[str]) -> None:
display_cmd = [Path(cmd[0]).stem, *cmd[1:]] display_cmd = [Path(cmd[0]).stem, *cmd[1:]]
logger.info("### %s", " ".join(display_cmd)) logger.info("### %s", " ".join(display_cmd))
subprocess.run(cmd, check=True, cwd=folder) # noqa: S603 subprocess.run(cmd, check=True, cwd=folder)
try: try:
run(install_cmd) run(install_cmd)
+77 -97
View File
@@ -1,111 +1,90 @@
# ruff: noqa: INP001
"""Utilities meant for devserver script, used only in source repository with dev deps.""" """Utilities meant for devserver script, used only in source repository with dev deps."""
from __future__ import annotations
import asyncio import asyncio
import subprocess
import sys import sys
from asyncio.subprocess import Process
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, Any, Self from subprocess import CalledProcessError
from typing import TYPE_CHECKING, Any
from urllib.parse import urlsplit from urllib.parse import urlsplit
from buildutil import find_dev_tool, find_install_tool, logger
from fastapi_vue.hostutil import parse_endpoint from fastapi_vue.hostutil import parse_endpoint
from buildutil import find_dev_tool, find_install_tool, logger
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Coroutine from collections.abc import Awaitable
class ProcessGroup: class ProcessGroup(asyncio.TaskGroup):
"""Manage async subprocesses with automatic cleanup, like TaskGroup for processes.""" """TaskGroup with structured ownership of async subprocesses."""
def __init__(self) -> None: def __init__(self, *, terminate_timeout: float = 10) -> None:
"""Initialize empty process tracking.""" """Set the grace period before terminate() escalates to kill()."""
self._procs: list[asyncio.subprocess.Process] = [] super().__init__()
self._cmds: dict[int, str] = {} # pid -> command name self._terminate_timeout = terminate_timeout
self._cmds: dict[Process, tuple[str, ...]] = {}
async def spawn( async def spawn(
self, self, *cmd: str, cwd: str | None = None, vital: bool = False
*cmd: str, ) -> Process:
cwd: str | None = None, """Spawn and own a subprocess. If a vital process exits, the group cancels."""
) -> asyncio.subprocess.Process:
"""Spawn a subprocess and track it."""
cmd_name = Path(cmd[0]).stem
logger.info(">>> %s", " ".join([cmd_name, *cmd[1:]]))
proc = await asyncio.create_subprocess_exec(*cmd, cwd=cwd)
self._procs.append(proc)
self._cmds[proc.pid] = cmd_name
return proc
async def wait( async def run() -> None:
self, name = Path(cmd[0]).stem
*waitables: "asyncio.subprocess.Process | Coroutine[Any, Any, Any]", logger.info(">>> %s", " ".join([name, *cmd[1:]]))
) -> None: try:
"""Wait for processes/coroutines to complete, raise SystemExit on failure.""" proc = await asyncio.create_subprocess_exec(*cmd, cwd=cwd)
self._cmds[proc] = cmd
started.set_result(proc)
except Exception as e: # noqa: BLE001
started.set_exception(e)
return
async def wait_proc(proc: asyncio.subprocess.Process) -> None: try:
returncode = await proc.wait() returncode = await proc.wait()
if returncode != 0: finally:
cmd_name = self._cmds.get(proc.pid, "unknown")
raise subprocess.CalledProcessError(returncode, cmd_name)
tasks = [
wait_proc(w) if isinstance(w, asyncio.subprocess.Process) else w
for w in waitables
]
try:
await asyncio.gather(*tasks)
except subprocess.CalledProcessError as e:
logger.warning("%s failed with exit status %d", e.cmd, e.returncode)
raise SystemExit(1) from None
async def __aenter__(self) -> Self:
"""Enter the async context manager."""
return self
async def __aexit__(self, exc_type: type[BaseException] | None, *_: object) -> None:
"""Wait for one process to exit, terminate others, then wait for all."""
await self._cleanup(immediate=exc_type is not None)
async def _cleanup(self, *, immediate: bool = False) -> None:
running = [p for p in self._procs if p.returncode is None]
if not running:
return
if not immediate:
# Wait for any one process to exit
with suppress(asyncio.CancelledError):
await asyncio.wait(
[asyncio.create_task(p.wait()) for p in running],
return_when=asyncio.FIRST_COMPLETED,
)
# Terminate remaining processes
for p in self._procs:
if p.returncode is None:
with suppress(ProcessLookupError): with suppress(ProcessLookupError):
p.terminate() proc.terminate()
# Wait for all to finish (with overall timeout), shielded from cancellation
still_running = [p for p in self._procs if p.returncode is None]
if still_running:
with suppress(asyncio.CancelledError):
try: try:
await asyncio.shield( await asyncio.wait_for(proc.wait(), self._terminate_timeout)
asyncio.wait_for(
asyncio.gather(*[p.wait() for p in still_running]),
timeout=10,
),
)
except TimeoutError: except TimeoutError:
for p in self._procs: with suppress(ProcessLookupError):
if p.returncode is None: proc.kill()
with suppress(ProcessLookupError): await proc.wait()
p.kill()
await p.wait() if vital:
logger.warning("Vital process %s exited", name)
raise CalledProcessError(returncode, cmd)
started = asyncio.get_running_loop().create_future()
self.create_task(run())
return await asyncio.shield(started)
async def wait(self, *waitables: Process | Awaitable) -> tuple[Any, ...]:
"""Wait concurrently and return results in argument order."""
async def task(w: Process | Awaitable) -> Any:
if not isinstance(w, Process):
return await w
if retcode := await w.wait():
cmd = self._cmds[w]
logger.warning(
"Process %s exited with status %d", Path(cmd[0]).stem, retcode
)
raise CalledProcessError(retcode, cmd)
return retcode
async with asyncio.TaskGroup() as group:
tasks = [group.create_task(task(w)) for w in waitables]
return tuple(task.result() for task in tasks)
async def http_get_server(url: str, timeout: float) -> str | None: # noqa: ASYNC109 async def http_get_server(url: str, timeout: float) -> str | None:
"""GET url with plain asyncio streams, return the response Server header. """GET url with plain asyncio streams, return the response Server header.
Returns an empty string when the server responds without a Server header, Returns an empty string when the server responds without a Server header,
@@ -128,31 +107,32 @@ async def http_get_server(url: str, timeout: float) -> str | None: # noqa: ASYN
writer.close() writer.close()
except OSError, EOFError, ValueError, TimeoutError: except OSError, EOFError, ValueError, TimeoutError:
return None return None
for line in data.decode("latin-1").split("\r\n"): for line in data.decode(errors="replace").split("\r\n"):
if line.lower().startswith("server:"): if line.lower().startswith("server:"):
return line.split(":", 1)[1].strip() return line[7:].strip()
return "" return ""
async def check_ports_free(*urls: str) -> None: async def check_ports_free(*urls: str) -> None:
"""Verify URLs are not responding (ports are free). Raise SystemExit if any respond.""" """Verify URLs are not responding (ports are free).
async def check(url: str) -> None: Meant to run as a task inside a TaskGroup. Logs the conflict and raises
server = await http_get_server(url, timeout=0.1) RuntimeError (handled like a failed process) if any URL responds.
"""
servers = await asyncio.gather(*(http_get_server(url, timeout=0.1) for url in urls))
for url, server in zip(urls, servers, strict=True):
if server is not None: if server is not None:
logger.warning( logger.error(
"Conflicting %s already running at %s", server or "server", url "Conflicting %s already running at %s", server or "server", url
) )
raise SystemExit(1) raise RuntimeError(url)
await asyncio.gather(*[check(url) for url in urls])
async def ready(url: str, path: str = "", max_attempts: int = 50) -> None: async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
"""Wait for the server to be ready by polling an endpoint. """Wait for the server to be ready by polling an endpoint.
Use empty path to disable the check and make this return immediately. Use empty path to disable the check and make this return immediately.
Raises SystemExit(1) if server doesn't start in time. Logs, then raises RuntimeError if the server doesn't start in time.
""" """
if not path: if not path:
return return
@@ -162,8 +142,8 @@ async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
logger.info("✓ Backend ready!") logger.info("✓ Backend ready!")
return return
if attempt == max_attempts - 1: if attempt == max_attempts - 1:
logger.warning("Backend didn't start in time") logger.error("Backend at %s didn't start in time", url)
raise SystemExit(1) raise RuntimeError(url)
await asyncio.sleep(0.1) await asyncio.sleep(0.1)