Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
33a4a76364 | ||
|
|
9fb4b5a681 | ||
|
|
0c1349b037 | ||
|
|
54f8c8e09b | ||
|
|
78f4ddb2f0 | ||
|
|
e6a0c57446 | ||
|
|
d6d07db2c5 | ||
|
|
38af57218a | ||
|
|
eb2e8f8273 | ||
|
|
792b9e7aa9 | ||
|
|
4c6ab3dde6 | ||
|
|
9d70f17587 | ||
|
|
029bfe105e | ||
|
|
62031fd5dd | ||
|
|
13cf716bb1 |
+218
-207
@@ -1,15 +1,15 @@
|
||||
# Analytics
|
||||
|
||||
Server-side visit analytics. Data lives in a plain JSON file — a msgspec
|
||||
Struct dumped to disk — separate from the kanta content database, path from
|
||||
`PAGERITE_ANALYTICS` (default: `analytics.json` in the per-site data
|
||||
directory, e.g. `localhost/analytics.json`).
|
||||
Server-side visit analytics built on a **raw access-log-style event store**.
|
||||
Data lives in a plain JSON file — a msgspec Struct dumped to disk — separate
|
||||
from the kanta content database, path from `PAGERITE_ANALYTICS` (default:
|
||||
`analytics.json` in the per-site data directory, e.g. `localhost/analytics.json`).
|
||||
|
||||
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
|
||||
`CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map,
|
||||
atomic JSON persistence).
|
||||
- `pagerite/pages.py` — entry-referer stashing in `show_page` (`_track_entry`,
|
||||
in `pagerite/tracking.py`), 404 recording.
|
||||
- `pagerite/analytics.py` — data model (`Analytics`, `Get`, `Msg`, `Client`,
|
||||
`Favicon`), the `Store` (raw log + atomic JSON persistence) and
|
||||
`Store.display()`, where **all** classification happens.
|
||||
- `pagerite/pages.py` — records every served document as one raw GET line
|
||||
(`_record_get`, in `pagerite/tracking.py`) with its true HTTP status.
|
||||
- `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
|
||||
`WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
|
||||
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
|
||||
@@ -18,172 +18,42 @@ directory, e.g. `localhost/analytics.json`).
|
||||
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
|
||||
into `#analytics-app` inside `#main`.
|
||||
|
||||
## What is collected
|
||||
## Raw records
|
||||
|
||||
The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
|
||||
whole browsing session and sends activity messages over it — JSON text
|
||||
frames matching the server's `Ping` msgspec struct with the fields `fr`
|
||||
(source path), `to` (navigation target), `read` (active seconds on `fr`
|
||||
since the last report) and `hide`; falsy fields are omitted. One channel
|
||||
follows the session, so the activity of a visit stays tied together, and
|
||||
while the user is active the accumulated reading time is flushed every few
|
||||
seconds: the trail times are cumulative, so a disconnection simply leaves
|
||||
the last reported time in place (no close beacon). After 5 minutes without
|
||||
any activity the client closes the socket itself — a sleeping browser tab
|
||||
would lose it anyway — and the next activity reconnects as a fresh session;
|
||||
reconnects are attempted only on user activity, with an exponential backoff
|
||||
between attempts so a failing endpoint is never hammered. Idle-time link preloads
|
||||
stay plain `fetch()` calls so the browser may cache the responses; the
|
||||
WebSocket reports actual navigations and active time spent on a page.
|
||||
The store is deliberately close to an access log: two append-only lists plus
|
||||
shared metadata. **Nothing is classified when recorded** — whether a client
|
||||
turns out to be a reader, a crawler or a scanner is decided by
|
||||
`Store.display()` from the raw events, so the stored data survives any future
|
||||
change to the classification rules.
|
||||
|
||||
- **Initial page load**: only `to` — the loaded path — is sent, never `fr`
|
||||
(an `fr` equal to `to` would log a bogus self-transition when a session
|
||||
already exists, e.g. a second tab). This message is what starts
|
||||
the visit and counts the entry page view — the document GET alone records
|
||||
nothing, so bots never register (admin browsing does register, but
|
||||
flagged `hide`; see **Admins** below). JS-running crawlers
|
||||
(Googlebot, GoogleOther, Applebot, ...) do connect and report, but their
|
||||
User-Agent gives them away: messages whose UA matches `_is_bot_ua`
|
||||
(anything calling
|
||||
itself a "bot", plus known exceptions such as GoogleOther) are ignored
|
||||
server-side, and their document GETs land in the crawler list instead.
|
||||
Real-browser bots whose UA does not match still register a visit, but
|
||||
their reported reading time stays under 5 seconds, so they are
|
||||
reclassified as crawler hits at display time (see **Crawler hits** below).
|
||||
No source-IP verification is done: a spoofed bot UA merely lands in the
|
||||
crawler stats, and scanners that probe telltale paths are caught by the
|
||||
abuse rules regardless. Reloads are not
|
||||
visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
|
||||
refresh neither counts a second view nor logs a self-transition. The GET
|
||||
handler stashes a cross-origin https `Referer` (origin part only —
|
||||
unavailable to JS once the page has loaded) and any
|
||||
`utm_*` query parameters in in-memory IP tables, consumed by the first
|
||||
message that
|
||||
starts the visit; internal or absent referers never touch the referer table.
|
||||
- **Internal fetch-navigations**: `to` is the target path, sent only after
|
||||
the swap actually happened (a failed swap falls back to a full load,
|
||||
whose initial message counts the view instead — no gap, no double count).
|
||||
- **External links** (`https` only): `to` is the link's full URL. This is the
|
||||
exit-link record; the user may continue navigating afterwards (new tab,
|
||||
back), so the exit URL is not necessarily the last trail entry. Outbound
|
||||
links are stored by full URL so several links to the same domain remain
|
||||
distinct.
|
||||
- **Excluded**: back/forward (popstate) navigations, navigating *to* the
|
||||
analytics page (`/_a` — its GET is untracked, and the server cannot
|
||||
record it as a navigation target anyway), and everything while the user has
|
||||
the editor
|
||||
open (`body.editing`). Admin noise, not visits. Navigating *away* from
|
||||
`/_a` does report: the fetch-navigation already GET-ed the target page
|
||||
without the preload header, and without the message that GET would flush to
|
||||
the crawler list.
|
||||
- **Admins**: when SSO is in use and the session is known to be an admin,
|
||||
the client still reports but adds `hide`. The activity is recorded as
|
||||
usual (navigations and all), but the `hide` flag is set on the **client
|
||||
record** — so it covers everything that client ever did: visits and
|
||||
crawler hits from before the login included. Hidden clients never appear
|
||||
in the viewer payload: `Store.display()` drops their visits, crawler
|
||||
hits, abuse hits and metadata, and computes every aggregate (site visits,
|
||||
page views, transitions) from the visible visits only, so nothing needs
|
||||
to be reversed or redacted. Pending crawler hits from a hidden client
|
||||
are discarded when they expire, so admin browsing never lands in the
|
||||
crawler list either. With no auth proxy
|
||||
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
|
||||
is recorded.
|
||||
- The server validates `to`: internal paths must be valid slug paths
|
||||
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
|
||||
https origin and accepted only when the client sent exactly that.
|
||||
- **External-site favicons**: for every external https origin seen as a visit
|
||||
referer, a crawler-hit referer or an exit link, the server fetches `{origin}/favicon.ico` in a
|
||||
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
|
||||
SVG is sniffed from the body when served without an image type) and stores
|
||||
the icon content-hashed on disk in the FileStore (served at `/_f/{name}`,
|
||||
extension matching the actual MIME). The origin → file name mapping is
|
||||
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
|
||||
recorded too and retried only after 7 days. Fetches are scheduled after
|
||||
each activity message and once at startup, which backfills icons for already-recorded
|
||||
data. The viewer payload carries `favicons` (origin → `/_f/...` path),
|
||||
and the viewer shows the icon wherever an external site is mentioned:
|
||||
referer/exit trail links in the visit table and the source/exit pills of
|
||||
the transition map (UTM-attributed source nodes without an https origin
|
||||
stay text-only).
|
||||
- **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
|
||||
`User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
|
||||
the first 6 bytes identify a shared `Client` record. The `Client` stores
|
||||
the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
|
||||
`country` from the language-region subtag, and asynchronously-filled
|
||||
`country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
|
||||
crawler hits and abuse hits all reference this record by its hash, so
|
||||
client metadata is stored once instead of repeated per event.
|
||||
- The visitor IP is stored in the `Client`. A reverse-DNS lookup is
|
||||
attempted for each new client and the result, when available, is stored as
|
||||
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
|
||||
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
|
||||
root, it is loaded at startup and used to look up `country`/`city`. These
|
||||
lookups run in background tasks after the event is stored, so WebSocket
|
||||
message handling is never delayed. The decompressed `dbip-*.mmdb` file is kept in
|
||||
the repository root and ignored by git. The CLI flag `--dbip`
|
||||
(`uv run pagerite --dbip`) downloads the latest
|
||||
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
|
||||
skipping the download when the local database is already current and
|
||||
removing older versions after an update; without the flag only an existing
|
||||
file is used.
|
||||
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
|
||||
hit — except idle-time link preloads from pagerite.js, which carry an
|
||||
`x-pagerite-preload` header and are not tracked at all (the navigation
|
||||
message sent when the user actually navigates to a preloaded page does
|
||||
the counting; forging
|
||||
the header only hides a GET from the crawler stats, the path-based abuse
|
||||
classification is unaffected). If a message
|
||||
from the same client arrives within 10 seconds the hit is discarded;
|
||||
otherwise it is written to `crawlers` — unless the client is hidden
|
||||
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as
|
||||
visits or views. Bots running real browsers can still slip past the UA
|
||||
check: a visit whose total reported reading time stays under 5 seconds
|
||||
(`_MIN_VISIT_READ`; durations are client-provided and trusted — such bots
|
||||
report 0–2 s) is reclassified as crawler hits at display time, one hit
|
||||
per internal trail page, and counts in no visit aggregate. The
|
||||
`Accept-Language` header is stored on the shared
|
||||
`Client` immediately; reverse-DNS host names and DB-IP geoip
|
||||
country/city are filled in asynchronously, just like for real visits. In
|
||||
the analytics viewer, crawler hits are grouped by client hash and shown as
|
||||
a trail of internal pages that crawler visited, preceded by its referer
|
||||
when there is one — spiders often advertise their own site as the
|
||||
referer, and it is rendered with its favicon like visit referers (crawler
|
||||
referers are included in the favicon fetch origins). The crawler table lists
|
||||
the most recent crawler first, with the most active as a tie-breaker.
|
||||
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
|
||||
starting with a dot (`/.env`, `/.git/config`) or ending in `.php` —
|
||||
classifies the source IP as abuse immediately, and ten plain 404s from one
|
||||
IP do too. Classification reclassifies history: all earlier crawler hits
|
||||
from that IP (persisted and pending) move to the `abuse` list, so a
|
||||
random-UA scanner no longer pollutes the crawler stats of the legitimate
|
||||
bot it impersonates. Once classified, every document GET and 404 from the
|
||||
IP is recorded as an abuse hit with the full request path (query string
|
||||
included), and its activity messages are ignored. The classified IP set (`abuse_ips`)
|
||||
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
|
||||
viewer, abuse hits are grouped by IP (never by client/UA — scanners
|
||||
randomize theirs) in a separate "Abuse" table. Identical paths are
|
||||
collapsed into one entry with their hit count. The 404 probes ("paths
|
||||
abused": flagged paths that triggered classification first, then other
|
||||
404s) are kept in a separate column from the real articles the abuser
|
||||
actually read ("articles read": document GETs that returned 200, not the
|
||||
404 fallback rendering — rendered as trail links like the visitor and
|
||||
crawler tables, with the query string stripped). Raw User-Agent strings are shown one
|
||||
per line with their occurrence counts, and the full lists are click-to-copy.
|
||||
Each `Get` record (one per served document):
|
||||
|
||||
## Visits and sessions
|
||||
- `t` — timestamp of the request,
|
||||
- `path` — full request path, query string included (e.g. `/.env?x=1`),
|
||||
- `status` — the true HTTP status of the response (200, or 404 for a category
|
||||
placeholder or a missing page),
|
||||
- `ref` — external https origin of the `Referer`, `""` for direct/internal
|
||||
(same-origin referers are dropped by the recorder),
|
||||
- `pre` — true for idle-time link preloads from pagerite.js
|
||||
(`x-pagerite-preload` header): never counted as a view, crawler hit or
|
||||
abuse — recorded only so a navigation later served from the in-memory page
|
||||
cache (which issues no GET at all) can be attributed this GET's status,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`.
|
||||
|
||||
There are no cookies. A visit is tied together by a client hash — the first
|
||||
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
|
||||
/64 network), the raw `User-Agent` string and the extracted
|
||||
`Accept-Language` tag. The first message from a client hash starts a new
|
||||
visit; subsequent messages extend it. Messages arriving with no known session
|
||||
(server restart) start a fresh visit from the first message — treated as
|
||||
missing data rather than dropped. The client-hash → visit map and the IP →
|
||||
entry-referer/UTM tables are in-memory only; client metadata is stored in
|
||||
`Analytics.clients` keyed by the client hash.
|
||||
304 revalidation responses return before recording and are not logged.
|
||||
|
||||
Each `Client` record:
|
||||
Each `Msg` record (one per pagerite.js activity message over `/_ws`):
|
||||
|
||||
- `t` — timestamp,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `fr` — path of the page the activity happened on (`""` for the initial
|
||||
load),
|
||||
- `to` — navigation target (validated at record time: internal slug path or
|
||||
external https URL; anything else is dropped — sanitation, not
|
||||
classification),
|
||||
- `read` — active seconds spent on `fr` since the previous report.
|
||||
|
||||
Each `Client` record (shared by every event, keyed by hash):
|
||||
|
||||
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
|
||||
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
|
||||
@@ -195,28 +65,174 @@ Each `Client` record:
|
||||
- `ua` — raw `User-Agent` string,
|
||||
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
|
||||
parsable, otherwise the raw string,
|
||||
- `hide` — true for admin clients (`hide` message field): all their visits,
|
||||
crawler hits and abuse hits are recorded but excluded from every
|
||||
statistic and from the viewer payload.
|
||||
- `hide` — true for admin clients (`hide` message field): everything this
|
||||
client ever did is recorded but excluded from every statistic and from the
|
||||
viewer payload. This is the one flag set at record time — it is a client
|
||||
property, not a classification.
|
||||
|
||||
Each `Visit` record:
|
||||
A reverse-DNS lookup is attempted for each new client and the result, when
|
||||
available, is stored as `host`; local/reserved/multicast addresses are
|
||||
skipped. If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present
|
||||
in the repository root, it is loaded at startup and used to look up
|
||||
`country`/`city`. These lookups run in background tasks after the event is
|
||||
stored, so WebSocket message handling is never delayed. The decompressed
|
||||
`dbip-*.mmdb` file is kept in the repository root and ignored by git. The
|
||||
CLI flag `--dbip` (`uv run pagerite --dbip`) downloads the latest
|
||||
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP at startup (in the app lifespan,
|
||||
before the MMDB is opened), skipping the download when the local database is
|
||||
already current and removing older versions after an update; without the flag
|
||||
only an existing file is used.
|
||||
|
||||
- `start` — timestamp of the first event,
|
||||
## What the client sends
|
||||
|
||||
The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
|
||||
whole browsing session and sends activity messages over it — JSON text
|
||||
frames matching the server's `Ping` msgspec struct with the fields `fr`
|
||||
(source path), `to` (navigation target), `read` (active seconds on `fr`
|
||||
since the last report) and `hide`; falsy fields are omitted. One channel
|
||||
follows the session, so the activity of a visit stays tied together, and
|
||||
while the user is active the accumulated reading time is flushed every few
|
||||
seconds: the times are incremental, so a disconnection simply leaves the
|
||||
last reported time in place (no close beacon). After 5 minutes without
|
||||
any activity the client closes the socket itself — a sleeping browser tab
|
||||
would lose it anyway — and the next activity reconnects; reconnects are
|
||||
attempted only on user activity, with an exponential backoff between
|
||||
attempts so a failing endpoint is never hammered. Idle-time link preloads
|
||||
stay plain `fetch()` calls so the browser may cache the responses; the
|
||||
WebSocket reports actual navigations and active time spent on a page.
|
||||
|
||||
- **Initial page load**: only `to` — the loaded path — is sent, never `fr`
|
||||
(an `fr` equal to `to` would log a bogus self-transition when a session
|
||||
already exists, e.g. a second tab). Reloads are not
|
||||
visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
|
||||
refresh neither counts a second view nor logs a self-transition.
|
||||
- **Internal fetch-navigations**: `to` is the target path, sent only after
|
||||
the swap actually happened (a failed swap falls back to a full load,
|
||||
whose initial message counts the view instead — no gap, no double count).
|
||||
- **External links** (`https` only): `to` is the link's full URL. This is the
|
||||
exit-link record; the user may continue navigating afterwards (new tab,
|
||||
back), so the exit URL is not necessarily the last trail entry. Outbound
|
||||
links are stored by full URL so several links to the same domain remain
|
||||
distinct.
|
||||
- **Excluded**: back/forward (popstate) navigations, navigating *to* the
|
||||
analytics page (`/_a` — its GET is untracked, and the server cannot
|
||||
record it as a navigation target anyway), and everything while the user has
|
||||
the editor open (`body.editing`). Admin noise, not visits. Navigating
|
||||
*away* from `/_a` does report.
|
||||
- **Admins**: when SSO is in use and the session is known to be an admin,
|
||||
the client still reports but adds `hide`. The activity is recorded as
|
||||
usual (navigations and all), but the `hide` flag is set on the **client
|
||||
record** — so it covers everything that client ever did, including the
|
||||
time before the login. Hidden clients never appear in the viewer payload:
|
||||
`Store.display()` drops their events and metadata, and computes every
|
||||
aggregate (site visits, page views, transitions) from the visible visits
|
||||
only, so nothing needs to be reversed or redacted. With no auth proxy
|
||||
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
|
||||
is recorded.
|
||||
- **External-site favicons**: for every external https origin seen as a GET
|
||||
referer or an exit link, the server fetches `{origin}/favicon.ico` in a
|
||||
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
|
||||
SVG is sniffed from the body when served without an image type) and stores
|
||||
the icon content-hashed on disk in the FileStore (served at `/_f/{name}`,
|
||||
extension matching the actual MIME). The origin → file name mapping is
|
||||
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
|
||||
recorded too and retried only after 7 days. Fetches are scheduled after
|
||||
each activity message and once at startup, which backfills icons for
|
||||
already-recorded data. The viewer payload carries `favicons` (origin →
|
||||
`/_f/...` path), and the viewer shows the icon wherever an external site
|
||||
is mentioned: referer/exit trail links in the visit table and the
|
||||
source/exit pills of the transition map (UTM-attributed source nodes
|
||||
without an https origin stay text-only).
|
||||
|
||||
## Display-time classification
|
||||
|
||||
`Store.display(in_menu)` derives the viewer payload from the raw events on
|
||||
every (debounced) broadcast — O(n log n) over the log, cheap enough for a
|
||||
small CMS. `in_menu(path)` resolves a path against the current menu (passed
|
||||
in from `tracking.py`, which owns the content database import) so 404
|
||||
responses for real menu nodes — category placeholders — are not mistaken
|
||||
for misses.
|
||||
|
||||
- **Visits and sessions**: a client's messages are grouped into visits
|
||||
chronologically; a new visit starts after 30 minutes of inactivity
|
||||
(`_SESSION_GAP`). A fresh page load with an already-open visit (second
|
||||
tab) extends it, logging a `(direct)` transition. The visit's trail holds
|
||||
first-seen targets in order; `read` updates accumulate active seconds on
|
||||
the trail item matching `fr`. Each trail item's HTTP status comes from
|
||||
the client's latest GET for that path — preloads included, which is what
|
||||
allows 404 pages to render red in the viewer even when the navigation
|
||||
itself was served from the page cache. The entry page's referer and
|
||||
`utm_*` tags come from the GET that loaded it (within 10 s before the
|
||||
first message).
|
||||
- **Crawler hits**: a document GET no activity message matched within
|
||||
`_CRAWLER_TIMEOUT` (10 s) is a crawler hit — plain bots that only fetch
|
||||
documents never register as visits. JS-running crawlers (Googlebot,
|
||||
GoogleOther, Applebot, ...) do connect and send messages, but their UA
|
||||
gives them away (`_is_bot_ua`): their messages are ignored at display
|
||||
time, so their GETs never match and land in the crawler list too. Real-
|
||||
browser bots whose UA does not match are caught by engagement: a visit
|
||||
whose total reported reading time is under 5 seconds (`_MIN_VISIT_READ`;
|
||||
durations are client-provided and trusted — such bots report 0–2 s) is
|
||||
reclassified as crawler hits, one per internal trail page, and counts in
|
||||
no visit aggregate. No source-IP verification is done: a spoofed bot UA
|
||||
merely lands in the crawler stats, and scanners that probe telltale paths
|
||||
are caught by the abuse rules regardless. In the viewer, crawler hits are
|
||||
grouped by client hash and shown as a trail of pages, preceded by the
|
||||
referer when there is one (rendered with its favicon like visit
|
||||
referers). The crawler table lists the most recent crawler first, with
|
||||
the most active as a tie-breaker.
|
||||
- **Abuse (scanner) hits**: a 404 on a telltale path — an empty URL segment
|
||||
(`//foo` — no real client generates those), any segment starting with a
|
||||
dot (`/.env`, `/.git/config`) or ending in `.php` — classifies the source
|
||||
IP as abuse, and ten plain 404s within one hour (`_ABUSE_404_WINDOW`) on
|
||||
paths that don't resolve to a menu node do too. Two exemptions keep
|
||||
legitimate traffic out: RFC 8615 well-known URIs (`/.well-known/…` —
|
||||
browsers and services probe them, e.g. Chrome's devtools fetch of
|
||||
`appspecific/com.chrome.devtools.json`) are never telltale and never
|
||||
count toward the threshold, and category placeholders return 404 but are
|
||||
real menu nodes, so they never count either. The window keeps a
|
||||
long-time reader's slowly accumulating misses from ever crossing the
|
||||
threshold — scanners spray in bursts. Hidden (admin) clients never
|
||||
trigger classification: editing means visiting not-found pages, since
|
||||
that is where the create pen lives. Once an IP is classified, **all** its document GETs are shown in the abuse list —
|
||||
including any that arrived before classification, since the raw log keeps
|
||||
everything — and its activity messages are ignored. In the viewer, abuse
|
||||
hits are grouped by IP (never by client/UA — scanners randomize theirs)
|
||||
in a separate "Abuse" table, split by the recorded status: the 404 probes
|
||||
("paths abused" — flagged paths that triggered classification first, then
|
||||
other 404s, shown verbatim with query strings) versus the real articles
|
||||
the abuser actually read ("articles read" — the 200 document GETs,
|
||||
rendered as trail links like the visitor and crawler tables, query string
|
||||
stripped). Raw User-Agent strings are shown one per line with their
|
||||
occurrence counts, and the full lists are click-to-copy.
|
||||
|
||||
In the visitor and crawler tables, internal paths that returned a 404 status
|
||||
are shown in red and the link title includes the status code, so it is easy
|
||||
to tell misses from real pages at a glance.
|
||||
|
||||
## Derived shapes (the viewer payload)
|
||||
|
||||
The `Display` payload contains the derived `visits`, `crawlers` and `abuse`
|
||||
rows (structs `Visit`/`Nav`/`TrailItem`, `CrawlerHit`, `AbuseHit` — display
|
||||
DTOs only, never persisted), the visible `clients`, the fetched `favicons`,
|
||||
and the aggregates below.
|
||||
|
||||
Each derived `Visit`:
|
||||
|
||||
- `start` — timestamp of the first activity,
|
||||
- `entry` — first page (path) seen,
|
||||
- `referer` — external https origin of the initial load, `""` for direct,
|
||||
- `referer` — external https origin of the entry GET, `""` for direct,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `trail` — the entry page and everything seen afterwards, keyed by the
|
||||
timestamp of first sight (insertion order = first-seen order). Each item
|
||||
holds `to` (page path or external exit URL), the accumulated active
|
||||
reading time in seconds (`read`) and the most recent HTTP status seen
|
||||
for the target (`status`). Re-visiting an already seen target updates
|
||||
its item instead of appending.
|
||||
- `navs` — every navigation message (`fr`, `to`), keyed by its timestamp,
|
||||
repeats included. The aggregates are computed from this log at display
|
||||
time.
|
||||
for the target (`status`),
|
||||
- `navs` — every navigation (`fr`, `to`), keyed by its timestamp, repeats
|
||||
included. The aggregates are computed from this log,
|
||||
- `utm` — `utm_*` query parameters from the landing URL, as a dict.
|
||||
|
||||
Each `CrawlerHit` record:
|
||||
Each derived `CrawlerHit`:
|
||||
|
||||
- `start` — timestamp of the document GET,
|
||||
- `entry` — page path requested,
|
||||
@@ -226,39 +242,31 @@ Each `CrawlerHit` record:
|
||||
- `status` — HTTP status of the served response (200 for a real page, 404
|
||||
for a category placeholder or missing page).
|
||||
|
||||
Each `AbuseHit` record:
|
||||
Each derived `AbuseHit`:
|
||||
|
||||
- `start` — timestamp of the request,
|
||||
- `path` — full request path including the query string (e.g. `/.env?x=1`),
|
||||
- `path` — full request path including the query string,
|
||||
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||
- `flag` — true for the path that triggered abuse classification (telltale
|
||||
path or the 404 that crossed the threshold),
|
||||
- `is_404` — true for 404 responses (probed paths and 404-fallback document
|
||||
GETs), false for real (200) document GETs — articles the abuser read.
|
||||
- `flag` — true for the paths that triggered abuse classification (telltale
|
||||
paths, or the 404 that crossed the threshold),
|
||||
- `is_404` — true for 404 responses, false for real (200) document GETs.
|
||||
|
||||
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
|
||||
are grouped by IP alone (resolved from the referenced `Client`). In the
|
||||
Abuse table identical paths are collapsed with their counts, split into the
|
||||
404 probes (flagged paths that triggered classification first, then other
|
||||
404s, shown verbatim) and the 200 document GETs shown as trail links in the
|
||||
separate articles column.
|
||||
Within each list paths are
|
||||
sorted by count descending, then by their earliest hit.
|
||||
|
||||
In the visitor and crawler tables, internal paths that returned a 404 status
|
||||
are shown in red and the link title includes the status code, so it is easy
|
||||
to tell misses from real pages at a glance.
|
||||
are grouped by IP alone (resolved from the referenced `Client`). In the
|
||||
Abuse table identical requests (same path and status class) are collapsed
|
||||
with their counts — a path's 404 probes and its later 200 reads never
|
||||
merge. Within each list paths are sorted by count descending, then by their
|
||||
earliest hit.
|
||||
|
||||
## Aggregates
|
||||
|
||||
Aggregates are **not stored**; they are computed at display time by
|
||||
`Store.display()` from the visit records (entry + `navs` log), skipping
|
||||
hidden clients' visits and short visits reclassified as crawler hits
|
||||
(under `_MIN_VISIT_READ` seconds of total reported reading time). This is
|
||||
what allows a client to become hidden after
|
||||
navigations were already logged: no counts need reversing. The computed
|
||||
shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
|
||||
`crawlers`, `abuse` and `clients`):
|
||||
`Store.display()` from the derived visits (entry + `navs` log), skipping
|
||||
hidden clients and short visits reclassified as crawler hits. This is
|
||||
what allows a client to become hidden after navigations were already
|
||||
logged: no counts need reversing. The computed shapes, part of the
|
||||
WebSocket payload (`Display` struct alongside `visits`, `crawlers`, `abuse`
|
||||
and `clients`):
|
||||
|
||||
- `transitions`: time series of page transitions, sparse nested dict
|
||||
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
|
||||
@@ -272,13 +280,16 @@ shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
|
||||
5-minute bucketing.
|
||||
|
||||
Sparseness keeps quiet sites small; dropping old data is a matter of deleting
|
||||
list entries (`visits` is a plain append-only list).
|
||||
list entries (`gets`/`msgs` are plain append-only lists).
|
||||
|
||||
## Persistence
|
||||
|
||||
The whole `Analytics` struct is JSON-encoded and written atomically
|
||||
(temp file + rename) on every recorded event. Traffic on a small CMS makes
|
||||
this cheap enough; batching can be added later without changing the format.
|
||||
A file written by the pre-redesign schema (stored `visits`/`crawlers`/`abuse`
|
||||
lists) is not convertible; it is renamed to `analytics.json.bak-legacy` and
|
||||
recording starts fresh.
|
||||
|
||||
## Viewing
|
||||
|
||||
|
||||
@@ -19,8 +19,8 @@ Pagerite is a single-user CMS/blog. This document records the initial high-level
|
||||
|
||||
- Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
|
||||
- **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
|
||||
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's left (the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
|
||||
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `{.right}` — `{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone left of the text on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
|
||||
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's start edge (left in LTR, right in RTL — the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
|
||||
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `{.right}` — `{.right}`, `{.left}` float at 30% of the text column, to its end/start edge following the text direction (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone at the text's start edge on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
|
||||
|
||||
## Page structure and navigation
|
||||
|
||||
|
||||
@@ -26,6 +26,7 @@ Vite builds ES-module `.js` outputs; in dev the backend links them as `<script t
|
||||
|
||||
All site data lives under `<hostname>/` in the cwd — `content.kantadb`,
|
||||
`analytics.json` and `files/` — where `<hostname>` is the CLI's first
|
||||
positional argument (default `localhost`, exported as `PAGERITE_HOSTNAME`;
|
||||
positional argument (default `localhost`, passed to the app as JSON in
|
||||
`PAGERITE_CONFIG`, see `pagerite/config.py`;
|
||||
`PAGERITE_DB`/`PAGERITE_ANALYTICS`/`PAGERITE_FILES` override individual
|
||||
paths). gitignored. Do not delete it without asking.
|
||||
|
||||
+34
-13
@@ -88,6 +88,10 @@ Region tags normalize to their base subtag (`fi-FI` → `fi`).
|
||||
### Rendering
|
||||
|
||||
- The translated Markdown goes through the same `markdown.render` pipeline.
|
||||
- Section anchors (`#hash` ids on h1/h2 headings) stay in the original
|
||||
language: render(anchors_from=...) pins the translated render's heading
|
||||
ids to the original text's slugs, matched by heading position, so links
|
||||
to sections don't break across languages.
|
||||
- Navigation/sidebar titles come from the translation's title map, with
|
||||
per-node fallback to the original title (a partially translated tree must
|
||||
still render).
|
||||
@@ -305,14 +309,15 @@ An external machine-translation service connects over WebSocket at
|
||||
`/_translate/{key}` — deliberately **not** under `/_api`: the SSO
|
||||
forward-auth does not cover that route, and the key in the path is the
|
||||
access control. Keys live in `Data.translate_keys` (key -> display name) —
|
||||
12 lowercase alphanumeric characters each, the first one generated at
|
||||
database bootstrap and multiple keys reserved for future management (e.g.
|
||||
a web UI). The full WS URL(s) are printed in the startup log
|
||||
(`ws://localhost:{port}/_translate/{key}` locally,
|
||||
`wss://{hostname}/_translate/{key}` on a public hostname) and the keys are
|
||||
surfaced to the admin in `GET /_api/settings` as `translate_keys`. An
|
||||
unknown or empty key rejects the handshake (close-before-accept → HTTP
|
||||
403). Transactions storing results record the connecting key as the kanta
|
||||
12 lowercase alphanumeric characters each; the first is generated at
|
||||
database bootstrap, further ones are managed in the editor's lang tab
|
||||
(add/rename/delete ride the `PUT /_api/settings` round-trip; the name is
|
||||
an inline display label only). The full WS URL(s) are printed in the
|
||||
startup log (`ws://localhost:{port}/_translate/{key}` locally,
|
||||
`wss://{hostname}/_translate/{key}` on a public hostname) and shown in the
|
||||
lang tab as click-to-copy links; the keys are also surfaced in
|
||||
`GET /_api/settings` as `translate_keys`. An unknown or empty key rejects
|
||||
the handshake (close-before-accept → HTTP 403). Transactions storing results record the connecting key as the kanta
|
||||
transaction `user`.
|
||||
|
||||
Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
|
||||
@@ -397,9 +402,14 @@ verbatim source substring — entity-decoded text, backslash escapes — is
|
||||
skipped and stays in the original language), and the returned translations
|
||||
are swapped in by offset. Markup corruption is therefore impossible by
|
||||
construction; the failure modes that remain are a wrong segment count, an
|
||||
empty segment, or markup injected INTO a segment (a `<br>` in a title
|
||||
translation would splice live HTML) — each returned segment must parse as
|
||||
pure prose, or the whole result is dropped and logged, and the (lang, key)
|
||||
empty segment, markup injected INTO a segment (a `<br>` in a title
|
||||
translation would splice live HTML), or a line that would start a new
|
||||
block where the segment lands (a ``` or ::: fence line would eat the rest
|
||||
of the block it splices into, closing fence included — segments are
|
||||
inline prose, so `pure_prose` alone cannot see this) — each returned
|
||||
segment must parse as
|
||||
pure prose with no block-starting line or blank line, or the whole result
|
||||
is dropped and logged, and the (lang, key)
|
||||
pair is skipped for the rest of the server run (generation is
|
||||
near-deterministic, so an immediate retry would re-fail; the fragment stays
|
||||
pending and gets another chance on restart or `DELETE /_api/translations`).
|
||||
@@ -450,14 +460,25 @@ stripped before the result goes back.
|
||||
|
||||
The same client-side enforcement covers markup bleed as a CLASS, not per
|
||||
artifact: `<` is the prose/markup boundary on the wire and never appears in
|
||||
a segment in either direction. Source pieces containing `<` are never
|
||||
dispatched (they stay in the original language — segments.py), and the
|
||||
a segment in either direction. A literal `<` in the source text (`<1MB` is
|
||||
text, not markup — a tag needs a letter or `/!?`) crosses encoded as the
|
||||
fullwidth `<` and is decoded on return, before the result is validated and
|
||||
spliced (segments.py) — the wire itself still never carries `<`, and the
|
||||
reference client cuts the model's output at the first `<`
|
||||
(scripts/translator.py) — echoed language tags, stray `<br>`s and any
|
||||
future variant are one handled case. (The cut is post-decode, not a
|
||||
generation stop string: Seed-X opens every generation with its `<s>`
|
||||
framing token, which would trip a `<` stop immediately.)
|
||||
|
||||
Server-side, a second layer covers what the inline parser cannot: ASCII
|
||||
punctuation that is plain prose on the wire but Markdown syntax in the
|
||||
splice context — quotes (a translated `"` would close the quoted image
|
||||
title it lands in), brackets (alt texts, re-inserted link texts), `|` in
|
||||
table rows, `\` escapes. Rather than rejecting such results, `join` swaps
|
||||
them for Unicode look-alikes before splicing (`_NEUTRAL` in
|
||||
segments.py — curly quotes, fullwidth brackets; the renderer's
|
||||
typographer curls straight quotes anyway).
|
||||
|
||||
Short fragments get more than a bare prompt: each segment may carry its
|
||||
surround in `Job.contexts` — a title carries the article's opening prose
|
||||
(its own block is just the title word), a segment carved out of a larger
|
||||
|
||||
@@ -37,3 +37,6 @@ __screenshots__/
|
||||
|
||||
# Playwright browser downloads (if ever installed locally)
|
||||
.pw-browsers/
|
||||
|
||||
# npm project config (audit/fund off: the audit endpoint stalls installs)
|
||||
!.npmrc
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
audit=false
|
||||
fund=false
|
||||
@@ -168,7 +168,8 @@ const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], c
|
||||
</script>
|
||||
|
||||
<template>
|
||||
<div class="analytics-view">
|
||||
<!-- Untranslated admin dashboard: always LTR, like the editor panel. -->
|
||||
<div class="analytics-view" lang="en" dir="ltr">
|
||||
<div class="analytics-panel">
|
||||
<header>
|
||||
<h1>Analytics</h1>
|
||||
|
||||
@@ -1,16 +1,21 @@
|
||||
<script setup>
|
||||
// Lang tab: the site-wide translation target languages (translate_langs)
|
||||
// and the translator service WebSocket URL(s) (translate_keys). ALL
|
||||
// languages are listed, English included — a page whose primary language
|
||||
// (Node.language, configured per row in the structure tab, inherited down
|
||||
// the hierarchy) differs can be translated INTO any other. Flag clicks
|
||||
// toggle and save immediately; the settings round-trip re-reads the
|
||||
// payload, so this tab only ever changes translate_langs. The settings
|
||||
// write's invalidation hook kicks the translation dispatcher. The refresh
|
||||
// button drops all machine translations (user patches are kept), making
|
||||
// the dispatcher re-translate everything.
|
||||
// and the translator service keys (translate_keys) with their WebSocket
|
||||
// URLs. ALL languages are listed, English included — a page whose primary
|
||||
// language (Node.language, configured per row in the structure tab,
|
||||
// inherited down the hierarchy) differs can be translated INTO any other.
|
||||
// Flag clicks toggle and save immediately; the settings round-trip
|
||||
// re-reads the payload, so this tab only ever changes translate_langs. The
|
||||
// settings write's invalidation hook kicks the translation dispatcher. The
|
||||
// refresh button drops all machine translations (user patches are kept),
|
||||
// making the dispatcher re-translate everything. Translator keys are
|
||||
// managed inline (➕ add, name edit, ✕ delete); new keys are generated
|
||||
// here in the server's format and everything rides the settings
|
||||
// round-trip. Clicking a key copies its full URL (following ws:// would
|
||||
// fail).
|
||||
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
|
||||
import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs'
|
||||
import { copyList } from './analytics/format.js'
|
||||
import { dropPageCache } from './swapdoc'
|
||||
|
||||
defineProps({ pagePath: { type: String, default: '' } })
|
||||
@@ -21,6 +26,16 @@ const saveError = ref('')
|
||||
const selected = ref(new Set())
|
||||
const keyUrls = ref([])
|
||||
|
||||
// Full WebSocket URL for a key. New keys are generated right here: 12
|
||||
// lowercase alphanumerics, the server-side format (state._KEY_ALPHABET).
|
||||
const wsUrl = (key) =>
|
||||
`${location.origin.replace(/^http/, 'ws')}/_translate/${key}`
|
||||
const KEY_ALPHABET = 'abcdefghijklmnopqrstuvwxyz0123456789'
|
||||
const newKey = () =>
|
||||
[...crypto.getRandomValues(new Uint8Array(12))]
|
||||
.map((b) => KEY_ALPHABET[b % KEY_ALPHABET.length])
|
||||
.join('')
|
||||
|
||||
// The toggleable targets: every translatable language, laid out in
|
||||
// geographic/cultural groups (one row each) rather than alphabetized —
|
||||
// related languages sit together (a node's own primary is excluded per
|
||||
@@ -52,9 +67,8 @@ onMounted(async () => {
|
||||
try {
|
||||
const s = await (await fetch('/_api/settings')).json()
|
||||
selected.value = new Set(s.translate_langs || [])
|
||||
const wsBase = location.origin.replace(/^http/, 'ws')
|
||||
keyUrls.value = Object.entries(s.translate_keys || {})
|
||||
.map(([key, name]) => ({ name, url: `${wsBase}/_translate/${key}` }))
|
||||
.map(([key, name]) => ({ key, name, url: wsUrl(key) }))
|
||||
} catch { /* keep defaults */ }
|
||||
})
|
||||
|
||||
@@ -100,6 +114,39 @@ async function refresh() {
|
||||
refreshing.value = false
|
||||
}
|
||||
}
|
||||
|
||||
// Key management rides the settings round-trip, like toggle() above:
|
||||
// mutate keyUrls, then PUT the whole settings payload with the new
|
||||
// translate_keys. ➕ adds a fresh unnamed key, names save on every
|
||||
// keystroke (@input — spamming the server is fine), ✕ deletes without
|
||||
// confirmation.
|
||||
async function saveKeys() {
|
||||
try {
|
||||
const s = await (await fetch('/_api/settings')).json()
|
||||
const res = await fetch('/_api/settings', {
|
||||
method: 'PUT',
|
||||
headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
...s,
|
||||
translate_keys: Object.fromEntries(keyUrls.value.map((k) => [k.key, k.name])),
|
||||
}),
|
||||
})
|
||||
saveError.value = res.ok ? '' : '⚠️ changes could not be saved'
|
||||
} catch {
|
||||
saveError.value = '⚠️ changes could not be saved'
|
||||
}
|
||||
}
|
||||
|
||||
function addKey() {
|
||||
const key = newKey()
|
||||
keyUrls.value.push({ key, name: '', url: wsUrl(key) })
|
||||
saveKeys()
|
||||
}
|
||||
|
||||
function removeKey(k) {
|
||||
keyUrls.value = keyUrls.value.filter((x) => x.key !== k.key)
|
||||
saveKeys()
|
||||
}
|
||||
</script>
|
||||
|
||||
<template>
|
||||
@@ -129,28 +176,37 @@ async function refresh() {
|
||||
|
||||
<section class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">translations</span>
|
||||
<small class="muted">deleting re-translates everything; user edits are kept</small>
|
||||
<span class="field-label">Translator API</span>
|
||||
</div>
|
||||
<button
|
||||
type="button"
|
||||
class="refresh-btn"
|
||||
:disabled="refreshing"
|
||||
title="delete all machine translations and let the translator re-fill them"
|
||||
@click="refresh"
|
||||
>
|
||||
{{ refreshing ? 'refreshing…' : 'refresh all translations' }}
|
||||
</button>
|
||||
</section>
|
||||
|
||||
<section v-if="keyUrls.length" class="block">
|
||||
<div class="block-head">
|
||||
<span class="field-label">translator service</span>
|
||||
<small class="muted">connect scripts/translator.py to</small>
|
||||
<div v-for="k in keyUrls" :key="k.key" class="key-row">
|
||||
<a
|
||||
:href="k.url"
|
||||
class="key-link"
|
||||
title="click to copy the URL"
|
||||
@click.prevent="copyList(k.url, $event)"
|
||||
>{{ k.key }}</a>
|
||||
<input
|
||||
v-model="k.name"
|
||||
type="text"
|
||||
class="edit key-name"
|
||||
title="display name"
|
||||
@input="saveKeys()"
|
||||
>
|
||||
<button type="button" class="act del" title="delete key" @click="removeKey(k)">✕</button>
|
||||
</div>
|
||||
<div v-for="k in keyUrls" :key="k.url" class="key-row">
|
||||
<code>{{ k.url }}</code>
|
||||
<small class="muted">{{ k.name }}</small>
|
||||
<div class="add-row">
|
||||
<button type="button" class="add" title="new translator key" @click="addKey()">➕ API key</button>
|
||||
</div>
|
||||
<p><small class="muted">AI translator agents can connect with the API keys to do machine translations to your selected languages. Click the button below to delete all translations and start over. User edits are kept.</small></p>
|
||||
<div class="refresh-row">
|
||||
<button
|
||||
type="button"
|
||||
class="refresh-btn"
|
||||
:disabled="refreshing"
|
||||
@click="refresh"
|
||||
>
|
||||
{{ refreshing ? 'Reseting…' : 'Reset' }}
|
||||
</button>
|
||||
</div>
|
||||
</section>
|
||||
</div>
|
||||
@@ -252,10 +308,83 @@ async function refresh() {
|
||||
gap: 0.6rem;
|
||||
}
|
||||
|
||||
.key-row code {
|
||||
/* Real links (handy for right-click/drag) showing just the key, but the
|
||||
click copies the full URL instead of following — ws:// would fail to
|
||||
navigate. Normal text color, not link-styled; position: relative
|
||||
anchors the "Copied!" popup (analytics/format.js). */
|
||||
.key-link {
|
||||
position: relative;
|
||||
color: var(--text);
|
||||
font-family: var(--font-code);
|
||||
user-select: all;
|
||||
}
|
||||
|
||||
.refresh-row {
|
||||
display: flex;
|
||||
align-items: baseline;
|
||||
gap: 0.6rem;
|
||||
}
|
||||
|
||||
/* Name input / ✕ / ➕ follow the structure tab's conventions: inputs stay
|
||||
borderless until interacted with, glyph buttons redden / solidify on
|
||||
hover. */
|
||||
.key-name {
|
||||
flex: 0 0 9rem;
|
||||
}
|
||||
|
||||
.edit {
|
||||
font: inherit;
|
||||
font-size: 0.85rem;
|
||||
padding: 0.1rem 0.4rem;
|
||||
background: transparent;
|
||||
color: var(--text);
|
||||
border: 1px solid transparent;
|
||||
border-radius: 4px;
|
||||
min-width: 0;
|
||||
}
|
||||
|
||||
.edit:hover {
|
||||
border-color: var(--line);
|
||||
}
|
||||
|
||||
.edit:focus {
|
||||
background: var(--bg);
|
||||
border-color: var(--accent);
|
||||
outline: none;
|
||||
}
|
||||
|
||||
.act {
|
||||
padding: 0 0.25rem;
|
||||
background: none;
|
||||
border: none;
|
||||
color: var(--muted);
|
||||
font-size: 0.8rem;
|
||||
cursor: pointer;
|
||||
white-space: nowrap;
|
||||
}
|
||||
|
||||
.del:hover {
|
||||
color: #e06c75;
|
||||
}
|
||||
|
||||
.add-row {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
}
|
||||
|
||||
.add {
|
||||
padding: 0 0.3rem;
|
||||
background: none;
|
||||
border: none;
|
||||
font-size: 0.9rem;
|
||||
cursor: pointer;
|
||||
opacity: 0.5;
|
||||
}
|
||||
|
||||
.add:hover {
|
||||
opacity: 1;
|
||||
}
|
||||
|
||||
.refresh-btn {
|
||||
align-self: flex-start;
|
||||
margin-bottom: 0.2rem;
|
||||
|
||||
@@ -434,7 +434,9 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
|
||||
|
||||
/**
|
||||
* Group abuse hits by IP and format each group as a row with the full paths
|
||||
* probed. Identical paths are collapsed into one entry with their hit count.
|
||||
* probed. Identical requests (same path and status class) are collapsed
|
||||
* into one entry with their hit count; a path's 404 probes and its real
|
||||
* (200) reads never merge.
|
||||
* The paths split into two lists: ``paths`` holds the 404 probes (flagged
|
||||
* paths — the ones that triggered abuse classification — first, then other
|
||||
* 404s) shown verbatim, query string included, and ``articles`` holds the
|
||||
@@ -465,7 +467,11 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
|
||||
g.lastClient = a.client
|
||||
}
|
||||
const path = a.path || ''
|
||||
const existing = g.pathCounts.get(path) || {
|
||||
// Collapse identical requests, but never merge a path's 404 probes with
|
||||
// its real (200) reads — a page probed while missing and later created
|
||||
// must show up in both columns, not flip to "articles read".
|
||||
const key = `${a.is_404 ? '4' : '2'}${path}`
|
||||
const existing = g.pathCounts.get(key) || {
|
||||
path,
|
||||
count: 0,
|
||||
firstStart: start,
|
||||
@@ -475,8 +481,7 @@ export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
|
||||
existing.count += 1
|
||||
if (start < existing.firstStart) existing.firstStart = start
|
||||
if (a.flag) existing.flag = true
|
||||
if (!a.is_404) existing.is_404 = false
|
||||
g.pathCounts.set(path, existing)
|
||||
g.pathCounts.set(key, existing)
|
||||
g.clientHashes.add(a.client)
|
||||
groups.set(ip, g)
|
||||
}
|
||||
|
||||
@@ -288,8 +288,10 @@ body {
|
||||
|
||||
/* Sidebar + main row. A symmetric grid: the article column is sized by the
|
||||
viewport alone (never by content), with equally sized flexible gutters
|
||||
on both sides. The sidebar sits in the left gutter, so it appearing or
|
||||
disappearing never shifts the article; the right gutter balances it.
|
||||
on both sides. The sidebar sits in the start gutter (grid columns are
|
||||
flow-relative: on RTL pages the whole composition mirrors), so it
|
||||
appearing or disappearing never shifts the article; the end gutter
|
||||
balances it.
|
||||
The outer tracks are minmax(0, 1fr) — a plain 1fr has an `auto` minimum,
|
||||
which let the 12rem sidebar expand its track at narrow widths and push
|
||||
the article off-center; now the sidebar overlays the gutter edge instead
|
||||
@@ -305,16 +307,16 @@ body {
|
||||
content length — code excluded) lift the 78rem cap: main takes the full
|
||||
width and the article composes itself inside it — fluid, bounded text
|
||||
lanes with the surplus left vacant (see the article layout rules
|
||||
below). The left track — main always sits in column 2 — collapses to
|
||||
below). The start track — main always sits in column 2 — collapses to
|
||||
zero when the page has no sidebar; the sidebar then simply overlays
|
||||
the vacant zone, as it does on single-column pages. */
|
||||
body:has(.multicol) #content {
|
||||
grid-template-columns: 0 minmax(0, 1fr);
|
||||
}
|
||||
|
||||
/* With a sidebar the left lane gets its own track at every width, so the
|
||||
/* With a sidebar the start lane gets its own track at every width, so the
|
||||
sidebar never overlaps the article and the article leans on the
|
||||
viewport's right edge (surplus extends the lane). The lane is flexible:
|
||||
viewport's end edge (surplus extends the lane). The lane is flexible:
|
||||
12rem when space is tight, growing up to 150% (18rem) once the viewport
|
||||
exceeds the article's 88.5rem (86rem + main's side padding). The --lane
|
||||
variable doubles as the measure for the margin boxes and the .wide
|
||||
@@ -398,7 +400,7 @@ body.editing #sidebar {
|
||||
|
||||
#sidebar {
|
||||
grid-column: 1;
|
||||
/* Pinned to the page's left edge (not the article's) and kept in view
|
||||
/* Pinned to the page's start edge (not the article's) and kept in view
|
||||
while scrolling. Translucent + blurred rather than an opaque box, so
|
||||
full-bleed .wide images can pass underneath without a hard edge. */
|
||||
justify-self: start;
|
||||
@@ -409,8 +411,11 @@ body.editing #sidebar {
|
||||
width: 12rem;
|
||||
max-height: 100vh;
|
||||
overflow-y: auto;
|
||||
padding: 1rem 1rem 1rem 1.25rem;
|
||||
border-radius: 0 0 0.5rem 0;
|
||||
/* The extra inline-start padding (the viewport-edge side, mirroring with
|
||||
the direction) matches main's side padding. */
|
||||
padding: 1rem;
|
||||
padding-inline: 1.25rem 1rem;
|
||||
border-end-start-radius: 0.5rem;
|
||||
background: color-mix(var(--bg) 75%, transparent);
|
||||
backdrop-filter: blur(0.5rem);
|
||||
}
|
||||
@@ -437,7 +442,7 @@ body.editing #sidebar {
|
||||
#sidebar ul ul li::before {
|
||||
content: "🔹";
|
||||
display: inline-block;
|
||||
margin-left: -1.3em;
|
||||
margin-inline-start: -1.3em;
|
||||
width: 1.3em;
|
||||
}
|
||||
|
||||
@@ -578,7 +583,7 @@ main {
|
||||
background: color-mix(var(--bg) 30%, transparent);
|
||||
color: var(--muted);
|
||||
font-size: 0.95rem;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
hyphens: none;
|
||||
}
|
||||
|
||||
@@ -679,19 +684,19 @@ article * + h6 {
|
||||
so the gap becomes the figure's margin plus the full --list-indent. */
|
||||
article :is(ul, ol:not([type])) {
|
||||
list-style: none;
|
||||
padding-left: 0;
|
||||
padding-inline-start: 0;
|
||||
--list-indent: 2em;
|
||||
display: flow-root;
|
||||
}
|
||||
|
||||
article :is(ul, ol:not([type])) > li {
|
||||
padding-left: var(--list-indent);
|
||||
padding-inline-start: var(--list-indent);
|
||||
}
|
||||
|
||||
article ul li::before {
|
||||
content: "🔹";
|
||||
display: inline-block;
|
||||
margin-left: calc(-1 * var(--list-indent));
|
||||
margin-inline-start: calc(-1 * var(--list-indent));
|
||||
width: var(--list-indent);
|
||||
text-align: center;
|
||||
}
|
||||
@@ -713,13 +718,13 @@ article ol:not([type]) > li::before {
|
||||
content: counter(item) ".";
|
||||
color: var(--muted);
|
||||
display: inline-block;
|
||||
margin-left: calc(-1 * var(--list-indent));
|
||||
margin-inline-start: calc(-1 * var(--list-indent));
|
||||
width: var(--list-indent);
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
}
|
||||
|
||||
/* Task lists: real clickable checkboxes. The checkbox stands in for the
|
||||
list marker — taken out of flow, left-aligned in the indent box and
|
||||
list marker — taken out of flow, start-aligned in the indent box and
|
||||
centered on the first line's middle, so the item text starts at the
|
||||
same edge as every other list item's. */
|
||||
article .task-list-item {
|
||||
@@ -732,7 +737,7 @@ article .task-list-item::before {
|
||||
|
||||
article .task-list-item-checkbox {
|
||||
position: absolute;
|
||||
left: 0;
|
||||
inset-inline-start: 0;
|
||||
top: calc(0.5lh - .15ex);
|
||||
translate: 0 -50%;
|
||||
font-size: inherit;
|
||||
@@ -772,7 +777,7 @@ article h1 .edit-link {
|
||||
position: static;
|
||||
font-size: 1.1rem;
|
||||
vertical-align: 0.3em;
|
||||
margin-left: 0.4rem;
|
||||
margin-inline-start: 0.4rem;
|
||||
}
|
||||
|
||||
/* Section pens sit at the end of anchored h2s, dimmer than the page pen
|
||||
@@ -781,7 +786,7 @@ article h2 .edit-section {
|
||||
position: static;
|
||||
font-size: 0.85rem;
|
||||
vertical-align: 0.35em;
|
||||
margin-left: 0.4rem;
|
||||
margin-inline-start: 0.4rem;
|
||||
opacity: 0.35;
|
||||
}
|
||||
|
||||
@@ -811,10 +816,10 @@ article dd {
|
||||
}
|
||||
|
||||
/* The long-article composition (.multicol): a fluid but bounded text
|
||||
lane with a 16rem side zone at the article's left, centered in main —
|
||||
surplus width becomes vacant space, never endless text (with a sidebar
|
||||
the article leans right instead and the sidebar's track is the left
|
||||
lane; see below). (The backend render splits the body into .colseg
|
||||
lane with a 16rem side zone at the article's start side, centered in
|
||||
main — surplus width becomes vacant space, never endless text (with a
|
||||
sidebar the article leans to the end edge instead and the sidebar's
|
||||
track is the start lane; see below). (The backend render splits the body into .colseg
|
||||
segments separated by full-width h2s and .wide elements, tags
|
||||
text-heavy segments of several paragraphs .cols — a ::: nocols
|
||||
container opts its section out, and column-filling paragraphs are
|
||||
@@ -832,7 +837,7 @@ article.multicol {
|
||||
@container (min-width: 45rem) {
|
||||
/* The side zone (not on phones): lane content indents 16rem; margin
|
||||
boxes ({.margin} / ::: margin blocks, ::: aside, {.margin} figures)
|
||||
are taken out of flow and placed against the article's left edge —
|
||||
are taken out of flow and placed against the article's start edge —
|
||||
the same region the nav sidebar overlays. The boxes stay in the
|
||||
column segment at their anchor point (the backend render no longer
|
||||
splits segments around them); absolute positioning off the article —
|
||||
@@ -842,14 +847,14 @@ article.multicol {
|
||||
article.multicol>.colseg,
|
||||
article.multicol>h1,
|
||||
article.multicol>h2 {
|
||||
margin-left: 16rem;
|
||||
margin-inline-start: 16rem;
|
||||
}
|
||||
|
||||
article.multicol .margin,
|
||||
article.multicol .aside,
|
||||
article.multicol figure.margin {
|
||||
position: absolute;
|
||||
left: 0;
|
||||
inset-inline-start: 0;
|
||||
width: 14rem;
|
||||
max-width: none;
|
||||
margin: 0.3rem 0 0;
|
||||
@@ -870,11 +875,11 @@ article.multicol {
|
||||
}
|
||||
}
|
||||
|
||||
/* With a sidebar, the sidebar's 12rem track IS the left lane at every
|
||||
/* With a sidebar, the sidebar's 12rem track IS the start lane at every
|
||||
width (see #content): no in-article zone, the text lane runs fluid (up
|
||||
to 86rem) and leans on main's right edge — surplus width extends the
|
||||
left lane instead of balancing out on the right — and margin boxes
|
||||
hang into the lane off the article's left border, sliding under the
|
||||
to 86rem) and leans on main's end edge — surplus width extends the
|
||||
start lane instead of balancing out at the end — and margin boxes
|
||||
hang into the lane off the article's start border, sliding under the
|
||||
translucent sticky nav, which only ever occupies its top. (Not below
|
||||
48rem: there the sidebar becomes a link strip above the article and
|
||||
there is no lane to fall into.) */
|
||||
@@ -884,8 +889,8 @@ article.multicol {
|
||||
margin-inline: auto 0;
|
||||
}
|
||||
|
||||
/* The sidebar fills the flexible lane (its left side stays on the
|
||||
viewport's left edge, growing rightward). */
|
||||
/* The sidebar fills the flexible lane (its start side stays on the
|
||||
viewport's start edge, growing toward the article). */
|
||||
body:has(#sidebar):has(.multicol):not(.editing) #sidebar {
|
||||
width: 100%;
|
||||
}
|
||||
@@ -893,29 +898,29 @@ article.multicol {
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>.colseg,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h1,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol>h2 {
|
||||
margin-left: 0;
|
||||
margin-inline-start: 0;
|
||||
}
|
||||
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .margin,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol .aside,
|
||||
body:has(#sidebar):has(.multicol):not(.editing) article.multicol figure.margin {
|
||||
position: absolute;
|
||||
/* Attached to the article's left border (1.25rem gap), hanging into
|
||||
the left lane and growing leftward with it: 12rem when the lane is
|
||||
/* Attached to the article's start border (1.25rem gap), hanging into
|
||||
the start lane and growing with it: 12rem when the lane is
|
||||
tight, up to 150% (18rem) when the track or the surplus has room
|
||||
(100cqw - 100% is the surplus left of the right-leaning article).
|
||||
(100cqw - 100% is the surplus beside the end-leaning article).
|
||||
The lane (track + main's padding) always guarantees the room. */
|
||||
--box-w: min(18rem, var(--lane) + 100cqw - 100% - 1.25rem);
|
||||
width: var(--box-w);
|
||||
max-width: none;
|
||||
left: calc(-1.25rem - var(--box-w));
|
||||
inset-inline-start: calc(-1.25rem - var(--box-w));
|
||||
margin: 0.3rem 0 0;
|
||||
}
|
||||
}
|
||||
|
||||
/* A shrink-wrapped figure (explicit image width) centers in the plain
|
||||
layout; inside a column the centering looks adrift — left-align.
|
||||
Floated figures keep their own margins (the text gap). */
|
||||
layout; inside a column the centering looks adrift — align to the start
|
||||
edge. Floated figures keep their own margins (the text gap). */
|
||||
.multicol .colseg.cols figure:has(img[width]):not(:has(.left), :has(.right), .margin) {
|
||||
margin-inline: 0;
|
||||
}
|
||||
@@ -988,13 +993,15 @@ article a:hover {
|
||||
|
||||
/* Blockquotes: spacing comes from the blockquote itself (bottom-only like
|
||||
everything else in articles); inner paragraphs keep only the gap between
|
||||
them. The negative left margin pushes the bar out past the text edge, so
|
||||
them. The negative start margin pushes the bar out past the text edge, so
|
||||
quoted text aligns with the surrounding paragraphs — same trick as code
|
||||
blocks. */
|
||||
blockquote {
|
||||
margin: 0 0 1rem -0.5rem;
|
||||
padding: 0 0 0 0.25rem;
|
||||
border-left: 0.25rem solid var(--accent2);
|
||||
margin: 0 0 1rem;
|
||||
margin-inline-start: -0.5rem;
|
||||
padding: 0;
|
||||
padding-inline-start: 0.25rem;
|
||||
border-inline-start: 0.25rem solid var(--accent2);
|
||||
color: var(--muted);
|
||||
}
|
||||
|
||||
@@ -1009,17 +1016,19 @@ blockquote p + p {
|
||||
/* Admonitions (markdown !!! note/warning/...) and GitHub-style alerts
|
||||
(> [!NOTE] ...): a lightweight callout in the blockquote idiom — accent
|
||||
bar and a faint wash, recolored per type, with a type emoji on the
|
||||
title. The negative left margin pushes bar and wash out past the text
|
||||
title. The negative start margin pushes bar and wash out past the text
|
||||
edge so the inner text aligns with surrounding paragraphs — same trick
|
||||
as blockquotes and code blocks (margin-left = border + padding-left).
|
||||
Bottom-only margins like everything else in articles; inner paragraphs
|
||||
carry no margins of their own. */
|
||||
as blockquotes and code blocks (margin-inline-start = border +
|
||||
padding-inline-start). Bottom-only margins like everything else in
|
||||
articles; inner paragraphs carry no margins of their own. */
|
||||
.admonition,
|
||||
.markdown-alert {
|
||||
margin: 0 0 1rem -1.15rem;
|
||||
margin: 0 0 1rem;
|
||||
margin-inline-start: -1.15rem;
|
||||
padding: 0.4rem 0.9rem;
|
||||
border-left: 0.25rem solid var(--admonition-color, var(--accent));
|
||||
border-radius: 0 0.3rem 0.3rem 0;
|
||||
border-inline-start: 0.25rem solid var(--admonition-color, var(--accent));
|
||||
border-start-end-radius: 0.3rem;
|
||||
border-end-end-radius: 0.3rem;
|
||||
background: color-mix(var(--admonition-color, var(--accent)) 7%, transparent);
|
||||
}
|
||||
|
||||
@@ -1037,7 +1046,7 @@ blockquote p + p {
|
||||
|
||||
.admonition-title::before,
|
||||
.markdown-alert-title::before {
|
||||
padding-right: 0.35em;
|
||||
padding-inline-end: 0.35em;
|
||||
}
|
||||
|
||||
.admonition.note .admonition-title::before,
|
||||
@@ -1099,18 +1108,19 @@ blockquote p + p {
|
||||
Where the layout has room for a side zone — multicol pages, the
|
||||
sidebar's track, the wide single-column gutter (see the article
|
||||
section and the figure rules below) — the boxes are taken out of flow
|
||||
and absolutely positioned into it, off the article's left border, each
|
||||
and absolutely positioned into it, off the article's start border, each
|
||||
at the vertical spot where it occurs in the text (boxes occurring
|
||||
closer together than their heights may overlap — keep them apart);
|
||||
otherwise they stay in-column left floats (consecutive floats stack
|
||||
via clear: left). Headings already clear floats, so in-column boxes
|
||||
never bleed into the next section. */
|
||||
otherwise they stay in-column start floats (consecutive floats stack
|
||||
via clear: inline-start). Headings already clear floats, so in-column
|
||||
boxes never bleed into the next section. */
|
||||
.aside {
|
||||
float: left;
|
||||
clear: left;
|
||||
float: inline-start;
|
||||
clear: inline-start;
|
||||
width: 30%;
|
||||
max-width: 20rem;
|
||||
margin: 0.3rem 1.2rem 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1.2rem;
|
||||
padding: 0.6rem 0.9rem;
|
||||
font-size: 0.9rem;
|
||||
color: var(--muted);
|
||||
@@ -1140,11 +1150,12 @@ blockquote p + p {
|
||||
}
|
||||
|
||||
.margin {
|
||||
float: left;
|
||||
clear: left;
|
||||
float: inline-start;
|
||||
clear: inline-start;
|
||||
width: 30%;
|
||||
max-width: 20rem;
|
||||
margin: 0.3rem 1.2rem 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1.2rem;
|
||||
font-size: 0.9rem;
|
||||
color: var(--muted);
|
||||
}
|
||||
@@ -1153,10 +1164,10 @@ pre {
|
||||
overflow-x: auto;
|
||||
padding: 0.5rem 0.8rem;
|
||||
/* Code text aligns with the surrounding paragraphs: the box extends
|
||||
past them by its own padding. Themes that add a left border must
|
||||
extend margin-left by the border width to keep this alignment. */
|
||||
margin-left: -0.8rem;
|
||||
margin-right: -0.8rem;
|
||||
past them by its own padding. Themes that add a leading border must
|
||||
extend margin-inline-start by the border width to keep this
|
||||
alignment. */
|
||||
margin-inline: -0.8rem;
|
||||
background: var(--code-bg);
|
||||
border-radius: 4px;
|
||||
position: relative;
|
||||
@@ -1196,7 +1207,7 @@ p code {
|
||||
}
|
||||
|
||||
code:not(pre code):first-child {
|
||||
padding-left: 0;
|
||||
padding-inline-start: 0;
|
||||
}
|
||||
|
||||
/* Click-to-copy button (added by pagerite.js) */
|
||||
@@ -1241,7 +1252,7 @@ td {
|
||||
}
|
||||
|
||||
th {
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
background: linear-gradient(180deg,
|
||||
var(--table-head-a, color-mix(var(--accent) 10%, var(--surface))),
|
||||
var(--table-head-b, color-mix(var(--accent) 18%, var(--surface))));
|
||||
@@ -1313,20 +1324,24 @@ figure img:not([width]) {
|
||||
width: 100%;
|
||||
}
|
||||
|
||||
/* Floated figures: {.right} / {.left}, defaulting to 30% of the column
|
||||
and capped at half of it. */
|
||||
/* Floated figures: {.right} / {.left} float to the text column's end/start
|
||||
edge — the class names are author-facing and fixed, but the sides follow
|
||||
the text direction (in RTL, .left floats right) — defaulting to 30% of
|
||||
the column and capped at half of it. */
|
||||
figure:has(.right) {
|
||||
float: right;
|
||||
float: inline-end;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 0 1rem 1em;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-start: 1em;
|
||||
}
|
||||
|
||||
figure:has(.left) {
|
||||
float: left;
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* The same floats for other blocks: ::: left / ::: right containers
|
||||
@@ -1334,17 +1349,19 @@ figure:has(.left) {
|
||||
tables all take the class directly ({.right} at the end of a
|
||||
paragraph's last line, a trailing {.left} line after a fence, ...). */
|
||||
:is(div, p, pre, blockquote, table).right {
|
||||
float: right;
|
||||
float: inline-end;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 0 1rem 1em;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-start: 1em;
|
||||
}
|
||||
|
||||
:is(div, p, pre, blockquote, table).left {
|
||||
float: left;
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* An image with an explicit width attribute shrink-wraps instead: the
|
||||
@@ -1356,13 +1373,15 @@ figure:has(img[width]) {
|
||||
}
|
||||
|
||||
/* {.margin} figures (the class moves onto the figure wrapper at render —
|
||||
see markdown.py) float left like {.left} ones — until they fall into
|
||||
the side zone (see the composition rules up in the article section). */
|
||||
see markdown.py) float to the start edge like {.left} ones — until
|
||||
they fall into the side zone (see the composition rules up in the
|
||||
article section). */
|
||||
figure.margin {
|
||||
float: left;
|
||||
float: inline-start;
|
||||
width: 30%;
|
||||
max-width: 50%;
|
||||
margin: 0.3rem 1em 1rem 0;
|
||||
margin: 0.3rem 0 1rem;
|
||||
margin-inline-end: 1em;
|
||||
}
|
||||
|
||||
/* Click-to-enlarge (pagerite.js): article figure images open in a
|
||||
@@ -1439,12 +1458,12 @@ article figure img {
|
||||
}
|
||||
}
|
||||
|
||||
/* Wide single-column pages: margin boxes lean into the vacant left
|
||||
/* Wide single-column pages: margin boxes lean into the vacant start-side
|
||||
gutter instead (below 104rem the gutter cannot hold the box, and while
|
||||
editing the docked panel reshapes the gutters — in both they stay
|
||||
plain floats). Out of flow like on multicol pages: the box hangs off
|
||||
the article's left border, growing with the gutter up to 150% (18rem),
|
||||
its right side 1.25rem off the border. */
|
||||
the article's start border, growing with the gutter up to 150% (18rem),
|
||||
its end side 1.25rem off the border. */
|
||||
@media (min-width: 104rem) {
|
||||
body:not(.editing):not(:has(.multicol)) article .margin,
|
||||
body:not(.editing):not(:has(.multicol)) article .aside,
|
||||
@@ -1453,14 +1472,14 @@ article figure img {
|
||||
--box-w: min(18rem, (100vw - 78rem) / 2 - 1.25rem);
|
||||
width: var(--box-w);
|
||||
max-width: none;
|
||||
left: calc(-1.25rem - var(--box-w));
|
||||
inset-inline-start: calc(-1.25rem - var(--box-w));
|
||||
margin: 0.3rem 0 0;
|
||||
}
|
||||
}
|
||||
|
||||
/* In the wide symmetric gutters (where the sidebar overlays the flexible
|
||||
left gutter rather than a reserved track) the sidebar flexes with the
|
||||
gutter up to 150% — its left side stays on the viewport's edge. */
|
||||
start gutter rather than a reserved track) the sidebar flexes with the
|
||||
gutter up to 150% — its start side stays on the viewport's edge. */
|
||||
@media (min-width: 102rem) {
|
||||
body:not(.editing) #sidebar {
|
||||
width: min(18rem, 100%);
|
||||
@@ -1509,7 +1528,7 @@ body.editing pre.wide {
|
||||
|
||||
/* Narrow single-column pages with a sidebar: below 102rem the symmetric
|
||||
gutters can no longer both hold the sidebar, so #content reserves it
|
||||
with a flexible left track (see the matching media query below) and the
|
||||
with a flexible start track (see the matching media query below) and the
|
||||
article always starts at the lane's width (+ main's 1.25rem padding) —
|
||||
the bleed margin measures off --lane. Scoped by :has(#sidebar) since
|
||||
the sidebar element is omitted entirely on pages without
|
||||
@@ -1536,9 +1555,9 @@ body:has(.multicol) pre.wide {
|
||||
}
|
||||
|
||||
/* Multicol with a sidebar track (≥48rem, see #content): main starts at
|
||||
the flexible lane's width and the article leans right, so the bleed
|
||||
extends left past the surplus and the lane to the true viewport edge —
|
||||
sliding under the translucent sidebar — and right past main's
|
||||
the flexible lane's width and the article leans to the end edge, so the
|
||||
bleed extends past the surplus and the lane to the true viewport start
|
||||
edge — sliding under the translucent sidebar — and past main's end
|
||||
padding. */
|
||||
@media (min-width: 48rem) {
|
||||
body:has(#sidebar):has(.multicol):not(.editing) figure:has(.wide),
|
||||
@@ -1556,7 +1575,7 @@ figcaption {
|
||||
hyphens: auto;
|
||||
-webkit-hyphens: auto;
|
||||
text-wrap: pretty;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
/* Never let a long caption stretch a shrink-to-fit figure wider than the
|
||||
image; the caption wraps at the figure's width instead. */
|
||||
width: 0;
|
||||
@@ -1582,7 +1601,7 @@ article h2 {
|
||||
|
||||
/* Narrow windows with a sidebar: below 102rem the symmetric gutters can no
|
||||
longer both hold the 12rem sidebar, so reserve its space with a flexible
|
||||
left track instead of letting it overlap the article (multicol pages use
|
||||
start track instead of letting it overlap the article (multicol pages use
|
||||
the same track at every width — see the #content rules above; their
|
||||
higher-specificity rule wins there). The lane is 12rem when space is
|
||||
tight, growing up to 150% (18rem) once the viewport exceeds the
|
||||
@@ -1702,10 +1721,10 @@ article h2 {
|
||||
width: fit-content;
|
||||
}
|
||||
|
||||
/* The single-column sidebar .wide margins assume a left sidebar column;
|
||||
with the sidebar on top the article is viewport-wide and the plain
|
||||
centered bleed applies again. (Multicol pages need no override: their
|
||||
cqw bleed is exact at any width.) */
|
||||
/* The single-column sidebar .wide margins assume a start-side sidebar
|
||||
column; with the sidebar on top the article is viewport-wide and the
|
||||
plain centered bleed applies again. (Multicol pages need no override:
|
||||
their cqw bleed is exact at any width.) */
|
||||
body:has(#sidebar):not(.editing):not(:has(.multicol)) figure:has(.wide) {
|
||||
margin-inline: calc(50% - 50vw);
|
||||
}
|
||||
@@ -1746,7 +1765,7 @@ article h2 {
|
||||
.dateline {
|
||||
color: var(--muted);
|
||||
font-size: 0.85rem;
|
||||
text-align: left;
|
||||
text-align: start;
|
||||
}
|
||||
|
||||
.footnotes {
|
||||
|
||||
+15
-77
@@ -1,80 +1,17 @@
|
||||
"""Command-line entry point for running the backend server."""
|
||||
|
||||
import argparse
|
||||
import gzip
|
||||
import os
|
||||
import sys
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
|
||||
import httpx
|
||||
import msgspec
|
||||
from fastapi_vue import server
|
||||
from fastapi_vue.hostutil import parse_endpoints
|
||||
|
||||
from pagerite.config import Config
|
||||
|
||||
DEFAULT_PORT = 8100
|
||||
DEVMODE = os.getenv("PAGERITE_DEV") == "1"
|
||||
|
||||
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
|
||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
|
||||
|
||||
|
||||
def _download_dbip() -> None:
|
||||
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
|
||||
today = date.today()
|
||||
months = [f"{today:%Y-%m}"]
|
||||
# The current month's file may not be published yet; fall back to last month.
|
||||
prev = (today.replace(day=1) - date.resolution).replace(day=1)
|
||||
months.append(f"{prev:%Y-%m}")
|
||||
|
||||
existing = sorted(
|
||||
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
|
||||
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
|
||||
)
|
||||
if existing and existing[-1] >= months[0]:
|
||||
print(
|
||||
f"pagerite: DB-IP database is current ({existing[-1]}), skipping download"
|
||||
)
|
||||
return
|
||||
|
||||
for month in months:
|
||||
url = DBIP_URL.format(month=month)
|
||||
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
|
||||
tmp = target.with_suffix(".mmdb.gz.tmp")
|
||||
print(f"pagerite: downloading {url}")
|
||||
try:
|
||||
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
|
||||
if r.status_code == 404:
|
||||
continue
|
||||
r.raise_for_status()
|
||||
with open(tmp, "wb") as f:
|
||||
for chunk in r.iter_bytes():
|
||||
f.write(chunk)
|
||||
except httpx.HTTPError as e:
|
||||
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
# Verify it is actually gzip data before installing it.
|
||||
try:
|
||||
with gzip.open(tmp, "rb") as f:
|
||||
f.read(1)
|
||||
except OSError:
|
||||
print(
|
||||
f"pagerite: DB-IP download for {month} was not valid gzip",
|
||||
file=sys.stderr,
|
||||
)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
os.replace(tmp, target)
|
||||
# Drop older databases so the app never picks up a stale one.
|
||||
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
|
||||
if old.name != target.name:
|
||||
old.unlink()
|
||||
print(f"pagerite: DB-IP database updated to {target.name}")
|
||||
return
|
||||
print("pagerite: could not download a DB-IP database", file=sys.stderr)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
"""Run the backend server with optional arguments."""
|
||||
@@ -101,23 +38,24 @@ def main() -> None:
|
||||
help="Download/update the DB-IP city lite database before starting.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
# Export the hostname before pagerite.app is imported: it derives the
|
||||
# data directory and public origin from it at import time.
|
||||
os.environ["PAGERITE_HOSTNAME"] = args.hostname
|
||||
# And the listen port: the app prints the translator WS URL at startup,
|
||||
# which for localhost includes the actual port.
|
||||
for endpoint in parse_endpoints(args.listen, DEFAULT_PORT):
|
||||
if "port" in endpoint:
|
||||
os.environ["PAGERITE_PORT"] = str(endpoint["port"])
|
||||
break
|
||||
if args.dbip:
|
||||
_download_dbip()
|
||||
# Hand configuration to the app as JSON in PAGERITE_CONFIG; it must be
|
||||
# set before pagerite.app is imported, as state.py reads it at import
|
||||
# time (data directory, public origin).
|
||||
os.environ["PAGERITE_CONFIG"] = msgspec.json.encode(
|
||||
Config(hostname=args.hostname, dbip=args.dbip)
|
||||
).decode()
|
||||
run_args: dict = {}
|
||||
if args.hostname != "localhost":
|
||||
# A public site sits behind TLS on its hostname; show that URL in the
|
||||
# startup box instead of the local listen address.
|
||||
run_args["startup_box"] = f"{{Name}} {{version}}\nhttps://{args.hostname}"
|
||||
server.run(
|
||||
"pagerite.app:app",
|
||||
listen=args.listen,
|
||||
default_port=DEFAULT_PORT,
|
||||
server_header=False,
|
||||
reload=Path(__file__).parent if DEVMODE else False,
|
||||
**run_args,
|
||||
)
|
||||
|
||||
|
||||
|
||||
+509
-512
File diff suppressed because it is too large
Load Diff
+16
-15
@@ -130,9 +130,7 @@ async def save_page(
|
||||
if i18n.add_patch(data, node, path, lang, page.markdown):
|
||||
_invalidate_pages()
|
||||
return
|
||||
with kanta.transaction(
|
||||
"page", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
|
||||
node = _ensure(data.menu, path)
|
||||
node.title = page.title
|
||||
node.chunks = store_chunks(data.chunks, page.markdown)
|
||||
@@ -242,9 +240,7 @@ async def update_structure(op: StructureOp, request: Request) -> None:
|
||||
if op.title is not None
|
||||
else "structure:reorder"
|
||||
)
|
||||
with kanta.transaction(
|
||||
action, user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
with kanta.transaction(action, user=request.headers.get("remote-user"), extra=path):
|
||||
if op.title is not None:
|
||||
node.title = op.title
|
||||
if target is not None and target != path:
|
||||
@@ -302,14 +298,13 @@ class SettingsIn(BaseModel):
|
||||
brand_html: str = ""
|
||||
transition: str = "cube"
|
||||
translate_langs: list[str] | None = None # None keeps the current set
|
||||
translate_keys: dict[str, str] | None = None # None keeps the current keys
|
||||
|
||||
|
||||
@router.put("/_api/settings", status_code=204)
|
||||
async def put_settings(settings: SettingsIn, request: Request) -> None:
|
||||
"""Update site-wide settings; invalidates cached pages and ETags."""
|
||||
with kanta.transaction(
|
||||
"settings", user=request.headers.get("remote-user")
|
||||
):
|
||||
with kanta.transaction("settings", user=request.headers.get("remote-user")):
|
||||
data.brand = settings.brand
|
||||
data.brand_html = settings.brand_html
|
||||
data.theme = settings.theme
|
||||
@@ -324,6 +319,8 @@ async def put_settings(settings: SettingsIn, request: Request) -> None:
|
||||
for lang in settings.translate_langs
|
||||
if (tag := i18n.base_tag(lang))
|
||||
}
|
||||
if settings.translate_keys is not None:
|
||||
data.translate_keys = settings.translate_keys
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@@ -335,9 +332,7 @@ async def delete_translations(request: Request) -> None:
|
||||
translators). User patches are kept; the availability index
|
||||
(node.langs) is rebuilt from them — patches alone still make a language
|
||||
exist on a page."""
|
||||
with kanta.transaction(
|
||||
"translate:reset", user=request.headers.get("remote-user")
|
||||
):
|
||||
with kanta.transaction("translate:reset", user=request.headers.get("remote-user")):
|
||||
i18n.clear_translations(data)
|
||||
_invalidate_pages()
|
||||
# Fragments rejected this run (segment validation) stay skipped no
|
||||
@@ -376,9 +371,7 @@ async def toggle_task_endpoint(body: ToggleTaskIn, request: Request) -> dict[str
|
||||
new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
|
||||
if new_markdown is None:
|
||||
raise HTTPException(400, "invalid task index")
|
||||
with kanta.transaction(
|
||||
"page", user=request.headers.get("remote-user"), extra=path
|
||||
):
|
||||
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
|
||||
# Re-chunk like any save: only the chunk containing the toggled
|
||||
# checkbox gets a new hash, the rest keep theirs.
|
||||
node.chunks = store_chunks(data.chunks, new_markdown)
|
||||
@@ -510,6 +503,14 @@ async def editor_ws(ws: WebSocket) -> None:
|
||||
# The title is injected as h1 when the markdown has
|
||||
# none; the editor's title field edits live-preview.
|
||||
title=msg.get("title") or (node.title if node else ""),
|
||||
# Pin section anchors to the original language so the
|
||||
# preview of a translation matches the served page
|
||||
# (no-op when the previewed markdown is the original).
|
||||
anchors_from=(
|
||||
(node_markdown(data, node) or "", node.title)
|
||||
if node
|
||||
else None
|
||||
),
|
||||
)
|
||||
await ws.send_json(
|
||||
{
|
||||
|
||||
+42
-4
@@ -30,18 +30,51 @@ walking the tree (``resolve``), moves are slot detach/attach
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from collections.abc import AsyncGenerator
|
||||
from contextlib import asynccontextmanager
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi import FastAPI, Request
|
||||
from fastapi.responses import Response
|
||||
from fastapi_vue import Frontend
|
||||
from starlette.types import ASGIApp, Receive, Scope, Send
|
||||
|
||||
from pagerite import api, files, pages, tracking
|
||||
from pagerite.__main__ import DEVMODE
|
||||
from pagerite.files import file_store
|
||||
from pagerite.state import analytics_store, frontend, kanta
|
||||
from collections.abc import AsyncGenerator
|
||||
from pagerite.state import analytics_store, config, kanta
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Vue build served at the site root, no SPA catch-all (assets only). The
|
||||
# build mirrors the URL space: hashed, immutable files live under
|
||||
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico.
|
||||
frontend = Frontend(
|
||||
Path(__file__).with_name("frontend-build"), spa=False, cached="/_assets/"
|
||||
)
|
||||
|
||||
|
||||
class _AccessLogExtraMiddleware:
|
||||
"""Fill the ``log_extra`` slot of fastapi_vue's access log.
|
||||
|
||||
Everything under ``/_api`` is gated by the SSO forward-auth, which names
|
||||
the authenticated user in the ``remote-user`` header; put that user on
|
||||
the access-log line, for plain requests and WebSocket open/close alike.
|
||||
The scope dict is shared with the outer AccessLogMiddleware, which reads
|
||||
the slot back at response/accept/close time.
|
||||
"""
|
||||
|
||||
def __init__(self, app: ASGIApp) -> None:
|
||||
self.app = app
|
||||
|
||||
async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
|
||||
if scope["type"] in ("http", "websocket") and scope["path"].startswith("/_api"):
|
||||
headers = dict(scope["headers"])
|
||||
user = headers.get(b"remote-user", b"").decode("latin-1")
|
||||
if user:
|
||||
scope.setdefault("state", {})["log_extra"] = user
|
||||
await self.app(scope, receive, send)
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(_app: FastAPI) -> AsyncGenerator:
|
||||
@@ -49,8 +82,11 @@ async def lifespan(_app: FastAPI) -> AsyncGenerator:
|
||||
async with kanta:
|
||||
await asyncio.to_thread(file_store.load)
|
||||
await frontend.load()
|
||||
# Decompress/open the DB-IP MMDB once at startup. Lookups are then
|
||||
# read-only and safe to run in background ``to_thread`` workers.
|
||||
# --dbip: update the DB-IP database first, then decompress/open the
|
||||
# MMDB once. Lookups are then read-only and safe to run in
|
||||
# background ``to_thread`` workers.
|
||||
if config.dbip:
|
||||
await asyncio.to_thread(tracking._download_dbip)
|
||||
await asyncio.to_thread(tracking._geoip._load)
|
||||
analytics_store.subscribe(tracking._schedule_analytics_broadcast)
|
||||
# Backfill favicons for external sites already in the recorded data.
|
||||
@@ -70,6 +106,8 @@ app = FastAPI(
|
||||
openapi_url=None,
|
||||
)
|
||||
|
||||
app.add_middleware(_AccessLogExtraMiddleware)
|
||||
|
||||
|
||||
@app.middleware("http")
|
||||
async def _headers(request: Request, call_next) -> Response:
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
"""CLI → app configuration, passed as JSON in the ``PAGERITE_CONFIG`` env var.
|
||||
|
||||
Kept dependency-free (msgspec only) so ``__main__`` can build and serialize
|
||||
the config before any app module is imported, and the app side parses the
|
||||
same struct back. Import-time safe: nothing here reads the environment
|
||||
until ``load()`` is called.
|
||||
"""
|
||||
|
||||
import os
|
||||
|
||||
import msgspec
|
||||
|
||||
|
||||
class Config(msgspec.Struct):
|
||||
"""Configuration passed from the CLI entry point to the app."""
|
||||
|
||||
#: Public hostname of the site; names the per-site data directory
|
||||
#: ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
|
||||
hostname: str = "localhost"
|
||||
#: Download/update the DB-IP city lite database at startup (--dbip).
|
||||
dbip: bool = False
|
||||
|
||||
|
||||
def load() -> Config:
|
||||
"""Parse ``PAGERITE_CONFIG``, or the defaults when unset."""
|
||||
if raw := os.getenv("PAGERITE_CONFIG"):
|
||||
return msgspec.json.decode(raw.encode(), type=Config)
|
||||
return Config()
|
||||
+2
-2
@@ -104,8 +104,8 @@ class Data(msgspec.Struct):
|
||||
#: API keys gating the translator service WebSocket (/_translate/{key};
|
||||
#: the external forward-auth does not cover that route): key -> display
|
||||
#: name. Keys are 12 lowercase alphanumeric characters; the first is
|
||||
#: generated at database bootstrap (see app.py), multiple keys are a
|
||||
#: future reservation (e.g. managed via a web interface).
|
||||
#: generated at database bootstrap, more are managed in the editor
|
||||
#: shell's lang tab (via /_api/settings).
|
||||
translate_keys: dict[str, str] = {}
|
||||
#: Wanted target languages for the translator service (presence-keys,
|
||||
#: value always True). The dispatcher offers jobs only in the
|
||||
|
||||
+55
-7
@@ -402,7 +402,9 @@ def _heading_ids(state) -> None:
|
||||
its self-link is ``href=""`` (back to the top of the page). An
|
||||
author-set `{#id}` always wins; auto ids slugify the heading text
|
||||
(python-slugify, mirroring the editor's slugify.js) and dedupe with
|
||||
-2/-3 suffixes per render. Headings that already contain a link are
|
||||
-2/-3 suffixes per render — unless env["anchor_ids"] presets them, as
|
||||
render(anchors_from=...) does for translated pages so section URLs
|
||||
stay in the original language. Headings that already contain a link are
|
||||
``data-line`` records the heading's markdown source line (0-based, after
|
||||
undoing the render(title=...) injection offset via ``env``) — the page
|
||||
editor uses it for section pens and piecewise-linear scroll sync.
|
||||
@@ -443,15 +445,25 @@ def _heading_ids(state) -> None:
|
||||
if len(heads) < ANCHOR_MIN_HEADINGS:
|
||||
return
|
||||
seen: set[str] = set()
|
||||
for i, token in heads:
|
||||
preset = state.env.get("anchor_ids")
|
||||
for k, (i, token) in enumerate(heads):
|
||||
inline = tokens[i + 1]
|
||||
hid = token.attrGet("id")
|
||||
if not isinstance(hid, str) or not hid:
|
||||
# Slug the visible text, not the raw markdown (`## [a](url)`).
|
||||
text = "".join(
|
||||
c.content for c in inline.children if c.type in ("text", "code_inline")
|
||||
)
|
||||
base = slugify(text) or "section"
|
||||
if preset is not None and k < len(preset):
|
||||
# Translated render: the original language's slug, matched
|
||||
# by heading position (a translation never adds, removes or
|
||||
# reorders headings; a patched one that does falls back to
|
||||
# slugging its own text past the end of the list).
|
||||
base = preset[k]
|
||||
else:
|
||||
# Slug the visible text, not the raw markdown (`## [a](url)`).
|
||||
text = "".join(
|
||||
c.content
|
||||
for c in inline.children
|
||||
if c.type in ("text", "code_inline")
|
||||
)
|
||||
base = slugify(text) or "section"
|
||||
hid, n = base, 2
|
||||
while hid in seen:
|
||||
hid = f"{base}-{n}"
|
||||
@@ -463,6 +475,36 @@ def _heading_ids(state) -> None:
|
||||
wrap(i, token, f"#{hid}")
|
||||
|
||||
|
||||
def anchor_ids(text: str, title: str | None = None) -> list[str]:
|
||||
"""The section anchor ids of text, in heading order.
|
||||
|
||||
render(anchors_from=...) feeds these to _heading_ids via
|
||||
env["anchor_ids"], pinning a translated render's anchors to the
|
||||
original language's slugs. The selection mirrors _heading_ids exactly
|
||||
(the same md instance assigns the ids during this parse, author-set
|
||||
{#id} included as-is); the in-body title h1 is excluded.
|
||||
"""
|
||||
if title and not has_h1(text):
|
||||
text = f"# {title}\n\n{text}"
|
||||
tokens = md.parse(text, {"page_path": ""})
|
||||
first_h1 = next(
|
||||
(
|
||||
i
|
||||
for i, t in enumerate(tokens)
|
||||
if t.type == "heading_open" and t.tag == "h1" and t.level == 0
|
||||
),
|
||||
None,
|
||||
)
|
||||
return [
|
||||
t.attrGet("id")
|
||||
for i, t in enumerate(tokens)
|
||||
if t.type == "heading_open"
|
||||
and t.tag in ("h1", "h2")
|
||||
and t.level == 0
|
||||
and i != first_h1
|
||||
]
|
||||
|
||||
|
||||
def make_md(*, verbatim: bool = False) -> MarkdownIt:
|
||||
"""A fully configured parser. The module-level ``md`` (below) is the
|
||||
render instance; ``verbatim=True`` builds the segmentation instance for
|
||||
@@ -618,12 +660,16 @@ def render(
|
||||
created: datetime | None = None,
|
||||
modified: datetime | None = None,
|
||||
title: str | None = None,
|
||||
anchors_from: tuple[str, str] | None = None,
|
||||
) -> Rendered:
|
||||
"""Render Markdown text to the article body's HTML and layout flags.
|
||||
|
||||
``title`` injects a ``# {title}`` line at the top when the markdown has
|
||||
no h1 of its own, so the implicit page title goes through the exact
|
||||
same pipeline as an explicit one (first-h1 anchor treatment included).
|
||||
``anchors_from`` is the (markdown, title) of the ORIGINAL language when
|
||||
rendering a translation: section anchors are pinned to its slugs so
|
||||
localized pages keep the original #hash URLs.
|
||||
|
||||
The top-level blocks are grouped into column segments: boundary blocks
|
||||
(h1/h2 headings, .wide — see _is_boundary) are rendered bare, the runs
|
||||
@@ -641,6 +687,8 @@ def render(
|
||||
right after the article's h1.
|
||||
"""
|
||||
env = {"page_path": page_path, "line_offset": 0}
|
||||
if anchors_from is not None:
|
||||
env["anchor_ids"] = anchor_ids(*anchors_from)
|
||||
if title and not has_h1(text):
|
||||
text = f"# {title}\n\n{text}"
|
||||
# The injected title shifts source lines by two; _heading_ids
|
||||
|
||||
+10
-34
@@ -3,11 +3,11 @@
|
||||
``GET /{path:path}`` resolves a slug path against the menu tree and renders
|
||||
the page (or a category placeholder, or 404); it must be registered AFTER
|
||||
the fastapi-vue asset routes so built frontend files win over content slugs
|
||||
(see app.py). Requests are recorded in analytics (crawler hits and 404s
|
||||
here, visits via the /_ws socket in tracking.py).
|
||||
(see app.py). Every served document is recorded raw in analytics (one
|
||||
access-log line with its true HTTP status; classification happens at
|
||||
display time — see pagerite/analytics.py).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from email.utils import format_datetime
|
||||
@@ -22,16 +22,9 @@ from pagerite.state import (
|
||||
SITE_URL,
|
||||
_html_response,
|
||||
_is_reserved,
|
||||
analytics_store,
|
||||
data,
|
||||
)
|
||||
from pagerite.tracking import (
|
||||
_client_ip,
|
||||
_enrich_client,
|
||||
_query_suffix,
|
||||
_schedule_client_enrichment,
|
||||
_track_entry,
|
||||
)
|
||||
from pagerite.tracking import _record_get
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -146,20 +139,13 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
placeholder page (nav links point straight at its first child).
|
||||
"""
|
||||
path = path.strip("/")
|
||||
ua = request.headers.get("user-agent", "")
|
||||
accept_language = request.headers.get("accept-language", "")
|
||||
if path and _is_reserved(path):
|
||||
# Invalid slug shape: not a content URL, let FastAPI return its
|
||||
# built-in 404 instead of rendering an editable article page.
|
||||
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
|
||||
# as abuse in analytics.
|
||||
client_hash = analytics_store.track_404(
|
||||
_client_ip(request),
|
||||
ua,
|
||||
f"/{path}{_query_suffix(request)}",
|
||||
accept_language,
|
||||
)
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
# Recorded like any other GET: telltale scanner paths (dotpaths
|
||||
# like /.env, *.php) classify the IP as abuse at display time.
|
||||
_record_get(request, status=404)
|
||||
raise HTTPException(404)
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
@@ -189,8 +175,7 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
if _is_trackable_path(path):
|
||||
flushed = _track_entry(path, request)
|
||||
_schedule_client_enrichment(flushed)
|
||||
_record_get(request)
|
||||
return _html_response(
|
||||
request,
|
||||
"page",
|
||||
@@ -220,8 +205,7 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
)
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
if _is_trackable_path(path):
|
||||
flushed = _track_entry(path, request, status=404)
|
||||
_schedule_client_enrichment(flushed)
|
||||
_record_get(request, status=404)
|
||||
return _html_response(
|
||||
request,
|
||||
"category",
|
||||
@@ -241,13 +225,5 @@ async def show_page(request: Request, path: str) -> Response:
|
||||
if item.published:
|
||||
return RedirectResponse(f"/{slug}")
|
||||
if _is_trackable_path(path):
|
||||
client_hash = analytics_store.track_404(
|
||||
_client_ip(request),
|
||||
ua,
|
||||
f"/{path}{_query_suffix(request)}",
|
||||
accept_language,
|
||||
)
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
flushed = _track_entry(path, request, status=404)
|
||||
_schedule_client_enrichment(flushed)
|
||||
_record_get(request, status=404)
|
||||
return _html_response(request, "not-found", path, 404)
|
||||
|
||||
+100
-22
@@ -20,8 +20,13 @@ each segment's source span was located at dispatch (``split``), and
|
||||
``join`` swaps in the translations. Markup therefore cannot break — it
|
||||
never left the server. A returned segment must still be pure prose itself
|
||||
(the model could inject markup INTO a segment); anything else — count
|
||||
mismatch, empty segment, markup tokens — rejects the whole result and the
|
||||
fragment stays pending.
|
||||
mismatch, empty segment, markup tokens, a line that would start a new
|
||||
block (a ``` or ::: fence would eat the rest of the block it lands in) —
|
||||
rejects the whole result and the
|
||||
fragment stays pending. Punctuation that is prose on the wire but syntax
|
||||
in the splice context (quotes in a title attribute, brackets in an alt
|
||||
text, "|" in a table row) is not worth a rejection either: it is swapped
|
||||
for Unicode look-alikes (``_NEUTRAL``) before splicing.
|
||||
|
||||
A block of plain text, prose links and paired text formatting
|
||||
(strong/em/s) crosses as ONE segment — link texts and formatted text
|
||||
@@ -42,9 +47,11 @@ snippets that don't fit together. Blocks with any other inline markup
|
||||
|
||||
Locating is best effort: a run that is not a verbatim source substring
|
||||
(entity-decoded text, backslash escapes) is skipped — it simply stays in
|
||||
the original language. So is any piece containing "<": "<" is the
|
||||
prose/markup boundary on the wire — translators cut their output there,
|
||||
so such pieces could not survive the round trip.
|
||||
the original language. A literal "<" in prose ("<1MB") is text, not
|
||||
markup, but cannot cross as-is — "<" is the prose/markup boundary on the
|
||||
wire, translators cut their output there — so it crosses encoded as the
|
||||
fullwidth "<" (``_encode``) and ``join`` decodes it back before
|
||||
validating and splicing.
|
||||
"""
|
||||
|
||||
import bisect
|
||||
@@ -69,6 +76,38 @@ _ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
|
||||
#: (inline attrs are consumed by the parser; a lone {dates} is not).
|
||||
_BRACES = re.compile(r"\{[^{}\n]*\}")
|
||||
|
||||
|
||||
def _encode(text: str) -> str:
|
||||
"""Wire form of a segment or context: a literal "<" as fullwidth "<".
|
||||
|
||||
A "<" in prose is text, not markup ("<1MB" — a tag needs a letter or
|
||||
/!?), but "<" is the prose/markup boundary on the wire (translators
|
||||
cut output at the first "<", scripts/translator.py), so it cannot
|
||||
cross as-is. join decodes it back before the pure_prose check and
|
||||
splicing — anything tag-like the model may have formed around it is
|
||||
still rejected there.
|
||||
"""
|
||||
return text.replace("<", "<")
|
||||
|
||||
#: ASCII punctuation that is plain prose to the inline parser (so
|
||||
#: pure_prose cannot catch it) but Markdown SYNTAX in a splice context:
|
||||
#: quotes close a quoted image/link title, brackets the [...] of alt and
|
||||
#: re-inserted link texts, "|" splits a table row, and "\" escapes the
|
||||
#: character after it (a trailing one eats a title's closing quote).
|
||||
#: Neutralized to Unicode look-alikes (join), which Markdown treats as
|
||||
#: plain text everywhere — the quotes are curled the way typographer=True
|
||||
#: renders them anyway.
|
||||
_NEUTRAL = str.maketrans(
|
||||
{
|
||||
'"': "”",
|
||||
"'": "’",
|
||||
"[": "[",
|
||||
"]": "]",
|
||||
"\\": "\",
|
||||
"|": "│",
|
||||
}
|
||||
)
|
||||
|
||||
#: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]",
|
||||
#: "[]" or a bare "]" (shortcut reference); the destination may nest one
|
||||
#: level of parens. Best effort — a mis-scan fails the span-reconstruction
|
||||
@@ -273,7 +312,7 @@ def _linked_block(
|
||||
raw = "".join(text for text, _ in pieces)
|
||||
lead = len(raw) - len(raw.lstrip())
|
||||
wire = raw.strip()
|
||||
if not _LETTER.search(wire) or "<" in wire or _BRACES.search(wire):
|
||||
if not _LETTER.search(wire) or _BRACES.search(wire):
|
||||
return None
|
||||
# Locate each piece verbatim, in order; the source slices between the
|
||||
# located pieces are then the link syntax, exact by construction.
|
||||
@@ -338,7 +377,7 @@ def _linked_block(
|
||||
rec.append(text_)
|
||||
if source[span_start:span_end] != "".join(rec):
|
||||
return None
|
||||
return Span(span_start, span_end, _weight(wire), marks), wire
|
||||
return Span(span_start, span_end, _weight(wire), marks), _encode(wire)
|
||||
|
||||
|
||||
def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
@@ -367,10 +406,8 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
def emit(run: str, at: int, ctx: str) -> None:
|
||||
"""Carve {...} spans out of the located run; emit the prose pieces,
|
||||
stripped — padding whitespace stays in the template, off the wire.
|
||||
Pieces containing "<" are never emitted: translators cut output at
|
||||
the first "<" (the prose/markup boundary, scripts/translator.py),
|
||||
so such a piece could not survive the round trip — it stays in the
|
||||
original language instead."""
|
||||
A literal "<" crosses encoded (``_encode``): it is text, not
|
||||
markup, but the wire keeps "<" as the prose/markup boundary."""
|
||||
pieces = []
|
||||
pos = 0
|
||||
for m in _BRACES.finditer(run):
|
||||
@@ -380,10 +417,10 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
for p0, p1 in pieces:
|
||||
raw = run[p0:p1]
|
||||
piece = raw.strip()
|
||||
if _LETTER.search(piece) and "<" not in piece:
|
||||
if _LETTER.search(piece):
|
||||
start = at + p0 + (len(raw) - len(raw.lstrip()))
|
||||
spans.append(Span(start, start + len(piece), 0, []))
|
||||
segments.append(piece)
|
||||
segments.append(_encode(piece))
|
||||
contexts.append(ctx)
|
||||
|
||||
tokens = _MD.parse(text)
|
||||
@@ -409,7 +446,7 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
cursor = span.end
|
||||
continue
|
||||
runs = _runs(kids)
|
||||
block = _block_text(kids).strip()
|
||||
block = _encode(_block_text(kids).strip())
|
||||
if alert and runs:
|
||||
run = _ALERT.sub("", runs[0], count=1)
|
||||
if _LETTER.search(run):
|
||||
@@ -417,7 +454,7 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
else:
|
||||
runs.pop(0)
|
||||
for run in runs:
|
||||
ctx = block if block and run.strip() != block else ""
|
||||
ctx = block if block and _encode(run.strip()) != block else ""
|
||||
pos = _locate(text, run, cursor)
|
||||
if pos != -1:
|
||||
emit(run, pos, ctx)
|
||||
@@ -435,6 +472,22 @@ def split(text: str) -> tuple[list[Span], list[str], list[str]]:
|
||||
return spans, segments, contexts
|
||||
|
||||
|
||||
#: Block-level Markdown a translation must not introduce: a segment is
|
||||
#: spliced INSIDE a block of the fragment, so a line starting a heading,
|
||||
#: quote, list, code/container fence or a setext/thematic-break underline
|
||||
#: would break the fragment's block structure — a ``` or ::: line eats the
|
||||
#: rest of the fence it lands in, closing fence included. pure_prose only
|
||||
#: parses inline and lets such lines through as softbreak prose, so join
|
||||
#: rejects them here. Blank lines split the host block and are rejected
|
||||
#: too (a faithful translation of a single block has none).
|
||||
_BLOCK = re.compile(
|
||||
r"^[ \t]*(?:#{1,6}(?:[ \t]|$)|>[ \t]?|(?:[-+*]|\d{1,9}[.)])[ \t]|`{3,}|~{3,}|:{3,}(?:[ \t]|$)"
|
||||
r"|-(?:[ \t]*-){2,}[ \t]*$|=[ =]*$|_(?:[ \t]*_){2,}[ \t]*$)",
|
||||
re.M,
|
||||
)
|
||||
_BLANK = re.compile(r"\n[ \t]*\n")
|
||||
|
||||
|
||||
def pure_prose(text: str) -> bool:
|
||||
"""True when the text parses as nothing but prose (text and softbreak
|
||||
tokens) — the acceptance test for a translated segment: the model may
|
||||
@@ -483,7 +536,9 @@ _GAP_S = 0.6
|
||||
_MATCH = 0.3
|
||||
|
||||
|
||||
def _find_mark(src: list[str], units: list[re.Match], start: int) -> tuple[int, int] | None:
|
||||
def _find_mark(
|
||||
src: list[str], units: list[re.Match], start: int
|
||||
) -> tuple[int, int] | None:
|
||||
"""Locate a mark's source words in the translation's units (from unit
|
||||
index ``start`` on), as the (start, end) unit-index span of the best
|
||||
fuzzy alignment; None when no alignment is convincing (the caller falls
|
||||
@@ -508,7 +563,10 @@ def _find_mark(src: list[str], units: list[re.Match], start: int) -> tuple[int,
|
||||
back[i][0] = (i - 1, 0)
|
||||
for j in range(1, m + 1):
|
||||
options = [
|
||||
(dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH, (i - 1, j - 1)),
|
||||
(
|
||||
dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH,
|
||||
(i - 1, j - 1),
|
||||
),
|
||||
(dp[i][j - 1] - _GAP_T, (i, j - 1)),
|
||||
(dp[i - 1][j] - _GAP_S, (i - 1, j)),
|
||||
]
|
||||
@@ -592,17 +650,37 @@ def _place_marks(translation: str, weight: int, marks: list[Mark]) -> str | None
|
||||
|
||||
def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
|
||||
"""Splice translated segments back into the original fragment; None on
|
||||
any validation failure (count mismatch, empty or non-prose segment) —
|
||||
the caller drops the result and the fragment stays pending. Segments
|
||||
with marks (a block that crossed as one piece) get their links
|
||||
re-inserted at weight-mapped positions after the prose check."""
|
||||
any validation failure (count mismatch, empty, non-prose or
|
||||
block-structure segment) — the caller drops the result and the fragment
|
||||
stays pending. Segments with marks (a block that crossed as one piece)
|
||||
get their links re-inserted at weight-mapped positions after the prose
|
||||
check.
|
||||
|
||||
Markdown-significant ASCII punctuation that pure_prose cannot see
|
||||
(plain text inline, syntax in the splice context — quoted titles, alt
|
||||
and link texts, table rows) is neutralized to Unicode look-alikes
|
||||
(``_NEUTRAL``) before splicing and mark placement (the swap is
|
||||
char-for-char, so unit alignment is unaffected); lines that would
|
||||
start a new block (a heading, a ``` or ::: fence — they would eat the
|
||||
rest of the block/fence they land in) reject the result outright
|
||||
(``_BLOCK``, ``_BLANK``)."""
|
||||
if len(texts) != len(spans):
|
||||
return None
|
||||
out: list[str] = []
|
||||
cursor = 0
|
||||
for span, translation in zip(spans, texts):
|
||||
if not translation.strip() or not pure_prose(translation):
|
||||
# Decode the wire form ("<" back to "<") first: pure_prose then
|
||||
# validates exactly what gets spliced — a "<" the model formed
|
||||
# into anything tag-like is markup and rejects the result.
|
||||
translation = translation.replace("<", "<")
|
||||
if (
|
||||
not translation.strip()
|
||||
or not pure_prose(translation)
|
||||
or _BLOCK.search(translation)
|
||||
or _BLANK.search(translation.strip())
|
||||
):
|
||||
return None
|
||||
translation = translation.translate(_NEUTRAL)
|
||||
if span.marks:
|
||||
translation = _place_marks(translation, span.weight, span.marks)
|
||||
if translation is None:
|
||||
|
||||
+10
-14
@@ -3,7 +3,7 @@
|
||||
Everything the route modules (files, api, tracking, pages) need that is not
|
||||
a route itself: environment-derived paths and tunables, the ``Data`` root
|
||||
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
|
||||
store, the fastapi-vue ``Frontend``, the page render cache
|
||||
store, the page render cache
|
||||
(``_render_html``/``_cached_body``/``_html_response`` plus the
|
||||
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
|
||||
content/settings write), the translator ``dispatcher``, the slug charset
|
||||
@@ -22,13 +22,13 @@ from pathlib import Path
|
||||
import blake3
|
||||
from fastapi import HTTPException, Request
|
||||
from fastapi.responses import Response
|
||||
from fastapi_vue import Frontend
|
||||
from kanta import Kanta
|
||||
from zstandard import ZstdCompressor
|
||||
|
||||
from pagerite import analytics, i18n, seed, translate, views
|
||||
from pagerite.__main__ import DEVMODE
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.config import load
|
||||
from pagerite.data import (
|
||||
Data,
|
||||
Node,
|
||||
@@ -39,10 +39,13 @@ from pagerite.data import (
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
|
||||
config = load()
|
||||
|
||||
# Site identity: the hostname comes from the CLI (first positional argument,
|
||||
# exported as PAGERITE_HOSTNAME) and names the per-site data directory
|
||||
# passed in PAGERITE_CONFIG) and names the per-site data directory
|
||||
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
|
||||
HOSTNAME = os.getenv("PAGERITE_HOSTNAME", "localhost")
|
||||
HOSTNAME = config.hostname
|
||||
SITE_DIR = Path(HOSTNAME)
|
||||
#: Public origin of the site, used for absolute social/canonical/sitemap
|
||||
#: URLs. Localhost serves varying ports, so it falls back to the request's
|
||||
@@ -76,12 +79,6 @@ FAVICON_MAXSIZE = 192
|
||||
data = Data()
|
||||
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
|
||||
|
||||
# Vue build served at the site root, no SPA catch-all (assets only). The
|
||||
# build mirrors the URL space: hashed, immutable files live under
|
||||
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico.
|
||||
BUILD_DIR = Path(__file__).with_name("frontend-build")
|
||||
frontend = Frontend(BUILD_DIR, spa=False, cached="/_assets/")
|
||||
|
||||
# Dynamic HTML is compressed per request at level 9 (static assets are
|
||||
# already pre-compressed by fastapi-vue's Frontend).
|
||||
_zstd = ZstdCompressor(9)
|
||||
@@ -345,6 +342,7 @@ def _seed(data: Data) -> None:
|
||||
|
||||
#: Translator key format: 12 lowercase alphanumeric characters — not
|
||||
#: brute-forceable over a WebSocket handshake, still human-manageable.
|
||||
#: The editor's lang tab generates further keys in the same format.
|
||||
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
|
||||
|
||||
|
||||
@@ -352,10 +350,8 @@ _KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
|
||||
def _translator_defaults(data: Data) -> None:
|
||||
"""Translator defaults on database creation: the first service key and
|
||||
the wanted target languages (Spanish and Chinese — English is the
|
||||
original language, never a translation target).
|
||||
|
||||
Keys are a dict (key -> display name) with the future reservation that
|
||||
multiple keys could be managed (e.g. via a web interface)."""
|
||||
original language, never a translation target). Further keys are
|
||||
managed in the editor shell's lang tab."""
|
||||
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
|
||||
data.translate_keys[key] = "default"
|
||||
data.translate_langs = {"es": True, "zh": True}
|
||||
|
||||
@@ -111,12 +111,13 @@ article h3 {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent);
|
||||
border-inline-start-color: var(--accent);
|
||||
background: color-mix(var(--accent) 6%, transparent);
|
||||
padding: 0.4rem 0.9rem;
|
||||
/* Keep the quoted text on the paragraph edge: the tinted box extends
|
||||
past it by its own border/padding, like code blocks. */
|
||||
margin: 0 -0.9rem 1rem calc(-0.25rem - 0.9rem);
|
||||
margin: 0 0 1rem;
|
||||
margin-inline: calc(-0.25rem - 0.9rem) -0.9rem;
|
||||
border-radius: 6px;
|
||||
}
|
||||
|
||||
@@ -125,9 +126,9 @@ blockquote {
|
||||
bar stays in both. */
|
||||
pre {
|
||||
border: 1px solid transparent;
|
||||
border-left: 0.25rem solid var(--accent);
|
||||
border-inline-start: 0.25rem solid var(--accent);
|
||||
/* Text on the paragraph edge: the box extends by padding + border. */
|
||||
margin-left: calc(-0.8rem - 0.25rem);
|
||||
margin-inline-start: calc(-0.8rem - 0.25rem);
|
||||
border-radius: 6px;
|
||||
}
|
||||
|
||||
|
||||
@@ -144,9 +144,9 @@ article h1 {
|
||||
font-weight: 700;
|
||||
padding-bottom: 0.5rem;
|
||||
/* The hazard-stripe underline breaks out of the page box: the negative
|
||||
right margin extends the h1's box (and thus its background) all the
|
||||
way to the viewport's right edge. */
|
||||
margin-right: calc((100% - 100vw) / 2);
|
||||
end margin extends the h1's box (and thus its background) all the
|
||||
way to the viewport's edge on that side. */
|
||||
margin-inline-end: calc((100% - 100vw) / 2);
|
||||
background:
|
||||
linear-gradient(-55deg,
|
||||
transparent 0 0.2rem,
|
||||
@@ -186,20 +186,21 @@ article ul ul ul li::before {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
background: color-mix(var(--accent2) 6%, transparent);
|
||||
padding: 0.25rem 0.75rem;
|
||||
/* Keep the quoted text on the paragraph edge: the tinted box extends
|
||||
past it by its own border/padding, like code blocks. */
|
||||
margin: 0 -0.75rem 1rem -1rem;
|
||||
margin: 0 0 1rem;
|
||||
margin-inline: -1rem -0.75rem;
|
||||
}
|
||||
|
||||
/* Code follows the color scheme; the dark-scheme well joins the violet
|
||||
family (--code-bg above). The orange side bar stays in both. */
|
||||
pre {
|
||||
border-left: 0.25rem solid var(--accent);
|
||||
border-inline-start: 0.25rem solid var(--accent);
|
||||
/* Text on the paragraph edge: the box extends by padding + border. */
|
||||
margin-left: calc(-0.8rem - 0.25rem);
|
||||
margin-inline-start: calc(-0.8rem - 0.25rem);
|
||||
border-radius: 3px;
|
||||
}
|
||||
|
||||
|
||||
@@ -66,7 +66,7 @@ article ul li::before {
|
||||
content: "◆";
|
||||
color: var(--accent);
|
||||
font-size: 0.8em;
|
||||
margin-left: calc(-1 * var(--list-indent) / 0.8);
|
||||
margin-inline-start: calc(-1 * var(--list-indent) / 0.8);
|
||||
width: calc(var(--list-indent) / 0.8);
|
||||
}
|
||||
|
||||
@@ -80,7 +80,7 @@ article ul ul ul li::before {
|
||||
}
|
||||
|
||||
blockquote {
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
}
|
||||
|
||||
/* Code panels sit slightly lighter than the page; the token colors come
|
||||
|
||||
@@ -160,7 +160,7 @@ main::before {
|
||||
top edge and a grassy shadow. */
|
||||
#sidebar {
|
||||
background: linear-gradient(160deg, #f4faddd9, #d9eec5cf);
|
||||
border-right: 1px solid #ffffff80;
|
||||
border-inline-end: 1px solid #ffffff80;
|
||||
border-bottom: 1px solid var(--line);
|
||||
box-shadow: 0 0.3rem 1rem #4f913b1f;
|
||||
border-radius: 1rem;
|
||||
@@ -225,7 +225,7 @@ article ul ul ul li::before {
|
||||
/* Quotes get a grassy edge and a wash of sunlight. */
|
||||
blockquote {
|
||||
color: #4d6849;
|
||||
border-left-color: var(--accent2);
|
||||
border-inline-start-color: var(--accent2);
|
||||
background: linear-gradient(90deg, #fff0a238, transparent 70%);
|
||||
padding-top: 0.25rem;
|
||||
padding-bottom: 0.25rem;
|
||||
|
||||
+120
-43
@@ -16,6 +16,7 @@ import os
|
||||
import re
|
||||
import shutil
|
||||
import socket
|
||||
from datetime import date
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
from urllib.parse import urlparse
|
||||
@@ -26,11 +27,16 @@ from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
|
||||
from fastapi.responses import Response
|
||||
|
||||
from pagerite import analytics
|
||||
from pagerite.data import resolve
|
||||
from pagerite.files import _hash_name, file_store
|
||||
from pagerite.state import SITE_URL, _html_response, analytics_store
|
||||
from pagerite.state import SITE_URL, _html_response, analytics_store, data
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# httpx logs every request at INFO (e.g. the favicon fetches below); our own
|
||||
# one-line summary in _schedule_favicon_fetch replaces that noise.
|
||||
logging.getLogger("httpx").setLevel(logging.WARNING)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
# Live WebSocket clients for the analytics stream.
|
||||
@@ -41,6 +47,59 @@ _analytics_broadcast_task: asyncio.Task | None = None
|
||||
# Repository root from this file's location (pagerite/tracking.py -> ..).
|
||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
|
||||
|
||||
|
||||
def _download_dbip() -> None:
|
||||
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
|
||||
today = date.today()
|
||||
months = [f"{today:%Y-%m}"]
|
||||
# The current month's file may not be published yet; fall back to last month.
|
||||
prev = (today.replace(day=1) - date.resolution).replace(day=1)
|
||||
months.append(f"{prev:%Y-%m}")
|
||||
|
||||
existing = sorted(
|
||||
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
|
||||
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
|
||||
)
|
||||
if existing and existing[-1] >= months[0]:
|
||||
logger.info("DB-IP database is current (%s), skipping download", existing[-1])
|
||||
return
|
||||
|
||||
for month in months:
|
||||
url = DBIP_URL.format(month=month)
|
||||
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
|
||||
tmp = target.with_suffix(".mmdb.gz.tmp")
|
||||
logger.info("Downloading %s", url)
|
||||
try:
|
||||
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
|
||||
if r.status_code == 404:
|
||||
continue
|
||||
r.raise_for_status()
|
||||
with open(tmp, "wb") as f:
|
||||
for chunk in r.iter_bytes():
|
||||
f.write(chunk)
|
||||
except httpx.HTTPError as e:
|
||||
logger.warning("DB-IP download failed: %s", e)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
# Verify it is actually gzip data before installing it.
|
||||
try:
|
||||
with gzip.open(tmp, "rb") as f:
|
||||
f.read(1)
|
||||
except OSError:
|
||||
logger.warning("DB-IP download for %s was not valid gzip", month)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
os.replace(tmp, target)
|
||||
# Drop older databases so the app never picks up a stale one.
|
||||
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
|
||||
if old.name != target.name:
|
||||
old.unlink()
|
||||
logger.info("DB-IP database updated to %s", target.name)
|
||||
return
|
||||
logger.warning("Could not download a DB-IP database")
|
||||
|
||||
|
||||
def _geoip_db_path() -> Path | None:
|
||||
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed
|
||||
@@ -253,9 +312,18 @@ async def _fetch_favicon(origin: str) -> None:
|
||||
|
||||
def _schedule_favicon_fetch() -> None:
|
||||
"""Start background favicon fetches for origins that need one."""
|
||||
for origin in analytics_store.favicon_origins_needed():
|
||||
if origin in _favicon_in_flight:
|
||||
continue
|
||||
origins = [
|
||||
origin
|
||||
for origin in analytics_store.favicon_origins_needed()
|
||||
if origin not in _favicon_in_flight
|
||||
]
|
||||
if not origins:
|
||||
return
|
||||
logger.info(
|
||||
"Fetching favicons: %s",
|
||||
", ".join(o.removeprefix("https://") for o in origins),
|
||||
)
|
||||
for origin in origins:
|
||||
_favicon_in_flight.add(origin)
|
||||
asyncio.create_task(_fetch_favicon(origin))
|
||||
|
||||
@@ -264,7 +332,7 @@ async def _broadcast_analytics() -> None:
|
||||
"""Send the current analytics snapshot to every connected WS client."""
|
||||
if not _analytics_ws_clients:
|
||||
return
|
||||
payload = analytics_store.display_json()
|
||||
payload = analytics_store.display_json(_in_menu)
|
||||
closed = set()
|
||||
for ws in _analytics_ws_clients:
|
||||
try:
|
||||
@@ -291,47 +359,51 @@ def _schedule_analytics_broadcast() -> None:
|
||||
)
|
||||
|
||||
|
||||
def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]:
|
||||
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
|
||||
def _in_menu(path: str) -> bool:
|
||||
"""True when ``path`` ("/a/b" or "/") resolves to a real menu node.
|
||||
|
||||
Nothing is counted on the GET itself — the client's first /_ws message
|
||||
starts the visit, so bots never register as visits (JS-running crawlers
|
||||
connect too, but the WebSocket handler ignores known bot UAs). (Admin
|
||||
clients report too, but with hide, which flags their visit hidden: it is
|
||||
recorded but excluded from all statistics and from the crawler list.)
|
||||
Category placeholders return 404 but are real nodes: their GETs must not
|
||||
count as misses in the display-time abuse classification.
|
||||
"""
|
||||
return resolve(data.menu, path.strip("/")) is not None
|
||||
|
||||
|
||||
def _record_get(request: Request, *, status: int = 200) -> None:
|
||||
"""Record the document GET as one raw access-log line in analytics.
|
||||
|
||||
Nothing is classified here — the true HTTP status, the full request path
|
||||
(query included), an external referer origin and the preload flag are
|
||||
stored, and visitor/crawler/abuse classification happens at display time
|
||||
(see analytics.Store.display). Idle-time preloads from pagerite.js
|
||||
(``x-pagerite-preload`` header) are recorded with ``pre=True``: never
|
||||
counted, but a navigation later served from the in-memory page cache is
|
||||
attributed this GET's status.
|
||||
|
||||
The devserver's health probe (``GET /?from=devserver.py`` from
|
||||
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
|
||||
logged as a crawler hit. The root-path and localhost checks prevent
|
||||
remote visitors from hiding traffic with the same query string.
|
||||
|
||||
Returns the client hashes of any pending crawler hits flushed to persistent
|
||||
storage, so callers can schedule async geoip and reverse-DNS enrichment.
|
||||
``127.0.0.1``) is ignored: it is not real traffic. The root-path and
|
||||
localhost checks prevent remote visitors from forging the same query.
|
||||
"""
|
||||
if request.headers.get("x-pagerite-preload"):
|
||||
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
|
||||
# activity message sent when the user actually navigates does the
|
||||
# counting.
|
||||
# (Forging the header only hides a GET from the crawler stats; the
|
||||
# path-based abuse classification is unaffected.)
|
||||
return []
|
||||
if (
|
||||
path == ""
|
||||
request.url.path == "/"
|
||||
and str(request.url.query) == "from=devserver.py"
|
||||
and _client_ip(request) == "127.0.0.1"
|
||||
):
|
||||
return []
|
||||
return
|
||||
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
|
||||
full_path = f"{request.url.path}{_query_suffix(request)}"
|
||||
return analytics_store.track_entry(
|
||||
request.headers.get("referer", ""),
|
||||
own_origin,
|
||||
referer = request.headers.get("referer", "")
|
||||
if analytics._origin(referer) in (None, own_origin):
|
||||
referer = ""
|
||||
client_hash = analytics_store.record_get(
|
||||
_client_ip(request),
|
||||
request.headers.get("user-agent", ""),
|
||||
full_path,
|
||||
request.headers.get("accept-language", ""),
|
||||
f"{request.url.path}{_query_suffix(request)}",
|
||||
status=status,
|
||||
referer=referer,
|
||||
accept_language=request.headers.get("accept-language", ""),
|
||||
pre=bool(request.headers.get("x-pagerite-preload")),
|
||||
)
|
||||
if client_hash is not None:
|
||||
_schedule_client_enrichment([client_hash])
|
||||
|
||||
|
||||
@router.get("/_a", response_model=None)
|
||||
@@ -358,14 +430,21 @@ async def activity_ws(ws: WebSocket) -> None:
|
||||
Public, like the pages themselves (only /_api is gated); one connection
|
||||
follows a browsing session. Messages are ``analytics.Ping`` structs as
|
||||
JSON text frames; ``to`` set is a navigation, ``read`` alone a
|
||||
reading-time update. The reverse-DNS and DB-IP geoip lookups happen in
|
||||
background tasks so message handling is never delayed by slow DNS or
|
||||
the first MMDB decompress.
|
||||
reading-time update. Everything is recorded raw — known bot UAs and
|
||||
abusive IPs are filtered at display time, not here. The reverse-DNS and
|
||||
DB-IP geoip lookups happen in background tasks so message handling is
|
||||
never delayed by slow DNS or the first MMDB decompress.
|
||||
"""
|
||||
await ws.accept()
|
||||
ip = _client_ip(ws)
|
||||
ua = ws.headers.get("user-agent", "")
|
||||
accept_language = ws.headers.get("accept-language", "")
|
||||
# Identify the visitor on the access-log open/close lines (the IP is
|
||||
# already printed there): compact UA plus the browser's language tag.
|
||||
lang, _country = analytics._parse_accept_language(accept_language)
|
||||
ws.scope.setdefault("state", {})["log_extra"] = " ".join(
|
||||
part for part in (analytics._compact_user_agent(ua), lang) if part
|
||||
)
|
||||
await ws.accept()
|
||||
try:
|
||||
while True:
|
||||
text = await ws.receive_text()
|
||||
@@ -373,7 +452,7 @@ async def activity_ws(ws: WebSocket) -> None:
|
||||
msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
|
||||
except msgspec.DecodeError:
|
||||
continue
|
||||
visit_index, flushed_clients = analytics_store.ping(
|
||||
new_client = analytics_store.record_msg(
|
||||
msg.fr,
|
||||
msg.to or None,
|
||||
ip,
|
||||
@@ -382,10 +461,8 @@ async def activity_ws(ws: WebSocket) -> None:
|
||||
hide=msg.hide,
|
||||
read=msg.read,
|
||||
)
|
||||
if visit_index is not None:
|
||||
visit = analytics_store.data.visits[visit_index]
|
||||
asyncio.create_task(_enrich_client(visit.client))
|
||||
_schedule_client_enrichment(flushed_clients)
|
||||
if new_client is not None:
|
||||
_schedule_client_enrichment([new_client])
|
||||
_schedule_favicon_fetch()
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
@@ -399,7 +476,7 @@ async def analytics_websocket(ws: WebSocket) -> None:
|
||||
endpoint. Powers the analytics viewer rendered at /_a.
|
||||
"""
|
||||
await ws.accept()
|
||||
await ws.send_text(analytics_store.display_json())
|
||||
await ws.send_text(analytics_store.display_json(_in_menu))
|
||||
_analytics_ws_clients.add(ws)
|
||||
try:
|
||||
while True:
|
||||
|
||||
+34
-28
@@ -277,34 +277,40 @@ class Dispatcher:
|
||||
job = None
|
||||
spans: list[Span] = []
|
||||
original = ""
|
||||
for lang in sorted(langs):
|
||||
# Titles first: a page's name in the menu is its most
|
||||
# visible string (stable: menu order kept within each kind).
|
||||
for item in sorted(
|
||||
pending_items(self.data, lang), key=lambda it: it.kind != "title"
|
||||
):
|
||||
if (lang, item.key) in inflight or (
|
||||
lang,
|
||||
item.key,
|
||||
) in self.validation_failures:
|
||||
continue
|
||||
spans, texts, contexts = split(item.text)
|
||||
if not texts:
|
||||
continue # prose that could not be located for splicing
|
||||
original = item.text
|
||||
if item.kind == "title" and item.context:
|
||||
# A title's surround is the article's opening prose
|
||||
# (TransItem.context), not its own one-word block.
|
||||
contexts = [item.context] * len(texts)
|
||||
job = Job(
|
||||
lang=lang,
|
||||
key=item.key,
|
||||
texts=texts,
|
||||
path=item.path,
|
||||
kind=item.kind,
|
||||
contexts=contexts,
|
||||
)
|
||||
break
|
||||
# Titles before articles — across languages too, so every menu
|
||||
# is named before any article body is worked on (a page's name
|
||||
# is its most visible string). pending_items emits in menu
|
||||
# order, a page's title before its chunks; filtering by kind
|
||||
# keeps that stable order within each kind.
|
||||
pending = {lang: pending_items(self.data, lang) for lang in sorted(langs)}
|
||||
for kind in ("title", "chunk"):
|
||||
for lang in sorted(langs):
|
||||
for item in pending[lang]:
|
||||
if (
|
||||
item.kind != kind
|
||||
or (lang, item.key) in inflight
|
||||
or (lang, item.key) in self.validation_failures
|
||||
):
|
||||
continue
|
||||
spans, texts, contexts = split(item.text)
|
||||
if not texts:
|
||||
continue # prose that could not be located for splicing
|
||||
original = item.text
|
||||
if item.kind == "title" and item.context:
|
||||
# A title's surround is the article's opening prose
|
||||
# (TransItem.context), not its own one-word block.
|
||||
contexts = [item.context] * len(texts)
|
||||
job = Job(
|
||||
lang=lang,
|
||||
key=item.key,
|
||||
texts=texts,
|
||||
path=item.path,
|
||||
kind=item.kind,
|
||||
contexts=contexts,
|
||||
)
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is not None:
|
||||
break
|
||||
if job is None:
|
||||
|
||||
+17
-5
@@ -25,6 +25,7 @@ from html5tagger import HTML, Document, E, Template
|
||||
from platformdirs import site_data_dir, user_data_path
|
||||
|
||||
from pagerite import i18n
|
||||
from pagerite.config import load as _load_config
|
||||
from pagerite.data import Data, Node, node_markdown, prettify, resolve, sorted_nodes
|
||||
from pagerite.i18n import Translation
|
||||
from pagerite.markdown import render
|
||||
@@ -45,6 +46,10 @@ def _data_roots() -> list[Path]:
|
||||
return [Path(r) for r in roots]
|
||||
|
||||
|
||||
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
|
||||
config = _load_config()
|
||||
|
||||
|
||||
def _theme_dirs() -> list[Path]:
|
||||
"""Theme search roots, most specific first; first match wins per file.
|
||||
|
||||
@@ -57,7 +62,7 @@ def _theme_dirs() -> list[Path]:
|
||||
"""
|
||||
return [
|
||||
Path("themes"),
|
||||
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "themes",
|
||||
Path(config.hostname) / "themes",
|
||||
*(root / "themes" for root in _data_roots()),
|
||||
Path(__file__).parent / "themes",
|
||||
]
|
||||
@@ -73,7 +78,7 @@ THEME_DIRS = _theme_dirs()
|
||||
# built-in --font-* variables.
|
||||
FONT_DIRS = [
|
||||
Path("fonts"),
|
||||
Path(os.getenv("PAGERITE_HOSTNAME", "localhost")) / "fonts",
|
||||
Path(config.hostname) / "fonts",
|
||||
*(root / "fonts" for root in _data_roots()),
|
||||
]
|
||||
|
||||
@@ -344,8 +349,9 @@ def _layout(
|
||||
property attributes, everything else (description, twitter:*) as name.
|
||||
|
||||
``lang`` is the served language for <html lang>; an RTL language (ar,
|
||||
fa, ...) also puts dir="rtl" on <html> (the editor panel carries its own
|
||||
lang="en" dir="ltr", so it is unaffected). ``canonical`` and
|
||||
fa, ...) also puts dir="rtl" on <html> (the editor panel and the
|
||||
analytics dashboard carry their own lang="en" dir="ltr", so they are
|
||||
unaffected). ``canonical`` and
|
||||
``alternates`` ((hreflang, href) pairs) are the page's language URLs
|
||||
(see docs/localization.md), emitted right after the viewport and before
|
||||
the social tags: canonical first, then the hreflang alternates.
|
||||
@@ -777,7 +783,11 @@ def page_content(
|
||||
node = resolve(menu, path)[-1]
|
||||
content = node_markdown(data, node) or ""
|
||||
title = node.title
|
||||
# The original text pins the section anchors: on a translated page the
|
||||
# heading slugs (and thus #hash URLs) stay in the original language.
|
||||
anchors_from = None
|
||||
if translation:
|
||||
anchors_from = (content, title)
|
||||
if translation.markdown is not None:
|
||||
content = translation.markdown
|
||||
title = (
|
||||
@@ -787,7 +797,9 @@ def page_content(
|
||||
)
|
||||
# The title is injected into the markdown (as # title when it has no
|
||||
# h1 of its own), so title and content render as one article.
|
||||
rendered = render(content, path, node.created, node.modified, title=title)
|
||||
rendered = render(
|
||||
content, path, node.created, node.modified, title=title, anchors_from=anchors_from
|
||||
)
|
||||
# Long articles get .multicol: the article column cap lifts (see the
|
||||
# #content grid in pagerite.css) and the .cols segments lay out in at
|
||||
# most two columns. The html is already segmented by render() — the
|
||||
|
||||
+1
-1
@@ -17,7 +17,7 @@ readme = "README.md"
|
||||
requires-python = ">=3.14"
|
||||
dependencies = [
|
||||
"blake3>=1.0.9",
|
||||
"fastapi-vue~=1.4.0",
|
||||
"fastapi-vue~=1.4.2",
|
||||
"fastapi[standard]>=0.141.1",
|
||||
"html5tagger>=2.0.0",
|
||||
"httpx>=0.28.1",
|
||||
|
||||
Reference in New Issue
Block a user