Cleanup: break up the massive app.py into separate modules of manageable size.
This commit is contained in:
@@ -10,11 +10,16 @@ Please instead ask the user to see from dev tools what you need, e.g. to look up
|
||||
Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes:
|
||||
|
||||
- `pagerite/` — Python backend package (hatchling build target).
|
||||
- `app.py` — FastAPI app and route registration.
|
||||
- `app.py` — thin FastAPI assembly: lifespan, `FastAPI(...)`, router includes (route ordering: api/tracking/files routers, then `frontend.route(app, "/")`, then the pages catch-all last).
|
||||
- `state.py` — shared core, no routes: env-derived site constants, `data`/`kanta`, `analytics_store`, the fastapi-vue `frontend`, the render cache (`_html_response`, `_invalidate_pages`), the translator `dispatcher`, slug helpers, `@kanta.bootstrap` hooks.
|
||||
- `files.py` — `FileStore` (content-addressed, RAM-cached), image derivative helpers (`store_image`), file routes (`/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, favicon settings).
|
||||
- `api.py` — editor REST + WS: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, `/_translate/{key}`.
|
||||
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetch, `/_ws`, `/_api/ws/analytics`, the `/_a` page (docs/analytics.md).
|
||||
- `pages.py` — public content pages: `/`, `/sitemap.xml`, `/robots.txt`, the `/{path:path}` catch-all.
|
||||
- `data.py` — msgspec Structs for the kanta database.
|
||||
- `chunks.py` — block-level Markdown chunking and content-hash keys for the chunk stores (docs/migrate.md).
|
||||
- `i18n.py` — language selection, translation assembly (chunks + patches) and translated-edit recording (user patches, per-language title overrides, refresh).
|
||||
- `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); app.py only registers the route.
|
||||
- `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); api.py only registers the route.
|
||||
- `segments.py` — the translation round trip: fragments split into pure-prose wire segments (via markdown.make_md's verbatim parser; link- and formatting-carrying blocks stay whole, link/formatted texts inline, Markdown stripped) and translations spliced back by source offset, link/formatting markdown re-inserted at weight-mapped positions (docs/localization.md).
|
||||
- `migrations.py` — kanta migrations (`migrate_vN`); ALL schema/storage upgrades live here (raw state dict before struct decoding), never in the app lifespan: v1 moves legacy in-db file blobs to the on-disk store and rebuilds the legacy flat `pages` as the menu tree, v2 rewrites `/_f/{hash}.ext` image links to the extension-less form, backfills AVIF/WebP/JPEG derivatives on disk and drops the obsolete `version` field.
|
||||
- `markdown.py` — markdown-it-py renderer.
|
||||
|
||||
+4
-3
@@ -8,9 +8,10 @@ directory, e.g. `localhost/analytics.json`).
|
||||
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
|
||||
`CrawlerHit`, `AbuseHit`, `Favicon`) and the `Store` (in-memory data + session map,
|
||||
atomic JSON persistence).
|
||||
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
|
||||
the `/_ws` activity WebSocket, and `WebSocket /_api/ws/analytics`
|
||||
(admin-gated like every `/_api` endpoint).
|
||||
- `pagerite/pages.py` — entry-referer stashing in `show_page` (`_track_entry`,
|
||||
in `pagerite/tracking.py`), 404 recording.
|
||||
- `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
|
||||
`WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
|
||||
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
|
||||
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
|
||||
normal site layout on the `/_a` analytics page.
|
||||
|
||||
+12
-4
@@ -4,13 +4,21 @@ The Python backend lives in `pagerite/`.
|
||||
|
||||
## `app.py`
|
||||
|
||||
The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything defined earlier wins. The one exception is the content catch-all `/{path:path}`, registered AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
|
||||
Thin FastAPI assembly: lifespan (open the kanta database, load the file store, the frontend build and GeoIP), the `FastAPI(...)` instance with built-in API docs disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content, the `server` header middleware, and router includes. The routes themselves live in specialized modules:
|
||||
|
||||
- `state.py` — shared core, no routes: the environment-derived site constants (`HOSTNAME`, `SITE_URL`, `DB_PATH`, `FILES_DIR`, image/favicon tunables), the `data` root and its `kanta` handle (`Kanta(..., migrations="pagerite.migrations")`), the `analytics_store`, the fastapi-vue `frontend`, the page render cache and `_html_response`, the translator `dispatcher`, the slug charset helpers, and the `@kanta.bootstrap` hooks (demo seed, translator defaults).
|
||||
- `files.py` — the `FileStore` and image derivative helpers, and the file routes: `/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, the favicon settings endpoints.
|
||||
- `api.py` — the editor REST API and WebSockets: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, and the translator channel `/_translate/{clientkey}`.
|
||||
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetching, debounced broadcasts, the `/_ws` activity socket, the admin stream `/_api/ws/analytics`, and the `/_a` viewer page.
|
||||
- `pages.py` — the public content pages: `/`, `/sitemap.xml`, `/robots.txt` and the `/{path:path}` catch-all.
|
||||
|
||||
Route ordering is load-bearing and lives in `app.py`: the api/tracking/files routers are included BEFORE `frontend.route(app, "/")` is called — fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything registered earlier wins. The content catch-all `/{path:path}` is included AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
|
||||
|
||||
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
|
||||
|
||||
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding — and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
|
||||
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `state.py`'s `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding — and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
|
||||
|
||||
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore`: content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
|
||||
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore` (in `files.py`): content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
|
||||
|
||||
## `data.py`
|
||||
|
||||
@@ -34,4 +42,4 @@ Any page with published children — a category page — lists them as a card gr
|
||||
|
||||
## `seed.py`
|
||||
|
||||
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `app.py`.
|
||||
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `state.py`.
|
||||
|
||||
@@ -10,11 +10,11 @@ The site structure is stored in the kanta database managed by `pagerite/data.py`
|
||||
|
||||
Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `pages` (pre-tree databases) migrates into `menu` via `migrate_v1`. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`.
|
||||
|
||||
Every content/settings write calls `_invalidate_pages()` in app.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
|
||||
Every content/settings write calls `_invalidate_pages()` in state.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
|
||||
|
||||
## Files
|
||||
|
||||
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in app.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
|
||||
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in files.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
|
||||
|
||||
## Banners
|
||||
|
||||
|
||||
@@ -347,7 +347,7 @@ patches, so the dispatcher re-translates everything from scratch; the
|
||||
run's validation skip-list is cleared with it, giving rejected fragments
|
||||
another chance.
|
||||
|
||||
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; app.py only
|
||||
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; api.py only
|
||||
registers the route):
|
||||
|
||||
- **One job at a time per connection** — the next job is sent only after
|
||||
|
||||
+2
-2
@@ -53,7 +53,7 @@ class Node(msgspec.Struct, omit_defaults=True):
|
||||
class Data(msgspec.Struct):
|
||||
...
|
||||
#: API keys gating the translator service WebSocket (/_translate/{key}):
|
||||
#: key -> display name; the first is generated at bootstrap (app.py).
|
||||
#: key -> display name; the first is generated at bootstrap (state.py).
|
||||
translate_keys: dict[str, str] = {}
|
||||
#: Wanted target languages for the translator service (presence-keys);
|
||||
#: jobs are offered only in these ∩ a connection's capabilities.
|
||||
@@ -147,7 +147,7 @@ translation data, in the same transaction:
|
||||
Chunking must be deterministic and shared with render/save, so
|
||||
`chunk_markdown` + `chunk_key` live in `pagerite/i18n.py` (or a small
|
||||
`pagerite/chunks.py`) and are imported by both `migrations.py` and
|
||||
`views.py`/`app.py`.
|
||||
`views.py`/`state.py`.
|
||||
|
||||
## Implementation notes (deviations from the plan above)
|
||||
|
||||
|
||||
+12
-5
@@ -33,7 +33,9 @@ def _download_dbip() -> None:
|
||||
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
|
||||
)
|
||||
if existing and existing[-1] >= months[0]:
|
||||
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download")
|
||||
print(
|
||||
f"pagerite: DB-IP database is current ({existing[-1]}), skipping download"
|
||||
)
|
||||
return
|
||||
|
||||
for month in months:
|
||||
@@ -58,7 +60,10 @@ def _download_dbip() -> None:
|
||||
with gzip.open(tmp, "rb") as f:
|
||||
f.read(1)
|
||||
except OSError:
|
||||
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr)
|
||||
print(
|
||||
f"pagerite: DB-IP download for {month} was not valid gzip",
|
||||
file=sys.stderr,
|
||||
)
|
||||
tmp.unlink(missing_ok=True)
|
||||
continue
|
||||
os.replace(tmp, target)
|
||||
@@ -78,9 +83,11 @@ def main() -> None:
|
||||
"hostname",
|
||||
nargs="?",
|
||||
default="localhost",
|
||||
help=("Public hostname of the site; names the data directory "
|
||||
"<hostname>/{content.kantadb, analytics.json, files} under the "
|
||||
"cwd (default: localhost)."),
|
||||
help=(
|
||||
"Public hostname of the site; names the data directory "
|
||||
"<hostname>/{content.kantadb, analytics.json, files} under the "
|
||||
"cwd (default: localhost)."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"-l",
|
||||
|
||||
+17
-8
@@ -362,6 +362,7 @@ def _is_bot_ua(ua: str) -> bool:
|
||||
"""True when the UA claims a crawler identity (bot or spider)."""
|
||||
return bool(_BOT_UA.search(ua))
|
||||
|
||||
|
||||
#: Plain-404 count per IP that classifies it as abuse even without a
|
||||
#: telltale path hit.
|
||||
_ABUSE_404_THRESHOLD = 10
|
||||
@@ -399,9 +400,7 @@ def _client_hash(ip: str, ua: str, lang: str) -> bytes:
|
||||
The key is the prettified IP (IPv6 /64), the raw UA string and the
|
||||
extracted language tag, separated by null bytes.
|
||||
"""
|
||||
return blake3.blake3(
|
||||
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
|
||||
).digest()[:6]
|
||||
return blake3.blake3(f"{_network_ip(ip)}\0{ua}\0{lang}".encode()).digest()[:6]
|
||||
|
||||
|
||||
class Store:
|
||||
@@ -538,7 +537,9 @@ class Store:
|
||||
if nav.to.startswith("/"):
|
||||
nav_views = display.views.setdefault(nav.to, {})
|
||||
nav_views[nb] = nav_views.get(nb, 0) + 1
|
||||
nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(nav.to, {})
|
||||
nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(
|
||||
nav.to, {}
|
||||
)
|
||||
nbuckets[nb] = nbuckets.get(nb, 0) + 1
|
||||
return display
|
||||
|
||||
@@ -664,16 +665,22 @@ class Store:
|
||||
self.data.abuse_ips[ip] = True
|
||||
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
|
||||
if moved:
|
||||
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip]
|
||||
self.data.crawlers = [
|
||||
h for h in self.data.crawlers if self._client_ip(h.client) != ip
|
||||
]
|
||||
for h in moved:
|
||||
self._abuse_hit(
|
||||
h.client,
|
||||
h.entry + (f"?{h.query}" if h.query else ""),
|
||||
start=h.start,
|
||||
)
|
||||
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip]
|
||||
pending = [
|
||||
h for h in self.pending_crawlers if self._client_ip(h.client) == ip
|
||||
]
|
||||
if pending:
|
||||
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip]
|
||||
self.pending_crawlers = [
|
||||
h for h in self.pending_crawlers if self._client_ip(h.client) != ip
|
||||
]
|
||||
for h in pending:
|
||||
self._abuse_hit(
|
||||
h.client,
|
||||
@@ -917,5 +924,7 @@ class Store:
|
||||
else:
|
||||
visit.trail[now] = TrailItem(to=target, status=target_status)
|
||||
self._save()
|
||||
visit_index = index if index is not None and index < len(self.data.visits) else None
|
||||
visit_index = (
|
||||
index if index is not None and index < len(self.data.visits) else None
|
||||
)
|
||||
return visit_index, flushed
|
||||
|
||||
+619
@@ -0,0 +1,619 @@
|
||||
"""Editor REST API and WebSocket sessions.
|
||||
|
||||
The management endpoints behind the SSO forward-auth gate: the site tree
|
||||
(``/_api/pages``), structure operations (``/_api/structure``), site-wide
|
||||
settings (``/_api/settings``), task-list toggles (``/_api/toggle-task``),
|
||||
the translations refresh (``/_api/translations``), the editor session
|
||||
socket (``/_api/ws/editor``), and the translator service channel
|
||||
(``/_translate/{clientkey}`` — deliberately NOT under ``/_api``: the
|
||||
server-generated key in the path is the access control).
|
||||
"""
|
||||
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
|
||||
from fastapi import (
|
||||
APIRouter,
|
||||
HTTPException,
|
||||
WebSocket,
|
||||
WebSocketDisconnect,
|
||||
)
|
||||
from pydantic import BaseModel
|
||||
|
||||
from pagerite import i18n, views
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.data import (
|
||||
Node,
|
||||
append_order,
|
||||
find_slot,
|
||||
node_markdown,
|
||||
resolve,
|
||||
sorted_nodes,
|
||||
)
|
||||
from pagerite.markdown import render, toggle_task
|
||||
from pagerite.state import (
|
||||
_check_reserved,
|
||||
_ensure,
|
||||
_invalidate_pages,
|
||||
_remove_page,
|
||||
data,
|
||||
dispatcher,
|
||||
kanta,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
class PageIn(BaseModel):
|
||||
"""Payload for creating or replacing a page."""
|
||||
|
||||
title: str
|
||||
markdown: str
|
||||
published: bool = True
|
||||
banner: str | None = None # None keeps the existing banner
|
||||
|
||||
|
||||
@router.get("/_api/pages")
|
||||
async def list_pages(lang: str | None = None) -> list[dict]:
|
||||
"""The site tree for the structure editor (all nodes, drafts included).
|
||||
|
||||
Nested by slug; each node carries its full path, menu order, flags and
|
||||
language settings (``language`` is the node's own primary-language
|
||||
setting, "" = inherit; ``primary`` is the resolved effective one).
|
||||
With a ``?lang=`` translation, titles come out in that language where a
|
||||
translation exists (``translated`` flags it — true trivially for rows
|
||||
whose primary language IS the selected one; other rows fall back to
|
||||
the original title, dimmed) — the structure itself (slugs, order,
|
||||
hierarchy) is language-independent.
|
||||
"""
|
||||
tag = i18n.base_tag(lang or "")
|
||||
titles = i18n.title_map(data, tag) if tag else {}
|
||||
|
||||
def dump(nodes: dict[str, Node], prefix: str, inherited: str) -> list[dict]:
|
||||
out = []
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
primary = node.language or inherited
|
||||
out.append(
|
||||
{
|
||||
"slug": slug,
|
||||
"path": path,
|
||||
"title": titles.get(path) or node.title,
|
||||
"translated": path in titles or (bool(tag) and primary == tag),
|
||||
"order": node.order,
|
||||
"published": node.published,
|
||||
"has_content": node.chunks is not None,
|
||||
"language": node.language,
|
||||
"primary": primary,
|
||||
"children": dump(node.children, path, primary),
|
||||
}
|
||||
)
|
||||
return out
|
||||
|
||||
return dump(data.menu, "", i18n.ORIGINAL_LANGUAGE)
|
||||
|
||||
|
||||
@router.put("/_api/pages/{path:path}", status_code=204)
|
||||
async def save_page(path: str, page: PageIn, lang: str | None = None) -> None:
|
||||
"""Create or replace the page at a slug path ("" or "/" = front page).
|
||||
|
||||
Missing ancestors are created as content-less category labels. Giving
|
||||
a category markdown turns it into a landing page. Empty markdown (after
|
||||
stripping) creates an empty page that renders with just its title —
|
||||
saving never deletes; use DELETE to remove a page (the page editor
|
||||
issues DELETE when you save empty text).
|
||||
|
||||
With a ``?lang=`` query (a translation, not the primary language) the
|
||||
save is a translated-view edit (docs/localization.md): the markdown is
|
||||
diffed against the currently served hybrid and the minimal diff is
|
||||
appended as a Patch under ``patches[f"{path}:{lang}"]`` — node.chunks
|
||||
and the original-language fields (title, published, banner) stay
|
||||
untouched.
|
||||
"""
|
||||
path = path.strip("/")
|
||||
_check_reserved(path)
|
||||
lang = i18n.base_tag(lang or "")
|
||||
if lang and lang != i18n.primary_lang(data.menu, path):
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
with kanta.transaction("save translation", extra=path):
|
||||
# Patches alone make the translated version exist.
|
||||
if i18n.add_patch(data, node, path, lang, page.markdown):
|
||||
_invalidate_pages()
|
||||
return
|
||||
with kanta.transaction("save page", extra=path):
|
||||
node = _ensure(data.menu, path)
|
||||
node.title = page.title
|
||||
node.chunks = store_chunks(data.chunks, page.markdown)
|
||||
node.published = page.published
|
||||
if page.banner is not None:
|
||||
node.banner = page.banner
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.delete("/_api/pages/{path:path}", status_code=204)
|
||||
async def delete_page(path: str) -> None:
|
||||
"""Delete a node by slug path.
|
||||
|
||||
A category (node with children) loses only its landing page and stays
|
||||
as a content-less label; a childless node is removed entirely.
|
||||
"""
|
||||
path = path.strip("/")
|
||||
_check_reserved(path)
|
||||
with kanta.transaction("delete page", extra=path):
|
||||
if not _remove_page(data.menu, path):
|
||||
raise HTTPException(404, "no such page")
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
class StructureOp(BaseModel):
|
||||
"""Rearrange the site tree: reorder, move/rename or retitle a node.
|
||||
|
||||
`order` is a fresh fractional key computed client-side from the node's
|
||||
new siblings (a value halfway between them); all other items keep
|
||||
theirs. `move_to` is the full target path — the parent must exist and
|
||||
the new slug be free. Moves carry the whole subtree. The front page is
|
||||
just the top-level node with slug "": renaming it away leaves no front
|
||||
page ("/" then redirects to the first nav item), and any childless
|
||||
top-level node can take the empty slug to become the front page.
|
||||
|
||||
With `lang` (a translation, not the node's primary language) a `title`
|
||||
edit writes a per-language title fragment instead of the original — the
|
||||
same storage as machine title translations (docs/localization.md);
|
||||
sending the original's text removes the override. Structural fields are
|
||||
not combinable with a translated title edit.
|
||||
|
||||
`language` sets the node's primary language (a BCP-47 base tag; "" =
|
||||
inherit from the nearest ancestor, the front page last, site default
|
||||
"en" final — Node.language), inherited by the whole subtree.
|
||||
"""
|
||||
|
||||
path: str
|
||||
order: float | None = None
|
||||
move_to: str | None = None
|
||||
title: str | None = None
|
||||
lang: str | None = None
|
||||
language: str | None = None
|
||||
|
||||
|
||||
@router.post("/_api/structure", status_code=204)
|
||||
async def update_structure(op: StructureOp) -> None:
|
||||
"""Apply one structure operation (see StructureOp)."""
|
||||
path = op.path.strip("/")
|
||||
chain = resolve(data.menu, path)
|
||||
if chain is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
node = chain[-1]
|
||||
lang = i18n.base_tag(op.lang or "")
|
||||
if op.language is not None:
|
||||
# Primary-language setting (inherited by the subtree): reselects
|
||||
# what "the original" means for the node — its language is part of
|
||||
# every render, so a change invalidates everywhere.
|
||||
language = i18n.base_tag(op.language)
|
||||
with kanta.transaction("set page language", extra=path):
|
||||
if language != node.language:
|
||||
node.language = language
|
||||
_invalidate_pages()
|
||||
return
|
||||
if op.title is not None and lang and lang != i18n.primary_lang(data.menu, path):
|
||||
# Translated title (i18n.set_title_translation): original title,
|
||||
# slugs and hierarchy stay untouched.
|
||||
with kanta.transaction("translate title", extra=path):
|
||||
if i18n.set_title_translation(data, node, lang, op.title):
|
||||
_invalidate_pages()
|
||||
return
|
||||
target = op.move_to.strip("/") if op.move_to is not None else None
|
||||
if target is not None and target != path:
|
||||
_check_reserved(target)
|
||||
if path and target.startswith(f"{path}/"):
|
||||
raise HTTPException(400, "cannot move a page under itself")
|
||||
slot = find_slot(data.menu, target)
|
||||
if slot is None:
|
||||
raise HTTPException(404, "target parent does not exist")
|
||||
tnodes, tslug = slot
|
||||
if tslug in tnodes:
|
||||
raise HTTPException(400, "target path exists")
|
||||
if not tslug and node.children:
|
||||
raise HTTPException(400, "the front page cannot have children")
|
||||
with kanta.transaction("update structure", extra=path):
|
||||
if op.title is not None:
|
||||
node.title = op.title
|
||||
if target is not None and target != path:
|
||||
snodes, sslug = find_slot(data.menu, path)
|
||||
del snodes[sslug]
|
||||
# A pure rename (same parent) keeps its position; only a move
|
||||
# to another level appends at the end (unless an order came
|
||||
# with the drop).
|
||||
same_level = path.rpartition("/")[0] == target.rpartition("/")[0]
|
||||
node.order = (
|
||||
op.order
|
||||
if op.order is not None
|
||||
else node.order
|
||||
if same_level
|
||||
else append_order(tnodes)
|
||||
)
|
||||
tnodes[tslug] = node
|
||||
elif op.order is not None:
|
||||
node.order = op.order
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.get("/_api/settings")
|
||||
async def get_settings() -> dict:
|
||||
"""Site-wide settings (brand, theme, custom CSS and favicon URL), plus
|
||||
the themes, banner designs and user fonts available on disk for the
|
||||
selectors, the translator service keys and the wanted translation
|
||||
languages (for the /_translate socket)."""
|
||||
return {
|
||||
"brand": data.brand,
|
||||
"brand_html": data.brand_html,
|
||||
"theme": data.theme,
|
||||
"custom_css": data.custom_css,
|
||||
"favicon": f"/_f/{data.favicon}" if data.favicon else "",
|
||||
"themes": views._theme_info(),
|
||||
"banner_designs": views._banner_design_names(),
|
||||
"fonts": views._user_fonts(),
|
||||
"transition": data.transition,
|
||||
"transitions": views._transition_names(),
|
||||
"translate_keys": data.translate_keys,
|
||||
# The site default primary language: the front page's resolved
|
||||
# setting (every page may override it, inherited down the tree).
|
||||
"primary_lang": i18n.primary_lang(data.menu, ""),
|
||||
"translate_langs": sorted(data.translate_langs),
|
||||
}
|
||||
|
||||
|
||||
class SettingsIn(BaseModel):
|
||||
"""Payload for updating site-wide settings."""
|
||||
|
||||
brand: str
|
||||
theme: str
|
||||
custom_css: str
|
||||
brand_html: str = ""
|
||||
transition: str = "cube"
|
||||
translate_langs: list[str] | None = None # None keeps the current set
|
||||
|
||||
|
||||
@router.put("/_api/settings", status_code=204)
|
||||
async def put_settings(settings: SettingsIn) -> None:
|
||||
"""Update site-wide settings; invalidates cached pages and ETags."""
|
||||
with kanta.transaction("update settings"):
|
||||
data.brand = settings.brand
|
||||
data.brand_html = settings.brand_html
|
||||
data.theme = settings.theme
|
||||
data.custom_css = settings.custom_css
|
||||
data.transition = settings.transition
|
||||
if settings.translate_langs is not None:
|
||||
# Any language may be a target — including the site default
|
||||
# (an article in another language can be translated INTO it);
|
||||
# a node's own primary is excluded per article, not here.
|
||||
data.translate_langs = {
|
||||
tag: True
|
||||
for lang in settings.translate_langs
|
||||
if (tag := i18n.base_tag(lang))
|
||||
}
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
@router.delete("/_api/translations", status_code=204)
|
||||
async def delete_translations() -> None:
|
||||
"""Drop all machine translations (Data.trans) so the dispatcher
|
||||
re-translates everything from scratch (a "refresh translations" action:
|
||||
the invalidation hook re-offers every fragment to connected
|
||||
translators). User patches are kept; the availability index
|
||||
(node.langs) is rebuilt from them — patches alone still make a language
|
||||
exist on a page."""
|
||||
with kanta.transaction("refresh translations"):
|
||||
i18n.clear_translations(data)
|
||||
_invalidate_pages()
|
||||
# Fragments rejected this run (segment validation) stay skipped no
|
||||
# longer: a refresh is precisely the "another chance" for them.
|
||||
dispatcher.reset_validation_failures()
|
||||
|
||||
|
||||
class ToggleTaskIn(BaseModel):
|
||||
"""Payload for toggling one task-list checkbox."""
|
||||
|
||||
path: str
|
||||
index: int
|
||||
markdown: str | None = None
|
||||
|
||||
|
||||
@router.post("/_api/toggle-task")
|
||||
async def toggle_task_endpoint(body: ToggleTaskIn) -> dict[str, str]:
|
||||
"""Toggle the Nth task-list checkbox in a page's Markdown source.
|
||||
|
||||
If ``markdown`` is provided the source is left untouched and the toggled
|
||||
Markdown is returned (used while the page editor is open, so the live
|
||||
CodeMirror document can be updated). Otherwise the stored page at
|
||||
``path`` is read, toggled, and saved.
|
||||
"""
|
||||
path = body.path.strip("/")
|
||||
_check_reserved(path)
|
||||
if body.markdown is not None:
|
||||
new_markdown = toggle_task(body.markdown, body.index)
|
||||
if new_markdown is None:
|
||||
raise HTTPException(400, "invalid task index")
|
||||
return {"markdown": new_markdown}
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is None or node.chunks is None:
|
||||
raise HTTPException(404, "no such page")
|
||||
new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
|
||||
if new_markdown is None:
|
||||
raise HTTPException(400, "invalid task index")
|
||||
with kanta.transaction("toggle task", extra=path):
|
||||
# Re-chunk like any save: only the chunk containing the toggled
|
||||
# checkbox gets a new hash, the rest keep theirs.
|
||||
node.chunks = store_chunks(data.chunks, new_markdown)
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
return {"markdown": new_markdown}
|
||||
|
||||
|
||||
# WebSocket API for external translation services (not under /_api: it is keyed
|
||||
# with Data.translate_keys instead of the SSO forward-auth). The dispatcher —
|
||||
# protocol, connected clients and the job pipeline — lives in translate.py.
|
||||
@router.websocket("/_translate/{clientkey}")
|
||||
async def translate_ws(ws: WebSocket, clientkey: str) -> None:
|
||||
"""Translator service channel (docs/localization.md).
|
||||
|
||||
Deliberately NOT under /_api/: the external forward-auth is skipped;
|
||||
the server-generated client key in the path is the access control
|
||||
(``Data.translate_keys``: key -> display name; the first is generated
|
||||
at bootstrap, all are shown in the admin's /_api/settings).
|
||||
"""
|
||||
await dispatcher.handle_ws(ws, clientkey)
|
||||
|
||||
|
||||
@router.websocket("/_api/ws/editor")
|
||||
async def editor_ws(ws: WebSocket) -> None:
|
||||
"""Editor session: open pages, render previews, save — over one socket.
|
||||
|
||||
Stateless protocol (each message carries the path):
|
||||
<- {"type": "open", "path", "lang"?}
|
||||
-> {"type": "doc", "path", "exists", "title", "markdown", "published",
|
||||
"banner", "banner_design", "lang", "primary_lang", "langs",
|
||||
"translate_langs"}
|
||||
<- {"type": "render", "path", "markdown"}
|
||||
-> {"type": "html", "path", "html"}
|
||||
<- {"type": "save", "path", "title"?, "markdown"?, "published"?,
|
||||
"banner"?, "banner_design"?, "move_from"?, "lang"?, "base"?}
|
||||
(absent fields keep their old values; move_from: rename/move a
|
||||
page, subtree included)
|
||||
-> {"type": "saved", "path"} | {"type": "error", "detail"}
|
||||
|
||||
With "lang" (a translation, not the primary language), open returns the
|
||||
effective hybrid Markdown and title for that language plus the language
|
||||
metadata the picker's UI needs; save diffs the submitted Markdown
|
||||
against "base" (the editor's shadow copy of the hybrid it started from
|
||||
— absent: the current hybrid) and stores it as a user Patch, and a
|
||||
changed title becomes a fragment in Data.trans — node.chunks and the
|
||||
other fields stay untouched (docs/localization.md).
|
||||
"""
|
||||
await ws.accept()
|
||||
try:
|
||||
while True:
|
||||
msg = await ws.receive_json()
|
||||
path = msg.get("path", "").strip("/")
|
||||
try:
|
||||
_check_reserved(path)
|
||||
except HTTPException:
|
||||
await ws.send_json({"type": "error", "detail": "reserved path"})
|
||||
continue
|
||||
match msg.get("type"):
|
||||
case "open":
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
# The article's primary language: its own setting,
|
||||
# inherited down the tree ("en" final fallback).
|
||||
node_lang = i18n.primary_lang(data.menu, path)
|
||||
lang = i18n.base_tag(str(msg.get("lang") or ""))
|
||||
if lang == node_lang:
|
||||
lang = ""
|
||||
markdown = ""
|
||||
title = node.title if node else ""
|
||||
if node is not None:
|
||||
markdown = node_markdown(data, node) or ""
|
||||
if lang and node.chunks is not None:
|
||||
# Translation view: the effective (hybrid)
|
||||
# Markdown and title for that language —
|
||||
# machine fragments + user patches over the
|
||||
# original (docs/localization.md editor flow).
|
||||
markdown = i18n.hybrid_markdown(data, node, path, lang)
|
||||
title = i18n.title_map(data, lang).get(path) or title
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "doc",
|
||||
"path": path,
|
||||
"exists": node is not None,
|
||||
"title": title,
|
||||
"markdown": markdown,
|
||||
"published": node.published if node else True,
|
||||
"banner": node.banner if node else "",
|
||||
# Own banner design setting: null = inherit,
|
||||
# "" = none, otherwise a design name.
|
||||
"banner_design": node.banner_design if node else None,
|
||||
# Which node's banner applies here ("" = front page,
|
||||
# null = default artwork); the site editor shows it
|
||||
# as the banner field's placeholder.
|
||||
"banner_from": views.banner_source(data.menu, path),
|
||||
# Which node's banner-design setting would apply on
|
||||
# inherit ("" = front page, null = the active
|
||||
# theme's default) and what design that resolves to.
|
||||
"banner_design_from": (
|
||||
src := views.banner_design_source(
|
||||
data.menu, path, data.theme
|
||||
)
|
||||
),
|
||||
"banner_design_inherited": (
|
||||
views.banner_design(data.menu, src, data.theme)
|
||||
if src is not None
|
||||
else views.theme_banner_design(data.theme)
|
||||
),
|
||||
# Language context for the editor's picker: the
|
||||
# language this Markdown represents ("" = primary),
|
||||
# the page's own primary language, the translations
|
||||
# this page already has, and the site-wide
|
||||
# configured target languages.
|
||||
"lang": lang,
|
||||
"primary_lang": node_lang,
|
||||
"langs": sorted(node.langs) if node else [],
|
||||
"translate_langs": sorted(data.translate_langs),
|
||||
}
|
||||
)
|
||||
case "render":
|
||||
markdown = msg.get("markdown", "")
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
rendered = render(
|
||||
markdown,
|
||||
path,
|
||||
node.created if node else None,
|
||||
node.modified if node else None,
|
||||
# The title is injected as h1 when the markdown has
|
||||
# none; the editor's title field edits live-preview.
|
||||
title=msg.get("title") or (node.title if node else ""),
|
||||
)
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "html",
|
||||
"path": path,
|
||||
"html": rendered.html,
|
||||
# Column-layout flag: the preview toggles the
|
||||
# article's .multicol class and swaps in the
|
||||
# segmented (.colseg/.cols) article html.
|
||||
"multicol": rendered.multicol,
|
||||
}
|
||||
)
|
||||
case "save":
|
||||
move_from = (msg.get("move_from") or path).strip("/")
|
||||
lang = i18n.base_tag(str(msg.get("lang") or ""))
|
||||
translated = bool(
|
||||
lang and lang != i18n.primary_lang(data.menu, move_from)
|
||||
)
|
||||
try:
|
||||
_check_reserved(move_from)
|
||||
except HTTPException:
|
||||
await ws.send_json({"type": "error", "detail": "reserved path"})
|
||||
continue
|
||||
old_chain = resolve(data.menu, move_from)
|
||||
old = old_chain[-1] if old_chain else None
|
||||
if old is None and move_from != path:
|
||||
move_from = path # nothing to carry over; plain save
|
||||
if move_from != path:
|
||||
# Rename/move: detach the node (subtree included)
|
||||
# and attach it at the new path. The target slug
|
||||
# must be free and the front page childless.
|
||||
if move_from and path.startswith(f"{move_from}/"):
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "cannot move a page under itself",
|
||||
}
|
||||
)
|
||||
continue
|
||||
tslug = path.rpartition("/")[2]
|
||||
if not tslug and old.children:
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "the front page cannot have children",
|
||||
}
|
||||
)
|
||||
continue
|
||||
tchain = resolve(data.menu, path)
|
||||
if tchain is not None:
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "target path exists",
|
||||
}
|
||||
)
|
||||
continue
|
||||
if translated and (
|
||||
move_from != path or old is None or old.chunks is None
|
||||
):
|
||||
# A translated-view save patches an existing
|
||||
# original; it cannot create or move pages.
|
||||
await ws.send_json({"type": "error", "detail": "no such page"})
|
||||
continue
|
||||
if translated and "markdown" in msg and not msg["markdown"].strip():
|
||||
# Saving never deletes; an emptied translation would
|
||||
# render as a blank page in that language.
|
||||
await ws.send_json(
|
||||
{
|
||||
"type": "error",
|
||||
"detail": "a translation cannot be emptied",
|
||||
}
|
||||
)
|
||||
continue
|
||||
with kanta.transaction("editor save", extra=path):
|
||||
if move_from != path:
|
||||
same_menu = (
|
||||
move_from.rpartition("/")[0] == path.rpartition("/")[0]
|
||||
)
|
||||
snodes, sslug = find_slot(data.menu, move_from)
|
||||
node = snodes.pop(sslug)
|
||||
parent = path.rpartition("/")[0]
|
||||
if parent:
|
||||
_ensure(data.menu, parent)
|
||||
tnodes, tslug = find_slot(data.menu, path)
|
||||
node.order = (
|
||||
node.order if same_menu else append_order(tnodes)
|
||||
)
|
||||
tnodes[tslug] = node
|
||||
else:
|
||||
node = old if old is not None else _ensure(data.menu, path)
|
||||
if translated:
|
||||
# node.chunks and the original-language fields
|
||||
# stay untouched: the markdown diff (against the
|
||||
# editor's shadow "base" — the hybrid it started
|
||||
# from; absent: the current hybrid) is appended
|
||||
# as a Patch, a changed title becomes a
|
||||
# per-language title override (i18n).
|
||||
changed = False
|
||||
if "markdown" in msg:
|
||||
base = msg.get("base")
|
||||
changed = i18n.add_patch(
|
||||
data,
|
||||
node,
|
||||
path,
|
||||
lang,
|
||||
msg["markdown"],
|
||||
base=base if isinstance(base, str) else None,
|
||||
)
|
||||
if "title" in msg and node.title:
|
||||
changed = (
|
||||
i18n.set_title_translation(
|
||||
data, node, lang, msg["title"]
|
||||
)
|
||||
or changed
|
||||
)
|
||||
if changed:
|
||||
_invalidate_pages()
|
||||
else:
|
||||
if "markdown" in msg:
|
||||
# Saving never deletes; empty markdown is an
|
||||
# empty page. Deletion is an explicit choice
|
||||
# by the page editor (REST DELETE).
|
||||
node.chunks = store_chunks(data.chunks, msg["markdown"])
|
||||
if "title" in msg:
|
||||
node.title = msg["title"]
|
||||
if "published" in msg:
|
||||
node.published = bool(msg["published"])
|
||||
if "banner" in msg:
|
||||
node.banner = msg["banner"]
|
||||
if "banner_design" in msg:
|
||||
node.banner_design = msg["banner_design"]
|
||||
node.modified = datetime.now(UTC)
|
||||
_invalidate_pages()
|
||||
await ws.send_json({"type": "saved", "path": path})
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
+50
-1789
File diff suppressed because it is too large
Load Diff
+4
-2
@@ -23,8 +23,10 @@ _FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
|
||||
#: end at the first blank line, which the generic blank-line split
|
||||
#: already does.
|
||||
_HTML_ATOMIC = (
|
||||
(re.compile(r"^ {0,3}<(?:script|pre|style|textarea)(?:\s|>|$)", re.I),
|
||||
re.compile(r"</(?:script|pre|style|textarea)\s*>", re.I)),
|
||||
(
|
||||
re.compile(r"^ {0,3}<(?:script|pre|style|textarea)(?:\s|>|$)", re.I),
|
||||
re.compile(r"</(?:script|pre|style|textarea)\s*>", re.I),
|
||||
),
|
||||
(re.compile(r"^ {0,3}<!--"), re.compile(r"-->")),
|
||||
(re.compile(r"^ {0,3}<\?"), re.compile(r"\?>")),
|
||||
(re.compile(r"^ {0,3}<!\[CDATA\["), re.compile(r"\]\]>")),
|
||||
|
||||
@@ -0,0 +1,375 @@
|
||||
"""Content-addressed file store, image derivatives, and file routes.
|
||||
|
||||
``FileStore`` keeps uploads, seed assets and fetched favicons on disk under
|
||||
hash-prefixed names, fully cached in RAM (uncompressed plus a zstd copy
|
||||
when compression shrinks the body), served immutable at ``/_f/``. Raster
|
||||
images and SVGs are recompressed into AVIF/WebP/JPEG derivatives
|
||||
(``store_image`` and helpers); the untouched original is kept alongside as
|
||||
``<hash>.orig<ext>`` (never served). Routes: upload/delete under
|
||||
``/_api/files``, the favicon settings endpoints, the ``/_f/`` server with
|
||||
Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import mimetypes
|
||||
import tempfile
|
||||
from contextlib import suppress
|
||||
from pathlib import Path
|
||||
|
||||
import blake3
|
||||
from fastapi import APIRouter, HTTPException, Request
|
||||
from fastapi.responses import Response
|
||||
from mediapreview import dispatch
|
||||
|
||||
from pagerite import views
|
||||
from pagerite.state import (
|
||||
FAVICON_MAXSIZE,
|
||||
FILES_DIR,
|
||||
IMAGE_JPG_QUALITY,
|
||||
IMAGE_MAXSIZE,
|
||||
IMAGE_QUALITY,
|
||||
IMAGE_WEBP_QUALITY,
|
||||
_invalidate_pages,
|
||||
_zstd,
|
||||
data,
|
||||
kanta,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# mediapreview logs pyvips noise ("VipsForeignSaveJpegTarget argument strip is
|
||||
# deprecated", "threadpool completed with N workers") at INFO; keep warnings.
|
||||
logging.getLogger("mediapreview").setLevel(logging.WARNING)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
class FileStore:
|
||||
"""Content-addressed files on disk, fully cached in RAM.
|
||||
|
||||
Every file is kept in RAM uncompressed and zstd-compressed (the
|
||||
compressed copy only when it actually shrinks the body), so ``/_f``
|
||||
serves both encodings without touching disk or re-compressing.
|
||||
"""
|
||||
|
||||
def __init__(self, path: Path) -> None:
|
||||
self.path = path
|
||||
#: name -> (uncompressed body, zstd body or None)
|
||||
self._cache: dict[str, tuple[bytes, bytes | None]] = {}
|
||||
|
||||
@staticmethod
|
||||
def _entry(body: bytes) -> tuple[bytes, bytes | None]:
|
||||
compressed = _zstd.compress(body)
|
||||
return body, compressed if len(compressed) < len(body) else None
|
||||
|
||||
def load(self) -> None:
|
||||
"""Read every stored file into the RAM cache (startup)."""
|
||||
try:
|
||||
entries = sorted(self.path.iterdir())
|
||||
except FileNotFoundError:
|
||||
return
|
||||
for f in entries:
|
||||
if f.is_file() and not f.name.startswith("."):
|
||||
self._cache.setdefault(f.name, self._entry(f.read_bytes()))
|
||||
|
||||
def get(self, name: str) -> tuple[bytes, bytes | None] | None:
|
||||
return self._cache.get(name)
|
||||
|
||||
def put(self, name: str, body: bytes) -> None:
|
||||
"""Store ``body`` under ``name`` on disk and in the RAM cache."""
|
||||
if name in self._cache:
|
||||
return
|
||||
self.path.mkdir(parents=True, exist_ok=True)
|
||||
(self.path / name).write_bytes(body)
|
||||
self._cache[name] = self._entry(body)
|
||||
|
||||
def delete(self, name: str) -> None:
|
||||
"""Delete a file plus its derivatives/original counterparts, if any.
|
||||
|
||||
An image upload is stored as a group sharing the hash prefix
|
||||
(``<hash>.orig.<ext>`` + ``<hash>.avif/.webp/.jpg``); deleting any
|
||||
of the names removes them all.
|
||||
"""
|
||||
stem = name.partition(".")[0]
|
||||
for key in [k for k in self._cache if k.partition(".")[0] == stem]:
|
||||
self._cache.pop(key, None)
|
||||
with suppress(FileNotFoundError):
|
||||
(self.path / key).unlink()
|
||||
|
||||
def __contains__(self, name: str) -> bool:
|
||||
return name in self._cache
|
||||
|
||||
|
||||
file_store = FileStore(FILES_DIR)
|
||||
|
||||
|
||||
def _ext(orig: str) -> str:
|
||||
"""Sanitized lowercase extension (with dot) of an original file name."""
|
||||
return "".join(c for c in Path(orig).suffix.lower() if c.isalnum() or c == ".")
|
||||
|
||||
|
||||
def _hash_name(body: bytes, orig: str) -> str:
|
||||
"""Content-addressed file name: blake3 hash prefix + original extension."""
|
||||
return blake3.blake3(body).hexdigest()[:12] + _ext(orig)
|
||||
|
||||
|
||||
def _to_avif(body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE) -> bytes | None:
|
||||
"""Recompress an image body to a thumbnailed AVIF via mediapreview's
|
||||
dispatch (pyvips for common formats, ffmpeg for HEIC/HEIF/AVIF), or
|
||||
None if the body is not a decodable image (stored as-is by the caller).
|
||||
Dispatch needs a real file for format routing, so the body goes
|
||||
through a temp file.
|
||||
"""
|
||||
with tempfile.NamedTemporaryFile(suffix=ext) as tmp:
|
||||
tmp.write(body)
|
||||
tmp.flush()
|
||||
try:
|
||||
avif, _resp = dispatch(
|
||||
Path(tmp.name),
|
||||
quality=IMAGE_QUALITY,
|
||||
maxsize=maxsize,
|
||||
maxzoom=1,
|
||||
)
|
||||
except Exception:
|
||||
return None
|
||||
return avif
|
||||
|
||||
|
||||
def _svg_to_png(body: bytes, maxsize: int) -> bytes | None:
|
||||
"""Rasterize an SVG to PNG via pyvips, scaled so the long side is
|
||||
``maxsize`` — SVGs often carry no meaningful intrinsic resolution, so
|
||||
we rasterize at full image size rather than the tiny nominal one."""
|
||||
import pyvips
|
||||
|
||||
try:
|
||||
img = pyvips.Image.new_from_buffer(body, "")
|
||||
scale = (
|
||||
maxsize / max(img.width, img.height)
|
||||
if img.width and img.height
|
||||
else maxsize
|
||||
)
|
||||
if scale != 1:
|
||||
img = pyvips.Image.new_from_buffer(body, "", scale=scale)
|
||||
return img.write_to_buffer(".png")
|
||||
except pyvips.Error:
|
||||
return None
|
||||
|
||||
|
||||
def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes:
|
||||
"""Re-encode the AVIF derivative into a fallback format (WebP/JPEG)
|
||||
via pyvips. JPEG has no alpha, so it is flattened onto white;
|
||||
``strip`` keeps metadata (EXIF) out of the fallbacks."""
|
||||
import pyvips
|
||||
|
||||
img = pyvips.Image.new_from_buffer(avif, "")
|
||||
if suffix == ".jpg" and img.hasalpha():
|
||||
img = img.flatten(background=[255, 255, 255])
|
||||
return img.write_to_buffer(suffix, Q=quality, strip=True)
|
||||
|
||||
|
||||
def _image_derivatives(
|
||||
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE
|
||||
) -> dict[str, bytes] | None:
|
||||
"""The served variants of an uploaded image: ``avif`` (primary,
|
||||
thumbnailed to ``maxsize``) plus ``webp`` and ``jpg`` fallbacks
|
||||
re-encoded from it. SVGs are rasterized first (they are vector, so
|
||||
the raster replaces nothing — the .svg itself stays servable).
|
||||
Returns None for non-decodable content (stored as-is by the caller).
|
||||
"""
|
||||
if ext == ".svg":
|
||||
png = _svg_to_png(body, maxsize)
|
||||
if png is None:
|
||||
return None
|
||||
body, ext = png, ".png"
|
||||
avif = _to_avif(body, ext, maxsize)
|
||||
if avif is None:
|
||||
return None
|
||||
return {
|
||||
"avif": avif,
|
||||
"webp": _avif_to_format(avif, ".webp", IMAGE_WEBP_QUALITY),
|
||||
"jpg": _avif_to_format(avif, ".jpg", IMAGE_JPG_QUALITY),
|
||||
}
|
||||
|
||||
|
||||
def store_image(
|
||||
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE, *, derive: bool = True
|
||||
) -> str:
|
||||
"""Store an image body content-addressed and return its file name.
|
||||
|
||||
Decodable images get AVIF/WebP/JPEG derivatives thumbnailed to
|
||||
``maxsize``; the original is kept as ``<hash>.orig<ext>`` (SVG
|
||||
originals as ``<hash>.svg``, still servable) and the bare ``<hash>``
|
||||
name is returned (the server negotiates the format by Accept header).
|
||||
Anything else — undecodable content, or ``derive=False`` (GIFs, whose
|
||||
animation recompression would lose) — is stored as-is and returned with
|
||||
its extension. Blocking (pyvips/ffmpeg); call via ``asyncio.to_thread``
|
||||
from async code.
|
||||
"""
|
||||
digest = blake3.blake3(body).hexdigest()[:12]
|
||||
derivatives = _image_derivatives(body, ext, maxsize) if derive else None
|
||||
if derivatives is None: # store the body as-is
|
||||
file_store.put(digest + ext, body)
|
||||
return digest + ext
|
||||
file_store.put(f"{digest}.svg" if ext == ".svg" else f"{digest}.orig{ext}", body)
|
||||
for fmt, variant in derivatives.items():
|
||||
file_store.put(f"{digest}.{fmt}", variant)
|
||||
return digest
|
||||
|
||||
|
||||
@router.put("/_api/files/{name}")
|
||||
async def upload_file(name: str, request: Request) -> dict[str, str]:
|
||||
"""Store an upload (image, video...) in the content-addressed store.
|
||||
|
||||
The stored name is a blake3 hash prefix + the original extension,
|
||||
served immutable at "/_f/{name}"; returns {"path": "/_f/..."}.
|
||||
|
||||
Raster images and SVGs are recompressed (SVGs rasterized) into AVIF
|
||||
(primary) plus WebP and JPEG fallbacks: the original goes to
|
||||
``<hash>.orig<ext>`` (kept for reprocessing, never served — it may
|
||||
carry EXIF data; SVG originals stay servable as ``<hash>.svg`` since
|
||||
vector carries no EXIF) and pages link the bare ``/_f/<hash>``, the
|
||||
server picking the format from the request's Accept header. GIFs are
|
||||
stored as-is (animation would be lost), as is other non-decodable
|
||||
content.
|
||||
"""
|
||||
if "/" in name or name in {".", ".."}:
|
||||
raise HTTPException(400, "bad file name")
|
||||
body = await request.body()
|
||||
if not body:
|
||||
raise HTTPException(400, "empty file")
|
||||
ext = _ext(name)
|
||||
stored = await asyncio.to_thread(store_image, body, ext, derive=ext != ".gif")
|
||||
return {"path": f"/_f/{stored}"}
|
||||
|
||||
|
||||
@router.delete("/_api/files/{name}", status_code=204)
|
||||
async def delete_file(name: str) -> None:
|
||||
"""Remove a file from the content-addressed store (no refcounting:
|
||||
other pages referencing the same content will 404)."""
|
||||
if name not in file_store:
|
||||
raise HTTPException(404, "no such file")
|
||||
file_store.delete(name)
|
||||
|
||||
|
||||
@router.put("/_api/settings/favicon")
|
||||
async def put_favicon(request: Request) -> dict[str, str]:
|
||||
"""Upload a favicon into the content-addressed store and activate it.
|
||||
|
||||
Raw image body (ico/png/svg...). Decodable images are thumbnailed to
|
||||
FAVICON_MAXSIZE (192px — browsers scale down from there themselves)
|
||||
and stored as AVIF/WebP/JPEG derivatives linked extension-less; SVG
|
||||
originals also stay servable under their ``.svg`` name. Undecodable
|
||||
bodies are stored as-is. Pages link it as <link rel="icon">. Returns
|
||||
{"path": "/_f/..."}.
|
||||
"""
|
||||
body = await request.body()
|
||||
if not body:
|
||||
raise HTTPException(400, "empty file")
|
||||
ext = _ext(request.headers.get("x-filename", "favicon.ico"))
|
||||
stored = await asyncio.to_thread(store_image, body, ext, FAVICON_MAXSIZE)
|
||||
with kanta.transaction("upload favicon"):
|
||||
data.favicon = stored
|
||||
_invalidate_pages()
|
||||
return {"path": f"/_f/{stored}"}
|
||||
|
||||
|
||||
@router.delete("/_api/settings/favicon", status_code=204)
|
||||
async def delete_favicon() -> None:
|
||||
"""Clear the custom favicon (back to the build's /favicon.ico).
|
||||
|
||||
The blob stays in the content-addressed store; only the reference goes.
|
||||
"""
|
||||
with kanta.transaction("clear favicon"):
|
||||
data.favicon = ""
|
||||
_invalidate_pages()
|
||||
|
||||
|
||||
async def _serve_user_file(path: Path | None, request: Request) -> Response:
|
||||
"""Serve a user-asset file resolved on disk, with mtime etag.
|
||||
|
||||
Read from disk on every request (etag by mtime+size): user assets are
|
||||
never built or content-hashed, so edits on disk show on the next page
|
||||
load, in prod as well as dev.
|
||||
"""
|
||||
if path is None:
|
||||
raise HTTPException(404)
|
||||
stat = path.stat()
|
||||
etag = f'"{stat.st_mtime_ns:x}-{stat.st_size:x}"'
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
|
||||
return Response(
|
||||
path.read_bytes(),
|
||||
media_type=mime,
|
||||
headers={"etag": etag, "cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/_themes/{name}/{filename}")
|
||||
async def theme_file(name: str, filename: str, request: Request) -> Response:
|
||||
"""Serve a theme/banner-design file, resolved across views.THEME_DIRS.
|
||||
|
||||
Stylesheets plus any extra assets the CSS references (like summer's
|
||||
grass.svg).
|
||||
"""
|
||||
return await _serve_user_file(views.theme_file(name, filename), request)
|
||||
|
||||
|
||||
@router.get("/_fonts/{name}/{filename}")
|
||||
async def user_font_file(name: str, filename: str, request: Request) -> Response:
|
||||
"""Serve a user font file, resolved across views.FONT_DIRS.
|
||||
|
||||
The folder's font.css (@font-face rules + --font-{name} stack variable)
|
||||
is linked on every page; the woff2 files it references come from here.
|
||||
"""
|
||||
return await _serve_user_file(views.font_file(name, filename), request)
|
||||
|
||||
|
||||
@router.get("/_f/{name}")
|
||||
async def stored_file(name: str, request: Request) -> Response:
|
||||
"""Serve a file from the content-addressed store (immutable: the name
|
||||
is its own hash, so cache forever). Bodies are served from the RAM
|
||||
cache, zstd-compressed when the client accepts it and compression
|
||||
actually shrank the file.
|
||||
|
||||
A bare ``/_f/{hash}`` (no extension, how pages link uploaded images)
|
||||
content-negotiates between the stored derivatives: a format is served
|
||||
only when the Accept header lists it explicitly — ``image/avif`` →
|
||||
AVIF, ``image/webp`` → WebP, anything else (including ``image/*`` and
|
||||
``*/*``) → JPEG. An explicit extension pins the format. ``.orig.``
|
||||
originals are internal (they may carry EXIF data) and never served."""
|
||||
if ".orig." in name:
|
||||
raise HTTPException(404)
|
||||
etag = name
|
||||
vary = ""
|
||||
entry = file_store.get(name)
|
||||
if entry is None and "." not in name:
|
||||
# Extension-less image link: negotiate avif/webp/jpg by Accept.
|
||||
vary = "accept"
|
||||
accept = request.headers.get("accept", "")
|
||||
if "image/avif" in accept:
|
||||
order = ("avif", "webp", "jpg")
|
||||
elif "image/webp" in accept:
|
||||
order = ("webp", "jpg", "avif")
|
||||
else:
|
||||
order = ("jpg", "webp", "avif")
|
||||
for ext in order:
|
||||
etag = f"{name}.{ext}"
|
||||
entry = file_store.get(etag)
|
||||
if entry is not None:
|
||||
break
|
||||
if entry is None:
|
||||
raise HTTPException(404)
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
body, compressed = entry
|
||||
headers = {"etag": etag, "cache-control": "public, max-age=31536000, immutable"}
|
||||
if compressed is not None and "zstd" in request.headers.get("accept-encoding", ""):
|
||||
headers["content-encoding"] = "zstd"
|
||||
vary = f"{vary}, accept-encoding".lstrip(", ")
|
||||
body = compressed
|
||||
if vary:
|
||||
headers["vary"] = vary
|
||||
mime = mimetypes.guess_type(etag)[0] or "application/octet-stream"
|
||||
return Response(body, media_type=mime, headers=headers)
|
||||
+14
-8
@@ -128,7 +128,9 @@ def make_patch(base: str, edited: str) -> Patch:
|
||||
"""
|
||||
a, b = chunk_markdown(base), chunk_markdown(edited)
|
||||
hunks: list[tuple[str, str]] = []
|
||||
for tag, i1, i2, j1, j2 in SequenceMatcher(None, a, b, autojunk=False).get_opcodes():
|
||||
for tag, i1, i2, j1, j2 in SequenceMatcher(
|
||||
None, a, b, autojunk=False
|
||||
).get_opcodes():
|
||||
if tag == "equal":
|
||||
continue
|
||||
search = "\n\n".join(a[i1:i2])
|
||||
@@ -155,12 +157,14 @@ def hybrid_markdown(data: Data, node: Node, path: str, lang: str) -> str:
|
||||
Not gated on ``node.langs`` (get_translation is the gated view): the
|
||||
editor save path diffs against this even for a language's first patch.
|
||||
"""
|
||||
hybrid = join_chunks([
|
||||
data.chunks.get(h, "")
|
||||
if h in node.no_trans
|
||||
else data.trans.get(h, {}).get(lang) or data.chunks.get(h, "")
|
||||
for h in node.chunks or []
|
||||
])
|
||||
hybrid = join_chunks(
|
||||
[
|
||||
data.chunks.get(h, "")
|
||||
if h in node.no_trans
|
||||
else data.trans.get(h, {}).get(lang) or data.chunks.get(h, "")
|
||||
for h in node.chunks or []
|
||||
]
|
||||
)
|
||||
for patch in data.patches.get(f"{path}:{lang}", []):
|
||||
hybrid = apply_patch(hybrid, patch)
|
||||
return hybrid
|
||||
@@ -175,7 +179,9 @@ def add_patch(
|
||||
translated version exist, so ``node.langs`` is set. Returns True when
|
||||
a patch was stored. Pure data ops — the caller wraps in a transaction
|
||||
and invalidates."""
|
||||
patch = make_patch(base if base is not None else hybrid_markdown(data, node, path, lang), edited)
|
||||
patch = make_patch(
|
||||
base if base is not None else hybrid_markdown(data, node, path, lang), edited
|
||||
)
|
||||
if not patch.hunks:
|
||||
return False
|
||||
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
|
||||
|
||||
@@ -422,7 +422,10 @@ def _heading_ids(state) -> None:
|
||||
heads = [
|
||||
(i, token)
|
||||
for i, token in enumerate(tokens)
|
||||
if token.type == "heading_open" and token.tag in ("h1", "h2") and token.level == 0 and i != first_h1
|
||||
if token.type == "heading_open"
|
||||
and token.tag in ("h1", "h2")
|
||||
and token.level == 0
|
||||
and i != first_h1
|
||||
]
|
||||
if len(heads) < ANCHOR_MIN_HEADINGS:
|
||||
return
|
||||
@@ -472,7 +475,9 @@ def make_md(*, verbatim: bool = False) -> MarkdownIt:
|
||||
.use(deflist_plugin)
|
||||
# label wrapping (render) puts the item text inside the checkbox
|
||||
# <label> html_inline; without it the text stays a plain token.
|
||||
.use(tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim)
|
||||
.use(
|
||||
tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim
|
||||
)
|
||||
.use(gfm_autolink_plugin)
|
||||
.use(sub_plugin)
|
||||
.use(superscript_plugin)
|
||||
|
||||
+19
-12
@@ -6,9 +6,9 @@ omitted) before it is decoded into ``Data`` structs, and runs exactly once
|
||||
per database based on its recorded version.
|
||||
|
||||
All storage/schema upgrades live here — including on-disk file work, which
|
||||
runs through app.py's file store (imported lazily: app.py owns the store
|
||||
and passes this module to Kanta; at migration time, during lifespan
|
||||
``kanta.open()``, the app module is fully loaded).
|
||||
runs through files.py's file store (imported lazily: files.py owns the store
|
||||
and state.py passes this module to Kanta; at migration time, during lifespan
|
||||
``kanta.open()``, both modules are fully loaded).
|
||||
"""
|
||||
|
||||
import base64
|
||||
@@ -25,7 +25,7 @@ def _append_order(nodes: dict) -> float:
|
||||
|
||||
|
||||
def _ensure(menu: dict, path: str) -> dict:
|
||||
"""Raw-dict equivalent of app._ensure: the node dict at ``path``,
|
||||
"""Raw-dict equivalent of state._ensure: the node dict at ``path``,
|
||||
creating it and any missing ancestors (content-less category labels)
|
||||
appended at the end of their level."""
|
||||
nodes = menu
|
||||
@@ -44,7 +44,7 @@ def migrate_v1(d: dict) -> None:
|
||||
and rebuild the legacy flat page store (``pages``) as the menu tree."""
|
||||
files = d.pop("files", None)
|
||||
if files:
|
||||
from pagerite.app import file_store
|
||||
from pagerite.files import file_store
|
||||
|
||||
for name, body in files.items():
|
||||
if isinstance(body, str): # JSON-level bytes are base64 strings
|
||||
@@ -74,9 +74,16 @@ def _backfill_derivatives() -> None:
|
||||
AVIF, and SVGs no raster variants at all). WebP/JPEG are re-encoded
|
||||
from an existing AVIF when available, everything else from the
|
||||
original (SVGs rasterized first)."""
|
||||
from pagerite import app
|
||||
from pagerite.files import (
|
||||
IMAGE_MAXSIZE,
|
||||
IMAGE_WEBP_QUALITY,
|
||||
IMAGE_JPG_QUALITY,
|
||||
_avif_to_format,
|
||||
_svg_to_png,
|
||||
_to_avif,
|
||||
file_store,
|
||||
)
|
||||
|
||||
file_store = app.file_store
|
||||
try:
|
||||
paths = [f for f in file_store.path.iterdir() if f.is_file()]
|
||||
except FileNotFoundError:
|
||||
@@ -96,22 +103,22 @@ def _backfill_derivatives() -> None:
|
||||
ext = source.suffix
|
||||
body = source.read_bytes()
|
||||
if ext == ".svg":
|
||||
png = app._svg_to_png(body, app.IMAGE_MAXSIZE)
|
||||
png = _svg_to_png(body, IMAGE_MAXSIZE)
|
||||
if png is None:
|
||||
continue
|
||||
body, ext = png, ".png"
|
||||
converted = app._to_avif(body, ext)
|
||||
converted = _to_avif(body, ext)
|
||||
if converted is None:
|
||||
continue
|
||||
file_store.put(f"{digest}.avif", converted)
|
||||
avif = file_store.get(f"{digest}.avif")
|
||||
for fmt, quality in (
|
||||
("webp", app.IMAGE_WEBP_QUALITY),
|
||||
("jpg", app.IMAGE_JPG_QUALITY),
|
||||
("webp", IMAGE_WEBP_QUALITY),
|
||||
("jpg", IMAGE_JPG_QUALITY),
|
||||
):
|
||||
if f"{digest}.{fmt}" not in names:
|
||||
file_store.put(
|
||||
f"{digest}.{fmt}", app._avif_to_format(avif[0], f".{fmt}", quality)
|
||||
f"{digest}.{fmt}", _avif_to_format(avif[0], f".{fmt}", quality)
|
||||
)
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,253 @@
|
||||
"""Public content pages: front page, sitemap, robots, and the catch-all.
|
||||
|
||||
``GET /{path:path}`` resolves a slug path against the menu tree and renders
|
||||
the page (or a category placeholder, or 404); it must be registered AFTER
|
||||
the fastapi-vue asset routes so built frontend files win over content slugs
|
||||
(see app.py). Requests are recorded in analytics (crawler hits and 404s
|
||||
here, visits via the /_ws socket in tracking.py).
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from email.utils import format_datetime
|
||||
from xml.sax.saxutils import escape as xml_escape
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Request
|
||||
from fastapi.responses import RedirectResponse, Response
|
||||
|
||||
from pagerite import i18n, state
|
||||
from pagerite.data import Node, resolve, sorted_nodes
|
||||
from pagerite.state import (
|
||||
SITE_URL,
|
||||
_html_response,
|
||||
_is_reserved,
|
||||
analytics_store,
|
||||
data,
|
||||
)
|
||||
from pagerite.tracking import (
|
||||
_client_ip,
|
||||
_enrich_client,
|
||||
_query_suffix,
|
||||
_schedule_client_enrichment,
|
||||
_track_entry,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
def _http_date(dt: datetime) -> str:
|
||||
"""RFC 7231 date for the Last-Modified header."""
|
||||
return format_datetime(dt.astimezone(UTC), usegmt=True)
|
||||
|
||||
|
||||
def _is_trackable_path(path: str) -> bool:
|
||||
"""Content URLs only: skip auth endpoints and reserved/machinery paths."""
|
||||
if not path:
|
||||
return True
|
||||
if path == "auth" or path.startswith("auth/"):
|
||||
return False
|
||||
return not _is_reserved(path)
|
||||
|
||||
|
||||
@router.get("/")
|
||||
async def front_page(request: Request) -> Response:
|
||||
"""Render the front page (slug path "")."""
|
||||
return await show_page(request, "")
|
||||
|
||||
|
||||
@router.get("/sitemap.xml")
|
||||
async def sitemap(request: Request) -> Response:
|
||||
"""Dynamically generate a sitemap of all published article pages."""
|
||||
base = SITE_URL or str(request.base_url).rstrip("/")
|
||||
entries: list[tuple[str, datetime, int]] = []
|
||||
|
||||
def walk(
|
||||
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
|
||||
) -> None:
|
||||
first_content_slug = next(
|
||||
(
|
||||
slug
|
||||
for slug, node in sorted_nodes(nodes)
|
||||
if node.published and node.chunks is not None
|
||||
),
|
||||
None,
|
||||
)
|
||||
for slug, node in sorted_nodes(nodes):
|
||||
path = f"{prefix}/{slug}" if prefix else slug
|
||||
depth = path.count("/") if path else 0
|
||||
if (
|
||||
not parent_has_content
|
||||
and slug == first_content_slug
|
||||
and node.published
|
||||
and node.chunks is not None
|
||||
and depth > 0
|
||||
):
|
||||
depth -= 1
|
||||
if node.published and node.chunks is not None:
|
||||
entries.append((path, node.modified, depth))
|
||||
if node.children:
|
||||
walk(node.children, path, node.chunks is not None)
|
||||
|
||||
walk(data.menu, "")
|
||||
|
||||
def priority(depth: int) -> float:
|
||||
return max(0.1, 1.0 - depth * 0.2)
|
||||
|
||||
lines = [
|
||||
'<?xml version="1.0" encoding="UTF-8"?>',
|
||||
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
|
||||
]
|
||||
for path, modified, depth in entries:
|
||||
loc = xml_escape(f"{base}/{path}" if path else base)
|
||||
lastmod = (
|
||||
modified.astimezone(UTC)
|
||||
.replace(microsecond=0)
|
||||
.isoformat()
|
||||
.replace("+00:00", "Z")
|
||||
)
|
||||
lines.append(
|
||||
f" <url>"
|
||||
f"<loc>{loc}</loc>"
|
||||
f"<lastmod>{lastmod}</lastmod>"
|
||||
f"<priority>{priority(depth):.1f}</priority>"
|
||||
f"</url>"
|
||||
)
|
||||
lines.append("</urlset>")
|
||||
|
||||
return Response(
|
||||
"\n".join(lines),
|
||||
media_type="application/xml",
|
||||
headers={"cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/robots.txt")
|
||||
async def robots_txt(request: Request) -> Response:
|
||||
"""Allow content crawling, keep the SSO login (/auth/) and the
|
||||
admin-gated API (/_api) out of search results, and point crawlers at
|
||||
the sitemap."""
|
||||
base = SITE_URL or str(request.base_url).rstrip("/")
|
||||
body = f"User-agent: *\nAllow: /\nDisallow: /auth/\nDisallow: /_api\nSitemap: {base}/sitemap.xml\n"
|
||||
return Response(
|
||||
body,
|
||||
media_type="text/plain",
|
||||
headers={"cache-control": "no-cache"},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/{path:path}", response_model=None)
|
||||
async def show_page(request: Request, path: str) -> Response:
|
||||
"""Render the content page at a slug path, or 404.
|
||||
|
||||
A node without content is a category label: its URL renders a
|
||||
placeholder page (nav links point straight at its first child).
|
||||
"""
|
||||
path = path.strip("/")
|
||||
ua = request.headers.get("user-agent", "")
|
||||
accept_language = request.headers.get("accept-language", "")
|
||||
if path and _is_reserved(path):
|
||||
# Invalid slug shape: not a content URL, let FastAPI return its
|
||||
# built-in 404 instead of rendering an editable article page.
|
||||
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
|
||||
# as abuse in analytics.
|
||||
client_hash = analytics_store.track_404(
|
||||
_client_ip(request),
|
||||
ua,
|
||||
f"/{path}{_query_suffix(request)}",
|
||||
accept_language,
|
||||
)
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
raise HTTPException(404)
|
||||
chain = resolve(data.menu, path)
|
||||
node = chain[-1] if chain else None
|
||||
if node is not None and node.published and node.chunks is not None:
|
||||
# Language selection (docs/localization.md): ?lang= wins when a
|
||||
# translation exists, else header logic. Analytics keep the raw
|
||||
# Accept-Language header regardless of the selection.
|
||||
query_lang = request.query_params.get("lang")
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda tag: tag in node.langs,
|
||||
original=i18n.primary_lang(data.menu, path),
|
||||
)
|
||||
# A ?lang= override is replicated onto the page's navigation links
|
||||
# (link_lang), so clicks and prefetches stay in the chosen language.
|
||||
# Query and header-selected renders of the same language differ in
|
||||
# their links, so link_lang is part of the ETag and body cache key.
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
# no-cache forbids serving a stored page without revalidation
|
||||
# (browsers would otherwise cache heuristically and serve stale
|
||||
# pages, e.g. after a theme change). In-session speed instead comes
|
||||
# from pagerite.js's in-memory page cache (preload everything, never
|
||||
# fetch on navigation); the ETag just makes those one-time preload
|
||||
# fetches and any revalidation cheap.
|
||||
etag = f'"{path}@{node.modified.timestamp()}g{state._render_gen}l{lang}q{link_lang}"'
|
||||
if request.headers.get("if-none-match") == etag:
|
||||
return Response(status_code=304)
|
||||
if _is_trackable_path(path):
|
||||
flushed = _track_entry(path, request)
|
||||
_schedule_client_enrichment(flushed)
|
||||
return _html_response(
|
||||
request,
|
||||
"page",
|
||||
path,
|
||||
headers={
|
||||
"etag": etag,
|
||||
"last-modified": _http_date(node.modified),
|
||||
"cache-control": "no-cache",
|
||||
},
|
||||
lang=lang,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if node is not None and node.published and node.chunks is None:
|
||||
# Category label without a landing page: placeholder with the pen
|
||||
# to create it (404 — no page here, but the node is real).
|
||||
# Language selection as on content pages, but over the whole
|
||||
# subtree's availability: the category has no chunks of its own —
|
||||
# its heading, the navigation and the cards' text localize from
|
||||
# the title map and the target articles' translations.
|
||||
query_lang = request.query_params.get("lang")
|
||||
subtree_langs = i18n.subtree_languages(node)
|
||||
lang = i18n.select_language(
|
||||
query_lang,
|
||||
accept_language,
|
||||
lambda tag: tag in subtree_langs,
|
||||
original=i18n.primary_lang(data.menu, path),
|
||||
)
|
||||
link_lang = i18n.base_tag(query_lang or "")
|
||||
if _is_trackable_path(path):
|
||||
flushed = _track_entry(path, request, status=404)
|
||||
_schedule_client_enrichment(flushed)
|
||||
return _html_response(
|
||||
request,
|
||||
"category",
|
||||
path,
|
||||
404,
|
||||
headers={
|
||||
"last-modified": _http_date(node.modified),
|
||||
"cache-control": "no-cache",
|
||||
},
|
||||
lang=lang,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if node is None and not path:
|
||||
# No front page (no top-level node with slug ""): "/" opens the
|
||||
# first item of the navigation instead.
|
||||
for slug, item in sorted_nodes(data.menu):
|
||||
if item.published:
|
||||
return RedirectResponse(f"/{slug}")
|
||||
if _is_trackable_path(path):
|
||||
client_hash = analytics_store.track_404(
|
||||
_client_ip(request),
|
||||
ua,
|
||||
f"/{path}{_query_suffix(request)}",
|
||||
accept_language,
|
||||
)
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
flushed = _track_entry(path, request, status=404)
|
||||
_schedule_client_enrichment(flushed)
|
||||
return _html_response(request, "not-found", path, 404)
|
||||
@@ -21,6 +21,7 @@ ASSETS = Path(__file__).with_name("seed-assets")
|
||||
def _asset(name: str) -> bytes:
|
||||
return (ASSETS / name).read_bytes()
|
||||
|
||||
|
||||
WELCOME = """\
|
||||
Welcome to your new **Pagerite** site. Everything you see is a page written in Markdown, served from a pretty URL, and editable right here in the browser.
|
||||
|
||||
|
||||
+15
-8
@@ -222,7 +222,9 @@ def _linked_block(
|
||||
reconstruction; anything not byte-exact (entities, escapes, an odd
|
||||
link tail) bails to the fallback.
|
||||
"""
|
||||
pieces: list[tuple[str, str]] = [] # (text, mark): "" plain, "link", else the delimiter
|
||||
pieces: list[
|
||||
tuple[str, str]
|
||||
] = [] # (text, mark): "" plain, "link", else the delimiter
|
||||
buf: list[str] = [] # current plain piece
|
||||
link: list[str] | None = None # current mark's text parts
|
||||
mark_kind = "" # the current mark's opener ("link" or the delimiter)
|
||||
@@ -236,7 +238,10 @@ def _linked_block(
|
||||
link = []
|
||||
mark_kind = "link" if tok.type == "link_open" else tok.markup
|
||||
elif tok.type in ("link_close", "strong_close", "em_close", "s_close"):
|
||||
if link is None or ("link" if tok.type == "link_close" else tok.markup) != mark_kind:
|
||||
if (
|
||||
link is None
|
||||
or ("link" if tok.type == "link_close" else tok.markup) != mark_kind
|
||||
):
|
||||
return None
|
||||
inner = "".join(link)
|
||||
if not _LETTER.search(inner):
|
||||
@@ -294,29 +299,31 @@ def _linked_block(
|
||||
# block-trailing one the scanned link tail or the close delimiter.
|
||||
if i == 0:
|
||||
opener = "[" if kind == "link" else kind
|
||||
if s < len(opener) or source[s - len(opener):s] != opener:
|
||||
if s < len(opener) or source[s - len(opener) : s] != opener:
|
||||
return None
|
||||
pre, span_start = opener, s - len(opener)
|
||||
elif pieces[i - 1][1]:
|
||||
pre = "" # the previous mark's post covers the whole gap
|
||||
else:
|
||||
pre = source[located[i - 1][1]:s]
|
||||
pre = source[located[i - 1][1] : s]
|
||||
if i + 1 < len(pieces):
|
||||
post = source[e:located[i + 1][0]]
|
||||
post = source[e : located[i + 1][0]]
|
||||
elif kind == "link":
|
||||
m = _LINK_TAIL.match(source, e)
|
||||
if m is None:
|
||||
return None
|
||||
post, span_end = m.group(), m.end()
|
||||
else:
|
||||
if source[e:e + len(kind)] != kind:
|
||||
if source[e : e + len(kind)] != kind:
|
||||
return None
|
||||
post, span_end = kind, e + len(kind)
|
||||
ps = min(max(offset - lead, 0), len(wire))
|
||||
pe = min(max(offset + len(text_) - lead, 0), len(wire))
|
||||
if pe <= ps:
|
||||
return None
|
||||
marks.append(Mark(_weight(wire[:ps]), _weight(wire[:pe]), pre, post, wire[ps:pe]))
|
||||
marks.append(
|
||||
Mark(_weight(wire[:ps]), _weight(wire[:pe]), pre, post, wire[ps:pe])
|
||||
)
|
||||
offset += len(text_)
|
||||
# Verify: the marks must reconstruct the source span exactly (the only
|
||||
# real risk is the guessed tail of a trailing link).
|
||||
@@ -491,7 +498,7 @@ def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
|
||||
translation = _place_marks(translation, span.weight, span.marks)
|
||||
if translation is None:
|
||||
return None
|
||||
out.append(original[cursor:span.start])
|
||||
out.append(original[cursor : span.start])
|
||||
out.append(translation)
|
||||
cursor = span.end
|
||||
out.append(original[cursor:])
|
||||
|
||||
@@ -0,0 +1,366 @@
|
||||
"""Shared core: site constants, the kanta database, and the render cache.
|
||||
|
||||
Everything the route modules (files, api, tracking, pages) need that is not
|
||||
a route itself: environment-derived paths and tunables, the ``Data`` root
|
||||
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
|
||||
store, the fastapi-vue ``Frontend``, the page render cache
|
||||
(``_render_html``/``_cached_body``/``_html_response`` plus the
|
||||
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
|
||||
content/settings write), the translator ``dispatcher``, the slug charset
|
||||
helpers, and the database bootstrap hooks (demo seed, translator defaults).
|
||||
Importable by every other pagerite module without cycles.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import secrets
|
||||
from datetime import UTC, datetime
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
|
||||
import blake3
|
||||
from fastapi import HTTPException, Request
|
||||
from fastapi.responses import Response
|
||||
from fastapi_vue import Frontend
|
||||
from kanta import Kanta
|
||||
from zstandard import ZstdCompressor
|
||||
|
||||
from pagerite import analytics, i18n, seed, translate, views
|
||||
from pagerite.__main__ import DEVMODE
|
||||
from pagerite.chunks import store_chunks
|
||||
from pagerite.data import (
|
||||
Data,
|
||||
Node,
|
||||
append_order,
|
||||
find_slot,
|
||||
prettify,
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Site identity: the hostname comes from the CLI (first positional argument,
|
||||
# exported as PAGERITE_HOSTNAME) and names the per-site data directory
|
||||
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
|
||||
HOSTNAME = os.getenv("PAGERITE_HOSTNAME", "localhost")
|
||||
SITE_DIR = Path(HOSTNAME)
|
||||
#: Public origin of the site, used for absolute social/canonical/sitemap
|
||||
#: URLs. Localhost serves varying ports, so it falls back to the request's
|
||||
#: own base URL instead.
|
||||
SITE_URL = f"https://{HOSTNAME}" if HOSTNAME != "localhost" else ""
|
||||
|
||||
DB_PATH = os.getenv("PAGERITE_DB", str(SITE_DIR / "content.kantadb"))
|
||||
|
||||
# Visit analytics go to their own JSON file, not the kanta database.
|
||||
ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json")))
|
||||
analytics_store = analytics.Store(ANALYTICS_PATH)
|
||||
|
||||
# Content-addressed file store (uploads, seed assets, fetched favicons):
|
||||
# files on disk under hash-prefixed names, cached in RAM, served at /_f/.
|
||||
FILES_DIR = Path(os.getenv("PAGERITE_FILES", str(SITE_DIR / "files")))
|
||||
|
||||
# Uploaded images are thumbnailed to this size and recompressed to AVIF
|
||||
# (primary), with WebP and JPEG fallbacks re-encoded from the AVIF at
|
||||
# somewhat lower quality (similar or smaller file size); the untouched
|
||||
# original is kept alongside as ``<hash>.orig<ext>`` (never served).
|
||||
IMAGE_MAXSIZE = 1920
|
||||
IMAGE_QUALITY = 60
|
||||
IMAGE_WEBP_QUALITY = 50
|
||||
IMAGE_JPG_QUALITY = 55
|
||||
|
||||
# Favicons get the same derivatives but thumbnailed much smaller — 192px
|
||||
# is plenty (browsers scale down for the 16x16 tab icon themselves).
|
||||
FAVICON_MAXSIZE = 192
|
||||
|
||||
# Our own data root; kanta edits it in place, reads are plain attribute access.
|
||||
data = Data()
|
||||
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
|
||||
|
||||
# Vue build served at the site root, no SPA catch-all (assets only). The
|
||||
# build mirrors the URL space: hashed, immutable files live under
|
||||
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico.
|
||||
BUILD_DIR = Path(__file__).with_name("frontend-build")
|
||||
frontend = Frontend(BUILD_DIR, spa=False, cached="/_assets/")
|
||||
|
||||
# Dynamic HTML is compressed per request at level 9 (static assets are
|
||||
# already pre-compressed by fastapi-vue's Frontend).
|
||||
_zstd = ZstdCompressor(9)
|
||||
|
||||
|
||||
def _render_html(
|
||||
kind: str,
|
||||
path: str,
|
||||
base_url: str,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> str:
|
||||
"""Render one of the generated pages (see _html_response)."""
|
||||
if kind == "page":
|
||||
# A selected language without an actual translation renders the
|
||||
# original (translation is None; see docs/localization.md).
|
||||
original = i18n.primary_lang(data.menu, path)
|
||||
translation = (
|
||||
i18n.get_translation(data, path, lang) if lang != original else None
|
||||
)
|
||||
return views.render_page(
|
||||
data.menu,
|
||||
data,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
base_url,
|
||||
transition=data.transition,
|
||||
lang=lang,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if kind == "category":
|
||||
# A category has no Markdown of its own; only the title map
|
||||
# localizes (heading, navigation, card text).
|
||||
original = i18n.primary_lang(data.menu, path)
|
||||
translation = (
|
||||
i18n.Translation(titles=i18n.title_map(data, lang))
|
||||
if lang != original
|
||||
else None
|
||||
)
|
||||
return views.render_category(
|
||||
data.menu,
|
||||
data,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
transition=data.transition,
|
||||
lang=lang,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
if kind == "not-found":
|
||||
return views.render_not_found(
|
||||
data.menu,
|
||||
path,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
transition=data.transition,
|
||||
)
|
||||
return views.render_analytics(
|
||||
data.menu,
|
||||
data.brand,
|
||||
data.custom_css,
|
||||
data.theme,
|
||||
data.favicon,
|
||||
data.brand_html,
|
||||
transition=data.transition,
|
||||
)
|
||||
|
||||
|
||||
# Render generation: bumped (and the body cache cleared) by every
|
||||
# content/settings write, so page ETags and cached copies invalidate when
|
||||
# navigation-affecting changes happen. In-memory only — not database state.
|
||||
_render_gen = 0
|
||||
|
||||
|
||||
def _invalidate_pages() -> None:
|
||||
"""Drop cached page bodies and bump the render generation (ETags);
|
||||
any content change also re-runs translation dispatch."""
|
||||
global _render_gen
|
||||
_render_gen += 1
|
||||
_cached_body.cache_clear()
|
||||
dispatcher.schedule()
|
||||
|
||||
|
||||
@lru_cache(maxsize=128)
|
||||
def _cached_body(
|
||||
kind: str,
|
||||
path: str,
|
||||
base_url: str,
|
||||
zstd: bool,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> bytes:
|
||||
"""Rendered page body; cleared by _invalidate_pages on any
|
||||
content/settings change. base_url feeds the social meta URLs, zstd
|
||||
selects the stored encoding (both variants are cached rather than
|
||||
re-compressed) and lang the selected language (not the raw
|
||||
Accept-Language header, which would blow up the cache key space).
|
||||
link_lang is the ?lang= override replicated onto the navigation links:
|
||||
a query render and a header-selected render of the same language differ
|
||||
in their links, so they are cached separately.
|
||||
"""
|
||||
body = _render_html(kind, path, base_url, lang, link_lang).encode()
|
||||
return _zstd.compress(body) if zstd else body
|
||||
|
||||
|
||||
def _html_response(
|
||||
request: Request,
|
||||
kind: str,
|
||||
path: str,
|
||||
status_code: int = 200,
|
||||
headers: dict | None = None,
|
||||
etag: bool = False,
|
||||
lang: str = i18n.ORIGINAL_LANGUAGE,
|
||||
link_lang: str = "",
|
||||
) -> Response:
|
||||
"""Response for a generated page, zstd-compressed when the client
|
||||
accepts it (no gzip fallback).
|
||||
|
||||
Done per handler rather than in middleware so that Frontend's
|
||||
already-compressed asset responses are never touched. The ETag stays
|
||||
identical across encodings (revalidation compares it before
|
||||
compression); ``vary: accept-encoding`` keeps caches from mixing the
|
||||
representations. In dev the cache is bypassed so theme/design edits on
|
||||
disk apply immediately.
|
||||
|
||||
``etag=True`` derives the validator from a blake3 hash of the
|
||||
(uncompressed) body — for pages like /_a that have no Node whose
|
||||
modified timestamp could serve as one — and answers matching
|
||||
if-none-match revalidations with a 304.
|
||||
"""
|
||||
zstd = "zstd" in request.headers.get("accept-encoding", "")
|
||||
# Absolute social/canonical URLs use the site's public origin; on
|
||||
# localhost (varying ports) fall back to the request's own base URL.
|
||||
base_url = SITE_URL or str(request.base_url).rstrip("/")
|
||||
if DEVMODE:
|
||||
identity = _render_html(kind, path, base_url, lang, link_lang).encode()
|
||||
body = _zstd.compress(identity) if zstd else identity
|
||||
else:
|
||||
identity = _cached_body(kind, path, base_url, False, lang, link_lang)
|
||||
body = (
|
||||
_cached_body(kind, path, base_url, True, lang, link_lang)
|
||||
if zstd
|
||||
else identity
|
||||
)
|
||||
h = dict(headers or {})
|
||||
# Content varies by language (Accept-Language selects a translation)
|
||||
# and by encoding; keep caches from mixing either representation.
|
||||
h["vary"] = "accept-language" + (", accept-encoding" if zstd else "")
|
||||
if etag:
|
||||
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
|
||||
h["etag"] = tag
|
||||
if request.headers.get("if-none-match") == tag:
|
||||
return Response(status_code=304, headers=h)
|
||||
if zstd:
|
||||
h["content-encoding"] = "zstd"
|
||||
return Response(body, status_code, h, media_type="text/html")
|
||||
|
||||
|
||||
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
|
||||
|
||||
|
||||
def _is_reserved(path: str) -> bool:
|
||||
"""Slug shape that content may never use: each segment must be lower-case
|
||||
ASCII letters, digits, hyphens and underscores (underscores may not be
|
||||
the first character), and dots are never allowed.
|
||||
"""
|
||||
if path == "":
|
||||
return False
|
||||
return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
|
||||
|
||||
|
||||
def _check_reserved(path: str) -> None:
|
||||
"""Reject paths that do not follow the slug charset."""
|
||||
if _is_reserved(path):
|
||||
raise HTTPException(
|
||||
400,
|
||||
'slugs may only use a-z, 0-9, "-" and "_" (not as the first character), and no dots',
|
||||
)
|
||||
|
||||
|
||||
def _ensure(menu: dict[str, Node], path: str) -> Node:
|
||||
"""Return the node at ``path``, creating it and any missing ancestors
|
||||
(content-less category labels) appended at the end of their level."""
|
||||
nodes = menu
|
||||
node = None
|
||||
for seg in path.split("/"):
|
||||
node = nodes.get(seg)
|
||||
if node is None:
|
||||
node = Node(title=prettify(seg), order=append_order(nodes))
|
||||
nodes[seg] = node
|
||||
nodes = node.children
|
||||
return node
|
||||
|
||||
|
||||
def _remove_page(menu: dict[str, Node], path: str) -> bool:
|
||||
"""Delete the node at ``path`` (inside a transaction).
|
||||
|
||||
A node with children becomes a content-less category label; a childless
|
||||
node is removed entirely. Returns False if the path does not exist.
|
||||
"""
|
||||
slot = find_slot(menu, path)
|
||||
node = slot[0].get(slot[1]) if slot else None
|
||||
if node is None:
|
||||
return False
|
||||
if node.children:
|
||||
node.chunks = None
|
||||
node.modified = datetime.now(UTC)
|
||||
else:
|
||||
del slot[0][slot[1]]
|
||||
return True
|
||||
|
||||
|
||||
def _store_seed_file(
|
||||
markdown: str, banner: str, orig: str, body: bytes
|
||||
) -> tuple[str, str]:
|
||||
"""Store a seed file content-addressed and point references at /_f/.
|
||||
|
||||
Images get the same AVIF/WebP/JPEG derivatives as uploads and are
|
||||
linked extension-less; other content is stored as-is with its
|
||||
extension."""
|
||||
from pagerite.files import _ext, store_image # lazy: files imports state
|
||||
|
||||
ext = _ext(orig)
|
||||
name = store_image(body, ext, derive=ext != ".gif")
|
||||
markdown = markdown.replace(f"]({orig}", f"](/_f/{name}")
|
||||
banner = banner.replace(f'src="/{orig}"', f'src="/_f/{name}"')
|
||||
banner = banner.replace(f'src="{orig}"', f'src="/_f/{name}"')
|
||||
return markdown, banner
|
||||
|
||||
|
||||
@kanta.bootstrap
|
||||
def _seed(data: Data) -> None:
|
||||
"""Write the demo pages on database creation (never on existing dbs)."""
|
||||
for path in seed.PAGES:
|
||||
title, markdown, files, banner, order, design = seed.PAGES[path]
|
||||
for orig, body in files.items():
|
||||
markdown, banner = _store_seed_file(markdown, banner, orig, body)
|
||||
node = _ensure(data.menu, path)
|
||||
node.title = title
|
||||
# Empty markdown means a pure category label (e.g. "showcase",
|
||||
# seeded only to carry a banner design): leave chunks as None so
|
||||
# the node renders the placeholder and nav points at its children.
|
||||
if markdown:
|
||||
node.chunks = store_chunks(data.chunks, markdown)
|
||||
node.banner = banner
|
||||
node.banner_design = design
|
||||
node.order = order
|
||||
|
||||
|
||||
#: Translator key format: 12 lowercase alphanumeric characters — not
|
||||
#: brute-forceable over a WebSocket handshake, still human-manageable.
|
||||
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
|
||||
|
||||
|
||||
@kanta.bootstrap
|
||||
def _translator_defaults(data: Data) -> None:
|
||||
"""Translator defaults on database creation: the first service key and
|
||||
the wanted target languages (Spanish and Chinese — English is the
|
||||
original language, never a translation target).
|
||||
|
||||
Keys are a dict (key -> display name) with the future reservation that
|
||||
multiple keys could be managed (e.g. via a web interface)."""
|
||||
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
|
||||
data.translate_keys[key] = "default"
|
||||
data.translate_langs = {"es": True, "zh": True}
|
||||
|
||||
|
||||
# The translator dispatcher — protocol, connected clients and the job
|
||||
# pipeline live in translate.py; its WebSocket route is in api.py.
|
||||
dispatcher = translate.Dispatcher(data, kanta, _invalidate_pages)
|
||||
@@ -0,0 +1,410 @@
|
||||
"""Visit analytics: collection sockets, geoip enrichment, favicon fetch.
|
||||
|
||||
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
|
||||
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
|
||||
enriched in background tasks with reverse DNS (cached PTR lookups) and the
|
||||
DB-IP city MMDB (``GeoIP``, decompressed and opened once at startup);
|
||||
external referrers get their favicon fetched and stored content-hashed.
|
||||
Snapshot broadcasts to connected admin sockets are debounced.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import gzip
|
||||
import ipaddress
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import socket
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
from urllib.parse import urlparse
|
||||
|
||||
import httpx
|
||||
import msgspec
|
||||
from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
|
||||
from fastapi.responses import Response
|
||||
|
||||
from pagerite import analytics
|
||||
from pagerite.files import _hash_name, file_store
|
||||
from pagerite.state import SITE_URL, _html_response, analytics_store
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
# Live WebSocket clients for the analytics stream.
|
||||
_analytics_ws_clients: set[WebSocket] = set()
|
||||
_analytics_broadcast_task: asyncio.Task | None = None
|
||||
|
||||
|
||||
# Repository root from this file's location (pagerite/tracking.py -> ..).
|
||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||
|
||||
|
||||
def _geoip_db_path() -> Path | None:
|
||||
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed
|
||||
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present.
|
||||
"""
|
||||
mmdb = sorted(_REPO_ROOT.glob("dbip-*.mmdb"))
|
||||
if mmdb:
|
||||
return mmdb[0]
|
||||
gz = sorted(_REPO_ROOT.glob("dbip-*.mmdb.gz"))
|
||||
if gz:
|
||||
return gz[0]
|
||||
return None
|
||||
|
||||
|
||||
class GeoIP:
|
||||
"""Lazy DB-IP MMDB reader. Call ``_load()`` once at startup before
|
||||
concurrent requests arrive; ``country()`` is read-only and safe to call
|
||||
from ``asyncio.to_thread`` workers afterwards.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._reader: object | None = None
|
||||
|
||||
def _decompress(self, source: Path, target: Path) -> None:
|
||||
if target.exists():
|
||||
return
|
||||
tmp = target.with_suffix(target.suffix + ".tmp")
|
||||
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
|
||||
shutil.copyfileobj(src, dst)
|
||||
os.replace(tmp, target)
|
||||
|
||||
def _load(self) -> None:
|
||||
if self._reader is not None:
|
||||
return
|
||||
source = _geoip_db_path()
|
||||
if source is None:
|
||||
return
|
||||
if source.suffix == ".gz":
|
||||
target = source.with_suffix("")
|
||||
self._decompress(source, target)
|
||||
source = target
|
||||
try:
|
||||
import maxminddb
|
||||
|
||||
self._reader = maxminddb.open_database(str(source))
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
def country(self, ip: str) -> str:
|
||||
"""Two-letter ISO country code for ``ip``, or "" when unavailable."""
|
||||
if not ip or self._reader is None:
|
||||
return ""
|
||||
try:
|
||||
rec = self._reader.get(ip)
|
||||
if rec:
|
||||
return (rec.get("country") or {}).get("iso_code", "")
|
||||
except Exception:
|
||||
pass
|
||||
return ""
|
||||
|
||||
def city(self, ip: str) -> str:
|
||||
"""City name for ``ip``, or "" when unavailable.
|
||||
|
||||
GeoIP sometimes appends district names in parentheses (e.g.
|
||||
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
|
||||
the value is stored.
|
||||
"""
|
||||
if not ip or self._reader is None:
|
||||
return ""
|
||||
try:
|
||||
rec = self._reader.get(ip)
|
||||
if rec:
|
||||
city = (rec.get("city") or {}).get("names", {}).get("en", "")
|
||||
if city:
|
||||
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
|
||||
return city
|
||||
except Exception:
|
||||
pass
|
||||
return ""
|
||||
|
||||
|
||||
_geoip = GeoIP()
|
||||
|
||||
|
||||
def _client_ip(request: Request | WebSocket) -> str:
|
||||
"""Client IP: first X-Forwarded-For hop (we sit behind a proxy), else
|
||||
the direct peer."""
|
||||
forwarded = request.headers.get("x-forwarded-for", "").split(",")[0].strip()
|
||||
return forwarded or (request.client.host if request.client else "")
|
||||
|
||||
|
||||
def _query_suffix(request: Request) -> str:
|
||||
"""The request's query string as a "?..." suffix, or "" when absent."""
|
||||
query = str(request.url.query)
|
||||
return f"?{query}" if query else ""
|
||||
|
||||
|
||||
@lru_cache(maxsize=4096)
|
||||
def _cached_ptr(ip: str) -> str:
|
||||
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
|
||||
if not ip:
|
||||
return ""
|
||||
try:
|
||||
addr = ipaddress.ip_address(ip)
|
||||
except ValueError:
|
||||
return ""
|
||||
if (
|
||||
addr.is_private
|
||||
or addr.is_loopback
|
||||
or addr.is_reserved
|
||||
or addr.is_multicast
|
||||
or addr.is_link_local
|
||||
):
|
||||
return ""
|
||||
try:
|
||||
host, _, _ = socket.gethostbyaddr(ip)
|
||||
except socket.herror:
|
||||
return ""
|
||||
return host
|
||||
|
||||
|
||||
async def _lookup_host(ip: str) -> str:
|
||||
"""Async wrapper around ``_cached_ptr``; runs the blocking lookup in a thread."""
|
||||
return await asyncio.to_thread(_cached_ptr, ip)
|
||||
|
||||
|
||||
async def _geoip_country(ip: str) -> str:
|
||||
"""Async wrapper around the DB-IP MMDB lookup."""
|
||||
return await asyncio.to_thread(_geoip.country, ip)
|
||||
|
||||
|
||||
async def _geoip_city(ip: str) -> str:
|
||||
"""Async wrapper around the DB-IP MMDB city lookup."""
|
||||
return await asyncio.to_thread(_geoip.city, ip)
|
||||
|
||||
|
||||
async def _enrich_client(client_hash: bytes) -> None:
|
||||
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
|
||||
client = analytics_store.data.clients.get(client_hash)
|
||||
if not client or not client.ip:
|
||||
return
|
||||
host = await _lookup_host(client.ip)
|
||||
country = await _geoip_country(client.ip)
|
||||
city = await _geoip_city(client.ip)
|
||||
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
|
||||
|
||||
|
||||
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
|
||||
"""Start background host/geoip enrichment for the given client hashes."""
|
||||
for client_hash in client_hashes:
|
||||
asyncio.create_task(_enrich_client(client_hash))
|
||||
|
||||
|
||||
#: Icon MIME -> file extension for the stored favicon name. The extension
|
||||
#: reflects the actual content, not the /favicon.ico request path.
|
||||
_FAVICON_EXT = {
|
||||
"image/x-icon": ".ico",
|
||||
"image/vnd.microsoft.icon": ".ico",
|
||||
"image/png": ".png",
|
||||
"image/gif": ".gif",
|
||||
"image/jpeg": ".jpg",
|
||||
"image/webp": ".webp",
|
||||
"image/avif": ".avif",
|
||||
"image/svg+xml": ".svg",
|
||||
}
|
||||
|
||||
_FAVICON_MAX_BYTES = 65536
|
||||
|
||||
#: Origins with a fetch task currently in flight.
|
||||
_favicon_in_flight: set[str] = set()
|
||||
|
||||
|
||||
async def _fetch_favicon(origin: str) -> None:
|
||||
"""Fetch ``{origin}/favicon.ico`` and store it content-hashed on disk.
|
||||
|
||||
The result (icon file name, or "" for a miss) is recorded in the
|
||||
analytics store; misses are retried after analytics._FAVICON_RETRY.
|
||||
Never raises: analytics must not break page serving.
|
||||
"""
|
||||
try:
|
||||
async with httpx.AsyncClient(follow_redirects=True, timeout=8) as client:
|
||||
r = await client.get(f"{origin}/favicon.ico")
|
||||
body = r.content
|
||||
if (
|
||||
not (200 <= r.status_code < 300)
|
||||
or not body
|
||||
or len(body) > _FAVICON_MAX_BYTES
|
||||
):
|
||||
analytics_store.record_favicon(origin)
|
||||
return
|
||||
mime = r.headers.get("content-type", "").split(";")[0].strip().lower()
|
||||
if not mime.startswith("image/"):
|
||||
# Served without an image type: sniff SVG, else assume ICO.
|
||||
if b"<svg" in body[:1024]:
|
||||
mime = "image/svg+xml"
|
||||
elif mime in ("", "application/octet-stream", "text/plain"):
|
||||
mime = "image/x-icon"
|
||||
else:
|
||||
analytics_store.record_favicon(origin)
|
||||
return
|
||||
ext = _FAVICON_EXT.get(mime, ".ico")
|
||||
name = _hash_name(body, f"favicon{ext}")
|
||||
file_store.put(name, body)
|
||||
analytics_store.record_favicon(origin, name)
|
||||
except httpx.HTTPError, OSError:
|
||||
analytics_store.record_favicon(origin)
|
||||
finally:
|
||||
_favicon_in_flight.discard(origin)
|
||||
|
||||
|
||||
def _schedule_favicon_fetch() -> None:
|
||||
"""Start background favicon fetches for origins that need one."""
|
||||
for origin in analytics_store.favicon_origins_needed():
|
||||
if origin in _favicon_in_flight:
|
||||
continue
|
||||
_favicon_in_flight.add(origin)
|
||||
asyncio.create_task(_fetch_favicon(origin))
|
||||
|
||||
|
||||
async def _broadcast_analytics() -> None:
|
||||
"""Send the current analytics snapshot to every connected WS client."""
|
||||
if not _analytics_ws_clients:
|
||||
return
|
||||
payload = analytics_store.display_json()
|
||||
closed = set()
|
||||
for ws in _analytics_ws_clients:
|
||||
try:
|
||||
await ws.send_text(payload)
|
||||
except Exception:
|
||||
closed.add(ws)
|
||||
for ws in closed:
|
||||
_analytics_ws_clients.discard(ws)
|
||||
|
||||
|
||||
async def _debounced_analytics_broadcast() -> None:
|
||||
"""Wait briefly, then broadcast the latest snapshot once."""
|
||||
await asyncio.sleep(0.2)
|
||||
await _broadcast_analytics()
|
||||
|
||||
|
||||
def _schedule_analytics_broadcast() -> None:
|
||||
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
|
||||
global _analytics_broadcast_task
|
||||
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
|
||||
return
|
||||
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
|
||||
_debounced_analytics_broadcast()
|
||||
)
|
||||
|
||||
|
||||
def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]:
|
||||
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
|
||||
|
||||
Nothing is counted on the GET itself — the client's first /_ws message
|
||||
starts the visit, so bots never register as visits (JS-running crawlers
|
||||
connect too, but the WebSocket handler ignores known bot UAs). (Admin
|
||||
clients report too, but with hide, which flags their visit hidden: it is
|
||||
recorded but excluded from all statistics and from the crawler list.)
|
||||
|
||||
The devserver's health probe (``GET /?from=devserver.py`` from
|
||||
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
|
||||
logged as a crawler hit. The root-path and localhost checks prevent
|
||||
remote visitors from hiding traffic with the same query string.
|
||||
|
||||
Returns the client hashes of any pending crawler hits flushed to persistent
|
||||
storage, so callers can schedule async geoip and reverse-DNS enrichment.
|
||||
"""
|
||||
if request.headers.get("x-pagerite-preload"):
|
||||
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
|
||||
# activity message sent when the user actually navigates does the
|
||||
# counting.
|
||||
# (Forging the header only hides a GET from the crawler stats; the
|
||||
# path-based abuse classification is unaffected.)
|
||||
return []
|
||||
if (
|
||||
path == ""
|
||||
and str(request.url.query) == "from=devserver.py"
|
||||
and _client_ip(request) == "127.0.0.1"
|
||||
):
|
||||
return []
|
||||
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
|
||||
full_path = f"{request.url.path}{_query_suffix(request)}"
|
||||
return analytics_store.track_entry(
|
||||
request.headers.get("referer", ""),
|
||||
own_origin,
|
||||
_client_ip(request),
|
||||
request.headers.get("user-agent", ""),
|
||||
full_path,
|
||||
request.headers.get("accept-language", ""),
|
||||
status=status,
|
||||
)
|
||||
|
||||
|
||||
@router.get("/_a", response_model=None)
|
||||
async def analytics_page(request: Request) -> Response:
|
||||
"""Render the analytics viewer as a normal site page at /_a.
|
||||
|
||||
The page itself is public, but the data stream (/_api/ws/analytics) stays
|
||||
admin-gated like the rest of /_api, so only authorized users see the
|
||||
statistics; others get the viewer with a "could not be loaded" message.
|
||||
"""
|
||||
return _html_response(
|
||||
request,
|
||||
"analytics",
|
||||
"",
|
||||
headers={"cache-control": "no-cache"},
|
||||
etag=True,
|
||||
)
|
||||
|
||||
|
||||
@router.websocket("/_ws")
|
||||
async def activity_ws(ws: WebSocket) -> None:
|
||||
"""Collect visitor activity: navigations and reading-time updates.
|
||||
|
||||
Public, like the pages themselves (only /_api is gated); one connection
|
||||
follows a browsing session. Messages are ``analytics.Ping`` structs as
|
||||
JSON text frames; ``to`` set is a navigation, ``read`` alone a
|
||||
reading-time update. The reverse-DNS and DB-IP geoip lookups happen in
|
||||
background tasks so message handling is never delayed by slow DNS or
|
||||
the first MMDB decompress.
|
||||
"""
|
||||
await ws.accept()
|
||||
ip = _client_ip(ws)
|
||||
ua = ws.headers.get("user-agent", "")
|
||||
accept_language = ws.headers.get("accept-language", "")
|
||||
try:
|
||||
while True:
|
||||
text = await ws.receive_text()
|
||||
try:
|
||||
msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
|
||||
except msgspec.DecodeError:
|
||||
continue
|
||||
visit_index, flushed_clients = analytics_store.ping(
|
||||
msg.fr,
|
||||
msg.to or None,
|
||||
ip,
|
||||
ua,
|
||||
accept_language,
|
||||
hide=msg.hide,
|
||||
read=msg.read,
|
||||
)
|
||||
if visit_index is not None:
|
||||
visit = analytics_store.data.visits[visit_index]
|
||||
asyncio.create_task(_enrich_client(visit.client))
|
||||
_schedule_client_enrichment(flushed_clients)
|
||||
_schedule_favicon_fetch()
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
|
||||
|
||||
@router.websocket("/_api/ws/analytics")
|
||||
async def analytics_websocket(ws: WebSocket) -> None:
|
||||
"""Stream the analytics snapshot, then push updates as they happen.
|
||||
|
||||
Admin-only via the /_api forward-auth gate, like every management
|
||||
endpoint. Powers the analytics viewer rendered at /_a.
|
||||
"""
|
||||
await ws.accept()
|
||||
await ws.send_text(analytics_store.display_json())
|
||||
_analytics_ws_clients.add(ws)
|
||||
try:
|
||||
while True:
|
||||
await ws.receive_text()
|
||||
except Exception:
|
||||
pass
|
||||
finally:
|
||||
_analytics_ws_clients.discard(ws)
|
||||
+45
-18
@@ -1,7 +1,7 @@
|
||||
"""Translator service protocol, dispatcher and its transport-independent core.
|
||||
|
||||
The external machine-translation service connects over WebSocket
|
||||
(``/_translate/<key>``, the route itself is in app.py) and exchanges JSON
|
||||
(``/_translate/<key>``, the route itself is in api.py) and exchanges JSON
|
||||
frames decoded into the tagged msgspec structs below (``bytes`` fields ride
|
||||
as base64 — no manual encoding anywhere). This module holds everything
|
||||
else: the message structs, the connected-client dispatcher (``Dispatcher``
|
||||
@@ -117,7 +117,9 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
|
||||
if key in seen or lang in data.trans.get(key, {}):
|
||||
return
|
||||
seen.add(key)
|
||||
items.append(TransItem(key=key, text=text, path=path, kind=kind, context=context))
|
||||
items.append(
|
||||
TransItem(key=key, text=text, path=path, kind=kind, context=context)
|
||||
)
|
||||
|
||||
def opening(node: Node) -> str:
|
||||
"""The article's opening prose (first segment, capped): the title
|
||||
@@ -137,8 +139,13 @@ def pending_items(data: Data, lang: str) -> list[TransItem]:
|
||||
node_lang = node.language or inherited
|
||||
if node.chunks is not None and node_lang != lang:
|
||||
if node.title:
|
||||
emit(chunk_key(node.title), node.title, path, "title",
|
||||
context=opening(node))
|
||||
emit(
|
||||
chunk_key(node.title),
|
||||
node.title,
|
||||
path,
|
||||
"title",
|
||||
context=opening(node),
|
||||
)
|
||||
for h in node.chunks:
|
||||
text = data.chunks.get(h)
|
||||
if (
|
||||
@@ -238,7 +245,7 @@ class Dispatcher:
|
||||
def __init__(self, data: Data, db: Kanta, invalidate) -> None:
|
||||
self.data = data
|
||||
self.db = db
|
||||
#: Sync content-change hook (app._invalidate_pages), called inside
|
||||
#: Sync content-change hook (state._invalidate_pages), called inside
|
||||
#: transactions; schedules the next dispatch pass.
|
||||
self.invalidate = invalidate
|
||||
#: Connected translator sockets and their per-connection state.
|
||||
@@ -247,6 +254,12 @@ class Dispatcher:
|
||||
#: (segment count, empty or non-prose segments, segments.py) this run.
|
||||
self.validation_failures: set[tuple[str, bytes]] = set()
|
||||
|
||||
def reset_validation_failures(self) -> None:
|
||||
"""Clear the skip list of fragments rejected this run (segment
|
||||
validation): a translations refresh is precisely the "another
|
||||
chance" for them."""
|
||||
self.validation_failures.clear()
|
||||
|
||||
def schedule(self) -> None:
|
||||
"""Schedule a dispatch pass, if any translator is connected.
|
||||
|
||||
@@ -265,9 +278,7 @@ class Dispatcher:
|
||||
async def _dispatch(self) -> None:
|
||||
"""Offer one pending item to every free capable connection."""
|
||||
wanted = {
|
||||
tag
|
||||
for lang in self.data.translate_langs
|
||||
if (tag := i18n.base_tag(lang))
|
||||
tag for lang in self.data.translate_langs if (tag := i18n.base_tag(lang))
|
||||
}
|
||||
if not wanted:
|
||||
return
|
||||
@@ -283,7 +294,10 @@ class Dispatcher:
|
||||
original = ""
|
||||
for lang in sorted(langs):
|
||||
for item in pending_items(self.data, lang):
|
||||
if (lang, item.key) in inflight or (lang, item.key) in self.validation_failures:
|
||||
if (lang, item.key) in inflight or (
|
||||
lang,
|
||||
item.key,
|
||||
) in self.validation_failures:
|
||||
continue
|
||||
spans, texts, contexts = split(item.text)
|
||||
if not texts:
|
||||
@@ -294,8 +308,12 @@ class Dispatcher:
|
||||
# (TransItem.context), not its own one-word block.
|
||||
contexts = [item.context] * len(texts)
|
||||
job = Job(
|
||||
lang=lang, key=item.key, texts=texts,
|
||||
path=item.path, kind=item.kind, contexts=contexts,
|
||||
lang=lang,
|
||||
key=item.key,
|
||||
texts=texts,
|
||||
path=item.path,
|
||||
kind=item.kind,
|
||||
contexts=contexts,
|
||||
)
|
||||
break
|
||||
if job is not None:
|
||||
@@ -339,9 +357,9 @@ class Dispatcher:
|
||||
if state is not None: # one Hello per connection
|
||||
await ws.close(code=1002)
|
||||
return
|
||||
state = _Connection({
|
||||
tag for lang in msg.langs if (tag := i18n.base_tag(lang))
|
||||
})
|
||||
state = _Connection(
|
||||
{tag for lang in msg.langs if (tag := i18n.base_tag(lang))}
|
||||
)
|
||||
self.clients[ws] = state
|
||||
self.schedule()
|
||||
else: # Result
|
||||
@@ -357,7 +375,11 @@ class Dispatcher:
|
||||
state.inflight = None
|
||||
state.spans = []
|
||||
state.original = ""
|
||||
text = join(original, spans, texts) if len(texts) == len(spans) else None
|
||||
text = (
|
||||
join(original, spans, texts)
|
||||
if len(texts) == len(spans)
|
||||
else None
|
||||
)
|
||||
if text is None:
|
||||
# The model broke the segment contract (count
|
||||
# mismatch, empty or non-prose segment): drop the
|
||||
@@ -367,11 +389,14 @@ class Dispatcher:
|
||||
self.validation_failures.add((lang, msg.key))
|
||||
logger.warning(
|
||||
"[%s] result for chunk %s rejected: invalid segments",
|
||||
lang, msg.key.hex(),
|
||||
lang,
|
||||
msg.key.hex(),
|
||||
)
|
||||
self.schedule()
|
||||
continue
|
||||
with self.db.transaction("translator results", user=clientkey, extra=lang):
|
||||
with self.db.transaction(
|
||||
"translator results", user=clientkey, extra=lang
|
||||
):
|
||||
paths = store_results(
|
||||
self.data, lang, [TransResult(key=msg.key, text=text)]
|
||||
)
|
||||
@@ -379,7 +404,9 @@ class Dispatcher:
|
||||
if paths:
|
||||
logger.info(
|
||||
"[%s] now available for %d page(s): %s",
|
||||
lang, len(paths), ", ".join(sorted(paths)),
|
||||
lang,
|
||||
len(paths),
|
||||
", ".join(sorted(paths)),
|
||||
)
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
|
||||
+144
-31
@@ -39,7 +39,9 @@ def _data_roots() -> list[Path]:
|
||||
roots = [user_data_path("pagerite", appauthor=False)]
|
||||
# site_data_dir keeps the multipath (site_data_path collapses it to the
|
||||
# first entry, since a Path cannot hold several).
|
||||
roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(os.pathsep)
|
||||
roots += site_data_dir("pagerite", appauthor=False, multipath=True).split(
|
||||
os.pathsep
|
||||
)
|
||||
return [Path(r) for r in roots]
|
||||
|
||||
|
||||
@@ -104,13 +106,15 @@ def _theme_color_schemes(theme: str) -> set[str]:
|
||||
return set()
|
||||
try:
|
||||
css = path.read_text()
|
||||
except (OSError, ValueError):
|
||||
except OSError, ValueError:
|
||||
return set()
|
||||
css = re.sub(r"/\*.*?\*/", "", css, flags=re.DOTALL)
|
||||
m = re.search(r"color-scheme\s*:\s*([^;]+);", css, re.IGNORECASE)
|
||||
if not m:
|
||||
return set()
|
||||
return {tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}}
|
||||
return {
|
||||
tok.lower() for tok in m.group(1).split() if tok.lower() in {"light", "dark"}
|
||||
}
|
||||
|
||||
|
||||
def _theme_mode(theme: str) -> str:
|
||||
@@ -346,7 +350,9 @@ def _layout(
|
||||
(see docs/localization.md), emitted right after the viewport and before
|
||||
the social tags: canonical first, then the hreflang alternates.
|
||||
"""
|
||||
doc = Document(E.Title, lang=lang, dir="rtl" if lang in i18n.RTL_LANGUAGES else "ltr")
|
||||
doc = Document(
|
||||
E.Title, lang=lang, dir="rtl" if lang in i18n.RTL_LANGUAGES else "ltr"
|
||||
)
|
||||
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs
|
||||
# the real device width, not the default 980px layout viewport.
|
||||
doc.meta(name="viewport", content="width=device-width, initial-scale=1")
|
||||
@@ -426,8 +432,7 @@ def _layout(
|
||||
if custom_css.strip():
|
||||
doc.style(custom_css, id="pagerite-user")
|
||||
body = (
|
||||
doc
|
||||
.header(
|
||||
doc.header(
|
||||
E.div(E.Banner, id="page-banner"),
|
||||
E.Brand,
|
||||
E.nav(E.Nav, id="nav"),
|
||||
@@ -463,10 +468,16 @@ def _brand_link(brand: str, brand_html: str = "", link_lang: str = "") -> HTML:
|
||||
when neither is set."""
|
||||
if brand_html.strip():
|
||||
return HTML(str(E.div(HTML(brand_html), id="brand")))
|
||||
return HTML(str(E.a(brand, href=_href("", link_lang), id="brand"))) if brand else HTML("")
|
||||
return (
|
||||
HTML(str(E.a(brand, href=_href("", link_lang), id="brand")))
|
||||
if brand
|
||||
else HTML("")
|
||||
)
|
||||
|
||||
|
||||
def _title(slug: str, node: Node, translation: Translation | None = None, path: str = "") -> str:
|
||||
def _title(
|
||||
slug: str, node: Node, translation: Translation | None = None, path: str = ""
|
||||
) -> str:
|
||||
"""Menu label: the configured title, prettified slug, "Home" fallback.
|
||||
|
||||
With a translation, its title map (keyed by node path) wins, falling
|
||||
@@ -486,8 +497,13 @@ def _href(path: str, link_lang: str = "") -> str:
|
||||
|
||||
|
||||
def _nav_link(
|
||||
doc, menu: dict[str, Node], node: Node, path: str, current: str,
|
||||
ancestors_current: bool = True, translation: Translation | None = None,
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
node: Node,
|
||||
path: str,
|
||||
current: str,
|
||||
ancestors_current: bool = True,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> None:
|
||||
"""Render one <li> linking the node. Category labels (no content of
|
||||
@@ -509,7 +525,12 @@ def _nav_link(
|
||||
)
|
||||
|
||||
|
||||
def nav_html(menu: dict[str, Node], current: str, translation: Translation | None = None, link_lang: str = "") -> HTML:
|
||||
def nav_html(
|
||||
menu: dict[str, Node],
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the contents of the #nav element for the current path.
|
||||
|
||||
Top-level items in menu order; the front page (slug "", href "/")
|
||||
@@ -520,11 +541,24 @@ def nav_html(menu: dict[str, Node], current: str, translation: Translation | Non
|
||||
with nav:
|
||||
for slug, node in sorted_nodes(menu):
|
||||
if node.published:
|
||||
_nav_link(nav, menu, node, slug, current, translation=translation, link_lang=link_lang)
|
||||
_nav_link(
|
||||
nav,
|
||||
menu,
|
||||
node,
|
||||
slug,
|
||||
current,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
return HTML(str(nav))
|
||||
|
||||
|
||||
def sidebar_html(menu: dict[str, Node], current: str, translation: Translation | None = None, link_lang: str = "") -> HTML:
|
||||
def sidebar_html(
|
||||
menu: dict[str, Node],
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the #sidebar element for the current path (empty when none).
|
||||
|
||||
The sidebar is the current main level section's sub-navigation: the
|
||||
@@ -558,20 +592,41 @@ def sidebar_html(menu: dict[str, Node], current: str, translation: Translation |
|
||||
nav = E.ul
|
||||
with nav:
|
||||
for slug, child in items:
|
||||
_sidebar_item(nav, menu, child, f"{section}/{slug}", current, translation, link_lang)
|
||||
_sidebar_item(
|
||||
nav, menu, child, f"{section}/{slug}", current, translation, link_lang
|
||||
)
|
||||
return HTML(str(E.aside(nav, id="sidebar")))
|
||||
|
||||
|
||||
def _sidebar_item(doc, menu: dict[str, Node], node: Node, path: str, current: str, translation: Translation | None = None, link_lang: str = "") -> None:
|
||||
def _sidebar_item(
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
node: Node,
|
||||
path: str,
|
||||
current: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
) -> None:
|
||||
"""One sidebar <li>: the node link, with its published children as a
|
||||
nested list (third level and deeper, recursively)."""
|
||||
_nav_link(doc, menu, node, path, current, ancestors_current=False, translation=translation, link_lang=link_lang)
|
||||
_nav_link(
|
||||
doc,
|
||||
menu,
|
||||
node,
|
||||
path,
|
||||
current,
|
||||
ancestors_current=False,
|
||||
translation=translation,
|
||||
link_lang=link_lang,
|
||||
)
|
||||
sub = [(s, c) for s, c in sorted_nodes(node.children) if c.published]
|
||||
if sub:
|
||||
# doc.li.a(...) above left the <li> open for nesting.
|
||||
with doc.ul:
|
||||
for slug, child in sub:
|
||||
_sidebar_item(doc, menu, child, f"{path}/{slug}", current, translation, link_lang)
|
||||
_sidebar_item(
|
||||
doc, menu, child, f"{path}/{slug}", current, translation, link_lang
|
||||
)
|
||||
|
||||
|
||||
def first_leaf(menu: dict[str, Node], path: str) -> str | None:
|
||||
@@ -703,7 +758,14 @@ def banner_source(menu: dict[str, Node], path: str) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
def page_content(menu: dict[str, Node], data: Data, path: str, translation: Translation | None = None, link_lang: str = "", lang: str = "") -> HTML:
|
||||
def page_content(
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> HTML:
|
||||
"""Render the contents of the #main element for a page.
|
||||
|
||||
A page with published children (a category page) lists them as cards
|
||||
@@ -718,7 +780,11 @@ def page_content(menu: dict[str, Node], data: Data, path: str, translation: Tran
|
||||
if translation:
|
||||
if translation.markdown is not None:
|
||||
content = translation.markdown
|
||||
title = _title(path.rpartition("/")[2], node, translation, path) if node.title else title
|
||||
title = (
|
||||
_title(path.rpartition("/")[2], node, translation, path)
|
||||
if node.title
|
||||
else title
|
||||
)
|
||||
# The title is injected into the markdown (as # title when it has no
|
||||
# h1 of its own), so title and content render as one article.
|
||||
rendered = render(content, path, node.created, node.modified, title=title)
|
||||
@@ -733,7 +799,16 @@ def page_content(menu: dict[str, Node], data: Data, path: str, translation: Tran
|
||||
return HTML(str(doc))
|
||||
|
||||
|
||||
def _cards(doc, menu: dict[str, Node], data: Data, node: Node, path: str, translation: Translation | None = None, link_lang: str = "", lang: str = "") -> None:
|
||||
def _cards(
|
||||
doc,
|
||||
menu: dict[str, Node],
|
||||
data: Data,
|
||||
node: Node,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> None:
|
||||
"""Card stacks of the node's published children (nothing when childless).
|
||||
|
||||
One column per direct child, all in a single full-width row (the .wide
|
||||
@@ -773,7 +848,15 @@ def _walk(node: Node, path: str):
|
||||
yield from _walk(child, f"{path}/{slug}")
|
||||
|
||||
|
||||
def _card(doc, data: Data, node: Node, path: str, translation: Translation | None = None, link_lang: str = "", lang: str = "") -> None:
|
||||
def _card(
|
||||
doc,
|
||||
data: Data,
|
||||
node: Node,
|
||||
path: str,
|
||||
translation: Translation | None = None,
|
||||
link_lang: str = "",
|
||||
lang: str = "",
|
||||
) -> None:
|
||||
"""One card in a stack: cover + title, plus the description when the
|
||||
page has no image (its card shows a gradient cover instead).
|
||||
|
||||
@@ -796,7 +879,9 @@ def _card(doc, data: Data, node: Node, path: str, translation: Translation | Non
|
||||
doc.span(class_="cover", style=f'background-image: url("{image}")')
|
||||
else:
|
||||
doc.span(class_="cover")
|
||||
doc.span(_title(path.rpartition("/")[2], node, translation, path), class_="title")
|
||||
doc.span(
|
||||
_title(path.rpartition("/")[2], node, translation, path), class_="title"
|
||||
)
|
||||
if description:
|
||||
doc.span(description, class_="desc")
|
||||
|
||||
@@ -823,7 +908,9 @@ def _description(html: str, limit: int = 200) -> str:
|
||||
return text
|
||||
# Prefer a clean cut: the last sentence ending within the limit, as
|
||||
# long as it does not reduce the description to a tiny fragment.
|
||||
if (end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)) > limit // 2:
|
||||
if (
|
||||
end := max((m.end() for m in _SENTENCE_END.finditer(text[:limit])), default=0)
|
||||
) > limit // 2:
|
||||
return text[:end]
|
||||
return text[:limit].rsplit(" ", 1)[0] + "…"
|
||||
|
||||
@@ -879,7 +966,12 @@ def _share_media(html: str, base_url: str) -> tuple[str, str]:
|
||||
|
||||
|
||||
def _social_meta(
|
||||
node: Node, path: str, title: str, html: str, brand: str, base_url: str,
|
||||
node: Node,
|
||||
path: str,
|
||||
title: str,
|
||||
html: str,
|
||||
brand: str,
|
||||
base_url: str,
|
||||
) -> dict[str, str]:
|
||||
"""Open Graph/Twitter/SEO meta tags for a content page.
|
||||
|
||||
@@ -896,9 +988,7 @@ def _social_meta(
|
||||
url = f"{base_url}/{path}" if base_url else ""
|
||||
text = _description(html)
|
||||
image, video = _share_media(html, base_url)
|
||||
twitter_image = (
|
||||
re.sub(r"(/_f/[0-9a-f]{12})$", r"\1.webp", image) if image else ""
|
||||
)
|
||||
twitter_image = re.sub(r"(/_f/[0-9a-f]{12})$", r"\1.webp", image) if image else ""
|
||||
return {
|
||||
"description": text,
|
||||
"og:type": "article",
|
||||
@@ -964,8 +1054,16 @@ def render_page(
|
||||
]
|
||||
return str(
|
||||
_layout(
|
||||
*_page_assets(), custom_css, theme, banner_design(menu, path, theme),
|
||||
transition, favicon, social, lang, canonical, alternates,
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
social,
|
||||
lang,
|
||||
canonical,
|
||||
alternates,
|
||||
)(
|
||||
Title=f"{title} – {brand}" if brand else title,
|
||||
Brand=_brand_link(brand, brand_html, link_lang),
|
||||
@@ -1015,7 +1113,15 @@ def render_category(
|
||||
else:
|
||||
doc.p("This section has no page of its own yet.")
|
||||
return str(
|
||||
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon, lang=lang)(
|
||||
_layout(
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
lang=lang,
|
||||
)(
|
||||
Title=f"{title} – {brand}" if brand else title,
|
||||
Brand=_brand_link(brand, brand_html, link_lang),
|
||||
Nav=nav_html(menu, path, translation, link_lang),
|
||||
@@ -1042,7 +1148,14 @@ def render_not_found(
|
||||
doc.h1("Not Found")
|
||||
doc.p(f"No article at /{path}. If there was before, it may have been deleted.")
|
||||
return str(
|
||||
_layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), transition, favicon)(
|
||||
_layout(
|
||||
*_page_assets(),
|
||||
custom_css,
|
||||
theme,
|
||||
banner_design(menu, path, theme),
|
||||
transition,
|
||||
favicon,
|
||||
)(
|
||||
Title=f"Not Found – {brand}" if brand else "Not Found",
|
||||
Brand=_brand_link(brand, brand_html),
|
||||
Nav=nav_html(menu, path),
|
||||
|
||||
+27
-19
@@ -91,22 +91,22 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
|
||||
CrawlerProfile(
|
||||
"googlebot",
|
||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
|
||||
"66.249.64.66", # US, Google
|
||||
"66.249.64.66", # US, Google
|
||||
),
|
||||
CrawlerProfile(
|
||||
"bingbot",
|
||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
|
||||
"40.77.167.0", # US, Microsoft
|
||||
"40.77.167.0", # US, Microsoft
|
||||
),
|
||||
CrawlerProfile(
|
||||
"duckduckbot",
|
||||
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
|
||||
"95.217.0.1", # Germany, Hetzner VPS
|
||||
"95.217.0.1", # Germany, Hetzner VPS
|
||||
),
|
||||
CrawlerProfile(
|
||||
"curl",
|
||||
"curl/8.5.0",
|
||||
"139.162.0.1", # Singapore, Linode VPS
|
||||
"139.162.0.1", # Singapore, Linode VPS
|
||||
),
|
||||
]
|
||||
|
||||
@@ -115,14 +115,14 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
|
||||
# host part; the host may rotate once mid-session.
|
||||
RESIDENTIAL_SOURCE_IPS: list[str] = [
|
||||
# Residential IPv4
|
||||
"91.154.140.209", # Finland, Elisa
|
||||
"84.143.145.207", # Germany, Deutsche Telekom
|
||||
"220.165.255.254", # China, Chinanet / China Telecom
|
||||
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
|
||||
"91.154.140.209", # Finland, Elisa
|
||||
"84.143.145.207", # Germany, Deutsche Telekom
|
||||
"220.165.255.254", # China, Chinanet / China Telecom
|
||||
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
|
||||
# Residential IPv6 /64 prefixes
|
||||
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
|
||||
"240e:45d:1e60:5b0::/64", # China, China Telecom
|
||||
"2409:8904:6720:4123::/64", # China, China Unicom
|
||||
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
|
||||
"240e:45d:1e60:5b0::/64", # China, China Telecom
|
||||
"2409:8904:6720:4123::/64", # China, China Unicom
|
||||
]
|
||||
|
||||
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
|
||||
@@ -130,9 +130,9 @@ RESIDENTIAL_SOURCE_IPS: list[str] = [
|
||||
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
|
||||
# and index 2 uses a fixed crawler UA.
|
||||
ABUSE_SOURCE_IPS: list[str] = [
|
||||
"45.63.0.12", # US, Vultr VPS
|
||||
"138.197.0.89", # US, DigitalOcean / Cloudways
|
||||
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
|
||||
"45.63.0.12", # US, Vultr VPS
|
||||
"138.197.0.89", # US, DigitalOcean / Cloudways
|
||||
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
|
||||
]
|
||||
|
||||
# Paths commonly probed by attackers looking for exposed config, admin panels,
|
||||
@@ -284,10 +284,16 @@ TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
|
||||
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
|
||||
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
|
||||
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
|
||||
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
|
||||
(
|
||||
"https://news.ycombinator.com/",
|
||||
{"utm_source": "hackernews", "utm_medium": "referral"},
|
||||
),
|
||||
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
|
||||
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
|
||||
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
|
||||
(
|
||||
"https://www.producthunt.com/",
|
||||
{"utm_source": "producthunt", "utm_medium": "referral"},
|
||||
),
|
||||
]
|
||||
|
||||
# Fraction of referered sessions that also carry UTM tags.
|
||||
@@ -437,7 +443,7 @@ def _random_ipv6_host(prefix: str) -> str:
|
||||
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
|
||||
if base.endswith("::"):
|
||||
base = base[:-2]
|
||||
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
|
||||
host = ":".join(f"{random.randint(0, 0xFFFF):04x}" for _ in range(4))
|
||||
return f"{base}:{host}"
|
||||
|
||||
|
||||
@@ -760,9 +766,11 @@ def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
|
||||
if ua_mode == 0:
|
||||
get_ua = _abuse_ua
|
||||
elif ua_mode == 1:
|
||||
|
||||
def get_ua() -> str:
|
||||
return BROWSER_PROFILES[0].user_agent
|
||||
else:
|
||||
|
||||
def get_ua() -> str:
|
||||
return CRAWLER_PROFILES[0].user_agent
|
||||
|
||||
@@ -804,8 +812,8 @@ def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
|
||||
nargs="?",
|
||||
default="http://localhost:8200",
|
||||
help="Base URL of the Pagerite site (default: http://localhost:8200). "
|
||||
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
|
||||
"missing scheme defaults to http://.",
|
||||
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
|
||||
"missing scheme defaults to http://.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-t",
|
||||
|
||||
+90
-31
@@ -50,13 +50,34 @@ SEED_X = "ByteDance-Seed/Seed-X-PPO-7B"
|
||||
|
||||
# Seed-X language tags (appended to the prompt; required by its PPO training)
|
||||
SEED_X_TAGS = {
|
||||
"arabic": "ar", "chinese": "zh", "czech": "cs", "danish": "da",
|
||||
"dutch": "nl", "english": "en", "finnish": "fi", "french": "fr",
|
||||
"german": "de", "greek": "el", "hungarian": "hu", "indonesian": "id",
|
||||
"italian": "it", "japanese": "ja", "korean": "ko", "malay": "ms",
|
||||
"norwegian": "no", "persian": "fa", "polish": "pl", "portuguese": "pt",
|
||||
"romanian": "ro", "russian": "ru", "spanish": "es", "swedish": "sv",
|
||||
"thai": "th", "turkish": "tr", "ukrainian": "uk", "vietnamese": "vi",
|
||||
"arabic": "ar",
|
||||
"chinese": "zh",
|
||||
"czech": "cs",
|
||||
"danish": "da",
|
||||
"dutch": "nl",
|
||||
"english": "en",
|
||||
"finnish": "fi",
|
||||
"french": "fr",
|
||||
"german": "de",
|
||||
"greek": "el",
|
||||
"hungarian": "hu",
|
||||
"indonesian": "id",
|
||||
"italian": "it",
|
||||
"japanese": "ja",
|
||||
"korean": "ko",
|
||||
"malay": "ms",
|
||||
"norwegian": "no",
|
||||
"persian": "fa",
|
||||
"polish": "pl",
|
||||
"portuguese": "pt",
|
||||
"romanian": "ro",
|
||||
"russian": "ru",
|
||||
"spanish": "es",
|
||||
"swedish": "sv",
|
||||
"thai": "th",
|
||||
"turkish": "tr",
|
||||
"ukrainian": "uk",
|
||||
"vietnamese": "vi",
|
||||
}
|
||||
SEED_X_NAMES = {v: k for k, v in SEED_X_TAGS.items()}
|
||||
|
||||
@@ -91,11 +112,11 @@ PROMPTS = {
|
||||
# separator in the output (the model merged them) → seed_x_chunk falls
|
||||
# back to the plain kind template.
|
||||
"title+context": "Translate the following {source_lang} title and the beginning of its article "
|
||||
"into {target_lang}:\n{text}\n\n{context} <{tag}>",
|
||||
"into {target_lang}:\n{text}\n\n{context} <{tag}>",
|
||||
# A segment carved out of a larger block (link text, partial run) with
|
||||
# its sentence as context — same mechanics as title+context.
|
||||
"chunk+context": "Translate the following {source_lang} text into {target_lang}:\n"
|
||||
"{text}\n\n{context} <{tag}>",
|
||||
"{text}\n\n{context} <{tag}>",
|
||||
}
|
||||
TERMINAL_PUNCT = ".,!?:;…。,!?;:、"
|
||||
|
||||
@@ -176,7 +197,8 @@ class SeedX:
|
||||
def _load(self):
|
||||
t0 = time.monotonic()
|
||||
self.model = AutoModelForCausalLM.from_pretrained(
|
||||
SEED_X, dtype=torch.bfloat16, device_map="auto")
|
||||
SEED_X, dtype=torch.bfloat16, device_map="auto"
|
||||
)
|
||||
print(f"[seed-x loaded in {time.monotonic() - t0:.0f}s]", file=sys.stderr)
|
||||
|
||||
def get(self):
|
||||
@@ -207,8 +229,16 @@ class SeedX:
|
||||
print(f"[seed-x unloaded after {IDLE_UNLOAD_S}s idle]", file=sys.stderr)
|
||||
|
||||
|
||||
def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str,
|
||||
kind: str = "chunk", context: str = "", source_lang: str = "English"):
|
||||
def seed_x_chunk(
|
||||
tokenizer,
|
||||
model,
|
||||
text: str,
|
||||
target_lang: str,
|
||||
tag: str,
|
||||
kind: str = "chunk",
|
||||
context: str = "",
|
||||
source_lang: str = "English",
|
||||
):
|
||||
"""Translate one segment; returns (translation, output_tokens, generation_seconds).
|
||||
|
||||
With context, the segment is translated together with its surround (a
|
||||
@@ -224,8 +254,13 @@ def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str,
|
||||
"""
|
||||
# No chat template on this model; the trailing language tag is required (trans/ style prompt).
|
||||
template = PROMPTS.get(f"{kind}+context" if context else kind, PROMPTS["chunk"])
|
||||
prompt = template.format(source_lang=source_lang, target_lang=target_lang,
|
||||
text=text, tag=tag, context=context)
|
||||
prompt = template.format(
|
||||
source_lang=source_lang,
|
||||
target_lang=target_lang,
|
||||
text=text,
|
||||
tag=tag,
|
||||
context=context,
|
||||
)
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||
t0 = time.monotonic()
|
||||
# The only stop string is the context separator. "<" must NOT be one:
|
||||
@@ -234,11 +269,17 @@ def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str,
|
||||
# decode; the post-decode cut at the first "<" then enforces the wire
|
||||
# invariant (prose only) against markup bleed.
|
||||
kwargs = {"stop_strings": ["\n\n"], "tokenizer": tokenizer} if context else {}
|
||||
out = model.generate(**inputs, max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]),
|
||||
do_sample=False, **kwargs)
|
||||
out = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=max(1024, 2 * inputs.input_ids.shape[1]),
|
||||
do_sample=False,
|
||||
**kwargs,
|
||||
)
|
||||
dt = time.monotonic() - t0
|
||||
n = out.shape[1] - inputs.input_ids.shape[1]
|
||||
decoded = tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
|
||||
decoded = tokenizer.decode(
|
||||
out[0][inputs.input_ids.shape[1] :], skip_special_tokens=True
|
||||
)
|
||||
translated = decoded.partition("<")[0]
|
||||
if not context:
|
||||
return translated.strip(), n, dt
|
||||
@@ -256,8 +297,9 @@ def seed_x_chunk(tokenizer, model, text: str, target_lang: str, tag: str,
|
||||
return out, n, dt
|
||||
# The model merged segment and context (no separator, or an empty first
|
||||
# part): retry without the context.
|
||||
again, n2, dt2 = seed_x_chunk(tokenizer, model, text, target_lang, tag,
|
||||
kind=kind, source_lang=source_lang)
|
||||
again, n2, dt2 = seed_x_chunk(
|
||||
tokenizer, model, text, target_lang, tag, kind=kind, source_lang=source_lang
|
||||
)
|
||||
return again, n + n2, dt + dt2
|
||||
|
||||
|
||||
@@ -272,14 +314,20 @@ async def do_job(ws, job: Job, seed_x: SeedX) -> None:
|
||||
tokens = dt = 0
|
||||
for i, text in enumerate(job.texts):
|
||||
ctx = job.contexts[i] if i < len(job.contexts) else ""
|
||||
translated, n, t = seed_x_chunk(tokenizer, model, text, lang_name, job.lang,
|
||||
kind=job.kind, context=ctx)
|
||||
translated, n, t = seed_x_chunk(
|
||||
tokenizer, model, text, lang_name, job.lang, kind=job.kind, context=ctx
|
||||
)
|
||||
texts.append(match_punctuation(text, translated))
|
||||
tokens += n
|
||||
dt += t
|
||||
print(f"[{job.lang} {job.kind} {job.path or '/'}: {len(texts)} segments, "
|
||||
f"{tokens} tokens in {dt:.1f}s = {tokens / dt:.1f} tok/s]", file=sys.stderr)
|
||||
await ws.send(msgspec.json.encode(Result(lang=job.lang, key=job.key, texts=texts)).decode())
|
||||
print(
|
||||
f"[{job.lang} {job.kind} {job.path or '/'}: {len(texts)} segments, "
|
||||
f"{tokens} tokens in {dt:.1f}s = {tokens / dt:.1f} tok/s]",
|
||||
file=sys.stderr,
|
||||
)
|
||||
await ws.send(
|
||||
msgspec.json.encode(Result(lang=job.lang, key=job.key, texts=texts)).decode()
|
||||
)
|
||||
seed_x.idle()
|
||||
|
||||
|
||||
@@ -291,23 +339,34 @@ async def serve(url: str, seed_x: SeedX) -> None:
|
||||
try:
|
||||
async with websockets.connect(url) as ws:
|
||||
backoff = 1
|
||||
await ws.send(msgspec.json.encode(Hello(langs=sorted(SEED_X_NAMES))).decode())
|
||||
print(f"[connected; announced {len(SEED_X_NAMES)} language capabilities]",
|
||||
file=sys.stderr)
|
||||
await ws.send(
|
||||
msgspec.json.encode(Hello(langs=sorted(SEED_X_NAMES))).decode()
|
||||
)
|
||||
print(
|
||||
f"[connected; announced {len(SEED_X_NAMES)} language capabilities]",
|
||||
file=sys.stderr,
|
||||
)
|
||||
async for raw in ws:
|
||||
await do_job(ws, msgspec.json.decode(raw, type=Job), seed_x)
|
||||
except websockets.exceptions.InvalidHandshake:
|
||||
sys.exit("handshake rejected; check the URL (including the key)")
|
||||
except (OSError, websockets.exceptions.ConnectionClosed) as e:
|
||||
print(f"[connection lost ({e}); reconnecting in {backoff}s]", file=sys.stderr)
|
||||
print(
|
||||
f"[connection lost ({e}); reconnecting in {backoff}s]", file=sys.stderr
|
||||
)
|
||||
await asyncio.sleep(backoff)
|
||||
backoff = min(backoff * 2, 60)
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("url", help="full translator WebSocket URL including the key, "
|
||||
"e.g. ws://localhost:8410/_translate/KEY")
|
||||
p = argparse.ArgumentParser(
|
||||
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
|
||||
)
|
||||
p.add_argument(
|
||||
"url",
|
||||
help="full translator WebSocket URL including the key, "
|
||||
"e.g. ws://localhost:8410/_translate/KEY",
|
||||
)
|
||||
args = p.parse_args()
|
||||
if not args.url.startswith(("ws://", "wss://")):
|
||||
p.error("url must start with ws:// or wss://")
|
||||
|
||||
Reference in New Issue
Block a user