Compare commits

...
195 Commits
Author SHA1 Message Date
LeoVasanko 13fecd2118 Hreflang alternates on category placeholder pages too 2026-09-04 18:12:48 +00:00
LeoVasanko 11a138e19f Translate category-label titles, not just page titles 2026-09-04 18:09:40 +00:00
LeoVasanko ebd5911a38 Store DBIP database in folder where the program is ran, not where it is installed. 2026-09-04 18:04:30 +00:00
LeoVasanko db57125953 Public language selector, linked with the editor's language selection. Shared popup menu operation with edit toolbar with consistent closing logic. 2026-09-04 17:52:51 +00:00
LeoVasanko 986e28c220 Serve /favicon.ico as a redirect to the configured site icon 2026-09-04 14:59:38 +00:00
LeoVasanko 33a4a76364 Skip npm's security audit on install (registry endpoint stalls for minutes) 2026-09-04 02:03:10 +00:00
LeoVasanko 9fb4b5a681 Reject translations that would splice block-level Markdown 2026-09-04 01:43:29 +00:00
LeoVasanko 0c1349b037 Keep section anchors in the original language on translated pages 2026-09-03 23:59:10 +00:00
LeoVasanko 54f8c8e09b Carry literal < through translation as fullwidth < 2026-09-03 23:52:39 +00:00
LeoVasanko 78f4ddb2f0 Dispatch translations titles-first across languages 2026-09-03 23:03:47 +00:00
LeoVasanko e6a0c57446 Mirror content layout with text direction; admin stays LTR 2026-09-03 20:01:23 +00:00
LeoVasanko d6d07db2c5 Access-log extras: remote-user on /_api, visitor info on /_ws 2026-09-03 19:38:15 +00:00
LeoVasanko 38af57218a Translator API key management. 2026-09-03 19:23:56 +00:00
LeoVasanko eb2e8f8273 Raw access-log analytics storage with display-time classification
Replace the pre-classified store (visits/crawlers/abuse lists written at
collection time, plus abuse_ips and in-memory pending/session tables) with
a raw append-only log: one Get record per document GET (full path, true
HTTP status, referer origin, preload flag) and one Msg per /_ws activity
message.  Visitor/crawler/abuse classification, visit grouping (30-minute
inactivity gap), status/referer/UTM attribution and all aggregates are
derived in Store.display(), so future rule changes never invalidate stored
data.  The viewer payload keeps its exact shape.

Fixes structurally:

- Abuser 404s on slug-format paths showed up as "articles read": the
  not-found branch recorded the request twice through separate status
  plumbing.  Each request is now recorded once with its true status.
- 404 trail links never rendered red: cache-served navigations issue no
  GET and the only real GET (the idle preload) was discarded before status
  recording.  Preloads are now recorded with pre=True, never counted, and
  used for status attribution.
- formatAbuseRows merged a path's 404 probes and 200 reads into one entry;
  the collapse is now keyed by (path, status class).

Rule improvements enabled by the redesign:

- The plain-404 abuse threshold counts within a 1-hour sliding window, so
  long-time readers accumulating misses never classify (scanners spray).
- Hidden (admin) clients never trigger abuse classification: editing means
  visiting not-found pages.
- /.well-known/ probes (RFC 8615, e.g. Chrome devtools) are never abuse
  evidence; //foo-style empty path segments are an instant telltale.
- Visits can no longer open on an external exit URL; favicon fetches skip
  hidden clients' referers/exits.

Legacy analytics.json files are set aside as .bak-legacy on startup.
2026-09-03 18:52:16 +00:00
LeoVasanko 792b9e7aa9 fastapi-vue-setup 1.4.2 upgrade 2026-09-03 17:40:55 +00:00
LeoVasanko 4c6ab3dde6 Ruff format 2026-09-03 17:40:51 +00:00
LeoVasanko 9d70f17587 Pass CLI config to the app as JSON in PAGERITE_CONFIG 2026-09-03 16:14:59 +00:00
LeoVasanko 029bfe105e Show public https URL in startup box for non-localhost sites 2026-09-03 16:09:08 +00:00
LeoVasanko 62031fd5dd Move DB-IP download into app lifespan with proper logging 2026-09-03 16:08:15 +00:00
LeoVasanko 13cf716bb1 Silence httpx request logs, log favicon fetches in one line 2026-09-03 16:03:55 +00:00
LeoVasanko cbd50cfece Clip chart plot curves to the chart area with a per-chart SVG clipPath.
Past-week overlays can run far above the autoscaled y range, and the svg
is overflow: visible for the axis labels — wrap the data paths (bars,
skyline, area/line curves) in a clipped group so they cannot paint outside
the plot rect.
2026-09-03 15:46:42 +00:00
LeoVasanko 85ef968cd6 Reclassify sub-5s visits as crawlers, show crawler referers and abuser articles.
Visits with under 5 s of total reported reading time are JS-running bots:
display() converts them to crawler hits (one per trail page, with referer
and UTM query) and excludes them from every aggregate.  Crawler referers
join the favicon fetch origins and render with their icon in the crawler
table.  Abuse hits now record the real response status, so the abuse table
splits 404 probes from the articles the abuser actually read (200 GETs),
shown as trail links like the visitor/crawler tables.
2026-09-03 15:43:43 +00:00
LeoVasanko 7631c9f0a3 Prioritize titles in the translator dispatch queue.
All pending titles are now offered before any article chunks (stable
sort, menu order kept within each kind) — a page's name in the menu is
its most visible string.
2026-09-03 15:26:36 +00:00
LeoVasanko d7d03754d1 Record the acting user and use noun-based transaction actions.
Every request-driven kanta transaction now passes user= from the
Remote-User header the SSO/forward-auth proxy sets (the editor socket
reads it from the WebSocket headers; the translator worker keeps its
client key). Action labels are short identifiers naming the object, not
sentences: page / page:{lang} / page:title / page:{lang}:title /
page:language / page:slug / page:delete / structure:reorder / settings /
translate:reset for admin actions, translate:{lang}[ :title ] for worker
submissions (title results identified via the job kind, now tracked in
the connection state).
2026-09-03 15:24:32 +00:00
LeoVasanko 65fec6c519 Drop the startup translator-URL log line; the key and URL are shown in the site settings UI. 2026-09-03 14:42:02 +00:00
LeoVasanko dc55445ae0 Find link/formatting mark boundaries in translations by fuzzy word alignment.
Weight-ratio mapping alone was routinely off by a word and could glue a
mark to its neighbor (losing the space between). Now each mark's source
words are aligned to the translation's words by form similarity
(sequence ratio + shared prefix, case-folded, capitalization bonus) with
cheap skip penalties, so inflection, dropped articles/prepositions and
reordering don't break the match; slices are cut exactly at word
boundaries. Alignments without an anchor pair fall back to the weight
ratio (still the CJK path).
2026-09-03 14:38:59 +00:00
LeoVasanko b133ad6dd6 Fix image .margin positioning where caption and image got that layout instead of the figure wrapper getting it. 2026-09-03 14:12:35 +00:00
LeoVasanko 4413c7efdf Cleanup: break up the massive app.py into separate modules of manageable size. 2026-09-03 13:36:29 +00:00
LeoVasanko 2f533eaf09 Suppress mediapreview info logs. 2026-09-03 03:57:12 +00:00
LeoVasanko 1d46843c76 Add missing proxy path to vite config. 2026-09-03 03:55:09 +00:00
LeoVasanko d3e2196c83 Localization (#1)
Implement comprehensive content localization, admin panels for editing each language, AI translation interface with automatic updates when base language version is changed.
- SEO tags for all language URLs
- Uses accept-language by default, ?lang=en overrides temporarily
- User edits patched on top of translations
- RTL language supportReviewed-on: #1
2026-09-03 03:54:27 +00:00
LeoVasanko b6e6e46cfb Update URLs 2026-08-31 22:33:42 +00:00
LeoVasanko 7a4544731d Fastapi-vue-setup updated to 1.4.0: prettier logs. 2026-08-31 20:50:38 +00:00
LeoVasanko fd1987c9b3 Slight color change for inline code text to make it stand out of paragraph text. Not applied to list items and headings. 2026-08-30 03:40:03 +00:00
LeoVasanko 85a306296f Never break inline code spans (nowrap)
word-break: keep-all does not suppress breaks at hard hyphens (an
explicit UAX #14 break opportunity), so `--arg` could split after the
dashes. Inline code is short — nowrap is the safe fix.
2026-08-30 02:20:54 +00:00
LeoVasanko 2b9c7635d3 Figure-wrap lone images even with trailing inline attrs
Space-separated brace attrs on a lone image's own line (e.g.
{style="max-width: 40em"}) are consumed onto the paragraph but leave
an empty text token in the inline children, which defeated the
lone-image check — the image stayed a bare inline <img> without the
figure wrapper (and without lightbox/zoom treatment). Strip empty text
tokens before the check.
2026-08-30 02:20:54 +00:00
LeoVasanko 7e8e0a2ea7 Click-to-enlarge lightbox for article figures
Clicking a figure image opens a full-viewport overlay: the image as
large as fits (100vw, flex-shrunk to leave exactly the caption's
height) with the caption below. Click or any key closes it; wheel and
touch scrolling stop at the overlay (overscroll-behavior: contain).

Styling lives in the base theme: blurred dark backdrop (theme-tunable
via --lightbox-bg / --lightbox-text), soft shadow, reduced-motion-aware
open animation.
2026-08-30 02:08:06 +00:00
LeoVasanko 8fc5dd7b9b Keep margin boxes in the column flow, position them out of flow
::: aside / .margin blocks no longer split the column segments: they
stay inside the .colseg at their anchor point, so the surrounding text
is measured and laid out as a single columned layout (split-off sides
previously lost .cols when too short on their own).

The zone placements (multicol side zone, sidebar track, wide gutter)
now position the boxes absolutely off the article's left border —
unaffected by any column layout inside — with the vertical spot coming
from the unset top (where the box occurs in the text). Narrow widths
keep the in-column float fallback. Trade-off: out-of-flow boxes no
longer stack via clear, so boxes anchored close together may overlap.
2026-08-30 01:41:47 +00:00
LeoVasanko fbebddeaa9 Style inline code; drop explicit colorspace from color-mix
Inline code gets padding-inline, keep-all word-break, and a color nudged
30% toward --muted from the inherited color, so code inside
accent-colored text keeps its hue.

All color-mix calls now rely on the default interpolation space (oklab)
instead of declaring srgb/oklab explicitly.
2026-08-30 01:22:22 +00:00
LeoVasanko 31895065f8 Collect analytics over a /_ws WebSocket instead of POST /_a pings 2026-08-29 21:20:04 +00:00
LeoVasanko 447a565b05 Disallow /auth/ and /_api in robots.txt 2026-08-29 20:45:39 +00:00
LeoVasanko 399aa95d44 Encapsulate all migrations in kanta migrate_vN, drop persisted Data.version
migrate_v1 now also rebuilds the legacy flat pages store as the menu
tree (moved from the app lifespan, raw-dict level); migrate_v2 now also
backfills missing AVIF/WebP/JPEG derivatives on disk (moved from the
lifespan) and drops the obsolete version field. The version render
counter was cache-invalidation state, not database state: replaced by an
in-memory render generation that clears the page-body LRU and feeds page
ETags. The legacy Page struct and Data.pages/version fields are removed;
old databases lose the stale keys on re-serialization.
2026-08-29 20:16:13 +00:00
LeoVasanko f793d21c5e Serve images extension-less at /_f/{hash} with Accept-negotiated AVIF/WebP/JPEG
Uploaded images (SVGs rasterized, GIFs excepted) are stored as the
original (hash.orig.ext, internal only, never served) plus AVIF primary
and WebP/JPEG fallback derivatives re-encoded from it. Pages link the
bare hash; the server serves a format only when Accept lists it
explicitly (image/avif -> AVIF, image/webp -> WebP, else JPEG) with
vary: accept, while an explicit extension pins the format. Favicons go
through the same pipeline at 192px. migrate_v2 rewrites old
/_f/{hash}.avif article links, a startup backfill creates missing
derivatives, and twitter:image pins the .webp variant for X's scraper.
2026-08-29 19:46:47 +00:00
LeoVasanko fb3e6d1a04 Use runtime-only Vue build, raise chunk warning to 1.2 MB 2026-08-29 06:27:16 +00:00
LeoVasanko fd75a260b5 Insert images and tables as block-level fresh lines
uploadImage and insertTable no longer inject at the cursor: on a
non-empty line (e.g. inside an existing image tag) the block goes on a
fresh blank-separated line after it, never into it.
2026-08-29 06:09:51 +00:00
LeoVasanko 6058853341 Recompress uploaded images to thumbnailed AVIF via mediapreview
PUT /_api/files now runs raster uploads through mediapreview.dispatch
(temp file for format routing: pyvips, ffmpeg for HEIC/HEIF/AVIF),
storing the untouched original as <hash>.orig<ext> and serving the
AVIF derivative <hash>.avif in links. SVG/GIF and failed conversions
fall back to plain <hash><ext> storage. FileStore.delete removes the
whole hash pair. Adds mediapreview[standard] dependency.
2026-08-29 05:59:37 +00:00
LeoVasanko 9adc48479f Scroll page on cursor move only when cursor leaves viewport
Cursor-driven editor→page scroll sync pinned the cursor's page position
at a fixed window height, so every cursor move dragged the page along.
Now the page scrolls only when the cursor's mapped position crosses a
viewport edge margin, and just enough to bring it back inside.
2026-08-29 05:08:52 +00:00
LeoVasanko 864492b897 Article editor toolbar: toggling fences/links, class pickers, sizes, Tab indent
- Image insert always adds an empty "" caption with the cursor inside.
- Fenced blocks (``` code, ::: aside) share one toggle: clicked inside
  one it is removed and the content selected; otherwise the selection
  (expanded to whole lines) is wrapped, cursor left on the opener line.
- Code button: inline wrap toggles (selection preserved, backtick runs),
  line-spanning selections make fenced blocks.
- Link button toggles: clicked inside [label](url) it unwraps.
- Block classes via two pickers (placement, AA size) with current-class
  indication and a "normal" reset; placement replaces ::: container
  names, fences get their own attribute line.
- New .small/.large/.huge text-size classes (0.7/1.5/3em); base body is
  exactly 1rem so the scale is uniform.
- .left/.right floats generalized from figures to any block.
- Toolbar polish: non-emoji glyphs, scaled-up symbols, borderless
  hover/active states, reordered (B/i and size last).
- Editor panel no longer closes on Escape.
- CodeMirror editors capture Tab/Shift-Tab for indent/dedent.
2026-08-29 05:04:01 +00:00
LeoVasanko 515c6e1435 Block attrs space-separated at end of a text line
_block_attrs now applies a trailing {...} to the block when it ends the
last text line after whitespace (some text {.small}), not only on a line
of its own; a space is what keeps the braces off an image/link ending the
line. Glued-to-text braces stay literal.
2026-08-29 05:03:30 +00:00
LeoVasanko daa1670653 Themeable selection color via --selection-bg, shared by page and CodeMirror
Nitro's orange accent selection fill clashed with accent-colored text.
Introduce --selection-bg (base: accent 30% mix, as before) used by both
::selection and the editors' selection layer; nitro overrides it with a
neutral grey. CodeMirror's selection needed a baseTheme with its exact
&light/&dark selectors to win, and is now hidden when unfocused like a
normal input.
2026-08-29 03:39:24 +00:00
LeoVasanko e3fbce8ece Fix slug input: free typing with live space→hyphen and lowercasing
Live-filtering through slugify ate hyphens and spaces (trailing hyphens
are stripped) and jumped the cursor to the end on mid-text edits. Now
only spaces→hyphens and lowercasing happen oninput (length-preserving,
cursor stays put); the full slugify runs at commit before the server
call, which keeps its own validation.
2026-08-29 03:14:57 +00:00
LeoVasanko 775dc3f65a Look up themes and theme per-file overrides in system, user and site directories. 2026-08-29 02:41:29 +00:00
LeoVasanko d21c3edbfc Indicate if themes support light or dark modes, or both 2026-08-29 01:41:16 +00:00
LeoVasanko 794e1ba26e Update README, more screenshots. 2026-08-29 01:26:58 +00:00
LeoVasanko b8d2b9f32c Add theme collage screenshot. 2026-08-29 00:49:17 +00:00
LeoVasanko ad03216e93 Pandoc-style brace attributes on code fences (```{.python .wide #id}) and nameless fenced divs (::: {.aside}) 2026-08-28 23:37:26 +00:00
LeoVasanko 1d64d1e653 Lists clear floated figures: flow-root on custom-marker lists so markers no longer overlap a left float 2026-08-28 23:30:44 +00:00
LeoVasanko 72e722c094 README rewrite with positioning and run instructions; production setup guide 2026-08-28 23:23:51 +00:00
LeoVasanko 22c4f0e752 Layout: breakable paragraphs across columns, sidebar gets its own flexible track at every width, cols requires multiple paragraphs; h3 size fix in corporate theme 2026-08-28 23:23:51 +00:00
LeoVasanko 5101a5c5bf Seed: rework gallery article (right-floated figure, aside box, design section), drop md-chase.jpg, default theme corporate 2026-08-28 23:23:51 +00:00
LeoVasanko c801fdc929 Make the page transition configurable, add crossfade/slide/reveal
The cube view-transition block moves out of pagerite.css into transition
designs (pagerite/themes/{name}/transition.css), handled like banner
designs: served from disk at /_themes/..., injected by the backend as
#pagerite-transition after the banner sheet, and picked site-wide in the
site settings (Data.transition, default cube; GET/PUT /_api/settings
carry transition + the designs found on disk). The site editor swaps the
sheet in place on change; the dev-mode order fixup knows the new id.

Designs: cube (unchanged), crossfade (all navigations), slide (sideways,
both pages moving together) and reveal (clip-path wipe over the
stationary old page) — the latter two mirrored on history-back and
crossfading within a section like cube.

Also compress the site settings form into a two-column grid: site name +
favicon on one row, theme + transition on the next.
2026-08-28 20:04:21 +00:00
LeoVasanko 8df6594432 Fix cube transition direction on forward history navigation
popstate always passed back=true, so going forward after back still
rotated the cube backwards. Track an incrementing idx in history.state
and only mirror the transition when the target entry is actually behind;
replaceState calls now preserve the state object instead of wiping it.
2026-08-28 19:35:07 +00:00
LeoVasanko 7743639812 List markers: balanced indent via a centered full-indent box (--list-indent), ordered-list counters, task checkboxes seated as markers with label-wrapped toggle text, summer ⚜️ bullets. 2026-08-28 19:26:51 +00:00
LeoVasanko f659f5f6be Sibling-combinator top margins for h1/h2: no gap when first in a container (aside, admonition, colseg), h1 + h2 stacks tighter, page-top headings keep their banner spacing. 2026-08-28 18:01:00 +00:00
LeoVasanko f962d70e3b Section edit pens, cursor-driven piecewise scroll sync, URL-following editor.
Anchored h1/h2s carry data-line (markdown source line, top-level
headings only) and h2s get dimmed section pens that open the page
editor at that section. Editor scroll sync is now piecewise-linear
keyed on those anchors: the page follows the cursor (fractional,
wrap-aware, fixed window anchor; editor scroll no longer drives it),
the editor follows page scroll with a progress-based viewport anchor,
so document ends line up exactly in both directions. The editor always
follows the URL — fetch-navigation retargets it — with unsaved text
stashed per path for the session and restored on return. Banner pen
removed (the tab stays in the shell); top-right order is now
analytics, site settings, login/logout.
2026-08-28 17:42:35 +00:00
LeoVasanko 6910ccad73 Navigable section anchors: slug ids + self-links on h1/h2, scroll-tracked location hash.
Long-enough articles (3+ in-body h1/h2) get slug ids (python-slugify,
mirroring the editor's slugify.js) and the heading text becomes a
self-link (a.anchor) for clean link copying; {#id} always wins,
duplicates get -2/-3 suffixes, h3+ never. The implicit page title is
now injected as '# {title}' into the markdown (render(title=...)), so
it takes the same first-h1 path as an explicit one: no id, href=""
self-link scrolling to top, not counted toward the threshold. The
.body wrapper is gone — segments are direct article children — and the
editor preview swaps the whole article in one go. pagerite.js tracks
the reading position in the location hash (last tagged heading above
the viewport middle; cleared at the top, never on unscrollable pages)
and scrolls to anchors after fetch-navigation.
2026-08-28 16:12:17 +00:00
LeoVasanko 2060dc815c Per-hostname data directory <hostname>/{content.kantadb, analytics.json, files} from CLI arg; drop learned site_url.
The first positional CLI argument (default localhost) names the site's
public hostname and its data directory under the cwd, replacing the
CWD-relative pagerite.* files. The public origin (https://<hostname>)
is now authoritative configuration instead of a value learned from
admin browsers: POST /_api/site-url, Data.site_url and the pagerite.js
reporter are removed, and page rendering, sitemap, robots.txt and
analytics own_origin all use SITE_URL consistently (localhost falls
back to the request's base URL).
2026-08-28 14:57:04 +00:00
LeoVasanko ffa7200a6e Automatic navigational cards on index pages, replacing the sidebar menu. 2026-08-28 00:25:10 +00:00
LeoVasanko bdac75e70a Extend initial edit time window to 48h. 2026-08-27 22:52:51 +00:00
LeoVasanko a116d78951 Flood fill images in aside. 2026-08-27 22:42:20 +00:00
LeoVasanko 825379ca2a Long articles become a bounded composition: fluid text lanes with a side zone for marginalia, centered in vacant space that grows with the viewport instead of stretching the content. 2026-08-27 22:15:17 +00:00
LeoVasanko ddf323bf41 Render the column layout structure on the backend, cap at two columns, add a left-margin breakout.
render() now segments the body into .colseg wrappers and flags .multicol
itself, replacing the fragile colseg injection in pagerite.js. Columns
are capped at two; shrink-wrapped figures left-align inside columns.
New {.margin} breakout (and ::: aside) drops blocks into the left gutter
on wide viewports, falling back to in-column floats.
2026-08-27 18:36:49 +00:00
LeoVasanko f708e1dbca Avoid breaks inside lists. 2026-08-27 17:24:08 +00:00
LeoVasanko 1d0bc3f59d Renewed analytics format. The hide flag moves to client, and we still track but hide more robustly, to avoid noise from admins checking out their own site. 2026-08-27 16:17:02 +00:00
LeoVasanko 60d5155fd2 Don't upper case lower headings. 2026-08-27 16:14:04 +00:00
LeoVasanko 3518ac9ac7 Fixes to markdown extensions/styling. Sort crawlers most recent first. 2026-08-27 01:27:49 +00:00
LeoVasanko 5c2433b766 Correct font scaling by x size. Fix code receiving twice a smaller font size. 2026-08-27 00:15:47 +00:00
LeoVasanko 2959f971bc Implement support for {} attrs on code blocks (line right after the closing fence) and containers (::: aside {...}). 2026-08-27 00:08:51 +00:00
LeoVasanko 8ee22de060 Many more Markdown extensions and formatting improvements. 2026-08-26 23:51:05 +00:00
LeoVasanko 9be1491f0e Improved mobile layout. Banner section gets smaller and editor goes full screen. 2026-08-26 18:44:17 +00:00
LeoVasanko 29c1dc64ed Transition graph shows per article read times in minutes and seconds. 2026-08-26 18:27:09 +00:00
LeoVasanko a27958117f Maintain nitro orange line in theme.css rather than banner.css so that banner changes don't remove it. 2026-08-26 17:40:23 +00:00
LeoVasanko 56fd5f854d Selection color from base theme. 2026-08-26 17:03:01 +00:00
LeoVasanko 932aee4404 Fixes to eyes banner positioning. 2026-08-26 16:20:58 +00:00
LeoVasanko ca1c0d5a6f Summer theme gets seamless transition between banner and page. Banners updated with support for this mode. 2026-08-26 15:58:02 +00:00
LeoVasanko 9cbc4ab59d Smoother scroll effects on banners using --pry. 2026-08-26 15:09:29 +00:00
LeoVasanko 7aa2abbf8f Analytics pings with query args, cleanup. 2026-08-26 14:53:09 +00:00
LeoVasanko eddb3f22f3 Adjusted connector thickness and bead animations to better deal with highly varying rates. 2026-08-26 14:34:10 +00:00
LeoVasanko 1702281dbf Fix edge condition for gaussian smoothing. 2026-08-26 14:00:29 +00:00
LeoVasanko 38b894a789 Layout transition lags fixed with better editor panel handling. This was causing .wide sections visibly lag behind on viewport width changes. 2026-08-26 13:47:51 +00:00
LeoVasanko 06409cdaac Drop page caches on edits that may cause changes to navigation, theming etc. 2026-08-26 13:19:17 +00:00
LeoVasanko 7ba6094360 Proper cube transition background colors via base CSS mixing in theme bg. 2026-08-26 13:01:42 +00:00
LeoVasanko c59ab2f058 Fix editor-page scroll syncing issues. 2026-08-26 02:29:33 +00:00
LeoVasanko 078e80591b Maintain normal column layout while editing. 2026-08-26 02:29:14 +00:00
LeoVasanko 6563eebdba Improve list bullet positioning. 2026-08-25 22:49:02 +00:00
LeoVasanko cac74ebb6b Remove uvicorn server header. 2026-08-25 22:40:36 +00:00
LeoVasanko 8b1bb6f994 Avoid column break right after heading. 2026-08-25 22:31:52 +00:00
LeoVasanko 3173a113b6 Improved theme banner consistency across themes. 2026-08-25 22:29:28 +00:00
LeoVasanko a21f919601 More robust admin check, only adding edit pens after probe completes. 2026-08-24 21:51:35 +00:00
LeoVasanko 158b5f1961 Editor panel layout fixes, use body scroll and bidirectional sync with editor. 2026-08-24 21:41:33 +00:00
LeoVasanko f9ce0a4f3b Smarter initial analytics page when there is no visitor data yet. 2026-08-24 21:13:20 +00:00
LeoVasanko 2d595f8c15 Reduce #brand to actual size, avoid clicks on empty banner space touching it. 2026-08-24 20:58:50 +00:00
LeoVasanko b1fc8e24d7 Eyes banner follows taps not just mouse. 2026-08-24 20:54:50 +00:00
LeoVasanko c5cf799f68 Better nav layout for portrait phones. 2026-08-24 20:43:18 +00:00
LeoVasanko a5edc7b3b6 Fix crawler misclassification from /_a and orphan counts on visit scrub
Two analytics corrections verified against the production capture:

- pagerite.js suppressed pings with fr == '/_a', but fetch-navigation
  away from the analytics page had already GET-ed the target without the
  preload header; the orphaned pending hit then flushed to the crawler
  list, classifying a real user as a crawler. Navigations away from /_a
  now ping normally (the server rejects /_a as a target regardless, and
  admin noise is already handled by hide=1).

- _remove_visit only reversed the visit's creation counts, leaving
  views/transitions from later pings behind as orphans on the graph with
  no matching row in the visitor table. An in-memory per-visit count log
  now tracks every count event, so an admin hide=1 scrub reverses the
  visit completely.
2026-08-24 20:18:13 +00:00
LeoVasanko 09ebc63690 Record response status per path; mark 404 trails red in the viewer
Document GETs now stash their status (200/404) in a pending table,
consumed by the matching ping: visits gain a per-path statuses map and
crawler hits a status field. Trail links with a 404 status render in
red with the status code in the tooltip, alongside the read time.
2026-08-24 19:43:23 +00:00
LeoVasanko 2868843028 Fix inverted client filter in admin hide ping
The hide=1 branch kept the admin's own pending crawler hits (==) instead
of discarding them (!=), so an admin's document GETs flushed to the
crawler list 10s later while every other client's pending hits were
wrongly dropped. This is why ordinary admin browsers showed up as
crawlers.
2026-08-24 19:40:12 +00:00
LeoVasanko 6fcebaea3f Transition map: svg-scaled fonts, border-clipped pill text, tight crop
- Fonts scale with the svg instead of the --u constant-screen-size
  compensation (ResizeObserver machinery removed)
- Larger node text (slug 19px, count 15px)
- Pill labels no longer ellipsis-truncated: text is clipped at the pill
  border via per-node clipPaths; captions center when they fit and anchor
  left on overflow so the title's beginning survives; count lines stay
  centered
- Bounding box crops to pill half extents plus the ribbon halo instead of
  the diagonal radius, removing the large top/bottom margins
2026-08-24 18:35:36 +00:00
LeoVasanko 4a48d08a19 Analytics layout: larger charts with natural-width cap, one-line totals
- Charts grow to a larger intrinsic size (1052x174) and never upscale
  past it; centered with equal side margins above the cap, full width
  below, svg always within page bounds; overflow visible so wider fonts
  don't clip at the viewBox edge
- Totals row aligns its left edge with the charts and shrinks (gap first,
  then font) via container units to always stay on one line
2026-08-24 18:21:58 +00:00
LeoVasanko e3e29251ca Transition map: cull invisible connectors, stable beads, lane labels
- Cull connections whose thin middle would render below ~0.8px
  (MIN_WMID); drop external source/exit nodes whose connectors are
  all culled, while site page nodes always stay
- Bead simulation persists across data reloads: emitters keyed per edge
  direction, beads tracked by progress, so unrelated count changes no
  longer reshuffle bead positions
- Bead speed relative to span length: constant 1.5s traversal per edge
- Top lane labeled with a house icon; all lane labels left-aligned just
  past the source pill (half height on near-vertical branch lanes), with
  guides running to the lane end so long slugs are never truncated
2026-08-24 17:58:02 +00:00
LeoVasanko 8affc41289 Tighter analytics chart chrome
- Unified chart text at 11px system-ui; fixed size independent of theme font
- Day view y axis reads "visits / 5 min" / "views / 5 min"
- Left margin and y tick spacing tightened (MARGIN_L 56 -> 40)
- Gap between visits and views charts removed
- Legend repositioned for the larger font
2026-08-24 17:26:30 +00:00
LeoVasanko 2de4717230 Fix week overlay alignment, in-plot ISO week legend, shorter charts
- weeklySeries shifts overlaid weeks onto the current week's time axis so
  they overlay inside the plot instead of overflowing left; oldest weeks
  paint first, current week on top
- Legend moved inside the visits chart's top right: current ISO week in
  accent, past weeks as a single muted "Week M" / "Week M–N" specimen
- Past week curves use the muted color instead of faded accent
- Chart height reduced ~30% (180 -> 126)
2026-08-24 17:03:49 +00:00
LeoVasanko 563e8fcaf2 Rework analytics chart scaling; self-contained SVG charts
- Charts render as single SVGs with axis labels inside the viewBox,
  replacing the stretched plot + HTML overlay labels
- Rolling ranges end at now, t0 aligned to UTC day; bucket size follows
  the window (6h up to 31 days) so "all" at its 30-day minimum renders
  identically to "month"
- X labels always centered on their true position; no edge-align shifting
- rangeWindow simplified to rolling spans ending at now
- TransitionGraph "all" visual scale floored at the 30-day plot minimum
2026-08-24 12:31:25 +00:00
LeoVasanko 921a5484a2 Fixed-sigma smoothing of traffic history plots. 2026-08-24 05:54:43 +00:00
LeoVasanko 0e2e52fa45 Use last 24h/7d/30d/365d/all analytics data. Previously some fields were unfiltered and weekly view was based on calendar weeks. 2026-08-24 05:42:32 +00:00
LeoVasanko 075848f782 Crawlers should include all sorts of spiders along with bots and googleother. 2026-08-24 05:28:05 +00:00
LeoVasanko 7b8899af92 Transition map: bounded node scaling, concentric branch lanes, exit row at bottom 2026-08-24 04:47:06 +00:00
LeoVasanko 6d2ae104d7 Fix analytics classification: ignore bot-UA pings, skip preload GETs.
JS-running crawlers (Googlebot, GoogleOther, Applebot) execute pagerite.js
and send navigation pings, registering as visitors. Pings whose User-Agent
matches _is_bot_ua (any "bot" token plus listed exceptions) are now
ignored, so their document GETs flush to the crawler list as intended. No
source verification: a spoofed bot UA merely lands in the crawler stats,
and path-based abuse classification catches scanners regardless.

Idle-time link preloads from pagerite.js were queued as pending crawler
hits and flushed to the crawler list whenever the user navigated more than
10s later, so real visitors' subpage loads showed up as crawler hits.
Preload fetches now carry an x-pagerite-preload header and the document
GET handler skips tracking for them; the ping sent on actual navigation
does the counting.
2026-08-22 18:23:38 +00:00
LeoVasanko b7fc543a83 Transition graph layout follows navigation. 2026-08-22 18:10:44 +00:00
LeoVasanko b868033ddc Pill shaped nodes 2026-08-22 15:48:33 +00:00
LeoVasanko 8aad64cced Analytics layout update, larger, consistent text sizing. 2026-08-22 14:45:56 +00:00
LeoVasanko 319163ee7e Page caching and zstd compression. Avoid useless fetching. Mobile layouts of navigation menus improved. 2026-08-22 14:05:12 +00:00
LeoVasanko 87b16b7144 Implement /robots.txt and /sitemap.xml. Update dev proxy to all-by-default. 2026-08-22 12:13:17 +00:00
LeoVasanko 20ae6501f2 Add dynamic /sitemap.xml and /robots.txt endpoints 2026-08-22 12:01:18 +00:00
LeoVasanko a16fe88114 Auto select day if less than 24h data for new sites. 2026-08-22 01:51:36 +00:00
LeoVasanko c77598adc7 Fine tuning date formatting. 2026-08-22 00:09:57 +00:00
LeoVasanko 29f8fac013 Slightly prettier analytics URL 2026-08-21 23:58:00 +00:00
LeoVasanko 6199e5a69e Support for UTM tags in transition graph as source sites. 2026-08-21 23:50:01 +00:00
LeoVasanko 375b4b6bdb analytics: unify visitor cell across visits, crawlers and abuse tables 2026-08-21 23:32:46 +00:00
LeoVasanko fdb3e42d6f analytics: shared Client struct, grouped abuse paths, unified visitor cell 2026-08-21 23:16:15 +00:00
LeoVasanko 51a6a16221 Neater abuse table formatting. 2026-08-21 22:31:37 +00:00
LeoVasanko 0798e24d24 Desaturated house emojis 2026-08-21 22:03:55 +00:00
LeoVasanko 3be2d08ac9 SI formatting of large visitor numbers. 2026-08-21 21:36:57 +00:00
LeoVasanko b7d5b23ae6 Cleaner formatting of utm tags in visitor table. 2026-08-21 21:22:16 +00:00
LeoVasanko 6eaa1c1a8b analytics: 24h day view with bar chart for precise realtime stats. Tables redesigned with cleaner layout. Tracking article read times. Adjust connection graph visualizations by time range. Other cleanup and supporting systems. 2026-08-21 20:14:04 +00:00
LeoVasanko f341d22aa0 Improved fake traffic generation with abuse bots, utm tags etc. 2026-08-21 20:10:41 +00:00
LeoVasanko 1a479ceb24 Add --dbip CLI flag to auto-download/update the DB-IP MMDB database.
Downloads the latest dbip-city-lite-YYYY-MM.mmdb.gz before starting the
server, skipping when the local database is current, falling back to the
previous month on 404, and removing older databases after an update.
Promotes httpx to a runtime dependency.
2026-08-21 03:09:04 +00:00
LeoVasanko ff553d018a Default scheme, host and port for fake_traffic script. 2026-08-21 02:52:53 +00:00
LeoVasanko c807d48a13 Add more external content in seed data. 2026-08-21 02:50:36 +00:00
LeoVasanko 9c383c1c8b Change default port mapping to 8100/8200/8210 (prod/vite/dev). Vite gets different port to avoid caching problems when switching between it and prod. 2026-08-21 02:49:18 +00:00
LeoVasanko 462e995adc Add external link (referer/outgoing) display on connection graph. 2026-08-21 02:44:54 +00:00
LeoVasanko ea069b98da Fix analytics app not mounting on fetch-navigation to /_a
load() queried the live document for the pagerite:analytics-src meta,
but the swap never touches <head> — the meta only exists in the fetched
doc, so the app never mounted unless /_a was loaded directly. Also cache
the fetched HTML so the post-swap preload doesn't re-GET the page we
just navigated to.
2026-08-21 01:53:39 +00:00
LeoVasanko deb5419c47 analytics improvements:
- keep visitor charts y-axis minimum range at 10
- keep 'all' chart x-axis minimum span at 30 days
- group crawler hits by (ip, ua) and list top pages visited, show crawler page load counts as N× prefix
- store and display geoip city, keep geoip country overwrite
- stream live updates over WebSocket /_api/ws/analytics
- include family ring arcs in transition map crop bounds
- remove top UA summary, limit crawlers to 10 and visits to 20
- human-readable relative timestamps with UTC tooltip
2026-08-21 01:36:58 +00:00
LeoVasanko 242b62784c Add fake traffic generator script
Uses Playwright to drive Chromium through real internal link clicks,
so pagerite.js analytics pings create normal visits. Also fires HTTP
GETs with crawler user-agents to record crawler hits.

Features:
- script-local deps via uv add --script (playwright, httpx)
- rotating pool of real public IPs via X-Forwarded-For
- Poisson inter-arrival delays between sessions/hits
- configurable browsers, crawlers, clicks, and dwell time
2026-08-20 23:37:02 +00:00
LeoVasanko f78229bd30 Refactor analytics to /_a instead of under article pages. 2026-08-20 23:16:51 +00:00
LeoVasanko 7f4bc8efa4 Ignore devserver health probe in analytics tracking
The devserver polls /?from=devserver.py to check backend readiness.
Without this exclusion each poll is recorded as a crawler hit. Only
exclude the exact case: front page, that query string, and 127.0.0.1,
so remote visitors cannot hide traffic by copying the parameter.
2026-08-20 22:16:11 +00:00
LeoVasanko 3100010335 Transition map: count-scaled edges, bead flows, external links
- Edge widths grow logarithmically with the connection count (~1 px at
  a single count, uncapped); connections below 1% of total traffic are
  pruned, bounding the graph to ~100 edges.
- Beads: per-direction flows emitted at a rate linear in the count,
  each bead simulated independently in JS (no in-flight limit), offset
  onto right-hand lanes so opposing flows don't collide, running under
  the node circles with a glow.
- External links: referer origins as a node row above the map, exit
  origins fanned outwards from their source page.
- Transitions are now stored per 5-minute bucket (sparse
  from -> to -> bucket -> count) so the graph filters by time range
  like the other series; legacy analytics files are discarded.
2026-08-20 22:06:29 +00:00
LeoVasanko 7556bb7f4f analytics: add crawler tracking, pretty UA/IP display and copy-to-clipboard 2026-08-20 20:31:04 +00:00
LeoVasanko 1a1a21712e Add country flags. 2026-08-20 19:47:56 +00:00
LeoVasanko f0e6162f02 Extended analytics data collection. 2026-08-20 19:40:05 +00:00
LeoVasanko b4e8fad090 Implement analytics feature
Add server-side visit analytics collection, a public-page ping endpoint,
and a full-screen AnalyticsView for admins.

Backend:
- Add pagerite/analytics.py: Analytics/Visit model, Store, and persistence
- Wire /_a ping endpoint and GET /_api/analytics into pagerite/app.py

Frontend:
- Add full-screen AnalyticsView with visitor charts and transition map
- Add VisitorCharts and TransitionGraph subcomponents
- Add analytics JS helpers in frontend/src/analytics/
- Send navigation pings from frontend/src/pagerite.js
- Mount AnalyticsView from frontend/src/main.js
- Document the feature in docs/analytics.md and update AGENTS.md
2026-08-20 18:43:57 +00:00
LeoVasanko 11f8de2df5 Fix pretty scrollbars not appearing in production. 2026-08-19 16:33:44 +00:00
LeoVasanko 81f08e7760 Updated docs 2026-08-19 16:11:42 +00:00
LeoVasanko 1d65a57fdf SEO/social meta for content pages; full-height site editor
- views.py: description, canonical, Open Graph and twitter:card tags
  from heuristics over the rendered article — first paragraph as
  description, share image prefers a {.hero} image, then first raster,
  then first SVG; first <video> becomes og:video; published/modified
  times from the node. Absolute URLs from the request base.
- SiteEditor: panel fills the full window height; the brand-HTML and
  custom-CSS CodeMirror windows grow to share leftover space equally
  instead of fixed max-heights.
2026-08-19 15:52:24 +00:00
LeoVasanko e175b39f35 Full height site editor panel. 2026-08-19 02:19:30 +00:00
LeoVasanko 178d22c05c Seed: eyes design on showcase, whale/wave art, richer long read
- The showcase category picks the eyes banner design (inherited by its
  pages; night-sky overrides with stars); the seeder now leaves
  content=None for entries with empty markdown so it stays a category
  label.
- Public-domain demo images in pagerite/seed-assets/ (Hokusai's Great
  Wave, Shute's 1892 Moby-Dick engravings), shipped via build artifacts.
- Welcome figure uses {width=420}; Gallery gets the Great Wave as its
  {.wide} piece; Loomings gains h2/h3 sections and the engravings.
2026-08-19 02:00:54 +00:00
LeoVasanko dfd3ef1dad Shorten seed doc titles; full names as h1 in page source
Menu entries read Basics / Extensions / Images and Layout; the pages
open with their own '# Markdown Basics' / '# Markdown Extensions'
headings (a markdown h1 suppresses the title h1, so nothing doubles).
2026-08-19 01:40:50 +00:00
LeoVasanko 512ed91b31 Bottom-anchor banner artwork, instant navigation via in-memory page cache, seed/stars/structure tweaks
- Fix banner artwork sizing: explicit 100% grid track so children
  stretch instead of resolving height:100% against a content-sized row
  (SVG intrinsic ratio bloated the row, cropping the artwork's bottom).
  Bottom-anchor via object-position, transform-origin and YMax slice.
- pagerite.js: in-memory page cache — preload every visible internal
  link once, serve navigation from memory without fetching; editors'
  loadPlain keeps the cache in sync (pagerite:page-fetched).
- Structure editor: delete pages directly, no two-step confirmation.
- New 'stars' banner design (drifting starfield) alongside 'eyes'.
- Rewrite seed content: welcome page, three-level docs section covering
  all Markdown features (source + rendered), showcase hierarchy with
  image positioning and a simple leaf-page banner example.
2026-08-19 01:40:26 +00:00
LeoVasanko 3b074c02ed Seed demo content only on database creation via @kanta.bootstrap
Previously startup appended any missing seed pages into existing
databases, resurrecting deleted content on running sites.
2026-08-19 00:57:35 +00:00
LeoVasanko 58e5d3ef2b Replace Paskia iframe auth flow with plain /auth/ links
Paskia does not support being iframed; the login/profile buttons are now
plain anchors, and a pageshow handler re-probes auth when history
navigation restores a cached page. Drops the paskia JS dependency.
2026-08-19 00:57:35 +00:00
LeoVasanko 5769747a95 Disable ligatures in CodeMirror editors
CodeMirror measures text per character; Fira Code's ligature glyphs
render wider than the measured sum of their parts, corrupting cursor
and selection rendering. The font is set on .cm-content with ligatures
off, and inner spans inherit it for consistent metrics.
2026-08-19 00:44:40 +00:00
LeoVasanko 180da893cb Hide the https:// scheme in autolinked URL text
Bare (GFM/linkified) and angle autolinks render without the https://
prefix; http:// and other schemes stay visible, and manually labelled
links keep their label.
2026-08-19 00:27:07 +00:00
LeoVasanko a9693e1dcb Nested sidebar nav, banner inherit label fix, editor UX polish
- Sidebar renders the section's whole subtree as nested lists (third
  level and deeper, article-list-style markers); the one-item sidebar
  rule yields when that item has children of its own; only the viewed
  page is highlighted (navbar keeps ancestor highlighting); first_leaf
  skips unpublished branches
- Banner design selector: inherit option names the resolved design and
  its true source (backend banner_design_source excludes the node's own
  setting; doc payload carries banner_design_inherited)
- Article editor: fenced code block on empty lines, table size picker
  popup, borderless save icon disabled by saturate(0) alone
- While editing, window scroll is locked and only #main scrolls; the
  panel exactly fills the available window height; editor scroll drives
  the article
2026-08-19 00:25:07 +00:00
LeoVasanko 4dac16a9ac Unify admin editors into a tabbed EditorShell (site, structure, article, banner)
- One mounted shell with four kept-alive tabs; pens open/switch tabs,
  closing only hides the shell so unsaved page-editor state survives
  until a real reload; admin panels never reload the page, regions are
  refreshed in place via the shared swapdoc.js helper
- SiteEditor split: site settings vs. the new StructureEditor tree tab;
  BannerEditor split out of the old site editor
- Article editor format bar (bold/italic/code/link/table/image) with a
  table size picker and Ctrl/Cmd-B/I/S bindings; fenced code block on an
  empty line; icon-only save button disabled (saturate(0)) when clean
- Media uploads use icon buttons; favicon upload by clicking the preview
  tile; placeholders reserved for actual defaults in effect
- Banner/site pens and auth buttons grouped in an .editor-pens container;
  page pen no longer shown to anonymous visitors
- While editing, window scroll is locked, the panel exactly fills the
  available window height, and only #main scrolls; editor scroll drives
  the article
- Banner design inherit label names the design and its true source;
  backend banner_design_source excludes the node's own setting and the
  doc payload carries banner_design_inherited
2026-08-19 00:14:05 +00:00
LeoVasanko 43b841ebb6 Add a mobile layout (48rem breakpoint) and the missing viewport meta
Pages now declare width=device-width, so phones render at real device
width instead of a scaled-down 980px layout viewport. Below 48rem the
content becomes a single column with the sidebar lifted above the
article as a wrapping link strip; floated .left/.right figures fall back
to plain centered figures (explicit img widths still shrink-wrap) and
.wide keeps its centered full-viewport bleed.
2026-08-18 22:33:11 +00:00
LeoVasanko f2715088d2 Add site-wide custom brand HTML with media upload
Data.brand_html replaces the brand link entirely (rendered in a #brand
div on top of the banner, next to the nav); editable in the site editor
via a small HTML CodeMirror with image/video upload and paste, live
preview, and debounced save through /_api/settings.
2026-08-18 22:20:38 +00:00
LeoVasanko f8cdfd53c0 Add admonitions plus autolink and sub/superscript markdown extensions
!!! note/warning/tip/... blocks (mdit-py-plugins admon) render as
lightweight callouts: accent bar + faint wash in the blockquote idiom,
recolored per type (tip green, warning pink/red via --admonition-color),
rounded and softly shadowed in the summer theme. Also enabled
gfm_autolink (bare URLs) and H~2~O / x^2^ sub/superscripts.
2026-08-18 22:00:00 +00:00
LeoVasanko e9297cac61 Add the summer theme: illustrated meadow banner, cohesive palette
A light, playful theme whose every color is sampled from its banner.svg
meadow scene (sky/grass/sun/flower pink): fixed daylit landscape page
background with a sun-glow echo, tilted gradient brand, flower bullets,
a dense grass strip (grass.svg, served via /_themes) under h1s, rounded
shadowed tables with a sunny hover sheen, and a layered-parallax banner
(sun rises, clouds drift, nearer hills move less) with idle animations
wrapped in prefers-reduced-motion.

The /_themes route now serves any theme asset (previously only
theme.css/banner.css) so CSS can reference sibling files like grass.svg.
2026-08-18 21:54:44 +00:00
LeoVasanko 2140f06501 Restyle tables and definition lists in the base theme
Tables drop the hard 1px grid: a soft accent-tinted vertical gradient
header (overridable via --table-head-a/b), a very faint diagonal
--table-tint wash lifting the cells apart, and only a whisper of a row
separator. Definition lists become a lightweight two-column grid
(minmax term column, explicit grid-column placement so runs of
dt/dd stack right) with accent terms and no borders or fills. Tables
also join the bottom-margin-only spacing convention.
2026-08-18 21:54:37 +00:00
LeoVasanko a75ad6f5da Overlay scrollbars via OverlayScrollbars
overflow: overlay is dead in current Chromium, so an appearing native
scrollbar shifted the layout again. OverlayScrollbars replaces it with
floating auto-hiding scrollbars that never reserve space (native window
scroll APIs unaffected on a body target). Themed via --os-* variables in
pagerite.css; the html scrollbar-color styling stays as no-JS fallback.
(package-lock.json is gitignored.)
2026-08-18 20:30:08 +00:00
LeoVasanko 7f140001fa Simplify image/figure styling: markdown always renders standalone images as figures
Standalone markdown images (captioned or not) now become block <figure>
elements; only inline-with-text images and raw author <img> HTML stay
plain. This collapses the img/figure selector duplication in pagerite.css
into a default + override structure: figure fills the column, floats go
30% (1em text gap), an explicit width attribute shrink-wraps the figure
(images with width are left untouched by CSS so the attribute hint
survives), and .wide re-anchors are grouped with the main rule.
2026-08-18 20:29:56 +00:00
LeoVasanko d87e5ae323 Update page title to indicate editing flows 2026-08-18 18:29:33 +00:00
LeoVasanko 14655ee4d7 Ship pagerite/themes in the built package
hatchling's only-packages=true drops directories without an __init__.py,
so pagerite/themes never made it into the wheel/sdist (this also means
the banner.svg artwork was never packaged). Force-include it as a build
artifact like frontend-build.
2026-08-18 06:47:58 +00:00
LeoVasanko 0ec9be330e Add the 'eyes' banner design; seed banners only on select sub pages
Banner designs can now ship arbitrary markup as banner.html (canvas +
style + script), taking precedence over banner.svg; the artwork is
inlined in a div[data-design] wrapper either way, which the editor's
live preview preserves (and never re-runs its scripts).

The bundled 'eyes' design (themes/eyes/banner.html + banner.css sizing
the stage) replaces the eyes canvas script that was embedded in the
notes-on-urls page's own banner — the page now just picks the design.
Seeds set banners only on select sub pages (the-long-read gradient,
canvas-nights stars): the front page no longer gets a banner image and
shows the theme's default design. Seed entries carry a banner_design
field.
2026-08-18 06:36:05 +00:00
LeoVasanko c5a91edac7 Fix theme re-selection after switching to none
Re-creating the #pagerite-theme link anchored to #pagerite-base, which
does not exist in dev (base CSS is a Vite-injected <style>), so the
fallback prepended it before the base styles and the theme lost the
cascade. Anchor to the following sheet (banner design / custom CSS) or
append at the end instead, keeping base < theme < design < custom order.
2026-08-18 06:23:49 +00:00
LeoVasanko 1c0ef315ff Refine dateline format; fix stale page caching breaking theme hot swap
Dateline is now '1 Jan 2026', with ' – edited 3 Jan 2026' appended only
when the last edit came at least 24h after publishing.

The Last-Modified header added heuristic browser caching of pages (no
Cache-Control was sent), so the site editor's re-fetch after a theme
change served the cached page with the old stylesheet link. Pages now
send 'Cache-Control: no-cache' — always revalidate, still cheap via the
ETag.
2026-08-18 06:20:38 +00:00
LeoVasanko 87c0aac3c0 Add Server/Last-Modified headers and a {dates} dateline tag
An http middleware sets 'Server: pagerite' (dropping uvicorn's versioned
default) and content pages return Last-Modified from Node.modified. In
markdown, a {dates} line expands to the article's published/updated
dateline (updated shown only when it falls on a later day); the editor
preview resolves it for pages that exist, unsaved pages show it literal.
Demonstrated in the long-read seed page.
2026-08-18 06:15:45 +00:00
LeoVasanko 3cefdec56e Scrap the right gutter on fluid (multicol) articles
Long articles now span from the left gutter (which holds the overlaying
sidebar) to the right viewport edge instead of reserving a symmetric
right gutter. The .wide breakout is re-anchored accordingly: left margin
is the gutter share (20vw of the 1fr+4fr grid, 1/5 of the post-editor
width while editing) plus main's padding; the <=102rem sidebar case keeps
its constant -13.25rem margin.
2026-08-18 06:09:03 +00:00
LeoVasanko e044ab86ee Unwrap seed markdown paragraphs to single lines
With markdown-it breaks:True, the 80ch hard wraps rendered as forced
<br>s: frozen wrapping at the source's line breaks, ragged unjustified
lines and stale manual hyphenation. One line per paragraph lets the
browser reflow, justify and hyphenate normally.
2026-08-18 06:04:39 +00:00
LeoVasanko 6303d5ae6d Fix multicol layout pinning two fixed columns at narrow widths
The 78rem minimum on the center track prevented the article from ever
shrinking, locking long articles into two fixed-width columns that
overflowed the window. Drop the floor: the 4fr center share now scales
with the window in both directions, so column widths flex continuously
and the count falls to 1 when space runs out.
2026-08-18 05:53:56 +00:00
LeoVasanko 7ec0d10155 Proxy /_themes to the backend in the Vite dev server 2026-08-18 05:49:00 +00:00
LeoVasanko 30947df16b Backend-served themes and selectable, inheritable banner designs
Themes move from Vite-built frontend assets to pagerite/themes/{name}/
folders holding theme.css and/or banner.css (+ banner.svg), served by the
backend at /_themes/{name}/... and re-read from disk per request (etag by
mtime), so on-disk edits show on the next page load even in prod and new
themes need no build or config. The theme and banner-design selectors
enumerate these folders via GET /_api/settings.

Banner designs: Node.banner_design picks a design per page (None inherits
from ancestors, then the front page, then the active theme's own design;
"" = none). The design's banner.css is linked in <head> (id
pagerite-banner, between theme and custom CSS) and its banner.svg inlined
into #page-banner first (marked svg[data-design]); the page's own
Node.banner HTML renders after it, so author code always wins. #page-banner
is now a stacking grid so artwork and author code overlay.

Dev/prod hot loading unified: the backend renders the theme/design links
in both modes; in dev pagerite.js only re-appends them (and the custom
CSS) after the Vite-injected base styles. Theme switches just swap the
link href. The pagerite:theme meta and Vite theme build entries are gone.
2026-08-18 05:46:02 +00:00
LeoVasanko c0817330e7 Add favicon upload to the site editor
Data.favicon names a blob in the content-addressed files store (no
migration: msgspec default). PUT/DELETE /_api/settings/favicon upload and
clear it; when set, every page links it as <link rel="icon">, otherwise
browsers fall back to the build's /favicon.ico. SiteEditor shows a preview
with upload/replace/remove and applies the change to the live page head.
2026-08-18 05:22:25 +00:00
LeoVasanko d03ee7fa8f Make article layout fluid: CSS-driven column count, uncapped width
Replace the fixed 100rem two-column breakpoint with 'columns: 30rem' so
CSS fits as many >=30rem columns as the article's width allows, and let
long (.multicol) articles grow past the 78rem cap (4:1 share against the
gutters, keeping them symmetric for the .wide breakout math). Also treat
mid-article h1s as full-width column separators like h2s.
2026-08-18 05:14:21 +00:00
LeoVasanko f9e314e5ea Enable markdown-it typographer and breaks; remove custom dash rule 2026-08-18 04:32:54 +00:00
LeoVasanko 0fcd7b16f5 Remove Vue favicon. 2026-08-18 04:08:47 +00:00
LeoVasanko f125966bb7 Fix various inconsistencies of sidebar and link handling on categories with only one child. 2026-08-18 03:57:41 +00:00
LeoVasanko 32f30dbdaa Don't let whitespace be considered a custom page banner. 2026-08-18 03:14:04 +00:00
LeoVasanko 124ef62b5f Fix flow from new page creation to page editor. 2026-08-18 03:10:26 +00:00
LeoVasanko a0990fc3f6 Fix brand text placing issue with purple theme. 2026-08-18 03:09:35 +00:00
LeoVasanko 87deaa380b Exclude script and style tags from page banner styling to avoid them becoming visible on page. 2026-08-18 02:59:49 +00:00
LeoVasanko f11b772493 Integrate Paskia auth UI flows, hide delete for empty pages. 2026-08-17 22:32:18 +00:00
LeoVasanko a45a575f72 Login button styling 2026-08-17 22:09:48 +00:00
LeoVasanko 9fcf4ee8e0 Stricter slug path processing, only serve site on 404 of missing article paths, not any other path. 2026-08-17 22:02:48 +00:00
LeoVasanko a4fab3dcc9 Ping our own API for auth check, not auth backend (forward auth). 2026-08-17 21:54:26 +00:00
101 changed files with 21663 additions and 3163 deletions
+2 -1
View File
@@ -1,7 +1,8 @@
.* .*
!.gitignore !.gitignore
*.lock *.lock
*.kantadb /localhost
dbip-*.mmdb*
/pagerite/frontend-build /pagerite/frontend-build
package-lock.json package-lock.json
+43 -220
View File
@@ -5,211 +5,53 @@
Please instead ask the user to see from dev tools what you need, e.g. to look up something in DOM or log. Use console.log for debugging where needed (and otherwise for permanently kept useful messages in the app). Please instead ask the user to see from dev tools what you need, e.g. to look up something in DOM or log. Use console.log for debugging where needed (and otherwise for permanently kept useful messages in the app).
## What this is
Pagerite: a single-user CMS/blog. FastAPI serves HTML rendered in Python
with html5tagger; content is persisted in a kanta database and rendered on
the fly per request. Vue is used only for interactive bits (editing tools),
not for the public pages. See `docs/design-principles.md` for the design.
## Layout ## Layout
- `pagerite/` — the Python backend package (hatchling build target). Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes:
- Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed)
- Dev mode `scripts/devserver.py` (which the user mostly uses for auto reloads, no build needed) - `pagerite/` — Python backend package (hatchling build target).
- Avoid running the server yourself, ask the user to test - `app.py` — thin FastAPI assembly: lifespan, `FastAPI(...)`, router includes (route ordering: api/tracking/files routers, then `frontend.route(app, "/")`, then the pages catch-all last).
- `app.py`the FastAPI app. FastAPI's built-in API docs are disabled - `state.py`shared core, no routes: env-derived site constants, `data`/`kanta`, `analytics_store`, the fastapi-vue `frontend`, the render cache (`_html_response`, `_invalidate_pages`), the translator `dispatcher`, slug helpers, `@kanta.bootstrap` hooks.
(`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to - `files.py``FileStore` (content-addressed, RAM-cached), image derivative helpers (`store_image`), file routes (`/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, favicon settings).
our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are - `api.py` — editor REST + WS: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, `/_translate/{key}`.
registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue - `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetch, `/_ws`, `/_api/ws/analytics`, the `/_a` page (docs/analytics.md).
inserts its file routes at the position where - `pages.py` — public content pages: `/`, `/sitemap.xml`, `/robots.txt`, the `/{path:path}` catch-all.
`route()` was called (during `load()` in the lifespan), so anything - `data.py` — msgspec Structs for the kanta database.
defined earlier wins. The one exception is the content catch-all - `chunks.py` — block-level Markdown chunking and content-hash keys for the chunk stores (docs/migrate.md).
`/{path:path}`, registered AFTER `frontend.route()` so that built - `i18n.py` — language selection, translation assembly (chunks + patches) and translated-edit recording (user patches, per-language title overrides, refresh).
frontend assets still take priority over content slugs. The `Frontend` - `translate.py` — translator service protocol (msgspec structs), the connected-client `Dispatcher` (job pipeline, result validation) and pending/store core for the `/_translate/{key}` WebSocket (docs/localization.md); api.py only registers the route.
is constructed with `spa=False` explicitly: it only serves the built - `segments.py` — the translation round trip: fragments split into pure-prose wire segments (via markdown.make_md's verbatim parser; link- and formatting-carrying blocks stay whole, link/formatted texts inline, Markdown stripped) and translations spliced back by source offset, link/formatting markdown re-inserted at weight-mapped positions (docs/localization.md).
files without a catch-all. The build mirrors the URL space — hashed - `migrations.py` — kanta migrations (`migrate_vN`); ALL schema/storage upgrades live here (raw state dict before struct decoding), never in the app lifespan: v1 moves legacy in-db file blobs to the on-disk store and rebuilds the legacy flat `pages` as the menu tree, v2 rewrites `/_f/{hash}.ext` image links to the extension-less form, backfills AVIF/WebP/JPEG derivatives on disk and drops the obsolete `version` field.
immutable assets under `/_assets/`, `favicon.ico` at the site root — - `markdown.py` — markdown-it-py renderer.
and an `index.html` in the build would become a `/` route, so leave it - `views.py` — shared page layout and rendering; theme/user-font resolution across `THEME_DIRS` / `FONT_DIRS` (cwd, site, platform data roots, then built-in `pagerite/themes/`, see `docs/themes-and-assets.md`).
out of the build to keep `/` ours. - `seed.py` — demo content, written only on first database creation.
- `data.py` — msgspec Structs for the kanta database. The site structure - `analytics.py` — visit analytics collection (see `docs/analytics.md`).
is a tree: `Data.menu` maps top-level slugs to `Node`s, each with - `frontend/src/` — Vue editor and public-page JS entries.
`children` keyed by slug — the URL path is the slug chain. The front - `main.js` — Vue editor app entry.
page is whichever top-level node has slug "" (parallel to the other - `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`).
main level pages, not their parent); it cannot have children, and - `langselect-main.js` + `LangSelector.vue` — public language selector, imported on demand by pagerite.js on pages with more than one hreflang alternate (the editors' `LangSelect` flag dropdown).
renaming its slug away leaves no front page ("/" redirects to the - `store.js` — the shared Pinia store (`useStore`, id `pagerite`) for cross-bundle UI state.
first nav item). `Node.content` is - `pagerite.js` — public page entry.
the Markdown page, or None for a pure category label whose URL renders - `editorLang.js` + `LangSelect.vue` — the editor shell's shared language selection and its selector component (page + structure tabs; drives the page preview while the panel is open, via `swapdoc.setLangOverride`).
a placeholder page (while nav links to it point at its first child); - `reconnect.js` — shared WebSocket pacing for all sockets (staggered connect slots, stuck-CONNECTING watchdog, exponential backoff): bursts and rapid retries trip the browser's WebSocket throttling.
every label's title and slug are editable. Siblings order by the fractional `Node.order` key: a moved - `assets/` — base CSS, Pygments styles, fonts.
item gets a fresh key relative to its new siblings, all others keep - `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test).
theirs. `resolve`/`find_slot` walk the tree by path; moves are slot - `scripts/translator.py` — Seed-X translator service client for the `/_translate/{key}` socket (reference client, runs in its own uv env via PEP 723); stays connected full time, unloads the model after 60 s idle and reloads on the next job.
detach/attach carrying the whole subtree. Legacy flat `Data.pages`
(pre-tree databases) migrates into `menu` on startup. The app owns Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed).
the `Data` object; reads are plain attribute access, writes in
`kanta.transaction(...)`.
`Data.files` is a content-addressed store (blake3[:12] + extension)
mapping file names to bytes, served at `/_f/{name}` with immutable
caching; pages reference files by absolute `/_f/` URLs so hierarchy
moves never break them. `Node.banner` is a raw trusted HTML snippet
for the header banner (img, styled div, canvas+script...); empty
inherits from the node's ancestors (front page last), then the active
theme's banner artwork: an inline SVG from
`pagerite/themes/{theme}/banner.svg`, inlined into `#page-banner` by
the backend only when no user banner applies (so it is recolorable
from the theme CSS via `var(...)` and never fights user designs; the
base stylesheet falls back to a plain gradient).
`Data.version` is bumped on every write
and embedded in page ETags so nav-affecting changes invalidate caches.
`Data.brand` is the site name (header link + `<title>` suffix), editable
in the site editor via `/_api/settings`; empty = no header link and
no `<title>` suffix. `Data.theme` is the active theme name (empty =
none/base only); themes live in `frontend/src/assets/themes/{theme}`.
`Data.custom_css` is raw trusted CSS injected inline in every page
`<head>` (id `pagerite-user`) and swapped during fetch-navigation;
editable in the site editor. Font picks (heading/body/brand) in the
site editor are stored as plain `:root` rows in `custom_css`
(`--font-body: var(--font-source-sans);` format — parsed out and
rewritten on change, the `:root` block added/removed as needed),
referencing the per-family variables (`--font-source-sans` etc.) from
pagerite.css;
the base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
- `markdown.py` — markdown-it-py renderer (html passthrough + attrs,
footnote, deflist, tasklists plugins). Custom image rule: relative srcs
resolve against the page path, titled images become figures.
- `views.py` — the shared page layout as an html5tagger `Template` with
placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav
rendering straight from the `Data.menu` tree (siblings sorted by
`Node.order`; nav links to content-less labels point at their first
child via `first_leaf`), and page/404 rendering. If the markdown contains its own h1, the page title
is NOT rendered as an additional h1 (it still supplies <title> and nav
labels). The navbar holds
top-level items only; the current section's subitems go to a left
`#sidebar`, which is rendered only when the section offers at least two
published items (no aside element at all on the front page, leaf pages
and one-page sections). Dynamic regions have stable ids
(`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps
(`#sidebar` may be absent on either side of a swap).
- `seed.py` — demo content written on startup for paths missing from the
database (never overwrites existing pages).
- `frontend/src/` — the Vue editor and public-page entries.
- `main.js` — Vue editor app entry, mounts PageEditor/SiteEditor.
- `pagerite.js` — public page entry; runs fetch-navigation, scroll-reveal,
brand shrink-to-fit (the themed size is the maximum; JS reduces the
font-size so a long brand or narrow viewport still fits one line),
code copy buttons, and the auth check: it fetches
`/auth/api/validate?perm=pagerite:admin` and only then injects the 🖊️
edit pens (asset URLs from the `pagerite:editor-src`/`-css` meta tags);
a 401 adds a "log in" link to `/auth/` in the banner corner, a 403
nothing, and any other result (no auth server, e.g. dev) leaves
editing open. Pages themselves render identically for everyone; the
real gate is the auth proxy in front of all of `/_api`. The backend links the shared CSS as two separate
stylesheets (base and theme) so they can be swapped or augmented.
- `assets/` — shared styles and data files built by Vite and served hashed
under `/_assets/`: `pagerite.css` (base layout + conservative variables),
`themes/{purple,corporate,nitro}/theme.css` (theme overrides and font
picks: `purple` = dark dusk palette with Fraunces/Literata and a tilted
oversized gradient brand; `corporate` = light-first with automatic
`prefers-color-scheme` dark mode, Montserrat/Inter and a huge solid
brand; `nitro` = racing/HUD style following `prefers-color-scheme`
(warm light-grey page, deep violet in dark), Montserrat/Literata,
black as an accent only, a straight orange blade under the banner, and
an orange racing-tab nav clipped with a bezier `shape()`), `pygments.css`,
and `fonts/` (self-hosted Source
Sans 3/Source Serif 4/Fraunces/Literata/Cormorant/Playfair
Display/Inter/Montserrat/Fira Code/Cause/Exo 2/New Rocker
variable woff2). The `::view-transition*` block at the end of `pagerite.css` (from
termotohtori.fi) is fragile — do not tweak.
- Vite builds ES-module `.js` outputs; the backend renders `<script
type="module">` for them (module scripts defer by default).
- The database file is `pagerite.kantadb` in the cwd (`PAGERITE_DB`
overrides); gitignored. Do not delete it without asking.
- `scripts/fastapi-vue/` — helper scripts from the fastapi-vue template
(build hook etc.), do not edit.
- `frontend/` — the Vue editor as **two separate apps** mounted in their
own host divs created inside the static document: `PageEditor.vue`
(CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`,
previewing into the visible article; editor scroll drives document
scroll) opened by the article pen — it edits content and title only,
never the path — and `SiteEditor.vue` (site brand + theme selector +
site-wide custom CSS + banner HTML edited in small CodeMirror windows;
banner previewed into `#page-banner`, CSS injected into
`<head id="pagerite-user">`) + vue-draggable structure tree with
always-editable title/slug inputs per row, opened by the banner pen —
everything saves immediately as you edit (brand/title/CSS debounced,
slug on commit since it renames the path), theme change swaps the
stylesheet in place, tree rows navigate in place without transitions when
focused, and the front page is a root-only row whose empty slug is
editable like any other. Every
non-empty list (and the root) ends with a non-draggable footer row
(vuedraggable `#footer` slot): clicking it starts a new pending page at
that level (its slug placeholder shows the slug derived live from the
title being typed), and while dragging it is the list's "end of list" drop
target. Committing a pending page PUTs it with empty markdown (creates
an empty page that renders with its title — saving never deletes;
deletion is the page editor's explicit choice: saving trimmed-empty
text issues a REST DELETE), then hands over to the page editor
(CodeMirror focuses on mount). Dropping ON the lower part of a row
moves the page under that
row (the child list's container invisibly overlaps its own row's bottom
via negative margin — Sortable inserts it as the first child natively),
while a row's exposed top edge inserts a sibling before it. Row
indentation is structural (each nested list margin-indents itself), so a
dragged row previews its whole subtree at the target list's depth. The
two pens swap the docked
panel for the other editor; clicking the open editor's own pen closes it. Normally dynamic-imported onto the content page by
pagerite.js when a 🖊️ edit pen is clicked (the pens are injected by
pagerite.js after the session validates; they carry
`data-editor-src`/`data-editor-css`/`data-editor-mode`).
In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`),
in prod from the hashed build assets resolved via
`frontend-build/.vite/manifest.json`. `vite.config.js` sets
`appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`,
`assetsDir: '_/assets'` (so the build mirrors the URL space;
`frontend/public/favicon.ico` lands at the build root and is served at
`/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus
`src/assets/pagerite.css` and every `src/assets/themes/*/theme.css` as
separate stylesheet entries (enumerated from the themes directory, so new
themes need no config change); there
is no `index.html` source (it would shadow `/` and turn missing dev paths
into an empty Vue shell). All outputs are ES modules. The build sets
`preserveEntrySignatures: 'exports-only'` because main.js is consumed
via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite
app builds otherwise strip unused entry exports, leaving dead edit pens.
In dev the backend links no stylesheets (Vite injects them from JS); the
active theme reaches the page as `<meta name="pagerite:theme">` and
pagerite.js imports that theme's CSS, while a theme switch in the site
editor swaps the Vite-injected `<style data-vite-dev-id>` tags (the
`<link>` sync used in prod is a no-op in dev).
vite-plugin-fastapi.js has an
auto-upgrade marker — edit `vite.config.js`, not the plugin.
- `docs/` — design documentation.
## Toolchain ## Toolchain
- Python >= 3.14, managed with **uv**. Dependencies: `fastapi[standard]`, - Python >= 3.14, managed with **uv**. Dependencies: `fastapi[standard]`, `fastapi-vue`, `html5tagger`, `kanta`, `markdown-it-py`, `mdit-py-plugins`, `platformdirs`, `pygments`, `tracerite`; dev group has `httpx`. Run anything via `uv run ...` (the venv is `.venv`).
`fastapi-vue`, `html5tagger`, `kanta`, `markdown-it-py`, `mdit-py-plugins`,
`pygments`, `tracerite`; dev group has `httpx`. Run anything via
`uv run ...` (the venv is `.venv`).
- Key libraries: - Key libraries:
- **html5tagger** — all HTML generation (`E`, `Document`, `Template`, - **html5tagger** — all HTML generation (`E`, `Document`, `Template`, `HTML` for trusted/raw HTML).
`HTML` for trusted/raw HTML).
- To create stand alone pages, begin with `doc = Document(...)` that gives a HTML5 page header - To create stand alone pages, begin with `doc = Document(...)` that gives a HTML5 page header
- Chain with `doc.p("text").br`: every attribute access creates element to doc (returning self), calls add content to current element. - Chain with `doc.p("text").br`: every attribute access creates element to doc (returning self), calls add content to current element.
- Closing tags are not used where optional, e.g. no `</p>` or `</li>` is ever included in output. Due to this proper "nesting" of content is NOT required and should be avoided. Where needed, () directly after tag define attributes and content INSIDE the element, then close the element. `with doc.ul:` and such may be used for larger chunks. - Closing tags are not used where optional, e.g. no `</p>` or `</li>` is ever included in output. Due to this proper "nesting" of content is NOT required and should be avoided. Where needed, () directly after tag define attributes and content INSIDE the element, then close the element. `with doc.ul:` and such may be used for larger chunks.
- Prefer building directly on one builder with `with` blocks (recursing - Prefer building directly on one builder with `with` blocks (recursing inside a with block for hierarchies) over preparing `E.` snippets into variables and composing them. Note `with doc.li:` alone fails (`li` has an optional end tag) — use `with doc.li.ul:` style chains, or `doc.li.a(...)` followed by a nested `with doc.ul:` block.
inside a with block for hierarchies) over preparing `E.` snippets into - `Template(builder)` freezes a builder with **Capitalized** attribute placeholders (e.g. `E.Title`, `doc.main(E.Main, id="main")`); calling it fills the slots with escaping — pass `HTML(...)` for raw HTML. Passing a list to a template slot expands it; passing a list to a normal builder call does NOT (spread it: `E.ul(*items)`).
variables and composing them. Note `with doc.li:` alone fails (`li`
has an optional end tag) — use `with doc.li.ul:` style chains, or
`doc.li.a(...)` followed by a nested `with doc.ul:` block.
- `Template(builder)` freezes a builder with **Capitalized** attribute
placeholders (e.g. `E.Title`, `doc.main(E.Main, id="main")`); calling
it fills the slots with escaping — pass `HTML(...)` for raw HTML.
Passing a list to a template slot expands it; passing a list to a
normal builder call does NOT (spread it: `E.ul(*items)`).
- To create plain HTML snippets use `E.div(E.p("content"))` etc using the `E` empty builder. - To create plain HTML snippets use `E.div(E.p("content"))` etc using the `E` empty builder.
- **kanta** — asyncio-native embedded database: `Kanta(filename, data)` - **kanta** — asyncio-native embedded database: `Kanta(filename, data)` root object, `transaction`, `flush`, snapshot/replay-log persistence.
root object, `transaction`, `flush`, snapshot/replay-log persistence.
- `async with Kanta(Data(),...) as kanta:` (or await kanta.open/close) - `async with Kanta(Data(),...) as kanta:` (or await kanta.open/close)
- `with kanta.transaction(...) as data:` - transactions only for writes - `with kanta.transaction(...) as data:` - transactions only for writes
- `data` may be referenced directly to read anywhere and to modify in transactions (`as data` is just a shorthand access) - `data` may be referenced directly to read anywhere and to modify in transactions (`as data` is just a shorthand access)
@@ -217,32 +59,13 @@ not for the public pages. See `docs/design-principles.md` for the design.
- We prefer objects rather than lists, as this works better in change diffs. E.g. `dict[str, True]` where the keys indicate presence and always have value `True`. - We prefer objects rather than lists, as this works better in change diffs. E.g. `dict[str, True]` where the keys indicate presence and always have value `True`.
- Maintaining and owning the app's own `Data` object is preferable; Kanta never copies this, only edits in place - Maintaining and owning the app's own `Data` object is preferable; Kanta never copies this, only edits in place
- Note: besides opening it every access is immediate direct variable access: no `await`, no locks, no delays - Note: besides opening it every access is immediate direct variable access: no `await`, no locks, no delays
- **fastapi-vue** — template glue for serving/building the Vue frontend; - **fastapi-vue** — template glue for serving/building the Vue frontend; keep its integration points (`Frontend`, build hook) intact.
keep its integration points (`Frontend`, build hook) intact. - **platformdirs** — platform user/system data dirs for the theme and font search roots (`views.THEME_DIRS` / `views.FONT_DIRS`; use `site_data_dir(..., multipath=True)`, not `site_data_path`, which collapses multipath).
- **markdown-it-py** — Markdown rendering with `html=True` raw - **markdown-it-py** — Markdown rendering with `html=True` raw passthrough; mdit-py-plugins for footnote/deflist/tasklists/attrs; in-body h1/h2 headings get auto slug ids + self-links when the body has 3+ of them (`python-slugify`, mirroring `slugify.js`); **Pygments** for server-side code highlighting (`nowrap` spans, styled by `frontend/src/assets/pygments.css` which maps token classes 1:1 onto the `--code-*` variables; light/dark palette sets live in `pagerite.css` and resolve via `light-dark()` from the theme's `color-scheme` — themes pick a set, not individual colors).
passthrough; mdit-py-plugins for footnote/deflist/tasklists/attrs;
**Pygments** for server-side code highlighting (`nowrap` spans, styled
by `frontend/src/assets/pygments.css` which maps token classes 1:1 onto
the `--code-*` variables; light/dark palette sets live in
`pagerite.css` and resolve via `light-dark()` from the theme's
`color-scheme` — themes pick a set, not individual colors).
## Conventions ## Conventions
- Keep dependencies minimal; add via `uv add` and mention it. - Keep dependencies minimal; add via `uv add` and mention it.
- The public URL space belongs to content (pretty slugs at root). Reserve - The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` (backend redirect to the configured site icon). Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus - No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session. The one keyed exception is `/_translate/{key}` (translator service; `Data.translate_keys`, see docs/localization.md).
`/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, - Update the relevant MarkDown files when architecture, tooling, or conventions change.
hyphens and underscores `[a-z0-9_-]` (the site editor filters input live
via `slugify.js`, built on the `transliteration` npm package — unicode
folds to ASCII, spaces become hyphens; an empty slug on a new page is
derived from its title), may not begin with `_` or `.`, and such URLs are
never looked up as content.
- No auth in core code; the SSO/reverse proxy gates all of `/_api`
(forward-auth) and owns `/auth/` (login/logout, session validation).
Pages render identically for everyone; pagerite.js adds the editing UI
only after the auth server validates the session. Never add output
sanitization "for safety" against the author — embedded HTML/scripts in
Markdown are passed through deliberately.
- Update this file and `docs/design-principles.md` when architecture,
tooling, or conventions change.
+32 -3
View File
@@ -1,5 +1,34 @@
![The same site in five themes](https://git.zi.fi/LeoVasanko/pagerite/raw/branch/main/docs/screenshots/themes.webp)
# Pagerite # Pagerite
A single-user CMS/blog. FastAPI serves HTML rendered in Python with html5tagger, A CMS for people who are done patching WordPress. There's no PHP or Node.js to exploit — the whole editing surface sits behind your own SSO proxy, so the server the internet can talk to just renders plain pages that search engines and social media can read too.
content is persisted in a kanta database and rendered on the fly per request.
Vue is used only for the interactive editing tools, not for the public pages. The articles have rich layout and don't look boxed in like with most web publishing platforms. The software is lightweight and fast enough to serve any number of visitors you have. We run our own site [vasanko.com](https://vasanko.com/) on it, in case you wish to have a quick look.
## Run it
```sh
uvx pagerite localhost
```
That serves a demo site on localhost using [uv](https://docs.astral.sh/uv/getting-started/installation/). When you take it to production, pass your domain name instead. Our [setup guide](https://git.zi.fi/LeoVasanko/pagerite/src/branch/main/docs/setup.md) walks through the whole production arrangement. **Read it before running this online.**
## What it's like
**You write, Pagerite renders.** Articles are Markdown with the extensions that matter — tables, footnotes, task lists, callouts, highlighted code, aside boxes — and raw HTML goes through untouched when Markdown runs out. Long articles reflow into a proper two-column composition on wide screens without you doing anything.
**Editing happens on the page.** Click the pen next to a heading and an editor docks beside the live article, previewing server-side as you type. Site name, theme, fonts, banner, custom CSS — changed in a panel, applied immediately. New pages grow from a in the structure tree; drag or rename rows to reorder your whole navigational hierarchy. Entirely custom or premade top banner designs per category or page are available, animations included.
Full scripting and styling is available for editors who wish to implement more complex functionality on their articles. This also means you should only let trusted users write on your site: this is by no means a public blog platform.
The worst case scenario when a hacker gains access to your admin accounts (say if you didn't read the setup guide): they can take over the entire site and run scripts on users' browsers, but the damage is limited to same domain. All your articles can be restored to the state prior to that hack or that one user's edits undone, and no data is irrecoverably lost. This is much better than other platforms that also let hackers run code on your server (WordPress).
![Live editing](https://git.zi.fi/LeoVasanko/pagerite/raw/branch/main/docs/screenshots/editor.webp)
**Theme just every part to your liking.** Themes, banner designs and page transitions are included — pick one from the site editor or copy a folder and make it yours. Several high quality fonts are included among with other assets: your site never phones a third party or us for anything. And if after all you need to customize, additional site and banner code may be provided by the admin panel.
**Search engines and social cards come free.** Every page gets a proper description, canonical link and Open Graph/Twitter card metadata derived from the article — including a share image picked from your own figures — without a single "SEO plugin". Category index pages, if you wish to have those, also get their sub pages shown automatically in card format.
![Graphs showing visitor stats and navigation across site branches.](https://git.zi.fi/LeoVasanko/pagerite/raw/branch/main/docs/screenshots/analytics.webp)
_You can see your readers. Built-in analytics need no cookies and no third-party tracker: visits, referers, reading time and a live map of how people move between your pages, plus separate ledgers for crawlers and the abusers probing for wordpress PHP files — who are, of course, wasting their time here._
+385
View File
@@ -0,0 +1,385 @@
# Analytics
Server-side visit analytics built on a **raw access-log-style event store**.
Data lives in a plain JSON file — a msgspec Struct dumped to disk — separate
from the kanta content database, path from `PAGERITE_ANALYTICS` (default:
`analytics.json` in the per-site data directory, e.g. `localhost/analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Get`, `Msg`, `Client`,
`Favicon`), the `Store` (raw log + atomic JSON persistence) and
`Store.display()`, where **all** classification happens.
- `pagerite/pages.py` — records every served document as one raw GET line
(`_record_get`, in `pagerite/tracking.py`) with its true HTTP status.
- `pagerite/tracking.py` — the `/_ws` activity WebSocket, and
`WebSocket /_api/ws/analytics` (admin-gated like every `/_api` endpoint).
- `frontend/src/pagerite.js` — the client activity channel and the 📊 pen.
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
normal site layout on the `/_a` analytics page.
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
into `#analytics-app` inside `#main`.
## Raw records
The store is deliberately close to an access log: two append-only lists plus
shared metadata. **Nothing is classified when recorded** — whether a client
turns out to be a reader, a crawler or a scanner is decided by
`Store.display()` from the raw events, so the stored data survives any future
change to the classification rules.
Each `Get` record (one per served document):
- `t` — timestamp of the request,
- `path` — full request path, query string included (e.g. `/.env?x=1`),
- `status` — the true HTTP status of the response (200, or 404 for a category
placeholder or a missing page),
- `ref` — external https origin of the `Referer`, `""` for direct/internal
(same-origin referers are dropped by the recorder),
- `pre` — true for idle-time link preloads from pagerite.js
(`x-pagerite-preload` header): never counted as a view, crawler hit or
abuse — recorded only so a navigation later served from the in-memory page
cache (which issues no GET at all) can be attributed this GET's status,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`.
304 revalidation responses return before recording and are not logged.
Each `Msg` record (one per pagerite.js activity message over `/_ws`):
- `t` — timestamp,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `fr` — path of the page the activity happened on (`""` for the initial
load),
- `to` — navigation target (validated at record time: internal slug path or
external https URL; anything else is dropped — sanitation, not
classification),
- `read` — active seconds spent on `fr` since the previous report.
Each `Client` record (shared by every event, keyed by hash):
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string,
- `hide` — true for admin clients (`hide` message field): everything this
client ever did is recorded but excluded from every statistic and from the
viewer payload. This is the one flag set at record time — it is a client
property, not a classification.
A reverse-DNS lookup is attempted for each new client and the result, when
available, is stored as `host`; local/reserved/multicast addresses are
skipped. If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present
in the working directory, it is loaded at startup and used to look up
`country`/`city`. These lookups run in background tasks after the event is
stored, so WebSocket message handling is never delayed. The decompressed
`dbip-*.mmdb` file is kept in the working directory and ignored by git. The
CLI flag `--dbip` (`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP at startup (in the app lifespan,
before the MMDB is opened), skipping the download when the local database is
already current and removing older versions after an update; without the flag
only an existing file is used.
## What the client sends
The client (`pagerite.js`) keeps a WebSocket connection to `/_ws` for the
whole browsing session and sends activity messages over it — JSON text
frames matching the server's `Ping` msgspec struct with the fields `fr`
(source path), `to` (navigation target), `read` (active seconds on `fr`
since the last report) and `hide`; falsy fields are omitted. One channel
follows the session, so the activity of a visit stays tied together, and
while the user is active the accumulated reading time is flushed every few
seconds: the times are incremental, so a disconnection simply leaves the
last reported time in place (no close beacon). After 5 minutes without
any activity the client closes the socket itself — a sleeping browser tab
would lose it anyway — and the next activity reconnects; reconnects are
attempted only on user activity, with an exponential backoff between
attempts so a failing endpoint is never hammered. Idle-time link preloads
stay plain `fetch()` calls so the browser may cache the responses; the
WebSocket reports actual navigations and active time spent on a page.
- **Initial page load**: only `to` — the loaded path — is sent, never `fr`
(an `fr` equal to `to` would log a bogus self-transition when a session
already exists, e.g. a second tab). Reloads are not
visits: the message is skipped (PerformanceNavigationTiming `reload`), so a
refresh neither counts a second view nor logs a self-transition.
- **Internal fetch-navigations**: `to` is the target path, sent only after
the swap actually happened (a failed swap falls back to a full load,
whose initial message counts the view instead — no gap, no double count).
- **External links** (`https` only): `to` is the link's full URL. This is the
exit-link record; the user may continue navigating afterwards (new tab,
back), so the exit URL is not necessarily the last trail entry. Outbound
links are stored by full URL so several links to the same domain remain
distinct.
- **Excluded**: back/forward (popstate) navigations, navigating *to* the
analytics page (`/_a` — its GET is untracked, and the server cannot
record it as a navigation target anyway), and everything while the user has
the editor open (`body.editing`). Admin noise, not visits. Navigating
*away* from `/_a` does report.
- **Admins**: when SSO is in use and the session is known to be an admin,
the client still reports but adds `hide`. The activity is recorded as
usual (navigations and all), but the `hide` flag is set on the **client
record** — so it covers everything that client ever did, including the
time before the login. Hidden clients never appear in the viewer payload:
`Store.display()` drops their events and metadata, and computes every
aggregate (site visits, page views, transitions) from the visible visits
only, so nothing needs to be reversed or redacted. With no auth proxy
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
is recorded.
- **External-site favicons**: for every external https origin seen as a GET
referer or an exit link, the server fetches `{origin}/favicon.ico` in a
background task (httpx, 8 s timeout, ≤ 64 KB, image content-types only —
SVG is sniffed from the body when served without an image type) and stores
the icon content-hashed on disk in the FileStore (served at `/_f/{name}`,
extension matching the actual MIME). The origin → file name mapping is
recorded in `Analytics.favicons` (`Favicon.file`/`fetched`); misses are
recorded too and retried only after 7 days. Fetches are scheduled after
each activity message and once at startup, which backfills icons for
already-recorded data. The viewer payload carries `favicons` (origin →
`/_f/...` path), and the viewer shows the icon wherever an external site
is mentioned: referer/exit trail links in the visit table and the
source/exit pills of the transition map (UTM-attributed source nodes
without an https origin stay text-only).
## Display-time classification
`Store.display(in_menu)` derives the viewer payload from the raw events on
every (debounced) broadcast — O(n log n) over the log, cheap enough for a
small CMS. `in_menu(path)` resolves a path against the current menu (passed
in from `tracking.py`, which owns the content database import) so 404
responses for real menu nodes — category placeholders — are not mistaken
for misses.
- **Visits and sessions**: a client's messages are grouped into visits
chronologically; a new visit starts after 30 minutes of inactivity
(`_SESSION_GAP`). A fresh page load with an already-open visit (second
tab) extends it, logging a `(direct)` transition. The visit's trail holds
first-seen targets in order; `read` updates accumulate active seconds on
the trail item matching `fr`. Each trail item's HTTP status comes from
the client's latest GET for that path — preloads included, which is what
allows 404 pages to render red in the viewer even when the navigation
itself was served from the page cache. The entry page's referer and
`utm_*` tags come from the GET that loaded it (within 10 s before the
first message).
- **Crawler hits**: a document GET no activity message matched within
`_CRAWLER_TIMEOUT` (10 s) is a crawler hit — plain bots that only fetch
documents never register as visits. JS-running crawlers (Googlebot,
GoogleOther, Applebot, ...) do connect and send messages, but their UA
gives them away (`_is_bot_ua`): their messages are ignored at display
time, so their GETs never match and land in the crawler list too. Real-
browser bots whose UA does not match are caught by engagement: a visit
whose total reported reading time is under 5 seconds (`_MIN_VISIT_READ`;
durations are client-provided and trusted — such bots report 02 s) is
reclassified as crawler hits, one per internal trail page, and counts in
no visit aggregate. No source-IP verification is done: a spoofed bot UA
merely lands in the crawler stats, and scanners that probe telltale paths
are caught by the abuse rules regardless. In the viewer, crawler hits are
grouped by client hash and shown as a trail of pages, preceded by the
referer when there is one (rendered with its favicon like visit
referers). The crawler table lists the most recent crawler first, with
the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 on a telltale path — an empty URL segment
(`//foo` — no real client generates those), any segment starting with a
dot (`/.env`, `/.git/config`) or ending in `.php` — classifies the source
IP as abuse, and ten plain 404s within one hour (`_ABUSE_404_WINDOW`) on
paths that don't resolve to a menu node do too. Two exemptions keep
legitimate traffic out: RFC 8615 well-known URIs (`/.well-known/…`
browsers and services probe them, e.g. Chrome's devtools fetch of
`appspecific/com.chrome.devtools.json`) are never telltale and never
count toward the threshold, and category placeholders return 404 but are
real menu nodes, so they never count either. The window keeps a
long-time reader's slowly accumulating misses from ever crossing the
threshold — scanners spray in bursts. Hidden (admin) clients never
trigger classification: editing means visiting not-found pages, since
that is where the create pen lives. Once an IP is classified, **all** its document GETs are shown in the abuse list —
including any that arrived before classification, since the raw log keeps
everything — and its activity messages are ignored. In the viewer, abuse
hits are grouped by IP (never by client/UA — scanners randomize theirs)
in a separate "Abuse" table, split by the recorded status: the 404 probes
("paths abused" — flagged paths that triggered classification first, then
other 404s, shown verbatim with query strings) versus the real articles
the abuser actually read ("articles read" — the 200 document GETs,
rendered as trail links like the visitor and crawler tables, query string
stripped). Raw User-Agent strings are shown one per line with their
occurrence counts, and the full lists are click-to-copy.
In the visitor and crawler tables, internal paths that returned a 404 status
are shown in red and the link title includes the status code, so it is easy
to tell misses from real pages at a glance.
## Derived shapes (the viewer payload)
The `Display` payload contains the derived `visits`, `crawlers` and `abuse`
rows (structs `Visit`/`Nav`/`TrailItem`, `CrawlerHit`, `AbuseHit` — display
DTOs only, never persisted), the visible `clients`, the fetched `favicons`,
and the aggregates below.
Each derived `Visit`:
- `start` — timestamp of the first activity,
- `entry` — first page (path) seen,
- `referer` — external https origin of the entry GET, `""` for direct,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `trail` — the entry page and everything seen afterwards, keyed by the
timestamp of first sight (insertion order = first-seen order). Each item
holds `to` (page path or external exit URL), the accumulated active
reading time in seconds (`read`) and the most recent HTTP status seen
for the target (`status`),
- `navs` — every navigation (`fr`, `to`), keyed by its timestamp, repeats
included. The aggregates are computed from this log,
- `utm``utm_*` query parameters from the landing URL, as a dict.
Each derived `CrawlerHit`:
- `start` — timestamp of the document GET,
- `entry` — page path requested,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `referer` — external https origin of the request, `""` for direct/none,
- `query` — raw query string of the request,
- `status` — HTTP status of the served response (200 for a real page, 404
for a category placeholder or missing page).
Each derived `AbuseHit`:
- `start` — timestamp of the request,
- `path` — full request path including the query string,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the paths that triggered abuse classification (telltale
paths, or the 404 that crossed the threshold),
- `is_404` — true for 404 responses, false for real (200) document GETs.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical requests (same path and status class) are collapsed
with their counts — a path's 404 probes and its later 200 reads never
merge. Within each list paths are sorted by count descending, then by their
earliest hit.
## Aggregates
Aggregates are **not stored**; they are computed at display time by
`Store.display()` from the derived visits (entry + `navs` log), skipping
hidden clients and short visits reclassified as crawler hits. This is
what allows a client to become hidden after navigations were already
logged: no counts need reversing. The computed shapes, part of the
WebSocket payload (`Display` struct alongside `visits`, `crawlers`, `abuse`
and `clients`):
- `transitions`: time series of page transitions, sparse nested dict
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
referer origin or `"(direct)"` for initial loads, a page path for
navigations.
- `views`: time series of page loads, `path -> bucket -> count`, sparse: only
non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp).
Every load counts, including repeats within a visit; external exit origins
are not page views and are not counted here.
- `site_visits`: `bucket -> count` of new visits started, same sparse
5-minute bucketing.
Sparseness keeps quiet sites small; dropping old data is a matter of deleting
list entries (`gets`/`msgs` are plain append-only lists).
## Persistence
The whole `Analytics` struct is JSON-encoded and written atomically
(temp file + rename) on every recorded event. Traffic on a small CMS makes
this cheap enough; batching can be added later without changing the format.
A file written by the pre-redesign schema (stored `visits`/`crawlers`/`abuse`
lists) is not convertible; it is renamed to `analytics.json.bak-legacy` and
recording starts fresh.
## Viewing
The 📊 pen in the banner corner (admins only, injected by pagerite.js next to
the edit pens) links to `/_a`, the analytics page. It is a normal site page:
the standard banner, navigation and footer stay in place, and the analytics
content is rendered inside `#main`. The page itself is public, but the data
stream comes from `WebSocket /_api/ws/analytics`, which remains admin-gated
like the rest of the management API; visitors without access see the viewer
with a "could not be loaded" message.
Because it is a real page, fetch-navigation handles it like any other internal
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
range selector updates the URL hash (`#week` etc.) so links to a specific
range can be shared. When the URL has no hash, the client derives the
default from the first analytics snapshot: `day` if the recorded history
spans less than 24 hours, otherwise `week`.
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
Charts are SVG curves (Catmull-Rom over an edge-aware Gaussian — a
change-point detector splits the series at traffic-level shifts, then each
segment is smoothed independently with a fixed sigma chosen so N events in
a single bucket peak at N events per unit. The raw series is drawn faint
underneath). Values are
**per-unit rates** — per hour on the week view (5-minute bucket counts × 12,
plotted at native 5-minute resolution), per day on the month+ ranges — and
the smoothing time scale follows the unit: the month+ sigmas are 24× the
hourly ones. The y max is derived from the smoothed curves so single-bucket
spikes don't blow up the scale, and raw spikes are clamped into the plot.
Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5
labeled intervals, minor lines at fifths when integral; the minimum y-axis
range is 10 so tiny values such as a single visit are not stretched to a
fractional scale).
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
weeks in the muted color at decreasing opacity (the current week keeps the
accent color and is
truncated at the current bucket, never drawing fake zeroes for the future);
a compact legend inside the top right of the visits chart marks the current
ISO week in accent and the overlaid past weeks as "Week M" or "Week MN" on
a muted specimen. Its x labels are weekday names centered at midday UTC, without
vertical grid
lines (day boundaries would be misleading in the viewer's timezone). The
month view labels days the same lineless way — day numbers at noon UTC,
with the month name substituted for the 1st. Month, year and all are
rolling windows ending at now, aligned to UTC day boundaries at the start
so the labels span the whole range; the bucket size follows the window —
6 hours up to 31 days, daily beyond — with boundary lines at months/years
on the longer ranges. All uses the full data reach, but keeps
at least the past 30 days (identical to the month view when the site is
younger than that, bucket size included) so the chart never collapses to a
tiny sliver when the site is young. Below the charts: a **transition map** (all pages from
`/_api/pages` — top-level menu items on a large-radius circular arc whose
bottom point is the last item (each earlier item a bit higher), connected
by a top lane labeled 🏠︎ beside the home pill (50% thicker than
the branch lanes, its label font and guide offset scaled along), each item's
subtree fanning out below it in menu order along a large-radius circular
arc that leaves heading
straight down and gradually bends right, index pages without views omitted
and their children promoted in their place. The submenu structure is drawn
as wide branch lanes: one per path prefix with at least two visible
nodes, running behind the branch's node pills as circle arcs concentric
with the fan (parent levels one radius step outward, so all lanes of a
group share exactly one form), each labeled with its branch slug
left-aligned just past the first pill and allowed to run along the lane to
its end, disappearing under later pills when long — so the lanes reflect
the path
structure even where index pages are omitted — opposite transition
directions joined into organic
tapered connections whose middle width grows logarithmically (base 2)
with the daily hit rate (uncapped), connections
carrying less than 1% of the total traffic
pruned, as are those whose thin middle would render below ~0.8 px —
fainter strands are invisible and only their wide end flares would show; beads are simulated one by one in JS (requestAnimationFrame) and
flow along each edge, persisting across data reloads (emitters are keyed
per edge direction and beads tracked by progress, so an unrelated count
change never reshuffles them), emitted at a rate linearly proportional
to the directional count with no in-flight limit, opposing directions
offset onto parallel lanes. External sources and exits whose connectors are
all culled by the width threshold are dropped from their rows themselves
(the site's own page nodes always stay, connected or not). External sources show as a node row above the
map: each visit is attributed to `utm_campaign`, then `utm_source`, then the
referer origin, then any other `utm_*` tag, so UTM-tagged visits are grouped
under their campaign/source value rather than the referer domain. A UTM
source node only links to its referer when every visit carrying that tag
came from the same origin. External exits are full-size nodes in a matching
row centered below the map, so the site itself stays in the middle), per-page view
counts, the top transitions and the 50 most recent visit trails. Data is
streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest
JSON snapshot on connect and again whenever the analytics file is updated
(with a small server-side debounce to avoid flooding under high traffic).
+45
View File
@@ -0,0 +1,45 @@
# Backend
The Python backend lives in `pagerite/`.
## `app.py`
Thin FastAPI assembly: lifespan (open the kanta database, load the file store, the frontend build and GeoIP), the `FastAPI(...)` instance with built-in API docs disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content, the `server` header middleware, and router includes. The routes themselves live in specialized modules:
- `state.py` — shared core, no routes: the environment-derived site constants (`HOSTNAME`, `SITE_URL`, `DB_PATH`, `FILES_DIR`, image/favicon tunables), the `data` root and its `kanta` handle (`Kanta(..., migrations="pagerite.migrations")`), the `analytics_store`, the fastapi-vue `frontend`, the page render cache and `_html_response`, the translator `dispatcher`, the slug charset helpers, and the `@kanta.bootstrap` hooks (demo seed, translator defaults).
- `files.py` — the `FileStore` and image derivative helpers, and the file routes: `/_api/files`, `/_f/`, `/_themes/`, `/_fonts/`, the favicon settings endpoints.
- `api.py` — the editor REST API and WebSockets: `/_api/pages`, `/_api/structure`, `/_api/settings`, `/_api/toggle-task`, `/_api/translations`, `/_api/ws/editor`, and the translator channel `/_translate/{clientkey}`.
- `tracking.py` — visit analytics: GeoIP, client enrichment, favicon fetching, debounced broadcasts, the `/_ws` activity socket, the admin stream `/_api/ws/analytics`, and the `/_a` viewer page.
- `pages.py` — the public content pages: `/`, `/sitemap.xml`, `/robots.txt` and the `/{path:path}` catch-all.
Route ordering is load-bearing and lives in `app.py`: the api/tracking/files routers are included BEFORE `frontend.route(app, "/")` is called — fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything registered earlier wins. The content catch-all `/{path:path}` is included AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `state.py`'s `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding — and cleared wholesale by `_invalidate_pages()` on every content/settings change, which also bumps the in-memory render generation. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and the render generation; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
Uploaded files, seed assets and fetched external-site favicons live in the `FileStore` (in `files.py`): content-addressed files on disk under `<hostname>/files/` (`PAGERITE_FILES`), fully cached in RAM at startup — both the raw body and a zstd-compressed copy (kept only when smaller). `GET /_f/{name}` serves from the RAM cache with immutable caching, answering the zstd variant when the client accepts it; the name is the ETag. Uploaded raster images (and rasterized SVGs) are stored as `<hash>.orig<ext>` (internal only, never served) plus AVIF, WebP and JPEG derivatives, and pages link the extension-less `/_f/{hash}`: the server serves a format only when the Accept header lists it explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise — including `*/*` — JPEG), with `vary: accept`; an explicit extension pins the format. `migrate_v2` rewrites old `/_f/{hash}.avif` article links to the bare form, backfills missing derivatives on disk, and drops the obsolete `version` field. Legacy databases that still carry blobs in a `files` kanta field or a flat `pages` store are migrated by `pagerite/migrations.py::migrate_v1` (kanta's `migrate_vN` mechanism, wired via `Kanta(..., migrations="pagerite.migrations")`), which rewrites the raw state before struct decoding — all schema/storage upgrades live in that module, none in the app lifespan.
## `data.py`
msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model.
## `markdown.py`
markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists, admon, gfm_autolink, sub/superscript plugins; typographer + breaks on). In bodies with at least three top-level h1/h2 headings (nested ones, e.g. inside `::: aside`, never participate), each gets a slug id (`python-slugify`, mirroring the editor's `slugify.js` — unicode folds to ASCII, separators become single hyphens) unless the author set `{#id}`, and their text is wrapped in a self-link (`a.anchor`) so section links are copyable; anchored headings also carry `data-line` with their markdown source line (the page editor's section pens and piecewise scroll sync key off it); the first in-body h1 is the article title — when the markdown has no h1, `render(title=...)` injects it as `# {title}` so implicit and explicit titles take the same path — it gets no id and doesn't count toward the three, its self-link is `href=""` (scroll to top); shorter articles stay anchor-free, h3+ is never navigable, and duplicates get `-2`/`-3` suffixes. Custom image rule: relative srcs resolve against the page path; an image standing alone in its paragraph becomes a figure (captioned when titled), while inline-with-text images and raw `<img>` HTML stay plain. A `{dates}` line expands to the article's published/updated dateline (`p.dateline`, from `Node.created`/`modified`; left literal in previews of unsaved pages). Code fences take pandoc-style brace attributes on the info line (` ```{.python .wide #id key=val} ` — the first class is the language when no bare language word precedes the braces) as well as a trailing `{...}` line; both land on the `<pre>`, the `<code>` keeps only the language class.
`render()` returns a `Rendered(html, multicol)`: the article content segmented for the column layout (there is no wrapper div — segments and bare blocks are direct `<article>` children) — h1/h2 headings and `.wide` blocks stand bare, the runs between them become `<div class="colseg">` (margin-breakout boxes — `.margin`, `::: aside` — stay inside the segment at their anchor point; the CSS positions them out of flow into the side zone) (plus `.cols` on segments with enough text in at least two paragraphs or one long enough to split across columns, `::: nocols` opting out; in column segments, paragraphs past `BREAKABLE_TEXT` visible characters are marked `.breakable` so they may split across columns), and `multicol` flags bodies long enough to columnize (visible-text thresholds, code excluded). `views.py` puts the class on the article; pagerite.css takes it from there (at most two columns, the left-margin breakout, all viewport adaptation).
## `views.py`
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`SITE_URL``https://<hostname>` from the CLI hostname argument; on localhost the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. Additionally `twitter:image` pins extension-less `/_f/{hash}` share images to the `.webp` variant — X only honors WebP via twitter:image (not og:image) and its scraper cannot be trusted to negotiate via Accept. The page title is injected as `# {title}` when the markdown has no h1 of its own, so it never appears twice (it always supplies `<title>` and nav labels).
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), rendered only from the second level down — main-level pages list their children as cards after the content instead. Below that, the sidebar renders when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, main-level pages, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
Any page with published children — a category page — lists them as a card grid (`nav.cards`) after the markdown content, as does the content-less category 404. Each card links to the child page (a content-less child to its first leaf) and shows the child's share image (the same hero → first raster → first SVG heuristics as `og:image`) as a full-card cover with the title overlaid.
## `seed.py`
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `state.py`.
+37
View File
@@ -0,0 +1,37 @@
# Content model
The site structure is stored in the kanta database managed by `pagerite/data.py`.
## Site tree
`Data.menu` maps top-level slugs to `Node`s, each with `children` keyed by slug — the URL path is the slug chain. The front page is whichever top-level node has slug "" (parallel to the other main level pages, not their parent); it cannot have children, and renaming its slug away leaves no front page ("/" redirects to the first nav item).
`Node.content` is the Markdown page, or None for a pure category label whose URL renders a 404 listing its children as cards (while nav links to it point at its first child); every label's title and slug are editable. A page with published children — a category page — lists them as cards after its markdown content; the sidebar sub-navigation renders only from the second level down, never on main-level pages.
Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `pages` (pre-tree databases) migrates into `menu` via `migrate_v1`. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`.
Every content/settings write calls `_invalidate_pages()` in state.py, which clears the rendered-body LRU and bumps an in-memory render generation embedded in page ETags, so nav-affecting changes invalidate caches. (This used to be a persisted `Data.version` counter — cache invalidation is not database state, so the field was dropped; old databases lose the key on re-serialization.)
## Files
Files are content-addressed (blake3[:12] + extension) and stored **on disk** under `<hostname>/files/` (path from `PAGERITE_FILES`), served at `/_f/{name}` with immutable caching. Uploaded raster images (except GIF) and SVGs (rasterized) get a set of derivatives: the untouched original under `<hash>.orig<ext>` (internal only — it may carry EXIF data and is never served; SVG originals stay servable as `<hash>.svg`), a mediapreview-recompressed AVIF (`<hash>.avif`, thumbnailed to `IMAGE_MAXSIZE` at `IMAGE_QUALITY`), and WebP/JPEG fallbacks re-encoded from the AVIF at lower quality (`IMAGE_WEBP_QUALITY`/`IMAGE_JPG_QUALITY`, chosen for similar-or-smaller file size). Pages link the bare `/_f/<hash>` and the server negotiates by Accept header: a format is served only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, anything else including `image/*` and `*/*` → JPEG); an explicit extension in the URL pins the format. Responses carry `vary: accept`. Favicons uploaded in settings go through the same pipeline at `FAVICON_MAXSIZE` (192px). Existing databases are updated by `migrate_v2` (link rewrite plus on-disk derivative backfill). Deleting any name of a hash removes the whole group. The `FileStore` in files.py caches every file in RAM, both uncompressed and zstd-compressed (the compressed copy only when smaller), so `/_f` answers both encodings without disk reads. Pages reference files by absolute `/_f/` URLs so hierarchy moves never break them. Pre-refactor databases kept the blobs in a `Data.files` kanta field; the kanta migration `pagerite/migrations.py::migrate_v1` writes them to disk on open and drops the field (removed from `Data`). Fetched favicons of external analytics sites live in the same store (see `docs/analytics.md`).
## Banners
`Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles.
`Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css lives in `<head>` (id `pagerite-banner`) between the theme and the custom CSS — a `<link>` in dev, an inline `<style>` in production.
## Site settings
`Data.brand` is the site name (header link + `<title>` suffix), editable in the site editor via `/_api/settings`; empty = no header link and no `<title>` suffix.
`Data.brand_html` is raw trusted HTML replacing the brand link entirely (rendered in a `#brand` div on top of the banner, next to the nav) — site-wide, not per-page like banners; edited in the site editor with image/video upload into the content-addressed file store.
`Data.theme` is the active theme name (empty = none/base only); themes are folders in `pagerite/themes/{name}` containing `theme.css` and/or `banner.css` (+ `banner.svg` artwork and any extra assets the CSS references, like summer's `grass.svg`), served by the backend at `/_themes/{name}/...` — read from disk per request (etag by mtime), never built, so on-disk edits show on the next page load even in prod. The theme selector and banner-design selector enumerate these folders via `GET /_api/settings`.
`Data.transition` is the page-transition design name (default `cube`): a theme folder shipping `transition.css`, injected as `#pagerite-transition` on every page and selected in the site editor (the selector enumerates `transition.css` folders via `GET /_api/settings`). See `docs/themes-and-assets.md`.
`Data.custom_css` is raw trusted CSS injected inline in every page `<head>` (id `pagerite-user`) and swapped during fetch-navigation; editable in the site editor. Font picks (heading/body/brand) in the site editor are stored as plain `:root` rows in `custom_css` (`--font-body: var(--font-source-sans);` format — parsed out and rewritten on change, the `:root` block added/removed as needed), referencing the per-family variables (`--font-source-sans` etc.) from `pagerite.css`; the base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
`Data.favicon` names a file in the content-addressed store (on disk under `<hostname>/files/`), uploaded/cleared in the site editor via `PUT`/`DELETE /_api/settings/favicon`; when set it is linked as `<link rel="icon">` on every page, otherwise browsers fall back to the build's `/favicon.ico` by convention.
+30 -209
View File
@@ -1,234 +1,55 @@
# Pagerite Design Principles # Pagerite Design Principles
Pagerite is a single-user CMS/blog. This document records the initial Pagerite is a single-user CMS/blog. This document records the initial high-level design decisions; it will be refined as the implementation evolves.
high-level design decisions; it will be refined as the implementation
evolves.
## Architecture ## Architecture
- **Server-side rendered.** FastAPI serves complete HTML pages, generated in - **Server-side rendered.** FastAPI serves complete HTML pages, generated in Python with **html5tagger**. There is no client-side templating or SPA for the public site.
Python with **html5tagger**. There is no client-side templating or SPA for - **Vue only where interactivity demands it.** Small interactive islands (editing tools mainly) are Vue components mounted into specific elements of the server-rendered pages. The public reading experience has no scripting requirement.
the public site. - **Persistence via kanta.** Content is stored in an asyncio-friendly kanta database. Rendering happens on the fly on each request — there are no pre-built static artifacts.
- **Vue only where interactivity demands it.** Small interactive islands
(editing tools mainly) are Vue components mounted into specific elements of
the server-rendered pages. The public reading experience has no scripting
requirement.
- **Persistence via kanta.** Content is stored in an asyncio-friendly kanta
database. Rendering happens on the fly on each request — there are no
pre-built static artifacts.
## Content model ## Content model
- Pages and blog articles are fundamentally the same kind of thing: named - Pages and blog articles are fundamentally the same kind of thing: named pieces of content. The blog/website distinction is blurred; an article is just a page (possibly with metadata such as a publication date and listing in a feed).
pieces of content. The blog/website distinction is blurred; an article is - **Pretty URLs.** Content is addressed by its name (slug), not by technical constructs — no `/cms/...` or `/blog/post1` prefixes. Slugs usually live directly at the site root; structured content may nest (`/docs/design-principles`-style). The URL space is the author's, so reserved prefixes must be kept few and deliberate: everything internal lives under `/_` (`/_api/`, `/_f/`, `/_assets/`). The only other reserved root path is `/favicon.ico`, served from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores (`[a-z0-9_-]`; input is transliterated and filtered as you type, and a new page's empty slug is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
just a page (possibly with metadata such as a publication date and - **Single user, trusted author.** No auth concerns in the core design. Everything published is public; only editing tools will later sit behind access control (external SSO when that time comes). The author is trusted to create well-meaning slugs and content — no sanitization for safety, only for correctness.
listing in a feed). - **Commenting** is not planned now but the model should not preclude it later.
- **Pretty URLs.** Content is addressed by its name (slug), not by technical
constructs — no `/cms/...` or `/blog/post1` prefixes. Slugs usually live
directly at the site root; structured content may nest
(`/docs/design-principles`-style). The URL space is the author's, so
reserved prefixes must be kept few and deliberate: everything internal
lives under `/_` (`/_api/`, `/_f/`, `/_assets/`). The only
other reserved root path is `/favicon.ico`, served from the build.
Slugs are lowercase ASCII letters, digits, hyphens and underscores
(`[a-z0-9_-]`; input is transliterated and filtered as you type, and a
new page's empty slug is derived from its title), may not begin with
`_` or `.`, and such URLs are never looked up as content.
- **Single user, trusted author.** No auth concerns in the core design.
Everything published is public; only editing tools will later sit behind
access control (external SSO when that time comes). The author is trusted
to create well-meaning slugs and content — no sanitization for safety,
only for correctness.
- **Commenting** is not planned now but the model should not preclude it
later.
## Authoring format ## Authoring format
- Content is written in **Markdown** with powerful extensions (tables, - Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
footnotes, code highlighting, etc.). - **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
- **Embedded HTML is passed through unfiltered**, including inline scripts - Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box and `{.margin}` / `::: margin` marks any block a margin note — on all but phone widths they are taken out of flow into the side zone at the article's start edge (left in LTR, right in RTL — the region the nav sidebar overlays, or the sidebar's own track when the layout reserves one) and the text never moves — and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
and other dynamic content the author wants to post. This is safe by the - **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored on disk (`<hostname>/files/`, RAM-cached uncompressed + zstd) by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/…`. Raster images (not GIF) and SVGs (rasterized) are recompressed via mediapreview: the original is kept as `{hash}.orig{ext}` (internal only, never served — it may carry EXIF data; SVG originals stay servable as `{hash}.svg`) while pages link the extension-less `/_f/{hash}` and the server picks from the derivatives (`{hash}.avif` / `{hash}.webp` / `{hash}.jpg`) by Accept header — a format only when listed explicitly (`image/avif` → AVIF, `image/webp` → WebP, otherwise JPEG), with `vary: accept`; an explicit extension in the URL pins the format. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/… "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column, to its end/start edge following the text direction (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.margin}` makes it a margin note, placed in the side zone at the text's start edge on all but phone widths, `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout, and space-separated at the end of a text line (`some text {.small}`) the braces likewise belong to the block — a space is what keeps them off an image or link ending the line, which keep their own directly-attached attrs; text size classes `{.small}` / `{.large}` / `{.huge}` (em-based) work on any block; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
single-trusted-author assumption above.
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition
lists, task lists, brace-attributes; tables and strikethrough from the
default preset), with `html=True` for raw passthrough. Fenced code blocks
are highlighted server-side with **Pygments** (`nowrap` spans styled by
`/_assets/pygments-*.css`, which maps every token class onto the `--code-*`
variables; the base stylesheet defines light and dark palette sets resolved
via `light-dark()`, so each theme gets the set matching its `color-scheme`
and may only retint `--code-bg` to keep the well in the page's color
family); a JS copy button appears on hover. Should this
prove limiting, we implement our own renderer on top of html5tagger,
which we already use for all HTML generation.
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`)
are stored by content hash — blake3, first 6 bytes hex + original
extension — and served immutable from `/_f/{hash}.ext`. Absolute URLs
that survive page renames and dedupe identical content; pages no longer
own files. An image with a title becomes a `<figure>` with
`<figcaption>`. Positioning is by attribute classes:
`![alt](/_f/….avif "Caption"){.right}``{.right}`, `{.left}` float at
30% of the text column (the caption wraps within it; an explicit
`width=300` overrides on uncaptioned images),
`{.wide}` goes full bleed (viewport edge to edge, or up to the docked
editor; the sidebar stacks on top of it); plain attributes like `width=300`
work too. Headings (h1/h2) clear floats, so images never overflow into the
next section.
## Page structure and navigation ## Page structure and navigation
- All pages share one static layout, defined once as an **html5tagger - All pages share one static layout, defined once as an **html5tagger Template** with capitalized placeholders (`Title`, `Banner`, `Nav`, `Sidebar`, `Main`) filled per request. The dynamic regions carry stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`).
Template** with capitalized placeholders (`Title`, `Banner`, `Nav`, - The page top is a **full-width banner header** with the site name and the navigation bar overlaid on it — no separate chrome header. The banner combines two layers, stacked in `#page-banner` (a grid, so they overlay): first the **banner design** — a named design living in a theme folder (`pagerite/themes/{name}/banner.css` plus artwork as `banner.html` — arbitrary markup like canvas + style + script — or `banner.svg`), chosen per page via `Node.banner_design` (a design name, "" for none, None to inherit from the nearest ancestor, then the front page, then the active theme's own design). The artwork is inlined into a `div[data-design]` wrapper: SVG artwork can be recolored from the theme stylesheet (corporate's single SVG serves both light and dark mode via `var()`-driven stops). Second, **per-page author code**: `Node.banner` holds an arbitrary trusted HTML snippet (an image, a styled div, canvas + script — anything), resolved by walking up the node's ancestors to the front page and rendered **after** the design artwork, so author styles always win over the design's own. The base stylesheet falls back to a plain gradient. There is deliberately no scrim fading the banner into the page background — any such fade would ruin user-supplied designs; themes that want one bake it into their SVG (purple does).
`Sidebar`, `Main`) filled per request. The dynamic regions carry stable - **Fetch-navigation.** Links are plain `<a href>`; a small script (`frontend/src/pagerite.js`) intercepts same-origin clicks, fetches the page, and swaps the `#page-banner`, `#nav`, `#sidebar` and `#main` regions, the document title, and the site-wide custom CSS (`<style id="pagerite-user">` in `<head>`), keeping the rest of `<head>` and the layout chrome. Without JS everything works as normal page loads. Scripts inside fetched banner and content regions are re-created so they execute. Swaps run inside `document.startViewTransition` for the page transition selected in the site settings (`Data.transition`; the `cube` design — CSS adapted from termotohtori.fi, fragile, do not tweak — rotates, mirrored on browser back; `crossfade` fades; both skipped under `prefers-reduced-motion`). With `cube`, navigation within the same top-level section crossfades instead of rotating.
ids (`#page-banner`, `#nav`, `#sidebar`, `#main`). - **The site structure is a tree of labels.** `Data.menu` holds the top-level items by slug, each with `children` keyed by slug — the URL path is the slug chain. The front page is a top-level node with slug "" (an item *parallel* to the other main level pages, not their parent) and cannot have children. The header navbar holds only the top level; a top-level item is highlighted when viewing any of its subpages. A page with published children lists them as **cards** after its content (the child page's share image as the cover, like the og tags, with the title overlaid); a **left sidebar** (`#sidebar`) with the section's sub-navigation appears only from the second level down, when there is something to navigate — main-level pages, sections with fewer than two published items, leaf pages and the front page render no aside element at all. Other sections' subitems are never shown without navigating into them first.
- The page top is a **full-width banner header** with the site name and the - **Landing pages are optional.** Every label can either have content (`Node.content`, a Markdown page) or none — a content-less label renders a 404 page listing its children as cards (with a pen to create the landing page) instead of redirecting, while nav links to it point straight at its first child, so categories need no filler content and normal navigation never sees the 404. Title and slug of every label are editable; renaming a slug moves the whole subtree. The sidebar never lists the section itself, avoiding title duplication with the navbar.
navigation bar overlaid on it — no separate chrome header. The banner is - **Menu order is manual.** Each node has a fractional `order` key among its siblings; reordering/moving writes only the moved node (it takes a fresh value halfway between its new siblings; all other items keep theirs). New pages append at the end of their menu. Structure edits (reorder, move/rename with the whole subtree, retitle) go through `POST /_api/structure` and the editor's structure panel.
**per-page configurable**: `Node.banner` holds an arbitrary trusted HTML
snippet (an image, a styled div, canvas + script — anything), resolved by
walking up the node's ancestors to the front page; when nothing in the
chain sets one, the active theme's banner artwork shows. That artwork is
an **inline SVG** (`pagerite/themes/{name}/banner.svg`) the backend
inlines into `#page-banner`: as markup it can be recolored from the theme
stylesheet (corporate's single SVG serves both light and dark mode via
`var()`-driven stops) and it is never rendered underneath a user banner.
The base stylesheet falls back to a plain gradient. There is deliberately
no scrim fading the banner into the page background — any such fade would
ruin user-supplied designs; themes that want one bake it into their SVG
(purple does).
- **Fetch-navigation.** Links are plain `<a href>`; a small script
(`frontend/src/pagerite.js`) intercepts same-origin clicks, fetches the
page, and swaps the `#page-banner`, `#nav`, `#sidebar` and `#main` regions,
the document title, and the site-wide custom CSS (`<style id="pagerite-user">`
in `<head>`), keeping the rest of `<head>` and the layout chrome. Without JS
everything works as normal page loads. Scripts inside fetched banner and
content regions are re-created so they execute. Swaps run inside `document.startViewTransition` for a rotating
cube page transition (CSS adapted from termotohtori.fi — the
`::view-transition*` block is fragile, do not tweak; skipped under
`prefers-reduced-motion`). Navigation within the same top-level section
crossfades instead of rotating; browser back navigation rotates in
reverse.
- **The site structure is a tree of labels.** `Data.menu` holds the
top-level items by slug, each with `children` keyed by slug — the URL
path is the slug chain. The front page is a top-level node with slug ""
(an item *parallel* to the other main level pages, not their parent) and
cannot have children. The header navbar holds only the top level; a
top-level item is highlighted when viewing any of its subpages. When the
current page is inside a main level section with children, those direct
children are listed in a **left sidebar** (`#sidebar`), one level deep.
The sidebar exists only when there is something to navigate — sections
with fewer than two published items, leaf pages and the front page render
no aside element at all. Other sections' subitems
are never shown without navigating into them first.
- **Landing pages are optional.** Every label can either have content
(`Node.content`, a Markdown page) or none — a content-less label renders
a placeholder page (404 with a pen to create it) instead of redirecting,
while nav links to it point straight at its first child, so categories
need no filler content and normal navigation never sees the placeholder.
Title and slug of every label are editable; renaming a
slug moves the whole subtree. The sidebar never lists the section
itself, avoiding title duplication with the navbar.
- **Menu order is manual.** Each node has a fractional `order` key among
its siblings; reordering/moving writes only the moved node (it takes a
fresh value halfway between its new siblings; all other items keep
theirs). New pages append at the end of their menu. Structure edits
(reorder, move/rename with the whole subtree, retitle) go through
`POST /_api/structure` and the editor's structure panel.
- Unpublished pages are hidden from both nav and URL access (404). - Unpublished pages are hidden from both nav and URL access (404).
## Reading experience ## Reading experience
- The article column is sized by the **viewport, never by content**: a - The article column is sized by the **viewport, never by content**: a symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps the layout stable across navigation. The sidebar occupies the left gutter, the right gutter balances it. Long articles (flagged `.multicol` by the backend render) lift the cap and become a bounded **composition**, centered in the available space with the surplus left vacant: a fluid text lane (up to 42rem) plus a 16rem **side zone at the article's left** — the region the nav sidebar overlays — which hosts margin boxes (`.margin`, `::: aside`, margin figures) at all but phone widths, without the text ever moving. On pages with a sidebar, the sidebar gets its own track at every width — flexible, 12rem when space is tight and growing up to 150% (18rem) once the viewport has room beyond the article, the sidebar keeping its left side on the viewport's edge — and the track is the left lane instead: no in-article zone, the text lane runs fluid up to 86rem leaning on the viewport's right edge (surplus extends the left lane), and the boxes hang into the lane off the article's left border (growing leftward with it, up to 18rem), sliding under the translucent sticky nav. Once two lanes fit beside the zone (≥96rem available in `main`), the text flows in two fluid lanes (36rem minimum, capped at 102rem total — technical content wants the wider lanes, and wider windows just add vacant space). The stages step by the space actually available in `main` (container queries + `cqw` units, so the docked editor's inset is automatic). `.wide` figures on multicol pages bleed to the viewport edges measured from `main` (`cqw`), sliding under the sidebar. The backend splits the body into `.colseg` segments at h1/h2 headings and `.wide` elements (full-width separators, never inside columns); margin boxes stay inside the segment at their anchor point and the CSS takes them out of flow — absolutely positioned off the article's left border into the zone, the columns flowing through unaffected — tagging segments that hold enough text in at least two paragraphs (or one long enough to split) with `.cols` — code blocks are excluded from that measure, a `::: nocols` container opts its whole section out, and column-filling paragraphs are marked `.breakable` so they may split across the column gap (shorter paragraphs stay whole). On wide single-column pages (≥104rem), margin boxes lean into the vacant left gutter as well, growing with it up to 18rem.
symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps - A gentle **scroll-reveal** of headings, figures and block-level elements (IntersectionObserver). It is layout-level: articles need no support for it, and `prefers-reduced-motion` disables all motion.
the layout stable across navigation. The sidebar occupies the left
gutter, the right gutter balances it; wide screens get columns inside
long articles without changing the article's width.
- A gentle **scroll-reveal** of headings, figures and block-level elements
(IntersectionObserver). It is layout-level: articles need no support
for it, and `prefers-reduced-motion` disables all motion.
## Styling ## Styling
- The base stylesheet `frontend/src/assets/pagerite.css` provides the layout, - The base stylesheet `frontend/src/assets/pagerite.css` provides the layout, typography and interaction rules with conservative CSS variables. A theme layer (`pagerite/themes/{name}/theme.css` — currently `purple`, `corporate` and `nitro`, served by the backend at `/_themes/{name}/theme.css` straight from disk, never built) overrides those variables and adds the visual styling; `Data.theme` selects the active theme (empty = none/base only) and the site editor can switch it, choosing from the theme folders found on disk. Vue may add per-component styles on top where needed. The corporate and nitro themes switch palettes automatically via `prefers-color-scheme` (corporate is light-first with a matching dark palette; nitro a warm light-grey page or, in dark mode, a deep violet one — its dark banner and orange accents carry over unchanged); purple (dusk) uses one fixed palette for everyone. Themes may restyle structural details the base leaves plain — heading colors and underlines, list markers, nav treatment, brand sizing. A theme folder may also ship a **banner design** (`banner.css` + `banner.svg`), selectable per page independently of the active theme. The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll parameter on `<html>` (event-driven, so it is still when the page is idle), the banner contents drift within their window (with scale overscan so no edge shows), and designs may key their own effects off the same parameter — purple's sun rises as you scroll.
typography and interaction rules with conservative CSS variables. A theme layer - Fonts, the shared stylesheet and pygments styles live under `frontend/src/assets/` and are emitted as hashed assets under `/_assets/` (Source Serif 4 for headings, Source Sans 3 for body, Fira Code for code by default; Fraunces, Literata, Cormorant, Playfair Display, Inter, Montserrat, Cause, Exo 2 and New Rocker kept as woff2 options with local `@font-face`, variable-weight where available). No third-party requests.
(`frontend/src/assets/themes/{name}/theme.css` — currently `purple`, `corporate`
and `nitro`) overrides those variables and
adds the visual styling; `Data.theme` selects the active theme (empty = none/base
only) and the site editor can switch it. Vue may add per-component styles on top
where needed. The corporate and nitro themes switch palettes automatically via
`prefers-color-scheme` (corporate is light-first with a matching dark palette;
nitro a warm light-grey page or, in dark mode, a deep violet one — its dark
banner and orange accents carry over unchanged); purple (dusk) uses one
fixed palette for everyone. Themes may restyle structural details the base
leaves plain — heading colors and underlines, list markers, nav treatment,
brand sizing. The banner artwork has scroll parallax: pagerite.js sets the
`--pry` scroll parameter on `<html>` (event-driven, so it is still when the
page is idle), the banner contents drift within their window (with scale
overscan so no edge shows), and themes may key their own effects off the
same parameter — purple's sun rises as you scroll.
- Fonts, the shared stylesheet and pygments styles
live under `frontend/src/assets/` and are emitted as hashed assets under
`/_assets/`
(Source Serif 4 for headings, Source Sans 3 for body, Fira Code for code
by default; Fraunces, Literata, Cormorant, Playfair Display, Inter,
Montserrat, Cause, Exo 2 and New Rocker kept as woff2 options with
local `@font-face`, variable-weight where available). No third-party
requests.
## Editing ## Editing
- Editing happens **in place**, in two modes opened by two pens: - Editing happens **in place**, in two modes opened by two pens:
- **Page mode** — the 🖊️ next to a page's heading (including 404s, which - **Page mode** — the 🖊️ next to a page's heading (including 404s, which is how new pages start) opens a CodeMirror Markdown editor docked to the left of the article: the panel is fixed to the viewport's left edge (its top tracks the banner's bottom until the banner scrolls away), the content shifts right and the sidebar hides while editing. Preview renders server-side per keystroke (no debouncing) and swaps the whole visible article content in one go (the edit pen and category cards survive the swap).
is how new pages start) opens a CodeMirror Markdown editor docked to - **Site mode** — the ⚙️ at the top right (after the 📊 analytics link, before login) opens a panel with the site **brand** (applied to the header live), a **theme** selector (swapping the theme stylesheet in place), a **page transition** selector (`cube`/`crossfade`, swapping `#pagerite-transition` in place), **font** picks (heading/body/brand — stored as plain `:root` rows inside the custom CSS, referencing the base stylesheet's per-family font variables), a **site-wide custom CSS** field (injected into `<style id="pagerite-user">` in the live page head and swapped during fetch-navigation), the page's **banner design** selector (inherit / none / any design found on disk, inherited by children), the page's **banner HTML** field (supplementing the design, previewed into the real banner region, so you see exactly which banner you're editing) and the **structure tree**. Everything saves immediately as you edit — no save button, no edit mode.
the left of the article: the host sits inside `#content` (below the - Clicking a pen again closes the editor (without saving; a dirty preview reloads the page). The pens are `<button>`s wired up by `pagerite.js` — editing is an action, not a navigation. The editor's WebSocket **reconnects automatically** with local text and pending saves preserved. (All users are trusted authors for now; access control later with SSO.)
banner, never over the footer), the content shifts right and the - **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published controls. Images can be pasted straight into the editor or chosen via a file input: they upload to the content store (`PUT /_api/files/...`) and insert `![alt](/_f/hash)` at the cursor.
sidebar hides while editing. Preview renders server-side per keystroke - The **structure panel** (vue-draggable tree of the whole site, in site mode) covers page management: reorder any menu level, drag across sections, add, delete (two clicks: the button arms, then deletes — no dialogs). Every node is a real label — content-less category rows offer a to give them a landing page. Deleting a category removes only its landing page (the label and its subpages stay). Every non-empty list ends with a row that starts a new page as a local-only tree row at that level; the row can be dragged into place before its title and slug are filled in and is persisted only on commit. While dragging, these rows double as "end of this list" drop targets; dropping ON the lower part of a row makes the page that row's first child (even a leaf's, creating a sublist), while a row's exposed top edge inserts a sibling before it. A dragged row's indentation previews the target list's depth. Rows are always editable: titles save while typing, slug edits commit on blur/Enter since they rename the path (moving the whole subtree). The front page is the root row with an empty slug — renaming it away leaves no front page ("/" redirects to the first nav item), and giving another top-level row the empty slug makes it the front page.
(no debouncing) straight into the visible article's heading and body. - Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a stateless JSON protocol (`open`/`render`/`save`; on save all fields are optional and absent ones keep their old values, `move_from` renames), avoiding REST polling and races. Rendering always stays server-side.
- **Site mode** — the 🖊️ on the banner opens a panel with the site - A REST API also exists for scripting, all under `/_api/`: `GET pages` (the full tree), `PUT/DELETE pages/{path}`, `GET/PUT settings` (site brand, theme and custom CSS), `POST structure` (reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`.
**brand** (applied to the header live), a **theme** selector (swapping - On startup, seed pages from `pagerite/seed.py` are added **only if missing** — existing user content is never overwritten.
the theme stylesheet in place), **font** picks (heading/body/brand —
stored as plain `:root` rows inside the custom CSS, referencing the base
stylesheet's per-family font variables), a **site-wide custom CSS** field (injected
into `<style id="pagerite-user">` in the live page head and swapped during
fetch-navigation), the page's **banner HTML** field (previewed into the
real banner region, so you see exactly which banner you're editing) and
the **structure tree**. Everything saves immediately as you edit — no
save button, no edit mode.
- Clicking a pen again closes the editor (without saving; a dirty preview
reloads the page). The pens are `<button>`s wired up by `pagerite.js`
editing is an action, not a navigation. The editor's WebSocket
**reconnects automatically** with local text and pending saves preserved.
(All users are trusted authors for now; access control later with
SSO.)
- **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published
controls.
Images can be pasted straight into the editor or chosen via a file
input: they upload to the content store (`PUT /_api/files/...`) and
insert `![alt](/_f/hash.ext)` at the cursor.
- The **structure panel** (vue-draggable tree of the whole site, in site
mode) covers page management: reorder any menu level, drag across
sections, add, delete (two clicks: the button arms, then deletes — no
dialogs). Every node is a real label — content-less category rows offer
a to give them a landing page.
Deleting a category removes only its landing page (the label and its
subpages stay). Every non-empty list ends with a row that starts a
new page as a local-only tree row at that level; the row can be dragged
into place before its title and slug are filled in and is persisted only
on commit. While dragging, these rows double as "end of this list"
drop targets; dropping ON the lower part of a row makes the page that
row's first child (even a leaf's, creating a sublist), while a row's
exposed top edge inserts a sibling before it. A dragged row's
indentation previews the target list's depth. Rows are always
editable: titles save while typing, slug edits commit on blur/Enter
since they rename the path (moving the whole subtree). The front page
is the root row with an empty slug — renaming it away leaves no front
page ("/" redirects to the first nav item), and giving another
top-level row the empty slug makes it the front page.
- Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a
stateless JSON protocol (`open`/`render`/`save`; on save all fields are
optional and absent ones keep their old values, `move_from` renames),
avoiding REST polling and races. Rendering always stays server-side.
- A REST API also exists for scripting, all under `/_api/`:
`GET pages` (the full tree), `PUT/DELETE pages/{path}`,
`GET/PUT settings` (site brand, theme and custom CSS), `POST structure`
(reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`.
- On startup, seed pages from `pagerite/seed.py` are added **only if
missing** — existing user content is never overwritten.
+35
View File
@@ -0,0 +1,35 @@
# Editing interface
The Vue editor is a single tabbed `EditorShell.vue` mounted in a host div created inside the static document.
## Tabs
The shell hosts five kept-alive tabs (ordered site-wide first — site, structure, localization — then, after a visual break, the per-page tabs — article, banner):
- `PageEditor.vue` — CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`, previewing into the visible article; editor and article scrolls are linked piecewise-linearly, keyed on the section anchors' `data-line` (markdown source line the backend stamps on top-level anchored h1/h2s): the page follows the cursor (fractional, wrap-aware, scrolling only when the cursor's page position leaves the viewport, with an edge margin), the editor follows page scroll with a progress-based viewport anchor, applied instantly (the window keeps scrolling normally while any editor is open — the panel is fixed to the viewport's left edge, its top tracking the banner's bottom edge until the banner scrolls away — and the panel scrolls internally); anchored h2s carry their own edit pens that open the editor scrolled to that section; a format bar offers Markdown helpers — bold/italic/code/link/table/image upload (always block-level on a fresh blank-separated line of its own — a cursor on a non-empty line, e.g. inside an existing image tag, inserts after that line, never into it; always with an empty `""` caption, cursor inside the quotes), toggling fences (` ``` ` code blocks and `::: aside` containers share the same machinery: clicked inside one they remove it and select the content, otherwise they wrap the selection or the cursor's line, keeping it selected), and `.left`/`.right`/`.wide`/`.margin` placement toggles plus `.small`/`.large`/`.huge` text-size toggles (brace attributes on the block at the cursor, mutually exclusive within each group; on `:::` containers a placement class replaces the container name instead — `::: aside``::: margin`), with Ctrl/Cmd-B/I/S bindings — for the hard-to-remember syntax. Edits content and title only, never the path.
- `BannerEditor.vue` — per-page banner HTML + banner design selector, previewed into `#page-banner`.
- `SiteEditor.vue` — site brand + optional custom brand HTML with image/video upload + theme selector + page-transition selector + font picker + favicon upload — clicking the preview tile picks a new one — + site-wide custom CSS, CSS injected into `<head id="pagerite-user">`.
- `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row, plus a per-row flag dropdown setting the page's primary language (`Node.language`, inherited by the subtree).
- `LocalizationEditor.vue` — the site-wide translation settings: target languages as a flag grid (toggles, grouped in geographic rows; see docs/localization.md), the refresh-all-translations button, and the translator service WebSocket URL(s) to connect `scripts/translator.py` to.
Media uploads everywhere use the image icon buttons (pasting into the editor works too). The article, banner and site-settings pens are shorthands that open the shell on the matching tab; once open, clicking a pen switches tabs (and retargets the editors to the current page) instead of closing/remounting. The close button in the tab bar closes the shell (deliberately NOT Escape — it fired too easily by accident); tabs have no close buttons of their own. Closing only HIDES the shell — the Vue app stays mounted, so page-editor state (unsaved text included) survives until a real page reload; the editor always follows the URL, so fetch-navigating with the shell open (or before re-opening it) retargets it to the new page — unsaved text is stashed per path for the session and restored when returning, cleared on save. Saving there is explicit (Ctrl+S) and refreshes the page regions in place. Admin panels never reload the page.
In-place page re-rendering shared by the banner/site/structure tabs lives in `swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic regions, replaceState). It also exports `dropPageCache`, which the editor tabs call after any save that can alter the rendered HTML of other pages (theme, headings, structure, banners, site brand/CSS, favicon). Dropping the cache while editing avoids re-fetching every page immediately; the public runtime re-preloads visible links once the editor panel closes.
The page and structure tabs share one language selector: `LangSelect.vue` (small flag + dropdown) v-modeled on the shell-wide selection in `editorLang.js` (`''` = primary). While the panel is open that selection overrides the page's normal language preferences: EditorShell calls `swapdoc.setLangOverride`, which pins every `loadPlain` fetch (`?lang=`, the primary by its own code) and pagerite.js's own fetches/prefetches (`pagerite:session-lang`), until the panel closes and the override clears.
All WebSockets (page/banner editors, analytics view, the pagerite.js activity channel) pace their connections through `reconnect.js`: new sockets are created a staggered slot apart (a page load opens Vite's HMR socket plus several of ours at the same moment, and such bursts — like rapid retries — trip the browser's WebSocket throttling, leaving every socket to the host "pending" for minutes), a watchdog closes sockets stuck CONNECTING so they reschedule instead of hanging forever, and retries follow an exponential backoff with jitter that only a healthy connection resets. While a socket is connecting or waiting to reconnect the panel says so (`ConnNote.vue`), and the CodeMirror editors stay locked until their document arrives (typing before the doc accept would be clobbered by it).
## Saving behavior
Everything saves immediately as you edit (brand/title/CSS debounced, slug on commit since it renames the path), theme change swaps the stylesheet in place, tree rows navigate in place without transitions when focused, and the front page is a root-only row whose empty slug is editable like any other. Saves that can affect other pages drop the prefetch cache; the cache is rebuilt when the editor panel closes so navigation stays instant.
Every non-empty list (and the root) ends with a non-draggable plus footer row (vuedraggable `#footer` slot): clicking it starts a new pending page at that level (its slug placeholder shows the slug derived live from the title being typed), and while dragging it is the list's "end of list" drop target. Committing a pending page PUTs it with empty markdown (creates an empty page that renders with its title — saving never deletes; deletion is the page editor's explicit choice: saving trimmed-empty text issues a REST DELETE), then switches to the page editor tab for the actual writing.
Dropping ON the lower part of a row moves the page under that row (the child list's container invisibly overlaps its own row's bottom via negative margin — Sortable inserts it as the first child natively), while a row's exposed top edge inserts a sibling before it. Row indentation is structural (each nested list margin-indents itself), so a dragged row previews its whole subtree at the target list's depth.
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js`, `src/pagerite.js` and `src/analytics-main.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme, banner-design and transition CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-transition`/`#pagerite-user` elements to restore the canonical order (base < theme < design < transition < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
+32
View File
@@ -0,0 +1,32 @@
# Frontend runtime
The public page runtime lives in `frontend/src/`.
## `main.js`
Vue editor app entry, mounts the tabbed `EditorShell`. See `docs/editing.md` for the editor UI.
## `pagerite.js`
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire). Editors can drop the entire cache with the `pagerite:drop-page-cache` event when site-wide or page changes (theme, headings, structure, banners, etc.) invalidate the cached HTML of other pages; `main.js` triggers a fresh `pagerite:preload-pages` pass when the editor panel closes so navigation is fast again. Navigation that starts while the editor is open bypasses the cache and fetches the target page on demand. Also runs scroll-reveal, a scroll-driven section hash (the location hash tracks the h1/h2 above the viewport middle via replaceState — removed above the first tagged heading and at the very top, never set on unscrollable pages), OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, click-to-enlarge on article figure images (a full-viewport lightbox with the caption, closed by click or Esc), and the auth check.
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`.
Asset wiring differs by mode. In dev the backend links the Vite dev-server URLs (`pagerite:editor-src`/`-css`/`pagerite:analytics-src` meta tags, `<link>` stylesheets) and Vite injects the entry CSS from JS for hot reloads. In production there are no pagerite meta tags: all page assets are inlined into the document — stylesheets as `<style>` elements in `<head>` (fixed order: base, theme, banner design, page transition, entry sheets, custom CSS last), module scripts as inline `<script>`s at the end of the body (relative chunk imports are rewritten to absolute `/_assets/` paths) — and the on-demand bundles' URLs ride in a `<script type="application/json" id="pagerite-assets">` config. The editor bundle always stays external, imported on demand when a pen is opened. Every stylesheet element carries a stable id so fetch-navigation and the site editor can sync `<head>` positionally across swaps (the analytics sheet exists on `/_a` only and is added/removed as you navigate). The analytics entry is inlined into the `/_a` page itself; pagerite.js re-creates that script element after fetch-navigating there (inline scripts don't execute on a DOM swap) and calls the module's exposed unmount before swapping away.
## `assets/`
Shared styles and data files built by Vite and served hashed under `/_assets/`: `pagerite.css` (base layout + conservative variables), `pygments.css`, and `fonts/` (self-hosted Source Sans 3/Source Serif 4/Fraunces/Literata/Cormorant/Playfair Display/Inter/Montserrat/Fira Code/Cause/Exo 2/New Rocker variable woff2).
The `::view-transition*` rules live in the page-transition designs (`pagerite/themes/{cube,crossfade}/transition.css`), not in the base stylesheet. Themes, banner designs and transitions are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details.
Vite builds ES-module `.js` outputs; in dev the backend links them as `<script type="module">` (module scripts defer by default), in production it inlines them at the end of the body.
## Data directory
All site data lives under `<hostname>/` in the cwd — `content.kantadb`,
`analytics.json` and `files/` — where `<hostname>` is the CLI's first
positional argument (default `localhost`, passed to the app as JSON in
`PAGERITE_CONFIG`, see `pagerite/config.py`;
`PAGERITE_DB`/`PAGERITE_ANALYTICS`/`PAGERITE_FILES` override individual
paths). gitignored. Do not delete it without asking.
+513
View File
@@ -0,0 +1,513 @@
# Localization
Pages are served in the visitor's language based on a `?lang=` query
parameter or the `Accept-Language` header.
- **Phase 1 (implemented):** negotiation, URL scheme, caching, rendering
plumbing. Translations are consumed through a stub interface; the database
still holds only the original language.
- **Phase 2 (implemented):** gettext-style fragment storage in the
database — machine-translated chunks plus user override patches, assembled
at render time. Storage details in `docs/migrate.md`.
## Phase 1: negotiation and URLs
### The primary language
Each article has a primary (original) language: `Node.language`, inherited
down the tree like `banner` — "" = the nearest ancestor's, the front page
last (it doubles as the site default), with `en` as the final fallback
(`ORIGINAL_LANGUAGE`, `primary_lang()` in `pagerite/i18n.py`). It is
configured per row in the structure editor. Everything per-article keys
off the resolved value: language selection, `<html lang>`, canonical URLs,
what counts as a translation, and the translation targets (a node's own
primary is never one — so the target set may include the site default, and
a page in another language can be translated into it).
### Language selection
Deliberately simple — **q-values are ignored**:
- All known `Accept-Language` implementations send the header **in order of
preference**, so we parse it as an ordered list and never reorder.
- Selection rule (`select_language` in `pagerite/i18n.py`):
1. If `?lang=<tag>` is present, use it (if a translation exists; otherwise
fall through to header logic).
2. If the article's original language appears anywhere in the header list,
use the **original**. Rationale: an AI translation is strictly worse
than the original for anyone who has that language configured at all
(e.g. `fi-FI, fi, en-US, en` gets English, not machine-translated
Finnish).
3. Otherwise walk the header list in order and use the first language for
which a translation exists.
4. Fall back to the original.
Region tags normalize to their base subtag (`fi-FI``fi`).
### URLs: pretty for users, indexable for search engines
- Canonical URLs stay pretty (`/some-page`). Each language version is
addressable as `/some-page?lang=fi` so search engines can index them.
- `<link rel="canonical">` names the **actually served language**: the plain
URL when serving the original (for SEO the non-query URL means the
article's own language), `?lang=xx` when serving a translation — however
the language was arrived at (query or header).
- `<link rel="alternate" hreflang="…">` entries follow the canonical
directly (before the social meta tags) and list the languages the page
is **actually available in**: `x-default` first, pointing at the plain
autodetecting URL, then every available language — the original again by
its plain URL, translations by `?lang=`. The public language selector
keys off these: pagerite.js mounts the editors' flag dropdown in the
top-right corner when the head advertises x-default plus more than one
language, loading its bundle (Vue + the flag SVG set) on demand.
- The override sticks for the session of clicks: a page requested with
`?lang=` replicates the query onto the navigation links it renders (nav,
sidebar, cards, brand — in-article links are content and stay as
authored), so plain clicks and no-JS navigation keep the language.
pagerite.js additionally strips the query from the address bar via
`history.replaceState` (pretty, shareable URLs), remembers the language,
and adds it to every internal fetch that lacks one (preloads,
fetch-navigations, history traversals); history entries stay query-less.
- The public selector's pick is the same override, pure JS state
(`pagerite:set-session-lang`): the session language changes and the page
swaps in place — no `?lang=` in the address bar, no reload. The choice is
linked with the editor panel's language dropdown both ways; closing the
panel keeps the chosen language instead of reverting.
- A full page refresh or a shared link resets to automatic selection (header
only). This gives a clean one-time override without cookies.
### Response correctness
- Content responses carry `Vary: accept-language` (added to the existing
`accept-encoding` vary).
- `_cached_body` and the page ETag include the **selected language** (not the
raw header, which would blow up the cache key space) and the **replicated
link language**: a `?lang=fi` render and a header-selected Finnish render
of the same page differ in their navigation links, so they are cached as
separate variants.
- `<html lang="…">` reflects the served language, and an RTL language
(`i18n.RTL_LANGUAGES` — ar, fa, he, ur) also sets `dir="rtl"` on `<html>`
(the editor panel carries its own `lang="en" dir="ltr"` so it stays LTR).
Client-side page swaps (fetch navigation in pagerite.js, editor re-renders
in swapdoc.js) copy both attributes from the fetched document, so a hot
switch into or out of an RTL page flips the layout without a reload.
### Rendering
- The translated Markdown goes through the same `markdown.render` pipeline.
- Section anchors (`#hash` ids on h1/h2 headings) stay in the original
language: render(anchors_from=...) pins the translated render's heading
ids to the original text's slugs, matched by heading position, so links
to sections don't break across languages.
- Navigation/sidebar titles come from the translation's title map, with
per-node fallback to the original title (a partially translated tree must
still render).
- Category placeholder pages (the 404s for content-less labels) select a
language like content pages, but over the **subtree's** combined
availability (`subtree_languages`) — they have no chunks of their own;
the heading, navigation and card text localize from the title map and
the target articles' translations. Their hreflang alternates are
computed exactly like a content page's (a translated title counts as
availability, so the language selector is offered there too).
- Card descriptions and cover picks run on the target article's hybrid
Markdown where that page is available in the served language, with
per-card fallback to the original.
- Fixed UI strings ("Not Found" etc.) and the editor UI stay English for now.
- The markdown typographer (SmartyPants) is English-centric; per-language
typographer options are a possible follow-up, not blocking.
## Phase 2: fragment-based translation storage (implemented)
Phase 1 assumed whole-page translated Markdown delivered from outside. The
refined model is gettext-style: an article has **one primary version** (its
`content`, in its own language) plus, per target language, **machine
fragments** (translated chunks of Markdown) and **user patches** (minimal
editor overrides). Both are stored in the database and assembled into the
served Markdown at render time.
### The scenario this must handle
1. Article written in English.
2. Machine-translated into Spanish → fragments stored.
3. Editor fixes one Spanish paragraph and changes a link elsewhere to point
at a Spanish resource → user patch hunks stored.
4. English article edited → the edited chunk's key changes; its Spanish
fragment no longer matches.
5. Page requested before the machine translation refreshes → served as a
**hybrid**: old fragments for unchanged chunks, plain English for the
edited chunk. User patches are attempted against this hybrid, best effort,
each hunk independently: the text fix is stale (its search text no longer
exists) and silently skipped; the link change still applies even though
the link sits in the now-English paragraph.
6. Machine translation refreshes → full Spanish again, with both patch hunks
applying.
### Chunks
`chunk_markdown(markdown)` splits the source into block-level chunks —
blank-line-separated blocks: headings, paragraphs, code fences (kept whole),
list blocks, tables, HTML blocks. A chunk's identity is its **source text**,
gettext-msgid style:
```python
chunk_key = blake3(normalize(chunk_text)).digest(9) # bytes; base64 at the JSON level
```
(`normalize`: strip trailing whitespace per line, collapse surrounding blank
lines — so whitespace-only source edits don't invalidate translations.)
Consequences:
- Editing the English source invalidates exactly the edited chunks; all
other fragments keep applying. Stale fragments are simply never referenced
again and can be garbage-collected lazily (or left; they are tiny).
- No explicit "source version" bookkeeping is needed — staleness falls out
of the keys.
### User patches
Editors always edit **full Markdown** in the existing editor UX — never
fragments. When editing a translated view (`?lang=es`), the editor is loaded
with the *current hybrid Markdown*; on save, the server computes a minimal
diff against that hybrid and stores it as a patch:
```python
class Patch(msgspec.Struct, omit_defaults=True):
"""One editing session's overrides, applied independently per hunk."""
hunks: list[tuple[str, str]] = [] # (search, replace) on hybrid Markdown
```
Hunks are produced from `difflib.SequenceMatcher` on the hybrid vs. the
edited text at block granularity: each `replace`/`delete`/`insert` opcode
becomes one `(search, replace)` pair, with the preceding block's tail as
left context for `insert` (pure inserts have empty search context otherwise).
Application is dead simple:
```python
def apply_patch(hybrid: str, patch: Patch) -> str:
for search, replace in patch.hunks:
if search and search in hybrid:
hybrid = hybrid.replace(search, replace, 1)
# missing search text = stale hunk -> silently skipped
return hybrid
```
Per-hunk independence is the robustness property from the scenario: a stale
text fix does not block a still-valid link change. Patches are stored as an
ordered list and applied in order.
### Storage
Full storage design and the `migrate_v3` restructuring live in
`docs/migrate.md`. The short version, as it concerns this document:
- Originals **and** translations are content-addressed text chunks in flat
stores: `Data.chunks: dict[bytes, str]` and
`Data.trans: dict[bytes, dict[str, str]]` (chunk hash → lang → text) —
path-independent, so repeated paragraphs and menu titles are translated
once and article moves touch nothing. `Node.chunks: list[bytes]` gives
each article its order.
- `Node` gains **`language: str = ""`**, inherited down the tree like
`banner` (empty = nearest ancestor, front page last, site default `en`
final). `select_language` and `<html lang>` use the resolved value instead
of the global `ORIGINAL_LANGUAGE` constant.
- **Known weakness:** changing a page's (or subtree's) `language` after
translations exist mis-keys everything — translations are keyed by
*source* chunks, so old entries silently stop matching and user patches
(searching for old-hybrid text) mostly go stale. That is acceptable:
the orphaned data is harmless and translations regenerate. We do not
migrate translations across a language change.
- Article paths are stored and keyed **without leading slashes**
(`"docs/setup"`, front page `""`); slashes are added only in hrefs.
### Render pipeline (the phase-1 `get_translation` stub, now real)
```python
def get_translation(data, path, lang) -> Translation | None:
if lang not in node.langs:
return None
hybrid = "\n\n".join(
chunks[h] if h in node.no_trans else trans.get(h, {}).get(lang, chunks[h])
for h in node.chunks
)
for patch in data.patches.get(f"{path}:{lang}", []):
hybrid = apply_patch(hybrid, patch)
return Translation(markdown=hybrid, titles=title_map(data, lang))
```
- Availability is an article-level index: `node.langs: dict[lang, True]`,
maintained by the translation writers (translator job, patch saves) in the
same transaction as their data writes — rendering and language selection
never probe the `trans` store chunk by chunk. A stale key is benign (the
"translation" just renders as the original).
- `titles` for nav/sidebar/cards: each node's translated title is
`trans.get(hash(node.title), {}).get(lang)` with per-node fallback — one
dict lookup per nav item at render time.
- Cache invalidation: writes to `chunks` / `trans` / `patches` (translator,
editor saves) call `_invalidate_pages()`, same as content writes.
### Editor flow
The page and structure editors share one language selector (`LangSelect.vue`:
a small flag button opening a dropdown; the same country-flag-icons set as
the analytics visitor cells), v-modeled on one shell-wide selection
(`editorLang.js`, `''` = the primary language). The page editor lists the
page's own primary language (`Node.language`, resolved through the
hierarchy and echoed in the WS doc as `primary_lang`) plus the union of
the page's translations (`node.langs`) and
the site-wide `translate_langs`; it always opens in the primary language,
even when the page itself was served in a translation. A note under the
toolbar states the blast radius:
edits to the primary language re-chunk the original (invalidating the
affected translation fragments everywhere); edits to a translation stay
local to that language.
While the editor panel is open, its language selection **overrides the
normal language preferences** for the page preview: EditorShell pins every
in-place re-render and pagerite.js fetch/prefetch to it (`?lang=` — a
primary selection pins by the current page's own resolved primary, which
`select_language` honors), and closing the panel restores the normal
preferences.
- WS `open` with a `lang` returns the effective **hybrid** Markdown and
title for that language (ungated by `node.langs` — a language without
any fragments yet starts from the original text), plus the language
metadata (`lang`, `primary_lang`, `langs`, `translate_langs`).
- The editor keeps a **shadow copy** of the Markdown it opened. WS `save`
with `lang` sends it as `base`; the server diffs `base` → submitted text
(`make_patch`) and appends a `Patch`. Diffing against the shadow (rather
than the current hybrid) keeps hunks correct when the original or the
machine translation moved under an open editor; application against the
then-current hybrid stays best-effort per hunk, as designed.
- A changed **title** on a translated save becomes a fragment in
`Data.trans` keyed by the original title's chunk hash — the same storage
as machine title translations. An untouched title field (holding the
served translation) is not sent, so saving never freezes a stale machine
title into an override.
- Saving never deletes; a translation additionally cannot be emptied (that
would render as a blank page in that language).
- The live preview renders the version being edited, whichever language
the page itself was loaded in (the render is just the edited Markdown +
title). A translated save keeps that preview in place — re-fetching the
page would come back in the header-selected language.
- Saving the primary-language version re-chunks the submitted Markdown and
updates `Data.chunks` / `node.chunks` — only genuinely new text lands in
the kanta change diff (see docs/migrate.md).
The **structure editor** selects from the same languages with the same
`LangSelect` (the selection is shared — switching in either tab switches
both, and the preview). It is also where a page's **primary language** is
configured: each row carries a small flag dropdown (the resolved flag,
dimmed while inherited) that sets `Node.language` via a structure op —
'' = inherit, so setting it on a section covers the whole subtree. The
tree it
lists (`GET /_api/pages?lang=`) comes back with per-language titles where a
translation exists (`translated` marks those rows; untranslated rows show
the original title, dimmed). Retitling in a non-primary language posts the
structure op with a `lang` and writes a per-language title fragment in
`Data.trans` (keyed by the original title's chunk hash, exactly like a
machine title translation — a user edit simply overwrites it); sending the
original's text drops the override. The structure itself — slugs,
hierarchy, order — is language-independent, so pending rows, slug edits,
drag-and-drop and deletes work identically in every language.
### Translator service API
An external machine-translation service connects over WebSocket at
`/_translate/{key}` — deliberately **not** under `/_api`: the SSO
forward-auth does not cover that route, and the key in the path is the
access control. Keys live in `Data.translate_keys` (key -> display name) —
12 lowercase alphanumeric characters each; the first is generated at
database bootstrap, further ones are managed in the editor's lang tab
(add/rename/delete ride the `PUT /_api/settings` round-trip; the name is
an inline display label only). The full WS URL(s) are printed in the
startup log (`ws://localhost:{port}/_translate/{key}` locally,
`wss://{hostname}/_translate/{key}` on a public hostname) and shown in the
lang tab as click-to-copy links; the keys are also surfaced in
`GET /_api/settings` as `translate_keys`. An unknown or empty key rejects
the handshake (close-before-accept → HTTP 403). Transactions storing results record the connecting key as the kanta
transaction `user`.
Frames are JSON-encoded tagged msgspec structs (`pagerite/translate.py`;
`bytes` fields ride as base64):
- `{"type": "hello", "langs": [...]}` — client greeting announcing its
**capabilities**: the language codes its model can produce (normalized
to base subtags; `en`/empty dropped).
- `{"type": "job", "lang", "key", "texts", "path", "kind", "contexts"}`
server push: ONE fragment to translate (an article title or a chunk), as
a list of **prose segments** (see Segmentation below). `contexts` is
parallel to `texts` ("" = none): the surround to translate the segment
in — for clients that translate better with context (see below).
Contexts are not part of the result.
- `{"type": "result", "lang", "key", "texts"}` — client reply: the
segments translated, same order and count, matching its job by (lang, key).
Which languages get translated is **server-configured**:
`Data.translate_langs` (presence-key dict, bootstrapped to Spanish and
Chinese — edited in the editor shell's localization tab, whose flag grid
lists every language including English, or set via `/_api/settings` as
`translate_langs`). A target equal to an article's own primary language is
skipped per article (its original already is that language), so the set
may freely contain the site default. The dispatcher offers a
connection jobs only in `wanted ∩ capable`; a connection without overlap
simply stays idle.
`DELETE /_api/translations` (the localization tab's "refresh all
translations" button) drops every machine translation (`Data.trans`) and
rebuilds the availability index (`node.langs`) from the surviving user
patches, so the dispatcher re-translates everything from scratch; the
run's validation skip-list is cleared with it, giving rejected fragments
another chance.
Dispatch semantics (the `Dispatcher` in `pagerite/translate.py`; api.py only
registers the route):
- **One job at a time per connection** — the next job is sent only after
the current one's result. Clients wanting parallelism open multiple
connections (e.g. several `scripts/translator.py` instances).
- Pending work is derived from the `trans` store
(`translate.pending_items`) minus the items in flight on any connection,
so a **disconnect requeues** that connection's in-flight item and it is
offered to any free capable connection.
- Dispatch re-runs on every relevant event: Hello, result, disconnect and
content change (`_invalidate_pages()` schedules it, so the pass runs
after the writing transaction commits).
- A result with no job in flight, a mismatched (lang, key), a duplicate
hello, or any malformed frame closes the socket with a protocol error.
Results are stored into `trans` in one transaction and set
`node.langs[lang]` on every article they touch (shared chunks make several
pages gain a language from one fragment). Unknown keys are stored anyway
and re-storing overwrites — results are idempotent.
#### Segmentation
Fragments cross the wire as **prose segments** (`pagerite/segments.py`): the
fragment is parsed with the project's own markdown-it setup
(`markdown.make_md(verbatim=True)` — all extensions, but no typographer or
tasklist label wrapping, so token text stays byte-identical to the source)
and split into the runs a model may touch: paragraph/heading/table-cell text
(merged across soft line breaks), image alt texts and captions, footnote
bodies. A block of plain text, inline **links and paired text formatting**
(strong/em/s) **stays whole** — link and formatted texts cross inline, in
sentence context, with the Markdown stripped (see below). Everything else
never leaves the server: code spans and
fences, URLs and autolinks, link/image *destinations*, `{...}` spans
(placeholders like `{dates}` as well as attrs), reference and footnote
labels, container fences, GFM alert markers (`[!NOTE]`), raw HTML — and the
remaining markup punctuation (`|`, `:::`), which is a run boundary.
Chunks with no segments (a lone `{dates}`, container fences, pure
code/HTML) are never dispatched at all (`needs_translation`); every
language renders them from the original chunk. Each segment is accompanied
by a context string (a segment carved out of a larger block carries the
block's plain text; a whole-block segment carries "") — context is a
prompt aid only, never spliced into the result.
Reassembly is offset splicing, not text the model produced: each segment's
source span was located at dispatch (sequential search; a run that is not a
verbatim source substring — entity-decoded text, backslash escapes — is
skipped and stays in the original language), and the returned translations
are swapped in by offset. Markup corruption is therefore impossible by
construction; the failure modes that remain are a wrong segment count, an
empty segment, markup injected INTO a segment (a `<br>` in a title
translation would splice live HTML), or a line that would start a new
block where the segment lands (a ``` or ::: fence line would eat the rest
of the block it splices into, closing fence included — segments are
inline prose, so `pure_prose` alone cannot see this) — each returned
segment must parse as
pure prose with no block-starting line or blank line, or the whole result
is dropped and logged, and the (lang, key)
pair is skipped for the rest of the server run (generation is
near-deterministic, so an immediate retry would re-fail; the fragment stays
pending and gets another chance on restart or `DELETE /_api/translations`).
`Data.trans` therefore only ever holds clean translated Markdown.
Link- and formatting-carrying blocks are the one place a segment is not
spliced verbatim: a label translated apart from its sentence comes back
grammatically incompatible with it (case government, particles, word
order), and shown the Markdown the model mangles it (Seed-X dropped the
`**` and the glued-on colon in `**Pagerite**: …`), so the block crosses
whole — all Markdown stripped — and the server re-inserts the link and
formatting syntax into the translated block. The boundaries are found by
**text processing alone** —
markers on the wire are hopeless (an earlier sentinel-masking design let
the model see and mangle exactly that punctuation: Seed-X renumbered the
tokens and turned `![` into `¡¡…!!`). Each mark's source words are aligned
to the translation's words by **form similarity** (`_find_mark` in
segments.py): a word-level alignment (sequence ratio plus shared prefix,
case-folded — inflection moves word endings, `banana` → `banaanilla`, and
articles or prepositions drop out; a mid-sentence capital on BOTH sides
earns a bonus, naming conventions being the likeliest shared cause) with
small penalties for skipped words, so reordering and dropped function
words don't break the match. An alignment is accepted only with an anchor
(one pair of similarity ≥ 0.7) and a decent average, and the slice is cut
exactly at word boundaries, so the whitespace between the mark and its
neighbors stays in the plain text. A mark with no convincing alignment
falls back to its weight ratio in the source block (word units before its
text boundaries over the block total) applied to the translation's units
— CJK ideographs count as one unit each, kana runs as one; for CJK targets
the fallback IS the path, cross-script form similarity being nil. Placement is
approximate and several reordered marks in one block can still cluster —
the accepted trade: better a coherent sentence with a slightly shifted link
than separately translated snippets that don't fit together. A boundary
that maps to an empty slice degrades to the source text rather than
emitting a broken `[](url)` or `**`. Blocks mixing in any other inline
markup (code spans, images, raw HTML) don't qualify and still split into
runs at those boundaries.
Punctuation is the translator's own job: Seed-X tends to "finish" short
labels (titles, nav items) with a comma or period the source never had.
Prompt wording is NOT the fix — a punctuation-instruction clause made
Seed-X slip into its `[COT]` reasoning mode (minutes-long generations with
reasoning text in the output, observed for Chinese). The reference client
enforces punctuation deterministically instead (`match_punctuation` in
scripts/translator.py): a translation of a segment without terminal
punctuation gets any added trailing marks (and a newly opened Spanish ¡/¿)
stripped before the result goes back.
The same client-side enforcement covers markup bleed as a CLASS, not per
artifact: `<` is the prose/markup boundary on the wire and never appears in
a segment in either direction. A literal `<` in the source text (`<1MB` is
text, not markup — a tag needs a letter or `/!?`) crosses encoded as the
fullwidth `` and is decoded on return, before the result is validated and
spliced (segments.py) — the wire itself still never carries `<`, and the
reference client cuts the model's output at the first `<`
(scripts/translator.py) — echoed language tags, stray `<br>`s and any
future variant are one handled case. (The cut is post-decode, not a
generation stop string: Seed-X opens every generation with its `<s>`
framing token, which would trip a `<` stop immediately.)
Server-side, a second layer covers what the inline parser cannot: ASCII
punctuation that is plain prose on the wire but Markdown syntax in the
splice context — quotes (a translated `"` would close the quoted image
title it lands in), brackets (alt texts, re-inserted link texts), `|` in
table rows, `\` escapes. Rather than rejecting such results, `join` swaps
them for Unicode look-alikes before splicing (`_NEUTRAL` in
segments.py — curly quotes, fullwidth brackets; the renderer's
typographer curls straight quotes anyway).
Short fragments get more than a bare prompt: each segment may carry its
surround in `Job.contexts` — a title carries the article's opening prose
(its own block is just the title word), a segment carved out of a larger
block (a partial run; a link text whose block didn't qualify for the
whole-block treatment) carries the block's plain text, and a
whole-block segment (a plain paragraph) is self-contextualizing and carries
"". The reference client translates segment and surround together, stops
generation at the blank line separating them, and keeps the segment's own
part of the output (its line resp. paragraph; a hard-break `␣␣\n` separator
works too). If the model merged them (no separator, or an empty first
part), it falls back to translating the segment alone. The surround fixes
context-free readings ("About" as "approximately" — with the opening it
becomes "Tietoa"/"Acerca de"; "here" as "就在这里" → the idiomatic
"点击这里") and, as a side effect, most stray trailing punctuation.
### Explicitly out of scope for phase 2
- The machine translation itself: the API above moves fragments in and out;
the translating is external. `scripts/translator.py` is the reference
client (Seed-X-PPO-7B only — its 28 languages are the ceiling).
- Garbage collection of orphaned chunks/translations (see docs/migrate.md).
- sitemap.xml per-language entries; translated UI chrome; per-language
typographer options; multi-locale date/number formatting.
+178
View File
@@ -0,0 +1,178 @@
# migrate_v3: content-addressed chunk storage
Status: **implemented**. `migrate_v3` restructures how article text and
translations are stored, motivated by the localization model in
`docs/localization.md` (phase 2). Since it is a full migration, it is free to
break the current `Node.content: str | None` layout.
## Goals
- **Minimal change diffs.** kanta persists change diffs; editing one
paragraph of a long article must not rewrite the whole article string, and
a translation refresh must touch only the re-translated chunks.
- **Fast, simple lookup.** Everything heavy lives in flat
`dict[hash, content]` stores; ordering lives in `list[hash]`. No large
nested structures, no deep paths.
- **Path-independent text.** Chunks and their translations are keyed by
content hash, not by article path — the same paragraph (or menu title)
appearing in several articles is stored and translated once. Moving or
renaming an article touches nothing.
## Design (chosen: global content-addressed stores)
Original articles are *also* stored as chunks; everything — originals and
translations — lives in flat hash-keyed dicts. Costs accepted: rendering does
one dict lookup per chunk (trivial), orphaned hashes need occasional garbage
collection, and the editor save path re-chunks server-side (it already
diffs). The rejected alternatives: per-article nested `LangVersion`
structures (churn, duplication, whole-string originals) and a hybrid with
whole originals plus global translations (keeps the worst change-diff
property).
## Target layout
```python
class Node(msgspec.Struct, omit_defaults=True):
...
#: Replaces `content: str | None`. None = pure category label;
#: a list (possibly empty) = a page, as ordered chunk hashes.
chunks: list[bytes] | None = None
#: Primary language of the article (BCP-47 base tag). "" = inherit
#: (nearest ancestor, front page last, site default "en" final).
language: str = ""
#: Chunk hashes the editor marked "do not translate" (always served
#: from the original). Presence-keys, value always True.
no_trans: dict[bytes, True] = {}
#: Languages this article is available in (besides its primary
#: language). Presence-keys, value always True — rendering, language
#: selection and hreflang alternates read this set instead of probing
#: the trans store chunk by chunk. Maintained by the writers (see
#: "Language index maintenance" below).
langs: dict[str, True] = {}
class Data(msgspec.Struct):
...
#: API keys gating the translator service WebSocket (/_translate/{key}):
#: key -> display name; the first is generated at bootstrap (state.py).
translate_keys: dict[str, str] = {}
#: Wanted target languages for the translator service (presence-keys);
#: jobs are offered only in these ∩ a connection's capabilities.
translate_langs: dict[str, True] = {}
#: All original-language text, content-addressed: blake3(normalized)
#: digest[:9] -> Markdown chunk. Shared by every article. Keys are
#: bytes; kanta/msgspec base64-encode them at the JSON level.
chunks: dict[bytes, str] = {}
#: Machine translations: chunk hash -> lang -> translated Markdown
#: (nested, not tuple keys: msgspec's JSON serializer rejects them).
#: Also used for node titles (hash of the title text).
trans: dict[bytes, dict[str, str]] = {}
#: User override patches per article and language:
#: f"{path}:{lang}" -> ordered patches (see localization.md).
patches: dict[str, list[Patch]] = {}
```
Notes:
- **Article paths never carry a leading slash** in the DB or in lookup keys
(`"docs/setup"`, front page `""`); the leading slash is added only when
building hrefs. `migrate_v3` audits existing stored paths (translation
keys, analytics references, any path-valued fields) and normalizes them.
- **Titles are chunks too**, by hash only: the nav renderer looks up
`trans.get(hash(node.title), {}).get(lang)`. No separate title storage;
editing a title invalidates its translations automatically.
- **Per-hunk options** live in two places: *inherent* options are derived at
chunking time (code fences, HTML blocks and prose-free chunks are
no-translate without storing anything — `needs_translation`, see
docs/localization.md "Masking"); *editor-set* flags are `node.no_trans`
(keyed by chunk
hash, so a heavy edit silently drops the flag — acceptable and
self-healing).
- **Patch payloads stay inline** in `Patch.hunks` — patches are small by
construction (minimal server-computed diffs). If a pathological case shows
up, hunks can be hash-stored later without schema pain.
## Language index maintenance (`node.langs`)
`node.langs` is a denormalized index over the `trans`/`patches` stores so
that article rendering, `select_language`'s availability check, and hreflang
alternate links never enumerate chunks. It is written by whoever writes
translation data, in the same transaction:
- **Translator service:** the WebSocket API at `/_translate/{key}` (see
docs/localization.md) offers pending fragments (titles + translatable
chunks lacking an entry for the language) as single-item jobs — one at
a time per connection, in `Data.translate_langs` ∩ the connection's
announced capabilities — and receives the matching result; storing it
writes the `trans[h][lang]` entry, sets `node.langs[lang] = True` on
every article that gained one and invalidates the page cache — all in
one transaction.
- **Translated-view save:** appending the first patch for `f"{path}:{lang}"`
sets `node.langs[lang] = True` (patches alone make the version exist).
- **Removals:** deleting a patch or GC'ing translations re-derives the key:
keep `lang` if any `trans` entry for the article's current chunks/title or
any patch remains, otherwise drop it. Stale `langs` keys are benign (an
advertised language that renders as the original), so removal can lag.
## Render / save pipeline (summary)
- **Render:** `text = "\n\n".join(chunks[h] for h in node.chunks)` for the
original; for language `L` (only ever attempted when `L in node.langs`),
per chunk `trans.get(h, {}).get(L)` unless missing or `h in node.no_trans`,
falling back to `chunks[h]`; then apply `patches.get(f"{path}:{L}", [])`
in order (per-hunk, best effort); then `markdown.render` as today. All of
this assembles the `Translation` the phase-1 plumbing already consumes.
- **Availability:** `node.langs` is the availability index; `?lang=`
handling uses exactly this set. (hreflang alternates are site-wide from
`translate_langs` instead — see docs/localization.md.)
- **Save (primary language):** server re-chunks the submitted Markdown,
inserts new hashes into `Data.chunks`, replaces `node.chunks`. Unchanged
chunks keep their hashes — only genuinely new text lands in the diff.
- **Save (translated view):** diff against the served hybrid, append a
`Patch` under `patches[f"{path}:{lang}"]`; `node.chunks` untouched.
- **Invalidate:** any write to `chunks` / `trans` / `patches` calls
`_invalidate_pages()`.
## migrate_v3 steps
1. Walk `menu`; for every node with a string `content`:
`chunks = chunk_markdown(content)`; write each into the new `chunks`
store; replace the field with the hash list (`None` stays `None`).
2. Initialize empty `chunks` / `trans` / `patches` stores.
3. Normalize stored paths: strip leading slashes anywhere paths are keys or
values.
4. `language`, `no_trans` and `langs` need nothing — struct defaults cover
them (`langs` starts empty; the translator job fills it as translations
land).
Chunking must be deterministic and shared with render/save, so
`chunk_markdown` + `chunk_key` live in `pagerite/i18n.py` (or a small
`pagerite/chunks.py`) and are imported by both `migrations.py` and
`views.py`/`state.py`.
## Implementation notes (deviations from the plan above)
- Chunking lives in `pagerite/chunks.py`; hashing uses the `blake3` package
(already a dependency), truncated to a 9-byte `bytes` digest (kanta's
JSON persistence base64-encodes bytes keys to 12-char strings).
- `trans` is keyed `hash -> lang -> text` (nested dict), not by
`f"{hash}:{lang}"` tuples: msgspec's JSON serializer only supports
str-like/number-like dict keys, and kanta persists as JSON lines.
- `Translation.titles` stayed keyed by node path (phase-1 shape, views
untouched): `get_translation` builds it by walking the menu with the same
per-title `trans.get(chunk_key(node.title), {}).get(lang)` lookups.
- Insert hunks anchor on the whole preceding block (not just its tail) —
a stronger, simpler search context.
- `make_patch` diffs with `SequenceMatcher(autojunk=False)` so patches are
deterministic (popular lines like blank separators never become junk).
- Step 3's path normalization is a no-op in practice: the only path-keyed
store (`patches`) starts empty at v3; analytics paths live outside the
kantadb. The code still strips leading slashes defensively.
## Garbage collection (later, manual or idle-time)
Orphaned entries accumulate: chunks no longer referenced by any
`node.chunks`/`node.title`, translations whose chunk hash is orphaned, patch
hunks that never match. All are harmless (never read). A GC pass is a single
tree walk collecting live hashes, then deleting the rest from `chunks` and
`trans`; patches whose every hunk is stale get pruned. Not part of
migrate_v3.
+5
View File
@@ -0,0 +1,5 @@
# Pagerite overview
Pagerite is a single-user CMS/blog. FastAPI serves HTML rendered in Python with html5tagger; content is persisted in a kanta database and rendered on the fly per request. Vue is used only for interactive bits (editing tools), not for the public pages.
See `docs/design-principles.md` for the high-level design and the other `docs/*.md` files for implementation details.
Binary file not shown.

After

Width:  |  Height:  |  Size: 56 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 104 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 74 KiB

+91
View File
@@ -0,0 +1,91 @@
# Production setup
From a local demo to a real site: run Pagerite as a systemd service behind
a reverse proxy that terminates HTTPS, with Paskia guarding the editing API.
The moving parts:
- **Pagerite** — serves the public site on `localhost:8100` and the editing
API under `/_api`.
- **Paskia** — the SSO server; owns `/auth/` and answers forward-auth
subrequests.
- **A reverse proxy** — Caddy below, but nginx or anything with
forward-auth support works the same way.
## Pagerite as a systemd service
Install [uv](https://docs.astral.sh/uv/getting-started/installation/) on the
system, create a user, and add a template unit:
```sh
sudo useradd --system --home-dir /srv/pagerite --create-home pagerite
curl -LsSf https://astral.sh/uv/install.sh | sudo env UV_INSTALL_DIR=/usr/local/bin sh
sudo systemctl edit --force --full pagerite.service
```
```ini
[Unit]
Description=Pagerite CMS
[Service]
Type=simple
User=pagerite
SyslogIdentifier=pagerite
WorkingDirectory=/srv/pagerite
ExecStart=uvx pagerite example.com --dbip
[Install]
WantedBy=multi-user.target
```
Replace `example.com` with your actual domain name. `--dbip` keeps the local GeoIP database up to date: leave out if you don't want DBIP data for analytics.
```sh
sudo systemctl enable --now pagerite
sudo journalctl -ocat -fu pagerite
```
## Running it on internet
We recommend Caddy for making your site publicly visible on the Internet. Presumably you already have some proxy, perhaps Nginx, but our setup is not much different of any other service you might already be running. ChatGPT and the likes can also help with the configuration because online documentation is limited. Note that Paskia also has extensive documentation on [running on various proxy servers](https://git.zi.fi/LeoVasanko/paskia/src/branch/main/docs/proxy/index.md)
Install [Caddy](https://caddyserver.com/) and follow the [Paskia setup guide](https://git.zi.fi/LeoVasanko/paskia) to get the SSO server running and its `auth` snippets copied to `/etc/caddy/auth` — that guide covers Paskia's own configuration and admin registration in detail.
Then the site config. Only the editing API needs gating; the site itself is public:
```caddyfile
example.com {
import auth/setup
reverse_proxy /auth/* localhost:4401
@api path /_api/*
handle @api {
import auth/require perm=pagerite:admin
reverse_proxy localhost:8100
}
handle {
reverse_proxy localhost:8100
}
}
```
Reload Caddy, then create a permission with scope `pagerite:admin` in the
Paskia admin panel (`/auth/admin/`) and assign it to yourself, as the Paskia
guide describes. Anonymous visitors now get 401 from `/_api`, logged-in
users without the permission get 403, and admins get the editing pens.
## nginx or another proxy
The shape is identical everywhere:
- `/auth/` proxies to Paskia (`localhost:4401`).
- `/_api` requires a forward-auth subrequest against Paskia — on nginx that
is `auth_request` against Paskia's verify endpoint — before proxying to
Pagerite (`localhost:8100`).
- Everything else proxies straight to Pagerite.
Paskia ships per-proxy forward-auth guides covering
[Caddy, nginx and others](https://git.zi.fi/LeoVasanko/paskia/src/branch/main/docs/proxy/index.md);
adapt the matcher to `/auth/` and `/_api` as above and leave the rest public.
+71
View File
@@ -0,0 +1,71 @@
# Themes and assets
## Built assets
Files under `frontend/src/assets/` are built by Vite and served hashed under `/_assets/`:
- `pagerite.css` — base layout + conservative variables.
- `pygments.css` — Pygments token styles mapped onto the `--code-*` variables.
- `fonts/` — self-hosted variable woff2 files for Source Sans 3, Source Serif 4, Fraunces, Literata, Cormorant, Playfair Display, Inter, Montserrat, Fira Code, Cause, Exo 2 and New Rocker.
The `::view-transition*` rules are not in the base stylesheet: they live in the page-transition designs (`pagerite/themes/{name}/transition.css`, see below).
## Themes
Themes are folders in `pagerite/themes/{name}/` containing `theme.css` and/or `banner.css` (+ `banner.svg` artwork and any extra assets the CSS references, like summer's `grass.svg`). They are served by the backend at `/_themes/{name}/...` — read from disk per request (etag by mtime), never built, so on-disk edits show on the next page load even in prod.
Themes are searched across several roots, most specific first (see `views.THEME_DIRS`):
1. `themes/` under the current working directory
2. `<site-dir>/themes/` (the site's own folder, e.g. `localhost/themes/`)
3. The platform user data dir's `pagerite/themes/` (Linux: `~/.local/share/pagerite/themes/`, Windows: `%LOCALAPPDATA%\pagerite\themes\`)
4. The platform system data dirs' `pagerite/themes/` (Linux: `/usr/local/share/pagerite/themes/`, `/usr/share/pagerite/themes/`, ...; Windows: `%PROGRAMDATA%\pagerite\themes\`) — several combine
5. `pagerite/themes/` (built-in package dir, fallback)
The data dirs come from `platformdirs` (`views._data_roots`).
All roots combine: listings are the union of folder names, and each file resolves from the first root that has it. So users can add completely new themes in any root, shadow a built-in file with their own (`<site>/themes/corporate/theme.css`), or extend a built-in theme with extra files (any files the user folder doesn't provide still come from the built-in). Since everything is read per request, new or changed folders take effect without a server restart.
`Data.theme` selects the active theme (empty = none/base only) and the site editor can switch it, choosing from the theme folders found on disk. Vue may add per-component styles on top where needed.
The site editor shows a light/dark-mode indicator in front of each theme name, read from the theme's `color-scheme` declaration in `theme.css`: ☀️ for light-only, 🌙 for dark-only, and 🌓 for themes that support both. The base theme (`none`) is light-only.
Current themes:
- `purple` — dark dusk palette with Fraunces/Literata and a tilted oversized gradient brand.
- `corporate` — light-first with automatic `prefers-color-scheme` dark mode, Montserrat/Inter and a huge solid brand.
- `nitro` — racing/HUD style following `prefers-color-scheme` (warm light-grey page, deep violet in dark), Montserrat/Literata, black as an accent only, a straight orange blade under the banner, and an orange racing-tab nav clipped with a bezier `shape()`.
- `summer` — light playful meadow, one palette sampled from its illustrated `banner.svg` (sky/grass/sun/flower pink), Fraunces/Literata, a tilted gradient brand, flower bullets, and a layered-parallax banner (sun rises, clouds drift, nearer hills move less) with idle animations (swaying flowers, floating clouds, breathing sun glow) wrapped in `prefers-reduced-motion: no-preference`.
## User fonts
Fonts are shared across themes, so user fonts live in `fonts/` folders next to the theme roots (same list minus the built-in fallback: `views.FONT_DIRS`, e.g. `localhost/fonts/` for the site; built-in fonts ship with the Vite build). A font is a folder `fonts/{name}/` with:
- `font.css``@font-face` rules with URLs relative to the folder (files served at `/_fonts/{name}/...`, per request like themes), plus a `:root { --font-{name}: "Family Name", serif; }` stack variable so themes and custom CSS reference it like the built-in `--font-*` variables.
- the font files the CSS references (e.g. `{name}.woff2`).
Every font.css is linked on all pages (after the base stylesheet, before the theme). The site editor's font picker lists user fonts too, with label and serif/sans grouping parsed from the `--font-{name}` stack. Everything is read per request, so new fonts appear without a restart.
## Banner designs
A theme folder may also ship a banner design (`banner.css` + `banner.html` arbitrary markup or `banner.svg`), selectable per page independently of the active theme. Standalone banner designs (no theme.css) ship as:
- `eyes` — a canvas critter in the grass.
- `stars` — a drifting starfield.
The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll parameter on `<html>` (event-driven, so it is still when the page is idle), the banner contents drift within their window (with scale overscan so no edge shows), and designs may key their own effects off the same parameter.
## Page transitions
A theme folder may ship a page transition (`transition.css`, `::view-transition*` rules), selected site-wide by `Data.transition` in the site settings and injected as `#pagerite-transition` (after the banner design). Standalone transition designs ship as:
- `cube` — rotating cube (from termotohtori.fi; the block is fragile — do not tweak), mirrored on history-back (`html.nav-back`), crossfading within a section (`html.nav-fade`).
- `slide` — plain sideways slide, old and new pages moving together; mirrored on history-back, crossfading within a section.
- `reveal` — clip-path wipe revealing the new page over the stationary old one; mirrored on history-back, crossfading within a section.
- `crossfade` — plain crossfade for all navigations.
pagerite.js toggles the `nav-back`/`nav-fade` classes on `<html>` around `document.startViewTransition` (skipped under `prefers-reduced-motion`); a transition design keys its `::view-transition*` rules off them as needed.
## Stylesheet order
The backend emits the stylesheets in a fixed order — base (Vite build), theme, banner design, page transition, entry sheets, custom CSS last — each with a stable id so fetch-navigation and the site editor can sync them in place. In dev they are `<link>`s (the base is Vite-injected from JS instead); in production they are inlined as `<style>` elements. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`. Code text (Fira Code by default) is optically matched to the body font by x-height: `font-size-adjust: ex-height var(--code-x-height)` scales whatever code font is in use, so a theme that switches its body font sets `--code-x-height` to that font's x-height ratio (base: 0.478 for Source Sans 3; themes ship values for Inter, Montserrat, Literata and Cause).
+3
View File
@@ -37,3 +37,6 @@ __screenshots__/
# Playwright browser downloads (if ever installed locally) # Playwright browser downloads (if ever installed locally)
.pw-browsers/ .pw-browsers/
# npm project config (audit/fund off: the audit endpoint stalls installs)
!.npmrc
+2
View File
@@ -0,0 +1,2 @@
audit=false
fund=false
+3
View File
@@ -18,6 +18,9 @@
"@codemirror/view": "^6.43.8", "@codemirror/view": "^6.43.8",
"@lezer/highlight": "^1.2.3", "@lezer/highlight": "^1.2.3",
"codemirror": "^6.0.2", "codemirror": "^6.0.2",
"country-flag-icons": "^1.6.20",
"overlayscrollbars": "^2.16.0",
"pinia": "^4.0.3",
"transliteration": "^2.6.1", "transliteration": "^2.6.1",
"vue": "^3.5.26", "vue": "^3.5.26",
"vuedraggable": "^4.1.0" "vuedraggable": "^4.1.0"
Binary file not shown.

Before

Width:  |  Height:  |  Size: 4.2 KiB

+538
View File
@@ -0,0 +1,538 @@
<script setup>
// Analytics viewer rendered as a normal page inside #main. Receives live
// analytics data over /_api/ws/analytics (admin-gated by the auth proxy) and
// renders totals, smoothed visit/views curves, a transition map, and recent
// visit/crawler tables. Read-only.
// See docs/analytics.md for the data format.
import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
import {
RANGES,
rangeWindow,
filterRecordsByRange,
filterTransitionsByRange,
filterViewsByRange,
} from './analytics/time.js'
import {
calcReadStats,
calcTotalViews,
copyIp,
copyList,
formatCount,
formatAbuseRows,
formatCrawlerRows,
formatVisitRows,
} from './analytics/format.js'
import TrailLink from './TrailLink.vue'
import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue'
import { VIEW_W } from './analytics/chart.js'
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
import ConnNote from './ConnNote.vue'
// Same centering margin as the charts, so the totals row's left edge
// aligns with the chart svg above the natural width.
const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const ABUSE_MAX_LINES = 5
const data = ref(null)
const pageTree = ref(null)
const error = ref('')
const now = ref(Date.now())
let ws = null
let reconnectTimeout = null
let connectWatchdog = null
const reconnects = reconnectPolicy()
let timeInterval = null
// The panel is live data over its socket: while it is connecting or waiting
// to reconnect, say so (ConnNote) instead of showing a silent stale view.
const conn = ref('connecting') // connecting | open | waiting
const retryIn = ref(0)
const connNote = computed(() =>
conn.value === 'connecting' ? 'connecting to the server…'
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
: '',
)
// The initial range comes from the URL hash (shareable links); without one,
// it is derived from the first analytics snapshot: day when the recorded
// history is shorter than 24 h, week otherwise.
const hashRange = location.hash.slice(1)
const range = ref(RANGES[hashRange] ? hashRange : 'week')
let rangePinned = Boolean(RANGES[hashRange])
function connectAnalytics() {
if (ws) return
conn.value = 'connecting'
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
clearTimeout(connectWatchdog)
connectWatchdog = watchConnecting(ws, 'analytics')
ws.onopen = () => {
conn.value = 'open'
reconnects.opened()
error.value = ''
}
ws.onmessage = (event) => {
try {
data.value = JSON.parse(event.data)
if (!rangePinned) {
rangePinned = true
const starts = (data.value?.visits || [])
.map((v) => Date.parse(v.start))
.filter((t) => !Number.isNaN(t))
if (starts.length && Date.now() - Math.min(...starts) < 24 * 3600 * 1000) {
range.value = 'day'
}
}
} catch {
error.value = 'analytics data could not be loaded'
}
}
ws.onerror = () => {
error.value = 'analytics data could not be loaded'
}
ws.onclose = () => {
ws = null
// The policy paces the retry: doubling backoff with jitter, reset only
// by a healthy connection — a fixed rapid loop trips the browser's
// WebSocket throttling (all sockets then sit "pending" for minutes).
const wait = reconnects.closed()
retryIn.value = Math.max(1, Math.round(wait / 1000))
conn.value = 'waiting'
reconnectTimeout = setTimeout(connectAnalytics, wait)
}
}
onMounted(async () => {
// The first connection takes a staggered slot (see ./reconnect).
reconnectTimeout = setTimeout(connectAnalytics, socketSlot())
now.value = Date.now()
timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
// The site tree for the transition map (all pages in menu order). Not
// fatal: without it the map just narrows to pages seen in transitions.
try {
const res = await fetch('/_api/pages')
if (res.ok) pageTree.value = await res.json()
} catch { /* map just narrows to pages seen in transitions */ }
})
onUnmounted(() => {
if (reconnectTimeout) clearTimeout(reconnectTimeout)
if (connectWatchdog) clearTimeout(connectWatchdog)
if (timeInterval) clearInterval(timeInterval)
if (ws) {
ws.onclose = null
ws.close()
ws = null
}
})
const window = computed(() => rangeWindow(range.value))
// All non-chart stats follow the selected range; the charts keep their own
// range-specific x windows (week overlays previous weeks aligned to Monday).
const rangeData = computed(() => {
if (!data.value) return null
const { t0, t1 } = window.value
return {
...data.value,
transitions: filterTransitionsByRange(data.value.transitions, t0, t1),
views: filterViewsByRange(data.value.views, t0, t1),
visits: filterRecordsByRange(data.value.visits, t0, t1),
crawlers: filterRecordsByRange(data.value.crawlers, t0, t1),
abuse: filterRecordsByRange(data.value.abuse, t0, t1),
}
})
const visits = computed(() => rangeData.value?.visits || [])
const totalViews = computed(() => calcTotalViews(rangeData.value?.views))
const readStats = computed(() => calcReadStats(visits.value))
// Keep the URL shareable when the range changes.
watch(range, (r) => {
const url = new URL(location.href)
url.hash = r
history.replaceState(history.state, '', url)
})
const clients = computed(() => data.value?.clients || {})
const favicons = computed(() => data.value?.favicons || {})
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
const crawlers = computed(() => rangeData.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, pageTree.value, now.value))
</script>
<template>
<!-- Untranslated admin dashboard: always LTR, like the editor panel. -->
<div class="analytics-view" lang="en" dir="ltr">
<div class="analytics-panel">
<header>
<h1>Analytics</h1>
<nav class="ranges">
<button v-for="(r, key) in RANGES" :key="key" type="button"
:class="{ active: range === key }" @click="range = key">
{{ r.label }}
</button>
</nav>
<a href="/" class="close" title="home"></a>
</header>
<ConnNote :text="connNote" />
<p v-if="error" class="error"> {{ error }}</p>
<p v-else-if="!data" class="loading">loading</p>
<template v-else>
<section class="totals" :style="{ marginLeft: CHART_MARGIN }">
<div><strong :title="String(visits.length)">{{ formatCount(visits.length) }}</strong> visits</div>
<div><strong :title="String(totalViews)">{{ formatCount(totalViews) }}</strong> page views</div>
<div><strong>{{ readStats.avgMinPerVisit }}</strong> min/visit</div>
<div><strong>{{ readStats.avgArticleMedianMin }}</strong> min/read</div>
</section>
<VisitorCharts :data="data" :range="range" />
<TransitionGraph :data="rangeData" :window="window" :page-tree="pageTree" :favicons="favicons" />
<section>
<h2>Recent visits</h2>
<div v-if="visitRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>trail</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(v, i) in visitRows" :key="i">
<td class="trail">
<TrailLink v-if="v.refererStep" :step="v.refererStep" :favicons="favicons" @close="$emit('close')" />
<span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span>
<TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" :favicons="favicons" @close="$emit('close')" />
</td>
<VisitorCell
:ip="v.ip"
:ip-display="v.ipDisplay"
:ua="v.ua"
:ua-raw="v.uaRaw"
:country="v.country"
:city="v.city"
:lang="v.lang"
:lang-display="v.langDisplay"
:is-host="v.isHost"
/>
<td class="last-seen muted"
:title="v.lastSeenLocal"
@click="copyList(v.lastSeenIso, $event)">{{ v.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
<p v-else class="empty">no visits recorded yet</p>
<div v-if="crawlerRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>pages crawled</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(c, i) in crawlerRows" :key="i">
<td class="trail">
<TrailLink v-if="c.refererStep" :step="c.refererStep" :favicons="favicons" @close="$emit('close')" />
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
</td>
<VisitorCell
:ip="c.ip"
:ip-display="c.ipDisplay"
:ua="c.ua"
:ua-raw="c.uaRaw"
:country="c.country"
:city="c.city"
:lang="c.lang"
:lang-display="c.langDisplay"
:is-host="c.isHost"
/>
<td class="last-seen muted"
:title="c.lastSeenLocal"
@click="copyList(c.lastSeenIso, $event)">{{ c.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
<div v-if="abuseRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>paths abused</th>
<th>articles read</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(a, i) in abuseRows" :key="i">
<td class="trail abuse-list clickable-list"
@click="copyList(a.allPaths, $event)">
<div class="abuse-items">
<span v-for="(p, pi) in a.paths.slice(0, ABUSE_MAX_LINES)" :key="pi"
class="inline-item">
<small v-if="p.count > 1" class="muted">{{ formatCount(p.count) }}×</small>{{ p.path }}
</span>
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
</div>
</td>
<td class="trail clickable-list"
@click="copyList(a.allArticles, $event)">
<TrailLink v-for="(s, si) in a.articles" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
<small v-if="!a.articles.length" class="muted"></small>
</td>
<VisitorCell
:ip="a.ip"
:ip-display="a.ipDisplay"
:ua="a.ua"
:ua-raw="a.uaRaw"
:country="a.country"
:city="a.city"
:lang="a.lang"
:lang-display="a.langDisplay"
:is-host="a.isHost"
:variant-count="a.clientCount"
/>
<td class="last-seen muted"
:title="a.lastSeenLocal"
@click="copyList(a.lastSeenIso, $event)">{{ a.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
</section>
</template>
</div>
</div>
</template>
<style scoped>
.analytics-view {
min-height: 100vh;
background: var(--bg, Canvas);
color: var(--text, CanvasText);
}
.analytics-panel {
margin: 0;
width: 100%;
/* Same 1.25rem side spacing as main's article padding. */
padding: 1.5rem 1.25rem 4rem;
/* Container for cqw-based shrink-to-fit (see .totals). */
container-type: inline-size;
}
.analytics-panel header {
display: flex;
align-items: center;
gap: 1rem;
}
.analytics-panel h1 {
margin: 0;
font-size: 1.4rem;
}
.ranges {
display: flex;
gap: 0.25rem;
margin-left: auto;
}
.ranges button {
padding: 0.2rem 0.7rem;
font: inherit;
font-size: 0.9rem;
color: var(--muted);
background: none;
border: 1px solid var(--line);
border-radius: 1rem;
cursor: pointer;
}
.ranges button:hover { color: var(--text); }
.ranges button.active {
color: var(--text);
border-color: var(--accent);
}
.close {
padding: 0 0.3rem;
background: none;
border: none;
color: var(--muted);
font-size: 1.2rem;
cursor: pointer;
}
.close:hover { color: var(--text); }
.analytics-panel h2 {
margin: 0 0 0.6rem;
font-size: 1rem;
color: var(--muted);
}
.analytics-panel section {
margin-top: 1.8rem;
}
.analytics-view a {
color: var(--text);
text-decoration: none;
}
.analytics-view a:hover { color: var(--accent); }
.analytics-view :deep(.muted) { color: var(--muted); }
.analytics-view :deep(.small) { font-size: 0.75em; }
/* One line at any width: the gap shrinks first, then the font (the number
scales along in em), both following the panel's container width. */
.totals {
display: flex;
gap: clamp(0.5rem, 3cqw, 2rem);
font-size: clamp(0.6rem, 2.2cqw, 1.1rem);
white-space: nowrap;
}
.totals strong { font-size: 1.36em; }
.visit-table-wrap {
overflow-x: auto;
}
.visit-table {
width: 100%;
border-collapse: collapse;
font-size: 0.9rem;
line-height: 1.3;
}
.visit-table th,
.visit-table td {
padding: 0.25rem 0.5rem;
border-bottom: 1px solid var(--line);
text-align: left;
vertical-align: top;
}
.visit-table th {
color: var(--muted);
font-weight: normal;
text-transform: lowercase;
position: sticky;
top: 0;
background: var(--bg, Canvas);
}
.visit-table .last-seen {
width: 6rem;
text-align: right;
white-space: nowrap;
cursor: pointer;
}
.visit-table .trail {
max-width: 20rem;
overflow-wrap: break-word;
}
.visit-table .trail a,
.visit-table .trail-link {
display: inline-block;
max-width: 8rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .trail > * + * {
margin-left: 0.5rem;
}
.analytics-view :deep(.trail-link.error),
.analytics-view :deep(.trail-link.error:hover) {
color: var(--error, #c00);
}
.visit-table .utm-tag {
display: inline-block;
max-width: 100%;
padding: 0.05rem 0.4rem;
border: 1px solid var(--line);
border-radius: 0.25rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .clickable-list {
cursor: pointer;
max-width: 22rem;
}
.visit-table .abuse-items {
display: flex;
flex-wrap: wrap;
gap: 0.15rem 0.5rem;
align-items: baseline;
}
.visit-table .inline-item {
max-width: 18rem;
min-width: 0;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
word-break: keep-all;
hyphens: none;
}
.visit-table :deep(.clickable-ip),
.visit-table .clickable-list,
.visit-table .last-seen {
cursor: pointer;
position: relative;
}
.visit-table :deep(.copy-popup) {
position: absolute;
bottom: calc(100% + 0.25rem);
left: 50%;
transform: translateX(-50%);
padding: 0.15rem 0.4rem;
background: var(--text, CanvasText);
color: var(--bg, Canvas);
border-radius: 0.25rem;
font-size: 0.75rem;
white-space: nowrap;
pointer-events: none;
z-index: 10;
}
.crawler-top-uas {
font-size: 0.9rem;
margin-bottom: 0.6rem;
}
.crawler-top-uas strong {
color: var(--muted);
}
.empty, .loading, .error { color: var(--muted); }
.error { color: var(--error, #c00); }
</style>
+458
View File
@@ -0,0 +1,458 @@
<script setup>
// Banner editor tab: per-page banner HTML and banner design, previewed into
// the real #page-banner region. Close and tab switching live in EditorShell.
import { computed, onActivated, onMounted, onUnmounted, ref, watch } from 'vue'
import { EditorView, basicSetup } from 'codemirror'
import { Compartment, EditorState } from '@codemirror/state'
import { keymap } from '@codemirror/view'
import { indentWithTab } from '@codemirror/commands'
import { html } from '@codemirror/lang-html'
import { cmHighlight, cmTheme } from './cmtheme'
import ConnNote from './ConnNote.vue'
import { reconnectPolicy, socketSlot, watchConnecting } from './reconnect'
import { dropPageCache, loadPlain, runScripts } from './swapdoc'
const props = defineProps({
pagePath: { type: String, default: '' },
})
// close/path-change are wired by EditorShell; this tab never emits them.
defineEmits(['close', 'pathChange'])
const path = ref('')
const banner = ref('')
const saveError = ref('')
const fileInput = ref(null)
const bannerEl = ref(null)
let ws = null
let pendingSave = null
let reconnectTimer = null
let connectWatchdog = null
// Reconnection pacing lives in ./reconnect (shared with the other sockets).
const reconnects = reconnectPolicy()
let everConnected = false
// Connection state drives the note at the top (ConnNote), and locks input
// until the banner's document has arrived (typing before it would be
// clobbered by the doc accept).
const conn = ref('connecting') // connecting | open | waiting
const retryIn = ref(0)
const docReady = ref(false)
const editable = new Compartment()
const connNote = computed(() =>
conn.value === 'connecting' ? 'connecting to the server…'
: conn.value === 'waiting' ? `connection lost — reconnecting in ~${retryIn.value} s…`
: docReady.value ? '' : 'loading the banner…',
)
let view = null // CodeMirror for the banner HTML
let syncing = false // set while replacing the document programmatically
// Banner design selector options come from the backend via settings.
const theme = ref('')
const bannerDesign = ref(null)
const bannerDesignFrom = ref(null)
const bannerDesigns = ref([])
// Where an empty banner code field falls back to (null = nowhere: the
// banner is just the design artwork), shown as a caption under the field.
const bannerFrom = ref(null)
// The design that "inherit" resolves to and where it comes from
// (bannerDesignFrom: an ancestor path, "" = the front page, null = the
// active theme's default).
const bannerDesignInherited = ref('')
// One-shot callback run on the next save ack (set by onBannerDesignChange,
// whose re-render must not race the save it triggers).
let refreshOnSave = null
// The inherit option names the design actually in effect and its source.
const inheritLabel = computed(() => {
if (bannerDesignFrom.value === null) {
return `— (from ${theme.value || 'none'})`
}
const where = bannerDesignFrom.value === ''
? 'the front page'
: `/${bannerDesignFrom.value}`
return `— (${bannerDesignInherited.value || 'none'} from ${where})`
})
function normPath(p) {
return p.trim().replace(/^\/+|\/+$/g, '')
}
function send(msg) {
if (ws && ws.readyState === WebSocket.OPEN) {
ws.send(JSON.stringify(msg))
} else {
ensureConnected()
}
}
function ensureConnected() {
if (!ws || ws.readyState === WebSocket.CLOSED || ws.readyState === WebSocket.CLOSING) {
connect()
}
}
const timers = {}
function debounce(key, fn, ms = 600) {
clearTimeout(timers[key])
timers[key] = setTimeout(fn, ms)
}
function save() {
const msg = { type: 'save', path: normPath(path.value), banner: banner.value }
pendingSave = msg
send(msg)
}
function openPath(p) {
path.value = p
// Lock input until the doc arrives (typing would be clobbered by it).
docReady.value = false
view?.dispatch({ effects: editable.reconfigure(EditorView.editable.of(false)) })
send({ type: 'open', path: p })
}
watch(() => props.pagePath, (p) => { openPath(normPath(p)) })
function updateWindowTitle() {
document.title = `banner: /${path.value} 🖊️`
}
onActivated(() => {
updateWindowTitle()
if (banner.value.trim()) previewBanner()
})
// The shell stays mounted while hidden: when it is re-shown with this tab
// active, restore the window title and banner preview.
function onEditorShown() {
if (document.body.dataset.editorMode !== 'banner') return
updateWindowTitle()
if (banner.value.trim()) previewBanner()
}
async function loadSettings() {
try {
const s = await (await fetch('/_api/settings')).json()
theme.value = s.theme || ''
bannerDesigns.value = s.banner_designs || []
} catch { /* keep default */ }
}
// Re-render the page from the server (after banner design changes), then
// re-overlay the page's own banner code if one is being edited.
async function rerender() {
if (await loadPlain(path.value)) {
if (banner.value.trim()) previewBanner()
}
}
// --- Banner design ---------------------------------------------------------
function onBannerDesignChange() {
const msg = {
type: 'save',
path: normPath(path.value),
banner_design: bannerDesign.value,
}
pendingSave = msg
send(msg)
refreshOnSave = rerender
}
// --- Banner editing ------------------------------------------------------
function setDocument(text) {
syncing = true
view.dispatch({ changes: { from: 0, to: view.state.doc.length, insert: text } })
syncing = false
banner.value = text
}
function previewBanner() {
const el = document.getElementById('page-banner')
if (!el) return
if (banner.value.trim()) {
// Own banner code supplements the design: the inlined design artwork
// (marked with [data-design]) is detached while the author code is
// swapped in (so runScripts never re-runs the design's own scripts),
// then put back first — author code stays last so its styles win.
const artwork = [...el.querySelectorAll('[data-design]')]
for (const a of artwork) a.remove()
el.innerHTML = banner.value
runScripts(el)
el.prepend(...artwork)
} else {
// No banner code of its own: the region must show the inherited/design
// banner — re-render from the server (an empty write here would wipe it).
loadPlain(path.value)
}
}
function onBannerInput() {
previewBanner()
debounce('banner-html', save, 400)
}
function stripBannerMedia(html) {
// A banner has one piece of media: uploading replaces earlier img/video
// tags instead of stacking them. (Other HTML, e.g. canvas+script, stays.)
const doc = new DOMParser().parseFromString(html, 'text/html')
for (const el of doc.querySelectorAll('img, video')) el.remove()
return doc.body.innerHTML.trim()
}
async function uploadBannerMedia(file) {
// Banner media goes to the shared content store, like article images.
if (!file || !/^(image|video)\//.test(file.type)) return
const name = file.name.replace(/[^\w.-]/g, '-')
const res = await fetch(`/_api/files/${encodeURIComponent(name)}`, { method: 'PUT', body: file })
if (!res.ok) return
const { path: stored } = await res.json()
const tag = file.type.startsWith('video/')
? `<video src="${stored}" autoplay muted loop playsinline></video>`
: `<img src="${stored}" alt="">`
const rest = stripBannerMedia(banner.value)
setDocument(rest ? `${tag}\n${rest}` : tag)
previewBanner()
save()
}
function onBannerPaste(ev) {
const file = [...(ev.clipboardData?.files || [])]
.find((f) => /^(image|video)\//.test(f.type))
if (file) {
ev.preventDefault()
uploadBannerMedia(file)
}
}
function onMessage(ev) {
const msg = JSON.parse(ev.data)
if (msg.type === 'doc' && msg.path === path.value) {
setDocument(msg.banner ?? '')
docReady.value = true
view.dispatch({ effects: editable.reconfigure(EditorView.editable.of(true)) })
bannerDesign.value = msg.banner_design ?? null
bannerDesignFrom.value = msg.banner_design_from ?? null
bannerDesignInherited.value = msg.banner_design_inherited ?? ''
bannerFrom.value = msg.banner_from ?? null
if (banner.value.trim()) previewBanner()
} else if (msg.type === 'saved') {
saveError.value = ''
pendingSave = null
// Banner HTML/design changes affect the rendered page; invalidate prefetches.
dropPageCache()
refreshOnSave?.()
refreshOnSave = null
} else if (msg.type === 'error') {
saveError.value = '⚠️ changes could not be saved'
}
}
function onKeydown(ev) {
if ((ev.ctrlKey || ev.metaKey) && ev.key === 's') {
ev.preventDefault()
save()
}
}
function connect() {
clearTimeout(reconnectTimer)
conn.value = 'connecting'
if (ws) {
// Replacing a stale socket: detach its handlers so its close is silent.
ws.onopen = ws.onmessage = ws.onclose = ws.onerror = null
if (ws.readyState !== WebSocket.CLOSED) ws.close()
}
ws = new WebSocket(
`${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/_api/ws/editor`,
)
ws.onmessage = onMessage
clearTimeout(connectWatchdog)
connectWatchdog = watchConnecting(ws, 'banner')
ws.onopen = () => {
conn.value = 'open'
reconnects.opened()
if (everConnected) {
if (pendingSave) send(pendingSave)
} else {
openPath(normPath(props.pagePath))
}
everConnected = true
}
ws.onclose = () => {
// The wait is the policy's: doubling backoff with jitter (./reconnect),
// reset only by a healthy connection — rapid retries trip the browser's
// WebSocket throttling (sockets stuck "pending" for minutes).
const wait = reconnects.closed()
retryIn.value = Math.max(1, Math.round(wait / 1000))
conn.value = 'waiting'
reconnectTimer = setTimeout(connect, wait)
}
}
onMounted(async () => {
// The first connection takes a staggered slot (see ./reconnect).
reconnectTimer = setTimeout(connect, socketSlot())
view = new EditorView({
state: EditorState.create({
doc: '',
extensions: [
basicSetup,
// Tab/Shift-Tab indent and dedent instead of moving focus.
keymap.of([indentWithTab]),
html(),
cmTheme,
cmHighlight,
EditorView.lineWrapping,
// Locked until the banner's document arrives (docReady/ConnNote).
editable.of(EditorView.editable.of(false)),
EditorView.updateListener.of((u) => {
if (u.docChanged && !syncing) {
banner.value = view.state.doc.toString()
onBannerInput()
}
}),
],
}),
parent: bannerEl.value,
})
addEventListener('keydown', onKeydown)
addEventListener('pagerite:editor-shown', onEditorShown)
await loadSettings()
})
onUnmounted(() => {
clearTimeout(reconnectTimer)
clearTimeout(connectWatchdog)
for (const t of Object.values(timers)) clearTimeout(t)
if (ws) {
ws.onclose = null // intentional close, no reconnect
ws.close()
}
view?.destroy()
removeEventListener('keydown', onKeydown)
removeEventListener('pagerite:editor-shown', onEditorShown)
})
</script>
<template>
<div class="banner-editor">
<div v-if="saveError">{{ saveError }}</div>
<ConnNote :text="connNote" />
<section class="block" @paste="onBannerPaste">
<div class="block-head">
<select
v-model="bannerDesign"
class="text-input design-select"
title="Banner design (artwork + its own styles)"
@change="onBannerDesignChange"
>
<option :value="null">{{ inheritLabel }}</option>
<option value="">none</option>
<option v-for="d in bannerDesigns" :key="d" :value="d">{{ d }}</option>
</select>
<button
type="button"
class="icon-btn"
title="upload banner image/video (replaces existing media) — pasting works too"
@click="fileInput.click()"
>🖼</button>
<input
ref="fileInput"
type="file"
accept="image/*,video/*"
hidden
@change="(ev) => { uploadBannerMedia(ev.target.files[0]); ev.target.value = '' }"
/>
</div>
<div v-if="bannerFrom !== null" class="note">
left empty, the banner code is inherited from /{{ bannerFrom }}
</div>
<div ref="bannerEl" class="banner-cm" />
</section>
</div>
</template>
<style scoped>
.banner-editor {
display: flex;
flex-direction: column;
}
.block {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.5rem 1rem;
background: var(--surface);
flex: 1;
min-height: 0;
}
.block-head {
display: flex;
align-items: center;
gap: 0.6rem;
}
.note {
color: var(--muted);
font-size: 0.8rem;
}
.block-head .icon-btn {
margin-left: auto;
padding: 0 0.2rem;
font-size: 1rem;
background: none;
border: none;
cursor: pointer;
opacity: 0.7;
}
.block-head .icon-btn:hover {
opacity: 1;
}
/* The banner design selector stays compact; the upload button is pushed
right by its auto margin. */
.design-select {
flex: 0 1 auto;
width: auto;
font-size: 0.85rem;
}
.text-input {
flex: 1;
min-width: 4rem;
font: inherit;
font-size: 0.9rem;
padding: 0.2rem 0.5rem;
background: var(--bg);
color: var(--text);
border: 1px solid var(--line);
border-radius: 4px;
}
/* CodeMirror window for the banner HTML; fills the tab and scrolls
internally. */
.banner-cm {
flex: 1;
min-height: 0;
border: 1px solid var(--line);
border-radius: 4px;
overflow: hidden;
}
.banner-cm :deep(.cm-editor) {
height: 100%;
font-size: 0.85rem;
}
.banner-cm :deep(.cm-scroller) {
overflow: auto;
}
.banner-cm :deep(.cm-gutters) {
display: none;
}
</style>
+21
View File
@@ -0,0 +1,21 @@
<script setup>
// Connection-state note for the WebSocket-backed panels (page/banner
// editors, analytics view): while the socket is connecting or waiting to
// reconnect the panel cannot load or save, and this says so. An empty
// text hides the note.
defineProps({ text: { type: String, default: '' } })
</script>
<template>
<div v-if="text" class="conn-note" role="status">{{ text }}</div>
</template>
<style scoped>
.conn-note {
padding: 0.2rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
color: var(--muted);
font-size: 0.8rem;
}
</style>
+222
View File
@@ -0,0 +1,222 @@
<script setup>
// Tabbed shell for the five admin editors. The individual pens are shorthands
// that open the shell on a given tab; once open, tabs switch instantly without
// closing the panel. Tabs are kept alive so switching preserves state.
import { onMounted, onUnmounted, provide, ref, watch } from 'vue'
import PageEditor from './PageEditor.vue'
import BannerEditor from './BannerEditor.vue'
import SiteEditor from './SiteEditor.vue'
import StructureEditor from './StructureEditor.vue'
import LocalizationEditor from './LocalizationEditor.vue'
import { editorLang, pagePrimary } from './editorLang'
import { loadPlain, setLangOverride } from './swapdoc'
const props = defineProps({
pagePath: { type: String, default: '' },
initialMode: { type: String, default: 'page' },
})
const emit = defineEmits(['close'])
const currentPath = ref(props.pagePath)
const activeMode = ref(props.initialMode)
// The shared language selection (./editorLang, v-modeled by the tabs'
// LangSelects) is linked to the whole-page language: while the shell is
// open it drives the page preview (overrides ?lang= / Accept-Language),
// and closing keeps the pick as the session language. The primary
// selection pins by the CURRENT PAGE's own primary language (pages may
// differ — Node.language is inherited down the tree).
let pinned = false
function pinPreviewLang() {
pinned = true
// '' pagePrimary = not yet learned: pin 'en', the server's final fallback
// (i18n.ORIGINAL_LANGUAGE).
setLangOverride(editorLang.value || pagePrimary.value || 'en')
loadPlain(currentPath.value)
}
// Opening the panel must not switch the page's language: adopt the
// session's chosen language (public selector / earlier pick) once, then
// pin. Runs only on (re)open — after that the selection is the user's.
function openShell() {
const session = window.__pageriteLang
if (!editorLang.value && session && session !== (pagePrimary.value || 'en'))
editorLang.value = session
pinPreviewLang()
}
function unpinPreviewLang() {
if (!pinned) return
pinned = false
setLangOverride(null)
loadPlain(currentPath.value)
}
watch(editorLang, () => { if (pinned) pinPreviewLang() })
// The page's primary may be (re)learned while pinned on it (doc accept,
// tree refresh, a language change on the row) — re-pin with the new code.
watch(pagePrimary, () => { if (pinned && !editorLang.value) pinPreviewLang() })
// Tab order: site-wide settings first (site, structure, localization), then
// — after a visual break — the per-page editors (article, banner).
const MODES = [
{ key: 'site', label: 'site', component: SiteEditor },
{ key: 'structure', label: 'structure', component: StructureEditor },
{ key: 'localization', label: 'lang', component: LocalizationEditor },
{ key: 'page', label: 'article', component: PageEditor, breakBefore: true },
{ key: 'banner', label: 'banner', component: BannerEditor },
]
function switchMode(mode) {
if (MODES.some((m) => m.key === mode)) activeMode.value = mode
}
function onPathChange(path) {
currentPath.value = path
}
function close() {
emit('close')
}
provide('editorShell', { switchMode })
watch(activeMode, (mode) => {
document.body.dataset.editorMode = mode
}, { immediate: true })
// A pen click while the shell is open (main.js) switches tabs, and also
// retargets the editors when the user fetch-navigated with the shell open.
function onSwitchEvent(ev) {
if (ev.detail?.path != null) currentPath.value = ev.detail.path
if (ev.detail?.mode) switchMode(ev.detail.mode)
}
// Closing the shell hides it but keeps it mounted (main.js); the tabs stay
// cached in KeepAlive the whole time, so no state is ever lost until a real
// page reload. On re-show each active tab re-applies its window title and
// preview via its own pagerite:editor-shown listener. No Escape-to-close:
// it fired too easily by accident (e.g. dismissing an editor popup).
onMounted(() => {
document.body.dataset.editorMode = activeMode.value
addEventListener('pagerite:switch-editor', onSwitchEvent)
addEventListener('pagerite:editor-shown', openShell)
addEventListener('pagerite:editor-hidden', unpinPreviewLang)
// The shell mounts visible (openEditor), so open immediately. The site
// default primary language comes from the settings — it only fills the
// unknown; the page/structure tabs refine pagePrimary per page as they
// learn it (their knowledge is strictly better).
openShell()
fetch('/_api/settings').then((r) => r.json()).then((s) => {
if (!pagePrimary.value) pagePrimary.value = s.primary_lang || 'en'
}).catch(() => { /* keep the fallback */ })
})
onUnmounted(() => {
removeEventListener('pagerite:switch-editor', onSwitchEvent)
removeEventListener('pagerite:editor-shown', openShell)
removeEventListener('pagerite:editor-hidden', unpinPreviewLang)
})
</script>
<template>
<div class="editor-root overlay" lang="en" dir="ltr">
<header class="editor-tabs">
<template v-for="m in MODES" :key="m.key">
<span v-if="m.breakBefore" class="tab-break" />
<button
type="button"
class="tab"
:class="{ active: activeMode === m.key }"
@click="switchMode(m.key)"
>
{{ m.label }}
</button>
</template>
<button type="button" class="close" title="close" @click="close"></button>
</header>
<div class="editor-tab-body">
<KeepAlive>
<component
:is="MODES.find((m) => m.key === activeMode).component"
:key="activeMode"
:page-path="currentPath"
@close="close"
@path-change="onPathChange"
/>
</KeepAlive>
</div>
</div>
</template>
<style scoped>
/* Docked-overlay positioning (sticky, height, slide-in) comes from the
global pagerite.css (.editor-root.overlay); the shell is a flex column so
the tab body fills what the tab bar leaves. */
.editor-root {
display: flex;
flex-direction: column;
}
.editor-tabs {
display: flex;
align-items: center;
gap: 0.25rem;
padding: 0.4rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
flex-shrink: 0;
}
.editor-tabs .tab {
padding: 0.25rem 0.8rem;
font: inherit;
font-size: 0.9rem;
color: var(--muted);
background: none;
border: none;
border-bottom: 2px solid transparent;
cursor: pointer;
}
.editor-tabs .tab:hover {
color: var(--text);
}
/* Visual break between the site-wide tabs and the per-page tabs. */
.editor-tabs .tab-break {
align-self: stretch;
margin: 0.2rem 0.5rem;
border-left: 1px solid var(--line);
}
.editor-tabs .tab.active {
color: var(--text);
border-bottom-color: var(--accent);
}
.editor-tabs .close {
margin-left: auto;
padding: 0 0.3rem;
background: none;
border: none;
color: var(--muted);
font-size: 1.05rem;
cursor: pointer;
}
.editor-tabs .close:hover {
color: var(--text);
}
.editor-tab-body {
flex: 1;
min-height: 0;
display: flex;
flex-direction: column;
}
/* The active tab component's root fills the body. */
.editor-tab-body > * {
flex: 1;
min-height: 0;
}
</style>
+168
View File
@@ -0,0 +1,168 @@
<script setup>
// The editor shell's one language selector (page + structure tabs): a small
// flag button opening a clean dropdown, v-modeled on the shared editorLang
// ('' = the primary language). The lang tab's flag grid is a different
// control (toggles, not a select) and stays as it is.
import { computed, nextTick, ref } from 'vue'
import { usePopup } from './dropdown'
const props = defineProps({
modelValue: { type: String, default: '' },
options: { type: Array, required: true }, // [{tag, code, name, flag, primary}]
title: { type: String, default: '' }, // toggle-button tooltip override
})
const emit = defineEmits(['update:modelValue'])
const open = ref(false)
const root = ref(null)
const toggleBtn = ref(null)
const pop = ref(null)
const popStyle = ref({})
// Closes on outside click / Escape (./dropdown), not on mouseleave.
usePopup(open, root)
const current = computed(
() => props.options.find((o) => o.tag === props.modelValue) ?? props.options[0],
)
function toggle() {
open.value = !open.value
if (open.value) {
// Position: fixed so the popup overflows the scrolling editor panel
// onto the page area instead of being clipped by it.
const r = toggleBtn.value.getBoundingClientRect()
popStyle.value = { top: `${r.bottom + 2}px`, left: `${r.left}px` }
// A toggle mounted near the right window edge (the public page
// selector sits top-right) opens the popup flush against that edge.
nextTick(() => {
const p = pop.value?.getBoundingClientRect()
if (p && p.right > innerWidth - 4) {
popStyle.value = {
...popStyle.value,
left: `${Math.max(4, innerWidth - 4 - p.width)}px`,
}
}
})
}
}
function select(tag) {
emit('update:modelValue', tag)
open.value = false
}
</script>
<template>
<span v-if="options.length > 1" ref="root" class="lang-select">
<button
ref="toggleBtn"
type="button"
class="lang-current"
:class="{ open }"
:title="title || (current
? `language: ${current.name}${current.primary ? ' (primary)' : ''}`
: '')"
@click="toggle"
><span v-if="current?.flag" class="flag" v-html="current.flag" /></button>
<span v-if="open" ref="pop" class="lang-pop" :style="popStyle">
<button
v-for="o in options"
:key="o.code"
type="button"
:class="{ active: o.tag === modelValue }"
:title="o.primary ? `${o.name} — the primary language` : `${o.name} — translation`"
@click="select(o.tag)"
><span v-if="o.flag" class="flag" v-html="o.flag" /> {{ o.name }}<small v-if="o.primary"> (primary)</small></button>
</span>
</span>
</template>
<style scoped>
.lang-select {
position: relative;
display: flex;
}
/* The closed state is just the small flag — no button chrome at all, on
hover either (it sits among borderless emoji-icon buttons); like them it
rests dimmed and brightens on hover. */
.lang-current {
display: flex;
align-items: center;
padding: 2px;
background: none;
border: none;
border-radius: 4px;
cursor: pointer;
opacity: 0.7;
}
.lang-current:hover,
.lang-current.open {
opacity: 1;
}
/* The dropdown matches the page's existing popups (.picker-pop look).
Fixed-positioned (anchored to the toggle's viewport rect on open) so it
is not clipped by the editor panel's scrolling overflow. */
.lang-pop {
position: fixed;
z-index: 20;
display: flex;
flex-direction: column;
align-items: stretch;
gap: 0.15rem;
padding: 0.3rem;
background: var(--bg);
border: 1px solid var(--line);
border-radius: 6px;
box-shadow: 0 4px 16px #0004;
white-space: nowrap;
}
.lang-pop button {
display: flex;
align-items: center;
gap: 0.4rem;
padding: 0.15rem 0.4rem;
font: inherit;
font-size: 0.9rem;
text-align: left;
color: var(--text);
background: none;
border: none;
border-radius: 4px;
cursor: pointer;
}
.lang-pop button:hover {
background: var(--surface);
}
.lang-pop button.active {
color: var(--accent);
}
.lang-pop small {
color: var(--muted);
}
/* em-sized so the chip matches the surrounding text/icon size in each
context; the hairline border delineates white-flagged countries (not
button chrome). */
.flag {
display: inline-flex;
width: 1.5em;
height: 1em;
flex: 0 0 auto;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+41
View File
@@ -0,0 +1,41 @@
<script setup>
// The public page's language selector: the editors' flag dropdown
// (LangSelect) as the first item of the banner's corner container, fed
// from the shared store (pagerite.js sets the page's hreflang alternates
// and served language per navigation). It binds the same store.lang the
// editor's dropdown binds, so both always show the same selection. A pick
// also dispatches pagerite:set-session-lang — pagerite.js swaps the page
// in place when the editor is closed (open, the editor reacts to the
// store and re-renders it).
import { computed } from 'vue'
import LangSelect from './LangSelect.vue'
import { flagFor, langName } from './langs'
import { useStore } from './store'
const store = useStore()
// The "(primary)" marker is admin-panel information; the public selector
// lists plain languages.
const options = computed(() =>
store.langAlternates.map((a) => ({
tag: a.tag,
code: a.tag,
name: langName(a.tag),
flag: flagFor(a.tag),
primary: false,
})),
)
const primaryTag = computed(() => store.langAlternates.find((a) => a.primary)?.tag ?? '')
// The explicit pick, else the served language (header-autodetected pages
// may have neither), else the primary.
const model = computed(() => store.lang || store.servedLang || primaryTag.value)
function go(tag) {
store.lang = tag === primaryTag.value ? '' : tag
dispatchEvent(new CustomEvent('pagerite:set-session-lang', { detail: { lang: tag } }))
}
</script>
<template>
<LangSelect :model-value="model" :options="options" @update:model-value="go" />
</template>
+410
View File
@@ -0,0 +1,410 @@
<script setup>
// Lang tab: the site-wide translation target languages (translate_langs)
// and the translator service keys (translate_keys) with their WebSocket
// URLs. ALL languages are listed, English included — a page whose primary
// language (Node.language, configured per row in the structure tab,
// inherited down the hierarchy) differs can be translated INTO any other.
// Flag clicks toggle and save immediately; the settings round-trip
// re-reads the payload, so this tab only ever changes translate_langs. The
// settings write's invalidation hook kicks the translation dispatcher. The
// refresh button drops all machine translations (user patches are kept),
// making the dispatcher re-translate everything. Translator keys are
// managed inline ( add, name edit, ✕ delete); new keys are generated
// here in the server's format and everything rides the settings
// round-trip. Clicking a key copies its full URL (following ws:// would
// fail).
import { computed, onActivated, onMounted, onUnmounted, ref } from 'vue'
import { LANG_GROUPS, TRANSLATABLE, flagFor, langName } from './langs'
import { copyList } from './analytics/format.js'
import { dropPageCache } from './swapdoc'
defineProps({ pagePath: { type: String, default: '' } })
// close/path-change are wired by EditorShell; this tab never emits them.
defineEmits(['close', 'pathChange'])
const saveError = ref('')
const selected = ref(new Set())
const keyUrls = ref([])
// Full WebSocket URL for a key. New keys are generated right here: 12
// lowercase alphanumerics, the server-side format (state._KEY_ALPHABET).
const wsUrl = (key) =>
`${location.origin.replace(/^http/, 'ws')}/_translate/${key}`
const KEY_ALPHABET = 'abcdefghijklmnopqrstuvwxyz0123456789'
const newKey = () =>
[...crypto.getRandomValues(new Uint8Array(12))]
.map((b) => KEY_ALPHABET[b % KEY_ALPHABET.length])
.join('')
// The toggleable targets: every translatable language, laid out in
// geographic/cultural groups (one row each) rather than alphabetized —
// related languages sit together (a node's own primary is excluded per
// article, server-side). Any code missing from LANG_GROUPS trails as an
// extra row.
const groups = computed(() => {
const tile = (code) => ({ code, name: langName(code), flag: flagFor(code) })
const rows = LANG_GROUPS.map((g) => g.filter((c) => c in TRANSLATABLE).map(tile))
const covered = new Set(LANG_GROUPS.flat())
const rest = Object.keys(TRANSLATABLE).filter((c) => !covered.has(c)).map(tile)
if (rest.length) rows.push(rest)
return rows.filter((r) => r.length)
})
function updateWindowTitle() {
document.title = 'lang 🖊️'
}
onActivated(updateWindowTitle)
// The shell stays mounted while hidden: when it is re-shown with this tab
// active, restore the window title.
function onEditorShown() {
if (document.body.dataset.editorMode === 'localization') updateWindowTitle()
}
onMounted(async () => {
addEventListener('pagerite:editor-shown', onEditorShown)
try {
const s = await (await fetch('/_api/settings')).json()
selected.value = new Set(s.translate_langs || [])
keyUrls.value = Object.entries(s.translate_keys || {})
.map(([key, name]) => ({ key, name, url: wsUrl(key) }))
} catch { /* keep defaults */ }
})
onUnmounted(() => removeEventListener('pagerite:editor-shown', onEditorShown))
async function toggle(code) {
const next = new Set(selected.value)
if (next.has(code)) next.delete(code)
else next.add(code)
selected.value = next
try {
const s = await (await fetch('/_api/settings')).json()
const res = await fetch('/_api/settings', {
method: 'PUT',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ ...s, translate_langs: [...next] }),
})
if (res.ok) {
saveError.value = ''
dropPageCache()
} else {
saveError.value = '⚠️ changes could not be saved'
}
} catch {
saveError.value = '⚠️ changes could not be saved'
}
}
// Delete all machine translations server-side; the dispatcher re-fills
// them (a connected translator starts getting jobs right away). User
// patches survive — they are edits, not machine output.
const refreshing = ref(false)
async function refresh() {
if (refreshing.value) return
refreshing.value = true
try {
const res = await fetch('/_api/translations', { method: 'DELETE' })
saveError.value = res.ok ? '' : '⚠️ translations could not be refreshed'
if (res.ok) dropPageCache()
} catch {
saveError.value = '⚠️ translations could not be refreshed'
} finally {
refreshing.value = false
}
}
// Key management rides the settings round-trip, like toggle() above:
// mutate keyUrls, then PUT the whole settings payload with the new
// translate_keys. adds a fresh unnamed key, names save on every
// keystroke (@input — spamming the server is fine), ✕ deletes without
// confirmation.
async function saveKeys() {
try {
const s = await (await fetch('/_api/settings')).json()
const res = await fetch('/_api/settings', {
method: 'PUT',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
...s,
translate_keys: Object.fromEntries(keyUrls.value.map((k) => [k.key, k.name])),
}),
})
saveError.value = res.ok ? '' : '⚠️ changes could not be saved'
} catch {
saveError.value = '⚠️ changes could not be saved'
}
}
function addKey() {
const key = newKey()
keyUrls.value.push({ key, name: '', url: wsUrl(key) })
saveKeys()
}
function removeKey(k) {
keyUrls.value = keyUrls.value.filter((x) => x.key !== k.key)
saveKeys()
}
</script>
<template>
<div class="localization-editor">
<div v-if="saveError">{{ saveError }}</div>
<section class="block">
<div class="block-head">
<span class="field-label">languages</span>
</div>
<div class="flags">
<div v-for="(row, ri) in groups" :key="ri" class="flag-row">
<button
v-for="o in row"
:key="o.code"
type="button"
class="flag-tile"
:class="{ selected: selected.has(o.code) }"
:title="`${o.name} (${o.code})`"
@click="toggle(o.code)"
>
<span class="flag" v-html="o.flag" />
</button>
</div>
</div>
</section>
<section class="block">
<div class="block-head">
<span class="field-label">Translator API</span>
</div>
<div v-for="k in keyUrls" :key="k.key" class="key-row">
<a
:href="k.url"
class="key-link"
title="click to copy the URL"
@click.prevent="copyList(k.url, $event)"
>{{ k.key }}</a>
<input
v-model="k.name"
type="text"
class="edit key-name"
title="display name"
@input="saveKeys()"
>
<button type="button" class="act del" title="delete key" @click="removeKey(k)"></button>
</div>
<div class="add-row">
<button type="button" class="add" title="new translator key" @click="addKey()"> API key</button>
</div>
<p><small class="muted">AI translator agents can connect with the API keys to do machine translations to your selected languages. Click the button below to delete all translations and start over. User edits are kept.</small></p>
<div class="refresh-row">
<button
type="button"
class="refresh-btn"
:disabled="refreshing"
@click="refresh"
>
{{ refreshing ? 'Reseting…' : 'Reset' }}
</button>
</div>
</section>
</div>
</template>
<style scoped>
.localization-editor {
overflow-y: auto;
background: var(--surface);
}
.block {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.5rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
}
.block-head {
display: flex;
align-items: center;
gap: 0.6rem;
}
.field-label {
color: var(--muted);
font-size: 0.85rem;
}
.muted {
color: var(--muted);
}
/* Flag grid: one geographic group per row. Deselected flags sit dimmed and
grayed; a click brings one to full color (selected = a translation
target) — the shading alone carries the state, no outline. */
.flags {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.2rem 0;
}
.flag-row {
display: flex;
flex-wrap: wrap;
gap: 0.5rem;
}
.flag-tile {
padding: 3px;
background: none;
border: 2px solid transparent;
border-radius: 5px;
cursor: pointer;
opacity: 0.4;
filter: grayscale(0.8);
transition: opacity 0.15s, filter 0.15s, border-color 0.15s;
}
.flag-tile:hover {
opacity: 0.8;
filter: none;
}
.flag-tile.selected {
opacity: 1;
filter: none;
}
/* Same flag chips as the PageEditor language picker / analytics cells. */
.flag {
display: inline-flex;
width: 18px;
height: 12px;
flex: 0 0 auto;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
}
.flag-tile .flag {
width: 36px;
height: 24px;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
.key-row {
display: flex;
align-items: baseline;
gap: 0.6rem;
}
/* Real links (handy for right-click/drag) showing just the key, but the
click copies the full URL instead of following — ws:// would fail to
navigate. Normal text color, not link-styled; position: relative
anchors the "Copied!" popup (analytics/format.js). */
.key-link {
position: relative;
color: var(--text);
font-family: var(--font-code);
user-select: all;
}
.refresh-row {
display: flex;
align-items: baseline;
gap: 0.6rem;
}
/* Name input / ✕ / follow the structure tab's conventions: inputs stay
borderless until interacted with, glyph buttons redden / solidify on
hover. */
.key-name {
flex: 0 0 9rem;
}
.edit {
font: inherit;
font-size: 0.85rem;
padding: 0.1rem 0.4rem;
background: transparent;
color: var(--text);
border: 1px solid transparent;
border-radius: 4px;
min-width: 0;
}
.edit:hover {
border-color: var(--line);
}
.edit:focus {
background: var(--bg);
border-color: var(--accent);
outline: none;
}
.act {
padding: 0 0.25rem;
background: none;
border: none;
color: var(--muted);
font-size: 0.8rem;
cursor: pointer;
white-space: nowrap;
}
.del:hover {
color: #e06c75;
}
.add-row {
display: flex;
align-items: center;
}
.add {
padding: 0 0.3rem;
background: none;
border: none;
font-size: 0.9rem;
cursor: pointer;
opacity: 0.5;
}
.add:hover {
opacity: 1;
}
.refresh-btn {
align-self: flex-start;
margin-bottom: 0.2rem;
padding: 0.3rem 0.8rem;
font: inherit;
font-size: 0.85rem;
color: var(--muted);
background: none;
border: 1px solid var(--line);
border-radius: 5px;
cursor: pointer;
}
.refresh-btn:hover:not(:disabled) {
color: var(--text);
border-color: var(--muted);
}
.refresh-btn:disabled {
opacity: 0.5;
cursor: default;
}
</style>
+1172 -104
View File
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+402
View File
@@ -0,0 +1,402 @@
<script setup>
// Structure tab: the draggable site structure tree. Focusing a page's row
// navigates to it in place (no transitions). Close and tab switching live
// in EditorShell.
//
// The tree comes from the server nested (GET /_api/pages); every node is
// real — a label with a title and slug, with content (landing page) or
// without (category whose URL renders a placeholder page). The front page
// is a top-level row with an empty slug, not the parent of the others.
//
// Languages: the LangSelect switches which language the TITLES are shown
// and edited in (rows without a translation show the original, dimmed) —
// the selection is shared shell-wide (./editorLang) with the page editor
// and the page preview. Translated title edits write a per-language
// fragment (POST /_api/structure with lang); the structure itself —
// slugs, order, hierarchy — is language-independent and always edits the
// same tree.
import { computed, inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
import StructureTree from './StructureTree.vue'
import LangSelect from './LangSelect.vue'
import { slugify } from './slugify'
import { flagFor, langName } from './langs'
import { editorLang, pagePrimary } from './editorLang'
import { dropPageCache, loadPlain } from './swapdoc'
const props = defineProps({
pagePath: { type: String, default: '' },
})
const emit = defineEmits(['close', 'pathChange'])
const shell = inject('editorShell', null)
const path = ref('')
const saveError = ref('')
const tree = ref([])
// The language the tree's titles are shown and edited in: "" = primary.
const lang = editorLang
const primaryLang = ref('en')
const siteLangs = ref([])
// The strip's options: the primary language first, then the configured
// translation targets (the lang tab manages that set).
const langOptions = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
.map((code) => ({
tag: code === primaryLang.value ? '' : code,
code,
name: langName(code),
flag: flagFor(code),
primary: code === primaryLang.value,
})),
)
const currentLang = computed(
() => langOptions.value.find((o) => o.tag === lang.value) ?? langOptions.value[0],
)
// The selection is shared (./editorLang): a change re-fetches the tree's
// titles in it (and EditorShell swaps the page preview into it).
watch(lang, () => refreshPages())
// Per-row primary language (Node.language, '' = inherit): the row's
// dropdown lists "inherit" first (naming what it resolves to), then every
// site language. Setting it on a section covers its whole subtree.
const rowLangChoices = computed(() =>
[primaryLang.value, ...siteLangs.value.filter((l) => l !== primaryLang.value)]
.map((code) => ({ tag: code, code, name: langName(code), flag: flagFor(code), primary: false })),
)
function rowLangOptions(el) {
const resolved = el.primary || primaryLang.value
return [
{ tag: '', code: '_inherit', name: `inherit (${langName(resolved)})`, flag: flagFor(resolved), primary: false },
...rowLangChoices.value,
]
}
async function setLanguage(node, tag) {
await postStructure({ path: node.path, language: tag })
}
function normPath(p) {
return p.trim().replace(/^\/+|\/+$/g, '')
}
// Debounce per key: text edits save while typing, without a request per
// keystroke.
const timers = {}
function debounce(key, fn, ms = 600) {
clearTimeout(timers[key])
timers[key] = setTimeout(fn, ms)
}
function updateWindowTitle() {
document.title = 'site structure 🖊️'
}
watch(() => props.pagePath, (p) => { path.value = normPath(p) })
onActivated(updateWindowTitle)
// The shell stays mounted while hidden: when it is re-shown with this tab
// active, restore the window title.
function onEditorShown() {
if (document.body.dataset.editorMode === 'structure') updateWindowTitle()
}
// Tree row focus: switch the edited page and show it, skipping transitions.
async function navigate(p) {
path.value = p
emit('pathChange', p)
await loadPlain(p)
}
// If the currently edited page moved (rename/move of itself or an
// ancestor), follow it to the new path.
function followMove(oldPath, newPath) {
if (path.value === oldPath) navigate(newPath)
else if (oldPath && path.value.startsWith(`${oldPath}/`)) {
navigate(newPath + path.value.slice(oldPath.length))
}
}
// --- New page flow -------------------------------------------------------
// The row at the end of any list adds a *pending* row there: a
// local-only item that can be dragged into place before anything is
// filled in. It is persisted only on commit (✓/Enter), at wherever it
// currently sits.
const pending = ref(null)
function newPage(list) {
if (pending.value) return // one at a time
pending.value = {
slug: '',
path: '',
title: '',
order: 0,
published: true,
has_content: true,
children: [],
pending: true,
}
list.push(pending.value)
}
// Where does the pending row currently sit? -> {parentPath, list, index}.
function locatePending(nodes, parentPath) {
const i = nodes.indexOf(pending.value)
if (i >= 0) return { parentPath, list: nodes, i }
for (const n of nodes) {
const found = locatePending(n.children, n.path)
if (found) return found
}
return null
}
function discardPending() {
const loc = locatePending(tree.value, '')
if (loc) loc.list.splice(loc.i, 1)
pending.value = null
}
async function commitPending() {
const node = pending.value
if (!node) return
// The typed slug is slugified at commit; empty derives one from the
// title (transliterated to ASCII).
const slug = slugify(node.slug.trim()) || slugify(node.title)
if (!slug) {
return
}
const loc = locatePending(tree.value, '')
const parentPath = loc?.parentPath ?? ''
const newPath = parentPath ? `${parentPath}/${slug}` : slug
const res = await fetch(`/_api/pages/${newPath}`, {
method: 'PUT',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
title: node.title.trim() || slug,
markdown: '', // empty markdown creates an empty page (never deletes)
published: true,
}),
})
if (!res.ok) {
// Show the server's reason (e.g. a reserved file name); the pending
// row stays so it can be edited and committed again.
saveError.value = `⚠️ ${await errorDetail(res)}`
return
}
// Place it exactly where the row was dropped: a fresh order key halfway
// between its new siblings (the PUT appended it at the end).
if (loc) {
const prev = loc.list[loc.i - 1]
const next = loc.list[loc.i + 1]
const order = prev && next ? (prev.order + next.order) / 2
: prev ? prev.order + 1
: next ? next.order - 1
: 1
await postStructure({ path: newPath, order })
} else {
saveError.value = ''
}
pending.value = null
await refreshPages()
dropPageCache()
await navigate(newPath)
// Hand over to the page editor tab for the actual writing.
shell?.switchMode('page')
}
// --- Site structure tree (drag-and-drop ordering/moving) ----------------
function findNode(nodes, p) {
for (const n of nodes) {
if (n.path === p) return n
const found = findNode(n.children, p)
if (found) return found
}
return null
}
async function refreshPages() {
try {
const q = lang.value ? `?lang=${lang.value}` : ''
tree.value = await (await fetch(`/_api/pages${q}`)).json()
// The tree carries each node's resolved primary language: publish the
// current page's (the shell pins the preview by it on '' selection).
pagePrimary.value = findNode(tree.value, path.value)?.primary || 'en'
} catch { /* list stays stale; not fatal */ }
}
// Human-readable reason from a failed API call (FastAPI errors carry a
// JSON {detail}), falling back to a generic message.
async function errorDetail(res) {
const body = await res.json().catch(() => null)
return body?.detail || 'changes could not be saved'
}
async function postStructure(op) {
const res = await fetch('/_api/structure', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(op),
})
if (res.ok) {
saveError.value = ''
// Structure changes alter navigation on every page; drop prefetches.
dropPageCache()
loadPlain(path.value) // refresh menus and content from the server
} else {
saveError.value = `⚠️ ${await errorDetail(res)}`
}
await refreshPages()
return res.ok
}
async function onReorder(parentPath, list, evt) {
// vuedraggable already mutated `list`; persist the moved item only: a
// fresh order key halfway between its new siblings (all other items
// keep theirs), plus the new path when the parent changed. The pending
// new-page row is local-only — its position is read at commit time.
const change = evt.moved || evt.added
if (!change) return
const el = change.element
if (el.pending) return
const i = change.newIndex
let prev, next
for (let j = i - 1; j >= 0 && !prev; j--) if (!list[j].pending) prev = list[j]
for (let j = i + 1; j < list.length && !next; j++) if (!list[j].pending) next = list[j]
const order = prev && next ? (prev.order + next.order) / 2
: prev ? prev.order + 1
: next ? next.order - 1
: 1
const newPath = parentPath ? `${parentPath}/${el.slug}` : el.slug
const op = { path: el.path, order }
if (newPath !== el.path) op.move_to = newPath
if (await postStructure(op) && op.move_to) followMove(el.path, op.move_to)
}
// Inline title/slug editing: rows are always editable. Title saves while
// typing (debounced) — in the selected language (a translation writes a
// title fragment, the primary language the original); the slug commits on
// blur/Enter, since it renames the path (moving the whole subtree with it).
// Slugs are language-independent.
function onTitleInput(node, ev) {
const title = ev.target.value.trim()
if (!title || title === node.title) return
debounce(`title:${node.path}`, async () => {
await postStructure({ path: node.path, title, lang: lang.value })
})
}
// Slug inputs are typed freely (spaces become hyphens live, see
// StructureTree onSlugInput); the value is slugified here at commit
// (blur/Enter) before talking to the server, which re-validates (e.g.
// reserved names) and its reason is shown.
async function commitSlug(node, ev) {
const slug = slugify(ev.target.value.trim())
ev.target.value = slug
if (slug === node.slug) return
const parent = node.path.split('/').slice(0, -1).join('/')
// Empty slug at top level = the front page (path "").
const moveTo = parent ? (slug ? `${parent}/${slug}` : parent) : slug
if (await postStructure({ path: node.path, move_to: moveTo })) {
followMove(node.path, moveTo)
} else {
ev.target.value = node.slug // rename failed: put the old slug back
}
}
// Deletion is immediate, no confirmation.
async function removePage(node) {
const res = await fetch(`/_api/pages/${node.path}`, { method: 'DELETE' })
if (res.ok) {
saveError.value = ''
refreshPages()
dropPageCache()
const p = node.path
if (p === path.value || (p && path.value.startsWith(`${p}/`))) {
// The current page was deleted — or reduced to a category, which now
// renders a placeholder page. Either way, re-render from the server.
if (node.children.length) loadPlain(path.value)
else navigate('')
} else {
loadPlain(path.value) // refresh menus
}
} else {
saveError.value = '⚠️ changes could not be saved'
}
}
provide('structureHandlers', {
current: () => path.value,
open: navigate,
removePage,
reorder: onReorder,
titleInput: onTitleInput,
commitSlug,
commitPending,
discardPending,
newPage,
langOptions: rowLangOptions,
setLanguage,
})
onMounted(() => {
path.value = normPath(props.pagePath)
refreshPages()
addEventListener('pagerite:editor-shown', onEditorShown)
// The language strip: site primary + configured targets.
fetch('/_api/settings').then((r) => r.json()).then((s) => {
primaryLang.value = s.primary_lang || 'en'
siteLangs.value = s.translate_langs || []
}).catch(() => { /* no strip */ })
})
onUnmounted(() => {
for (const t of Object.values(timers)) clearTimeout(t)
removeEventListener('pagerite:editor-shown', onEditorShown)
})
</script>
<template>
<div class="structure-editor">
<div v-if="saveError">{{ saveError }}</div>
<div v-if="langOptions.length > 1" class="block lang-block">
<div><LangSelect v-model="lang" :options="langOptions" /></div>
<small v-if="lang" class="muted">
viewing {{ currentLang.name }} titles dimmed rows are untranslated
(shown in the primary language); slugs never translate
</small>
</div>
<section class="block structure">
<StructureTree :nodes="tree" :lang="lang" />
</section>
</div>
</template>
<style scoped>
.structure-editor {
display: flex;
flex-direction: column;
}
.block {
display: flex;
flex-direction: column;
gap: 0.4rem;
padding: 0.5rem 1rem;
border-bottom: 1px solid var(--line);
background: var(--surface);
}
.structure {
flex: 1;
overflow-y: auto;
min-height: 0;
}
/* The language selector is LangSelect.vue — its styles live there. */
.muted {
color: var(--muted);
}
</style>
+70 -29
View File
@@ -1,7 +1,13 @@
<script setup> <script setup>
// Recursive site-structure tree with drag-and-drop ordering (vue-draggable). // Recursive site-structure tree with drag-and-drop ordering (vue-draggable).
// Nodes come from the server (GET /_api/pages via SiteEditor.vue) as // Nodes come from the server (GET /_api/pages via StructureEditor.vue) as
// {slug, path, title, order, published, has_content, children}. // {slug, path, title, translated, order, published, has_content, language,
// primary, children}. The row's flag (LangSelect) sets the node's primary
// language (language; '' = inherit — dimmed, showing the resolved flag);
// the setting covers the whole subtree.
// With a `lang` prop (StructureEditor's language strip) the titles shown
// are that language's; `translated` marks rows with an actual translation
// (untranslated rows show the original title, dimmed).
// Every node is real: a label whose title and slug are always editable // Every node is real: a label whose title and slug are always editable
// inline — the title saves while typing (and focusing it opens the page), // inline — the title saves while typing (and focusing it opens the page),
// the slug commits on blur/Enter since it renames the path, moving the // the slug commits on blur/Enter since it renames the path, moving the
@@ -19,31 +25,37 @@
// The front page is the root row with an empty slug: renaming it away // The front page is the root row with an empty slug: renaming it away
// leaves no front page, and giving another top-level row the empty slug // leaves no front page, and giving another top-level row the empty slug
// makes it the front page. Delete is a two-step inline button (no dialog). // makes it the front page. Delete is a two-step inline button (no dialog).
// Actions are injected from SiteEditor.vue to avoid per-level event // Actions are injected from the parent editor (StructureEditor.vue) to
// forwarding. // avoid per-level event forwarding.
import { inject } from 'vue' import { inject } from 'vue'
import draggable from 'vuedraggable' import draggable from 'vuedraggable'
import { slugify } from './slugify' import { slugify } from './slugify'
import LangSelect from './LangSelect.vue'
defineOptions({ name: 'StructureTree' }) defineOptions({ name: 'StructureTree' })
const props = defineProps({ const props = defineProps({
nodes: { type: Array, required: true }, nodes: { type: Array, required: true },
parentPath: { type: String, default: '' }, parentPath: { type: String, default: '' },
depth: { type: Number, default: 0 }, depth: { type: Number, default: 0 },
// StructureEditor's selected language ('' = original). Only used for the
// untranslated-title styling here; the fetch and title edits live in the
// parent (handlers.titleInput posts the lang with the op).
lang: { type: String, default: '' },
}) })
const handlers = inject('structureHandlers') const handlers = inject('structureHandlers')
// Live-filter the slug inputs as they are typed (oninput): invalid // Slug inputs accept free typing; the only live rewrites are turning
// characters are simply not accepted, spaces become hyphens and unicode // spaces into hyphens and lowercasing (both keep the length for ASCII,
// folds to ASCII (see slugify.js). Existing rows commit on change, the // so the cursor stays put). Anything else (unicode folding, stripping,
// pending row is v-modeled. // collapsing) is left for commit time, where the value is run through
function onSlugInput(ev) { // slugify before talking to the server (StructureEditor). `element` is
ev.target.value = slugify(ev.target.value) // the pending row (v-modeled), null for existing rows (plain :value
} // binding, read back on commit).
function onSlugInput(element, ev) {
function onPendingSlugInput(element, ev) { const v = ev.target.value.replace(/\s/g, '-').toLowerCase()
element.slug = slugify(ev.target.value) ev.target.value = v
if (element) element.slug = v
} }
// Focus the title input of a fresh pending row. // Focus the title input of a fresh pending row.
@@ -104,7 +116,7 @@ function onEnd() {
v-model="element.title" v-model="element.title"
v-focus v-focus
class="edit title-edit" class="edit title-edit"
placeholder="Title" title="Page title"
@keyup.enter="handlers.commitPending()" @keyup.enter="handlers.commitPending()"
@keyup.esc="handlers.discardPending()" @keyup.esc="handlers.discardPending()"
/> />
@@ -113,7 +125,7 @@ function onEnd() {
class="edit slug-edit" class="edit slug-edit"
:placeholder="slugify(element.title)" :placeholder="slugify(element.title)"
title="Slug (last path segment) — empty: derived from the title" title="Slug (last path segment) — empty: derived from the title"
@input="onPendingSlugInput(element, $event)" @input="onSlugInput(element, $event)"
@keyup.enter="handlers.commitPending()" @keyup.enter="handlers.commitPending()"
@keyup.esc="handlers.discardPending()" @keyup.esc="handlers.discardPending()"
/> />
@@ -125,9 +137,11 @@ function onEnd() {
<template v-else> <template v-else>
<input <input
class="edit title-edit" class="edit title-edit"
:class="{ untranslated: lang && !element.translated }"
:value="element.title" :value="element.title"
placeholder="Title" :title="lang && !element.translated
title="Label in the navigation — saves while typing; click opens the page" ? 'No translation yet showing the original; typing creates the translated title'
: 'Label in the navigation saves while typing; click opens the page'"
@input="handlers.titleInput(element, $event)" @input="handlers.titleInput(element, $event)"
@focus="handlers.open(element.path)" @focus="handlers.open(element.path)"
/> />
@@ -136,20 +150,31 @@ function onEnd() {
:value="element.slug" :value="element.slug"
placeholder="front page" placeholder="front page"
title="Slug (last path segment) — renames move the whole subtree. Empty at top level = front page" title="Slug (last path segment) — renames move the whole subtree. Empty at top level = front page"
@input="onSlugInput" @input="onSlugInput(null, $event)"
@change="handlers.commitSlug(element, $event)" @change="handlers.commitSlug(element, $event)"
/> />
<span class="acts"> <span class="acts">
<span
class="row-lang"
:class="{ inherited: !element.language }"
><LangSelect
:model-value="element.language"
:options="handlers.langOptions(element)"
:title="element.language
? `primary language: set on this page (subtree inherits)`
: `primary language: inherited — set it here (subtree inherits)`"
@update:model-value="handlers.setLanguage(element, $event)"
/></span>
<span v-if="!element.published" class="draft">draft</span> <span v-if="!element.published" class="draft">draft</span>
<button <button
v-if="element.has_content || !element.children.length"
type="button" type="button"
class="act del" class="act del"
:class="{ armed: handlers.arming() === element.path }"
:title="element.children.length :title="element.children.length
? 'delete the landing page (the category keeps its subpages)' ? 'delete the landing page (the category keeps its subpages)'
: 'delete page'" : 'delete page'"
@click="handlers.armRemove(element)" @click="handlers.removePage(element)"
>{{ handlers.arming() === element.path ? 'delete?' : '' }}</button> ></button>
</span> </span>
</template> </template>
</div> </div>
@@ -158,6 +183,7 @@ function onEnd() {
:nodes="element.children" :nodes="element.children"
:parent-path="element.path" :parent-path="element.path"
:depth="depth + 1" :depth="depth + 1"
:lang="lang"
/> />
</div> </div>
</template> </template>
@@ -223,7 +249,7 @@ body.tree-dragging .treelist {
level, not across levels). */ level, not across levels). */
.row { .row {
display: grid; display: grid;
grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem 5rem; grid-template-columns: 1.2em minmax(3rem, 1fr) 7rem auto;
align-items: baseline; align-items: baseline;
gap: 0.35rem; gap: 0.35rem;
/* Vertical spacing widens the drop zones: the exposed top strip is the /* Vertical spacing widens the drop zones: the exposed top strip is the
@@ -274,6 +300,13 @@ body.tree-dragging .treelist {
cursor: text; cursor: text;
} }
/* With a language selected (StructureEditor's strip), rows without an
actual translation show the original title dimmed and italic. */
.title-edit.untranslated {
color: var(--muted);
font-style: italic;
}
.slug-edit { .slug-edit {
font-family: var(--font-code); font-family: var(--font-code);
} }
@@ -285,6 +318,20 @@ body.tree-dragging .treelist {
justify-content: end; justify-content: end;
} }
/* Row language selector (LangSelect): the effective primary language's
flag; dimmed while the setting is inherited rather than set on the row. */
.row-lang {
display: inline-flex;
}
.row-lang.inherited :deep(.lang-current) {
opacity: 0.45;
}
.row-lang.inherited:hover :deep(.lang-current) {
opacity: 0.85;
}
.draft { .draft {
color: var(--muted); color: var(--muted);
font-size: 0.75rem; font-size: 0.75rem;
@@ -300,12 +347,6 @@ body.tree-dragging .treelist {
white-space: nowrap; white-space: nowrap;
} }
/* Two-step delete: the first click arms the button, the second deletes. */
.act.armed {
color: #e06c75;
font-weight: 600;
}
.del:hover { .del:hover {
color: #e06c75; color: #e06c75;
} }
+51
View File
@@ -0,0 +1,51 @@
<script setup>
import { computed } from 'vue'
import { formatCount, formatReadTime } from './analytics/format.js'
const props = defineProps({
step: { type: Object, required: true },
count: { type: Number, default: 0 },
favicons: { type: Object, default: null },
})
defineEmits(['close'])
const hasError = computed(() => props.step.status >= 400)
const favicon = computed(() =>
props.step.external && props.step.origin ? props.favicons?.[props.step.origin] : null,
)
const title = computed(() => {
const parts = [props.step.title]
if (props.step.readSeconds > 0) {
parts.push(formatReadTime(props.step.readSeconds))
}
if (hasError.value) {
parts.push(`${props.step.status}`)
}
return parts.filter(Boolean).join(' — ')
})
</script>
<template>
<a class="trail-link"
:class="{ error: hasError }"
:href="step.path"
:title="title"
:target="step.external ? '_blank' : undefined"
:rel="step.external ? 'noopener' : undefined"
@click="(e) => { if (!step.external) $emit('close') }">
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
<img v-if="favicon" class="favicon" :src="favicon" alt="" />
<span>{{ step.slug }}</span>
</a>
</template>
<style scoped>
.favicon {
width: 1em;
height: 1em;
margin-right: 0.25em;
vertical-align: -0.1em;
}
</style>
+307
View File
@@ -0,0 +1,307 @@
<script setup>
/**
* Radial transition map for a pre-filtered time range.
*
* The parent filters transitions, views and visits to the selected range
* before passing them in; `window` carries the absolute [t0, t1) window
* so the visual scale can normalize against a one-week reference.
*/
import { computed, onBeforeUnmount, onMounted, shallowRef, watch } from 'vue'
import { DAY } from './analytics/time.js'
import { formatCount, formatReadTime } from './analytics/format.js'
import {
TNODE_W,
TNODE_H,
BEAD_R,
buildTransitionGraph,
} from './analytics/transitions.js'
const props = defineProps({
data: { type: Object, default: null },
window: { type: Object, required: true },
pageTree: { type: Array, default: null },
favicons: { type: Object, default: null },
})
// origin -> /_f/... icon URL, keyed by the node's origin (source/exit
// pills only; UTM-tagged source nodes without an https origin stay
// text-only).
const extFavicon = (x) => {
if (!x.path?.startsWith('https://')) return null
try {
return props.favicons?.[new URL(x.path).origin] || null
} catch {
return null
}
}
const dayScale = computed(() => {
const { t0, t1 } = props.window
// Convert raw counts to a daily hit rate (hits/day).
if (t0 != null && t1 != null) return DAY / (t1 - t0)
// 'all': scale by the actual data span, but never less than the 30-day
// minimum the plot enforces, so sparse young data is not over-amplified.
const times = new Set()
for (const buckets of Object.values(props.data?.views || {})) {
for (const k of Object.keys(buckets)) times.add(Date.parse(k))
}
const arr = [...times]
if (arr.length < 2) return 1
const span = Math.max(...arr) - Math.min(...arr)
return DAY / Math.max(span, 30 * DAY)
})
const graph = computed(() =>
props.data
? buildTransitionGraph(props.data, props.pageTree, props.data.visits || [], dayScale.value)
: null,
)
// Bead animation: every bead is simulated independently in JS. Each flow
// (one per edge direction) emits a bead every `interval` seconds; beads
// cross their segment in a constant TRAVERSAL_S seconds (speed relative
// to span length) and are dropped at the end.
// There is deliberately no cap on beads in flight.
// Emitters persist across data reloads, keyed by flow.key: an unchanged
// link keeps its emission phase and in-flight beads (tracked by progress,
// not absolute time), so a count change elsewhere never reshuffles them.
const beads = shallowRef([])
let rafId = 0
const emitters = new Map() // flow.key -> { flow, interval, next, alive }
const live = [] // { e, p } — beads in flight, p = progress 0..1
let lastTick = 0
const MAX_BEAD_RATE = 120 // upper bound on total beads per second
const TRAVERSAL_S = 0.4 // seconds to cross any segment, end to end
const syncBeads = (flows) => {
const reduced = matchMedia('(prefers-reduced-motion: reduce)').matches
if (!flows?.length || reduced) {
emitters.clear()
live.length = 0
beads.value = []
return
}
// Cap the total bead emission rate so a busy range cannot spawn enough
// beads to kill the page. Existing per-range time scaling is preserved;
// this is only a proportional emergency throttle when the limit is hit.
const totalRate = flows.reduce((s, f) => s + 1 / f.interval, 0)
const scale = totalRate > MAX_BEAD_RATE ? MAX_BEAD_RATE / totalRate : 1
const now = performance.now()
const seen = new Set()
for (const flow of flows) {
seen.add(flow.key)
const interval = (flow.interval / scale) * 1000
const e = emitters.get(flow.key)
if (e) {
e.flow = flow // pick up new geometry/rate, keep the phase
e.interval = interval
continue
}
// New emitter: pre-fill the traversal with evenly spaced beads (random
// phase), so the flow appears already running instead of empty.
const phase = Math.random() * interval
const dp = interval / 1000 / TRAVERSAL_S
const ne = { flow, interval, next: now + phase, alive: true }
for (let p = 1 - phase / 1000 / TRAVERSAL_S; p > 0; p -= dp) {
live.push({ e: ne, p })
}
emitters.set(flow.key, ne)
}
for (const [key, e] of emitters) {
if (!seen.has(key)) {
e.alive = false
emitters.delete(key)
}
}
for (let i = live.length - 1; i >= 0; i--) {
if (!live[i].e.alive) live.splice(i, 1)
}
}
const tick = (t) => {
const dt = lastTick ? (t - lastTick) / 1000 : 0
lastTick = t
for (const e of emitters.values()) {
while (e.next <= t) {
live.push({ e, p: 0 })
e.next += e.interval
}
}
const out = []
for (let i = live.length - 1; i >= 0; i--) {
const b = live[i]
b.p += dt / TRAVERSAL_S
if (b.p >= 1) {
live.splice(i, 1)
continue
}
const f = b.e.flow
out.push({ x: f.x1 + (f.x2 - f.x1) * b.p, y: f.y1 + (f.y2 - f.y1) * b.p })
}
beads.value = out
rafId = requestAnimationFrame(tick)
}
watch(() => graph.value?.flows, syncBeads, { immediate: true })
onMounted(() => {
if (!matchMedia('(prefers-reduced-motion: reduce)').matches) {
rafId = requestAnimationFrame(tick)
}
})
onBeforeUnmount(() => cancelAnimationFrame(rafId))
// The svg never renders larger than its natural size (1 viewBox unit = 1
// px, max-width below): the layout geometry is designed in pixel-like
// units, and upscaling would blow up the pills around their text. Narrow
// panels scale the graph down to fit (width: 100%), text along with it.
// Pill text is not truncated: text is clipped at the pill's rounded border
// (clipPath per node, inset a few units for padding). Captions center when
// they fit; overlong ones anchor left so their beginning (not their
// middle) survives the clip. Width estimate: ~0.52 em per glyph.
const fitsPill = (label, fontPx = 19) => label.length * 0.52 * fontPx <= TNODE_W - 16
// With a favicon the label leaves room for the icon at the pill's left
// and is always left-anchored past it.
const labelX = (x) =>
extFavicon(x) ? x.x - TNODE_W / 2 + 36 : fitsPill(x.label) ? x.x : x.x - TNODE_W / 2 + 8
const labelAnchor = (x) => (!extFavicon(x) && fitsPill(x.label) ? 'middle' : 'start')
const countLabel = (n) =>
n.readSec ? `${formatCount(n.views)}×${formatReadTime(n.readSec)}` : formatCount(n.views)
</script>
<template>
<section v-if="graph">
<svg class="tmap" :style="{ maxWidth: `${graph.bounds.x1 - graph.bounds.x0}px` }" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
role="img" aria-label="map of transitions between pages">
<path v-for="(a, i) in graph.arcs" :key="'a' + i"
:id="`tarc${i}`" :d="a.d" :class="['tarc', a.top && 'tarc-top']" />
<template v-for="(a, i) in graph.arcs" :key="'t' + i">
<path v-if="a.ld" :id="`tarcl${i}`" :d="a.ld" fill="none" stroke="none" />
<text v-if="a.ld" class="tarclabel" :class="{ 'tarclabel-top': a.top }"><textPath :href="`#tarcl${i}`" startOffset="0">{{ a.label }}</textPath></text>
</template>
<path v-for="(e, i) in graph.edges" :key="'e' + i"
:d="e.d" :class="['tconn', e.external && 'tconn-exit']">
<title>{{ e.title }}</title>
</path>
<circle v-for="(b, i) in beads" :key="'b' + i"
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
<g v-for="(x, i) in graph.extNodes" :key="'x' + i">
<clipPath :id="`xclip${i}`">
<rect :x="x.x - TNODE_W/2 + 6" :y="x.y - TNODE_H/2" :width="TNODE_W - 12"
:height="TNODE_H" :rx="TNODE_H/2 - 4" />
</clipPath>
<a v-if="x.href" :href="x.href" target="_blank" rel="noopener">
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<image v-if="extFavicon(x)" :href="extFavicon(x)" :x="x.x - TNODE_W/2 + 12" :y="x.y - TNODE_H*0.16 - 11" width="22" height="22" />
<text :x="labelX(x)" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: labelAnchor(x) }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</a>
<g v-else>
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<image v-if="extFavicon(x)" :href="extFavicon(x)" :x="x.x - TNODE_W/2 + 12" :y="x.y - TNODE_H*0.16 - 11" width="22" height="22" />
<text :x="labelX(x)" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: labelAnchor(x) }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</g>
</g>
<g v-for="(n, i) in graph.nodes" :key="n.path">
<clipPath :id="`nclip${i}`">
<rect :x="n.x - TNODE_W/2 + 6" :y="n.y - TNODE_H/2" :width="TNODE_W - 12"
:height="TNODE_H" :rx="TNODE_H/2 - 4" />
</clipPath>
<a :href="n.path">
<title>{{ n.title }}</title>
<rect :x="n.x - TNODE_W/2" :y="n.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2" class="tnode" />
<g :clip-path="`url(#nclip${i})`">
<text :x="fitsPill(n.label) ? n.x : n.x - TNODE_W/2 + 8" :y="n.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(n.label) ? 'middle' : 'start' }">{{ n.label }}</text>
<text :x="n.x" :y="n.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">
{{ countLabel(n) }}
</text>
</g>
</a>
</g>
</svg>
</section>
</template>
<style scoped>
/* Transition map: radial graph of internal page-to-page transitions. */
.tmap {
display: block;
width: 100%;
/* max-width is set inline to the natural content width (px = viewBox
units), so wide panels never upscale the graph beyond 1:1. */
margin: 0 auto;
}
.tmap .tconn {
fill: var(--accent);
opacity: 0.4; /* uniform, not strength-encoded: width carries that */
}
.tmap .tconn-exit {
fill: var(--text);
}
.tmap .tbead {
fill: var(--accent);
opacity: 0.85;
filter: drop-shadow(0 0 2.5px var(--accent));
}
.tmap .txnode {
fill: var(--text);
stroke: none;
}
.tmap .txnode-source { fill: var(--text); }
.tmap .txnode-exit { fill: var(--text); }
/* Branch lanes: one wide concentric arc per path prefix, running behind
the node pills around the fan's circle center; parent levels sit one
indent (radius step) outward. Each lane's label follows a short guide
arc across the first inter-node gap (the part pills never cover). */
.tmap .tarc {
fill: none;
stroke: var(--muted);
stroke-width: 16;
opacity: 0.25;
}
.tmap .tarc-top { stroke-width: 24; }
/* Lane labels are left-aligned: each guide arc starts just past the source
pill's edge, the earliest point where the text is visible. */
.tmap .tarclabel {
fill: var(--muted);
font-size: 13px;
text-anchor: start;
}
/* The top lane is 50% thicker; its 🏠︎ label scales along. */
.tmap .tarclabel-top {
font-size: 19.5px;
}
.tmap .tnode {
fill: var(--accent);
stroke: none;
}
/* Text sizes are viewBox units: they shrink along with the graph on
narrow panels. Overlong labels are clipped at the pill border. */
.tmap .tnodeslug {
fill: var(--bg, Canvas);
font-size: 19px;
text-anchor: start;
}
.tmap a { cursor: pointer; }
.tmap .tnodecount {
fill: var(--bg, Canvas);
opacity: 0.75;
font-size: 15px;
text-anchor: middle;
}
section { margin-top: 1.8rem; }
</style>
+156
View File
@@ -0,0 +1,156 @@
<script setup>
// Visitor metadata cell shared by the recent-visits, crawlers, and abuse tables.
// Displays IP/network/host, country flag/city, UA, and language when available.
// Clicking the IP copies the full address to the clipboard.
// ``variantCount`` overrides the UA line to warn when multiple client
// fingerprints share the same IP (e.g. a scanner rotating UAs).
import { computed } from 'vue'
import * as flagSvgs from 'country-flag-icons/string/3x2'
import { copyIp, formatLang } from './analytics/format.js'
const props = defineProps({
ip: { type: String, default: '' },
ipDisplay: { type: String, default: '—' },
ua: { type: String, default: '' },
uaRaw: { type: String, default: '' },
country: { type: String, default: '' },
city: { type: String, default: '' },
lang: { type: String, default: '' },
langDisplay: { type: String, default: '' },
isHost: { type: Boolean, default: false },
variantCount: { type: Number, default: 1 },
})
const hasCountry = computed(() => !!(props.country && props.country !== '—'))
const hasCity = computed(() => !!(props.city && props.city !== '—'))
const hasLocale = computed(() => hasCountry.value || hasCity.value)
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
const showLang = computed(() => langValue.value && langValue.value !== '—')
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script>
<template>
<td class="visitor-cell" :class="{ 'host-cell': isHost }">
<div class="visitor-rows">
<div class="visitor-row">
<div class="locale-line">
<span v-if="flagSvg(country)" class="flag" v-html="flagSvg(country)" :title="countryName(country) || country"></span>
<template v-if="hasCity"><small class="city-name muted">{{ city }}</small></template>
<template v-else-if="!hasLocale"></template>
</div>
<div class="ip-line">
<span class="clickable-ip small muted"
:title="ip"
@click="copyIp(ip, $event)">{{ ipDisplay }}</span>
</div>
</div>
<div class="visitor-row">
<div class="ua-line">
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
</div>
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
</div>
</div>
</td>
</template>
<style scoped>
.visitor-cell {
width: 18em;
max-width: 18em;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: top;
}
.visitor-cell.host-cell {
text-align: right;
}
.visitor-rows {
display: flex;
flex-direction: column;
gap: 0.15rem;
}
.visitor-row {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.5rem;
}
.visitor-row > * {
min-width: 0;
}
.locale-line,
.ip-line,
.ua-line {
flex: 1 1 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.locale-line {
text-align: left;
display: flex;
align-items: center;
gap: 0.3rem;
}
.ip-line {
text-align: right;
}
.ua-line {
text-align: left;
}
.locale-lang {
flex: 0 0 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
text-align: right;
}
.city-name {
display: inline-block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
vertical-align: middle;
}
.flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+209
View File
@@ -0,0 +1,209 @@
<script setup>
/**
* Visitor and page-view smoothed curves for a single shared time range.
*/
import { computed, onMounted, onUnmounted, ref } from 'vue'
import { makeSeries } from './analytics/time.js'
import {
CHART_H,
CHART_W,
MARGIN_B,
MARGIN_L,
VIEW_H,
VIEW_W,
buildChart,
} from './analytics/chart.js'
const DAY_REFRESH_MS = 15000
// Keep the whole svg within page bounds: full width below the natural
// size, centered with equal side margins above it (max() clamps the
// centering margin to 0 at the breakpoint, so the rule is continuous).
const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const props = defineProps({
data: { type: Object, default: null },
range: { type: String, required: true },
})
// Views across all pages combined into one raw bucket map.
const allViews = computed(() => {
const all = {}
for (const buckets of Object.values(props.data?.views || {})) {
for (const [k, c] of Object.entries(buckets)) all[k] = (all[k] || 0) + c
}
return all
})
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
const viewSeries = computed(() => makeSeries(allViews.value, props.range))
function freqLabel(unit) {
return unit === '5min' ? '5 min' : unit === 'hour' ? 'hourly' : 'daily'
}
/** Vertical axis caption: "visits / 5 min" on the day view, else "hourly visits" style. */
function axisLabel(unit, ylabel) {
return unit === '5min' ? `${ylabel} / 5 min` : `${freqLabel(unit)} ${ylabel}`
}
/** Legend label for the overlaid past weeks: "Week M" or "Week MN". */
function pastLabel(series) {
const oldest = series.at(-1).label.slice(5) // strip "Week "
return series.length > 2 ? `Week ${oldest}${series[1].label.slice(5)}` : `Week ${oldest}`
}
const now = ref(Date.now())
let refreshInterval = null
onMounted(() => {
refreshInterval = setInterval(() => { now.value = Date.now() }, DAY_REFRESH_MS)
})
onUnmounted(() => {
if (refreshInterval) clearInterval(refreshInterval)
})
const visitChart = computed(() => buildChart(visitSeries.value, now.value))
const viewChart = computed(() => buildChart(viewSeries.value, now.value))
</script>
<template>
<section v-for="c in [
{ ylabel: 'visits', chart: visitChart, legend: true },
{ ylabel: 'views', chart: viewChart, legend: false },
]" :key="c.ylabel">
<template v-if="c.chart">
<svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`"
:style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }"
role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)">
<!-- Clip the plot curves to the chart area: past-week overlays can
run far above the autoscaled y range, and the svg itself is
overflow: visible for the axis labels. -->
<clipPath :id="`plot-${c.ylabel}`">
<rect x="0" y="0" :width="CHART_W" :height="CHART_H" />
</clipPath>
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
<template v-for="t in c.chart.xticks" :key="'t' + t.x">
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
class="minor vertical" />
</template>
<g :clip-path="`url(#plot-${c.ylabel})`">
<template v-if="c.chart.bars">
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
<path :d="c.chart.skyline" class="line" />
</template>
<template v-else>
<!-- Oldest overlay weeks first so the current week paints on top. -->
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
<path v-if="s.area" :d="s.area" class="area" />
<path :d="s.line" class="line" :class="{ past: s.past }"
:style="{ opacity: s.opacity }" />
</template>
</template>
</g>
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
class="axis" />
<text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y"
text-anchor="end" dominant-baseline="middle" class="ylab">{{ g.label }}</text>
<text :x="-(MARGIN_L - 10)" :y="CHART_H / 2" text-anchor="middle"
:transform="`rotate(-90 ${-(MARGIN_L - 10)} ${CHART_H / 2})`"
class="yaxis-label">{{ axisLabel(c.chart.unit, c.ylabel) }}</text>
<text v-for="t in c.chart.xticks" :key="'x' + t.x" :x="t.x" :y="CHART_H + MARGIN_B - 8"
text-anchor="middle" class="xlab">{{ t.label }}</text>
<!-- Week overlay legend, top right inside the plot: current week in
accent, one muted specimen for the whole past range. -->
<g v-if="c.legend && c.chart.series.length > 1">
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="10" y2="10" class="line" />
<text :x="CHART_W - 72" y="10" dominant-baseline="middle"
class="leglab">{{ c.chart.series[0].label }}</text>
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="25" y2="25"
class="line past" style="opacity: 0.6" />
<text :x="CHART_W - 72" y="25" dominant-baseline="middle"
class="leglab">{{ pastLabel(c.chart.series) }}</text>
</g>
</svg>
</template>
</section>
</template>
<style scoped>
/* Each chart is a self-contained SVG: the viewBox includes the axis label
margins, so nothing is positioned with HTML overlays. Never upscale past
the natural size (1 viewBox unit = 1 px, max-width set inline) — that
would blow up the constant-size text; smaller panels still scale the
chart down to fit. The margin-left (set inline) centers the chart above
its natural width; the svg always stays within page bounds.
overflow: visible lets wider fonts extend past the viewBox instead of
clipping. */
.chart {
display: block;
width: 100%;
height: auto;
overflow: visible;
}
.chart .ylab,
.chart .xlab,
.chart .yaxis-label,
.chart .leglab {
font-family: system-ui, sans-serif; /* theme fonts can be overly styled */
font-size: 11px;
fill: var(--muted);
}
.chart .ylab {
font-variant-numeric: tabular-nums;
}
.chart .minor {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
opacity: 0.35;
}
.chart .minor.vertical {
opacity: 0.25;
}
.chart .major {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
stroke-dasharray: 3 4;
opacity: 0.8;
}
.chart .axis {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
}
.chart .area {
fill: var(--accent);
opacity: 0.15;
}
.chart .bar {
fill: var(--accent);
opacity: 0.15;
}
.chart .line {
fill: none;
stroke: var(--accent);
stroke-width: 2;
vector-effect: non-scaling-stroke;
stroke-linejoin: round;
stroke-linecap: round;
}
/* Past overlay weeks contrast with the current week's accent color. */
.chart .line.past {
stroke: var(--muted);
}
.empty { color: var(--muted); }
</style>
+30
View File
@@ -0,0 +1,30 @@
// Analytics page entry: mounts AnalyticsView inside the normal page layout.
// In production the backend inlines this module into the /_a page (and
// pagerite.js re-creates the script element after fetch-navigations there);
// in dev pagerite.js imports it from the Vite dev server on demand. Either
// way it auto-mounts on #analytics-app when it evaluates, and unmounts when
// pagerite.js announces a swap away from /_a.
import { createApp } from 'vue'
import AnalyticsView from './AnalyticsView.vue'
let app = null
export function mount(container) {
if (app) return
app = createApp(AnalyticsView)
app.mount(container)
}
export function unmount() {
app?.unmount()
app = null
}
// pagerite.js calls this before swapping away from /_a; each evaluation
// (the inlined production module evaluates fresh on every visit) replaces
// the handle.
window.__pageriteAnalyticsUnmount = unmount
// Auto-mount when the page holding #analytics-app is present.
const container = document.getElementById('analytics-app')
if (container) mount(container)
+388
View File
@@ -0,0 +1,388 @@
/**
* Chart geometry, smoothing, and SVG path generation for analytics charts.
*
* Fixed 720x180 plot area inside a larger viewBox that also holds the axis
* labels, so each chart SVG is self-contained; values are per-unit rates
* (hour on the week view, day on month+).
*/
import { DAY, HOUR, MIN5, WEEK, mondayUTC } from './time.js'
import { formatCount } from './format.js'
export const CHART_W = 1000
export const CHART_H = 150
export const PAD_TOP = 14 // room above the highest point
export const MARGIN_L = 40 // y tick labels + vertical axis label
export const MARGIN_B = 24 // x tick labels
export const VIEW_W = MARGIN_L + CHART_W + 8
export const VIEW_H = CHART_H + MARGIN_B
/**
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
* most 5 intervals, so labeled ticks are always round and evenly divided.
* A minimum range of 10 keeps tiny near-zero values (e.g. a single visit)
* from being enlarged to a fractional scale; minor lines subdivide each
* major step in five when that yields integers.
*/
export function yScale(maxValue) {
let step = 1
outer: for (let exp = -3; exp < 8; exp++) {
for (const base of [1, 2, 5]) {
step = base * 10 ** exp
if (Math.ceil(maxValue / step) <= 5) break outer
}
}
let max = Math.ceil(maxValue / step) * step
if (max < 10) {
max = 10
step = 2
}
const minor = step >= 5 && step % 5 === 0 ? step / 5 : null
return { max, step, minor }
}
/**
* Edge-aware Gaussian smoothing with a fixed bandwidth. A change-point
* detector first finds traffic-level shifts (two-unit totals compared on
* both sides of each bucket; strong ratio + significance marks a candidate,
* and each run of candidates keeps only its best-scoring bucket as an
* edge). Each edge-delimited segment is then smoothed independently: every
* bucket spreads its count with a fixed Gaussian sigma chosen so N events
* in a single bucket peak at N events per unit. Mass past a detected change
* point is dropped (kernel renormalized); mass past a true series edge is
* mirrored back, so the curve doesn't fall where data simply ends. Either
* way total visitor count is preserved exactly. The unit is
* one hour on the week view and one day on the month+ views, so the
* smoothing time scale follows the range. The raw series is drawn faintly
* behind the curve for reference. Operates on raw counts.
*/
export function smooth(counts, binMinutes, unitMinutes, {
detectorWindowMinutes = 2 * unitMinutes,
// Count thresholds are defined per hour and scale with the unit, so
// "low traffic" means the same thing on hourly and daily views
// (5-20 events/hour = 120-480/day on the month+ ranges).
highTrafficEvents = 10 * unitMinutes / 60,
minRatio = 2.5,
minSignificance = 4,
} = {}) {
const n = counts.length
if (!n) return counts
const detectorWindowBins = Math.max(1, Math.round(detectorWindowMinutes / binMinutes))
const cumsum = new Float64Array(n + 1)
for (let i = 0; i < n; i++) cumsum[i + 1] = cumsum[i] + counts[i]
// Detect abrupt regime changes from aggregated traffic on both sides.
// Individual bins are deliberately ignored because even high traffic
// produces many 0-1 count bins at five-minute resolution.
const score = new Float64Array(n)
const candidate = new Uint8Array(n)
for (let i = detectorWindowBins; i < n - detectorWindowBins; i++) {
const left = cumsum[i] - cumsum[i - detectorWindowBins]
const right = cumsum[i + detectorWindowBins] - cumsum[i]
const high = Math.max(left, right)
const low = Math.min(left, right)
if (high < highTrafficEvents) continue
const ratio = (high + 1) / (low + 1)
const significance = (high - low) / Math.sqrt(high + low + 1)
if (ratio >= minRatio && significance >= minSignificance) {
candidate[i] = 1
score[i] = significance * Math.log(ratio)
}
}
// Collapse each continuous detector region to its strongest boundary.
const edges = []
for (let i = 0; i < n;) {
if (!candidate[i]) { i++; continue }
let j = i + 1
while (j < n && candidate[j]) j++
let best = i
for (let k = i + 1; k < j; k++) {
if (score[k] > score[best]) best = k
}
edges.push(best)
i = j
}
// Fixed sigma: N events in one bucket peak at N events per unit.
// sigma_bins * sqrt(2*pi) = rate = unitMinutes / binMinutes.
const sigmaBins = unitMinutes / (binMinutes * Math.sqrt(2 * Math.PI))
const radius = Math.ceil(4 * sigmaBins)
// Process each discontinuity-delimited regime independently so the
// Gaussian cannot see through a detected boundary. Each input bin spreads
// its count with the fixed sigma. Mass that would fall past a detected
// change point is dropped and the kernel renormalized; mass that would
// fall past a true series edge (first/last bin) is mirrored back into the
// segment, as if the data continued as its own reflection, so constant or
// rising data doesn't produce a spurious falling edge. Total visitor count
// is preserved apart from floating-point error.
const bounds = [0, ...edges, n]
const smoothed = new Float64Array(n)
for (let b = 0; b < bounds.length - 1; b++) {
const lo = bounds[b]
const length = bounds[b + 1] - lo
const mirrorLeft = lo === 0
const mirrorRight = lo + length === n
const segment = counts.slice(lo, lo + length)
for (let j = 0; j < length; j++) {
const count = segment[j]
if (!count) continue
// Collect (target bin, weight) pairs over the full kernel, folding
// mirrored mass at series edges and dropping mass past change points.
const spread = new Map()
let weightSum = 0
for (let i = j - radius; i <= j + radius; i++) {
let k = i
// Fold repeatedly for segments shorter than the kernel radius.
while (k < 0 || k >= length) {
if (k < 0 && mirrorLeft) k = -k - 1
else if (k >= length && mirrorRight) k = 2 * length - 1 - k
else { k = null; break }
}
if (k === null) continue
const w = Math.exp(-0.5 * ((i - j) / sigmaBins) ** 2)
spread.set(k, (spread.get(k) || 0) + w)
weightSum += w
}
for (const [k, w] of spread) {
smoothed[lo + k] += count * w / weightSum
}
}
}
return [...smoothed]
}
/**
* Catmull-Rom spline through the (smoothed) points, control points clamped
* to the plot area so the curve can never dip below zero or above the max.
*/
export function spline(pts) {
if (pts.length < 3) {
return `M${pts.map((p) => `${p.x},${p.y}`).join('L')}`
}
const clampY = (y) => Math.min(CHART_H, Math.max(PAD_TOP, y))
let d = `M${pts[0].x},${pts[0].y}`
for (let i = 0; i < pts.length - 1; i++) {
const p0 = pts[i - 1] || pts[i]
const p1 = pts[i]
const p2 = pts[i + 1]
const p3 = pts[i + 2] || p2
const c1y = clampY(p1.y + (p2.y - p0.y) / 6)
const c2y = clampY(p2.y - (p3.y - p1.y) / 6)
d += `C${p1.x + (p2.x - p0.x) / 6},${c1y} `
+ `${p2.x - (p3.x - p1.x) / 6},${c2y} ${p2.x},${p2.y}`
}
return d
}
/** Build a full chart model from a series descriptor produced by time.js. */
export function buildChart(input, now = Date.now()) {
if (!input || !input.series.length) return null
if (input.unit === '5min') return buildDayChart(input, now)
const { series, t0, t1, rate, binMinutes, unitMinutes, unit } = input
// Values are per-unit rates (hour on the week view, day on month+); the
// y max is derived from the *smoothed* curves so random single-bucket
// spikes don't blow up the scale. Smoothing works on raw counts (its edge
// detector thresholds are count-based), the result is scaled back to rates.
const smoothed = series.map((s) =>
smooth(s.points.map((p) => p.count), binMinutes, unitMinutes).map((v) => v * rate))
// Scale from the current/primary series only; older overlay weeks are drawn
// with the same scale and allowed to overflow if they are busier.
const highest = Math.max(0, ...smoothed[0])
const { max, step, minor } = yScale(highest)
const x = (t) => ((t - t0) / (t1 - t0)) * CHART_W
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const drawn = series.map((s, si) => {
const pts = s.points.map((p, i) => ({ x: x(p.t), y: y(smoothed[si][i]) }))
const line = spline(pts)
const first = pts[0]
const last = pts.at(-1)
return {
...s,
line,
area: s.area ? `${line}L${last.x},${CHART_H}L${first.x},${CHART_H}Z` : null,
}
})
// Major (labeled) and minor (hairline) y grid ticks.
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), label: fmtY(v) })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
// X ticks. Week view: weekday labels centered at midday UTC, no vertical
// lines (day boundaries would be misleading in the viewer's timezone).
// Month view: likewise lineless, day numbers at noon UTC with the month
// name substituted for the 1st (marking the month change). Longer
// ranges: boundary lines at Mondays / months / years.
const isWeek = t1 - t0 === WEEK
const isMonth = !isWeek && t1 - t0 <= 31 * DAY
let xticks
if (isWeek) {
xticks = Array.from({ length: 7 }, (_, d) => {
const t = t0 + d * DAY + 12 * HOUR
return {
x: x(t),
label: new Date(t).toLocaleDateString(undefined, {
weekday: 'short', timeZone: 'UTC',
}),
line: false,
}
})
} else if (isMonth) {
// t0 is day-aligned; label every day whose noon falls inside the range.
xticks = []
for (let day = t0; day + 12 * HOUR < t1; day += DAY) {
const date = new Date(day)
const t = day + 12 * HOUR
xticks.push({
x: x(t),
label: date.getUTCDate() === 1
? date.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' })
: String(date.getUTCDate()),
line: false,
})
}
} else {
xticks = xticksFor(t0, t1).map((t) => ({
x: x(t), label: fmtTick(t, t1 - t0), line: true,
}))
}
return { max, majors, minors, series: drawn, xticks, unit }
}
/**
* Day view: 5-minute bars for the last 24 hours. Bars are drawn at raw
* counts; the skyline uses a projected full-bucket value for the still-open
* final bucket. The y scale is derived from the projected skyline maximum.
*/
export function buildDayChart(input, now = Date.now()) {
const { series, t0, t1 } = input
const points = series[0]?.points || []
const n = points.length
if (!n) return null
const bucketMs = (t1 - t0) / n
const bucketWidth = CHART_W / n
const gap = 0.2
const barWidth = Math.max(0.2, bucketWidth - gap)
const x = (i) => i * bucketWidth + gap / 2
const prevRaw = n > 1 ? points[n - 2].count : 0
const projected = points.map((p, i) => {
if (i !== n - 1) return p.count
const bucketStart = t0 + i * bucketMs
const elapsed = Math.max(1, Math.min(bucketMs, now - bucketStart))
// Blend the observed partial bucket with the previous full bucket:
// the longer the current bucket has run, the less we borrow from it.
const share = elapsed / bucketMs
return p.count + prevRaw * (1 - share)
})
const highest = Math.max(0, ...projected)
const { max, step, minor } = yScale(highest)
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const bars = points.map((p, i) => {
const bx = x(i)
const by = y(p.count)
return {
x: bx,
y: by,
width: barWidth,
height: CHART_H - by,
raw: p.count,
projected: projected[i],
}
})
let skyline = ''
for (let i = 0; i < bars.length; i++) {
const b = bars[i]
const top = y(b.projected)
if (i === 0) {
skyline += `M${b.x},${top} H${b.x + b.width}`
} else {
skyline += ` V${top} H${b.x + b.width}`
}
}
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), label: fmtY(v) })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
const xticks = []
const tickStep = 3 * HOUR
const firstTick = Math.ceil(t0 / tickStep) * tickStep
for (let t = firstTick; t < t1; t += tickStep) {
if (t < t0) continue
const d = new Date(t)
xticks.push({
x: ((t - t0) / (t1 - t0)) * CHART_W,
label: `${String(d.getUTCHours()).padStart(2, '0')}:00`,
line: false,
})
}
return { bars, skyline: skyline.trim(), max, majors, minors, xticks, unit: '5min', series: [] }
}
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month
* boundaries up to a few years, then years. */
export function xticksFor(t0, t1) {
const span = t1 - t0
const ticks = []
if (span <= 100 * DAY) {
for (let t = mondayUTC(t0); t <= t1; t += WEEK) {
if (t >= t0) ticks.push(t)
}
return ticks
}
if (span <= 4 * 365 * DAY) {
const d = new Date(t0)
let t = Date.UTC(d.getUTCFullYear(), d.getUTCMonth() + 1, 1)
for (; t <= t1;) {
ticks.push(t)
const m = new Date(t)
t = Date.UTC(m.getUTCFullYear(), m.getUTCMonth() + 1, 1)
}
return ticks
}
const d = new Date(t0)
for (let yr = d.getUTCFullYear() + 1; Date.UTC(yr, 0, 1) <= t1; yr++) {
ticks.push(Date.UTC(yr, 0, 1))
}
return ticks
}
export function fmtTick(t, span) {
const d = new Date(t)
if (span <= 100 * DAY) {
return d.toLocaleDateString(undefined, { month: 'short', day: 'numeric', timeZone: 'UTC' })
}
if (span <= 4 * 365 * DAY) {
return d.getUTCMonth() === 0
? d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
: d.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' })
}
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
}
/** Y labels use the same compact formatter as text labels. */
export function fmtY(v) {
return formatCount(v)
}
+591
View File
@@ -0,0 +1,591 @@
/**
* Formatters and aggregators for summary sections: totals and the recent
* visit trail.
*/
/**
* IPv4 unchanged, IPv6 returns the /64 network prefix in compact form.
* Falls back to the original value when parsing fails.
*/
export const hostIP = (ip) => {
try {
if (!ip || !ip.includes(':')) return ip
const strip = (s) => s.replace(/^\[|\]$/g, '')
const norm = strip(new URL(`http://[${ip}]/`).hostname)
const [l, r] = norm.split('::').map((s) => (s ? s.split(':') : []))
const full = r
? [...l, ...Array(8 - l.length - r.length).fill('0'), ...r]
: l
return strip(
new URL(`http://[${full.slice(0, 4).join(':')}::]/`).hostname,
).replace(/::$/, '')
} catch (e) {
console.error('hostIP processing failed for:', ip, e)
return ip
}
}
function showCopiedFeedback(el) {
if (!el || typeof document === 'undefined') return
const popup = document.createElement('span')
popup.textContent = 'Copied!'
popup.className = 'copy-popup'
popup.style.cssText =
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
'pointer-events:none;z-index:10;'
el.classList.add('has-copy-popup')
el.appendChild(popup)
setTimeout(() => {
popup.remove()
el.classList.remove('has-copy-popup')
}, 1200)
}
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
export async function copyIp(ip, event) {
if (!ip) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(ip)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
export async function copyList(text, event) {
if (!text) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(text)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Total page views across every page and every bucket. */
export function calcTotalViews(views) {
let n = 0
for (const buckets of Object.values(views || {})) {
for (const c of Object.values(buckets)) n += c
}
return n
}
// Very short reads are navigation/skims, not real reading time.
export const MIN_READ_SECONDS = 10
/** path -> accumulated read seconds for a visit, derived from its trail. */
export function readMapOf(v) {
const map = {}
for (const item of Object.values(v.trail || {})) {
if (item.read) map[item.to] = (map[item.to] || 0) + item.read
}
return map
}
/** Average minutes per visit and average of per-article median read minutes. */
export function calcReadStats(visits) {
const perArticle = {}
let totalVisitSeconds = 0
let visitCount = 0
for (const v of visits || []) {
const read = readMapOf(v)
const secs = Object.values(read).filter((s) => s >= MIN_READ_SECONDS)
if (!secs.length) continue
visitCount++
totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
for (const [path, s] of Object.entries(read)) {
if (s >= MIN_READ_SECONDS) {
; (perArticle[path] || (perArticle[path] = [])).push(s)
}
}
}
const avgMinPerVisit = visitCount
? Math.max(1, Math.round(totalVisitSeconds / visitCount / 60))
: 0
let articleMedianSum = 0
const articleCount = Object.keys(perArticle).length
for (const arr of Object.values(perArticle)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median = arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
articleMedianSum += Math.max(MIN_READ_SECONDS, median)
}
const avgArticleMedianMin = articleCount
? Math.max(1, Math.round(articleMedianSum / articleCount / 60))
: 0
return { avgMinPerVisit, avgArticleMedianMin }
}
/** Build a path -> page title lookup from the site tree. */
function buildTitleMap(pageTree) {
const titles = new Map()
const walk = (items) => {
for (const item of items || []) {
titles.set(`/${item.path}`, item.title)
walk(item.children)
}
}
walk(pageTree)
return titles
}
/** Last path segment for display; front page becomes a house icon. */
function slugOf(path) {
return path === '/' ? '🏠︎' : path.split('/').pop()
}
/** Host name of an external https origin, with scheme and www. stripped. */
function externalSlug(origin) {
try {
return new URL(origin).host.replace(/^www\./, '')
} catch {
return origin.replace(/^https?:\/\//, '').replace(/^www\./, '')
}
}
/** Origin (scheme://host) of an external https URL, for favicon lookup. */
function externalOrigin(url) {
try {
return new URL(url).origin
} catch {
return ''
}
}
/** Format one trail step: an internal page or an external https origin. */
function stepOf(path, titles) {
if (path?.startsWith('/')) {
return { path, slug: slugOf(path), title: titles.get(path) || '', external: false, home: path === '/' }
}
if (path?.startsWith('https://')) {
return {
path,
slug: externalSlug(path),
title: 'External site',
external: true,
origin: externalOrigin(path),
}
}
return null
}
/**
* Human-readable relative timestamp. Adapted from cista-storage: uses
* ``Intl.RelativeTimeFormat`` for short intervals and a compact date for
* anything older than a week.
*/
export function formatWhen(ts, now = Date.now()) {
const date = new Date(ts)
const diff = date.getTime() - now
const adiff = Math.abs(diff)
const formatter = new Intl.RelativeTimeFormat('en', { numeric: 'auto' })
if (adiff <= 5000) return 'now'
if (adiff <= 60000) {
return formatter
.format(Math.round(diff / 1000), 'second')
.replace(' ago', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 3600000) {
return formatter
.format(Math.round(diff / 60000), 'minute')
.replace('utes', '')
.replace('ute', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 86400000) {
return formatter
.format(Math.round(diff / 3600000), 'hour')
.replace('hours', 'h')
.replace('hour', 'h')
.replaceAll(' ', '\u202F')
}
if (adiff <= 604800000) {
return formatter
.format(Math.round(diff / 86400000), 'day')
.replaceAll(' ', '\u202F')
}
let d = date
.toLocaleDateString('en-ie', {
weekday: 'short',
year: 'numeric',
month: 'short',
day: 'numeric',
})
.replace('Sept', 'Sep')
if (d.length === 14) d = d.replace(' ', ' \u2007')
d = d.replaceAll(' ', '\u202F').replace('\u202F', '\u00A0')
d = d.slice(0, -4) + d.slice(-2)
return d
}
/** Full UTC timestamp for tooltips, e.g. "2026-08-21 00:20:48 UTC". */
export function formatWhenTooltip(ts) {
return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC')
}
/** Full local timestamp for tooltips, e.g. "21 Aug 2026, 17:38:48". */
export function formatWhenLocal(ts) {
return new Date(ts).toLocaleString('en-ie', {
year: 'numeric',
month: 'short',
day: 'numeric',
hour: '2-digit',
minute: '2-digit',
second: '2-digit',
})
}
/** Preserve locale case with the region/country subtag upper-cased. */
export function formatLang(value) {
if (!value || value === '—') return value
const parts = value.split('-')
if (parts.length > 1) {
parts[parts.length - 1] = parts[parts.length - 1].toUpperCase()
}
return parts.join('-')
}
/** ISO 8601 UTC timestamp without subseconds, e.g. "2026-08-21T00:20:48Z". */
export function formatWhenIso(ts) {
return `${new Date(ts).toISOString().split('.')[0]}Z`
}
/**
* Compact read time for tooltips: "50s" under a minute, "1m23s" otherwise.
*/
export function formatReadTime(seconds) {
if (seconds < 60) return `${seconds}s`
return `${Math.floor(seconds / 60)}m${seconds % 60}s`
}
/**
* Compact visitor counts: plain below 1k, then 1.2k / 10k / 1.2M.
* Truncated, not rounded.
*/
export function formatCount(n) {
if (n < 1000) return String(n)
if (n < 10000) return `${Math.trunc(n / 1000)}.${Math.trunc((n % 1000) / 100)}k`
if (n < 1_000_000) return `${Math.trunc(n / 1000)}k`
return `${Math.trunc(n / 1_000_000)}.${Math.trunc((n % 1_000_000) / 100_000)}M`
}
/**
* Format recent visits for display, newest first. Each step is a linked slug
* pointing to its article; external referers/origins are shown as their
* domain name with the full origin as the link href. The link title shows the
* article heading when known, or "External site" for origins.
*/
export function formatRecentVisits(visits, pageTree, limit = 50) {
const titles = buildTitleMap(pageTree)
return [...visits]
.reverse()
.map((v) => ({
when: new Date(v.start).toLocaleString(),
steps: [v.referer, ...Object.values(v.trail || {}).map((t) => t.to)]
.map((p) => stepOf(p, titles))
.filter(Boolean),
}))
.filter((v) => v.steps.length)
.slice(0, limit)
}
/**
* Count distinct values of a visit field, sorted most-common first.
* Returns an array of [value, count] pairs.
*/
export function countByField(visits, field) {
const counts = {}
for (const v of visits || []) {
const value = v[field]
if (!value) continue
counts[value] = (counts[value] || 0) + 1
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/**
* Count UTM parameter occurrences across visits. Each distinct
* ``parameter: value`` pair is counted separately. Returns [pair, count].
*/
export function countUtmTags(visits) {
const counts = {}
for (const v of visits || []) {
for (const [key, value] of Object.entries(v.utm || {})) {
const label = `${key}: ${value}`
counts[label] = (counts[label] || 0) + 1
}
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/** Format a list of [value, count] pairs for inline display. */
export function formatCounts(entries) {
return entries.map(([value, count]) => `${value} (${count})`).join(', ')
}
/**
* Count distinct User-Agent strings among crawler hits, most common first.
* Returns an array of [ua, count] pairs. ``clients`` maps client hashes to
* client records.
*/
export function countCrawlerUas(crawlers, clients) {
const counts = {}
for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {}
const value = client.ua_pretty || client.ua || '(no UA)'
counts[value] = (counts[value] || 0) + 1
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/**
* Reduce a reverse-DNS hostname to its right-most components that fit
* within ``limit`` characters. This keeps the meaningful main domain
* while avoiding absurdly long subdomains like ``xxx.yyy.zzz...provider.net``.
*/
export function mainDomain(host, limit = 24) {
if (!host) return host
const labels = host.split('.').filter(Boolean)
if (!labels.length) return host
const parts = [labels.pop()]
while (labels.length) {
const next = labels[labels.length - 1]
const candidate = `${next}.${parts.join('.')}`
if (candidate.length > limit) break
parts.unshift(labels.pop())
}
return parts.join('.')
}
/**
* Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by most recent hit
* first, with total hits as a tie-breaker. The group's ``refererStep`` is the
* latest external referer seen for the crawler — spiders often advertise
* their own site there — rendered with its favicon like visit referers.
* ``clients`` maps client hashes to client records.
*/
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
const groups = new Map()
for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {}
const g = groups.get(c.client) || {
clientHash: c.client,
client,
lastStart: 0,
referer: '',
pages: new Map(),
}
const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start
if (c.referer) g.referer = c.referer
if (c.entry?.startsWith('/')) {
const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
existing.count += 1
if (c.status != null) existing.status = c.status
g.pages.set(c.entry, existing)
}
groups.set(c.client, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pages.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart || totalHits(b) - totalHits(a))
.slice(0, 10)
.map((g) => {
const client = g.client || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
refererStep: stepOf(g.referer, titles),
pages: [...g.pages.entries()]
.sort((a, b) => b[1].count - a[1].count)
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical requests (same path and status class) are collapsed
* into one entry with their hit count; a path's 404 probes and its real
* (200) reads never merge.
* The paths split into two lists: ``paths`` holds the 404 probes (flagged
* paths — the ones that triggered abuse classification — first, then other
* 404s) shown verbatim, query string included, and ``articles`` holds the
* real (200) document GETs as trail steps resolved against the page tree
* (query string stripped), rendered like the visitor/crawler trails. Within
* each list paths are sorted by count descending, then earliest first.
* Rows are sorted by most recent hit first. Visitor metadata comes from the
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
* how many distinct client variations the IP produced.
* ``clients`` maps client hashes to client records.
*/
export function formatAbuseRows(abuse, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
const groups = new Map()
for (const a of abuse || []) {
const client = (clients || {})[a.client] || {}
const ip = client.ip || ''
const g = groups.get(ip) || {
ip,
pathCounts: new Map(),
clientHashes: new Set(),
lastStart: 0,
lastClient: a.client,
}
const start = new Date(a.start).getTime()
if (start > g.lastStart) {
g.lastStart = start
g.lastClient = a.client
}
const path = a.path || ''
// Collapse identical requests, but never merge a path's 404 probes with
// its real (200) reads — a page probed while missing and later created
// must show up in both columns, not flip to "articles read".
const key = `${a.is_404 ? '4' : '2'}${path}`
const existing = g.pathCounts.get(key) || {
path,
count: 0,
firstStart: start,
flag: a.flag || false,
is_404: a.is_404 || false,
}
existing.count += 1
if (start < existing.firstStart) existing.firstStart = start
if (a.flag) existing.flag = true
g.pathCounts.set(key, existing)
g.clientHashes.add(a.client)
groups.set(ip, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pathCounts.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart)
.slice(0, 10)
.map((g) => {
const all = [...g.pathCounts.values()]
const byCount = (a, b) => b.count - a.count || a.firstStart - b.firstStart
const paths = all
.filter((p) => p.flag || p.is_404)
.sort((a, b) => (a.flag ? 0 : 1) - (b.flag ? 0 : 1) || byCount(a, b))
const articles = all.filter((p) => !p.flag && !p.is_404).sort(byCount)
const pathList = (list) =>
list.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path)).join('\n')
const client = (clients || {})[g.lastClient] || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
paths: paths.map((p) => ({
path: p.path,
count: p.count,
flag: p.flag,
is_404: p.is_404,
})),
allPaths: pathList(paths),
articles: articles
.map((p) => {
const step = stepOf(p.path.split('?')[0], titles)
return step ? { ...step, count: p.count } : null
})
.filter(Boolean),
allArticles: pathList(articles),
clientCount: g.clientHashes.size,
ip: client.ip || g.ip,
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Format raw visit records as rows for a technical table. Returns objects
* with display strings; missing values become "—". ``trail`` starts with the
* external referer (when present), then the entry page and any further internal
* pages or external exit origins. Only the 20 most recent visits are shown.
* ``clients`` maps client hashes to client records.
*/
export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const client = (clients || {})[v.client] || {}
const trail = Object.values(v.trail || {})
.map((item) => {
const step = stepOf(item.to, titles)
if (step) {
if (item.read) step.readSeconds = item.read
if (item.status) step.status = item.status
}
return step
})
.filter(Boolean)
const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean)
const utm = utmValues.length ? utmValues.join(' · ') : ''
const utmTitle = Object.entries(v.utm || {})
.map(([k, value]) => `${k}=${value}`)
.join(', ')
const dash = (s) => (s || '—')
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(v.start, now),
lastSeenIso: formatWhenIso(v.start),
lastSeenLocal: formatWhenLocal(v.start),
langDisplay: formatLang(client.lang),
trail,
refererStep: stepOf(v.referer, titles),
referer: dash(v.referer),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
lang: dash(client.lang),
country: dash(client.country),
city: dash(client.city),
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
utm: utm || '—',
utmTitle,
}
})
}
+267
View File
@@ -0,0 +1,267 @@
/**
* Time ranges, week alignment and re-bucketing for analytics charts.
*
* Raw data comes as sparse 5-minute buckets; the range picks the x window
* and a coarser bucket size to keep point counts sane. The week range is
* aligned to Monday 00:00 UTC and overlays previous weeks' curves (fading
* with age), so weekly patterns compare directly.
*/
export const MIN5 = 5 * 60e3
export const HOUR = 3600e3
export const DAY = 86400e3
export const WEEK = 7 * DAY
export const RANGES = {
day: { label: 'day', span: DAY, bucket: MIN5 },
week: { label: 'week' },
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
year: { label: 'year', span: 365 * DAY, bucket: DAY },
all: { label: 'all', span: null, bucket: DAY, minSpan: 30 * DAY },
}
/** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */
export function mondayUTC(t) {
const d = Math.floor(t / DAY)
return (d - ((d + 3) % 7)) * DAY
}
/** ISO 8601 week number of the week containing t (via its Thursday). */
export function isoWeek(t) {
const d = new Date(t)
d.setUTCHours(0, 0, 0, 0)
d.setUTCDate(d.getUTCDate() + 4 - (d.getUTCDay() || 7))
const yearStart = Date.UTC(d.getUTCFullYear(), 0, 1)
return Math.ceil(((d - yearStart) / DAY + 1) / 7)
}
/** Parse sparse timestamp buckets into a { epochMs: count } map. */
export function rawTimes(buckets) {
const raw = {}
// Key by parsed timestamp: Python writes "+00:00", JS ISO uses "Z".
for (const [k, c] of Object.entries(buckets || {})) raw[Date.parse(k)] = c
return raw
}
/** Sum counts from raw 5-minute buckets between t0 (inclusive) and t1 (exclusive). */
export function sumRange(raw, t0, t1) {
let n = 0
for (let s = t0; s < t1; s += MIN5) n += raw[s] || 0
return n
}
/**
* One series per overlaid week: [this week, 1 week ago, ...], at native
* 5-minute resolution, up to 8 weeks back (and only weeks that overlap the
* recorded data at all). Each older week's timestamps are shifted forward
* onto the current week's axis so all curves overlay inside the plot.
* The current week is truncated at the current bucket
* — no fake zeroes drawn for the future. Counts are rates per hour
* (bucket count * 12): a lone visit in a 5-minute bucket reads as "12/h".
* The coarser ranges use per-day rates instead (unitMinutes = 24*60).
*/
export function weeklySeries(buckets) {
const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number)
const now = Date.now()
const thisMonday = mondayUTC(now)
if (!times.length) {
const points = []
const end = Math.min(thisMonday + WEEK, Math.floor(now / MIN5) * MIN5 + MIN5)
for (let t = thisMonday; t < end; t += MIN5) {
points.push({ t, count: 0 })
}
return {
series: [{ points, label: `Week ${isoWeek(thisMonday)}`, opacity: 1, area: true }],
t0: thisMonday,
t1: thisMonday + WEEK,
rate: HOUR / MIN5,
binMinutes: 5,
unitMinutes: 60,
unit: 'hour',
}
}
const oldest = Math.min(...times)
// Weeks back as far as the data reaches: difference in Monday indices.
const available = (thisMonday - mondayUTC(oldest)) / WEEK + 1
const count = Math.min(available, 8)
const out = []
for (let back = 0; back < count; back++) {
const start = thisMonday - back * WEEK
const end = back === 0
? Math.min(start + WEEK, Math.floor(now / MIN5) * MIN5 + MIN5)
: start + WEEK
const points = []
for (let t = start; t < end; t += MIN5) {
points.push({ t: t + back * WEEK, count: raw[t] || 0 })
}
out.push({
points,
label: `Week ${isoWeek(start)}`,
opacity: Math.max(0.15, 1 - back * 0.25),
past: back > 0,
area: back === 0,
})
}
return {
series: out,
t0: thisMonday,
t1: thisMonday + WEEK,
rate: HOUR / MIN5,
binMinutes: 5,
unitMinutes: 60,
unit: 'hour',
}
}
/**
* Rolling window for the non-week ranges (x max = now), counts converted
* to per-day rates (the unit the month+ charts are read in).
* Ranges without a fixed span use the full data reach, but never less than
* their configured minSpan so the chart keeps a readable minimum x scale.
* t0 is aligned to the UTC day so the x labels cover the whole range;
* t1 is now, so the scale never extends into the future. The bucket size
* follows the resulting window (6h up to 31 days, daily beyond), so ranges
* covering the same window — "all" at its 30-day minimum vs "month" —
* render the identical curve.
*/
export function rollingSeries(buckets, rangeKey) {
const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number)
const { span, bucket, minSpan = 0 } = RANGES[rangeKey]
const t1 = Date.now()
const t0 = Math.floor((span != null
? t1 - span
: Math.min(times.length ? Math.min(...times) : Infinity, t1 - minSpan)) / DAY) * DAY
if (!times.length) {
const bucketMs = t1 - t0 <= 31 * DAY ? Math.min(bucket, 6 * HOUR) : bucket
const points = []
for (let t = t0; t < t1; t += bucketMs) {
points.push({ t, count: 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: true }],
t0,
t1,
rate: DAY / bucketMs,
binMinutes: bucketMs / 60e3,
unitMinutes: 24 * 60,
unit: 'day',
}
}
// The bucket follows the actual window length, not the range key: when
// "all" is capped to its 30-day minimum it covers the very window "month"
// does, and daily bins would draw a different curve over the same data
// (coarser edge detection, points a day apart plotted at bin starts, the
// last point stuck at today's midnight instead of reaching now).
const bucketMs = t1 - t0 <= 31 * DAY ? Math.min(bucket, 6 * HOUR) : bucket
const points = []
for (let t = t0; t < t1; t += bucketMs) {
points.push({ t, count: sumRange(raw, t, t + bucketMs) })
}
return {
series: [{ points, label: '', opacity: 1, area: true }],
t0,
t1,
rate: DAY / bucketMs,
binMinutes: bucketMs / 60e3,
unitMinutes: 24 * 60,
unit: 'day',
}
}
/**
* Day view: raw 5-minute bucket counts for the current 24-hour window.
* No smoothing or rate conversion is applied; counts are used as-is.
*/
export function daySeries(buckets) {
const raw = rawTimes(buckets)
const now = Date.now()
const { span, bucket } = RANGES.day
const t1 = Math.floor(now / bucket) * bucket + bucket
const t0 = t1 - span
const points = []
for (let t = t0; t < t1; t += bucket) {
points.push({ t, count: raw[t] || 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: false }],
t0,
t1,
rate: 1,
binMinutes: bucket / 60e3,
unitMinutes: bucket / 60e3,
unit: '5min',
}
}
/** Dispatch to daily, weekly or rolling series based on the selected range. */
export function makeSeries(buckets, rangeKey) {
if (rangeKey === 'day') return daySeries(buckets)
if (rangeKey === 'week') return weeklySeries(buckets)
return rollingSeries(buckets, rangeKey)
}
/**
* Absolute UTC time window for a given range key. Used to filter visits,
* transitions and views for the non-chart stats on the analytics page.
* Every bounded range is a rolling span ending at now; the charts instead
* align week to Monday 00:00 UTC (overlaying previous weeks) and month+
* to UTC day boundaries, so their x windows differ from the stats range
* on purpose.
* Returns { t0, t1 } where null means unbounded.
*/
export function rangeWindow(rangeKey) {
const now = Date.now()
if (rangeKey === 'all') {
return { t0: null, t1: null }
}
const span = rangeKey === 'week' ? WEEK : RANGES[rangeKey].span
return { t0: now - span, t1: now }
}
/**
* Sum the bucketed transition matrix (from -> to -> bucket ISO -> count)
* into a plain from -> to -> count matrix for the window [t0, t1).
*/
export function filterTransitionsByRange(transitions, t0, t1) {
const out = {}
for (const [fr, tos] of Object.entries(transitions || {})) {
for (const [to, buckets] of Object.entries(tos)) {
let n = 0
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) n += c
}
if (n) {
out[fr] = out[fr] || {}
out[fr][to] = n
}
}
}
return out
}
/** Keep only the 5-minute view buckets that fall inside [t0, t1). */
export function filterViewsByRange(views, t0, t1) {
const filtered = {}
for (const [path, buckets] of Object.entries(views || {})) {
const out = {}
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out[k] = c
}
if (Object.keys(out).length) filtered[path] = out
}
return filtered
}
/** Keep only records whose start time falls inside [t0, t1). */
export function filterRecordsByRange(records, t0, t1) {
const out = []
for (const r of records || []) {
const t = Date.parse(r.start)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out.push(r)
}
return out
}
+943
View File
@@ -0,0 +1,943 @@
/**
* Radial transition map and helpers.
*
* Site map following the menu structure: top-level items in a row at the
* top (below the external source row), each item's subtree fanning out
* below it in menu order along a slightly circular downward arc. Index
* pages with no views are omitted, their children moving up in their
* place. All pages of the site are shown (from /_api/pages), plus any
* extra paths seen in transitions (deleted pages); these form their own
* top-level groups. Internal path -> path transitions join opposite
* directions into straight connections (middle width = total
* count; connectors flare into the node pills at both ends and wrap
* around their backs, surrounding them; the pills are drawn on top). Connection width grows
* logarithmically with the daily hit rate (base-2 log, one hit/day
* renders zero width, each doubling adds a fixed step, uncapped);
* connections carrying less than 1% of the total
* traffic are pruned, which naturally keeps the graph under ~100
* connections. Animated beads flow along every edge in each direction,
* emitted at time intervals inversely proportional (linear) to the
* directional count.
* External sources appear as nodes in a row above the map. Sources are
* identified from visit records in this order: utm_campaign, utm_source,
* referer, then other utm_* tags. Visits with a UTM tag are grouped under
* that tag's value, not under the referer domain. A UTM source node only
* becomes a clickable link when every visit carrying that UTM tag came
* from the same referer. External exits are full-size nodes in a row below
* the map, mirroring the source row, so the site itself stays in the
* middle. Each distinct full exit URL is its own node. Self-loops (reload
* pings) are skipped.
*/
import { MIN_READ_SECONDS, readMapOf } from './format.js'
// Nodes are constant-size pills (stadium rects) holding the slug and the
// view count on two centered lines. TNODE_BOUND is the pill's bounding
// radius, used for layout clearance and placement; connectors and flows
// use the exact outline geometry instead (pillContact below).
export const TNODE_W = 160
export const TNODE_H = 54
const TNODE_BOUND = Math.hypot(TNODE_W, TNODE_H) / 2
const PILL_R = TNODE_H / 2 // cap radius and straight-section half-height
const PILL_OFF = TNODE_W / 2 - PILL_R // x offset of the cap centers
/**
* Where the ray from a node center along (ux, uy) exits the pill outline
* (a capsule: straight top/bottom plus semicircular caps), enlarged by
* `margin`. Returns the distance `t` to the contact point and the outline
* arc position `s` of that point (see pillPointAt).
*/
const pillContact = (ux, uy, margin = 0) => {
const r = PILL_R + margin
const off = PILL_OFF + margin
const q = (Math.PI / 2) * r
// Straight top/bottom: valid when the crossing lands on the flat section.
let tf = Infinity
if (Math.abs(uy) > 1e-9) {
const t = r / Math.abs(uy)
if (Math.abs(t * ux) <= off + 1e-9) tf = t
}
// Rounded cap on the side the ray points to.
const cx = off * (ux >= 0 ? 1 : -1)
const disc = r * r - (cx * uy) ** 2
const tc = disc >= 0 ? cx * ux + Math.sqrt(disc) : Infinity
if (tf <= tc) {
const x = tf * ux
return { t: tf, s: uy > 0 ? q + off - x : q + 2 * off + Math.PI * r + x + off }
}
if (tc < Infinity) {
let th = Math.atan2(tc * uy, tc * ux - cx)
if (th < 0) th += 2 * Math.PI
const s = cx > 0
? th <= Math.PI / 2
? th * r
: q + 4 * off + Math.PI * r + (th - (3 * Math.PI) / 2) * r
: q + 2 * off + (th - Math.PI / 2) * r
return { t: tc, s }
}
return { t: TNODE_BOUND + margin, s: 0 }
}
/** Total perimeter of the (margined) pill outline. */
const pillPerimeter = (margin = 0) =>
4 * (PILL_OFF + margin) + 2 * Math.PI * (PILL_R + margin)
/**
* Point on the pill outline at arc position `s`, counterclockwise from the
* right cap tip: right cap up, top flat right-to-left, left cap down,
* bottom flat left-to-right, right cap up to the tip. Pills are never
* rotated, so the returned offset from the node center is in absolute
* coordinates.
*/
const pillPointAt = (s, margin = 0) => {
const r = PILL_R + margin
const off = PILL_OFF + margin
const P = pillPerimeter(margin)
const q = (Math.PI / 2) * r
s = ((s % P) + P) % P
if (s < q) {
const th = s / r
return [off + r * Math.cos(th), r * Math.sin(th)]
}
s -= q
if (s < 2 * off) return [off - s, r]
s -= 2 * off
if (s < Math.PI * r) {
const th = Math.PI / 2 + s / r
return [-off + r * Math.cos(th), r * Math.sin(th)]
}
s -= Math.PI * r
if (s < 2 * off) return [-off + s, -r]
s -= 2 * off
const th = (3 * Math.PI) / 2 + s / r
return [off + r * Math.cos(th), r * Math.sin(th)]
}
/** Unit tangent to the pill outline at arc position `s`, in the direction
* of increasing `s` (numeric; exact on both flats and caps). */
const pillTangent = (s, margin = 0) => {
const [x1, y1] = pillPointAt(s - 0.5, margin)
const [x2, y2] = pillPointAt(s + 0.5, margin)
const m = Math.hypot(x2 - x1, y2 - y1) || 1
return [(x2 - x1) / m, (y2 - y1) / m]
}
// Edge width: half-width = WIDTH_GROWTH * log2(daily / DAILY_REF),
// where `daily` is the connection's hit rate in hits/day (callers scale
// raw counts by DAY / range). DAILY_REF hits/day renders zero width;
// WIDTH_GROWTH is the half-width added per doubling of the rate.
// Connections whose thin middle would render below MIN_WMID are culled
// entirely (fainter strands are practically invisible), as are those
// carrying less than PRUNE_FRACTION of the total traffic (this also
// keeps the graph under ~100 connections).
const WIDTH_GROWTH = 1.4 // half-width px per doubling of the daily rate
const DAILY_REF = 0.8 // hits/day at which the width is zero
const PRUNE_FRACTION = 0.01
// ~0.8 px full width at natural size (1 viewBox unit = 1 px).
const MIN_WMID = 0.4
// Beads: each edge direction emits beads at dailyRate * BEAD_RATE beads
// per second (linear in the daily hit rate). The rate is much reduced
// from real time to keep the animation lightweight. The component
// simulates every bead independently in JS with a constant traversal
// time per edge (speed relative to span length), with no limit on beads
// in flight.
export const BEAD_R = 3.2
const BEAD_RATE = 0.0084 // beads per second per hit/day
const FLOW_OFFSET = 3 // lane offset to the right of the travel direction
const MAX_EXT_IN = 8 // referer nodes in the top row
const MAX_EXT_OUT = 12 // exit nodes in the bottom row
const EXT_GAP = 12 // vertical margin of the source/exit rows to the map
/** Flatten the site tree into navigation order via DFS. */
function buildNavigationOrder(pageTree) {
const order = new Map()
const walk = (items) => {
for (const item of items || []) {
const p = `/${item.path}`
if (!order.has(p)) order.set(p, order.size)
walk(item.children)
}
}
walk(pageTree)
return order
}
/** Map page paths to their article titles from the site tree. */
function buildTitleMap(pageTree) {
const titles = new Map()
const walk = (items) => {
for (const item of items || []) {
titles.set(`/${item.path}`, item.title)
walk(item.children)
}
}
walk(pageTree)
return titles
}
/** Extract internal page-to-page transitions, excluding self-loops. */
function collectInternalTransitions(transitions) {
const internal = []
for (const [fr, tos] of Object.entries(transitions || {})) {
if (!fr.startsWith('/')) continue
for (const [to, count] of Object.entries(tos)) {
if (to.startsWith('/') && to !== fr) internal.push({ fr, to, count })
}
}
return internal
}
/** Domain-only label for an external origin (path and www. removed).
* Pills clip the text at their border; no length cap needed. */
function extLabel(ext) {
try {
return new URL(ext).hostname.replace(/^www\./, '')
} catch {
return ext.replace(/^https?:\/\//, '').replace(/^www\./, '').split('/')[0]
}
}
/**
* Collect outgoing external transitions: page path -> full exit URL.
* Aggregated per (URL, page) pair. Incoming external links are now derived
* from visit records (which carry UTM tags), so only exits remain here.
*/
function collectExitPairs(transitions) {
const pairs = new Map() // `${ext} ${page}` -> {ext, page, out}
for (const [fr, tos] of Object.entries(transitions || {})) {
if (!fr.startsWith('/')) continue // ignore external -> anything
for (const [to, count] of Object.entries(tos)) {
if (!to.startsWith('http')) continue
const k = `${to} ${fr}`
const p = pairs.get(k) || { ext: to, page: fr, out: 0 }
p.out += count
pairs.set(k, p)
}
}
return [...pairs.values()]
}
/** Build nodes with depth and a path lookup map; children are wired to parents. */
function buildNodeTree(internal, navOrder) {
const paths = new Set(['/', ...navOrder.keys()])
for (const e of internal) { paths.add(e.fr); paths.add(e.to) }
const depth = (p) => (p === '/' ? 0 : p.split('/').length - 1)
const nodes = [...paths].map((p) => ({
path: p, depth: depth(p), angle: 0, children: [],
}))
const byPath = new Map(nodes.map((n) => [n.path, n]))
// Parent is the nearest ancestor present in the map, front page last.
const parentOf = (p) => {
let q = p
while (q !== '/') {
q = q.slice(0, q.lastIndexOf('/')) || '/'
if (byPath.has(q)) return byPath.get(q)
}
return byPath.get('/')
}
for (const n of nodes) {
if (n.path !== '/') parentOf(n.path).children.push(n)
}
return { nodes, byPath, root: byPath.get('/') }
}
/** Sort each node's children by navigation order, recursively. */
function sortByNav(root, navOrder) {
const byNav = (a, b) =>
(navOrder.get(a.path) ?? Infinity) - (navOrder.get(b.path) ?? Infinity)
|| a.path.localeCompare(b.path)
const walk = (n) => {
n.children.sort(byNav)
n.children.forEach(walk)
}
walk(root)
}
/** Compute median reading time per article in seconds. */
function buildReadSeconds(visits) {
const times = {}
for (const v of visits || []) {
for (const [path, sec] of Object.entries(readMapOf(v))) {
if (sec >= MIN_READ_SECONDS) {
; (times[path] || (times[path] = [])).push(sec)
}
}
}
const seconds = {}
for (const [path, arr] of Object.entries(times)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median =
arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
seconds[path] = Math.round(median)
}
return seconds
}
/** Compute view counts, labels and hidden flags for each node. */
function annotateNodes(nodes, viewsData, titles, readSeconds) {
const viewCount = (p) => {
let n = 0
for (const c of Object.values(viewsData?.[p] || {})) n += c
return n
}
for (const n of nodes) {
n.views = viewCount(n.path)
n.readSec = readSeconds[n.path] || 0
// Article title inside the pill (clipped at the pill border on
// render), slug as fallback for pages missing from the site tree.
n.label = titles.get(n.path) || (n.path === '/' ? '🏠︎' : n.path.split('/').pop())
n.title = titles.get(n.path) || ''
// Category (non-leaf) pages with no views in this window are omitted:
// their children move up in their place (see layoutGroups).
n.hidden = n.children.length > 0 && n.views === 0
}
}
/**
* Top-down layout following the menu structure: top-level items in an
* equally spaced row at the top (right below the external source row),
* the row following a shallow circular sag (center lowest) so connections
* between neighbors do not overlap the pills in between. Each top item's
* whole subtree fans out from it in menu (DFS preorder) order along a
* large-radius circular arc that leaves the parent heading straight down
* and gradually bends to the right — no horizontal space is reserved for fans, they
* extend under the slots to their right. Hidden index pages are omitted
* from the fan; when the top item itself is hidden, the fan shifts one
* slot up, the first visible child taking the top position. Branch lanes
* labeled with the branch slug (see the branch-lane pass at the end)
* keep the omitted menu levels visible.
*/
function layoutGroups(root) {
// Top slots are spaced well over one pill width apart regardless of
// fan sizes.
const SLOT = TNODE_W + 100
const CLEAR = TNODE_W * 0.8 // fan spacing per member along the curve
// First pass: visible members per group, in menu order. Hidden index
// pages are skipped, but their children still appear. The front page
// forms its own group. groupRoots keeps each group's subtree root for
// the branch-curve pass below.
const groups = []
const groupRoots = []
for (const g of [root, ...root.children]) {
const members = []
if (g === root) {
if (!g.hidden) members.push(g)
} else {
const walk = (n) => {
if (!n.hidden) members.push(n)
n.children.forEach(walk)
}
walk(g)
}
if (members.length) {
groups.push(members)
groupRoots.push(g)
}
}
// Top row on a large-radius circular arc whose bottom point is the
// LAST top item: each earlier item sits a bit higher (drop = 15% of
// the row span). Flat row when there is a single group.
const half = ((groups.length - 1) * SLOT) / 2 || 1
const span = (groups.length - 1) * SLOT
const topD = span * 0.15
const R_T = span ? (span * span + topD * topD) / (2 * topD) : 0
const topY = span
? (x) => topD - R_T + Math.sqrt(R_T * R_T - (x - half) * (x - half))
: () => 0
// Second pass: place groups. Fan members follow a circular arc of
// large radius FAN_R centered at (gx + FAN_R, y0): the trail leaves
// the top node heading straight down (vertical tangent) and bends
// right gently, member i at arc angle π i·CLEAR/FAN_R (spaced by
// arc length CLEAR). A circle — not a spline — so the branch lanes
// below can be concentric arcs: identical forms, only radii differ.
const FAN_R = 1000
groups.forEach((members, gi) => {
const gx = gi * SLOT - half
const y0 = topY(gx)
members[0].x = gx
members[0].y = y0
for (let i = 1; i < members.length; i++) {
const th = Math.PI - (i * CLEAR) / FAN_R
members[i].x = gx + FAN_R * (1 + Math.cos(th))
members[i].y = y0 + FAN_R * Math.sin(th)
}
})
// Branch lanes: one wide arc per path prefix (slug depth ≥ 1) whose
// subtree holds at least two visible nodes (a branch's visible nodes
// form one contiguous run in the fan's DFS preorder). Every lane of a
// group is an arc around the group's fan center with a radius one
// INDENT larger per parent level — concentric circles, so all lanes
// share exactly one form. Lanes span their branch's nodes plus a
// little extra tucked under the first/last pill (so the line caps are
// never visible) and run behind the pills. A label arc carries the
// branch slug, left-aligned just past the first pill and free to run
// to the lane's end — longer text simply passes under later pills,
// which are drawn on top. Hidden (unplaced) index pages still
// define a lane: it follows their promoted children, so lanes reflect
// the path structure rather than page existence.
const INDENT = 20 // lane spacing (radius) per nesting level (> lane width)
const END_TUCK = 22 // arc units tucked under the first/last pill
const GAP_TRIM = 32 // label arc clearance from the pills
// Labels are left-aligned on their guide: the guide starts just past the
// source pill's edge, the earliest point where the text is visible.
const LABEL_PAD = 6
// The label guide rides GUIDE_OFF outward of the lane centerline: the
// text's alphabetic baseline sits on the guide, so this puts the
// glyph middle (not the baseline) on the lane center at any zoom —
// dominant-baseline tricks are em-based and break under downscale.
const GUIDE_OFF = 3.5
const branches = []
groups.forEach((members, gi) => {
const g = groupRoots[gi]
if (g === root || members.length < 2) return
const idx = new Map(members.map((n, i) => [n, i]))
const C = [gi * SLOT - half + FAN_R, topY(gi * SLOT - half)]
const walk = (n) => {
let first = Infinity
let last = -1
const span = (m) => {
const k = idx.get(m)
if (k !== undefined) {
first = Math.min(first, k)
last = Math.max(last, k)
}
m.children.forEach(span)
}
span(n)
if (n.depth >= 1 && last > first) {
branches.push({ depth: n.depth, name: n.path.split('/').pop(), C, first, last })
}
n.children.forEach(walk)
}
walk(g)
})
const depthMax = branches.reduce((d, b) => Math.max(d, b.depth), 1)
let arcLeft = Infinity // leftmost lane point, for the bounding box
const arcs = branches.map(({ depth, name, C, first, last }) => {
const R = FAN_R + (depthMax - depth) * INDENT
const th = (i) => Math.PI - (i * CLEAR) / FAN_R
const pt = (a, r) => [C[0] + r * Math.cos(a), C[1] + r * Math.sin(a)]
// Arc from angle a down to angle b (a > b; visually counterclockwise
// from the west point downward, hence sweep flag 0).
const arc = (a, b, r) => {
const [x0, y0] = pt(a, r)
const [x1, y1] = pt(b, r)
return `M ${x0.toFixed(2)} ${y0.toFixed(2)} A ${r.toFixed(2)} ${r.toFixed(2)} 0 0 0 ${x1.toFixed(2)} ${y1.toFixed(2)}`
}
const d = arc(th(first) + END_TUCK / R, th(last) - END_TUCK / R, R)
// Start just past the first pill: lanes leave the source node nearly
// vertically, so the pill's extent along the arc is its half height.
// The guide runs to the lane's end so long slugs are never cut off.
const ld = arc(th(first) - (TNODE_H / 2 + LABEL_PAD) / R,
th(last) - END_TUCK / R, R + GUIDE_OFF)
arcLeft = Math.min(arcLeft, pt(th(first) + END_TUCK / R, R)[0])
return { d, ld, label: name }
})
// Top lane: an arc along the top row's own circle, connecting
// the top nodes of all groups and tucked under the first and last of
// them (the arc bottoms at the last item, so it continues rightward
// under its pill). Drawn 50% thicker than branch lanes. A 🏠︎ label
// marks the lane right after the home pill, on a guide arc like
// the branch labels but with the offset and clearance scaled up by the
// same 50% to keep the glyph centered on the wider lane.
if (span) {
const d = `M ${(-half - END_TUCK).toFixed(2)} ${topY(-half - END_TUCK).toFixed(2)} `
+ `A ${R_T.toFixed(2)} ${R_T.toFixed(2)} 0 0 0 ${(half + END_TUCK).toFixed(2)} ${topY(half + END_TUCK).toFixed(2)}`
const rG = R_T + GUIDE_OFF * 1.5
const ptG = (x) => [x, topD - R_T + Math.sqrt(rG * rG - (x - half) ** 2)]
// Left-aligned like the branch labels: the guide starts just past the
// home pill's edge (scaled with the lane thickness).
const g0 = -half + TNODE_W / 2 + LABEL_PAD * 1.5
const g1 = SLOT - half - TNODE_W / 2 - GAP_TRIM * 1.5
const [gx0, gy0] = ptG(g0)
const [gx1, gy1] = ptG(g1)
arcs.unshift({
d,
ld: `M ${gx0.toFixed(2)} ${gy0.toFixed(2)} A ${rG.toFixed(2)} ${rG.toFixed(2)} 0 0 0 ${gx1.toFixed(2)} ${gy1.toFixed(2)}`,
label: '🏠︎',
top: true,
})
}
return { arcs, arcLeft }
}
/** Collapse opposite transition directions into one unordered pair per page pair. */
function aggregatePairs(internal) {
const pairs = new Map() // unordered pair key -> [countAB, countBA]
for (const e of internal) {
const forward = e.fr < e.to
const k = forward ? `${e.fr} ${e.to}` : `${e.to} ${e.fr}`
const c = pairs.get(k) || [0, 0]
c[forward ? 0 : 1] += e.count
pairs.set(k, c)
}
return pairs
}
const fmtPt = (p) => `${p[0].toFixed(2)} ${p[1].toFixed(2)}`
/**
* Build one ribbon edge between two nodes with counts ab and ba.
* `wMid` is the half-width of the thin middle (already strength-scaled by
* the caller). Each end flares into the node's pill surround (the outline
* enlarged by margin S): the flare contact points follow the pill outline
* a constant arc distance to each side of the direct contact point, and
* the back of the ribbon wraps all the way around the pill between them,
* surrounding the node. The pills themselves are drawn on top.
*/
function buildRibbon(a, b, ab, ba, wMid, external = false) {
const count = ab + ba
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len
const uy = (b.y - a.y) / len
const nx = -uy
const ny = ux
// Direct contact: where the centerline exits each pill's surround.
const S = 4
const cA = pillContact(ux, uy, S)
const cB = pillContact(-ux, -uy, S)
// Flares take a fair share of the free span while leaving the
// count-scaled thin middle a visible share of the connection length.
// The maximum flare length scales with the contact distance so wide
// approach angles still show a wide connector end.
const free = Math.max(0, len - cA.t - cB.t)
const FLARE = Math.min(Math.max(cA.t, cB.t) * 1.2, free * 0.4)
// Flare endpoints: walk the outline a constant arc distance to each
// side of the direct contact point (spanning flats and caps alike).
const D = (Math.PI / 4) * (PILL_R + S)
// Per node: endpoints for the +n (left) and -n (right) flare sides,
// each with its arc position, absolute point, and an outline tangent
// oriented back toward the direct contact point.
const ends = (cx, cy, contact) => {
const pick = (s) => {
const [px, py] = pillPointAt(s, S)
// Outline tangent oriented back toward the direct contact point
// (the flare side sweeps from the contact point around to its
// endpoint and into the connection), so it can never fork outward.
const tan = pillTangent(s, S)
if (s > contact.s) { tan[0] = -tan[0]; tan[1] = -tan[1] }
return { s, p: [cx + px, cy + py], tan, side: px * nx + py * ny }
}
const plus = pick(contact.s + D)
const minus = pick(contact.s - D)
return plus.side >= 0 ? [plus, minus] : [minus, plus]
}
const [aLeftEnd, aRightEnd] = ends(a.x, a.y, cA)
const [bLeftEnd, bRightEnd] = ends(b.x, b.y, cB)
// Point on the connection centerline at distance t from A, offset s
// perpendicular to it.
const P = (t, s) => [
a.x + t * ux + s * nx,
a.y + t * uy + s * ny,
]
// One side of a flare: from the outline endpoint, leaving tangent to
// the pill outline, to the connection middle arriving parallel with
// the centerline. The tangent pull is clamped so the control point
// stays well on its own side of the centerline — otherwise a long
// flare on a rounded cap crosses the opposite side.
const flarePoints = (end, midT, s, dir) => {
let hEnd = FLARE * 0.65
const hMid = FLARE * 0.4
const tanS = end.tan[0] * nx + end.tan[1] * ny // inward rate
if (tanS * end.side < 0) {
hEnd = Math.min(hEnd, (Math.abs(end.side) * 0.6) / Math.abs(tanS))
}
return {
pEnd: end.p,
cEnd: [end.p[0] + end.tan[0] * hEnd, end.p[1] + end.tan[1] * hEnd],
cMid: P(midT - dir * hMid, s * wMid),
pMid: P(midT, s * wMid),
}
}
// Emit a cubic in either traversal direction. Reversing a cubic requires
// swapping its control points, rather than recalculating the geometry.
const curve = (f, reverse = false) => {
if (!reverse) {
return `C ${fmtPt(f.cEnd)} ${fmtPt(f.cMid)} ${fmtPt(f.pMid)} `
}
return `C ${fmtPt(f.cMid)} ${fmtPt(f.cEnd)} ${fmtPt(f.pEnd)} `
}
// Trace the surround outline the long way around (behind the node) from
// arc s1 to arc s2. Sampled as a polyline: the visible result is a thin
// halo hugging the pill, so exact arc segments are unnecessary.
const outlineWrap = (cx, cy, s1, s2) => {
const per = pillPerimeter(S)
const dPlus = ((s2 - s1) % per + per) % per
const total = dPlus > per / 2 ? dPlus : per - dPlus
const dir = dPlus > per / 2 ? 1 : -1
const n = Math.max(4, Math.ceil(total / 6))
let out = ''
for (let i = 1; i <= n; i++) {
const [x, y] = pillPointAt(s1 + (dir * total * i) / n, S)
out += `L ${(cx + x).toFixed(2)} ${(cy + y).toFixed(2)} `
}
return out
}
const aLeft = flarePoints(aLeftEnd, cA.t + FLARE, 1, 1)
const bLeft = flarePoints(bLeftEnd, len - cB.t - FLARE, 1, -1)
const bRight = flarePoints(bRightEnd, len - cB.t - FLARE, -1, -1)
const aRight = flarePoints(aRightEnd, cA.t + FLARE, -1, 1)
// Each end wraps the full back of the node pill between its two flare
// contact points (bLeft -> bRight around B, aRight -> aLeft around A).
const d = `M ${fmtPt(aLeft.pEnd)} `
+ curve(aLeft)
+ `L ${fmtPt(bLeft.pMid)} `
+ curve(bLeft, true)
+ outlineWrap(b.x, b.y, bLeftEnd.s, bRightEnd.s)
+ curve(bRight)
+ `L ${fmtPt(aRight.pMid)} `
+ curve(aRight, true)
+ outlineWrap(a.x, a.y, aRightEnd.s, aLeftEnd.s)
+ 'Z'
return {
d,
title: `${a.path}${b.path}: ${count} (${ab} / ${ba})`,
external,
}
}
/**
* Flow descriptors for the bead animation, one per edge direction with a
* nonzero count: a straight segment running from inside the source node
* to inside the target node (beads render under the node pills, so
* they emerge from and vanish beneath the nodes rather than popping in
* at the surround), plus the emission interval (seconds between beads,
* inverse of the daily hit rate * BEAD_RATE). Each segment is offset to the
* right-hand side of its travel direction, so opposing flows on the same
* edge run on parallel lanes instead of colliding. The component turns
* these into independently simulated beads.
*/
function buildFlows(a, b, ab, ba, dayScale = 1) {
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len
const uy = (b.y - a.y) / len
const rA = pillContact(ux, uy).t
const rB = pillContact(-ux, -uy).t
const t0 = rA / 3
const t1 = len - rB / 3
if (t1 - t0 < 12) return []
// Unit normal pointing to the visual right of the A -> B direction.
const rx = -uy
const ry = ux
const span = t1 - t0
const flow = (count, fromT, toT) => {
// Each direction shifts to its own right, away from the opposing lane.
const s = fromT < toT ? FLOW_OFFSET : -FLOW_OFFSET
return {
x1: a.x + fromT * ux + s * rx,
y1: a.y + fromT * uy + s * ry,
x2: a.x + toT * ux + s * rx,
y2: a.y + toT * uy + s * ry,
len: span,
interval: 1 / (count * BEAD_RATE * dayScale),
}
}
const flows = []
// Stable key per edge direction so the component's bead simulation can
// match flows across data reloads and keep bead phases/positions.
if (ab) flows.push({ ...flow(ab, t0, t1), key: `${a.path} ${b.path}` })
if (ba) flows.push({ ...flow(ba, t1, t0), key: `${b.path} ${a.path}` })
return flows
}
/**
* Half-width for a connection middle: base-2 logarithmic in the daily
* hit rate, zero at DAILY_REF hits/day, uncapped. Absolute on purpose —
* cool routes stay visible regardless of how hot the hottest connection
* is. Callers cull results below MIN_WMID.
*/
const scaledWidth = (daily) => {
if (daily <= 0) return 0
return WIDTH_GROWTH * Math.log2(daily / DAILY_REF)
}
/**
* Build ribbon edges and bead flows for every aggregated page-to-page
* pair. Pairs carrying less than PRUNE_FRACTION of the total internal
* traffic are pruned (this naturally bounds the graph to ~100 edges).
*/
function buildInternalEdges(pairs, byPath, dayScale = 1) {
let total = 0
for (const [, [ab, ba]] of pairs) total += ab + ba
const minCount = total * PRUNE_FRACTION
const edges = []
const flows = []
for (const [k, [ab, ba]] of pairs) {
if (ab + ba < minCount) continue
const [pf, pt] = k.split(' ')
const a = byPath.get(pf)
const b = byPath.get(pt)
if (a.hidden || b.hidden) continue // unplaced index pages are omitted
const wMid = scaledWidth((ab + ba) * dayScale)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(a, b, ab, ba, wMid))
flows.push(...buildFlows(a, b, ab, ba, dayScale))
}
return { edges, flows }
}
const UTM_PRIORITY = ['utm_campaign', 'utm_source']
const UTM_FALLBACK = ['utm_medium', 'utm_content', 'utm_term', 'utm_id']
/** Identify the source of a visit according to the requested priority. */
function identifySource(visit) {
const utm = visit.utm || {}
for (const k of UTM_PRIORITY) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
if (visit.referer?.startsWith('http')) {
return { value: visit.referer, isUtm: false }
}
for (const k of UTM_FALLBACK) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
return null
}
/**
* Collect source -> entry page pairs from visit records. Sources are
* identified by UTM campaign/source (then referer, then other UTM tags).
* A UTM source only gets a link href when every visit using that source
* came from the same referer; referer sources always link to their origin.
*/
function collectSourcePairs(visits) {
const groups = new Map() // `${source}\0${page}` -> pair
for (const v of visits || []) {
const src = identifySource(v)
if (!src) continue
const k = `${src.value}\0${v.entry}`
const p = groups.get(k) || {
source: src.value,
page: v.entry,
in: 0,
refs: new Set(),
missingRef: false,
href: null,
isUtm: src.isUtm,
}
p.in += 1
if (v.referer?.startsWith('http')) {
p.refs.add(v.referer)
} else {
p.missingRef = true
}
groups.set(k, p)
}
for (const p of groups.values()) {
if (p.isUtm && !p.missingRef && p.refs.size === 1) {
const ref = [...p.refs][0]
if (ref.startsWith('http')) p.href = ref
} else if (!p.isUtm && p.source.startsWith('http')) {
p.href = p.source
}
}
return [...groups.values()]
}
/**
* Place external source and exit nodes and build their edges and bead
* flows.
* Sources (incoming links) are derived from visit UTM/referer data and form
* a row centered above the map, hottest first; exits come from the
* transition matrix and form a matching row centered below the map, so
* the site itself stays in the middle. Both rows sit EXT_GAP beyond the
* map's bounds.
* Widths and pruning use the same log scale and traffic-share rule as
* internal connections.
*/
function buildExternal({ sources, exits }, byPath, innerBounds, dayScale = 1) {
const extNodes = []
const edges = []
const flows = []
let extTotal = 0
for (const p of sources) extTotal += p.in
for (const p of exits) extTotal += p.out
const minCount = extTotal * PRUNE_FRACTION
const liveSources = sources.filter((p) => byPath.has(p.page))
const liveExits = exits.filter((p) => byPath.has(p.page))
if (!liveSources.length && !liveExits.length) return { extNodes, edges, flows }
const width = (count) => scaledWidth(count * dayScale)
// Incoming: one source node per identified source, in a row centered
// above the map, with an edge to each page that source led to. A source
// whose connectors are all culled (below MIN_WMID) is dropped itself.
const bySource = new Map() // source -> pairs, sorted by total incoming count
for (const p of liveSources.filter((p) => p.in >= minCount)) {
const g = bySource.get(p.source) || []
g.push(p)
bySource.set(p.source, g)
}
const origins = [...bySource]
.map(([source, ps]) => ({
source,
ps,
total: ps.reduce((s, p) => s + p.in, 0),
href: ps[0].href,
isUtm: ps[0].isUtm,
}))
.sort((a, b) => b.total - a.total)
.slice(0, MAX_EXT_IN)
.filter(({ ps }) =>
ps.some((p) => !byPath.get(p.page).hidden && width(p.in) >= MIN_WMID))
if (origins.length) {
const cx = (innerBounds.x0 + innerBounds.x1) / 2
const y = innerBounds.y0 - TNODE_BOUND - EXT_GAP
const spacing = TNODE_W + 44
const x0 = cx - ((origins.length - 1) * spacing) / 2
origins.forEach(({ source, ps, total, href, isUtm }, i) => {
const label = isUtm ? source : extLabel(source)
const xn = {
path: source,
href,
label, // clipped at the pill border on render
x: x0 + i * spacing,
y,
count: total,
kind: 'source',
}
extNodes.push(xn)
for (const p of ps) {
const page = byPath.get(p.page)
if (page.hidden) continue
const wMid = width(p.in)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(xn, page, p.in, 0, wMid, true))
flows.push(...buildFlows(xn, page, p.in, 0, dayScale))
}
})
}
// Outgoing: one exit node per distinct full URL (so several links to
// the same domain stay distinct), showing the total count across all
// pages linking to it, in a row centered below the map (hottest
// first), mirroring the source row above. Each (URL, page) pair
// contributes an edge from that page. An exit whose connectors are all
// culled (below MIN_WMID) is dropped itself.
const byExt = new Map() // full URL -> { ext, out, pairs }
for (const p of liveExits.filter((p) => p.out >= minCount)) {
const g = byExt.get(p.ext) || { ext: p.ext, out: 0, pairs: [] }
g.out += p.out
g.pairs.push(p)
byExt.set(p.ext, g)
}
const targets = [...byExt.values()]
.sort((a, b) => b.out - a.out)
.slice(0, MAX_EXT_OUT)
.filter(({ pairs }) =>
pairs.some((p) => !byPath.get(p.page).hidden && width(p.out) >= MIN_WMID))
if (targets.length) {
const cx = (innerBounds.x0 + innerBounds.x1) / 2
const y = innerBounds.y1 + TNODE_BOUND + EXT_GAP
const spacing = TNODE_W + 44
const x0 = cx - ((targets.length - 1) * spacing) / 2
targets.forEach(({ ext, out, pairs }, i) => {
const xn = {
path: ext,
href: ext,
label: extLabel(ext),
x: x0 + i * spacing,
y,
count: out,
kind: 'exit',
}
extNodes.push(xn)
for (const p of pairs) {
const page = byPath.get(p.page)
if (page.hidden) continue
const wMid = width(p.out)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(page, xn, p.out, 0, wMid, true))
flows.push(...buildFlows(page, xn, p.out, 0, dayScale))
}
})
}
return { extNodes, edges, flows }
}
/**
* Build the transition map model.
* Returns { nodes, edges, flows, extNodes, arcs, bounds } or null when
* there is nothing to show. `arcs` holds the branch curves; `nodes` only
* contains placed (visible) nodes. `dayScale` converts raw counts to a
* daily hit rate (DAY / range ms) for edge widths and bead rates.
*/
export function buildTransitionGraph(data, pageTree, visits = [], dayScale = 1) {
const internal = collectInternalTransitions(data?.transitions)
const sources = collectSourcePairs(visits)
const exits = collectExitPairs(data?.transitions)
const navOrder = buildNavigationOrder(pageTree)
const titles = buildTitleMap(pageTree)
const readSeconds = buildReadSeconds(visits)
if (!internal.length && !navOrder.size) return null
const { nodes, byPath, root } = buildNodeTree(internal, navOrder)
sortByNav(root, navOrder)
annotateNodes(nodes, data?.views, titles, readSeconds)
const { arcs, arcLeft } = layoutGroups(root)
const placed = nodes.filter((n) => !n.hidden)
const pairs = aggregatePairs(internal)
const { edges, flows } = buildInternalEdges(pairs, byPath, dayScale)
// Tight bounding box of the placed page nodes, extended to cover the
// branch curves running left of the pills; external nodes extend it.
// Margins cover just the ribbon surround (pill outline + flare margin
// S=4) plus a small pad: the pill's half width/height, not its diagonal
// radius, so the map crops tight especially at top and bottom.
const pad = 8
const MX = TNODE_W / 2 + 4 + pad
const MY = TNODE_H / 2 + 4 + pad
const xs = placed.map((n) => n.x)
const ys = placed.map((n) => n.y)
const bounds = {
x0: Math.min(Math.min(...xs) - MX, arcLeft - pad),
y0: Math.min(...ys) - MY,
x1: Math.max(...xs) + MX,
y1: Math.max(...ys) + MY,
}
const ext = buildExternal({ sources, exits }, byPath, bounds, dayScale)
for (const xn of ext.extNodes) {
bounds.x0 = Math.min(bounds.x0, xn.x - MX)
bounds.y0 = Math.min(bounds.y0, xn.y - MY)
bounds.x1 = Math.max(bounds.x1, xn.x + MX)
bounds.y1 = Math.max(bounds.y1, xn.y + MY)
}
return {
nodes: placed,
edges: [...edges, ...ext.edges],
flows: [...flows, ...ext.flows],
extNodes: ext.extNodes,
arcs,
bounds,
}
}
File diff suppressed because it is too large Load Diff
+2 -2
View File
@@ -25,12 +25,12 @@ pre code .cp { color: var(--code-comment); font-weight: bold; font-style: italic
pre code .cpf { color: var(--code-comment); font-style: italic } /* Comment.PreprocFile */ pre code .cpf { color: var(--code-comment); font-style: italic } /* Comment.PreprocFile */
pre code .c1 { color: var(--code-comment); font-style: italic } /* Comment.Single */ pre code .c1 { color: var(--code-comment); font-style: italic } /* Comment.Single */
pre code .cs { color: var(--code-comment); font-weight: bold; font-style: italic } /* Comment.Special */ pre code .cs { color: var(--code-comment); font-weight: bold; font-style: italic } /* Comment.Special */
pre code .gd { color: var(--code-error); background-color: color-mix(in oklab, var(--code-error) 25%, var(--code-bg)) } /* Generic.Deleted */ pre code .gd { color: var(--code-error); background-color: color-mix(var(--code-error) 25%, var(--code-bg)) } /* Generic.Deleted */
pre code .ge { color: var(--code-text); font-style: italic } /* Generic.Emph */ pre code .ge { color: var(--code-text); font-style: italic } /* Generic.Emph */
pre code .ges { color: var(--code-text); font-weight: bold; font-style: italic } /* Generic.EmphStrong */ pre code .ges { color: var(--code-text); font-weight: bold; font-style: italic } /* Generic.EmphStrong */
pre code .gr { color: var(--code-error) } /* Generic.Error */ pre code .gr { color: var(--code-error) } /* Generic.Error */
pre code .gh { color: var(--code-builtin); font-weight: bold } /* Generic.Heading */ pre code .gh { color: var(--code-builtin); font-weight: bold } /* Generic.Heading */
pre code .gi { color: var(--code-added); background-color: color-mix(in oklab, var(--code-added) 25%, var(--code-bg)) } /* Generic.Inserted */ pre code .gi { color: var(--code-added); background-color: color-mix(var(--code-added) 25%, var(--code-bg)) } /* Generic.Inserted */
pre code .go { color: var(--code-muted) } /* Generic.Output */ pre code .go { color: var(--code-muted) } /* Generic.Output */
pre code .gp { color: var(--code-muted) } /* Generic.Prompt */ pre code .gp { color: var(--code-muted) } /* Generic.Prompt */
pre code .gs { color: var(--code-text); font-weight: bold } /* Generic.Strong */ pre code .gs { color: var(--code-text); font-weight: bold } /* Generic.Strong */
+33 -4
View File
@@ -8,21 +8,50 @@ import { tags } from '@lezer/highlight'
// The base theme sets monospace on .cm-scroller, so the font must be set // The base theme sets monospace on .cm-scroller, so the font must be set
// there, not on "&". // there, not on "&".
export const cmTheme = EditorView.theme({ const cmEditorTheme = EditorView.theme({
"&": { "&": {
backgroundColor: "var(--bg)", backgroundColor: "var(--bg)",
color: "var(--text)", color: "var(--text)",
}, },
".cm-scroller": { fontFamily: '"Fira Code", monospace' }, ".cm-scroller": { fontFamily: '"Fira Code", monospace' },
".cm-content": { caretColor: "var(--text)" }, // Fira Code in CodeMirror: set the font on .cm-content (not only the
// scroller) and force every span inside to inherit it, so highlighting
// spans can't drift to a different font/metrics. Ligatures are disabled
// entirely — CodeMirror measures per character, and ligature glyphs
// render wider than the measured sum of their parts.
".cm-content": {
caretColor: "var(--text)",
fontFamily: '"Fira Code", monospace',
fontVariantLigatures: "none",
fontFeatureSettings: '"calt" 0',
letterSpacing: "normal",
},
".cm-content *": {
fontFamily: "inherit",
letterSpacing: "inherit",
},
".cm-cursor": { borderLeftColor: "var(--text)" }, ".cm-cursor": { borderLeftColor: "var(--text)" },
// basicSetup's active-line highlight assumes a dark theme. // basicSetup's active-line highlight assumes a dark theme.
".cm-activeLine": { backgroundColor: "transparent" }, ".cm-activeLine": { backgroundColor: "transparent" },
"&.cm-focused .cm-selectionBackground, .cm-selectionBackground":
{ backgroundColor: "var(--line)" },
"&.cm-focused": { outline: "none" }, "&.cm-focused": { outline: "none" },
}) })
// Selection color needs a baseTheme: only base themes support the
// &light/&dark selectors, and @codemirror/view's own selection rules use
// them — we must match its selectors exactly (equal specificity) and rely
// on mounting later to win. Focused: the page's --selection-bg (the base
// accents tint; themes may override it). Unfocused: hidden, like a normal
// input (CodeMirror greys it by default).
const cmSelection = EditorView.baseTheme({
"&light .cm-selectionBackground, &dark .cm-selectionBackground":
{ backgroundColor: "transparent" },
"&light.cm-focused > .cm-scroller > .cm-selectionLayer .cm-selectionBackground, &dark.cm-focused > .cm-scroller > .cm-selectionLayer .cm-selectionBackground":
{ backgroundColor: "var(--selection-bg)" },
})
// Exported as one extension so the editors just list `cmTheme`.
export const cmTheme = [cmEditorTheme, cmSelection]
export const cmHighlight = syntaxHighlighting(HighlightStyle.define([ export const cmHighlight = syntaxHighlighting(HighlightStyle.define([
{ tag: tags.heading, fontWeight: "600", color: "var(--accent)" }, { tag: tags.heading, fontWeight: "600", color: "var(--accent)" },
{ tag: tags.strong, fontWeight: "700" }, { tag: tags.strong, fontWeight: "700" },
+24
View File
@@ -0,0 +1,24 @@
// Shared popup open-state behavior: while `open` (a ref, truthy = open)
// is set, a pointerdown outside `root` (a template ref covering both the
// toggle button and the popup) or Escape resets it to null. One logic for
// every dropdown (LangSelect, the page editor's class/table pickers), so
// they can't drift apart.
import { onBeforeUnmount, watch } from 'vue'
export function usePopup(open, root) {
let off = null
const stop = watch(open, (v) => {
off?.()
off = null
if (!v) return
const down = (ev) => { if (!root.value?.contains(ev.target)) open.value = null }
const key = (ev) => { if (ev.key === 'Escape') open.value = null }
addEventListener('pointerdown', down, true)
addEventListener('keydown', key)
off = () => {
removeEventListener('pointerdown', down, true)
removeEventListener('keydown', key)
}
})
onBeforeUnmount(() => { off?.(); stop() })
}
+20
View File
@@ -0,0 +1,20 @@
// The editor shell's shared language selection ('' = the primary language):
// backed by the app-wide store (./store), so the editor tabs' LangSelects
// and the public corner selector bind the same value. Linked to the
// whole-page language: while the panel is open it drives the page preview
// (EditorShell applies it as the fetch-time language override, swapdoc).
import { computed, ref } from 'vue'
import { pinia, useStore } from './store'
export const editorLang = computed({
get: () => useStore(pinia).lang,
set: (v) => { useStore(pinia).lang = v },
})
// The CURRENT PAGE's primary language ('' = not yet learned): the shell's
// settings fetch fills it with the site default; the page/structure tabs
// then refine it per page (doc accept / tree rows — strictly better
// sources, so they overwrite freely while the settings fetch only fills
// the unknown). EditorShell pins the preview by it when the selection is
// '' (the primary).
export const pagePrimary = ref('')
+57
View File
@@ -0,0 +1,57 @@
// Language helpers shared by the editors (the PageEditor language picker,
// the localization settings tab). Flags come from the country-flag-icons
// set, same as the analytics visitor cells.
import * as flagSvgs from 'country-flag-icons/string/3x2'
// The Seed-X reference translator's languages (scripts/translator.py) — the
// translation-target ceiling — each mapped to the language's home country
// flag (England for English, Portugal for Portuguese — not the most
// populous variant). Internal tags are the bare 2-letter base subtags; the
// translator decides the variant. A variant tag (en-US, pt-BR) is still a
// valid explicit selection for a future translator that distinguishes them
// — flagFor shows its own region then.
export const TRANSLATABLE = {
ar: 'EG', cs: 'CZ', da: 'DK', de: 'DE', el: 'GR', en: 'GB', es: 'ES',
fa: 'IR', fi: 'FI', fr: 'FR', hu: 'HU', id: 'ID', it: 'IT', ja: 'JP',
ko: 'KR', ms: 'MY', nl: 'NL', no: 'NO', pl: 'PL', pt: 'PT', ro: 'RO',
ru: 'RU', sv: 'SE', th: 'TH', tr: 'TR', uk: 'UA', vi: 'VN', zh: 'CN',
}
// The languages in geographic/cultural groups (the lang tab's flag grid
// lays them out one group per row, in this order): English with the
// Nordics, then Western/Central and Eastern Europe, Southern Europe with
// the Middle East, and Asia.
export const LANG_GROUPS = [
['en', 'nl', 'da', 'no', 'sv', 'fi', 'ru'],
['fr', 'de', 'pl', 'cs', 'hu', 'ro', 'uk'],
['es', 'pt', 'it', 'el', 'tr', 'ar', 'fa'],
['zh', 'ja', 'ko', 'vi', 'th', 'id', 'ms'],
]
const displayNames = new Intl.DisplayNames(['en'], { type: 'language' })
// English display name for a language tag ("fi" -> "Finnish").
export function langName(tag) {
try {
return displayNames.of(tag) || tag
} catch {
return tag
}
}
// Flag SVG string for a language tag: an explicit region variant (en-US)
// gets its own region's flag; a bare base tag maps to the language's home
// country (en → GB, pt → PT); languages outside the list fall back to the
// tag's most likely region.
export function flagFor(tag) {
tag = tag || ''
if (!tag.includes('-')) {
const country = TRANSLATABLE[tag.split('-')[0].toLowerCase()]
if (country) return flagSvgs[country] || ''
}
try {
return flagSvgs[new Intl.Locale(tag).maximize().region] || ''
} catch {
return ''
}
}
+50
View File
@@ -0,0 +1,50 @@
// Public language-selector entry: imported on demand by pagerite.js on
// pages advertising more than one language in their hreflang alternates.
// Vue, Pinia and the flag SVG set live in this chunk only — untranslated
// pages never pay for them. The selector's state lives in the shared
// store (./store), not the DOM: the corner container is rebuilt freely
// and ensureMounted re-mounts from the store.
import { createApp } from 'vue'
import LangSelector from './LangSelector.vue'
import { pinia, useStore } from './store'
let app = null
function store() {
return useStore(pinia)
}
// The current page's languages (called on every navigation).
export function setLanguages(alternates, current) {
Object.assign(store(), {
langAlternates: alternates,
servedLang: current,
langSelectorActive: true,
})
}
// The current page is single-language: the selector goes away.
export function hide() {
store().langSelectorActive = false
app?.unmount()
app = null
}
// Mount the selector as the container's first item; re-mount when its
// element went away with a container rebuild (a live app updates from the
// store reactively).
export function ensureMounted(host) {
if (!store().langSelectorActive || !host) {
app?.unmount()
app = null
return
}
if (app && host.contains(app._container)) return
app?.unmount()
const el = document.createElement('div')
el.id = 'lang-selector'
host.prepend(el)
app = createApp(LangSelector)
app.use(pinia)
app.mount(el)
}
+132 -25
View File
@@ -1,9 +1,14 @@
// Pagerite editor entries. Two separate apps, mounted in their own // Pagerite editor entry. A single tabbed EditorShell is mounted in a
// dynamically created host divs inside the static document: // dynamically created host div inside the static document. The shell hosts
// - PageEditor ("page" mode): pen next to an article heading — Markdown // four tabs: PageEditor (Markdown + preview), BannerEditor (per-page banner
// editing with the preview rendered into the visible article. // HTML/design), SiteEditor (site-wide config), and StructureEditor (the
// - SiteEditor ("site" mode): pen on the banner — banner HTML editing // site structure tree).
// (previewed into the real banner) and the site structure tree. //
// Individual pens are shorthands that open the shell on a particular tab;
// once the shell is open, pens switch tabs instead of closing/remounting.
// Closing hides the shell but keeps the Vue app mounted, so editor state —
// including unsaved page text — survives and editing can continue on the
// next pen click; it is lost only on a real page reload.
if (import.meta.env.DEV) { if (import.meta.env.DEV) {
// Base styles only; the theme CSS is imported by pagerite.js (which // Base styles only; the theme CSS is imported by pagerite.js (which
// always runs first — the editor opens from public pages). // always runs first — the editor opens from public pages).
@@ -11,25 +16,94 @@ if (import.meta.env.DEV) {
} }
import { createApp } from 'vue' import { createApp } from 'vue'
import PageEditor from './PageEditor.vue' import EditorShell from './EditorShell.vue'
import SiteEditor from './SiteEditor.vue'
let host = null let host = null
let app = null
let savedTitle = null
let visible = false
let slideAnimation = null
const SLIDE_MS = 250 // keep in sync with the panel slide in pagerite.css
// The layout switches instantly when .editing toggles — no margin/width
// transitions anywhere, so viewport resizes (and the vw-based .wide bleed)
// always stay instant. The visible slide is a compositor-only FLIP
// transform on #content, running in sync with the panel's own slide
// (editor-slide-in / .closing in pagerite.css): both move by --editor-w
// over the same duration and easing, so .wide's left edge tracks the
// panel's right edge exactly throughout.
function setEditingClass(enable) {
const content = document.getElementById('content')
const before = content.getBoundingClientRect().left
document.body.classList.toggle('editing', enable)
const delta = before - content.getBoundingClientRect().left
slideAnimation?.cancel()
if (delta) {
slideAnimation = content.animate(
{ transform: [`translateX(${delta}px)`, 'translateX(0)'] },
{ duration: SLIDE_MS, easing: 'ease' }
)
}
}
// The panel is fixed to the viewport's left edge (pagerite.css) but tracks
// the page: its top is the banner's bottom edge while the banner is visible
// (= #content's top edge), and the viewport top once the banner has
// scrolled away. The window keeps scrolling normally while editing.
// Below 48rem the panel covers the entire viewport (pagerite.css), so its
// top stays 0 regardless of the banner.
const narrow = matchMedia('(max-width: 48rem)')
function trackPanelTop() {
const content = document.getElementById('content')
if (host && content) {
host.style.top = narrow.matches
? '0px'
: `${Math.max(0, content.getBoundingClientRect().top)}px`
}
}
function startTrackingPanel() {
trackPanelTop()
addEventListener('scroll', trackPanelTop, { passive: true })
addEventListener('resize', trackPanelTop)
}
function stopTrackingPanel() {
removeEventListener('scroll', trackPanelTop)
removeEventListener('resize', trackPanelTop)
}
export function openEditor(path, { mode = 'page' } = {}) { export function openEditor(path, { mode = 'page' } = {}) {
closeEditor() if (app) {
// Shell already mounted: re-show it if hidden, switch to the requested
// tab and retarget the editors to the current page (the shell survives
// fetch-navigation).
if (!visible) showEditor()
document.body.dataset.editorMode = mode
dispatchEvent(new CustomEvent('pagerite:switch-editor', { detail: { mode, path } }))
return
}
savedTitle = document.title
host = document.createElement('div') host = document.createElement('div')
host.className = 'editor-host' host.className = 'editor-host'
// Docked inside #content: below the banner, next to the article only. // Appended to <body>, not #content: the open/close slide transforms
document.getElementById('content').prepend(host) // #content (setEditingClass), and a transformed element becomes the
document.body.classList.add('editing') // containing block for fixed-position descendants — the panel would be
// Which kind of editor is open; pagerite.js uses this to decide // dragged along with the content instead of sliding on its own.
// whether a pen click closes the panel or swaps in the other editor. document.body.append(host)
setEditingClass(true)
startTrackingPanel()
// Which tab is active; pagerite.js uses this to decide whether a pen click
// closes the panel or switches tabs.
document.body.dataset.editorMode = mode document.body.dataset.editorMode = mode
createApp(mode === 'site' ? SiteEditor : PageEditor, { visible = true
app = createApp(EditorShell, {
pagePath: path, pagePath: path,
initialMode: mode,
onClose: closeEditor, onClose: closeEditor,
}).mount(host) })
app.mount(host)
// The slide-in (editor-slide-in in pagerite.css) is a one-shot open // The slide-in (editor-slide-in in pagerite.css) is a one-shot open
// effect; once finished, drop it so that later stylesheet swaps (theme // effect; once finished, drop it so that later stylesheet swaps (theme
// change re-creating @keyframes) cannot restart it. // change re-creating @keyframes) cannot restart it.
@@ -41,13 +115,46 @@ export function openEditor(path, { mode = 'page' } = {}) {
}) })
} }
export function closeEditor() { function showEditor() {
if (!host) return savedTitle = document.title
document.body.classList.remove('editing') host.style.display = ''
delete document.body.dataset.editorMode host.firstElementChild?.classList.remove('closing')
// Slide the panel out in sync with the page shifting back. setEditingClass(true)
host.firstElementChild?.classList.add('closing') startTrackingPanel()
const old = host visible = true
host = null dispatchEvent(new CustomEvent('pagerite:editor-shown'))
setTimeout(() => old.remove(), 250)
} }
export function closeEditor() {
if (!visible) return
visible = false
stopTrackingPanel()
setEditingClass(false)
// dataset.editorMode is kept while hidden: the tabs use it to tell whether
// a pagerite:editor-shown event targets them.
// Slide the panel out in sync with the page shifting back, then hide it.
host.firstElementChild?.classList.add('closing')
const h = host
setTimeout(() => { h.style.display = 'none' }, 250)
dispatchEvent(new CustomEvent('pagerite:editor-hidden'))
// The editor may have dropped the prefetch cache; warm it again for the
// now-final page so navigation stays instant.
dispatchEvent(new CustomEvent('pagerite:preload-pages'))
// Restore the server-rendered title for the current URL. Re-fetching makes
// sure a brand change in the site editor or an in-place navigation leaves
// the correct public title behind.
const restoreTitle = savedTitle
savedTitle = null
fetch(location.pathname)
.then((r) => r.text())
.then((html) => {
if (visible) return // reopened meanwhile; the editor owns the title
const doc = new DOMParser().parseFromString(html, 'text/html')
if (doc.title) document.title = doc.title
})
.catch(() => {
if (!visible && restoreTitle != null) document.title = restoreTitle
})
}
+864 -160
View File
File diff suppressed because it is too large Load Diff
+56
View File
@@ -0,0 +1,56 @@
// Shared reconnect policy for the WebSockets (page/banner editors,
// analytics view, the activity channel). Two things trip a browser's
// WebSocket throttling, after which every socket to the host sits
// "pending" (never opens, never closes) for minutes:
//
// 1. A burst of simultaneous attempts — page load opens Vite's HMR
// socket plus several of ours at the same moment, and every refresh
// repeats the burst. socketSlot() spaces new sockets out.
// 2. Too-frequent retries — so failed attempts back off exponentially
// (a few seconds, doubling to half a minute), reset only after a
// connection stayed open long enough to count as healthy. A socket
// that closes right after opening must NOT reset the backoff.
export function reconnectPolicy({ min = 2000, max = 30000, healthyAfter = 30000 } = {}) {
let delay = min
let openedAt = 0
return {
// Stamp a socket that just opened.
opened() {
openedAt = Date.now()
},
// The socket closed: the wait before the next attempt (up to 50%
// jitter; the base doubles per failure). A healthy streak resets it.
closed() {
if (openedAt && Date.now() - openedAt >= healthyAfter) delay = min
openedAt = 0
const wait = Math.round(delay * (1 + Math.random() * 0.5))
delay = Math.min(delay * 2, max)
return wait
},
}
}
// Sockets created at the same moment (page load: Vite's HMR socket plus
// ours) read as one burst to the browser's throttling. Space new sockets
// out: each call reserves a slot a beat after the previous one.
let nextSlot = 0
export function socketSlot() {
const now = Date.now()
const wait = Math.max(0, nextSlot - now)
nextSlot = Math.max(now, nextSlot) + 300
return wait
}
// A socket still CONNECTING after this long counts as a failed attempt:
// browser throttling leaves sockets "pending" (no open, no close) for
// minutes, and without a watchdog the app would wait on one forever (the
// recurring empty editor). Closing it fires onclose, which reschedules
// through the policy's backoff — it never reconnects aggressively itself.
export function watchConnecting(ws, label) {
return setTimeout(() => {
if (ws.readyState === WebSocket.CONNECTING) {
console.warn(`[pagerite] ${label} socket stuck connecting — closing it, retrying with backoff`)
ws.close()
}
}, 10_000)
}
+26
View File
@@ -0,0 +1,26 @@
// The app's shared Pinia store — cross-bundle UI state lives here. Every
// entry chunk imports its own copy of this module, so the Pinia instance
// is parked on window (Vue itself is a shared chunk, so reactivity works
// across the copies). Pass `pinia` explicitly when calling useStore
// outside a component (module code, no active instance).
import { createPinia, defineStore } from 'pinia'
export const pinia = (window.__pageritePinia ??= createPinia())
export const useStore = defineStore('pagerite', {
state: () => ({
// The ONE language selection, v-modeled by both dropdowns (editor
// tabs, public corner selector): '' = no explicit pick (the page's
// primary / autodetect), else a concrete tag. A pick from either
// dropdown is visible to everyone immediately.
lang: '',
// The language the current page was actually served in (set by
// pagerite.js per navigation) — the selector's highlight fallback
// when there is no explicit pick.
servedLang: '',
// The public selector's page data: hreflang alternates
// ([{tag, href, primary}]) and whether to show at all.
langAlternates: [],
langSelectorActive: false,
}),
})
+164
View File
@@ -0,0 +1,164 @@
// Shared in-place page re-rendering for the editor tabs: fetch a page
// without transitions and swap its dynamic regions into the live document.
// Used by BannerEditor (banner design changes), SiteEditor (theme changes)
// and StructureEditor (tree navigation).
// Drop the public page runtime's in-memory prefetch cache. Editors call this
// whenever a site-wide or page change invalidates the cached HTML of other
// pages (theme, headings, structure, banner, etc.). The cache is rebuilt by
// re-preloading visible links once the editor panel closes.
export function dropPageCache() {
dispatchEvent(new CustomEvent('pagerite:drop-page-cache'))
}
// The editor's language override (set by EditorShell): while the panel is
// open, its language selection wins over the normal preferences — every
// in-place re-render asks for that language explicitly, and pagerite.js
// applies it to its own fetches and prefetches (pagerite:session-lang).
// The primary selection pins by its code: ?lang=<primary> selects the
// original explicitly (i18n.select_language). Panel closed, the session's
// chosen language (window.__pageriteLang) takes over — the pick stays.
let overrideLang = null // the ?lang= value in force, null = the session's
export function setLangOverride(queryLang) {
overrideLang = queryLang || null
dispatchEvent(new CustomEvent('pagerite:session-lang', { detail: { lang: overrideLang } }))
}
export function runScripts(root) {
// Scripts injected via innerHTML do not execute; re-create them.
if (!root) return
for (const old of root.querySelectorAll('script')) {
const s = document.createElement('script')
for (const a of old.attributes) s.setAttribute(a.name, a.value)
s.textContent = old.textContent
old.replaceWith(s)
}
}
function swapRegions(doc) {
for (const id of ['page-banner', 'nav', 'main']) {
const fresh = doc.getElementById(id)
const el = document.getElementById(id)
if (fresh && el) el.replaceWith(document.importNode(fresh, true))
}
// #sidebar is omitted entirely when the section has no sub-navigation,
// so it may be absent on either side: replace, insert, or remove.
const freshSidebar = doc.getElementById('sidebar')
const curSidebar = document.getElementById('sidebar')
if (freshSidebar && curSidebar) {
curSidebar.replaceWith(document.importNode(freshSidebar, true))
} else if (freshSidebar) {
document.getElementById('main')?.before(document.importNode(freshSidebar, true))
} else if (curSidebar) {
curSidebar.remove()
}
// The brand lives in the header, outside the swappable regions, and is
// absent entirely when neither a brand nor custom brand HTML is set. A
// plain link keeps its element (text swap only, preserving the shrink-
// to-fit observers); anything else (custom HTML wrapper) is replaced.
const freshBrand = doc.getElementById('brand')
const curBrand = document.getElementById('brand')
if (freshBrand && curBrand && freshBrand.tagName === 'A' && curBrand.tagName === 'A') {
curBrand.textContent = freshBrand.textContent
} else if (freshBrand && curBrand) {
curBrand.replaceWith(document.importNode(freshBrand, true))
runScripts(document.getElementById('brand'))
} else if (curBrand) {
curBrand.remove()
} else if (freshBrand) {
document.getElementById('nav')?.before(document.importNode(freshBrand, true))
runScripts(document.getElementById('brand'))
}
// Site-wide custom CSS is in <head> and must be swapped too.
const freshUserStyle = doc.getElementById('pagerite-user')
const curUserStyle = document.getElementById('pagerite-user')
if (freshUserStyle && curUserStyle) {
curUserStyle.textContent = freshUserStyle.textContent
} else if (freshUserStyle) {
document.head.appendChild(document.importNode(freshUserStyle, true))
} else if (curUserStyle) {
curUserStyle.remove()
}
// Theme and other public stylesheets live in <head>, rendered with stable
// ids by the backend (links in dev, inline <style> elements in prod);
// sync them positionally so the custom CSS (rendered last) always keeps
// winning by order. Diff-based: unchanged sheets keep their elements, so
// their @keyframes are never torn down (re-creating keyframes would
// replay the editor's slide-in animation).
const sel = 'link[rel="stylesheet"][id], style[id]'
const freshEls = [...doc.head.querySelectorAll(sel)]
const freshIds = new Set(freshEls.map((el) => el.id))
for (const el of [...document.head.querySelectorAll(sel)]) {
if (!freshIds.has(el.id)) el.remove()
}
// Insert missing sheets in the fresh document's order, each right after
// its predecessor's element. The first sheet rendered is always the base
// CSS, so its element doubles as the fallback anchor when nothing matched
// yet (e.g. no theme was selected before and the position is otherwise
// lost).
let anchor = null
for (const el of freshEls) {
const cur = el.id && document.getElementById(el.id)
if (cur && cur.outerHTML === el.outerHTML) {
anchor = cur
continue
}
const imported = document.importNode(el, true)
// Same id, new content (theme switch): replace in place, keeping position.
if (cur) cur.replaceWith(imported)
else if (anchor) anchor.after(imported)
else {
const base = document.getElementById('pagerite-base')
if (base) base.after(imported)
else document.head.append(imported)
}
anchor = imported
}
// The served language rides on <html> (lang + dir, rtl for e.g. Arabic):
// follow the swapped page. The editor panel carries its own lang="en"
// dir="ltr", so it is unaffected.
document.documentElement.lang = doc.documentElement.lang
document.documentElement.dir = doc.documentElement.dir
// The editor keeps its own title while open; only inherit the server title
// when navigating outside the editor (e.g. fetch-navigation swaps).
if (!document.body.classList.contains('editing')) {
document.title = doc.title
}
}
// Fetch /p, swap its regions into the live page and replaceState to it.
// Returns the final URL (after redirects), or null when the fetch did not
// yield a page. Category and missing URLs render a placeholder 404 page —
// fine to swap in (new pages are created by editing them). The fetch pins
// the editor's language override, or — panel closed — the session's chosen
// language (window.__pageriteLang).
export async function loadPlain(p) {
let doc
let finalUrl = `/${p}`
let html
try {
const pin = overrideLang || window.__pageriteLang
const res = await fetch(pin ? `${finalUrl}?lang=${pin}` : finalUrl)
const type = res.headers.get('content-type') || ''
if (!type.includes('text/html')) return null
if (res.redirected) finalUrl = res.url
html = await res.text()
doc = new DOMParser().parseFromString(html, 'text/html')
} catch { return null }
if (!doc.getElementById('main')) return null
swapRegions(doc)
// The address bar keeps the pretty URL: a language query is a fetch
// detail, never shown (pagerite.js's initial ?lang= works the same).
const pretty = new URL(finalUrl, location.href)
pretty.searchParams.delete('lang')
history.replaceState(history.state, '', pretty)
runScripts(document.getElementById('page-banner'))
runScripts(document.getElementById('main'))
// Keep pagerite.js's in-memory page cache in sync with the fresh copy.
// The URL is announced as fetched: a language-pinned copy caches under
// its own ?lang= key, where navigation with the same pin finds it.
dispatchEvent(new CustomEvent('pagerite:page-fetched', { detail: { url: finalUrl, html } }))
dispatchEvent(new CustomEvent('pagerite:preview')) // re-inject + re-tuck the edit pens
return finalUrl
}
+3 -1
View File
@@ -5,13 +5,14 @@
* Configures Vite for FastAPI backend integration: * Configures Vite for FastAPI backend integration:
* - Proxies /api/* requests to the FastAPI backend * - Proxies /api/* requests to the FastAPI backend
* - Builds to the Python module's frontend-build directory * - Builds to the Python module's frontend-build directory
* - Disables Vite's screen clearing on startup
* *
* Options: * Options:
* paths - Array of paths to proxy (default: ["/api"]) * paths - Array of paths to proxy (default: ["/api"])
*/ */
export default function fastapiVue({ paths = ["/api"] } = {}) { export default function fastapiVue({ paths = ["/api"] } = {}) {
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:3200" const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:8210"
// Build proxy configuration for each path // Build proxy configuration for each path
const proxy = {} const proxy = {}
@@ -26,6 +27,7 @@ export default function fastapiVue({ paths = ["/api"] } = {}) {
return { return {
name: "vite-plugin-fastapi-pagerite", name: "vite-plugin-fastapi-pagerite",
config: () => ({ config: () => ({
clearScreen: false,
server: { proxy }, server: { proxy },
build: { build: {
outDir: "../pagerite/frontend-build", outDir: "../pagerite/frontend-build",
+20 -17
View File
@@ -1,5 +1,4 @@
import { fileURLToPath, URL } from 'node:url' import { fileURLToPath, URL } from 'node:url'
import { readdirSync } from 'node:fs'
import fastapiVue from './vite-plugin-fastapi.js' import fastapiVue from './vite-plugin-fastapi.js'
import { defineConfig } from 'vite' import { defineConfig } from 'vite'
@@ -8,24 +7,16 @@ import vueDevTools from 'vite-plugin-vue-devtools'
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200' const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200'
// Every theme directory ships its theme.css as a separate build entry, so // Proxy everything except Vite's own dev-time paths and the backend machinery
// the backend can link base and theme stylesheets independently. // to the FastAPI backend in dev. /_api, /_f, /_themes, /_fonts and /_a are
const themesDir = fileURLToPath(new URL('./src/assets/themes', import.meta.url)) // handled by the fastapi-vue plugin, and /@..., /src, /node_modules, /__...
const themeInputs = Object.fromEntries( // stay with Vite.
readdirSync(themesDir, { withFileTypes: true }) const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
.filter((d) => d.isDirectory())
.map((d) => [`theme_${d.name}`, `${themesDir}/${d.name}/theme.css`]),
)
// Proxy content pages (/slug, /path/to/slug) to the FastAPI backend in dev.
// Excludes Vite internals (/@..., /src, /node_modules, /__...) and the
// backend's /_ prefix. /_api and /_f are handled by the fastapi-vue plugin.
const CONTENT_PROXY = '^\\/(?!_|@|src|node_modules|__)(?:[^./?]+(?:\\/[^./?]+)*)?(?:\\?.*)?$'
// https://vite.dev/config/ // https://vite.dev/config/
export default defineConfig({ export default defineConfig({
plugins: [ plugins: [
fastapiVue({ paths: ["/_api", "/_f"] }), fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_fonts", "/_a", "/_ws", "/_translate"] }),
vue(), vue(),
vueDevTools(), vueDevTools(),
], ],
@@ -35,10 +26,19 @@ export default defineConfig({
}, },
}, },
appType: 'mpa', // no SPA fallback; every HTML page is served by FastAPI appType: 'mpa', // no SPA fallback; every HTML page is served by FastAPI
resolve: {
alias: {
// All components are precompiled SFCs — drop the runtime template
// compiler (~60 kB min) from the bundle.
vue: 'vue/dist/vue.runtime.esm-bundler.js',
},
},
build: { build: {
// The main editor bundle (CodeMirror + Vue) is intentionally one chunk.
chunkSizeWarningLimit: 1200,
// Mirror the URL space in the build output: hashed files land under // Mirror the URL space in the build output: hashed files land under
// frontend-build/_assets/ and the Frontend serves the build directory // frontend-build/_assets/ and the Frontend serves the build directory
// at the site root (frontend/public/favicon.ico -> /favicon.ico). // at the site root.
manifest: true, manifest: true,
assetsDir: '_assets', assetsDir: '_assets',
rollupOptions: { rollupOptions: {
@@ -49,8 +49,11 @@ export default defineConfig({
input: { input: {
main: fileURLToPath(new URL('./src/main.js', import.meta.url)), main: fileURLToPath(new URL('./src/main.js', import.meta.url)),
pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)), pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)),
analytics: fileURLToPath(new URL('./src/analytics-main.js', import.meta.url)),
langselect: fileURLToPath(new URL('./src/langselect-main.js', import.meta.url)),
// Only the base CSS is built; theme/banner-design stylesheets live
// in pagerite/themes/{name}/ and are served by the backend as-is.
pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)), pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)),
...themeInputs,
}, },
}, },
}, },
+34 -4
View File
@@ -1,31 +1,61 @@
# auto-upgrade@fastapi-vue-setup - remove this if you modify this file
"""Command-line entry point for running the backend server.""" """Command-line entry point for running the backend server."""
import argparse import argparse
import os import os
from pathlib import Path
import msgspec
from fastapi_vue import server from fastapi_vue import server
DEFAULT_PORT = 3100 from pagerite.config import Config
DEFAULT_PORT = 8100
DEVMODE = os.getenv("PAGERITE_DEV") == "1" DEVMODE = os.getenv("PAGERITE_DEV") == "1"
def main() -> None: def main() -> None:
"""Run the backend server with optional arguments.""" """Run the backend server with optional arguments."""
parser = argparse.ArgumentParser(description="Run the pagerite server.") parser = argparse.ArgumentParser(description="Run the pagerite server.")
parser.add_argument(
"hostname",
nargs="?",
default="localhost",
help=(
"Public hostname of the site; names the data directory "
"<hostname>/{content.kantadb, analytics.json, files} under the "
"cwd (default: localhost)."
),
)
parser.add_argument( parser.add_argument(
"-l", "-l",
"--listen", "--listen",
action="append", action="append",
help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."), help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."),
) )
parser.add_argument(
"--dbip",
action="store_true",
help="Download/update the DB-IP city lite database before starting.",
)
args = parser.parse_args() args = parser.parse_args()
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {} # Hand configuration to the app as JSON in PAGERITE_CONFIG; it must be
# set before pagerite.app is imported, as state.py reads it at import
# time (data directory, public origin).
os.environ["PAGERITE_CONFIG"] = msgspec.json.encode(
Config(hostname=args.hostname, dbip=args.dbip)
).decode()
run_args: dict = {}
if args.hostname != "localhost":
# A public site sits behind TLS on its hostname; show that URL in the
# startup box instead of the local listen address.
run_args["startup_box"] = f"{{Name}} {{version}}\nhttps://{args.hostname}"
server.run( server.run(
"pagerite.app:app", "pagerite.app:app",
listen=args.listen, listen=args.listen,
default_port=DEFAULT_PORT, default_port=DEFAULT_PORT,
**dev, server_header=False,
reload=Path(__file__).parent if DEVMODE else False,
**run_args,
) )
+972
View File
@@ -0,0 +1,972 @@
"""Server-side visit analytics (collection only; see docs/analytics.md).
Raw recording, display-time classification. Every document GET is appended
to ``Analytics.gets`` as a raw access-log line (path with query string, true
HTTP status, external referer origin, preload flag) and every pagerite.js
activity message from the /_ws WebSocket is appended to ``Analytics.msgs``
(navigations ``fr`` -> ``to`` and active reading-time updates). Nothing is
classified when it is recorded: whether a client turns out to be a reader,
a crawler or a scanner is decided by ``Store.display()`` from the raw
events, so the stored data survives any future change to the classification
rules.
At display time:
- An IP with a 404 on a telltale scanner path (empty segment like
``//foo``, dot segment, *.php) or ten plain 404s within an hour on paths
that don't resolve to a menu node is classified as abuse; every document
GET from that IP is shown in the abuse list, split by status into the 404
probes and the real articles read. Its activity messages are ignored.
Hidden (admin) clients never trigger classification — editing means
visiting not-found pages. RFC 8615 well-known URIs (``/.well-known/…``)
are mostly legitimate browser/service probes and never count as abuse
evidence.
- A document GET never followed by an activity message (within
``_CRAWLER_TIMEOUT``) is a crawler hit. Messages whose UA claims a
crawler identity (``_is_bot_ua``) are ignored, so JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) land in the crawler list too.
- The remaining messages are grouped into visits per client: a new visit
starts after ``_SESSION_GAP`` of inactivity. A visit whose total
reported reading time is under ``_MIN_VISIT_READ`` seconds is a
real-browser bot and is reclassified as crawler hits (durations are
client-provided and trusted — such bots report 02 s).
- Admin clients (``hide`` message field, set at collection time on the
client record) are excluded from every list and aggregate.
Idle-time link preloads from pagerite.js (``x-pagerite-preload`` header)
are recorded raw but flag ``pre``: they never count as views, crawler hits
or abuse — they exist so a navigation served from the in-memory page cache
can still be attributed the HTTP status of its preload GET.
Aggregates (site visits, page views, transitions) are not stored; they are
computed at display time from the derived visits. Client metadata (IP, UA,
language, country/city, host) is stored once per unique client hash and
referenced from every event.
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
rewritten atomically on every recorded event.
"""
import ipaddress
import os
import re
import tempfile
from bisect import bisect_left, bisect_right
from collections.abc import Callable
from contextlib import suppress
from datetime import UTC, datetime, timedelta
from pathlib import Path
from urllib.parse import parse_qs, urlencode, urlparse
import blake3
import msgspec
from ua_parser import parse
def _compact_user_agent(ua: str) -> str:
"""Format a User-Agent string into a compact display form.
Returns the original UA when the parser cannot identify the browser/OS.
"""
if not ua or not ua.strip() or ua == "-":
return ""
r = parse(ua)
browser = r.user_agent.family if r.user_agent else None
ver = r.user_agent.major if r.user_agent else ""
os_name = r.os.family if r.os else None
dev = r.device.family if r.device else None
if browser in (None, "Other") and os_name in (None, "Other"):
return ua
if browser and browser != "Other":
browser = browser.split()[0]
else:
browser = ""
os_name = os_name if os_name and os_name != "Other" else ""
if dev in (None, "Other") or dev == browser:
dev = ""
parts = [f"{browser}/{ver}" if browser else "", os_name, dev]
return " ".join(p for p in parts if p).strip()
class Ping(msgspec.Struct, omit_defaults=True):
"""One client message on the /_ws activity WebSocket (wire format).
Sent as a JSON text frame (msgspec-encoded, decoded to str for the
wire). ``to`` set: a navigation — internal page path or external https
exit URL. ``read`` alone (with ``fr``): an active reading-time update
for the page ``fr``; these arrive frequently while the user is active.
``hide`` flags the client as an admin: everything it ever did is
excluded from the statistics.
"""
#: Path of the page the activity happened on ("" for the initial load).
fr: str = ""
#: Navigation target: internal path or external https exit URL.
to: str = ""
#: Active reading time (seconds) spent on ``fr`` since the last report.
read: int = 0
#: Admin client: record but hide everything from the statistics.
hide: bool = False
class Get(msgspec.Struct, omit_defaults=True):
"""One document GET — the raw access-log line.
Recorded once per served document (200, or 404 for a category
placeholder or a missing page); never classified at this point.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
t: datetime
#: Full request path including the query string (e.g. "/.env?x=1").
path: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: True HTTP status of the response.
status: int = 200
#: External https origin of the Referer, "" for direct/internal.
ref: str = ""
#: Idle-time cache warm-up by pagerite.js (x-pagerite-preload header):
#: never counted as a view/crawler/abuse hit; recorded only so a later
#: cache-served navigation can be attributed this GET's status.
pre: bool = False
class Msg(msgspec.Struct, omit_defaults=True):
"""One raw activity message from pagerite.js over /_ws.
``to`` set: a navigation — validated internal page path or external
https exit URL. ``read`` alone (with ``fr``): an active reading-time
update for the page ``fr`` (seconds since the previous report).
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
t: datetime
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: Path of the page the activity happened on ("" for the initial load).
fr: str = ""
#: Navigation target: internal path or external https exit URL.
to: str = ""
#: Active reading time (seconds) spent on ``fr`` since the last report.
read: int = 0
class Client(msgspec.Struct, omit_defaults=True):
"""Client metadata shared by the raw events and every derived row.
Identified by a 6-byte blake3 hash of the IPv4 address or IPv6 /64
network, the full User-Agent string and the extracted language tag.
Country/city/host are filled in asynchronously after the first event.
"""
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
ip: str = ""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter country code from the DB-IP geoip lookup, or "".
country: str = ""
#: City name from the DB-IP geoip lookup, or "".
city: str = ""
#: Raw User-Agent header.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
#: True for admin clients (hide=1 ping): everything this client ever did
#: is excluded from all statistics and from the viewer payload.
hide: bool = False
# --- Display DTOs -------------------------------------------------------
# The structs below are never persisted; Store.display() builds them from
# the raw events. They define the viewer payload shape consumed by
# frontend/src/analytics/*.
class Nav(msgspec.Struct, omit_defaults=True):
"""One navigation inside a visit: from ``fr`` to ``to``.
``to`` is an internal page path or an external https exit URL. Every
navigation is logged (repeats included), keyed by its timestamp in
``Visit.navs``, so display-time aggregates can count views and
transitions; ``Visit.trail`` keeps the first-seen order.
"""
fr: str
to: str
class TrailItem(msgspec.Struct, omit_defaults=True):
"""One first-seen target in a visit trail: a page or external exit URL.
``read`` accumulates active reading time (seconds) across the whole
visit; ``status`` is the most recent HTTP status seen for the target.
"""
to: str
#: Accumulated active reading time in seconds.
read: int = 0
#: Most recent HTTP status of the response (200 or 404).
status: int = 200
class Visit(msgspec.Struct, omit_defaults=True):
"""One visit: a client's activity since ``_SESSION_GAP`` of inactivity.
``trail`` holds the entry page and everything seen afterwards, keyed by
the timestamp of first sight (insertion order = first-seen order);
re-visiting an already seen target updates its item instead of
appending. Client metadata is held in ``Analytics.clients`` keyed by
``client``.
"""
start: datetime
entry: str
#: External https origin of the initial load, "" for direct visits.
referer: str = ""
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: First-seen targets keyed by their timestamp (entry included).
trail: dict[datetime, TrailItem] = {}
#: Every navigation (repeats included) keyed by its timestamp; the
#: aggregates are computed from this log at display time.
navs: dict[datetime, Nav] = {}
#: UTM query parameters from the landing URL, keyed by parameter name.
utm: dict[str, str] = {}
class CrawlerHit(msgspec.Struct, omit_defaults=True):
"""A document GET that was never followed by an activity message.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime
entry: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: External https origin of the initial load, "" for direct/none.
referer: str = ""
#: Raw query string of the landing URL (UTM tags can be parsed from it).
query: str = ""
#: HTTP status of the served response (200 or 404 for content pages).
status: int = 200
class AbuseHit(msgspec.Struct, omit_defaults=True):
"""A document GET from an IP classified as a scanner/abuser.
Unlike crawler hits the full request path (query string included) is
kept: the interesting part is exactly which paths were probed.
``flag`` marks the paths that triggered classification; ``is_404``
distinguishes 404 responses (probed paths and 404-fallback document
GETs) from real 200 document GETs — the abuser actually reading
articles. Client metadata is held in ``Analytics.clients`` keyed by
``client``.
"""
start: datetime
#: Full request path including the query string (e.g. "/.env?x=1").
path: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: True when this path triggered abuse classification (telltale path
#: or the 404 that crossed the threshold).
flag: bool = False
#: True for 404 responses; false for real (200) document GETs.
is_404: bool = False
class Favicon(msgspec.Struct, omit_defaults=True):
"""Favicon fetch record for one external https origin.
The icon itself is stored on disk under a content-hashed name (like
uploads, but outside the kanta db), referenced here by ``file``; an
empty ``file`` is a known miss, retried after ``_FAVICON_RETRY``.
"""
#: Content-hashed file name of the stored icon, "" when the fetch failed.
file: str = ""
#: When the fetch was last attempted.
fetched: datetime | None = None
class Analytics(msgspec.Struct, omit_defaults=True):
"""Root of the analytics JSON file: the raw event log, append-only by
design. Old data is dropped by deleting list entries."""
#: Every document GET, in arrival order (see Get).
gets: list[Get] = []
#: Every activity message from pagerite.js, in arrival order (see Msg).
msgs: list[Msg] = []
#: Favicon fetch records keyed by external https origin.
favicons: dict[str, Favicon] = {}
#: Client metadata keyed by 6-byte blake3 hash.
clients: dict[bytes, Client] = {}
class Display(msgspec.Struct, omit_defaults=True):
"""The viewer payload: visible derived data plus display-time aggregates.
Hidden clients are excluded everywhere: their events are dropped, and
the aggregates are computed from the visible visits only.
The aggregate shapes match what the viewer consumes: sparse 5-minute
buckets keyed by their floored ISO timestamp.
"""
visits: list[Visit] = []
crawlers: list[CrawlerHit] = []
abuse: list[AbuseHit] = []
clients: dict[bytes, Client] = {}
#: origin -> URL path of the stored favicon ("/_f/<file>"),
#: only for origins whose icon was fetched successfully.
favicons: dict[str, str] = {}
#: Page transitions per 5-minute bucket (sparse):
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
#: "(direct)" for initial loads, a page path for pings.
transitions: dict[str, dict[str, dict[str, int]]] = {}
#: Page views per 5-minute bucket: path -> bucket ISO -> count (sparse).
views: dict[str, dict[str, int]] = {}
#: New visits per 5-minute bucket: bucket ISO -> count (sparse).
site_visits: dict[str, int] = {}
def _bucket(now: datetime) -> str:
"""Start of the 5-minute interval containing ``now``, as ISO string."""
return now.replace(minute=now.minute // 5 * 5, second=0, microsecond=0).isoformat()
def _origin(url: str) -> str | None:
"""The origin part of an https URL (scheme://host[:port]), else None."""
try:
parsed = urlparse(url)
except ValueError:
return None
if parsed.scheme != "https" or not parsed.netloc:
return None
return f"https://{parsed.netloc}"
def _external_target(url: str) -> str | None:
"""A valid https URL (origin or full page), else None."""
try:
parsed = urlparse(url)
except ValueError:
return None
if parsed.scheme != "https" or not parsed.netloc:
return None
return url
_SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*")
def _internal_path(to: str) -> str | None:
"""A valid internal page path ("/" or slug segments), else None."""
path = to.split("?")[0].split("#")[0].strip("/")
if not path:
return "/"
if all(_SEGMENT.fullmatch(seg) for seg in path.split("/")):
return f"/{path}"
return None
def _parse_accept_language(value: str) -> tuple[str, str]:
"""First Accept-Language tag and the region/country subtag if present.
``en-US, fr;q=0.9`` -> ("en-us", "US"). Wildcards and missing regions
produce an empty country. The region is intentionally approximate:
it reflects the browser's language preference, not geo-location.
"""
if not value:
return "", ""
tag = value.split(",")[0].split(";")[0].strip()
if not tag or tag == "*":
return "", ""
lang = tag.lower()
country = ""
# Region subtags follow the initial language tag (en-US, zh-Hans-CN).
# A bare two-letter tag such as "fr" is a language code, not a region.
for part in reversed(tag.split("-")[1:]):
if len(part) == 2 and part.isalpha():
country = part.upper()
break
return lang, country
def _utm_tags(query: str) -> dict[str, str]:
"""UTM parameters from a query string, keeping only the first value."""
if not query:
return {}
parsed = parse_qs(query, keep_blank_values=True)
return {k: v[0] for k, v in parsed.items() if k.startswith("utm_")}
_CRAWLER_TIMEOUT = timedelta(seconds=10)
#: Inactivity after which a client's next navigation starts a new visit.
_SESSION_GAP = timedelta(minutes=30)
#: Minimum total reported reading time (seconds, summed over the trail) for
#: a session to count as a visit; shorter sessions are JS-running bots and
#: are shown as crawler hits instead.
_MIN_VISIT_READ = 5
#: How long a failed favicon fetch suppresses retries for the same origin.
_FAVICON_RETRY = timedelta(days=7)
#: UAs of JS-running crawlers, which would register as visitors on their
#: activity messages. Anything calling itself a "bot" or "spider" matches;
#: known crawlers without those tokens (GoogleOther) are listed as extra
#: alternates. No source verification: a spoofed bot UA just lands in the
#: crawler list, and scanners that probe telltale paths are caught by the
#: abuse rules anyway.
_BOT_UA = re.compile(r"bot|spider|googleother", re.IGNORECASE)
def _is_bot_ua(ua: str) -> bool:
"""True when the UA claims a crawler identity (bot or spider)."""
return bool(_BOT_UA.search(ua))
#: Plain-404 count per IP within ``_ABUSE_404_WINDOW`` that classifies it as
#: abuse even without a telltale path hit. Windowed so a long-time reader
#: slowly accumulating misses (deleted articles over months) never crosses
#: it — scanners spray their probes in bursts.
_ABUSE_404_THRESHOLD = 10
#: Sliding window the plain-404 threshold is counted over.
_ABUSE_404_WINDOW = timedelta(hours=1)
#: Paths that instantly classify an IP as abuse when they 404: an empty
#: segment ("//foo" — no real client requests those), any segment starting
#: with a dot ("/.env", "/.git/config") or ending in ".php".
_ABUSE_PATH = re.compile(r"/{2,}|(^|/)\.|\.php$", re.IGNORECASE)
def _is_abuse_path(path: str) -> bool:
"""Telltale scanner path: empty segment, dot segment or *.php."""
return bool(_ABUSE_PATH.search(path.split("?")[0]))
def _is_well_known(path: str) -> bool:
"""RFC 8615 well-known URI ("/.well-known/...").
Mostly legitimate (browsers and services probe them, e.g. Chrome's
devtools fetch of appspecific/com.chrome.devtools.json): never telltale
and never counted toward the plain-404 threshold.
"""
return path.split("?")[0].lower().startswith("/.well-known/")
def _network_ip(ip: str) -> str:
"""IPv4 address unchanged, IPv6 collapsed to its /64 network address.
We hash the network rather than the full address so that clients in the
same /64 (a typical end-user allocation) are treated as one visitor.
"""
if not ip:
return ip
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ip
if isinstance(addr, ipaddress.IPv6Address):
return str(ipaddress.IPv6Network(f"{ip}/64", strict=False).network_address)
return ip
def _client_hash(ip: str, ua: str, lang: str) -> bytes:
"""6-byte blake3 digest identifying a visitor/client tuple.
The key is the prettified IP (IPv6 /64), the raw UA string and the
extracted language tag, separated by null bytes.
"""
return blake3.blake3(f"{_network_ip(ip)}\0{ua}\0{lang}".encode()).digest()[:6]
class Store:
"""In-memory analytics data: the raw event log plus its JSON persistence."""
def __init__(self, path: Path) -> None:
self.path = path
self.data = Analytics()
if path.exists():
try:
raw = path.read_bytes()
except OSError:
raw = b""
# The pre-redesign schema (stored visits/crawlers/abuse lists) is
# not convertible: set it aside and start fresh.
legacy = b'"visits"' in raw
if raw and not legacy:
try:
self.data = msgspec.json.decode(raw, type=Analytics)
except msgspec.DecodeError:
legacy = True # corrupt file: start fresh
if legacy:
with suppress(OSError):
path.rename(path.with_name(path.name + ".bak-legacy"))
#: Callables to notify when persisted data changes. Registered by the
#: analytics WebSocket broadcaster.
self._on_change: list[Callable[[], None]] = []
def subscribe(self, callback: Callable[[], None]) -> None:
"""Register a callback to be called after every persisted change."""
if callback not in self._on_change:
self._on_change.append(callback)
def unsubscribe(self, callback: Callable[[], None]) -> None:
"""Remove a previously registered change callback."""
with suppress(ValueError):
self._on_change.remove(callback)
def _notify(self) -> None:
for callback in self._on_change:
callback()
def _save(self) -> None:
"""Rewrite the JSON file atomically (temp file + rename)."""
try:
fd, tmp = tempfile.mkstemp(
dir=self.path.parent, prefix=self.path.name, suffix=".tmp"
)
with os.fdopen(fd, "wb") as f:
f.write(msgspec.json.encode(self.data))
os.replace(tmp, self.path)
except OSError:
pass # analytics must never break page serving
else:
self._notify()
def _hidden(self, client_hash: bytes) -> bool:
"""True when the client record is flagged hidden (admin)."""
client = self.data.clients.get(client_hash)
return client is not None and client.hide
def _ensure_client(
self,
ip: str,
ua: str,
lang: str,
*,
country: str = "",
) -> bytes:
"""Get or create a ``Client`` record; return its 6-byte hash."""
h = _client_hash(ip, ua, lang)
if h not in self.data.clients:
self.data.clients[h] = Client(
ip=ip,
ua=ua,
ua_pretty=_compact_user_agent(ua),
lang=lang,
country=country,
)
self._save()
return h
def enrich_client(
self,
client_hash: bytes,
*,
host: str = "",
country: str = "",
city: str = "",
) -> None:
"""Fill in host/geoip fields on a client record after async lookups."""
client = self.data.clients.get(client_hash)
if client is None:
return
changed = False
if host and not client.host:
client.host = host
changed = True
if country:
client.country = country
changed = True
if city:
client.city = city
changed = True
if changed:
self._save()
def record_get(
self,
ip: str,
ua: str,
path: str,
*,
status: int = 200,
referer: str = "",
accept_language: str = "",
pre: bool = False,
) -> bytes | None:
"""Append one document GET to the raw log.
``path`` is the full request path, query string included; ``status``
the true HTTP status of the response; ``referer`` the raw Referer
header (reduced here to an external https origin, "" when internal
or absent); ``pre`` marks idle-time preloads from pagerite.js.
Returns the client hash when the client record was just created (so
the caller can schedule async enrichment), else None.
"""
lang, country = _parse_accept_language(accept_language)
client_hash = _client_hash(ip, ua, lang)
new = client_hash not in self.data.clients
if new:
self._ensure_client(ip, ua, lang, country=country)
self.data.gets.append(
Get(
t=datetime.now(UTC),
path=path,
client=client_hash,
status=status,
ref=_origin(referer) or "",
pre=pre,
)
)
self._save()
return client_hash if new else None
def record_msg(
self,
fr: str,
to: str | None,
ip: str,
ua: str,
accept_language: str = "",
hide: bool = False,
read: int = 0,
) -> bytes | None:
"""Append one client activity message (``Ping`` from pagerite.js) to
the raw log.
``to`` is validated here — internal slug path or external https URL,
anything else is dropped; that is sanitation, not classification.
No abuse/bot filtering happens at record time: those messages are
stored raw and filtered at display time, so future rule changes lose
nothing. ``hide`` flags the client record as an admin; the message
itself is recorded normally and hidden at display time like
everything else the client ever did.
Returns the client hash when the client record was just created (so
the caller can schedule async enrichment), else None.
"""
lang, country = _parse_accept_language(accept_language)
client_hash = _client_hash(ip, ua, lang)
new = client_hash not in self.data.clients
if new:
self._ensure_client(ip, ua, lang, country=country)
if hide:
self.data.clients[client_hash].hide = True
fr = (_internal_path(fr) or "") if fr else ""
target = ""
if to:
if to.startswith("/") and not to.startswith("//"):
target = _internal_path(to) or ""
else:
target = _external_target(to) or ""
if target or read > 0:
self.data.msgs.append(
Msg(t=datetime.now(UTC), client=client_hash, fr=fr, to=target, read=read)
)
if target or read > 0 or hide:
self._save()
return client_hash if new else None
def favicon_origins_needed(self) -> list[str]:
"""External https origins seen in the raw events whose favicon needs
fetching.
Covers referer origins of document GETs and external exit targets of
activity messages — spiders often advertise their own site as the
referer, so the icon identifies them in the crawler table. Hidden
(admin) clients' events never trigger fetches. Origins with a
stored icon, or a miss younger than ``_FAVICON_RETRY``, are skipped.
"""
origins: set[str] = set()
for g in self.data.gets:
if g.ref and not self._hidden(g.client):
origins.add(g.ref)
for m in self.data.msgs:
origin = _origin(m.to)
if origin is not None and not self._hidden(m.client):
origins.add(origin)
now = datetime.now(UTC)
return [
origin
for origin in origins
if (f := self.data.favicons.get(origin)) is None
or (not f.file and (f.fetched is None or now - f.fetched > _FAVICON_RETRY))
]
def record_favicon(self, origin: str, file: str = "") -> None:
"""Store the favicon fetch result for ``origin`` ("" = miss)."""
self.data.favicons[origin] = Favicon(file=file, fetched=datetime.now(UTC))
self._save()
def display(self, in_menu: Callable[[str], bool] | None = None) -> Display:
"""Build the viewer payload from the raw events.
All classification happens here, so the stored data is independent
of the rules:
- abuse IPs: a 404 on a telltale path (empty segment ``//foo``, dot
segment or ``*.php`` — but never ``/.well-known/…``, which is
mostly legitimate browser probing and counts neither as telltale
nor toward the threshold), or ``_ABUSE_404_THRESHOLD`` plain 404s
within ``_ABUSE_404_WINDOW`` on paths that don't resolve to a
menu node (``in_menu``; category placeholders return 404 but are
real nodes and never count). Hidden clients never classify.
Every non-preload GET from such an IP becomes an abuse row; its
messages are ignored.
- crawler hits: non-preload GETs no activity message matched within
``_CRAWLER_TIMEOUT``, plus visits reclassified as real-browser
bots (under ``_MIN_VISIT_READ`` seconds of total reading time).
Messages from bot-UAs are ignored, so their GETs never match.
- visits: the remaining messages, grouped per client with a new
visit after ``_SESSION_GAP`` of inactivity. Trail statuses come
from the client's GETs (preloads included — a cache-served
navigation's only GET is its preload); the entry referer and UTM
tags from the GET that loaded the entry page.
Hidden (admin) clients are excluded from every list and aggregate.
"""
in_menu = in_menu or (lambda path: False)
data = self.data
ip_of = {h: c.ip for h, c in data.clients.items()}
# --- abuse classification: each IP's document GETs, chronologically.
# Hidden (admin) clients never classify: editing means visiting
# not-found pages (that is where the create pen lives).
gets_by_ip: dict[str, list[Get]] = {}
for g in data.gets:
if not g.pre and not self._hidden(g.client):
gets_by_ip.setdefault(ip_of.get(g.client, ""), []).append(g)
menu_cache: dict[str, bool] = {}
def real_node(path: str) -> bool:
if path not in menu_cache:
menu_cache[path] = in_menu(path)
return menu_cache[path]
abuse_ips: set[str] = set()
flag_ids: set[int] = set()
for ip, gets in gets_by_ip.items():
gets.sort(key=lambda g: g.t)
telltale = False
recent: list[Get] = [] # plain 404s inside the sliding window
for g in gets:
if g.status != 404:
continue
path = g.path.split("?")[0]
if _is_well_known(path):
continue # legitimate probes, never abuse evidence
if _is_abuse_path(path):
telltale = True
flag_ids.add(id(g))
elif not real_node(path):
recent.append(g)
while g.t - recent[0].t > _ABUSE_404_WINDOW:
recent.pop(0)
if len(recent) == _ABUSE_404_THRESHOLD:
flag_ids.add(id(g))
if telltale or len(recent) >= _ABUSE_404_THRESHOLD:
abuse_ips.add(ip)
# --- per-client event lists, chronological
msgs_by_client: dict[bytes, list[Msg]] = {}
for m in data.msgs:
msgs_by_client.setdefault(m.client, []).append(m)
for msgs in msgs_by_client.values():
msgs.sort(key=lambda m: m.t)
gets_by_client: dict[bytes, list[Get]] = {}
for g in data.gets:
gets_by_client.setdefault(g.client, []).append(g)
for gets in gets_by_client.values():
gets.sort(key=lambda g: g.t)
def status_at(client_hash: bytes, path: str, t: datetime) -> int:
"""Latest status served to the client for ``path`` at or before ``t``."""
gets = gets_by_client.get(client_hash)
if not gets:
return 200
i = bisect_right(gets, t, key=lambda g: g.t)
for g in reversed(gets[:i]):
if g.path.split("?")[0] == path:
return g.status
return 200
def entry_get(client_hash: bytes, path: str, t: datetime) -> Get | None:
"""The GET that loaded ``path`` just before the message at ``t``."""
gets = gets_by_client.get(client_hash)
if not gets:
return None
i = bisect_right(gets, t, key=lambda g: g.t)
for g in reversed(gets[:i]):
if g.t < t - _CRAWLER_TIMEOUT:
break
if not g.pre and g.path.split("?")[0] == path:
return g
return None
# --- visits: group each visible, non-abuse, non-bot client's messages
visits: list[Visit] = []
for h, msgs in msgs_by_client.items():
if self._hidden(h) or ip_of.get(h, "") in abuse_ips:
continue
client = data.clients.get(h)
if client is not None and _is_bot_ua(client.ua):
continue
visit: Visit | None = None
last_t: datetime | None = None
for m in msgs:
if m.to:
if visit is None or (
last_t is not None and m.t - last_t > _SESSION_GAP
):
# First navigation ever, or after a long silence:
# start a visit. The entry referer/UTM tags come
# from the GET that loaded the entry page. Only an
# internal page load can open a visit — an external
# exit without an open visit is dropped.
if not m.to.startswith("/"):
last_t = m.t
continue
visit = Visit(start=m.t, entry=m.to, client=h)
g = entry_get(h, m.to, m.t)
if g is not None:
visit.referer = g.ref
visit.utm = _utm_tags(
g.path.split("?", 1)[1] if "?" in g.path else ""
)
visit.trail[m.t] = TrailItem(
to=m.to, status=status_at(h, m.to, m.t)
)
visits.append(visit)
else:
fr = m.fr or "(direct)"
visit.navs[m.t] = Nav(fr=fr, to=m.to)
status = status_at(h, m.to, m.t)
# First-seen only: repeat pages and repeated exits
# update the existing trail item instead of appending.
for item in visit.trail.values():
if item.to == m.to:
item.status = status
break
else:
visit.trail[m.t] = TrailItem(to=m.to, status=status)
if m.read > 0 and m.fr and visit is not None:
for item in visit.trail.values():
if item.to == m.fr:
item.read += m.read
break
last_t = m.t
# --- crawler hits: document GETs no message matched
crawlers: list[CrawlerHit] = []
msg_times = {h: [m.t for m in msgs] for h, msgs in msgs_by_client.items()}
for g in data.gets:
if g.pre or self._hidden(g.client):
continue
if ip_of.get(g.client, "") in abuse_ips:
continue
client = data.clients.get(g.client)
if client is None or not _is_bot_ua(client.ua):
msgs = msgs_by_client.get(g.client, [])
times = msg_times.get(g.client, [])
stripped = g.path.split("?")[0]
matched = False
for m in msgs[bisect_left(times, g.t) :]:
if m.t - g.t > _CRAWLER_TIMEOUT:
break
if m.to == stripped:
matched = True
break
if matched:
continue
entry, _, query = g.path.partition("?")
crawlers.append(
CrawlerHit(
start=g.t,
entry=entry,
client=g.client,
referer=g.ref,
query=query,
status=g.status,
)
)
# --- real-browser bots: visits with too little reading time become
# crawler hits (one per internal trail page) and count nowhere
kept: list[Visit] = []
for visit in visits:
if sum(item.read for item in visit.trail.values()) >= _MIN_VISIT_READ:
kept.append(visit)
continue
query = urlencode(visit.utm)
first = True
for t, item in visit.trail.items():
if not item.to.startswith("/"):
continue
crawlers.append(
CrawlerHit(
start=t,
entry=item.to,
client=visit.client,
referer=visit.referer if first else "",
query=query if first else "",
status=item.status,
)
)
first = False
display = Display(
visits=kept,
crawlers=crawlers,
abuse=[
AbuseHit(
start=g.t,
path=g.path,
client=g.client,
flag=id(g) in flag_ids,
is_404=g.status != 200,
)
for g in data.gets
if not g.pre
and not self._hidden(g.client)
and ip_of.get(g.client, "") in abuse_ips
],
clients={h: c for h, c in data.clients.items() if not c.hide},
favicons={
origin: f"/_f/{f.file}"
for origin, f in data.favicons.items()
if f.file
},
)
for visit in kept:
bucket = _bucket(visit.start)
site = display.site_visits
site[bucket] = site.get(bucket, 0) + 1
entry_views = display.views.setdefault(visit.entry, {})
entry_views[bucket] = entry_views.get(bucket, 0) + 1
fr = visit.referer or "(direct)"
buckets = display.transitions.setdefault(fr, {}).setdefault(visit.entry, {})
buckets[bucket] = buckets.get(bucket, 0) + 1
for t, nav in visit.navs.items():
nb = _bucket(t)
if nav.to.startswith("/"):
nav_views = display.views.setdefault(nav.to, {})
nav_views[nb] = nav_views.get(nb, 0) + 1
nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(
nav.to, {}
)
nbuckets[nb] = nbuckets.get(nb, 0) + 1
return display
def display_json(self, in_menu: Callable[[str], bool] | None = None) -> str:
"""The ``display()`` payload as a JSON string for the WebSocket."""
return msgspec.json.encode(self.display(in_menu)).decode()
+654
View File
@@ -0,0 +1,654 @@
"""Editor REST API and WebSocket sessions.
The management endpoints behind the SSO forward-auth gate: the site tree
(``/_api/pages``), structure operations (``/_api/structure``), site-wide
settings (``/_api/settings``), task-list toggles (``/_api/toggle-task``),
the translations refresh (``/_api/translations``), the editor session
socket (``/_api/ws/editor``), and the translator service channel
(``/_translate/{clientkey}`` — deliberately NOT under ``/_api``: the
server-generated key in the path is the access control).
"""
import logging
from datetime import UTC, datetime
from fastapi import (
APIRouter,
HTTPException,
Request,
WebSocket,
WebSocketDisconnect,
)
from pydantic import BaseModel
from pagerite import i18n, views
from pagerite.chunks import store_chunks
from pagerite.data import (
Node,
append_order,
find_slot,
node_markdown,
resolve,
sorted_nodes,
)
from pagerite.markdown import render, toggle_task
from pagerite.state import (
_check_reserved,
_ensure,
_invalidate_pages,
_remove_page,
data,
dispatcher,
kanta,
)
logger = logging.getLogger(__name__)
router = APIRouter()
class PageIn(BaseModel):
"""Payload for creating or replacing a page."""
title: str
markdown: str
published: bool = True
banner: str | None = None # None keeps the existing banner
@router.get("/_api/pages")
async def list_pages(lang: str | None = None) -> list[dict]:
"""The site tree for the structure editor (all nodes, drafts included).
Nested by slug; each node carries its full path, menu order, flags and
language settings (``language`` is the node's own primary-language
setting, "" = inherit; ``primary`` is the resolved effective one).
With a ``?lang=`` translation, titles come out in that language where a
translation exists (``translated`` flags it — true trivially for rows
whose primary language IS the selected one; other rows fall back to
the original title, dimmed) — the structure itself (slugs, order,
hierarchy) is language-independent.
"""
tag = i18n.base_tag(lang or "")
titles = i18n.title_map(data, tag) if tag else {}
def dump(nodes: dict[str, Node], prefix: str, inherited: str) -> list[dict]:
out = []
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
primary = node.language or inherited
out.append(
{
"slug": slug,
"path": path,
"title": titles.get(path) or node.title,
"translated": path in titles or (bool(tag) and primary == tag),
"order": node.order,
"published": node.published,
"has_content": node.chunks is not None,
"language": node.language,
"primary": primary,
"children": dump(node.children, path, primary),
}
)
return out
return dump(data.menu, "", i18n.ORIGINAL_LANGUAGE)
@router.put("/_api/pages/{path:path}", status_code=204)
async def save_page(
path: str, page: PageIn, request: Request, lang: str | None = None
) -> None:
"""Create or replace the page at a slug path ("" or "/" = front page).
Missing ancestors are created as content-less category labels. Giving
a category markdown turns it into a landing page. Empty markdown (after
stripping) creates an empty page that renders with just its title —
saving never deletes; use DELETE to remove a page (the page editor
issues DELETE when you save empty text).
With a ``?lang=`` query (a translation, not the primary language) the
save is a translated-view edit (docs/localization.md): the markdown is
diffed against the currently served hybrid and the minimal diff is
appended as a Patch under ``patches[f"{path}:{lang}"]`` — node.chunks
and the original-language fields (title, published, banner) stay
untouched.
"""
path = path.strip("/")
_check_reserved(path)
lang = i18n.base_tag(lang or "")
if lang and lang != i18n.primary_lang(data.menu, path):
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None:
raise HTTPException(404, "no such page")
with kanta.transaction(
f"page:{lang}", user=request.headers.get("remote-user"), extra=path
):
# Patches alone make the translated version exist.
if i18n.add_patch(data, node, path, lang, page.markdown):
_invalidate_pages()
return
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
node = _ensure(data.menu, path)
node.title = page.title
node.chunks = store_chunks(data.chunks, page.markdown)
node.published = page.published
if page.banner is not None:
node.banner = page.banner
node.modified = datetime.now(UTC)
_invalidate_pages()
@router.delete("/_api/pages/{path:path}", status_code=204)
async def delete_page(path: str, request: Request) -> None:
"""Delete a node by slug path.
A category (node with children) loses only its landing page and stays
as a content-less label; a childless node is removed entirely.
"""
path = path.strip("/")
_check_reserved(path)
with kanta.transaction(
"page:delete", user=request.headers.get("remote-user"), extra=path
):
if not _remove_page(data.menu, path):
raise HTTPException(404, "no such page")
_invalidate_pages()
class StructureOp(BaseModel):
"""Rearrange the site tree: reorder, move/rename or retitle a node.
`order` is a fresh fractional key computed client-side from the node's
new siblings (a value halfway between them); all other items keep
theirs. `move_to` is the full target path — the parent must exist and
the new slug be free. Moves carry the whole subtree. The front page is
just the top-level node with slug "": renaming it away leaves no front
page ("/" then redirects to the first nav item), and any childless
top-level node can take the empty slug to become the front page.
With `lang` (a translation, not the node's primary language) a `title`
edit writes a per-language title fragment instead of the original — the
same storage as machine title translations (docs/localization.md);
sending the original's text removes the override. Structural fields are
not combinable with a translated title edit.
`language` sets the node's primary language (a BCP-47 base tag; "" =
inherit from the nearest ancestor, the front page last, site default
"en" final — Node.language), inherited by the whole subtree.
"""
path: str
order: float | None = None
move_to: str | None = None
title: str | None = None
lang: str | None = None
language: str | None = None
@router.post("/_api/structure", status_code=204)
async def update_structure(op: StructureOp, request: Request) -> None:
"""Apply one structure operation (see StructureOp)."""
path = op.path.strip("/")
chain = resolve(data.menu, path)
if chain is None:
raise HTTPException(404, "no such page")
node = chain[-1]
lang = i18n.base_tag(op.lang or "")
if op.language is not None:
# Primary-language setting (inherited by the subtree): reselects
# what "the original" means for the node — its language is part of
# every render, so a change invalidates everywhere.
language = i18n.base_tag(op.language)
with kanta.transaction(
"page:language", user=request.headers.get("remote-user"), extra=path
):
if language != node.language:
node.language = language
_invalidate_pages()
return
if op.title is not None and lang and lang != i18n.primary_lang(data.menu, path):
# Translated title (i18n.set_title_translation): original title,
# slugs and hierarchy stay untouched.
with kanta.transaction(
f"page:{lang}:title", user=request.headers.get("remote-user"), extra=path
):
if i18n.set_title_translation(data, node, lang, op.title):
_invalidate_pages()
return
target = op.move_to.strip("/") if op.move_to is not None else None
if target is not None and target != path:
_check_reserved(target)
if path and target.startswith(f"{path}/"):
raise HTTPException(400, "cannot move a page under itself")
slot = find_slot(data.menu, target)
if slot is None:
raise HTTPException(404, "target parent does not exist")
tnodes, tslug = slot
if tslug in tnodes:
raise HTTPException(400, "target path exists")
if not tslug and node.children:
raise HTTPException(400, "the front page cannot have children")
# One structure call can combine a title set, a move/rename and a
# reorder; the action names the most significant of them.
action = (
"page:slug"
if target is not None and target != path
else "page:title"
if op.title is not None
else "structure:reorder"
)
with kanta.transaction(action, user=request.headers.get("remote-user"), extra=path):
if op.title is not None:
node.title = op.title
if target is not None and target != path:
snodes, sslug = find_slot(data.menu, path)
del snodes[sslug]
# A pure rename (same parent) keeps its position; only a move
# to another level appends at the end (unless an order came
# with the drop).
same_level = path.rpartition("/")[0] == target.rpartition("/")[0]
node.order = (
op.order
if op.order is not None
else node.order
if same_level
else append_order(tnodes)
)
tnodes[tslug] = node
elif op.order is not None:
node.order = op.order
node.modified = datetime.now(UTC)
_invalidate_pages()
@router.get("/_api/settings")
async def get_settings() -> dict:
"""Site-wide settings (brand, theme, custom CSS and favicon URL), plus
the themes, banner designs and user fonts available on disk for the
selectors, the translator service keys and the wanted translation
languages (for the /_translate socket)."""
return {
"brand": data.brand,
"brand_html": data.brand_html,
"theme": data.theme,
"custom_css": data.custom_css,
"favicon": f"/_f/{data.favicon}" if data.favicon else "",
"themes": views._theme_info(),
"banner_designs": views._banner_design_names(),
"fonts": views._user_fonts(),
"transition": data.transition,
"transitions": views._transition_names(),
"translate_keys": data.translate_keys,
# The site default primary language: the front page's resolved
# setting (every page may override it, inherited down the tree).
"primary_lang": i18n.primary_lang(data.menu, ""),
"translate_langs": sorted(data.translate_langs),
}
class SettingsIn(BaseModel):
"""Payload for updating site-wide settings."""
brand: str
theme: str
custom_css: str
brand_html: str = ""
transition: str = "cube"
translate_langs: list[str] | None = None # None keeps the current set
translate_keys: dict[str, str] | None = None # None keeps the current keys
@router.put("/_api/settings", status_code=204)
async def put_settings(settings: SettingsIn, request: Request) -> None:
"""Update site-wide settings; invalidates cached pages and ETags."""
with kanta.transaction("settings", user=request.headers.get("remote-user")):
data.brand = settings.brand
data.brand_html = settings.brand_html
data.theme = settings.theme
data.custom_css = settings.custom_css
data.transition = settings.transition
if settings.translate_langs is not None:
# Any language may be a target — including the site default
# (an article in another language can be translated INTO it);
# a node's own primary is excluded per article, not here.
data.translate_langs = {
tag: True
for lang in settings.translate_langs
if (tag := i18n.base_tag(lang))
}
if settings.translate_keys is not None:
data.translate_keys = settings.translate_keys
_invalidate_pages()
@router.delete("/_api/translations", status_code=204)
async def delete_translations(request: Request) -> None:
"""Drop all machine translations (Data.trans) so the dispatcher
re-translates everything from scratch (a translate:reset action:
the invalidation hook re-offers every fragment to connected
translators). User patches are kept; the availability index
(node.langs) is rebuilt from them — patches alone still make a language
exist on a page."""
with kanta.transaction("translate:reset", user=request.headers.get("remote-user")):
i18n.clear_translations(data)
_invalidate_pages()
# Fragments rejected this run (segment validation) stay skipped no
# longer: a refresh is precisely the "another chance" for them.
dispatcher.reset_validation_failures()
class ToggleTaskIn(BaseModel):
"""Payload for toggling one task-list checkbox."""
path: str
index: int
markdown: str | None = None
@router.post("/_api/toggle-task")
async def toggle_task_endpoint(body: ToggleTaskIn, request: Request) -> dict[str, str]:
"""Toggle the Nth task-list checkbox in a page's Markdown source.
If ``markdown`` is provided the source is left untouched and the toggled
Markdown is returned (used while the page editor is open, so the live
CodeMirror document can be updated). Otherwise the stored page at
``path`` is read, toggled, and saved.
"""
path = body.path.strip("/")
_check_reserved(path)
if body.markdown is not None:
new_markdown = toggle_task(body.markdown, body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
return {"markdown": new_markdown}
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None:
raise HTTPException(404, "no such page")
new_markdown = toggle_task(node_markdown(data, node) or "", body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
with kanta.transaction("page", user=request.headers.get("remote-user"), extra=path):
# Re-chunk like any save: only the chunk containing the toggled
# checkbox gets a new hash, the rest keep theirs.
node.chunks = store_chunks(data.chunks, new_markdown)
node.modified = datetime.now(UTC)
_invalidate_pages()
return {"markdown": new_markdown}
# WebSocket API for external translation services (not under /_api: it is keyed
# with Data.translate_keys instead of the SSO forward-auth). The dispatcher —
# protocol, connected clients and the job pipeline — lives in translate.py.
@router.websocket("/_translate/{clientkey}")
async def translate_ws(ws: WebSocket, clientkey: str) -> None:
"""Translator service channel (docs/localization.md).
Deliberately NOT under /_api/: the external forward-auth is skipped;
the server-generated client key in the path is the access control
(``Data.translate_keys``: key -> display name; the first is generated
at bootstrap, all are shown in the admin's /_api/settings).
"""
await dispatcher.handle_ws(ws, clientkey)
@router.websocket("/_api/ws/editor")
async def editor_ws(ws: WebSocket) -> None:
"""Editor session: open pages, render previews, save — over one socket.
Stateless protocol (each message carries the path):
<- {"type": "open", "path", "lang"?}
-> {"type": "doc", "path", "exists", "title", "markdown", "published",
"banner", "banner_design", "lang", "primary_lang", "langs",
"translate_langs"}
<- {"type": "render", "path", "markdown"}
-> {"type": "html", "path", "html"}
<- {"type": "save", "path", "title"?, "markdown"?, "published"?,
"banner"?, "banner_design"?, "move_from"?, "lang"?, "base"?}
(absent fields keep their old values; move_from: rename/move a
page, subtree included)
-> {"type": "saved", "path"} | {"type": "error", "detail"}
With "lang" (a translation, not the primary language), open returns the
effective hybrid Markdown and title for that language plus the language
metadata the picker's UI needs; save diffs the submitted Markdown
against "base" (the editor's shadow copy of the hybrid it started from
— absent: the current hybrid) and stores it as a user Patch, and a
changed title becomes a fragment in Data.trans — node.chunks and the
other fields stay untouched (docs/localization.md).
"""
await ws.accept()
try:
while True:
msg = await ws.receive_json()
path = msg.get("path", "").strip("/")
try:
_check_reserved(path)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
match msg.get("type"):
case "open":
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
# The article's primary language: its own setting,
# inherited down the tree ("en" final fallback).
node_lang = i18n.primary_lang(data.menu, path)
lang = i18n.base_tag(str(msg.get("lang") or ""))
if lang == node_lang:
lang = ""
markdown = ""
title = node.title if node else ""
if node is not None:
markdown = node_markdown(data, node) or ""
if lang and node.chunks is not None:
# Translation view: the effective (hybrid)
# Markdown and title for that language —
# machine fragments + user patches over the
# original (docs/localization.md editor flow).
markdown = i18n.hybrid_markdown(data, node, path, lang)
title = i18n.title_map(data, lang).get(path) or title
await ws.send_json(
{
"type": "doc",
"path": path,
"exists": node is not None,
"title": title,
"markdown": markdown,
"published": node.published if node else True,
"banner": node.banner if node else "",
# Own banner design setting: null = inherit,
# "" = none, otherwise a design name.
"banner_design": node.banner_design if node else None,
# Which node's banner applies here ("" = front page,
# null = default artwork); the site editor shows it
# as the banner field's placeholder.
"banner_from": views.banner_source(data.menu, path),
# Which node's banner-design setting would apply on
# inherit ("" = front page, null = the active
# theme's default) and what design that resolves to.
"banner_design_from": (
src := views.banner_design_source(
data.menu, path, data.theme
)
),
"banner_design_inherited": (
views.banner_design(data.menu, src, data.theme)
if src is not None
else views.theme_banner_design(data.theme)
),
# Language context for the editor's picker: the
# language this Markdown represents ("" = primary),
# the page's own primary language, the translations
# this page already has, and the site-wide
# configured target languages.
"lang": lang,
"primary_lang": node_lang,
"langs": sorted(node.langs) if node else [],
"translate_langs": sorted(data.translate_langs),
}
)
case "render":
markdown = msg.get("markdown", "")
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
rendered = render(
markdown,
path,
node.created if node else None,
node.modified if node else None,
# The title is injected as h1 when the markdown has
# none; the editor's title field edits live-preview.
title=msg.get("title") or (node.title if node else ""),
# Pin section anchors to the original language so the
# preview of a translation matches the served page
# (no-op when the previewed markdown is the original).
anchors_from=(
(node_markdown(data, node) or "", node.title)
if node
else None
),
)
await ws.send_json(
{
"type": "html",
"path": path,
"html": rendered.html,
# Column-layout flag: the preview toggles the
# article's .multicol class and swaps in the
# segmented (.colseg/.cols) article html.
"multicol": rendered.multicol,
}
)
case "save":
move_from = (msg.get("move_from") or path).strip("/")
lang = i18n.base_tag(str(msg.get("lang") or ""))
translated = bool(
lang and lang != i18n.primary_lang(data.menu, move_from)
)
try:
_check_reserved(move_from)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
old_chain = resolve(data.menu, move_from)
old = old_chain[-1] if old_chain else None
if old is None and move_from != path:
move_from = path # nothing to carry over; plain save
if move_from != path:
# Rename/move: detach the node (subtree included)
# and attach it at the new path. The target slug
# must be free and the front page childless.
if move_from and path.startswith(f"{move_from}/"):
await ws.send_json(
{
"type": "error",
"detail": "cannot move a page under itself",
}
)
continue
tslug = path.rpartition("/")[2]
if not tslug and old.children:
await ws.send_json(
{
"type": "error",
"detail": "the front page cannot have children",
}
)
continue
tchain = resolve(data.menu, path)
if tchain is not None:
await ws.send_json(
{
"type": "error",
"detail": "target path exists",
}
)
continue
if translated and (
move_from != path or old is None or old.chunks is None
):
# A translated-view save patches an existing
# original; it cannot create or move pages.
await ws.send_json({"type": "error", "detail": "no such page"})
continue
if translated and "markdown" in msg and not msg["markdown"].strip():
# Saving never deletes; an emptied translation would
# render as a blank page in that language.
await ws.send_json(
{
"type": "error",
"detail": "a translation cannot be emptied",
}
)
continue
with kanta.transaction(
f"page:{lang}" if translated else "page",
user=ws.headers.get("remote-user"),
extra=path,
):
if move_from != path:
same_menu = (
move_from.rpartition("/")[0] == path.rpartition("/")[0]
)
snodes, sslug = find_slot(data.menu, move_from)
node = snodes.pop(sslug)
parent = path.rpartition("/")[0]
if parent:
_ensure(data.menu, parent)
tnodes, tslug = find_slot(data.menu, path)
node.order = (
node.order if same_menu else append_order(tnodes)
)
tnodes[tslug] = node
else:
node = old if old is not None else _ensure(data.menu, path)
if translated:
# node.chunks and the original-language fields
# stay untouched: the markdown diff (against the
# editor's shadow "base" — the hybrid it started
# from; absent: the current hybrid) is appended
# as a Patch, a changed title becomes a
# per-language title override (i18n).
changed = False
if "markdown" in msg:
base = msg.get("base")
changed = i18n.add_patch(
data,
node,
path,
lang,
msg["markdown"],
base=base if isinstance(base, str) else None,
)
if "title" in msg and node.title:
changed = (
i18n.set_title_translation(
data, node, lang, msg["title"]
)
or changed
)
if changed:
_invalidate_pages()
else:
if "markdown" in msg:
# Saving never deletes; empty markdown is an
# empty page. Deletion is an explicit choice
# by the page editor (REST DELETE).
node.chunks = store_chunks(data.chunks, msg["markdown"])
if "title" in msg:
node.title = msg["title"]
if "published" in msg:
node.published = bool(msg["published"])
if "banner" in msg:
node.banner = msg["banner"]
if "banner_design" in msg:
node.banner_design = msg["banner_design"]
node.modified = datetime.now(UTC)
_invalidate_pages()
await ws.send_json({"type": "saved", "path": path})
except WebSocketDisconnect:
pass
+85 -543
View File
@@ -1,145 +1,98 @@
"""FastAPI application: server-rendered content pages plus Vue assets. """FastAPI application assembly: server-rendered content pages plus Vue assets.
Route ordering matters: our routes are defined before The routes live in specialized modules, included below as APIRouters:
``frontend.route(app, "/")`` is called, so they take priority over
the asset routes that fastapi-vue inserts at that position during ``load()``. - ``pagerite.state`` — shared core, no routes: site constants, the kanta
The content catch-all (``/{path:path}``) is defined last, so built database, the analytics store, the fastapi-vue frontend, the render
frontend assets still win over content slugs; anything unmatched falls cache, the translator dispatcher, and the database bootstrap hooks.
through to content (and 404 if no page exists there). - ``pagerite.files`` — the content-addressed file store and its routes
(``/_api/files``, ``/_f/``, ``/_themes/``, ``/_fonts/``, favicon).
- ``pagerite.api`` — the editor REST API and WebSocket sessions
(``/_api/*``, ``/_translate/{clientkey}``).
- ``pagerite.tracking`` — visit analytics (``/_ws``, ``/_api/ws/analytics``,
the ``/_a`` viewer page).
- ``pagerite.pages`` — the public content pages: ``/``, ``/sitemap.xml``,
``/robots.txt`` and the ``/{path:path}`` catch-all.
Route ordering matters: our own routers are included before
``frontend.route(app, "/")`` is called. That call only records the current
route-table length; the actual asset routes are spliced in at that position
later, when ``frontend.load()`` runs inside the lifespan — so they take
priority over anything registered after this point but never shadow our
own routes. The content catch-all (``/{path:path}``) is included last, so
built frontend assets still win over content slugs; anything unmatched
falls through to content (and 404 if no page exists there).
The site structure is a tree of Nodes (see data.py); URL paths resolve by The site structure is a tree of Nodes (see data.py); URL paths resolve by
walking the tree (``resolve``), moves are slot detach/attach walking the tree (``resolve``), moves are slot detach/attach
(``find_slot``) with a fresh order key from the new siblings. (``find_slot``) with a fresh order key from the new siblings.
""" """
import mimetypes import asyncio
import os import logging
import re from collections.abc import AsyncGenerator
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from datetime import UTC, datetime
from pathlib import Path from pathlib import Path
import blake3 from fastapi import FastAPI, Request
from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect from fastapi.responses import Response
from fastapi.responses import HTMLResponse, RedirectResponse, Response
from fastapi_vue import Frontend from fastapi_vue import Frontend
from kanta import Kanta from starlette.types import ASGIApp, Receive, Scope, Send
from pydantic import BaseModel
from pagerite import seed, views from pagerite import api, files, pages, tracking
from pagerite.__main__ import DEVMODE from pagerite.__main__ import DEVMODE
from pagerite.data import ( from pagerite.files import file_store
Data, from pagerite.state import analytics_store, config, kanta
Node,
append_order,
find_slot,
prettify,
resolve,
sorted_nodes,
)
from pagerite.markdown import has_h1, render, toggle_task
DB_PATH = os.getenv("PAGERITE_DB", "pagerite.kantadb") logger = logging.getLogger(__name__)
# Our own data root; kanta edits it in place, reads are plain attribute access.
data = Data()
kanta = Kanta(DB_PATH, data)
# Vue build served at the site root, no SPA catch-all (assets only). The # Vue build served at the site root, no SPA catch-all (assets only). The
# build mirrors the URL space: hashed, immutable files live under # build mirrors the URL space: hashed, immutable files live under
# /_assets/ (assetsDir: '_/assets'), the favicon at /favicon.ico. # /_assets/ (assetsDir: '_/assets').
BUILD_DIR = Path(__file__).with_name("frontend-build") frontend = Frontend(
frontend = Frontend(BUILD_DIR, spa=False, cached="/_assets/") Path(__file__).with_name("frontend-build"), spa=False, cached="/_assets/"
)
def _hash_name(body: bytes, orig: str) -> str: class _AccessLogExtraMiddleware:
"""Content-addressed file name: blake3 hash prefix + original extension.""" """Fill the ``log_extra`` slot of fastapi_vue's access log.
ext = "".join(c for c in Path(orig).suffix.lower() if c.isalnum() or c == ".")
return blake3.blake3(body).hexdigest()[:12] + ext
Everything under ``/_api`` is gated by the SSO forward-auth, which names
def _store_seed_file(markdown: str, banner: str, orig: str, body: bytes) -> tuple[str, str]: the authenticated user in the ``remote-user`` header; put that user on
"""Store a seed file content-addressed and point references at /_f/.""" the access-log line, for plain requests and WebSocket open/close alike.
name = _hash_name(body, orig) The scope dict is shared with the outer AccessLogMiddleware, which reads
data.files.setdefault(name, body) the slot back at response/accept/close time.
markdown = markdown.replace(f"]({orig}", f"](/_f/{name}")
banner = banner.replace(f'src="/{orig}"', f'src="/_f/{name}"')
banner = banner.replace(f'src="{orig}"', f'src="/_f/{name}"')
return markdown, banner
def _ensure(menu: dict[str, Node], path: str) -> Node:
"""Return the node at ``path``, creating it and any missing ancestors
(content-less category labels) appended at the end of their level."""
nodes = menu
node = None
for seg in path.split("/"):
node = nodes.get(seg)
if node is None:
node = Node(title=prettify(seg), order=append_order(nodes))
nodes[seg] = node
nodes = node.children
return node
def _remove_page_content(menu: dict[str, Node], path: str) -> None:
"""Delete a page's markdown content.
A node with children becomes a content-less category label; a childless
node is removed entirely. Does nothing if the path does not exist.
""" """
slot = find_slot(menu, path)
if slot is None:
return
node = slot[0].get(slot[1])
if node is None:
return
if node.children:
node.content = None
node.modified = datetime.now(UTC)
else:
del slot[0][slot[1]]
def __init__(self, app: ASGIApp) -> None:
self.app = app
def _migrate_legacy() -> None: async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
"""Rebuild the legacy flat page store as a tree (one-time migration).""" if scope["type"] in ("http", "websocket") and scope["path"].startswith("/_api"):
if not data.pages: headers = dict(scope["headers"])
return user = headers.get(b"remote-user", b"").decode("latin-1")
with kanta.transaction("migrate pages to tree"): if user:
for path, page in data.pages.items(): scope.setdefault("state", {})["log_extra"] = user
node = _ensure(data.menu, path) await self.app(scope, receive, send)
node.title = page.title
node.content = page.markdown
node.banner = page.banner
node.published = page.published
node.order = page.order
node.created = page.created
node.modified = page.modified
data.pages.clear()
data.version += 1
@asynccontextmanager @asynccontextmanager
async def lifespan(_app: FastAPI) -> AsyncIterator[None]: async def lifespan(_app: FastAPI) -> AsyncGenerator:
"""Open the database, migrate/seed content, load assets.""" """Open the database (migrations run inside kanta.open), load assets, load GeoIP."""
await kanta.open() async with kanta:
_migrate_legacy() await asyncio.to_thread(file_store.load)
missing = [p for p in seed.PAGES if resolve(data.menu, p) is None] await frontend.load()
if missing: # --dbip: update the DB-IP database first, then decompress/open the
with kanta.transaction("seed missing pages"): # MMDB once. Lookups are then read-only and safe to run in
for path in missing: # background ``to_thread`` workers.
title, markdown, files, banner, order = seed.PAGES[path] if config.dbip:
for orig, body in files.items(): await asyncio.to_thread(tracking._download_dbip)
markdown, banner = _store_seed_file(markdown, banner, orig, body) await asyncio.to_thread(tracking._geoip._load)
node = _ensure(data.menu, path) analytics_store.subscribe(tracking._schedule_analytics_broadcast)
node.title = title # Backfill favicons for external sites already in the recorded data.
node.content = markdown tracking._schedule_favicon_fetch()
node.banner = banner yield
node.order = order analytics_store.unsubscribe(tracking._schedule_analytics_broadcast)
await frontend.load()
yield
await kanta.close()
# docs_url/openapi_url disabled: /docs belongs to our content, and the API # docs_url/openapi_url disabled: /docs belongs to our content, and the API
@@ -153,438 +106,27 @@ app = FastAPI(
openapi_url=None, openapi_url=None,
) )
app.add_middleware(_AccessLogExtraMiddleware)
class PageIn(BaseModel):
"""Payload for creating or replacing a page."""
title: str
markdown: str
published: bool = True
banner: str | None = None # None keeps the existing banner
@app.get("/_api/pages") @app.middleware("http")
async def list_pages() -> list[dict]: async def _headers(request: Request, call_next) -> Response:
"""The site tree for the structure editor (all nodes, drafts included). """Replace uvicorn's default Server header with ours (no version)."""
response = await call_next(request)
Nested by slug; each node carries its full path, menu order and flags. response.headers["server"] = "pagerite"
""" return response
def dump(nodes: dict[str, Node], prefix: str) -> list[dict]:
out = []
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
out.append({
"slug": slug,
"path": path,
"title": node.title,
"order": node.order,
"published": node.published,
"has_content": node.content is not None,
"children": dump(node.children, path),
})
return out
return dump(data.menu, "")
@app.put("/_api/pages/{path:path}", status_code=204) # Our own routes first: the editor API and translator socket, the analytics
async def save_page(path: str, page: PageIn) -> None: # machinery, and the file store/user assets.
"""Create or replace the page at a slug path ("" or "/" = front page). app.include_router(api.router)
app.include_router(tracking.router)
Missing ancestors are created as content-less category labels. Giving app.include_router(files.router)
a category markdown turns it into a landing page. Empty markdown (after
stripping) creates an empty page that renders with just its title —
saving never deletes; use DELETE to remove a page (the page editor
issues DELETE when you save empty text).
"""
path = path.strip("/")
_check_reserved(path)
with kanta.transaction("save page", extra=path):
node = _ensure(data.menu, path)
node.title = page.title
node.content = page.markdown
node.published = page.published
if page.banner is not None:
node.banner = page.banner
node.modified = datetime.now(UTC)
data.version += 1
class StructureOp(BaseModel):
"""Rearrange the site tree: reorder, move/rename or retitle a node.
`order` is a fresh fractional key computed client-side from the node's
new siblings (a value halfway between them); all other items keep
theirs. `move_to` is the full target path — the parent must exist and
the new slug be free. Moves carry the whole subtree. The front page is
just the top-level node with slug "": renaming it away leaves no front
page ("/" then redirects to the first nav item), and any childless
top-level node can take the empty slug to become the front page.
"""
path: str
order: float | None = None
move_to: str | None = None
title: str | None = None
@app.post("/_api/structure", status_code=204)
async def update_structure(op: StructureOp) -> None:
"""Apply one structure operation (see StructureOp)."""
path = op.path.strip("/")
chain = resolve(data.menu, path)
if chain is None:
raise HTTPException(404, "no such page")
node = chain[-1]
target = op.move_to.strip("/") if op.move_to is not None else None
if target is not None and target != path:
_check_reserved(target)
if path and target.startswith(f"{path}/"):
raise HTTPException(400, "cannot move a page under itself")
slot = find_slot(data.menu, target)
if slot is None:
raise HTTPException(404, "target parent does not exist")
tnodes, tslug = slot
if tslug in tnodes:
raise HTTPException(400, "target path exists")
if not tslug and node.children:
raise HTTPException(400, "the front page cannot have children")
with kanta.transaction("update structure", extra=path):
if op.title is not None:
node.title = op.title
if target is not None and target != path:
snodes, sslug = find_slot(data.menu, path)
del snodes[sslug]
# A pure rename (same parent) keeps its position; only a move
# to another level appends at the end (unless an order came
# with the drop).
same_level = path.rpartition("/")[0] == target.rpartition("/")[0]
node.order = (
op.order
if op.order is not None
else node.order if same_level else append_order(tnodes)
)
tnodes[tslug] = node
elif op.order is not None:
node.order = op.order
node.modified = datetime.now(UTC)
data.version += 1
@app.get("/_api/settings")
async def get_settings() -> dict[str, str]:
"""Site-wide settings (brand, theme and custom CSS)."""
return {"brand": data.brand, "theme": data.theme, "custom_css": data.custom_css}
class SettingsIn(BaseModel):
"""Payload for updating site-wide settings."""
brand: str
theme: str
custom_css: str
@app.put("/_api/settings", status_code=204)
async def put_settings(settings: SettingsIn) -> None:
"""Update site-wide settings; bumps the version so ETags invalidate."""
with kanta.transaction("update settings"):
data.brand = settings.brand
data.theme = settings.theme
data.custom_css = settings.custom_css
data.version += 1
class ToggleTaskIn(BaseModel):
"""Payload for toggling one task-list checkbox."""
path: str
index: int
markdown: str | None = None
@app.post("/_api/toggle-task")
async def toggle_task_endpoint(body: ToggleTaskIn) -> dict[str, str]:
"""Toggle the Nth task-list checkbox in a page's Markdown source.
If ``markdown`` is provided the source is left untouched and the toggled
Markdown is returned (used while the page editor is open, so the live
CodeMirror document can be updated). Otherwise the stored page at
``path`` is read, toggled, and saved.
"""
path = body.path.strip("/")
_check_reserved(path)
if body.markdown is not None:
new_markdown = toggle_task(body.markdown, body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
return {"markdown": new_markdown}
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.content is None:
raise HTTPException(404, "no such page")
new_markdown = toggle_task(node.content, body.index)
if new_markdown is None:
raise HTTPException(400, "invalid task index")
with kanta.transaction("toggle task", extra=path):
node.content = new_markdown
node.modified = datetime.now(UTC)
data.version += 1
return {"markdown": new_markdown}
@app.put("/_api/files/{name}")
async def upload_file(name: str, request: Request) -> dict[str, str]:
"""Store an upload (image, video...) in the content-addressed store.
The stored name is a blake3 hash prefix + the original extension,
served immutable at "/_f/{name}"; returns {"path": "/_f/..."}.
"""
if "/" in name or name in {".", ".."}:
raise HTTPException(400, "bad file name")
body = await request.body()
stored = _hash_name(body, name)
with kanta.transaction("upload file", extra=name):
data.files[stored] = body
data.version += 1
return {"path": f"/_f/{stored}"}
@app.delete("/_api/files/{name}", status_code=204)
async def delete_file(name: str) -> None:
"""Remove a file from the content-addressed store (no refcounting:
other pages referencing the same content will 404)."""
if name not in data.files:
raise HTTPException(404, "no such file")
with kanta.transaction("delete file", extra=name):
del data.files[name]
data.version += 1
@app.get("/_f/{name}")
async def stored_file(name: str, request: Request) -> Response:
"""Serve a file from the content-addressed store (immutable: the name
is its own hash, so cache forever)."""
body = data.files.get(name)
if body is None:
raise HTTPException(404)
if request.headers.get("if-none-match") == name:
return Response(status_code=304)
mime = mimetypes.guess_type(name)[0] or "application/octet-stream"
return Response(
body,
media_type=mime,
headers={"etag": name, "cache-control": "public, max-age=31536000, immutable"},
)
@app.delete("/_api/pages/{path:path}", status_code=204)
async def delete_page(path: str) -> None:
"""Delete a node by slug path.
A category (node with children) loses only its landing page and stays
as a content-less label; a childless node is removed entirely.
"""
path = path.strip("/")
_check_reserved(path)
slot = find_slot(data.menu, path)
node = slot[0].get(slot[1]) if slot else None
if node is None:
raise HTTPException(404, "no such page")
with kanta.transaction("delete page", extra=path):
if node.children:
node.content = None
node.modified = datetime.now(UTC)
else:
del slot[0][slot[1]]
data.version += 1
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
def _is_reserved(path: str) -> bool:
"""Slug shape that content may never use: each segment must be lower-case
ASCII letters, digits, hyphens and underscores (underscores may not be
the first character), and dots are never allowed.
"""
if path == "":
return False
return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
def _check_reserved(path: str) -> None:
"""Reject paths that do not follow the slug charset."""
if _is_reserved(path):
raise HTTPException(
400,
'slugs may only use a-z, 0-9, "-" and "_" (not as the first character), and no dots',
)
@app.websocket("/_api/ws/editor")
async def editor_ws(ws: WebSocket) -> None:
"""Editor session: open pages, render previews, save — over one socket.
Stateless protocol (each message carries the path):
<- {"type": "open", "path"}
-> {"type": "doc", "path", "exists", "title", "markdown", "published",
"banner"}
<- {"type": "render", "path", "markdown"}
-> {"type": "html", "path", "html"}
<- {"type": "save", "path", "title"?, "markdown"?, "published"?,
"banner"?, "move_from"?} (absent fields keep their old values;
move_from: rename/move a page, subtree included)
-> {"type": "saved", "path"} | {"type": "error", "detail"}
"""
await ws.accept()
try:
while True:
msg = await ws.receive_json()
path = msg.get("path", "").strip("/")
try:
_check_reserved(path)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
match msg.get("type"):
case "open":
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
await ws.send_json({
"type": "doc",
"path": path,
"exists": node is not None,
"title": node.title if node else "",
"markdown": node.content if node and node.content is not None else "",
"published": node.published if node else True,
"banner": node.banner if node else "",
# Which node's banner applies here ("" = front page,
# null = default artwork); the site editor shows it
# as the banner field's placeholder.
"banner_from": views.banner_source(data.menu, path),
})
case "render":
markdown = msg.get("markdown", "")
await ws.send_json({
"type": "html",
"path": path,
"html": render(markdown, path),
"has_h1": has_h1(markdown),
})
case "save":
move_from = (msg.get("move_from") or path).strip("/")
try:
_check_reserved(move_from)
except HTTPException:
await ws.send_json({"type": "error", "detail": "reserved path"})
continue
old_chain = resolve(data.menu, move_from)
old = old_chain[-1] if old_chain else None
if old is None and move_from != path:
move_from = path # nothing to carry over; plain save
if move_from != path:
# Rename/move: detach the node (subtree included)
# and attach it at the new path. The target slug
# must be free and the front page childless.
if move_from and path.startswith(f"{move_from}/"):
await ws.send_json({
"type": "error",
"detail": "cannot move a page under itself",
})
continue
tslug = path.rpartition("/")[2]
if not tslug and old.children:
await ws.send_json({
"type": "error",
"detail": "the front page cannot have children",
})
continue
tchain = resolve(data.menu, path)
if tchain is not None:
await ws.send_json({
"type": "error",
"detail": "target path exists",
})
continue
with kanta.transaction("editor save", extra=path):
if move_from != path:
same_menu = (
move_from.rpartition("/")[0] == path.rpartition("/")[0]
)
snodes, sslug = find_slot(data.menu, move_from)
node = snodes.pop(sslug)
parent = path.rpartition("/")[0]
if parent:
_ensure(data.menu, parent)
tnodes, tslug = find_slot(data.menu, path)
node.order = (
node.order if same_menu else append_order(tnodes)
)
tnodes[tslug] = node
else:
node = old if old is not None else _ensure(data.menu, path)
if "markdown" in msg:
# Saving never deletes; empty markdown is an
# empty page. Deletion is an explicit choice by
# the page editor (REST DELETE).
node.content = msg["markdown"]
if "title" in msg:
node.title = msg["title"]
if "published" in msg:
node.published = bool(msg["published"])
if "banner" in msg:
node.banner = msg["banner"]
node.modified = datetime.now(UTC)
data.version += 1
await ws.send_json({"type": "saved", "path": path})
except WebSocketDisconnect:
pass
@app.get("/")
async def front_page(request: Request) -> Response:
"""Render the front page (slug path "")."""
return await show_page(request, "")
# Vue build asset routes are inserted at this position during load(): the # Vue build asset routes are inserted at this position during load(): the
# build mirrors the URL space (/_assets/*, /favicon.ico at the root). # build mirrors the URL space (/_assets/*).
frontend.route(app, "/") frontend.route(app, "/")
# The content catch-all goes last: built assets win over content slugs,
@app.get("/{path:path}", response_model=None) # anything unmatched falls through to content (and 404).
async def show_page(request: Request, path: str) -> HTMLResponse | Response: app.include_router(pages.router)
"""Render the content page at a slug path, or 404.
A node without content is a category label: its URL renders a
placeholder page (nav links point straight at its first child).
"""
path = path.strip("/")
if path and _is_reserved(path):
# Reserved slug shape: never content — no tree lookup.
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme), 404)
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is not None and node.published and node.content is not None:
# ETag on content + render version; clients revalidate cheaply,
# which keeps prefetched pages warm and current.
etag = f'"{path}@{node.modified.timestamp()}v{data.version}"'
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
return HTMLResponse(
views.render_page(data.menu, path, data.brand, data.custom_css, data.theme),
headers={"etag": etag},
)
if node is not None and node.published and node.content is None:
# Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real).
return HTMLResponse(views.render_category(data.menu, path, data.brand, data.custom_css, data.theme), 404)
if node is None and not path:
# No front page (no top-level node with slug ""): "/" opens the
# first item of the navigation instead.
for slug, item in sorted_nodes(data.menu):
if item.published:
return RedirectResponse(f"/{slug}")
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme), 404)
+164
View File
@@ -0,0 +1,164 @@
"""Block-level Markdown chunking for content-addressed storage.
A page's Markdown is split into deterministic block-level chunks, each
stored once under its content hash in ``Data.chunks`` (docs/migrate.md).
Shared by the render/save pipeline (app.py, views.py, i18n.py) and the
schema migration (migrations.py), so a chunk's key is stable no matter
where the split happens.
"""
import re
import blake3
from pagerite.segments import has_prose
#: Fenced code block opener/closer: up to 3 spaces indent, then 3+
#: backticks or tildes (CommonMark).
_FENCE_OPEN = re.compile(r"^ {0,3}(`{3,}|~{3,})")
#: HTML block openers that may span blank lines (CommonMark types 1-5:
#: script/pre/style/textarea, comments, processing instructions,
#: declarations, CDATA) with their closing condition. Other HTML blocks
#: end at the first blank line, which the generic blank-line split
#: already does.
_HTML_ATOMIC = (
(
re.compile(r"^ {0,3}<(?:script|pre|style|textarea)(?:\s|>|$)", re.I),
re.compile(r"</(?:script|pre|style|textarea)\s*>", re.I),
),
(re.compile(r"^ {0,3}<!--"), re.compile(r"-->")),
(re.compile(r"^ {0,3}<\?"), re.compile(r"\?>")),
(re.compile(r"^ {0,3}<!\[CDATA\["), re.compile(r"\]\]>")),
(re.compile(r"^ {0,3}<![A-Za-z]"), re.compile(r">")),
)
#: First line of a generic HTML block (a block-level tag).
_HTML_TAG = re.compile(r"^ {0,3}</?[A-Za-z][^>]*>")
def _fence_close(line: str, opener: str) -> bool:
"""True when ``line`` closes a code fence opened by ``opener``: the
same marker char, at least as many, and nothing else on the line."""
stripped = line.strip()
return (
len(stripped) >= len(opener)
and stripped[0] == opener[0]
and set(stripped) == {opener[0]}
)
def chunk_markdown(markdown: str) -> list[str]:
"""Split Markdown into block-level chunks, deterministically.
Blocks are separated by blank lines; fenced code blocks and the
multi-line HTML blocks (comments, script/pre/style, CDATA...) are
kept atomic, even across blank lines, and end at their closing
condition. Chunks carry no surrounding blank lines and no trailing
newline; rejoining with ``join_chunks`` reproduces the source modulo
blank-line normalization.
"""
chunks: list[str] = []
buf: list[str] = []
fence = "" # opener marker of the code fence we are in ("" = outside)
html_end: re.Pattern | None = None # closes the atomic HTML block we are in
def flush() -> None:
text = "\n".join(buf).strip("\n")
if text.strip():
chunks.append(text)
buf.clear()
for line in markdown.split("\n"):
if fence:
buf.append(line)
if _fence_close(line, fence):
fence = ""
flush()
continue
if html_end is not None:
buf.append(line)
if html_end.search(line):
html_end = None
flush()
continue
if not line.strip():
flush()
continue
if m := _FENCE_OPEN.match(line):
# Fences interrupt paragraphs (CommonMark): start a new block.
flush()
fence = m.group(1)
buf.append(line)
continue
if not buf:
for open_re, close_re in _HTML_ATOMIC:
if open_re.match(line):
buf.append(line)
if close_re.search(line): # opens and closes on one line
flush()
else:
html_end = close_re
break
else:
buf.append(line)
continue
buf.append(line)
flush() # an unterminated fence/HTML block runs to EOF, kept as code/HTML
return chunks
def _normalize(text: str) -> str:
"""Whitespace-insensitive chunk identity: strip trailing whitespace
per line and collapse surrounding blank lines, so whitespace-only
source edits don't invalidate translations."""
return "\n".join(line.rstrip() for line in text.split("\n")).strip("\n")
def chunk_key(text: str) -> bytes:
"""Content key of a chunk: the first 9 bytes of the blake3 digest of
the normalized text (72 bits — a site's chunk count stays far below
the birthday bound), using the same hasher as app.py's file store.
Keys are bytes: kanta/msgspec base64-encode them at the JSON
persistence level, so the raw database dicts carry 12-char strings.
"""
return blake3.blake3(_normalize(text).encode()).digest(9)
def needs_translation(chunk: str) -> bool:
"""False for chunks without prose: pure code fences, HTML blocks, and
anything that yields no translatable segments (pagerite/segments.py) —
container fences, lone {placeholders}, reference definitions.
These are inherently no-translate (docs/migrate.md): derived from the
chunk text itself, nothing is stored. Every language renders them from
the original chunk via the hybrid fallback.
"""
if _FENCE_OPEN.match(chunk):
return False
first = chunk.split("\n", 1)[0]
if any(open_re.match(first) for open_re, _ in _HTML_ATOMIC):
return False
if _HTML_TAG.match(first):
return False
return has_prose(chunk)
def join_chunks(chunks: list[str]) -> str:
"""The stored page form of chunks: blocks joined by a blank line,
with a trailing newline ("" for no chunks)."""
return "\n\n".join(chunks) + "\n" if chunks else ""
def store_chunks(store: dict[bytes, str], markdown: str) -> list[bytes]:
"""Chunk ``markdown`` into ``store`` (hash -> text); return the ordered
hashes. Unchanged chunks keep their hashes, so only genuinely new text
lands in the kanta change diff. First writer wins: variants sharing a
key differ only in insignificant whitespace (see chunk_key)."""
hashes = []
for chunk in chunk_markdown(markdown):
key = chunk_key(chunk)
store.setdefault(key, chunk)
hashes.append(key)
return hashes
+28
View File
@@ -0,0 +1,28 @@
"""CLI → app configuration, passed as JSON in the ``PAGERITE_CONFIG`` env var.
Kept dependency-free (msgspec only) so ``__main__`` can build and serialize
the config before any app module is imported, and the app side parses the
same struct back. Import-time safe: nothing here reads the environment
until ``load()`` is called.
"""
import os
import msgspec
class Config(msgspec.Struct):
"""Configuration passed from the CLI entry point to the app."""
#: Public hostname of the site; names the per-site data directory
#: ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
hostname: str = "localhost"
#: Download/update the DB-IP city lite database at startup (--dbip).
dbip: bool = False
def load() -> Config:
"""Parse ``PAGERITE_CONFIG``, or the defaults when unset."""
if raw := os.getenv("PAGERITE_CONFIG"):
return msgspec.json.decode(raw.encode(), type=Config)
return Config()
+88 -39
View File
@@ -2,8 +2,9 @@
The site structure is a tree of Nodes. Every node is a menu label with a The site structure is a tree of Nodes. Every node is a menu label with a
configurable title and slug (its key in the parent's ``children``); the configurable title and slug (its key in the parent's ``children``); the
URL path is the chain of slugs from the top level. ``content`` is the URL path is the chain of slugs from the top level. ``chunks`` is the
node's Markdown page, or None for a pure category label, whose URL renders node's Markdown page as ordered content-hash keys into ``Data.chunks``
(docs/migrate.md), or None for a pure category label, whose URL renders
a placeholder page while nav links point at its first child. a placeholder page while nav links point at its first child.
""" """
@@ -11,6 +12,16 @@ from datetime import UTC, datetime
import msgspec import msgspec
from pagerite.chunks import join_chunks
class Patch(msgspec.Struct, omit_defaults=True):
"""One editing session's overrides on a translated view, applied
independently per hunk (docs/localization.md)."""
#: (search, replace) pairs on the served hybrid Markdown.
hunks: list[tuple[str, str]] = []
class Node(msgspec.Struct, omit_defaults=True): class Node(msgspec.Struct, omit_defaults=True):
"""One item of the site hierarchy. """One item of the site hierarchy.
@@ -28,12 +39,30 @@ class Node(msgspec.Struct, omit_defaults=True):
title: str = "" title: str = ""
order: float = 0 order: float = 0
#: Markdown source of the node's page; None = pure category label #: Ordered chunk hashes (9-byte keys into ``Data.chunks``); None =
#: (its URL renders a placeholder page). #: pure category label (its URL renders a placeholder page), a list
content: str | None = None #: (possibly empty) = a page.
#: Raw HTML for the header banner (img, styled div, canvas+script...). chunks: list[bytes] | None = None
#: Primary language of the article (BCP-47 base tag). "" = inherit
#: (nearest ancestor, front page last, site default "en" final).
language: str = ""
#: Chunk hashes the editor marked "do not translate" (always served
#: from the original). Presence-keys, value always True.
no_trans: dict[bytes, bool] = {}
#: Languages this article is available in (besides its primary
#: language). Presence-keys, value always True — the availability
#: index for rendering and language selection; maintained by whoever
#: writes translation data (docs/migrate.md).
langs: dict[str, bool] = {}
#: Raw HTML for the header banner (img, styled div, canvas+script...),
#: rendered after the banner design's artwork so author code always
#: wins over the design's own styles.
#: Empty inherits the nearest ancestor's banner, front page last. #: Empty inherits the nearest ancestor's banner, front page last.
banner: str = "" banner: str = ""
#: Banner design: a theme folder name (its banner.css/banner.svg),
#: "" = explicitly no design, None = inherit (nearest ancestor, front
#: page last, then the active theme's own design).
banner_design: str | None = None
published: bool = True published: bool = True
children: dict[str, "Node"] = {} children: dict[str, "Node"] = {}
created: datetime = msgspec.field( created: datetime = msgspec.field(
@@ -44,51 +73,71 @@ class Node(msgspec.Struct, omit_defaults=True):
) )
class Page(msgspec.Struct, omit_defaults=True):
"""Legacy flat page record, from before the tree model.
Kept only so old databases still decode; app.py migrates any entries
into ``Data.menu`` on startup and clears this.
"""
title: str
markdown: str
published: bool = True
order: float = 0
banner: str = ""
created: datetime = msgspec.field(
default_factory=lambda: datetime.now(UTC),
)
modified: datetime = msgspec.field(
default_factory=lambda: datetime.now(UTC),
)
class Data(msgspec.Struct): class Data(msgspec.Struct):
"""Root object of the kanta database. Owned and edited in place by us.""" """Root object of the kanta database. Owned and edited in place by us."""
#: Top-level menu items by slug; "" is the front page. #: Top-level menu items by slug; "" is the front page.
menu: dict[str, Node] = {} menu: dict[str, Node] = {}
#: Content-addressed file store: name (blake3 hash prefix + extension)
#: -> bytes, served immutable at "/_f/{name}". Absolute URLs that stay
#: valid when pages move.
files: dict[str, bytes] = {}
#: Bumped on every structure/content write, so page ETags (which embed
#: it) invalidate cached copies when navigation-affecting changes happen.
version: int = 0
#: Site name shown in the header and <title> suffix; editable in the #: Site name shown in the header and <title> suffix; editable in the
#: site editor. Empty = no brand link in the header, no title suffix. #: site editor. Empty = no brand link in the header, no title suffix.
brand: str = "Pagerite" brand: str = "Pagerite"
#: Raw trusted HTML replacing the brand link entirely (a logo image,
#: styled markup, canvas+script...), site-wide — not per-page
#: overridable like banners. Rendered in the header on top of the
#: banner artwork, next to the nav. Empty = the plain brand link.
brand_html: str = ""
#: Active theme name (empty = none/base only). Themes live in #: Active theme name (empty = none/base only). Themes live in
#: frontend/src/assets/themes/{theme}/theme.css, with their banner #: pagerite/themes/{theme}/ (theme.css and/or banner.css/banner.svg/
#: artwork at pagerite/themes/{theme}/banner.svg (inlined server-side). #: banner.html), served by the backend from disk.
theme: str = "purple" theme: str = "corporate"
#: Page transition design name (cube, crossfade, ...). Designs live in
#: pagerite/themes/{name}/transition.css and are injected as
#: #pagerite-transition on every page.
transition: str = "cube"
#: Raw site-wide custom CSS, injected inline in every page <head>. #: Raw site-wide custom CSS, injected inline in every page <head>.
#: Trusted author content; not sanitized. #: Trusted author content; not sanitized.
custom_css: str = "" custom_css: str = ""
#: Legacy flat page store (pre-tree databases); migrated into `menu` #: Favicon: content-addressed file name (served at "/_f/{name}"),
#: on startup, then cleared. Never written otherwise. #: linked as <link rel="icon"> on every page; /favicon.ico redirects
pages: dict[str, Page] = {} #: to it. Empty = no icon (and /favicon.ico 404s).
favicon: str = ""
#: API keys gating the translator service WebSocket (/_translate/{key};
#: the external forward-auth does not cover that route): key -> display
#: name. Keys are 12 lowercase alphanumeric characters; the first is
#: generated at database bootstrap, more are managed in the editor
#: shell's lang tab (via /_api/settings).
translate_keys: dict[str, str] = {}
#: Wanted target languages for the translator service (presence-keys,
#: value always True). The dispatcher offers jobs only in the
#: intersection of these and a connection's announced capabilities.
#: Bootstrapped to es+zh; edited in the editor shell's localization
#: tab (or via /_api/settings).
translate_langs: dict[str, bool] = {}
#: All original-language page text, content-addressed:
#: chunk_key (9 bytes; base64 at the JSON level) -> Markdown chunk.
#: Shared by every article.
chunks: dict[bytes, str] = {}
#: Machine translations: chunk hash -> lang -> translated Markdown
#: (a nested dict rather than tuple keys, which msgspec's JSON
#: serializer does not support). Also used for node titles (hash of
#: the title text).
trans: dict[bytes, dict[str, str]] = {}
#: User override patches per article and language:
#: f"{path}:{lang}" -> ordered patches (paths without leading slash).
patches: dict[str, list[Patch]] = {}
def node_markdown(data: Data, node: Node) -> str | None:
"""The node's original Markdown assembled from the chunk store.
None for category labels (chunks is None); an empty page gives "".
Hashes missing from the store (shouldn't happen) are skipped.
"""
if node.chunks is None:
return None
return join_chunks(
[t for h in node.chunks if (t := data.chunks.get(h)) is not None]
)
def prettify(slug: str) -> str: def prettify(slug: str) -> str:
+391
View File
@@ -0,0 +1,391 @@
"""Content-addressed file store, image derivatives, and file routes.
``FileStore`` keeps uploads, seed assets and fetched favicons on disk under
hash-prefixed names, fully cached in RAM (uncompressed plus a zstd copy
when compression shrinks the body), served immutable at ``/_f/``. Raster
images and SVGs are recompressed into AVIF/WebP/JPEG derivatives
(``store_image`` and helpers); the untouched original is kept alongside as
``<hash>.orig<ext>`` (never served). Routes: upload/delete under
``/_api/files``, the favicon settings endpoints, the /favicon.ico
redirect to the configured icon, the ``/_f/`` server with
Accept-negotiated formats, and the user assets (``/_themes/``, ``/_fonts/``).
"""
import asyncio
import logging
import mimetypes
import tempfile
from contextlib import suppress
from pathlib import Path
import blake3
from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import RedirectResponse, Response
from mediapreview import dispatch
from pagerite import views
from pagerite.state import (
FAVICON_MAXSIZE,
FILES_DIR,
IMAGE_JPG_QUALITY,
IMAGE_MAXSIZE,
IMAGE_QUALITY,
IMAGE_WEBP_QUALITY,
_invalidate_pages,
_zstd,
data,
kanta,
)
logger = logging.getLogger(__name__)
# mediapreview logs pyvips noise ("VipsForeignSaveJpegTarget argument strip is
# deprecated", "threadpool completed with N workers") at INFO; keep warnings.
logging.getLogger("mediapreview").setLevel(logging.WARNING)
router = APIRouter()
class FileStore:
"""Content-addressed files on disk, fully cached in RAM.
Every file is kept in RAM uncompressed and zstd-compressed (the
compressed copy only when it actually shrinks the body), so ``/_f``
serves both encodings without touching disk or re-compressing.
"""
def __init__(self, path: Path) -> None:
self.path = path
#: name -> (uncompressed body, zstd body or None)
self._cache: dict[str, tuple[bytes, bytes | None]] = {}
@staticmethod
def _entry(body: bytes) -> tuple[bytes, bytes | None]:
compressed = _zstd.compress(body)
return body, compressed if len(compressed) < len(body) else None
def load(self) -> None:
"""Read every stored file into the RAM cache (startup)."""
try:
entries = sorted(self.path.iterdir())
except FileNotFoundError:
return
for f in entries:
if f.is_file() and not f.name.startswith("."):
self._cache.setdefault(f.name, self._entry(f.read_bytes()))
def get(self, name: str) -> tuple[bytes, bytes | None] | None:
return self._cache.get(name)
def put(self, name: str, body: bytes) -> None:
"""Store ``body`` under ``name`` on disk and in the RAM cache."""
if name in self._cache:
return
self.path.mkdir(parents=True, exist_ok=True)
(self.path / name).write_bytes(body)
self._cache[name] = self._entry(body)
def delete(self, name: str) -> None:
"""Delete a file plus its derivatives/original counterparts, if any.
An image upload is stored as a group sharing the hash prefix
(``<hash>.orig.<ext>`` + ``<hash>.avif/.webp/.jpg``); deleting any
of the names removes them all.
"""
stem = name.partition(".")[0]
for key in [k for k in self._cache if k.partition(".")[0] == stem]:
self._cache.pop(key, None)
with suppress(FileNotFoundError):
(self.path / key).unlink()
def __contains__(self, name: str) -> bool:
return name in self._cache
file_store = FileStore(FILES_DIR)
def _ext(orig: str) -> str:
"""Sanitized lowercase extension (with dot) of an original file name."""
return "".join(c for c in Path(orig).suffix.lower() if c.isalnum() or c == ".")
def _hash_name(body: bytes, orig: str) -> str:
"""Content-addressed file name: blake3 hash prefix + original extension."""
return blake3.blake3(body).hexdigest()[:12] + _ext(orig)
def _to_avif(body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE) -> bytes | None:
"""Recompress an image body to a thumbnailed AVIF via mediapreview's
dispatch (pyvips for common formats, ffmpeg for HEIC/HEIF/AVIF), or
None if the body is not a decodable image (stored as-is by the caller).
Dispatch needs a real file for format routing, so the body goes
through a temp file.
"""
with tempfile.NamedTemporaryFile(suffix=ext) as tmp:
tmp.write(body)
tmp.flush()
try:
avif, _resp = dispatch(
Path(tmp.name),
quality=IMAGE_QUALITY,
maxsize=maxsize,
maxzoom=1,
)
except Exception:
return None
return avif
def _svg_to_png(body: bytes, maxsize: int) -> bytes | None:
"""Rasterize an SVG to PNG via pyvips, scaled so the long side is
``maxsize`` — SVGs often carry no meaningful intrinsic resolution, so
we rasterize at full image size rather than the tiny nominal one."""
import pyvips
try:
img = pyvips.Image.new_from_buffer(body, "")
scale = (
maxsize / max(img.width, img.height)
if img.width and img.height
else maxsize
)
if scale != 1:
img = pyvips.Image.new_from_buffer(body, "", scale=scale)
return img.write_to_buffer(".png")
except pyvips.Error:
return None
def _avif_to_format(avif: bytes, suffix: str, quality: int) -> bytes:
"""Re-encode the AVIF derivative into a fallback format (WebP/JPEG)
via pyvips. JPEG has no alpha, so it is flattened onto white;
``strip`` keeps metadata (EXIF) out of the fallbacks."""
import pyvips
img = pyvips.Image.new_from_buffer(avif, "")
if suffix == ".jpg" and img.hasalpha():
img = img.flatten(background=[255, 255, 255])
return img.write_to_buffer(suffix, Q=quality, strip=True)
def _image_derivatives(
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE
) -> dict[str, bytes] | None:
"""The served variants of an uploaded image: ``avif`` (primary,
thumbnailed to ``maxsize``) plus ``webp`` and ``jpg`` fallbacks
re-encoded from it. SVGs are rasterized first (they are vector, so
the raster replaces nothing — the .svg itself stays servable).
Returns None for non-decodable content (stored as-is by the caller).
"""
if ext == ".svg":
png = _svg_to_png(body, maxsize)
if png is None:
return None
body, ext = png, ".png"
avif = _to_avif(body, ext, maxsize)
if avif is None:
return None
return {
"avif": avif,
"webp": _avif_to_format(avif, ".webp", IMAGE_WEBP_QUALITY),
"jpg": _avif_to_format(avif, ".jpg", IMAGE_JPG_QUALITY),
}
def store_image(
body: bytes, ext: str, maxsize: int = IMAGE_MAXSIZE, *, derive: bool = True
) -> str:
"""Store an image body content-addressed and return its file name.
Decodable images get AVIF/WebP/JPEG derivatives thumbnailed to
``maxsize``; the original is kept as ``<hash>.orig<ext>`` (SVG
originals as ``<hash>.svg``, still servable) and the bare ``<hash>``
name is returned (the server negotiates the format by Accept header).
Anything else — undecodable content, or ``derive=False`` (GIFs, whose
animation recompression would lose) — is stored as-is and returned with
its extension. Blocking (pyvips/ffmpeg); call via ``asyncio.to_thread``
from async code.
"""
digest = blake3.blake3(body).hexdigest()[:12]
derivatives = _image_derivatives(body, ext, maxsize) if derive else None
if derivatives is None: # store the body as-is
file_store.put(digest + ext, body)
return digest + ext
file_store.put(f"{digest}.svg" if ext == ".svg" else f"{digest}.orig{ext}", body)
for fmt, variant in derivatives.items():
file_store.put(f"{digest}.{fmt}", variant)
return digest
@router.put("/_api/files/{name}")
async def upload_file(name: str, request: Request) -> dict[str, str]:
"""Store an upload (image, video...) in the content-addressed store.
The stored name is a blake3 hash prefix + the original extension,
served immutable at "/_f/{name}"; returns {"path": "/_f/..."}.
Raster images and SVGs are recompressed (SVGs rasterized) into AVIF
(primary) plus WebP and JPEG fallbacks: the original goes to
``<hash>.orig<ext>`` (kept for reprocessing, never served — it may
carry EXIF data; SVG originals stay servable as ``<hash>.svg`` since
vector carries no EXIF) and pages link the bare ``/_f/<hash>``, the
server picking the format from the request's Accept header. GIFs are
stored as-is (animation would be lost), as is other non-decodable
content.
"""
if "/" in name or name in {".", ".."}:
raise HTTPException(400, "bad file name")
body = await request.body()
if not body:
raise HTTPException(400, "empty file")
ext = _ext(name)
stored = await asyncio.to_thread(store_image, body, ext, derive=ext != ".gif")
return {"path": f"/_f/{stored}"}
@router.delete("/_api/files/{name}", status_code=204)
async def delete_file(name: str) -> None:
"""Remove a file from the content-addressed store (no refcounting:
other pages referencing the same content will 404)."""
if name not in file_store:
raise HTTPException(404, "no such file")
file_store.delete(name)
@router.get("/favicon.ico", include_in_schema=False)
async def favicon_ico() -> Response:
"""The conventional /favicon.ico: redirect to the configured site icon.
Browsers request this path on their own (tabs, bookmarks, feeds and
other non-HTML contexts) regardless of the <link rel="icon"> pages
carry. Redirect to the icon's store URL, which negotiates the format
and caches immutably; 404 when no custom icon is configured.
"""
if not data.favicon:
raise HTTPException(404)
return RedirectResponse(f"/_f/{data.favicon}")
@router.put("/_api/settings/favicon")
async def put_favicon(request: Request) -> dict[str, str]:
"""Upload a favicon into the content-addressed store and activate it.
Raw image body (ico/png/svg...). Decodable images are thumbnailed to
FAVICON_MAXSIZE (192px — browsers scale down from there themselves)
and stored as AVIF/WebP/JPEG derivatives linked extension-less; SVG
originals also stay servable under their ``.svg`` name. Undecodable
bodies are stored as-is. Pages link it as <link rel="icon">. Returns
{"path": "/_f/..."}.
"""
body = await request.body()
if not body:
raise HTTPException(400, "empty file")
ext = _ext(request.headers.get("x-filename", "favicon.ico"))
stored = await asyncio.to_thread(store_image, body, ext, FAVICON_MAXSIZE)
with kanta.transaction("settings", user=request.headers.get("remote-user")):
data.favicon = stored
_invalidate_pages()
return {"path": f"/_f/{stored}"}
@router.delete("/_api/settings/favicon", status_code=204)
async def delete_favicon(request: Request) -> None:
"""Clear the custom favicon (/favicon.ico goes back to 404, pages drop
the <link rel="icon">).
The blob stays in the content-addressed store; only the reference goes.
"""
with kanta.transaction("settings", user=request.headers.get("remote-user")):
data.favicon = ""
_invalidate_pages()
async def _serve_user_file(path: Path | None, request: Request) -> Response:
"""Serve a user-asset file resolved on disk, with mtime etag.
Read from disk on every request (etag by mtime+size): user assets are
never built or content-hashed, so edits on disk show on the next page
load, in prod as well as dev.
"""
if path is None:
raise HTTPException(404)
stat = path.stat()
etag = f'"{stat.st_mtime_ns:x}-{stat.st_size:x}"'
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
return Response(
path.read_bytes(),
media_type=mime,
headers={"etag": etag, "cache-control": "no-cache"},
)
@router.get("/_themes/{name}/{filename}")
async def theme_file(name: str, filename: str, request: Request) -> Response:
"""Serve a theme/banner-design file, resolved across views.THEME_DIRS.
Stylesheets plus any extra assets the CSS references (like summer's
grass.svg).
"""
return await _serve_user_file(views.theme_file(name, filename), request)
@router.get("/_fonts/{name}/{filename}")
async def user_font_file(name: str, filename: str, request: Request) -> Response:
"""Serve a user font file, resolved across views.FONT_DIRS.
The folder's font.css (@font-face rules + --font-{name} stack variable)
is linked on every page; the woff2 files it references come from here.
"""
return await _serve_user_file(views.font_file(name, filename), request)
@router.get("/_f/{name}")
async def stored_file(name: str, request: Request) -> Response:
"""Serve a file from the content-addressed store (immutable: the name
is its own hash, so cache forever). Bodies are served from the RAM
cache, zstd-compressed when the client accepts it and compression
actually shrank the file.
A bare ``/_f/{hash}`` (no extension, how pages link uploaded images)
content-negotiates between the stored derivatives: a format is served
only when the Accept header lists it explicitly — ``image/avif`` →
AVIF, ``image/webp`` → WebP, anything else (including ``image/*`` and
``*/*``) → JPEG. An explicit extension pins the format. ``.orig.``
originals are internal (they may carry EXIF data) and never served."""
if ".orig." in name:
raise HTTPException(404)
etag = name
vary = ""
entry = file_store.get(name)
if entry is None and "." not in name:
# Extension-less image link: negotiate avif/webp/jpg by Accept.
vary = "accept"
accept = request.headers.get("accept", "")
if "image/avif" in accept:
order = ("avif", "webp", "jpg")
elif "image/webp" in accept:
order = ("webp", "jpg", "avif")
else:
order = ("jpg", "webp", "avif")
for ext in order:
etag = f"{name}.{ext}"
entry = file_store.get(etag)
if entry is not None:
break
if entry is None:
raise HTTPException(404)
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
body, compressed = entry
headers = {"etag": etag, "cache-control": "public, max-age=31536000, immutable"}
if compressed is not None and "zstd" in request.headers.get("accept-encoding", ""):
headers["content-encoding"] = "zstd"
vary = f"{vary}, accept-encoding".lstrip(", ")
body = compressed
if vary:
headers["vary"] = vary
mime = mimetypes.guess_type(etag)[0] or "application/octet-stream"
return Response(body, media_type=mime, headers=headers)
+282
View File
@@ -0,0 +1,282 @@
"""Localization: language selection, translation storage and assembly.
See docs/localization.md and docs/migrate.md. Each article's primary
language is ``Node.language``, inherited down the hierarchy (front page =
site default, ORIGINAL_LANGUAGE as the final fallback). The database holds
the original language as content-addressed chunks (``Data.chunks``); per
target language there are machine-translated fragments (``Data.trans``)
and user override patches (``Data.patches``), assembled into the served
Markdown at render time, with per-node fallback to the original titles.
"""
from collections.abc import Callable
from difflib import SequenceMatcher
import msgspec
from pagerite.chunks import chunk_key, chunk_markdown, join_chunks
from pagerite.data import Data, Node, Patch, resolve
#: Final fallback for a page's primary language when neither it nor any
#: ancestor (up to the front page) sets one (Node.language, "" = inherit).
ORIGINAL_LANGUAGE = "en"
#: Languages written right-to-left; pages served in one get dir="rtl" on
#: <html> (views._layout).
RTL_LANGUAGES = frozenset({"ar", "fa", "he", "ur"})
def primary_lang(menu: dict[str, Node], path: str) -> str:
"""The primary language of the article at ``path``: its own
``language`` setting, else the nearest ancestor's (the front page
last — it doubles as the site default), falling back to
ORIGINAL_LANGUAGE. Missing tail segments (a page being created)
resolve to the nearest existing ancestor."""
p = path.strip("/")
while True:
chain = resolve(menu, p)
if chain:
for node in reversed(chain):
if node.language:
return node.language
if not p:
return ORIGINAL_LANGUAGE
p = p.rpartition("/")[0]
class Translation(msgspec.Struct, omit_defaults=True):
"""Translated content for one page and language.
``markdown`` is the translated page source in the same format as the
original (None = keep the original Markdown); ``titles`` maps node paths
(top-level slug, then slash-joined) to translated navigation titles, so a
partially translated tree still renders with per-node English fallback.
"""
markdown: str | None = None
titles: dict[str, str] = {}
def base_tag(tag: str) -> str:
"""The lowercase base subtag of a language tag (fi-FI -> fi)."""
return tag.strip().lower().partition("-")[0]
def parse_accept_language(header: str) -> list[str]:
"""Accept-Language header as an ordered, deduped list of base subtags.
q-values are deliberately ignored: all known implementations send the
header in order of preference. Region tags normalize to their base
subtag (fi-FI -> fi); "*" and empties are dropped.
"""
langs = []
for part in header.split(","):
tag = base_tag(part.split(";", 1)[0])
if tag and tag != "*" and tag not in langs:
langs.append(tag)
return langs
def select_language(
query_lang: str | None,
accept_language: str | None,
is_available: Callable[[str], bool],
original: str = ORIGINAL_LANGUAGE,
) -> str:
"""The language to serve (see docs/localization.md).
1. ``?lang=`` wins when a translation exists for it (otherwise falls
through to the header logic).
2. The original language anywhere in the header list wins — an AI
translation is strictly worse than the original for anyone who has
English configured at all.
3. Otherwise the first header language with an available translation.
4. Fall back to the original.
"""
if query_lang:
tag = base_tag(query_lang)
if tag == original or (tag and is_available(tag)):
return tag
langs = parse_accept_language(accept_language or "")
if original in langs:
return original
for lang in langs:
if lang != original and is_available(lang):
return lang
return original
def apply_patch(hybrid: str, patch: Patch) -> str:
"""Apply one patch to the hybrid Markdown, best effort, each hunk
independently: a hunk whose search text no longer exists is stale and
silently skipped (docs/localization.md)."""
for search, replace in patch.hunks:
if search and search in hybrid:
hybrid = hybrid.replace(search, replace, 1)
return hybrid
def make_patch(base: str, edited: str) -> Patch:
"""The minimal diff of ``edited`` against the served ``base`` hybrid as
(search, replace) hunks at block granularity (docs/localization.md).
Blocks are the chunk_markdown split, so hunks align with translation
units and code fences never straddle a hunk boundary. Pure inserts
anchor on the preceding block (an empty search would never match);
inserts at the very top anchor on the first block. autojunk is off:
the diff must be deterministic, and pages are small.
"""
a, b = chunk_markdown(base), chunk_markdown(edited)
hunks: list[tuple[str, str]] = []
for tag, i1, i2, j1, j2 in SequenceMatcher(
None, a, b, autojunk=False
).get_opcodes():
if tag == "equal":
continue
search = "\n\n".join(a[i1:i2])
replace = "\n\n".join(b[j1:j2])
if tag == "insert":
if i1:
search = a[i1 - 1]
replace = f"{a[i1 - 1]}\n\n{replace}"
elif a:
search = a[0]
replace = f"{replace}\n\n{a[0]}"
# else: base is empty — the hunk is inert (empty search is
# skipped by apply_patch); saving a translation of an empty
# page records nothing applicable.
hunks.append((search, replace))
return Patch(hunks=hunks)
def hybrid_markdown(data: Data, node: Node, path: str, lang: str) -> str:
"""The served Markdown for ``lang``: per chunk the translation from
``Data.trans``, unless missing or marked no-translate (fallback to the
original chunk), then the language's user patches applied in order.
Not gated on ``node.langs`` (get_translation is the gated view): the
editor save path diffs against this even for a language's first patch.
"""
hybrid = join_chunks(
[
data.chunks.get(h, "")
if h in node.no_trans
else data.trans.get(h, {}).get(lang) or data.chunks.get(h, "")
for h in node.chunks or []
]
)
for patch in data.patches.get(f"{path}:{lang}", []):
hybrid = apply_patch(hybrid, patch)
return hybrid
def add_patch(
data: Data, node: Node, path: str, lang: str, edited: str, base: str | None = None
) -> bool:
"""Record a translated-view edit as a user Patch: the minimal diff of
``edited`` against ``base`` (default: the currently served hybrid),
appended to the language's patch list. Patches alone make the
translated version exist, so ``node.langs`` is set. Returns True when
a patch was stored. Pure data ops — the caller wraps in a transaction
and invalidates."""
patch = make_patch(
base if base is not None else hybrid_markdown(data, node, path, lang), edited
)
if not patch.hunks:
return False
data.patches.setdefault(f"{path}:{lang}", []).append(patch)
node.langs[lang] = True
return True
def set_title_translation(data: Data, node: Node, lang: str, title: str) -> bool:
"""Record (or drop) a per-language title override: a fragment in
``Data.trans`` keyed by the ORIGINAL title's chunk hash — the same
storage machine title translations use, overriding them. Sending the
original's text drops the override. Returns True when anything changed.
Pure data ops — the caller wraps in a transaction and invalidates."""
key = chunk_key(node.title)
current = data.trans.get(key, {}).get(lang)
if title == node.title:
if current is None:
return False
del data.trans[key][lang]
return True
if current == title:
return False
data.trans.setdefault(key, {})[lang] = title
node.langs[lang] = True
return True
def clear_translations(data: Data) -> None:
"""Drop all machine translations (``Data.trans``) and rebuild the
availability index (``node.langs``) from the surviving user patches —
patches alone make a language exist on a page. Pure data ops — the
caller wraps in a transaction and invalidates."""
data.trans.clear()
patch_langs: dict[str, set[str]] = {}
for key in data.patches:
path, _, lang = key.rpartition(":")
patch_langs.setdefault(path, set()).add(lang)
def walk(nodes: dict[str, Node], prefix: str) -> None:
for slug, node in nodes.items():
path = f"{prefix}/{slug}" if prefix else slug
node.langs = {lang: True for lang in patch_langs.get(path, ())}
walk(node.children, path)
walk(data.menu, "")
def title_map(data: Data, lang: str) -> dict[str, str]:
"""path -> translated title for every node that has one.
Titles are chunks too (docs/migrate.md): keyed by the hash of the
title text, so editing a title invalidates its translations. Nodes
without an entry fall back to their original title in views — as do
nodes whose primary language IS ``lang`` (their original title already
is in that language).
"""
titles = {}
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in nodes.items():
path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited
if node.title and node_lang != lang:
t = data.trans.get(chunk_key(node.title), {}).get(lang)
if t:
titles[path] = t
walk(node.children, path, node_lang)
walk(data.menu, "", ORIGINAL_LANGUAGE)
return titles
def subtree_languages(node: Node) -> set[str]:
"""Languages available anywhere in the node's subtree (the union of the
``langs`` indexes). Category placeholder pages select their language
from this: they have no chunks of their own, but their title,
navigation and card text localize wherever a translation exists."""
langs = set(node.langs)
for child in node.children.values():
langs |= subtree_languages(child)
return langs
def get_translation(data: Data, path: str, lang: str) -> Translation | None:
"""The translation of the page at ``path`` for ``lang``, or None.
None when the page does not exist or is not available in ``lang``:
``node.langs`` is the availability index (a stale key is benign — the
"translation" then just renders as the original).
"""
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is None or node.chunks is None or lang not in node.langs:
return None
return Translation(
markdown=hybrid_markdown(data, node, path, lang),
titles=title_map(data, lang),
)
+649 -28
View File
@@ -2,30 +2,85 @@
Raw HTML (including inline scripts) is passed through unfiltered: the Raw HTML (including inline scripts) is passed through unfiltered: the
single author is trusted. Extensions: tables and strikethrough (from the single author is trusted. Extensions: tables and strikethrough (from the
"default" preset), footnotes, definition lists, task lists and "default" preset), footnotes, definition lists, task lists,
brace-attributes (`{.class width=300}` on any element, images in brace-attributes (`{.class width=300}` on any element, images in
particular). particular), admonitions (``!!! note Title`` with an indented body —
note/tip/warning/etc., the title optional) and GitHub-style alerts
(``> [!NOTE]`` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the
same callout styling). ``::: name`` opens a generic container rendered
as ``<div class="name">`` and closed by a matching ``:::`` (nest by
giving the outer container more colons, e.g. `::::`); the name may be
followed by brace attributes (``::: aside {.right}``), or omitted for a
pandoc-style nameless div (``::: {.aside}``). ``::: aside``
floats as a muted side box, floating in the side zone at the article's
left on all but phone widths — the same margin float ``{.margin}`` (or
``::: margin``) gives any block — and ``::: nocols`` opts its section out
of the column layout. A brace-attribute
line as a block's last line (no blank line between) applies to the whole
block, e.g. a paragraph ending with ``{.wide}`` breaks out of the column
layout as a full-width element; written after a block (code fence,
heading, container, ...) it applies to that preceding block. Bare URLs autolink (GFM), with
the ``https://`` scheme hidden in the link text (``http://`` and other
schemes stay visible; manually labelled links are untouched), and
``H~2~O`` / ``x^2^`` give sub/superscripts.
render() also builds the layout structure: the top-level blocks are
segmented for the column layout — h1/h2 headings and ``.wide`` blocks
stand on their own, the runs between them are wrapped in
``<div class="colseg">`` (tagged
``.cols`` when the segment holds enough text — COLS_TEXT — in at least
COLS_PARAS paragraphs or one paragraph long enough to turn .breakable,
unless a ``::: nocols`` container opts it out;
in column segments, paragraphs past BREAKABLE_TEXT are marked
``.breakable`` so they may split across columns). Margin-breakout boxes
(``.margin``, ``::: aside``) stay inside the segment at their anchor
point; pagerite.css takes them out of flow (absolute, off the article's
left border, into the side zone), so the columns flow through as if the
box wasn't there. The result carries
``multicol`` when the whole body justifies columns (views.py puts the
class on the article); how many columns (never more than two), whether
the margin breakout applies and every other viewport adaptation is then
pagerite.css's call. The thresholds measure visible text, code blocks
excluded.
markdown-it's typographer is enabled, so body text gets SmartyPants-style
replacements: straight quotes become curly, ``--`` / ``---`` become en / em
dashes, ``...`` becomes an ellipsis, ``(c)`` becomes ©, and so on. Single
line breaks inside paragraphs become ``<br>`` (``breaks: True``). Code
spans/blocks and raw HTML are left untouched.
Images get special treatment: a relative `src` is resolved against the Images get special treatment: a relative `src` is resolved against the
page's own path (so `![alt](photo.avif)` in `/docs/design` is served from page's own path (so `![alt](photo.avif)` in `/docs/design` is served from
`/docs/design/photo.avif`), and an image with a title becomes a `/docs/design/photo.avif`), and an image standing alone in its paragraph
`<figure>` with `<figcaption>`. Positioning is done with attribute becomes a block `<figure>` with `<figcaption>` when it has a title.
Images inline with other content stay plain inline `<img>`, as does raw
`<img>` HTML written by the author. Positioning is done with attribute
classes, e.g. `![alt](photo.avif "Caption"){.right}`. classes, e.g. `![alt](photo.avif "Caption"){.right}`.
""" """
import re import re
from datetime import datetime, timedelta
from typing import NamedTuple
from markdown_it import MarkdownIt from markdown_it import MarkdownIt
from markdown_it.common.utils import escapeHtml from markdown_it.common.utils import escapeHtml
from markdown_it.renderer import RendererHTML from markdown_it.renderer import RendererHTML
from markdown_it.token import Token
from mdit_py_plugins.admon import admon_plugin
from mdit_py_plugins.attrs import attrs_plugin from mdit_py_plugins.attrs import attrs_plugin
from mdit_py_plugins.attrs.parse import ParseError, parse as parse_attrs
from mdit_py_plugins.container import container_plugin
from mdit_py_plugins.deflist import deflist_plugin from mdit_py_plugins.deflist import deflist_plugin
from mdit_py_plugins.footnote import footnote_plugin from mdit_py_plugins.footnote import footnote_plugin
from mdit_py_plugins.gfm_autolink import gfm_autolink_plugin
from mdit_py_plugins.subscript import sub_plugin
from mdit_py_plugins.superscript import superscript_plugin
from mdit_py_plugins.tasklists import tasklists_plugin from mdit_py_plugins.tasklists import tasklists_plugin
from pygments import highlight from pygments import highlight
from pygments.formatters import HtmlFormatter from pygments.formatters import HtmlFormatter
from pygments.lexers import get_lexer_by_name from pygments.lexers import get_lexer_by_name
from pygments.util import ClassNotFound from pygments.util import ClassNotFound
from slugify import slugify
# Styles in /_assets/pygments-*.css match this formatter (regenerate: # Styles in /_assets/pygments-*.css match this formatter (regenerate:
# HtmlFormatter(style="github-dark").get_style_defs("pre code")) # HtmlFormatter(style="github-dark").get_style_defs("pre code"))
@@ -48,6 +103,45 @@ def _highlight(text: str, lang: str, _attrs: str) -> str:
return highlight(text, lexer, _formatter) return highlight(text, lexer, _formatter)
def _fence_rule(
self: RendererHTML,
tokens,
idx: int,
options,
env: dict,
) -> str:
"""Render a fenced code block.
Like the default fence rule, but block attributes go on the <pre> —
the block element — instead of the <code>, which keeps only the
language class. Attributes are accepted both pandoc-style on the
info line (```{.python .wide #id key=val} — the first class is the
language when no bare language word precedes the braces) and as a
trailing `{...}` line applied by _block_attrs. This is what makes
e.g. `{.wide}` or `{style="..."}` style the block itself.
"""
token = tokens[idx]
info = token.info.strip() if token.info else ""
lang, _, brace = info.partition("{")
lang = lang.split(maxsplit=1)[0] if lang.strip() else ""
if brace:
try:
_, attrs = parse_attrs("{" + brace)
except ParseError:
attrs = {}
classes = attrs.pop("class", "").split()
if not lang and classes:
lang = classes.pop(0)
if classes:
_apply_attrs(token, {"class": " ".join(classes)})
_apply_attrs(token, attrs)
highlighted = _highlight(token.content, lang, "") or escapeHtml(token.content)
code_class = f' class="{options.langPrefix}{lang}"' if lang else ""
return (
f"<pre{self.renderAttrs(token)}><code{code_class}>{highlighted}</code></pre>\n"
)
def _image_rule( def _image_rule(
self: RendererHTML, self: RendererHTML,
tokens, tokens,
@@ -62,17 +156,36 @@ def _image_rule(
page = env.get("page_path", "") page = env.get("page_path", "")
token.attrs["src"] = f"/{page}/{src}" if page else f"/{src}" token.attrs["src"] = f"/{page}/{src}" if page else f"/{src}"
token.attrs["alt"] = self.renderInlineAsText(token.children, options, env) token.attrs["alt"] = self.renderInlineAsText(token.children, options, env)
if len(tokens) == 1:
# The only inline content of its paragraph: render as a block
# figure, captioned when titled. (The <p> wrapper is dropped by
# _unwrap_lone_figures below.) {.margin} positions the whole
# figure, so it moves from the img onto the figure wrapper — left
# on the img, the margin-breakout CSS would pull the image out of
# the figure (and mostly off-screen), leaving the caption behind.
classes = (token.attrs.get("class") or "").split()
figure_class = ""
if "margin" in classes:
classes.remove("margin")
if classes:
token.attrs["class"] = " ".join(classes)
else:
del token.attrs["class"]
figure_class = ' class="margin"'
img = self.renderToken(tokens, idx, options, env)
title = token.attrs.get("title")
caption = f"<figcaption>{escapeHtml(title)}</figcaption>" if title else ""
return f"<figure{figure_class}>{img}{caption}</figure>"
img = self.renderToken(tokens, idx, options, env) img = self.renderToken(tokens, idx, options, env)
if title := token.attrs.get("title"): # Inline with other content: a plain inline image.
return f"<figure>{img}<figcaption>{escapeHtml(title)}</figcaption></figure>"
return img return img
def _unwrap_lone_figures(state) -> None: def _unwrap_lone_figures(state) -> None:
"""Drop the <p> wrapper around a lone titled image. """Drop the <p> wrapper around a lone image.
markdown-it wraps inline content in a paragraph, but our image rule markdown-it wraps inline content in a paragraph, but our image rule
turns titled images into <figure> — a block element that is invalid turns lone images into <figure> — a block element that is invalid
inside <p>. Browsers hoist it out, leaving an empty paragraph whose inside <p>. Browsers hoist it out, leaving an empty paragraph whose
margins disturb the layout. margins disturb the layout.
""" """
@@ -80,10 +193,22 @@ def _unwrap_lone_figures(state) -> None:
for i, token in enumerate(tokens): for i, token in enumerate(tokens):
if token.type != "inline" or not token.children: if token.type != "inline" or not token.children:
continue continue
[child] = token.children if len(token.children) == 1 else [None] # Attrs consumed out of the text (e.g. {style=...} space-separated
if child and child.type == "image" and child.attrs.get("title"): # on the image's own line) leave empty text tokens behind — strip
if (tokens[i - 1].type == "paragraph_open" # them so the lone-image check is not thrown off by user styling.
and tokens[i + 1].type == "paragraph_close"): children = [c for c in token.children if c.type != "text" or c.content]
if children:
token.children = children
[child] = children if len(children) == 1 else [None]
if child and child.type == "image":
if (
tokens[i - 1].type == "paragraph_open"
and tokens[i + 1].type == "paragraph_close"
):
# A lone image becomes a <figure> (see _image_rule); block
# attrs on the paragraph (e.g. a trailing {.wide} line) move
# onto the image so they survive the unwrap.
_apply_attrs(child, tokens[i - 1].attrs or {})
tokens[i - 1].hidden = True tokens[i - 1].hidden = True
tokens[i + 1].hidden = True tokens[i + 1].hidden = True
@@ -109,29 +234,525 @@ def _tag_task_checkboxes(state) -> None:
index += 1 index += 1
md = ( def _shorten_autolinks(state) -> None:
MarkdownIt("default", {"html": True, "highlight": _highlight}) """Hide the https:// scheme in the text of bare autolinked URLs.
.use(attrs_plugin)
.use(footnote_plugin) GFM linkify sets the link text to the URL itself; only those links
.use(deflist_plugin) (markup "autolink") are shortened. http:// and other schemes stay
.use(tasklists_plugin, enabled=True) visible, and manually labelled links keep whatever label was written.
) """
md.add_render_rule("image", _image_rule) for token in state.tokens:
md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures) if token.type != "inline" or not token.children:
md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes) continue
for i, child in enumerate(token.children):
if child.type == "link_open" and child.markup == "autolink":
text = token.children[i + 1]
if text.type == "text" and text.content.startswith("https://"):
text.content = text.content.removeprefix("https://")
def render(text: str, page_path: str = "") -> str: _CONTAINER_NAME_RE = re.compile(r"[a-zA-Z][\w-]*")
"""Render Markdown text to an HTML string."""
return md.render(text, {"page_path": page_path})
def _apply_attrs(token, attrs: dict) -> None:
"""Join/set parsed brace attributes (`{.class key=value}`) on a token."""
for key, value in attrs.items():
if key == "class":
token.attrJoin("class", value)
else:
token.attrSet(key, value)
def _container_validate(params: str, _markup: str) -> bool:
"""`::: name`, optionally followed by brace attrs (`::: aside {.right}`).
Pandoc-style nameless divs (`::: {.aside}`) are accepted too — the
attrs alone give the container its classes.
"""
name, _, rest = params.strip().partition(" ")
if name.startswith("{"):
name, rest = "", params.strip()
elif not _CONTAINER_NAME_RE.fullmatch(name):
return False
rest = rest.strip()
if not rest:
return bool(name) # a nameless container needs the attrs
try:
pos, _ = parse_attrs(rest)
except ParseError:
return False
# parse() stops at (returns the index of) the closing brace.
return pos == len(rest) - 1
def _container_attrs(state) -> None:
"""Apply `::: name {attrs}` classes to container tokens at parse time.
The container plugin's default render is a plain renderToken, so the
name and brace attributes must live on the token itself — and being a
core rule (rather than a render rule) lets the segmentation in
render() see the classes (the ::: nocols opt-out, {.wide}
containers).
"""
for token in state.tokens:
if token.type != "container_block_open":
continue
info = token.info.strip()
name, _, rest = info.partition(" ")
if name.startswith("{"):
name, rest = "", info
if name:
token.attrJoin("class", name)
if rest.strip():
_, attrs = parse_attrs(rest.strip())
_apply_attrs(token, attrs)
def _block_attrs(state) -> None:
"""Apply `{.class key=value}` on a block's last line to the block.
The inline attrs plugin only covers attributes right after an image,
code span or link; this extends the same brace syntax to whole blocks.
A paragraph takes them at the end of its last line, either directly
(a trailing `{.wide}` line, no blank line between) or space-separated
at the end of the text (`some text {.small}`) — a space means the
braces belong to the block, not to an image or link before them.
A lone `{...}` paragraph applies to the previous block instead (this
is how headings take attributes, since a heading's next line always
starts a new paragraph). Runs before the typographer so quotes inside
attributes stay straight.
"""
tokens = state.tokens
for i, token in enumerate(tokens):
if token.type != "inline" or not token.children:
continue
text = token.children[-1]
if text.type != "text":
continue
m = re.search(r"(\{[^{}]*\})\s*$", text.content)
if not m:
continue
start = m.start(1)
if start and not text.content[start - 1].isspace():
continue # glued to the text — literal, or inline attrs
try:
_, attrs = parse_attrs(m.group(1))
except ParseError:
continue
standalone = len(token.children) == 1
if not standalone and start == 0 and token.children[-2].type != "softbreak":
continue
# The target: the enclosing block for a trailing attrs line, or the
# previous same-level block for a standalone attrs paragraph —
# including self-contained blocks like code fences and <hr>. Never
# a hidden token (tight-list paragraphs render no tag to hold the
# attributes) — in that case leave the text untouched instead of
# silently swallowing it.
own = i - 1 # standalone: the attrs paragraph's own opening token
j = i - 1
while j >= 0:
target = tokens[j]
if target.hidden:
pass
elif standalone:
if (
j != own
and target.level == tokens[own].level
and (
target.nesting == 1
or target.type in ("fence", "code_block", "hr")
)
):
break
elif target.nesting == 1:
break
j -= 1
if j < 0:
continue
_apply_attrs(tokens[j], attrs)
if standalone:
tokens[own].hidden = True
token.children = []
tokens[i + 1].hidden = True
elif start == 0:
del token.children[-2:]
else:
# Braces space-separated at the end of a text line: strip them
# (a whitespace-only remainder means they were on a line of
# their own after all — drop the softbreak too).
text.content = text.content[:start].rstrip()
if not text.content and token.children[-2].type == "softbreak":
del token.children[-2:]
#: Minimum number of in-body h1/h2 headings for section anchors to be
#: useful — shorter articles get no ids/self-links at all.
ANCHOR_MIN_HEADINGS = 3
def _heading_ids(state) -> None:
"""Anchor the in-body h1/h2 headings of long-enough articles.
The markdown body's own h1 and h2 headings get a slug id and their
text is wrapped in a self-link (``<a class="anchor" href="#id">``) so
section links are copyable by click or right-click — but only when the
body has at least ANCHOR_MIN_HEADINGS of them; shorter articles stay
anchor-free. The FIRST h1 is the article title: like the implicit
page-title h1 it gets no id, does not count toward the threshold, and
its self-link is ``href=""`` (back to the top of the page). An
author-set `{#id}` always wins; auto ids slugify the heading text
(python-slugify, mirroring the editor's slugify.js) and dedupe with
-2/-3 suffixes per render — unless env["anchor_ids"] presets them, as
render(anchors_from=...) does for translated pages so section URLs
stay in the original language. Headings that already contain a link are
``data-line`` records the heading's markdown source line (0-based, after
undoing the render(title=...) injection offset via ``env``) — the page
editor uses it for section pens and piecewise-linear scroll sync.
"""
tokens = state.tokens
line_offset = state.env.get("line_offset", 0)
def wrap(i: int, token, href: str) -> None:
inline = tokens[i + 1]
if not inline.children or any(c.type == "link_open" for c in inline.children):
return
anchor = Token("link_open", "a", 1)
anchor.attrs = {"href": href, "class": "anchor"}
inline.children = [anchor, *inline.children, Token("link_close", "a", -1)]
# The first in-body h1 is the title: href="" self-link, never an id.
# Only TOP-LEVEL headings participate — h1/h2 nested in ::: containers
# or asides (level > 0) get no anchors, data-lines or pens.
first_h1 = next(
(
i
for i, t in enumerate(tokens)
if t.type == "heading_open" and t.tag == "h1" and t.level == 0
),
None,
)
if first_h1 is not None:
wrap(first_h1, tokens[first_h1], "")
heads = [
(i, token)
for i, token in enumerate(tokens)
if token.type == "heading_open"
and token.tag in ("h1", "h2")
and token.level == 0
and i != first_h1
]
if len(heads) < ANCHOR_MIN_HEADINGS:
return
seen: set[str] = set()
preset = state.env.get("anchor_ids")
for k, (i, token) in enumerate(heads):
inline = tokens[i + 1]
hid = token.attrGet("id")
if not isinstance(hid, str) or not hid:
if preset is not None and k < len(preset):
# Translated render: the original language's slug, matched
# by heading position (a translation never adds, removes or
# reorders headings; a patched one that does falls back to
# slugging its own text past the end of the list).
base = preset[k]
else:
# Slug the visible text, not the raw markdown (`## [a](url)`).
text = "".join(
c.content
for c in inline.children
if c.type in ("text", "code_inline")
)
base = slugify(text) or "section"
hid, n = base, 2
while hid in seen:
hid = f"{base}-{n}"
n += 1
token.attrSet("id", hid)
seen.add(hid)
if token.map:
token.attrSet("data-line", str(max(0, token.map[0] - line_offset)))
wrap(i, token, f"#{hid}")
def anchor_ids(text: str, title: str | None = None) -> list[str]:
"""The section anchor ids of text, in heading order.
render(anchors_from=...) feeds these to _heading_ids via
env["anchor_ids"], pinning a translated render's anchors to the
original language's slugs. The selection mirrors _heading_ids exactly
(the same md instance assigns the ids during this parse, author-set
{#id} included as-is); the in-body title h1 is excluded.
"""
if title and not has_h1(text):
text = f"# {title}\n\n{text}"
tokens = md.parse(text, {"page_path": ""})
first_h1 = next(
(
i
for i, t in enumerate(tokens)
if t.type == "heading_open" and t.tag == "h1" and t.level == 0
),
None,
)
return [
t.attrGet("id")
for i, t in enumerate(tokens)
if t.type == "heading_open"
and t.tag in ("h1", "h2")
and t.level == 0
and i != first_h1
]
def make_md(*, verbatim: bool = False) -> MarkdownIt:
"""A fully configured parser. The module-level ``md`` (below) is the
render instance; ``verbatim=True`` builds the segmentation instance for
segments.py, where token text must stay byte-identical to the source so
prose spans can be spliced back by offset: no typographer (quotes and
dashes stay straight), no tasklist label wrapping (the item text stays
a plain text token), and soft line breaks (wrapped prose merges into
one segment instead of splitting at hardbreaks)."""
parser = (
MarkdownIt(
"default",
{
"html": True,
"highlight": _highlight,
"typographer": not verbatim,
"breaks": not verbatim,
},
)
.use(attrs_plugin)
.use(admon_plugin)
.use(container_plugin, "block", validate=_container_validate)
.use(footnote_plugin)
.use(deflist_plugin)
# label wrapping (render) puts the item text inside the checkbox
# <label> html_inline; without it the text stays a plain token.
.use(
tasklists_plugin, enabled=True, label=not verbatim, label_after=not verbatim
)
.use(gfm_autolink_plugin)
.use(sub_plugin)
.use(superscript_plugin)
)
parser.add_render_rule("image", _image_rule)
parser.add_render_rule("fence", _fence_rule)
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
parser.options["alerts"] = True
# Block attrs must be stripped before the typographer curlifies their quotes.
parser.core.ruler.before("replacements", "block_attrs", _block_attrs)
parser.core.ruler.push("container_attrs", _container_attrs)
parser.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
parser.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
parser.core.ruler.push("shorten_autolinks", _shorten_autolinks)
parser.core.ruler.push("heading_ids", _heading_ids)
return parser
md = make_md()
# Text-length thresholds (visible characters, code blocks excluded) for the
# column layout: the article goes .multicol past MULTICOL_TEXT, and a column
# segment gets .cols past COLS_TEXT — provided it also has at least
# COLS_PARAS paragraphs or a paragraph long enough to turn .breakable: a
# lone unbreakable paragraph would fill a column on its own and strand the
# rest (e.g. a floated figure) in the other, leaving a mostly empty column.
MULTICOL_TEXT = 1800
COLS_TEXT = 600
COLS_PARAS = 2
#: Paragraphs past this visible length are marked .breakable, letting them
#: split across columns (shorter ones stay unbreakable so a paragraph never
#: straddles the column gap).
BREAKABLE_TEXT = 800
_PRE_BLOCK_RE = re.compile(r"<pre\b.*?</pre>", re.S)
_TAG_RE = re.compile(r"<[^>]+>")
_PARA_OPEN_RE = re.compile(r"<p[\s>]")
_PARA_RE = re.compile(r"<p((?:\s[^>]*)?)>(.*?)</p>", re.S)
# Classes that take their block out of the column flow: .wide is a
# full-width separator that splits the column segments. Margin-breakout
# boxes (.margin/.aside) are NOT boundaries: they stay inside the segment
# at their anchor point, and CSS positions them absolutely out of the
# article's left border (the zone rules anchor off the article), so the
# column flow is unaffected.
_WIDE = "wide"
class Rendered(NamedTuple):
"""render() result: the segmented body HTML, and whether the article
should carry .multicol (enough visible text to justify columns)."""
html: str
multicol: bool
def _classes(token) -> set[str]:
return set((token.attrGet("class") or "").split())
def _text_len(html: str) -> int:
"""Visible-text length of rendered HTML, code blocks excluded."""
return len(_TAG_RE.sub("", _PRE_BLOCK_RE.sub("", html)).strip())
def _breakable_paras(html: str) -> str:
"""Mark column-filling paragraphs .breakable so they may split.
Columns keep paragraphs whole (break-inside: avoid-column), but a
paragraph long enough to fill a column would strand everything after
it in a column of its own — these get .breakable, and pagerite.css
lets them split across the column gap. Only applied to .cols segments.
"""
def repl(m: re.Match[str]) -> str:
attrs, body = m.group(1), m.group(2)
if _text_len(body) <= BREAKABLE_TEXT:
return m.group(0)
if 'class="' in attrs:
attrs = attrs.replace('class="', 'class="breakable ', 1)
else:
attrs = f'{attrs} class="breakable"'
return f"<p{attrs}>{body}</p>"
return _PARA_RE.sub(repl, html)
def _top_level_blocks(tokens: list) -> list[list]:
"""Split the token stream into its top-level blocks.
A new block starts at each level-0 opening/self-contained token;
closing and nested tokens (inline children, sub-containers) belong to
the current block, so every slice is balanced and renders on its own.
"""
blocks = []
for token in tokens:
if token.level == 0 and token.nesting >= 0:
blocks.append([token])
elif blocks:
blocks[-1].append(token)
return blocks
def _is_boundary(block: list) -> bool:
"""True for blocks that never go inside a column segment (see the
_WIDE comment above): h1/h2 headings and anything carrying .wide."""
first = block[0]
if first.type == "heading_open" and first.tag in ("h1", "h2"):
return True
for token in block:
if _WIDE in _classes(token):
return True
if token.type == "inline":
children = token.children or []
if any(_WIDE in _classes(c) for c in children):
return True
return False
def render(
text: str,
page_path: str = "",
created: datetime | None = None,
modified: datetime | None = None,
title: str | None = None,
anchors_from: tuple[str, str] | None = None,
) -> Rendered:
"""Render Markdown text to the article body's HTML and layout flags.
``title`` injects a ``# {title}`` line at the top when the markdown has
no h1 of its own, so the implicit page title goes through the exact
same pipeline as an explicit one (first-h1 anchor treatment included).
``anchors_from`` is the (markdown, title) of the ORIGINAL language when
rendering a translation: section anchors are pinned to its slugs so
localized pages keep the original #hash URLs.
The top-level blocks are grouped into column segments: boundary blocks
(h1/h2 headings, .wide — see _is_boundary) are rendered bare, the runs
between them wrapped in <div class="colseg">.
A segment is tagged .cols when it holds enough text (COLS_TEXT) in at
least two paragraphs (COLS_PARAS) or one breakable-length paragraph,
and no ::: nocols container; its long paragraphs are marked .breakable;
the article is .multicol when the whole body exceeds MULTICOL_TEXT.
pagerite.css keys all column and margin-breakout layout off these
classes.
A ``{dates}`` line expands to the article's published/updated dateline
(needs ``created``/``modified``; left as-is in contexts without them,
e.g. the editor preview). Position is the author's choice — typically
right after the article's h1.
"""
env = {"page_path": page_path, "line_offset": 0}
if anchors_from is not None:
env["anchor_ids"] = anchor_ids(*anchors_from)
if title and not has_h1(text):
text = f"# {title}\n\n{text}"
# The injected title shifts source lines by two; _heading_ids
# subtracts this from its data-line attributes.
env["line_offset"] = 2
blocks = _top_level_blocks(md.parse(text, env))
# Group consecutive non-boundary blocks into segments (is_segment,
# flat tokens); boundary blocks stand on their own between them.
groups: list[tuple[bool, list]] = []
for block in blocks:
if _is_boundary(block):
groups.append((False, block))
elif groups and groups[-1][0]:
groups[-1][1].extend(block)
else:
groups.append((True, list(block)))
parts = []
total = 0
for is_segment, group in groups:
html = md.renderer.render(group, md.options, env)
if not html.strip():
continue # e.g. a consumed standalone-attrs paragraph
text_len = _text_len(html)
total += text_len
if not is_segment:
parts.append(html)
continue
nocols = any(
"nocols" in _classes(t) for t in group if t.type == "container_block_open"
)
marked = _breakable_paras(html)
cols = (
" cols"
if text_len > COLS_TEXT
and not nocols
and (len(_PARA_OPEN_RE.findall(html)) >= COLS_PARAS or marked != html)
else ""
)
if cols:
html = marked
parts.append(f'<div class="colseg{cols}">{html}</div>')
html = "".join(parts)
if created is not None and "<p>{dates}</p>" in html:
html = html.replace("<p>{dates}</p>", _dateline(created, modified))
return Rendered(html, total > MULTICOL_TEXT)
def _dateline(created: datetime, modified: datetime | None) -> str:
"""Dateline for the ``{dates}`` tag: "1 Jan 2026", plus
" edited 3 Jan 2026" when the last edit came >= 48h after
publishing (quick fixes right after posting stay unmentioned)."""
out = f'<time datetime="{created.isoformat()}">{created.day} {created:%b %Y}</time>'
if modified is not None and modified - created >= timedelta(hours=48):
out += f' edited <time datetime="{modified.isoformat()}">{modified.day} {modified:%b %Y}</time>'
return f'<p class="dateline">{out}</p>'
def has_h1(text: str) -> bool: def has_h1(text: str) -> bool:
"""True if the Markdown source itself contains an h1 heading. """True if the Markdown source itself contains an h1 heading.
When it does, the article owns its heading and the page title is not When it does, the article owns its heading and render(title=...) does
rendered as an additional h1 (the title is still used for the document not inject the page title as an h1 (the title is still used for the
<title> and navigation labels). document <title> and navigation labels).
""" """
return any(t.type == "heading_open" and t.tag == "h1" for t in md.parse(text)) return any(t.type == "heading_open" and t.tag == "h1" for t in md.parse(text))
+178
View File
@@ -0,0 +1,178 @@
"""Kanta schema migrations, discovered by name (``migrate_vN``).
Each function receives the raw state dict (JSON-level: bytes are base64
strings, datetimes RFC 3339 strings, struct fields with default values
omitted) before it is decoded into ``Data`` structs, and runs exactly once
per database based on its recorded version.
All storage/schema upgrades live here — including on-disk file work, which
runs through files.py's file store (imported lazily: files.py owns the store
and state.py passes this module to Kanta; at migration time, during lifespan
``kanta.open()``, both modules are fully loaded).
"""
import base64
import re
from pathlib import Path
from pagerite.chunks import chunk_key, chunk_markdown
from pagerite.data import prettify
def _append_order(nodes: dict) -> float:
"""Raw-dict equivalent of data.append_order (order keys may be absent)."""
return max((n.get("order", 0) for n in nodes.values()), default=0) + 1
def _ensure(menu: dict, path: str) -> dict:
"""Raw-dict equivalent of state._ensure: the node dict at ``path``,
creating it and any missing ancestors (content-less category labels)
appended at the end of their level."""
nodes = menu
node = None
for seg in path.split("/"):
node = nodes.get(seg)
if node is None:
node = {"title": prettify(seg), "order": _append_order(nodes)}
nodes[seg] = node
nodes = node.setdefault("children", {})
return node
def migrate_v1(d: dict) -> None:
"""Move in-database file blobs to the on-disk content-addressed store,
and rebuild the legacy flat page store (``pages``) as the menu tree."""
files = d.pop("files", None)
if files:
from pagerite.files import file_store
for name, body in files.items():
if isinstance(body, str): # JSON-level bytes are base64 strings
body = base64.b64decode(body)
file_store.put(name, body)
pages = d.pop("pages", None)
if not pages:
return
menu = d.setdefault("menu", {})
for path, page in pages.items():
node = _ensure(menu, path)
node["title"] = page["title"]
node["content"] = page["markdown"]
for key in ("banner", "published", "order", "created", "modified"):
if key in page:
node[key] = page[key]
#: Extension-less file links: uploaded images are linked as /_f/<hash>
#: and the server negotiates avif/webp/jpg from the Accept header.
_DERIVATIVE_LINK = re.compile(r"(/_f/[0-9a-f]{12})\.(?:avif|webp)\b")
def _backfill_derivatives() -> None:
"""Create missing AVIF/WebP/JPEG derivatives for files stored before
they were introduced (older uploads may have only the original plus
AVIF, and SVGs no raster variants at all). WebP/JPEG are re-encoded
from an existing AVIF when available, everything else from the
original (SVGs rasterized first)."""
from pagerite.files import (
IMAGE_MAXSIZE,
IMAGE_WEBP_QUALITY,
IMAGE_JPG_QUALITY,
_avif_to_format,
_svg_to_png,
_to_avif,
file_store,
)
try:
paths = [f for f in file_store.path.iterdir() if f.is_file()]
except FileNotFoundError:
return
groups: dict[str, list[Path]] = {}
for p in paths:
groups.setdefault(p.name.partition(".")[0], []).append(p)
for digest, files in groups.items():
names = {p.name for p in files}
source = next(
(p for p in files if ".orig." in p.name or p.suffix == ".svg"), None
)
if source is None:
continue # plain as-is file, no derivatives to make
avif = file_store.get(f"{digest}.avif")
if avif is None:
ext = source.suffix
body = source.read_bytes()
if ext == ".svg":
png = _svg_to_png(body, IMAGE_MAXSIZE)
if png is None:
continue
body, ext = png, ".png"
converted = _to_avif(body, ext)
if converted is None:
continue
file_store.put(f"{digest}.avif", converted)
avif = file_store.get(f"{digest}.avif")
for fmt, quality in (
("webp", IMAGE_WEBP_QUALITY),
("jpg", IMAGE_JPG_QUALITY),
):
if f"{digest}.{fmt}" not in names:
file_store.put(
f"{digest}.{fmt}", _avif_to_format(avif[0], f".{fmt}", quality)
)
def migrate_v2(d: dict) -> None:
"""Extension-less image links: strip .avif/.webp extensions from /_f/
links in page content and banners (the server now negotiates the format
by Accept header), backfill missing AVIF/WebP/JPEG derivatives on disk,
and drop the obsolete render-counter field ``version`` (invalidation is
an in-memory concern now, not database state)."""
def walk(nodes: dict) -> None:
for node in nodes.values():
for field in ("content", "banner"):
if isinstance(node.get(field), str):
node[field] = _DERIVATIVE_LINK.sub(r"\1", node[field])
walk(node.get("children") or {})
walk(d.get("menu") or {})
d.pop("version", None)
_backfill_derivatives()
def migrate_v3(d: dict) -> None:
"""Content-addressed chunk storage (docs/migrate.md): split every
node's string ``content`` into block chunks stored once per content
hash in the new ``chunks`` store; the node keeps the ordered hash
list as ``chunks`` (an absent content stays absent, i.e. None = a
pure category label; "" chunks to an empty list = an empty page).
Chunk keys are 9-byte blake3 digests; at this raw JSON level they are
base64 strings (decoding into the structs restores ``bytes`` keys).
``trans``/``patches`` start empty; the translator job fills them and
maintains the ``langs`` index as translations land. ``language``,
``no_trans`` and ``langs`` need nothing — struct defaults cover them.
"""
store = d.setdefault("chunks", {})
d.setdefault("trans", {})
patches = d.setdefault("patches", {})
def walk(nodes: dict) -> None:
for node in nodes.values():
content = node.pop("content", None)
if isinstance(content, str):
hashes = []
for chunk in chunk_markdown(content):
key = base64.b64encode(chunk_key(chunk)).decode()
store.setdefault(key, chunk)
hashes.append(key)
node["chunks"] = hashes
walk(node.get("children") or {})
walk(d.get("menu") or {})
# Article paths never carry a leading slash in keys (docs/migrate.md).
# The only path-keyed store starts empty here, so this is defensive
# for databases that went through a downgrade/upgrade cycle.
for key in [k for k in patches if k.startswith("/")]:
patches[key.lstrip("/")] = patches.pop(key)
+229
View File
@@ -0,0 +1,229 @@
"""Public content pages: front page, sitemap, robots, and the catch-all.
``GET /{path:path}`` resolves a slug path against the menu tree and renders
the page (or a category placeholder, or 404); it must be registered AFTER
the fastapi-vue asset routes so built frontend files win over content slugs
(see app.py). Every served document is recorded raw in analytics (one
access-log line with its true HTTP status; classification happens at
display time — see pagerite/analytics.py).
"""
import logging
from datetime import UTC, datetime
from email.utils import format_datetime
from xml.sax.saxutils import escape as xml_escape
from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import RedirectResponse, Response
from pagerite import i18n, state
from pagerite.data import Node, resolve, sorted_nodes
from pagerite.state import (
SITE_URL,
_html_response,
_is_reserved,
data,
)
from pagerite.tracking import _record_get
logger = logging.getLogger(__name__)
router = APIRouter()
def _http_date(dt: datetime) -> str:
"""RFC 7231 date for the Last-Modified header."""
return format_datetime(dt.astimezone(UTC), usegmt=True)
def _is_trackable_path(path: str) -> bool:
"""Content URLs only: skip auth endpoints and reserved/machinery paths."""
if not path:
return True
if path == "auth" or path.startswith("auth/"):
return False
return not _is_reserved(path)
@router.get("/")
async def front_page(request: Request) -> Response:
"""Render the front page (slug path "")."""
return await show_page(request, "")
@router.get("/sitemap.xml")
async def sitemap(request: Request) -> Response:
"""Dynamically generate a sitemap of all published article pages."""
base = SITE_URL or str(request.base_url).rstrip("/")
entries: list[tuple[str, datetime, int]] = []
def walk(
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
) -> None:
first_content_slug = next(
(
slug
for slug, node in sorted_nodes(nodes)
if node.published and node.chunks is not None
),
None,
)
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
depth = path.count("/") if path else 0
if (
not parent_has_content
and slug == first_content_slug
and node.published
and node.chunks is not None
and depth > 0
):
depth -= 1
if node.published and node.chunks is not None:
entries.append((path, node.modified, depth))
if node.children:
walk(node.children, path, node.chunks is not None)
walk(data.menu, "")
def priority(depth: int) -> float:
return max(0.1, 1.0 - depth * 0.2)
lines = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
]
for path, modified, depth in entries:
loc = xml_escape(f"{base}/{path}" if path else base)
lastmod = (
modified.astimezone(UTC)
.replace(microsecond=0)
.isoformat()
.replace("+00:00", "Z")
)
lines.append(
f" <url>"
f"<loc>{loc}</loc>"
f"<lastmod>{lastmod}</lastmod>"
f"<priority>{priority(depth):.1f}</priority>"
f"</url>"
)
lines.append("</urlset>")
return Response(
"\n".join(lines),
media_type="application/xml",
headers={"cache-control": "no-cache"},
)
@router.get("/robots.txt")
async def robots_txt(request: Request) -> Response:
"""Allow content crawling, keep the SSO login (/auth/) and the
admin-gated API (/_api) out of search results, and point crawlers at
the sitemap."""
base = SITE_URL or str(request.base_url).rstrip("/")
body = f"User-agent: *\nAllow: /\nDisallow: /auth/\nDisallow: /_api\nSitemap: {base}/sitemap.xml\n"
return Response(
body,
media_type="text/plain",
headers={"cache-control": "no-cache"},
)
@router.get("/{path:path}", response_model=None)
async def show_page(request: Request, path: str) -> Response:
"""Render the content page at a slug path, or 404.
A node without content is a category label: its URL renders a
placeholder page (nav links point straight at its first child).
"""
path = path.strip("/")
accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page.
# Recorded like any other GET: telltale scanner paths (dotpaths
# like /.env, *.php) classify the IP as abuse at display time.
_record_get(request, status=404)
raise HTTPException(404)
chain = resolve(data.menu, path)
node = chain[-1] if chain else None
if node is not None and node.published and node.chunks is not None:
# Language selection (docs/localization.md): ?lang= wins when a
# translation exists, else header logic. Analytics keep the raw
# Accept-Language header regardless of the selection.
query_lang = request.query_params.get("lang")
lang = i18n.select_language(
query_lang,
accept_language,
lambda tag: tag in node.langs,
original=i18n.primary_lang(data.menu, path),
)
# A ?lang= override is replicated onto the page's navigation links
# (link_lang), so clicks and prefetches stay in the chosen language.
# Query and header-selected renders of the same language differ in
# their links, so link_lang is part of the ETag and body cache key.
link_lang = i18n.base_tag(query_lang or "")
# no-cache forbids serving a stored page without revalidation
# (browsers would otherwise cache heuristically and serve stale
# pages, e.g. after a theme change). In-session speed instead comes
# from pagerite.js's in-memory page cache (preload everything, never
# fetch on navigation); the ETag just makes those one-time preload
# fetches and any revalidation cheap.
etag = f'"{path}@{node.modified.timestamp()}g{state._render_gen}l{lang}q{link_lang}"'
if request.headers.get("if-none-match") == etag:
return Response(status_code=304)
if _is_trackable_path(path):
_record_get(request)
return _html_response(
request,
"page",
path,
headers={
"etag": etag,
"last-modified": _http_date(node.modified),
"cache-control": "no-cache",
},
lang=lang,
link_lang=link_lang,
)
if node is not None and node.published and node.chunks is None:
# Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real).
# Language selection as on content pages, but over the whole
# subtree's availability: the category has no chunks of its own —
# its heading, the navigation and the cards' text localize from
# the title map and the target articles' translations.
query_lang = request.query_params.get("lang")
subtree_langs = i18n.subtree_languages(node)
lang = i18n.select_language(
query_lang,
accept_language,
lambda tag: tag in subtree_langs,
original=i18n.primary_lang(data.menu, path),
)
link_lang = i18n.base_tag(query_lang or "")
if _is_trackable_path(path):
_record_get(request, status=404)
return _html_response(
request,
"category",
path,
404,
headers={
"last-modified": _http_date(node.modified),
"cache-control": "no-cache",
},
lang=lang,
link_lang=link_lang,
)
if node is None and not path:
# No front page (no top-level node with slug ""): "/" opens the
# first item of the navigation instead.
for slug, item in sorted_nodes(data.menu):
if item.published:
return RedirectResponse(f"/{slug}")
if _is_trackable_path(path):
_record_get(request, status=404)
return _html_response(request, "not-found", path, 404)
Binary file not shown.

After

Width:  |  Height:  |  Size: 410 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 59 KiB

+412 -332
View File
@@ -1,385 +1,437 @@
"""Seed content written to the database on first run (when it is empty). """Seed content written to the database on first creation only
(``@kanta.bootstrap`` in ``app.py``).
Demonstrates the formatting options: images attached to pages and served A "welcome to your new site" starter: a structured docs section (three
from the page path, figures with captions, attribute classes for menu levels deep) covering editing and one long article that walks the
positioning, footnotes, definition lists, task lists, tables and raw HTML. full Markdown feature set — each feature shown as its Markdown source in
a code block followed by the rendered result — and a showcase section
with image positioning, long-form layout and banner designs.
Binary seed images live in ``seed-assets/`` (public domain, from
Wikimedia Commons: the two whale engravings are Augustus Burnham
Shute's illustrations for an 1892 edition of Moby-Dick; the wave is
Hokusai's "The Great Wave off Kanagawa").
""" """
from pathlib import Path
ASSETS = Path(__file__).with_name("seed-assets")
def _asset(name: str) -> bytes:
return (ASSETS / name).read_bytes()
WELCOME = """\ WELCOME = """\
Welcome to your new **Pagerite** site. Pages are written in Markdown — Welcome to your new **Pagerite** site. Everything you see is a page written in Markdown, served from a pretty URL, and editable right here in the browser.
including raw HTML — and served from pretty URLs.
Have a look around: Where to go next:
- The [docs](/docs) section explains [how to write content](/docs/editing), - The [docs](/docs/editing) section explains how to edit this site and walks through every supported Markdown feature, source and result side by side.
including images and positioning. - The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read.
- [The Long Read](/blog/the-long-read) demonstrates a longer article with - Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree.
scroll effects.
- The [about](/about) page shows off assorted formatting.
![Abstract waves](waves.svg "Generated SVG artwork, attached to this page") ![Abstract waves](waves.svg "Generated SVG artwork, attached to this page"){width=420}
*Delete or rewrite any of these pages — they are only here to get you started.*
""" """
ABOUT = """\ EDITING = """\
This site runs on **Pagerite**: FastAPI + html5tagger + kanta, with content Everything on the site is editable in place. Log in, and pens appear: 🖊️ on the page and banner, ⚙️ in the banner corner for site settings.
written in Markdown.
Some formatting samples: ## The editor
- [x] Write content in Markdown The 🖊️ pens open a tabbed editor over the page you are viewing:
- [x] Attach images to pages
- [ ] Add editing UI
Term - **Article** — the page's title and Markdown, with a live preview. The format bar inserts the harder-to-remember syntax (links, tables, images); Ctrl/Cmd-B, I and S do what you expect. Saving is explicit: 💾 or Ctrl+S.
: A definition list entry, rendered by the deflist plugin. - **Banner** — per-page banner HTML and a banner design picker. Banners are raw HTML (an image, a styled div, a canvas with a script) and subpages inherit the nearest banner up their path.
- **Site** — brand, theme, fonts, favicon and custom CSS, all applied immediately.
- **Structure** — the page tree. Drag rows to reorder or nest, rename titles and slugs inline, adds a page, ✕ deletes one.
And a table: ## URLs and structure
The URL is the structure: a page at `docs/markdown` lives under `docs`, and the menus are derived from that. Slugs are lowercase ASCII (`a-z 0-9 - _`). A node without content is a category label — it renders a placeholder and its menu link points at its first child page. This site's own `docs` label demonstrates that, and the sidebar on this page shows the two submenu levels below it.
Images and files uploaded anywhere land in a content-addressed store served from `/_f/{hash}`, so links survive page moves. The server picks AVIF, WebP or JPEG from your browser's Accept header. The article editor's format bar and copy-paste both upload images for you.
{dates}
"""
# The full feature walkthrough: every supported extension in one long
# article, each shown as Markdown source followed by the rendered result.
MD_ARTICLE = """\
# Markdown
Everything Pagerite's renderer supports, on one long page — each feature shown first as Markdown source, then rendered. This page is also the live demo of the reading layout: on a wide screen the text flows in columns, and side boxes lean into the margin.
## Text and headings
```markdown
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a
[link to the front page](/). A hard line break
is just a newline.
Straight quotes become "curly", dashes -- and --- come out
properly, and ... becomes an ellipsis, all automatically.
```
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). A hard line break
is just a newline.
Straight quotes become "curly", dashes -- and --- come out properly, and ... becomes an ellipsis, all automatically.
Headings from `##` down organize the article. On pages with at least three of them, each h1/h2 gets an anchor id and a self-link, so sections are linkable (try hovering a heading here) — and the editor's section pens and scroll sync key off the same anchors.
## Lists
```markdown
- One
- Two
- Nested
1. First
2. Second
- [x] Task lists with real checkboxes
- [x] Clickable on the rendered page
- [ ] Like this one
```
- One
- Two
- Nested
1. First
2. Second
- [x] Task lists with real checkboxes
- [x] Clickable on the rendered page
- [ ] Like this one
## Quotes and alerts
```markdown
> A blockquote. Newlines inside it are kept,
> and a blank `>` line starts a new paragraph.
> [!NOTE]
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
> render as callout boxes.
```
> A blockquote. Newlines inside it are kept,
> and a blank `>` line starts a new paragraph.
> [!NOTE]
> GitHub-style alerts — `NOTE`, `TIP`, `IMPORTANT`, `WARNING`, `CAUTION` —
> render as callout boxes.
## Code
Fenced blocks get server-side syntax highlighting, and a copy button on hover:
````markdown
```python
def greet(name: str) -> str:
return f"Hello, {name}!"
```
````
```python
def greet(name: str) -> str:
return f"Hello, {name}!"
```
## Tables
```markdown
| Feature | Status |
|---------|--------|
| Pages | done |
| Images | done |
```
| Feature | Status | | Feature | Status |
|---------|--------| |---------|--------|
| Pages | done | | Pages | done |
| Images | done | | Images | done |
| Comments| later |
Footnotes work too.[^1] ## Footnotes
```markdown
Footnotes work inline.[^1]
[^1]: Rendered at the bottom of the page, with a back-reference. [^1]: Rendered at the bottom of the page, with a back-reference.
```
Footnotes work inline.[^1]
[^1]: Rendered at the bottom of the page, with a back-reference.
## Definition lists
```markdown
Term
: A definition list entry.
Another term
: With its definition.
```
Term
: A definition list entry.
Another term
: With its definition.
## Sub- and superscript
```markdown
H~2~O and x^2^ + y^2^ = z^2^.
```
H~2~O and x^2^ + y^2^ = z^2^.
## Admonitions
```markdown
!!! note
An admonition block for notes, warnings, tips...
!!! warning "Mind the whale"
With an optional custom title.
```
!!! note
An admonition block for notes, warnings, tips...
!!! warning "Mind the whale"
With an optional custom title.
## Containers and margin notes
`::: name` wraps its contents in a `<div class="name">` — brace attributes allowed. Three names are built in: `aside` floats a muted side box, `margin` marks a block as a margin note, and `nocols` opts its section out of the column layout. The `{.margin}` attribute does the same for a single block, written on its last line:
````markdown
::: aside
A side box. On all but phone widths it floats in the side zone at
the article's left, and the text never moves.
:::
This paragraph is a margin note.
{.margin}
::: nocols
This section never flows into columns, however long the article.
:::
````
::: aside
A side box. On all but phone widths it floats in the side zone at the article's left, and the text never moves.
:::
This paragraph is a margin note.
{.margin}
::: nocols
This section never flows into columns, however long the article.
:::
## Raw HTML
HTML passes through untouched — useful for `<kbd>` keys, `<details>` sections, embedded media:
```html
<details><summary>Click to expand</summary>Hidden content.</details>
```
<details><summary>Click to expand</summary>Hidden content.</details>
## Datelines
A `{dates}` line on its own expands to the article's published/updated dateline:
```markdown
{dates}
```
{dates}
## Images and layout
An image standing alone in its paragraph becomes a `<figure>`; its title becomes the caption; brace attributes control placement — `{.right}`, `{.left}`, `{.margin}`, `{.wide}`, or plain ones like `width=280`. That deserves its own page: [Images and Layout](/docs/markdown/images-and-layout).
## The page title
If your Markdown contains its own `# heading`, the page title is not repeated as a second h1 — it still supplies the `<title>` and the menu labels. This page is an example: its `# Markdown` heading *is* the title.
""" """
EDITING = """\ MD_LAYOUT = """\
Pages are written in Markdown with extensions. Everything below is plain ## Images and figures
Markdown source — no special support from the article is needed for the
site's layout or scroll effects.
## Images An image standing alone in its paragraph becomes a `<figure>`; its title becomes the caption. Inline images within text stay plain.
Upload a file (`PUT /_api/files/{filename}`) and it lands in the ```markdown
content-addressed store, served immutable from `/_f/{hash}.ext` — an ![Abstract shapes](shapes.svg "A captioned figure")
absolute URL that survives page moves:
```
![Abstract shapes](/_f/....svg "A captioned figure"){.right width=280}
``` ```
![Abstract shapes](shapes.svg "A captioned figure, floated right with an attribute class"){.right width=280} ![Abstract shapes](shapes.svg "A captioned figure")
The title becomes a `<figcaption>`, and brace attributes (the attrs ## Positioning with attributes
plugin) control positioning: `{.right}`, `{.left}`, `{.wide}`, plus plain
attributes like `width=280`. Absolute and external URLs pass through
unchanged.
## Text Brace attributes (the attrs plugin) control placement: `{.right}` and `{.left}` float, `{.margin}` moves a figure into the side zone, `{.wide}` breaks out of the text column, and plain attributes like `width=280` pass through.
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and ```markdown
[links](/about) as usual. Blockquotes: ![Abstract shapes](shapes.svg "Floated right"){.right width=280}
> The URL space is the author's. Pretty slugs at the root, nesting only
> where the content is genuinely structured.
## Code
```python
def render(text: str, page_path: str) -> str:
return md.render(text, {"page_path": page_path})
``` ```
![Abstract shapes](shapes.svg "Floated right — text wraps around it"){.right width=280}
Floated images let the text wrap around them, like this paragraph does. Relative image paths resolve against the page's own path, so attached files travel with the page. Uploaded files get content-addressed `/_f/` URLs that never break, no matter where the page moves.
```markdown
![Abstract waves](waves.svg "A figure in the margin"){.margin}
```
![Abstract waves](waves.svg "A figure in the margin, with {.margin}"){.margin}
The same figure as a margin note: it leans into the side zone left of the text on all but phone widths, alongside the text it belongs to.
{.wide} artwork spans the full content width:
```markdown
![Dunes](dunes.svg "Full-width artwork"){.wide}
```
![Dunes](dunes.svg "Full-width artwork between sections"){.wide}
""" """
LONG_READ = """\ GALLERY = """\
*An essay long enough to scroll, to demonstrate the gentle reveal of This page's banner is the **eyes** design — a critter in the grass in — picked from banner menu (⚙️ in the top right corner). The selection applies to current page and all its children, allowing differently themed sections be created. [Night Sky](night-sky) picked its own. You should also find the theme settings, which allow choosing overall site theme, fonts and transitions. You may wish to try the more playful **summer** theme which the eyes theme builds on.
headings, figures and code blocks as they enter the viewport.*
![Layered dunes](dunes.svg "Full-width artwork between sections"){.wide} ## Break out of the box!
## Chapter one ![Dunes](dunes.svg){.wide}
The distinction between a blog and a website is largely an accident of ::: aside
history. Early content management systems filed everything under "posts", ![Abstract waves](waves.svg)
stamped them with a date, and arranged them in reverse chronological order
under a `/blog/` prefix. Anything else was a "page", which lived somewhere
else entirely, often in a separate editing interface with separate rules.
But readers do not think in these terms. A reader follows a link, reads ## Aside boxes
what is there, and follows another link. The URL is a promise about where
something lives, not about which database table it came from. Pagerite
therefore treats every piece of content as a page: named, addressable, and
rendered on the fly.
## Chapter two When you have to sideline a bit with something important to say, use `::: aside` and end with `:::`, markdown between.
Consider what happens to URLs when the tooling leads the design. You get On larger screens they break outside the normal page bounds. `{.margin}` can be used to a similar effect without a box.
addresses like `/cms/frontpage` or `/blog/post1` — the name of the machine :::
leaking into the name of the thing. The slug should be chosen by the
author, the way a book's title is chosen, and it should sit at the root of
the site like the title sits on the cover.
Nesting still has its place. Structured content — documentation, a series, Images and text boxes can also be positioned for a more lively layout.
a portfolio — benefits from paths that mirror the structure. The
navigation on this very site is derived from the paths: open a section,
and you see what it contains. No menu editor, no duplication of structure
in two places.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod ![Shapes](shapes.svg "Floated with {.right}"){.right}
tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim
veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea
commodo consequat. Duis aute irure dolor in reprehenderit in voluptate
velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat
cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id
est laborum.
Sed ut perspiciatis unde omnis iste natus error sit voluptatem accusantium This text wraps around a left or right floated figure. The caption comes from the image title, with additional styling like `{.left width=240}` — brace attributes on the image itself.
doloremque laudantium, totam rem aperiam, eaque ipsa quae ab illo
inventore veritatis et quasi architecto beatae vitae dicta sunt explicabo.
Nemo enim ipsam voluptatem quia voluptas sit aspernatur aut odit aut
fugit, sed quia consequuntur magni dolores eos qui ratione voluptatem
sequi nesciunt.
## Chapter three Note how the layout may take different forms from a phone in portrait to widest of desktop browsers, not leaving large empty areas nor being constrained to a classic container box model.
On the reading experience itself: motion on the web is usually either Lifting off elements here and there makes a great difference to how your site is received!
absent or obnoxious. The interesting middle ground is motion that
acknowledges the reader's own movement — the scroll. Elements that fade
in as they enter the viewport give the page a sense of depth, as if the
content were arriving just in time.
Crucially, none of this may depend on the article. The author writes ### Design matters
Markdown; the effects come from the layout. And when the reader prefers
reduced motion, everything must hold still.
Neque porro quisquam est, qui dolorem ipsum quia dolor sit amet, Good graphical design gives a website a clear visual structure and makes information easy to understand at a glance. Layout, spacing, typography, color, and imagery should work together to establish hierarchy and guide attention naturally through the page. Consistency between sections also helps users quickly learn how the interface is organized.
consectetur, adipisci velit, sed quia non numquam eius modi tempora
incidunt ut labore et dolore magnam aliquam quaerat voluptatem. Ut enim ad
minima veniam, quis nostrum exercitationem ullam corporis suscipit
laboriosam, nisi ut aliquid ex ea commodi consequatur?
```text A strong website layout balances visual character with usability. Content should have enough space to remain readable, while navigation and important actions should be easy to find without dominating the design. Responsive layouts should preserve these relationships across different screen sizes rather than simply shrinking the desktop arrangement.
Quis autem vel eum iure reprehenderit
qui in ea voluptate velit esse quam nihil
molestiae consequatur, vel illum qui
dolorem eum fugiat quo voluptas nulla pariatur?
```
At vero eos et accusamus et iusto odio dignissimos ducimus qui blanditiis
praesentium voluptatum deleniti atque corrupti quos dolores et quas
molestias excepturi sint occaecati cupiditate non provident, similique
sunt in culpa qui officia deserunt mollitia animi, id est laborum et
dolorum fuga. Et harum quidem rerum facilis est et expedita distinctio.
## Chapter four
Nam libero tempore, cum soluta nobis est eligendi optio cumque nihil
impedit quo minus id quod maxime placeat facere possimus, omnis voluptas
assumenda est, omnis dolor repellendus. Temporibus autem quibusdam et aut
officiis debitis aut rerum necessitatibus saepe eveniet ut et voluptates
repudiandae sint et molestiae non recusandae.
Itaque earum rerum hic tenetur a sapiente delectus, ut aut reiciendis
voluptatibus maiores alias consequatur aut perferendis doloribus
asperiores repellat. And so we arrive back where we started: the blog and
the website were one thing all along. [Return to the front page](/).
""" """
NOTES_ON_URLS = """\ NIGHT_SKY = """\
A URL is part of the content. A few rules of thumb I keep coming back to: This page's banner is not an image or a code snippet — it's the **stars** banner design, picked from a dropdown in the banner editor (🖊️ in the banner corner). Nothing is stored in the page beyond that choice.
- Pick slugs like book titles, not like database keys. Banner designs are folders in `pagerite/themes/{name}/` — a `banner.css` plus a `banner.html` or `banner.svg` — so a design can be anything from a static gradient to an animated canvas like the starfield above. This site ships `stars` and `eyes` (a critter in the grass, seen on the [gallery](/showcase/gallery)), and themes can bring their own.
- Nest only when the structure is real.
- Once published, a URL is a promise. Redirect if you must break it.
> Cool URIs don't change; uncool ones at least apologise. Subpages inherit the nearest banner and design up their path, so a whole section can share one look — set one on a category and every page under it gets it, until a page overrides with its own. This page is a leaf: the design chosen here affects nothing else.
That's all. Short posts are posts too. A page can also carry its own banner HTML — an `<img>`, a styled div, a canvas with a script — which renders *on top of* the design's artwork, so author code always wins. But most of the time, picking a design is all you need.
""" """
CANVAS_NIGHTS = """\ # Moby-Dick; or, The Whale (1851), Herman Melville — public domain.
This post's banner is not an image at all — it's a `<canvas>` animated by # Chapter 1, abridged and headed. A real long-read: flowing sections,
a few lines of JavaScript embedded in the page's banner HTML. # figures, a list, side notes — not a feature showcase. (Engravings:
# Augustus Burnham Shute's illustrations for the 1892 edition, public
# domain; the wave is Hokusai, public domain.)
LOOMINGS = """\
*The opening of Herman Melville's Moby-Dick (1851), abridged — here to show what a longer article feels like: the multi-column layout on wide screens, images breaking up the text, and the gentle reveal as sections scroll into view.*
Banners on this site are arbitrary markup: an image, a gradient div, or a {dates}
small animated scene like the one above. Subpages inherit the nearest
banner up their path, so a whole section can share one look.
```js ![The Great Wave off Kanagawa](great-wave.jpg "Hokusai, c. 1831 — the sea, full-bleed with {.wide}"){.wide}
// the essence of the banner above
stars.forEach(s => { s.x = (s.x + s.speed * dt) % 1 })
```
No build step, no framework — the snippet is stored with the page and ## The watery part of the world
dropped into the header as-is.
Call me Ishmael. Some years ago — never mind how long precisely — having little or no money in my purse, and nothing particular to interest me on shore, I thought I would sail about a little and see the watery part of the world. It is a way I have of driving off the spleen and regulating the circulation. Whenever I find myself growing grim about the mouth; whenever it is a damp, drizzly November in my soul; whenever I find myself involuntarily pausing before coffin warehouses, and bringing up the rear of every funeral I meet; and especially whenever my hypos get such an upper hand of me, that it requires a strong moral principle to prevent me from deliberately stepping into the street, and methodically knocking people's hats off — then, I account it high time to get to sea as soon as I can. This is my substitute for pistol and ball. With a philosophical flourish Cato throws himself upon his sword; I quietly take to the ship. There is nothing surprising in this. If they but knew it, almost all men in their degree, some time or other, cherish very nearly the same feelings towards the ocean with me.
::: aside
Melville interrupts his story often — whole chapters on cetology, rope and chowder. Abridgments drop most of them, but notes like this one are where they would have gone.
:::
There now is your insular city of the Manhattoes, belted round by wharves as Indian isles by coral reefs — commerce surrounds it with her surf. Right and left, the streets take you waterward. Its extreme downtown is the battery, where that noble mole is washed by waves, and cooled by breezes, which a few hours previous were out of sight of land. Look at the crowds of water-gazers there.
Circumambulate the city of a dreamy Sabbath afternoon. Go from Corlears Hook to Coenties Slip, and from thence, by Whitehall, northward. What do you see? — Posted like silent sentinels all around the town, stand thousands upon thousands of mortal men fixed in ocean reveries. Some leaning against the spiles; some seated upon the pier-heads; some looking over the bulwarks of ships from China; some high aloft in the rigging, as if striving to get a still better seaward peep. But these are all landsmen; of week days pent up in lath and plaster — tied to counters, nailed to benches, clinched to desks. How then is this? Are the green fields gone? What do they here?
But look! here come more crowds, pacing straight for the water, and seemingly bound for a dive. Strange! Nothing will content them but the extremest limit of the land; loitering under the shady lee of yonder warehouses will not suffice. No. They must get just as nigh the water as they possibly can without falling in. And there they stand — miles of them — leagues. Inlanders all, they come from lanes and alleys, streets and avenues — north, east, south, and west. Yet here they all unite. Tell me, does the magnetic virtue of the needles of the compasses of all those ships attract them thither?
## Meditation and water
Once more. Say you are in the country; in some high land of lakes. Take almost any path you please, and ten to one it carries you down in a dale, and leaves you there by a pool in the stream. There is magic in it. Let the most absent-minded of men be plunged in his deepest reveries — stand that man on his legs, set his feet a-going, and he will infallibly lead you to water, if water there be in all that region. Should you ever be athirst in the great American desert, try this experiment, if your caravan happen to be supplied with a metaphysical professor. Yes, as every one knows, meditation and water are wedded for ever.
Ishmael sails from New Bedford, the whaling port south of Boston — Nantucket was the older, prouder whaling town, and he briefly considers it first.
{.margin}
### The artist's problem
But here is an artist. He desires to paint you the dreamiest, shadiest, quietest, most enchanting bit of romantic landscape in all the valley of the Saco. What is the chief element he employs? There stand his trees, each with a hollow trunk, as if a hermit and a crucifix were within; and here sleeps his meadow, and there sleep his cattle; and up from yonder cottage goes a sleepy smoke. Deep into distant woodlands winds a mazy way, reaching to overlapping spurs of mountains bathed in their hill-side blue. But though the picture lies thus tranced, and though this pine-tree shakes down its sighs like leaves upon this shepherd's head, yet all were vain, unless the shepherd's eye were fixed upon the magic stream before him.
Why did the poor poet of Tennessee, upon suddenly receiving two handfuls of silver, deliberate whether to buy him a coat, which he sadly needed, or invest his money in a pedestrian trip to Rockaway Beach? Why is almost every robust healthy boy with a robust healthy soul in him, at some time or other crazy to go to sea? Why upon your first voyage as a passenger, did you yourself feel such a mystical vibration, when you were first told that you and your ship were now out of sight of land? Why did the old Persians hold the sea holy? Why did the Greeks give it a separate deity, and own brother of Jove? Surely all this is not without meaning. And still deeper the meaning of that story of Narcissus, who because he could not grasp the tormenting, mild image he saw in the fountain, plunged into it and was drowned. But that same image, we ourselves see in all rivers and oceans. It is the image of the ungraspable phantom of life; and this is the key to it all.
## A simple sailor, right before the mast
Now, when I say that I am in the habit of going to sea whenever I begin to grow hazy about the eyes, and begin to be over conscious of my lungs, I do not mean to have it inferred that I ever go to sea as a passenger. For to go as a passenger you must needs have a purse, and a purse is but a rag unless you have something in it. Besides, passengers get sea-sick — grow quarrelsome — don't sleep of nights — do not enjoy themselves much, as a general thing; — no, I never go as a passenger; nor, though I am something of a salt, do I ever go to sea as a Commodore, or a Captain, or a Cook. I abandon the glory and distinction of such offices to those who like them. For my part, I abominate all honorable respectable toils, trials, and tribulations of every kind whatsoever. It is quite as much as I can do to take care of myself, without taking care of ships, barques, brigs, schooners, and what not.
![Moby Dick breeches a whaleboat](md-whale.jpg "Augustus Burnham Shute, 1892"){.right width=400}
No, when I go to sea, I go as a simple sailor, right before the mast, plumb down into the forecastle, aloft there to the royal mast-head. True, they rather order me about some, and make me jump from spar to spar, like a grasshopper in a May meadow. And at first, this sort of thing is unpleasant enough. It touches one's sense of honor, particularly if you come of an old established family in the land, the Van Rensselaers, or Randolphs, or Hardicanutes. And more than all, if just previous to putting your hand into the tar-pot, you have been lording it as a country schoolmaster, making the tallest boys stand in awe of you. The transition is a keen one, I assure you, from a schoolmaster to a sailor, and requires a strong decoction of Seneca and the Stoics to enable you to grin and bear it. But even this wears off in time.
What of it, if some old hunks of a sea-captain orders me to get a broom and sweep down the decks? What does that indignity amount to, weighed, I mean, in the scales of the New Testament? Do you think the archangel Gabriel thinks anything the less of me, because I promptly and respectfully obey that old hunks in that particular instance? Who ain't a slave? Tell me that. Well, then, however the old sea-captains may order me about — however they may thump and punch me about, I have the satisfaction of knowing that it is all right; that everybody else is one way or other served in much the same way — either in a physical or metaphysical point of view, that is; and so the universal thump is passed round, and all hands should rub each other's shoulder-blades, and be content.
And finally, what shall I say of the reasons for going a-whaling? Chief among them:
- The overwhelming idea of the great whale himself — such a portentous and mysterious monster roused all my curiosity.
- The undeliverable, nameless perils of the whale, and the attendants of the wondrous world of waters.
- The tormenting, mild image of the ungraspable phantom of life, seen in all rivers and oceans.
These were the things that finally drew me to the sea — and if they but knew it, almost all men cherish very nearly the same feelings towards the ocean with me.
""" """
SMALL_RELEASES = """\ SMALL_RELEASES = """\
Software wants to be shipped. The longer a change sits unmerged, the more Software wants to be shipped. The longer a change sits unmerged, the more it rots: context fades, conflicts accumulate, and the diff grows teeth.
it rots: context fades, conflicts accumulate, and the diff grows teeth.
1. Cut the scope until it fits in a day. 1. Cut the scope until it fits in a day.
2. Ship it behind whatever door you like. 2. Ship it behind whatever door you like.
3. Let real use argue with your assumptions. 3. Let real use argue with your assumptions.
A release is a conversation with reality. Small releases keep the A release is a conversation with reality. Small releases keep the conversation lively — and small *pieces* keep the whole thing standing, as [the comic on the About page](/about) illustrates all too well.
conversation lively.
""" """
CANVAS_BANNER = """\ ABOUT = """\
<canvas id="stars"></canvas> This site runs on **Pagerite**: FastAPI + html5tagger + kanta, with content written in Markdown and rendered on the fly.
<script>
(() => { - [How to edit this site](/docs/editing)
const c = document.getElementById("stars"); - [Everything Markdown can do](/docs/markdown)
const ctx = c.getContext("2d"); - [The showcase](/showcase/gallery)
const fit = () => { c.width = c.clientWidth; c.height = c.clientHeight; };
fit(); Pagerite keeps its dependency list short and knows every entry on it. Modern software in general builds on taller towers of other people's work:
addEventListener("resize", fit);
const stars = Array.from({ length: 110 }, () => ({ [![xkcd 2347: Dependency](https://imgs.xkcd.com/comics/dependency.png "xkcd 2347: Dependency"){width=280}](https://xkcd.com/2347/)
x: Math.random(), y: Math.random(),
r: Math.random() * 1.4 + 0.3, v: Math.random() * 0.05 + 0.01, *Replace this page with whatever your site is about.*
}));
let prev = performance.now();
(function frame(now) {
if (!c.isConnected) return;
const dt = Math.min(now - prev, 100); prev = now;
ctx.fillStyle = "#0b0e1d";
ctx.fillRect(0, 0, c.width, c.height);
ctx.fillStyle = "#cdd6ff";
for (const s of stars) {
s.x = (s.x + s.v * dt / 1000) % 1;
ctx.beginPath();
ctx.arc(s.x * c.width, s.y * c.height, s.r, 0, 7);
ctx.fill();
}
requestAnimationFrame(frame);
})(prev);
})();
</script>
""" """
EYES_BANNER = """\
<canvas id="eyes"></canvas>
<script>
(() => {
const c = document.getElementById("eyes");
const ctx = c.getContext("2d");
const BG = "#f3e9d7";
const fit = () => { c.width = c.clientWidth; c.height = c.clientHeight; };
fit();
addEventListener("resize", fit);
// Mouse in canvas coordinates; pupils wander idly when it goes stale.
let mx = 0, my = 0, lastMove = 0;
addEventListener("mousemove", (e) => {
const r = c.getBoundingClientRect();
mx = e.clientX - r.left;
my = e.clientY - r.top;
lastMove = performance.now();
});
// The pair of eyes is one critter: it wanders around the banner, and
// every so often ducks below the bottom edge, then pops back up.
let gx = 0.5, gy = 0.5; // group position (fractions of the canvas)
let tx = 0.5, ty = 0.5; // wander target
let yoff = 0, vy = 0; // vertical hide/pop spring (px)
let hidePhase = 0; // 0 = up, 1 = ducking, 2 = down, waiting
let nextMove = 0, nextHide = 4000 + Math.random() * 5000, resurfaceAt = 0;
// Per-eye pupil state: spring physics for goofy lag and overshoot.
const eyes = [{ x: 0, y: 0, vx: 0, vy: 0, pr: 0.3 }, { x: 0, y: 0, vx: 0, vy: 0, pr: 0.3 }];
let prev = performance.now();
(function frame(now) {
if (!c.isConnected) return;
const dt = Math.min(now - prev, 100) / 16.7; prev = now;
ctx.fillStyle = BG;
ctx.fillRect(0, 0, c.width, c.height);
const R = Math.min(c.height * 0.32, 70);
// Wander: ease toward a spot, pick a new one every few seconds.
if (now > nextMove && !hidePhase) {
tx = 0.15 + Math.random() * 0.7;
ty = 0.3 + Math.random() * 0.4;
nextMove = now + 2500 + Math.random() * 3500;
}
gx += (tx - gx) * 0.02 * dt;
gy += (ty - gy) * 0.02 * dt;
// Duck down, wait hidden, then spring back (underdamped = pops past
// the resting point and wobbles). Resurfaces at a new spot.
if (hidePhase === 0 && now > nextHide) hidePhase = 1;
if (hidePhase === 1 && yoff > c.height * 0.9) {
hidePhase = 2;
resurfaceAt = now + 500 + Math.random() * 900;
}
if (hidePhase === 2 && now > resurfaceAt) {
hidePhase = 0;
nextHide = now + 5000 + Math.random() * 7000;
tx = 0.15 + Math.random() * 0.7;
gx = tx;
nextMove = now + 3000 + Math.random() * 3000;
}
const yTarget = hidePhase ? c.height : 0;
vy += (yTarget - yoff) * 0.06 * dt;
vy *= 0.85;
yoff += vy * dt;
const cy0 = gy * c.height + yoff;
const cx0 = gx * c.width;
const watching = now - lastMove < 4000;
eyes.forEach((e, i) => {
const cx = cx0 + (i ? 1.3 : -1.3) * R;
// Pupil target: toward the cursor, or a slow idle drift.
let ptx, pty;
if (watching) {
const dx = mx - cx, dy = my - cy0;
const d = Math.hypot(dx, dy) || 1;
const reach = R * 0.45 * Math.min(1, d / 200);
ptx = (dx / d) * reach; pty = (dy / d) * reach;
} else {
ptx = Math.sin(now / 900 + i * 2) * R * 0.3;
pty = Math.cos(now / 1300 + i * 3) * R * 0.2;
}
// Spring toward the target (underdamped: overshoots, wobbles).
e.vx += (ptx - e.x) * 0.08 * dt; e.vy += (pty - e.y) * 0.08 * dt;
e.vx *= 0.82; e.vy *= 0.82;
e.x += e.vx * dt; e.y += e.vy * dt;
// Pupils dilate when the cursor comes close to the eye.
const near = Math.hypot(mx - cx, my - cy0) < R * 2.5;
e.pr += ((near ? 0.42 : 0.3) - e.pr) * 0.1 * dt;
// Sclera.
ctx.fillStyle = "#fff";
ctx.beginPath();
ctx.ellipse(cx, cy0, R, R * 1.15, 0, 0, 7);
ctx.fill();
// Iris + pupil + glint, clipped to the sclera.
ctx.save();
ctx.beginPath();
ctx.ellipse(cx, cy0, R, R * 1.15, 0, 0, 7);
ctx.clip();
ctx.fillStyle = "#7c5cff";
ctx.beginPath();
ctx.arc(cx + e.x, cy0 + e.y, R * 0.55, 0, 7);
ctx.fill();
ctx.fillStyle = "#1d1730";
ctx.beginPath();
ctx.arc(cx + e.x, cy0 + e.y, R * e.pr, 0, 7);
ctx.fill();
ctx.fillStyle = "#fff";
ctx.beginPath();
ctx.arc(cx + e.x - R * 0.15, cy0 + e.y - R * 0.18, R * 0.09, 0, 7);
ctx.fill();
ctx.restore();
ctx.strokeStyle = "#2b2440";
ctx.lineWidth = 2;
ctx.beginPath();
ctx.ellipse(cx, cy0, R, R * 1.15, 0, 0, 7);
ctx.stroke();
});
requestAnimationFrame(frame);
})(prev);
})();
</script>
"""
BLOG_BANNER = '<div style="background: linear-gradient(100deg, #14243d, #3d2b6b 45%, #7c5cff 75%, #ff5c8a)"></div>'
FRONT_BANNER = '<img src="/waves.svg" alt="">'
WAVES_SVG = """\ WAVES_SVG = """\
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 800 400"> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 800 400">
<defs> <defs>
@@ -422,28 +474,56 @@ DUNES_SVG = """\
</svg> </svg>
""" """
#: path -> (title, markdown, {filename: bytes}, banner HTML, menu order). #: path -> (title, markdown, {filename: bytes}, banner HTML, menu order,
#: Note there are deliberately no "docs" or "blog" landing pages: those #: banner design). Designs demonstrate per-page choice: the gallery
#: labels are created without content, so they render a placeholder page #: picks "eyes" (its page alone), night-sky picks "stars"; everything
#: and their nav links point at the first child (see views.first_leaf). #: else inherits the active theme's own design.
PAGES: dict[str, tuple[str, str, dict[str, bytes], str, float]] = { #: Note there are deliberately no "docs" or "showcase" landing pages:
"": ("Welcome", WELCOME, {"waves.svg": WAVES_SVG.encode()}, FRONT_BANNER, 1), #: those labels are created without content, so they render a placeholder
"about": ("About", ABOUT, {}, "", 2), #: page and their nav links point at the first child (see
"docs/editing": ( #: views.first_leaf). "showcase" is seeded explicitly (empty markdown,
"Writing Content", #: which the seeder leaves as content=None) purely to fix its menu order.
EDITING, PAGES: dict[str, tuple[str, str, dict[str, bytes], str, float, str | None]] = {
{"shapes.svg": SHAPES_SVG.encode()}, "": ("Welcome", WELCOME, {"waves.svg": WAVES_SVG.encode()}, "", 1, None),
"about": ("About", ABOUT, {}, "", 3, None),
"docs/editing": ("Editing This Site", EDITING, {}, "", 1, None),
"docs/markdown": ("Markdown", MD_ARTICLE, {}, "", 2, None),
"docs/markdown/images-and-layout": (
"Images and Layout",
MD_LAYOUT,
{
"shapes.svg": SHAPES_SVG.encode(),
"waves.svg": WAVES_SVG.encode(),
"dunes.svg": DUNES_SVG.encode(),
},
"", "",
1, 1,
None,
), ),
"blog/the-long-read": ( "showcase": ("Showcase", "", {}, "", 4, None),
"The Long Read", "showcase/gallery": (
LONG_READ, "Gallery",
{"dunes.svg": DUNES_SVG.encode()}, GALLERY,
BLOG_BANNER, {
"dunes.svg": DUNES_SVG.encode(),
"shapes.svg": SHAPES_SVG.encode(),
"waves.svg": WAVES_SVG.encode(),
},
"",
1, 1,
"eyes",
), ),
"blog/notes-on-urls": ("Notes on URLs", NOTES_ON_URLS, {}, EYES_BANNER, 2), "showcase/loomings": (
"blog/canvas-nights": ("Canvas Nights", CANVAS_NIGHTS, {}, CANVAS_BANNER, 3), "Loomings — a Long Read",
"blog/small-releases": ("Small Releases", SMALL_RELEASES, {}, "", 4), LOOMINGS,
{
"great-wave.jpg": _asset("great-wave.jpg"),
"md-whale.jpg": _asset("md-whale.jpg"),
},
"",
2,
None,
),
"showcase/night-sky": ("Night Sky", NIGHT_SKY, {}, "", 3, "stars"),
"showcase/small-releases": ("Small Releases", SMALL_RELEASES, {}, "", 4, None),
} }
+700
View File
@@ -0,0 +1,700 @@
"""Segmented translation round trip: prose out, translations back in.
A translator model mangles anything that is not plain prose — sentinels get
renumbered, ``![`` becomes sentence punctuation, stray ``<br>`` tags appear.
So the model is never shown any of it: a fragment (a Markdown chunk or a
node title) is parsed with the project's own markdown-it setup
(``markdown.make_md(verbatim=True)`` — extensions included, so container,
attrs, footnote and tasklist syntax never leaks into text tokens) and split
into **prose segments**: the merged text runs, plus image alt texts and
link/image titles. Only those cross the wire, as a plain list of strings
(Job.texts / Result.texts in translate.py) — accompanied, per segment, by
a CONTEXT (Job.contexts): a segment carved out of a larger block (a link
text, a partial run) carries the block's plain text, so the model sees the
sentence it lives in; whole-block segments are self-contextualizing and
carry "". Title fragments carry the article's opening instead (assigned by
the dispatcher from TransItem.context).
Reassembly is server-side offset splicing, not text the model produced:
each segment's source span was located at dispatch (``split``), and
``join`` swaps in the translations. Markup therefore cannot break — it
never left the server. A returned segment must still be pure prose itself
(the model could inject markup INTO a segment); anything else — count
mismatch, empty segment, markup tokens, a line that would start a new
block (a ``` or ::: fence would eat the rest of the block it lands in) —
rejects the whole result and the
fragment stays pending. Punctuation that is prose on the wire but syntax
in the splice context (quotes in a title attribute, brackets in an alt
text, "|" in a table row) is not worth a rejection either: it is swapped
for Unicode look-alikes (``_NEUTRAL``) before splicing.
A block of plain text, prose links and paired text formatting
(strong/em/s) crosses as ONE segment — link texts and formatted text
inline, in sentence context, with the Markdown stripped (the model
mangles it: sentinels get renumbered, ``**`` gets dropped or moved) —
because a label translated apart from its sentence comes back
grammatically incompatible with it (case government, particles, word
order). ``join`` re-inserts the link/formatting markdown into the
translated block at fuzzily matched positions (``_place_marks``): no
markers on the wire, the boundaries are found by aligning the mark's
source words to the translation's words by form similarity (``_find_mark``
— inflection, dropped articles and reordering tolerated), with the
source/translation weight ratio as fallback (the CJK path, where
cross-script form similarity is nil). Placement is approximate: better a
coherent sentence with a slightly shifted link than separately translated
snippets that don't fit together. Blocks with any other inline markup
(code, images, HTML) still split into runs at those boundaries.
Locating is best effort: a run that is not a verbatim source substring
(entity-decoded text, backslash escapes) is skipped — it simply stays in
the original language. A literal "<" in prose ("<1MB") is text, not
markup, but cannot cross as-is — "<" is the prose/markup boundary on the
wire, translators cut their output there — so it crosses encoded as the
fullwidth "" (``_encode``) and ``join`` decodes it back before
validating and splicing.
"""
import bisect
import difflib
import re
from typing import NamedTuple
from pagerite.markdown import make_md
#: The segmentation parser: the project's own markdown-it, verbatim flavor
#: (see make_md). Never used for rendering.
_MD = make_md(verbatim=True)
#: Any Unicode letter (digits and underscore are not prose).
_LETTER = re.compile(r"[^\W\d_]")
#: A GFM alert marker ([!NOTE] etc.) at the start of a blockquote's first
#: paragraph: syntax, not prose — stripped from the first segment.
_ALERT = re.compile(r"^\[![A-Za-z]+\][ \t]*")
#: Any {...} span: {placeholders} and attrs that ended up inside prose
#: (inline attrs are consumed by the parser; a lone {dates} is not).
_BRACES = re.compile(r"\{[^{}\n]*\}")
def _encode(text: str) -> str:
"""Wire form of a segment or context: a literal "<" as fullwidth "".
A "<" in prose is text, not markup ("<1MB" — a tag needs a letter or
/!?), but "<" is the prose/markup boundary on the wire (translators
cut output at the first "<", scripts/translator.py), so it cannot
cross as-is. join decodes it back before the pure_prose check and
splicing — anything tag-like the model may have formed around it is
still rejected there.
"""
return text.replace("<", "")
#: ASCII punctuation that is plain prose to the inline parser (so
#: pure_prose cannot catch it) but Markdown SYNTAX in a splice context:
#: quotes close a quoted image/link title, brackets the [...] of alt and
#: re-inserted link texts, "|" splits a table row, and "\" escapes the
#: character after it (a trailing one eats a title's closing quote).
#: Neutralized to Unicode look-alikes (join), which Markdown treats as
#: plain text everywhere — the quotes are curled the way typographer=True
#: renders them anyway.
_NEUTRAL = str.maketrans(
{
'"': "",
"'": "",
"[": "",
"]": "",
"\\": "",
"|": "",
}
)
#: A link's tail after its text: "](dest)", "](dest \"title\")", "][ref]",
#: "[]" or a bare "]" (shortcut reference); the destination may nest one
#: level of parens. Best effort — a mis-scan fails the span-reconstruction
#: check in _linked_block and the block falls back to per-run segments.
_LINK_TAIL = re.compile(r"\](?:\((?:\\.|[^()\\]|\([^()]*\))*\)|\[(?:\\.|[^\]])*\])?")
#: Weight units for mapping link boundaries from source to translation:
#: a word counts 1 and so does every single CJK ideograph (kana runs count
#: as one) — CJK has no spaces to count words by. Punctuation and
#: whitespace count nothing, so mapped boundaries always land on unit
#: starts.
_UNIT = re.compile(
r"[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]" # CJK ideographs: one unit each
r"|[\u3040-\u309f\u30a0-\u30ff]+" # kana runs: one unit each
r"|\w+" # anything else word-like (Latin, Cyrillic, Hangul, digits)
)
class Mark(NamedTuple):
"""One inline link or paired formatting (strong/em/s) inside a
whole-block segment: the source weight (unit count, see _UNIT) at the
inner text's start and end (fallback for mapping the boundaries into
the translation when fuzzy word alignment finds nothing, _find_mark),
the exact source syntax around the text ("[" / "](url)", "**" / "**",
...) and the source text itself — the words fuzzy alignment looks for,
and the fallback when the mapped slice comes out empty (better an
untranslated label than a broken "[](url)")."""
w_start: int
w_end: int
pre: str
post: str
inner: str
class Span(NamedTuple):
"""A segment's source span in the fragment: offsets for splicing the
translation back, the segment's source weight and the links to
re-insert into its translation (empty = a plain prose segment)."""
start: int
end: int
weight: int
marks: list[Mark]
def _weight(text: str) -> int:
"""The text's weight in translation-mapping units (see _UNIT)."""
return len(_UNIT.findall(text))
def _runs(children: list) -> list[str]:
"""Prose runs of an inline token's children, in order.
Text tokens merge across soft breaks into one run; every markup token
(emphasis, links, code, images, HTML, footnote refs, hard breaks) is a
run boundary. Link and image *text* is prose; autolink text (the URL
itself) is not. Image tokens contribute their alt-text children and
their title attribute.
"""
runs: list[str] = []
cur: list[str] = []
def flush() -> None:
if cur:
s = "".join(cur)
cur.clear()
if _LETTER.search(s):
runs.append(s)
skip = 0 # inside an autolink (its text is the URL — not prose)
for t in children:
if skip:
if t.type == "link_close":
skip -= 1
continue
if t.type == "text":
cur.append(t.content)
elif t.type == "softbreak":
cur.append("\n")
elif t.type == "link_open" and t.markup == "autolink":
flush()
skip = 1
elif t.type == "image":
flush()
if t.children:
runs.extend(_runs(t.children))
title = t.attrGet("title")
if title and _LETTER.search(title):
runs.append(title)
else:
flush()
if t.children:
runs.extend(_runs(t.children))
flush()
return runs
def _block_text(children: list) -> str:
"""The block's text as a reader sees it: text runs and link texts
merged (softbreaks as newlines); image alts, autolink URLs, code and
other markup content excluded. Used as the translation CONTEXT for
segments carved out of the block (link texts, partial runs): a lone
word translates differently than the same word inside its sentence."""
parts: list[str] = []
skip = 0 # inside an autolink (its text is the URL)
for t in children:
if skip:
if t.type == "link_close":
skip -= 1
continue
if t.type == "text":
parts.append(t.content)
elif t.type == "softbreak":
parts.append("\n")
elif t.type == "link_open" and t.markup == "autolink":
skip = 1
elif t.type == "image":
continue
elif t.children:
parts.append(_block_text(t.children))
return "".join(parts)
def _locate(source: str, needle: str, cursor: int) -> int:
"""The needle's offset in source at/after cursor, -1 when absent.
An occurrence preceded by a backslash is an escaped character, not the
token's source: keep looking (failing that, the run is skipped — it
stays in the original language).
"""
pos = source.find(needle, cursor)
while pos > 0 and source[pos - 1] == "\\":
pos = source.find(needle, pos + 1)
return pos
def _linked_block(
source: str, kids: list, cursor: int, strip_alert: bool
) -> tuple[Span, str] | None:
"""A whole-block segment for an inline of plain text, prose links and
paired text formatting (strong/em/s): (Span, wire text) with the links
and formatting as marks, or None when the block has any other shape —
the caller then falls back to per-run segments.
The block crosses the wire as one prose piece, link texts and formatted
text inline (the model is never shown any Markdown — it mangles it),
so a translation that inflects or reorders around them stays coherent;
join re-inserts the link/formatting syntax at weight-mapped positions.
The source span is located piece by piece and verified by
reconstruction; anything not byte-exact (entities, escapes, an odd
link tail) bails to the fallback.
"""
pieces: list[
tuple[str, str]
] = [] # (text, mark): "" plain, "link", else the delimiter
buf: list[str] = [] # current plain piece
link: list[str] | None = None # current mark's text parts
mark_kind = "" # the current mark's opener ("link" or the delimiter)
for tok in kids:
if tok.type in ("link_open", "strong_open", "em_open", "s_open"):
if link is not None or tok.markup == "autolink":
return None
if buf:
pieces.append(("".join(buf), ""))
buf = []
link = []
mark_kind = "link" if tok.type == "link_open" else tok.markup
elif tok.type in ("link_close", "strong_close", "em_close", "s_close"):
if (
link is None
or ("link" if tok.type == "link_close" else tok.markup) != mark_kind
):
return None
inner = "".join(link)
if not _LETTER.search(inner):
return None
pieces.append((inner, mark_kind))
link = None
elif tok.type in ("text", "softbreak"):
(link if link is not None else buf).append(
"\n" if tok.type == "softbreak" else tok.content
)
else: # code, images, HTML, footnote refs: run boundaries
return None
if link is not None:
return None # unbalanced (the parser should not do this)
if buf:
pieces.append(("".join(buf), ""))
if not any(mark for _, mark in pieces):
return None
if strip_alert and pieces and not pieces[0][1]:
# A GFM alert marker leading the blockquote's first paragraph is
# syntax; strip it from the wire text (it stays out of the span).
first = _ALERT.sub("", pieces[0][0], count=1)
if first.strip():
pieces[0] = (first, "")
else:
pieces.pop(0)
if not pieces:
return None
raw = "".join(text for text, _ in pieces)
lead = len(raw) - len(raw.lstrip())
wire = raw.strip()
if not _LETTER.search(wire) or _BRACES.search(wire):
return None
# Locate each piece verbatim, in order; the source slices between the
# located pieces are then the link syntax, exact by construction.
located: list[tuple[int, int]] = []
pos = cursor
for text_, _ in pieces:
at = _locate(source, text_, pos)
if at == -1:
return None
located.append((at, at + len(text_)))
pos = at + len(text_)
span_start, span_end = located[0][0], located[-1][1]
marks: list[Mark] = []
offset = 0 # raw (pre-strip) plain-text offset of the current piece
for i, ((text_, kind), (s, e)) in enumerate(zip(pieces, located)):
if not kind:
offset += len(text_)
continue
# The syntax around the text: the gap between pieces goes to the
# mark on its left as post (so between two marks the whole "](u)["
# or "**" is the first's post); a block-leading mark takes its
# opener in front of its text ("[" or the delimiter), a
# block-trailing one the scanned link tail or the close delimiter.
if i == 0:
opener = "[" if kind == "link" else kind
if s < len(opener) or source[s - len(opener) : s] != opener:
return None
pre, span_start = opener, s - len(opener)
elif pieces[i - 1][1]:
pre = "" # the previous mark's post covers the whole gap
else:
pre = source[located[i - 1][1] : s]
if i + 1 < len(pieces):
post = source[e : located[i + 1][0]]
elif kind == "link":
m = _LINK_TAIL.match(source, e)
if m is None:
return None
post, span_end = m.group(), m.end()
else:
if source[e : e + len(kind)] != kind:
return None
post, span_end = kind, e + len(kind)
ps = min(max(offset - lead, 0), len(wire))
pe = min(max(offset + len(text_) - lead, 0), len(wire))
if pe <= ps:
return None
marks.append(
Mark(_weight(wire[:ps]), _weight(wire[:pe]), pre, post, wire[ps:pe])
)
offset += len(text_)
# Verify: the marks must reconstruct the source span exactly (the only
# real risk is the guessed tail of a trailing link).
rec: list[str] = []
mi = 0
for text_, kind in pieces:
if kind:
mark = marks[mi]
mi += 1
rec += [mark.pre, text_, mark.post]
else:
rec.append(text_)
if source[span_start:span_end] != "".join(rec):
return None
return Span(span_start, span_end, _weight(wire), marks), _encode(wire)
def split(text: str) -> tuple[list[Span], list[str], list[str]]:
"""Split a fragment into (spans, segments, contexts): prose segments to
translate, their source spans in ``text`` for splicing the translations
back, and per-segment translation context.
A block of plain text, prose links and paired formatting (strong/em/s)
becomes ONE segment (link/formatted text inline, in context, Markdown
stripped), the links and formatting recorded as marks on its Span for
weight-mapped re-insertion in join. Other blocks split into text runs
at markup boundaries; runs containing {...} spans are carved further —
the braces stay out of the wire text. A run that cannot be located
verbatim in the source contributes no segment. A segment's context is
its block's plain text when the segment was carved OUT of a larger
block (a partial run); a segment that IS the whole block (a plain
paragraph, a heading, a linked block) is self-contextualizing and gets
"".
"""
spans: list[Span] = []
segments: list[str] = []
contexts: list[str] = []
cursor = 0
blockquote_fresh = 0 # blockquote depth whose first inline is upcoming
def emit(run: str, at: int, ctx: str) -> None:
"""Carve {...} spans out of the located run; emit the prose pieces,
stripped — padding whitespace stays in the template, off the wire.
A literal "<" crosses encoded (``_encode``): it is text, not
markup, but the wire keeps "<" as the prose/markup boundary."""
pieces = []
pos = 0
for m in _BRACES.finditer(run):
pieces.append((pos, m.start()))
pos = m.end()
pieces.append((pos, len(run)))
for p0, p1 in pieces:
raw = run[p0:p1]
piece = raw.strip()
if _LETTER.search(piece):
start = at + p0 + (len(raw) - len(raw.lstrip()))
spans.append(Span(start, start + len(piece), 0, []))
segments.append(_encode(piece))
contexts.append(ctx)
tokens = _MD.parse(text)
for t in tokens:
if t.type == "blockquote_open":
blockquote_fresh += 1
elif t.type == "blockquote_close":
blockquote_fresh -= 1
elif t.type == "inline":
kids = t.children or []
# An alert marker ([!NOTE]) leading a blockquote's first
# paragraph is syntax; both paths strip it. (Only the first
# inline of the blockquote can carry it — the flag clears on
# the first inline seen.)
alert = bool(blockquote_fresh)
blockquote_fresh = 0
linked = _linked_block(text, kids, cursor, strip_alert=alert)
if linked is not None:
span, wire = linked
spans.append(span)
segments.append(wire)
contexts.append("")
cursor = span.end
continue
runs = _runs(kids)
block = _encode(_block_text(kids).strip())
if alert and runs:
run = _ALERT.sub("", runs[0], count=1)
if _LETTER.search(run):
runs[0] = run
else:
runs.pop(0)
for run in runs:
ctx = block if block and _encode(run.strip()) != block else ""
pos = _locate(text, run, cursor)
if pos != -1:
emit(run, pos, ctx)
cursor = pos + len(run)
elif "\n" in run:
# Indented continuation lines etc. break the verbatim
# match: locate each line separately instead.
for part in run.split("\n"):
if not _LETTER.search(part):
continue
pos = _locate(text, part, cursor)
if pos != -1:
emit(part, pos, ctx)
cursor = pos + len(part)
return spans, segments, contexts
#: Block-level Markdown a translation must not introduce: a segment is
#: spliced INSIDE a block of the fragment, so a line starting a heading,
#: quote, list, code/container fence or a setext/thematic-break underline
#: would break the fragment's block structure — a ``` or ::: line eats the
#: rest of the fence it lands in, closing fence included. pure_prose only
#: parses inline and lets such lines through as softbreak prose, so join
#: rejects them here. Blank lines split the host block and are rejected
#: too (a faithful translation of a single block has none).
_BLOCK = re.compile(
r"^[ \t]*(?:#{1,6}(?:[ \t]|$)|>[ \t]?|(?:[-+*]|\d{1,9}[.)])[ \t]|`{3,}|~{3,}|:{3,}(?:[ \t]|$)"
r"|-(?:[ \t]*-){2,}[ \t]*$|=[ =]*$|_(?:[ \t]*_){2,}[ \t]*$)",
re.M,
)
_BLANK = re.compile(r"\n[ \t]*\n")
def pure_prose(text: str) -> bool:
"""True when the text parses as nothing but prose (text and softbreak
tokens) — the acceptance test for a translated segment: the model may
not return markup of its own (a `<br>` here would splice live HTML into
the fragment)."""
children = _MD.parseInline(text)[0].children or []
return all(t.type in ("text", "softbreak") for t in children)
def _word_sim(a: str, b: str) -> float:
"""How likely two words are the same term across a translation, 0..1.
A case-folded exact match is 1; otherwise the better of the sequence
ratio and the shared-prefix ratio — inflection and derivational change
mostly move the ending ("banana" -> "banaanilla") or drop an article or
preposition around it. Case-folded so capitalization differences across
languages don't hide a term, with a small bonus when BOTH sides are
capitalized: a mid-sentence capital on both sides is likely the same
name (capitalization conventions differ per language, so its absence
proves nothing).
"""
bonus = 0.1 if a[:1].isupper() and b[:1].isupper() else 0.0
a, b = a.casefold(), b.casefold()
if a == b:
return 1.0
prefix = 0
for ca, cb in zip(a, b):
if ca != cb:
break
prefix += 1
sim = max(
difflib.SequenceMatcher(None, a, b).ratio(),
prefix / max(len(a), len(b)),
)
return min(1.0, sim + bonus)
#: Alignment costs for _find_mark: skipping a translation word (an article
#: or preposition the target language added) is cheap, skipping a source
#: word (one the translation dropped) costs more — a mark whose words
#: mostly vanished is no match at all. Every matched pair pays _MATCH, so
#: aligning a word to a lookalike-nothing (similarity below _MATCH) is
#: worse than skipping it.
_GAP_T = 0.25
_GAP_S = 0.6
_MATCH = 0.3
def _find_mark(
src: list[str], units: list[re.Match], start: int
) -> tuple[int, int] | None:
"""Locate a mark's source words in the translation's units (from unit
index ``start`` on), as the (start, end) unit-index span of the best
fuzzy alignment; None when no alignment is convincing (the caller falls
back to the weight ratio).
Word-for-word alignment with skips (_word_sim per pair, _GAP_T/_GAP_S
per skipped word): reordering is handled by the search itself, an added
or dropped article/preposition by the skip penalties. Accepted only
with an anchor — one pair of similarity >= 0.7 — and a decent average,
so a fully reworded label doesn't snap onto chance lookalikes.
"""
tgt = [u.group() for u in units[start:]]
n, m = len(src), len(tgt)
if not n or not m:
return None
# dp[i][j]: best score aligning src[:i] to tgt[:j]; a free tail (the
# answer is the best dp[n][j] over j) keeps trailing words costless.
dp = [[0.0] * (m + 1) for _ in range(n + 1)]
back: list[list[tuple[int, int]]] = [[(0, 0)] * (m + 1) for _ in range(n + 1)]
for i in range(1, n + 1):
dp[i][0] = dp[i - 1][0] - _GAP_S
back[i][0] = (i - 1, 0)
for j in range(1, m + 1):
options = [
(
dp[i - 1][j - 1] + _word_sim(src[i - 1], tgt[j - 1]) - _MATCH,
(i - 1, j - 1),
),
(dp[i][j - 1] - _GAP_T, (i, j - 1)),
(dp[i - 1][j] - _GAP_S, (i - 1, j)),
]
dp[i][j], back[i][j] = max(options, key=lambda o: o[0])
j_end = max(range(m + 1), key=lambda j: dp[n][j])
pairs: list[tuple[int, int]] = [] # matched (source, target) indices
i, j = n, j_end
while i > 0:
pi, pj = back[i][j]
if (pi, pj) == (i - 1, j - 1):
pairs.append((i - 1, j - 1))
i, j = pi, pj
if not pairs:
return None
pairs.reverse() # backtracking collected them last-first
sims = [_word_sim(src[a], tgt[t]) for a, t in pairs]
# Weak pairs at the span's ends are not part of the label (a declined
# neighbor the DP matched for a pittance) — trim them off.
while len(sims) > 1 and sims[0] < 0.5:
pairs.pop(0)
sims.pop(0)
while len(sims) > 1 and sims[-1] < 0.5:
pairs.pop()
sims.pop()
if max(sims) < 0.7 or sum(sims) / len(sims) < 0.45:
return None
return start + pairs[0][1], start + pairs[-1][1] + 1
def _place_marks(translation: str, weight: int, marks: list[Mark]) -> str | None:
"""Re-insert a whole-block segment's links into its translation.
Each mark's boundaries are found by fuzzy word-form alignment
(_find_mark): the mark's source words are matched against the
translation's units by form similarity — no markers on the wire
(sentinels never survived the model), no assumption that word order or
count survived either. Slicing exactly at unit boundaries keeps the
whitespace between the mark and its neighbors in the plain text, where
it belongs. A mark with no convincing alignment falls back to its
source weight ratio (units before the boundary / total applied to the
translation's units) — the pre-fuzz heuristic, still the CJK path,
where form similarity across scripts is nil. A boundary landing empty
degrades to the source link text: better an untranslated label than a
broken "[](url)". None when the translation has no units to map onto
(the caller rejects the result).
"""
units = list(_UNIT.finditer(translation))
total = len(units)
if not total or not weight:
return None
starts = [u.start() for u in units]
bounds = starts + [len(translation)]
out: list[str] = []
cur = 0 # char cursor: never before the previous mark's end
ucur = 0 # unit cursor, the same monotonicity in unit indices
for mark in marks:
found = _find_mark(_UNIT.findall(mark.inner), units, ucur)
if found is not None:
u1, u2 = found
x1, x2 = units[u1].start(), units[u2 - 1].end()
else:
x1 = bounds[min(round(mark.w_start / weight * total), total)]
x2 = bounds[min(round(mark.w_end / weight * total), total)]
x1 = max(x1, cur)
x2 = max(x2, x1)
# The slice ends at the next unit's start, so the whitespace
# and punctuation before that unit is inside it — but it
# belongs BETWEEN the mark and the following word, not in the
# inner text: end the inner text at its last unit and leave
# the rest for the following slice.
raw = translation[x1:x2]
inner_units = list(_UNIT.finditer(raw))
x2 = x1 + inner_units[-1].end() if inner_units else x1
inner = translation[x1:x2].strip() or mark.inner
out += [translation[cur:x1], mark.pre, inner, mark.post]
cur = x2
ucur = bisect.bisect_left(starts, x2)
out.append(translation[cur:])
return "".join(out)
def join(original: str, spans: list[Span], texts: list[str]) -> str | None:
"""Splice translated segments back into the original fragment; None on
any validation failure (count mismatch, empty, non-prose or
block-structure segment) — the caller drops the result and the fragment
stays pending. Segments with marks (a block that crossed as one piece)
get their links re-inserted at weight-mapped positions after the prose
check.
Markdown-significant ASCII punctuation that pure_prose cannot see
(plain text inline, syntax in the splice context — quoted titles, alt
and link texts, table rows) is neutralized to Unicode look-alikes
(``_NEUTRAL``) before splicing and mark placement (the swap is
char-for-char, so unit alignment is unaffected); lines that would
start a new block (a heading, a ``` or ::: fence — they would eat the
rest of the block/fence they land in) reject the result outright
(``_BLOCK``, ``_BLANK``)."""
if len(texts) != len(spans):
return None
out: list[str] = []
cursor = 0
for span, translation in zip(spans, texts):
# Decode the wire form ("" back to "<") first: pure_prose then
# validates exactly what gets spliced — a "" the model formed
# into anything tag-like is markup and rejects the result.
translation = translation.replace("", "<")
if (
not translation.strip()
or not pure_prose(translation)
or _BLOCK.search(translation)
or _BLANK.search(translation.strip())
):
return None
translation = translation.translate(_NEUTRAL)
if span.marks:
translation = _place_marks(translation, span.weight, span.marks)
if translation is None:
return None
out.append(original[cursor : span.start])
out.append(translation)
cursor = span.end
out.append(original[cursor:])
return "".join(out)
def has_prose(text: str) -> bool:
"""True when the fragment yields at least one translatable segment.
Chunks that are all markup, code, placeholders or reference definitions
have no business reaching the model: every language renders them from
the original chunk."""
return bool(split(text)[1])
+363
View File
@@ -0,0 +1,363 @@
"""Shared core: site constants, the kanta database, and the render cache.
Everything the route modules (files, api, tracking, pages) need that is not
a route itself: environment-derived paths and tunables, the ``Data`` root
with its ``Kanta`` handle (migrations in pagerite.migrations), the analytics
store, the page render cache
(``_render_html``/``_cached_body``/``_html_response`` plus the
``_render_gen`` ETag generation, bumped by ``_invalidate_pages`` on every
content/settings write), the translator ``dispatcher``, the slug charset
helpers, and the database bootstrap hooks (demo seed, translator defaults).
Importable by every other pagerite module without cycles.
"""
import logging
import os
import re
import secrets
from datetime import UTC, datetime
from functools import lru_cache
from pathlib import Path
import blake3
from fastapi import HTTPException, Request
from fastapi.responses import Response
from kanta import Kanta
from zstandard import ZstdCompressor
from pagerite import analytics, i18n, seed, translate, views
from pagerite.__main__ import DEVMODE
from pagerite.chunks import store_chunks
from pagerite.config import load
from pagerite.data import (
Data,
Node,
append_order,
find_slot,
prettify,
)
logger = logging.getLogger(__name__)
#: The CLI-passed configuration (PAGERITE_CONFIG) for this process.
config = load()
# Site identity: the hostname comes from the CLI (first positional argument,
# passed in PAGERITE_CONFIG) and names the per-site data directory
# ``<hostname>/{content.kantadb, analytics.json, files}`` under the cwd.
HOSTNAME = config.hostname
SITE_DIR = Path(HOSTNAME)
#: Public origin of the site, used for absolute social/canonical/sitemap
#: URLs. Localhost serves varying ports, so it falls back to the request's
#: own base URL instead.
SITE_URL = f"https://{HOSTNAME}" if HOSTNAME != "localhost" else ""
DB_PATH = os.getenv("PAGERITE_DB", str(SITE_DIR / "content.kantadb"))
# Visit analytics go to their own JSON file, not the kanta database.
ANALYTICS_PATH = Path(os.getenv("PAGERITE_ANALYTICS", str(SITE_DIR / "analytics.json")))
analytics_store = analytics.Store(ANALYTICS_PATH)
# Content-addressed file store (uploads, seed assets, fetched favicons):
# files on disk under hash-prefixed names, cached in RAM, served at /_f/.
FILES_DIR = Path(os.getenv("PAGERITE_FILES", str(SITE_DIR / "files")))
# Uploaded images are thumbnailed to this size and recompressed to AVIF
# (primary), with WebP and JPEG fallbacks re-encoded from the AVIF at
# somewhat lower quality (similar or smaller file size); the untouched
# original is kept alongside as ``<hash>.orig<ext>`` (never served).
IMAGE_MAXSIZE = 1920
IMAGE_QUALITY = 60
IMAGE_WEBP_QUALITY = 50
IMAGE_JPG_QUALITY = 55
# Favicons get the same derivatives but thumbnailed much smaller — 192px
# is plenty (browsers scale down for the 16x16 tab icon themselves).
FAVICON_MAXSIZE = 192
# Our own data root; kanta edits it in place, reads are plain attribute access.
data = Data()
kanta = Kanta(DB_PATH, data, migrations="pagerite.migrations")
# Dynamic HTML is compressed per request at level 9 (static assets are
# already pre-compressed by fastapi-vue's Frontend).
_zstd = ZstdCompressor(9)
def _render_html(
kind: str,
path: str,
base_url: str,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> str:
"""Render one of the generated pages (see _html_response)."""
if kind == "page":
# A selected language without an actual translation renders the
# original (translation is None; see docs/localization.md).
original = i18n.primary_lang(data.menu, path)
translation = (
i18n.get_translation(data, path, lang) if lang != original else None
)
return views.render_page(
data.menu,
data,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
base_url,
transition=data.transition,
lang=lang,
translation=translation,
link_lang=link_lang,
)
if kind == "category":
# A category has no Markdown of its own; only the title map
# localizes (heading, navigation, card text).
original = i18n.primary_lang(data.menu, path)
translation = (
i18n.Translation(titles=i18n.title_map(data, lang))
if lang != original
else None
)
return views.render_category(
data.menu,
data,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
base_url,
transition=data.transition,
lang=lang,
translation=translation,
link_lang=link_lang,
)
if kind == "not-found":
return views.render_not_found(
data.menu,
path,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
transition=data.transition,
)
return views.render_analytics(
data.menu,
data.brand,
data.custom_css,
data.theme,
data.favicon,
data.brand_html,
transition=data.transition,
)
# Render generation: bumped (and the body cache cleared) by every
# content/settings write, so page ETags and cached copies invalidate when
# navigation-affecting changes happen. In-memory only — not database state.
_render_gen = 0
def _invalidate_pages() -> None:
"""Drop cached page bodies and bump the render generation (ETags);
any content change also re-runs translation dispatch."""
global _render_gen
_render_gen += 1
_cached_body.cache_clear()
dispatcher.schedule()
@lru_cache(maxsize=128)
def _cached_body(
kind: str,
path: str,
base_url: str,
zstd: bool,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> bytes:
"""Rendered page body; cleared by _invalidate_pages on any
content/settings change. base_url feeds the social meta URLs, zstd
selects the stored encoding (both variants are cached rather than
re-compressed) and lang the selected language (not the raw
Accept-Language header, which would blow up the cache key space).
link_lang is the ?lang= override replicated onto the navigation links:
a query render and a header-selected render of the same language differ
in their links, so they are cached separately.
"""
body = _render_html(kind, path, base_url, lang, link_lang).encode()
return _zstd.compress(body) if zstd else body
def _html_response(
request: Request,
kind: str,
path: str,
status_code: int = 200,
headers: dict | None = None,
etag: bool = False,
lang: str = i18n.ORIGINAL_LANGUAGE,
link_lang: str = "",
) -> Response:
"""Response for a generated page, zstd-compressed when the client
accepts it (no gzip fallback).
Done per handler rather than in middleware so that Frontend's
already-compressed asset responses are never touched. The ETag stays
identical across encodings (revalidation compares it before
compression); ``vary: accept-encoding`` keeps caches from mixing the
representations. In dev the cache is bypassed so theme/design edits on
disk apply immediately.
``etag=True`` derives the validator from a blake3 hash of the
(uncompressed) body — for pages like /_a that have no Node whose
modified timestamp could serve as one — and answers matching
if-none-match revalidations with a 304.
"""
zstd = "zstd" in request.headers.get("accept-encoding", "")
# Absolute social/canonical URLs use the site's public origin; on
# localhost (varying ports) fall back to the request's own base URL.
base_url = SITE_URL or str(request.base_url).rstrip("/")
if DEVMODE:
identity = _render_html(kind, path, base_url, lang, link_lang).encode()
body = _zstd.compress(identity) if zstd else identity
else:
identity = _cached_body(kind, path, base_url, False, lang, link_lang)
body = (
_cached_body(kind, path, base_url, True, lang, link_lang)
if zstd
else identity
)
h = dict(headers or {})
# Content varies by language (Accept-Language selects a translation)
# and by encoding; keep caches from mixing either representation.
h["vary"] = "accept-language" + (", accept-encoding" if zstd else "")
if etag:
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
h["etag"] = tag
if request.headers.get("if-none-match") == tag:
return Response(status_code=304, headers=h)
if zstd:
h["content-encoding"] = "zstd"
return Response(body, status_code, h, media_type="text/html")
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
def _is_reserved(path: str) -> bool:
"""Slug shape that content may never use: each segment must be lower-case
ASCII letters, digits, hyphens and underscores (underscores may not be
the first character), and dots are never allowed.
"""
if path == "":
return False
return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
def _check_reserved(path: str) -> None:
"""Reject paths that do not follow the slug charset."""
if _is_reserved(path):
raise HTTPException(
400,
'slugs may only use a-z, 0-9, "-" and "_" (not as the first character), and no dots',
)
def _ensure(menu: dict[str, Node], path: str) -> Node:
"""Return the node at ``path``, creating it and any missing ancestors
(content-less category labels) appended at the end of their level."""
nodes = menu
node = None
for seg in path.split("/"):
node = nodes.get(seg)
if node is None:
node = Node(title=prettify(seg), order=append_order(nodes))
nodes[seg] = node
nodes = node.children
return node
def _remove_page(menu: dict[str, Node], path: str) -> bool:
"""Delete the node at ``path`` (inside a transaction).
A node with children becomes a content-less category label; a childless
node is removed entirely. Returns False if the path does not exist.
"""
slot = find_slot(menu, path)
node = slot[0].get(slot[1]) if slot else None
if node is None:
return False
if node.children:
node.chunks = None
node.modified = datetime.now(UTC)
else:
del slot[0][slot[1]]
return True
def _store_seed_file(
markdown: str, banner: str, orig: str, body: bytes
) -> tuple[str, str]:
"""Store a seed file content-addressed and point references at /_f/.
Images get the same AVIF/WebP/JPEG derivatives as uploads and are
linked extension-less; other content is stored as-is with its
extension."""
from pagerite.files import _ext, store_image # lazy: files imports state
ext = _ext(orig)
name = store_image(body, ext, derive=ext != ".gif")
markdown = markdown.replace(f"]({orig}", f"](/_f/{name}")
banner = banner.replace(f'src="/{orig}"', f'src="/_f/{name}"')
banner = banner.replace(f'src="{orig}"', f'src="/_f/{name}"')
return markdown, banner
@kanta.bootstrap
def _seed(data: Data) -> None:
"""Write the demo pages on database creation (never on existing dbs)."""
for path in seed.PAGES:
title, markdown, files, banner, order, design = seed.PAGES[path]
for orig, body in files.items():
markdown, banner = _store_seed_file(markdown, banner, orig, body)
node = _ensure(data.menu, path)
node.title = title
# Empty markdown means a pure category label (e.g. "showcase",
# seeded only to carry a banner design): leave chunks as None so
# the node renders the placeholder and nav points at its children.
if markdown:
node.chunks = store_chunks(data.chunks, markdown)
node.banner = banner
node.banner_design = design
node.order = order
#: Translator key format: 12 lowercase alphanumeric characters — not
#: brute-forceable over a WebSocket handshake, still human-manageable.
#: The editor's lang tab generates further keys in the same format.
_KEY_ALPHABET = "abcdefghijklmnopqrstuvwxyz0123456789"
@kanta.bootstrap
def _translator_defaults(data: Data) -> None:
"""Translator defaults on database creation: the first service key and
the wanted target languages (Spanish and Chinese — English is the
original language, never a translation target). Further keys are
managed in the editor shell's lang tab."""
key = "".join(secrets.choice(_KEY_ALPHABET) for _ in range(12))
data.translate_keys[key] = "default"
data.translate_langs = {"es": True, "zh": True}
# The translator dispatcher — protocol, connected clients and the job
# pipeline live in translate.py; its WebSocket route is in api.py.
dispatcher = translate.Dispatcher(data, kanta, _invalidate_pages)
+74
View File
@@ -0,0 +1,74 @@
/* Corporate banner design: sizing and colors for the geometric artwork
(inlined by the backend into #page-banner). The cb-* classes recolor the
SVG from the active palette (var(--accent)), so one SVG serves light
and dark — and other themes too. */
/* Artwork colors, light mode */
.cb-bg0 {
stop-color: #ffffff;
}
.cb-bg1 {
stop-color: #e6eefe;
}
.cb-r0 {
stop-color: var(--accent);
}
.cb-r1 {
stop-color: #00b3ff;
}
.cb-g0,
.cb-g1 {
stop-color: var(--accent);
}
.cb-dot {
fill: var(--accent);
}
.cb-orbit {
stroke: var(--accent);
}
.cb-spark {
fill: var(--accent);
}
/* Artwork colors, dark mode */
@media (prefers-color-scheme: dark) {
.cb-bg0 {
stop-color: #0d1830;
}
.cb-bg1 {
stop-color: #0a1122;
}
.cb-r0 {
stop-color: #2f7bff;
}
.cb-r1 {
stop-color: #00d0ff;
}
.cb-g0,
.cb-g1 {
stop-color: #2f7bff;
}
.cb-dot {
fill: #4d8dff;
}
.cb-orbit {
stroke: #4d8dff;
}
.cb-spark {
fill: #6ea8ff;
}
}
+1 -1
View File
@@ -1,4 +1,4 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1600 360" preserveAspectRatio="xMidYMid slice"> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1600 360" preserveAspectRatio="xMidYMax slice">
<defs> <defs>
<linearGradient id="cbg" x1="0" y1="0" x2="0" y2="1"> <linearGradient id="cbg" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" class="cb-bg0"/> <stop offset="0" class="cb-bg0"/>

Before

Width:  |  Height:  |  Size: 2.0 KiB

After

Width:  |  Height:  |  Size: 2.0 KiB

@@ -1,9 +1,9 @@
/* Corporate theme: bright and bold professional. Saturated royal-blue /* Corporate theme: bright and bold professional. Saturated royal-blue
gradients on white, geometric Montserrat display type over Inter body, gradients on white, geometric Montserrat display type over Inter body,
and a genuinely large brand with a soft blue overlap shadow. Automatic and a genuinely large brand with a soft blue overlap shadow. Automatic
dark mode keeps the same saturated blue identity on deep navy; the dark mode keeps the same saturated blue identity on deep navy. The
banner artwork (inlined by the backend) is recolored from here via the companion banner design (banner.css, artwork inlined by the backend)
cb-* classes, so one SVG serves both modes. */ recolors its SVG from the active palette via the cb-* classes. */
:root { :root {
color-scheme: light dark; color-scheme: light dark;
@@ -16,40 +16,7 @@
--line: #12203f14; --line: #12203f14;
--font-body: var(--font-inter); --font-body: var(--font-inter);
--font-heading: var(--font-montserrat); --font-heading: var(--font-montserrat);
} --code-x-height: 0.546; /* Inter's x-height ratio */
/* Banner artwork colors, light mode */
.cb-bg0 {
stop-color: #ffffff;
}
.cb-bg1 {
stop-color: #e6eefe;
}
.cb-r0 {
stop-color: var(--accent);
}
.cb-r1 {
stop-color: #00b3ff;
}
.cb-g0,
.cb-g1 {
stop-color: var(--accent);
}
.cb-dot {
fill: var(--accent);
}
.cb-orbit {
stroke: var(--accent);
}
.cb-spark {
fill: var(--accent);
} }
@media (prefers-color-scheme: dark) { @media (prefers-color-scheme: dark) {
@@ -64,45 +31,6 @@
/* Code wells stay navy in dark mode (light mode uses --surface). */ /* Code wells stay navy in dark mode (light mode uses --surface). */
--code-bg: #0d1b3e; --code-bg: #0d1b3e;
} }
/* Banner artwork colors, dark mode */
.cb-bg0 {
stop-color: #0d1830;
}
.cb-bg1 {
stop-color: #0a1122;
}
.cb-r0 {
stop-color: #2f7bff;
}
.cb-r1 {
stop-color: #00d0ff;
}
.cb-g0,
.cb-g1 {
stop-color: #2f7bff;
}
.cb-dot {
fill: #4d8dff;
}
.cb-orbit {
stroke: #4d8dff;
}
.cb-spark {
fill: #6ea8ff;
}
}
::selection {
background: var(--accent);
color: #fff;
} }
/* Genuinely large solid brand with a soft blue shadow overlapping the /* Genuinely large solid brand with a soft blue shadow overlapping the
@@ -111,7 +39,7 @@
font-size: clamp(4rem, 11vw, 8.5rem); font-size: clamp(4rem, 11vw, 8.5rem);
font-weight: 800; font-weight: 800;
letter-spacing: -0.04em; letter-spacing: -0.04em;
line-height: 1; line-height: 1.1;
white-space: nowrap; white-space: nowrap;
color: var(--accent2); color: var(--accent2);
filter: drop-shadow(0 0.4rem 1.4rem rgb(10 92 255 / 0.3)); filter: drop-shadow(0 0.4rem 1.4rem rgb(10 92 255 / 0.3));
@@ -124,11 +52,6 @@
} }
} }
#banner {
min-height: 15rem;
border-bottom: none;
}
#nav { #nav {
font-size: 1.05em; font-size: 1.05em;
font-weight: 600; font-weight: 600;
@@ -183,19 +106,18 @@ article h2 {
article h3 { article h3 {
font-weight: 700; font-weight: 700;
font-size: 0.95rem;
letter-spacing: 0.08em; letter-spacing: 0.08em;
text-transform: uppercase;
color: var(--muted); color: var(--muted);
} }
blockquote { blockquote {
border-left-color: var(--accent); border-inline-start-color: var(--accent);
background: color-mix(in oklab, var(--accent) 6%, transparent); background: color-mix(var(--accent) 6%, transparent);
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
margin: 0 -0.9rem 1rem calc(-0.25rem - 0.9rem); margin: 0 0 1rem;
margin-inline: calc(-0.25rem - 0.9rem) -0.9rem;
border-radius: 6px; border-radius: 6px;
} }
@@ -204,16 +126,12 @@ blockquote {
bar stays in both. */ bar stays in both. */
pre { pre {
border: 1px solid transparent; border: 1px solid transparent;
border-left: 0.25rem solid var(--accent); border-inline-start: 0.25rem solid var(--accent);
/* Text on the paragraph edge: the box extends by padding + border. */ /* Text on the paragraph edge: the box extends by padding + border. */
margin-left: calc(-0.8rem - 0.25rem); margin-inline-start: calc(-0.8rem - 0.25rem);
border-radius: 6px; border-radius: 6px;
} }
img { img {
border-radius: 4px; border-radius: 4px;
} }
::view-transition {
background: var(--bg);
}
+20
View File
@@ -0,0 +1,20 @@
/* Crossfade page transition. Injected by the backend as #pagerite-transition
when the "crossfade" transition is selected in the site settings. The old
snapshot stays fully opaque underneath while the new one fades in on top —
never a dip to black. Direction-neutral, so html.nav-back needs no
mirroring (and same-section html.nav-fade changes nothing). */
@keyframes nav-fade-in {
from {
opacity: 0;
}
}
::view-transition-old(root),
::view-transition-new(root) {
mix-blend-mode: normal;
animation: none;
}
::view-transition-new(root) {
animation: 200ms ease-in-out nav-fade-in;
}
+98
View File
@@ -0,0 +1,98 @@
/* Rotating-cube page transition (from termotohtori.fi). FRAGILE — do not
tweak. Injected by the backend as #pagerite-transition when the "cube"
transition is selected in the site settings. pagerite.js toggles
html.nav-back for history-back navigation and html.nav-fade for
same-section navigation. */
::view-transition {
perspective: 1000px;
inset: 0;
background: color-mix(var(--bg) 50%, black 50%);
}
::view-transition-group(root),
::view-transition-image-pair(root) {
transform-style: preserve-3d;
isolation: auto;
}
::view-transition-old(root),
::view-transition-new(root) {
mix-blend-mode: normal;
backface-visibility: hidden;
animation: none;
}
@keyframes group-rotate {
to {
transform: rotateY(-90deg);
}
}
@keyframes fade-out-a-bit {
to {
opacity: 0.5;
}
}
@keyframes fade-in-a-bit {
from {
opacity: 0.5;
}
}
::view-transition-group(root) {
transform-origin: 50% 50% -50vw;
animation: 300ms ease-in-out forwards group-rotate;
}
::view-transition-old(root) {
animation: 300ms ease-in-out forwards fade-out-a-bit;
}
::view-transition-new(root) {
transform-origin: 0 0;
transform: rotateY(90deg);
inset: 0 auto 0 100%;
animation: 300ms ease-in-out forwards fade-in-a-bit;
}
/* Reverse direction for browser back navigation (same geometry, mirrored). */
@keyframes group-rotate-back {
to {
transform: rotateY(90deg);
}
}
html.nav-back::view-transition-group(root) {
animation-name: group-rotate-back;
}
html.nav-back::view-transition-new(root) {
transform-origin: 100% 0;
transform: rotateY(-90deg);
inset: 0 100% 0 auto;
}
/* Same-section navigation: a plain crossfade instead of the cube. These
rules only override animation/geometry, leaving the block above's
perspective and layering untouched. The old snapshot stays fully opaque
underneath while the new one fades in on top — never a dip to black. */
@keyframes nav-fade-in {
from {
opacity: 0;
}
}
html.nav-fade::view-transition-group(root) {
animation: none;
}
html.nav-fade::view-transition-old(root) {
animation: none;
}
html.nav-fade::view-transition-new(root) {
transform: none;
inset: 0;
animation: 200ms ease-in-out nav-fade-in;
}
+13
View File
@@ -0,0 +1,13 @@
/* Eyes banner design: a canvas critter watching the cursor from the
grass (banner.html — markup + styles + script inlined by the backend
into #page-banner). The canvas fills the banner; the scene composes
against the 13rem layout box and extends flat grass into any overflow
a theme adds below (summer's cross-fade strip). */
/* Opt out of the base parallax (scale overscan + --pry drift): it lands on
the design wrapper the backend puts around banner.html, and scaling from
the bottom edge would crop the top of the composed scene — pushing the
critter out of view — while this canvas animates on its own. */
#page-banner>[data-design="eyes"] {
transform: none;
}
+482
View File
@@ -0,0 +1,482 @@
<canvas id="eyes"></canvas>
<style>
#eyes {
width: 100%;
/* 13rem — the banner's layout height, correct from the first frame,
before any external stylesheet has sized #page-banner. Themes that
extend the banner past the layout box (summer's overflow fade) raise
--eyes-h to 100% so the canvas follows the taller stage; the script
still composes the scene against the 13rem box and only extends the
meadow, so the scene itself never shifts. */
height: var(--eyes-h, 13rem);
display: block;
}
</style>
<script><!--
(() => {
const c = document.getElementById('eyes')
const ctx = c.getContext('2d')
// Sync the backing store to the canvas' laid-out size. Checked every
// frame: this inline script runs before the stylesheets that size
// #page-banner, so observers/load events can still miss the transition.
// Assigning width/height also clears the canvas. DPR is read here, not
// captured: it changes with browser zoom.
const syncSize = () => {
const DPR = devicePixelRatio || 1
const w = Math.round(Math.max(1, c.clientWidth) * DPR)
const h = Math.round(Math.max(1, c.clientHeight) * DPR)
if (c.width !== w || c.height !== h) {
c.width = w
c.height = h
}
ctx.setTransform(DPR, 0, 0, DPR, 0, 0)
}
let mx = 0
let my = 0
let lastMove = 0
// Mouse and touch tracked with separate listeners (pointer events arrive
// too late on some mobile browsers). Passive listeners: a drag on the
// banner still scrolls the page — on browsers that stop delivering
// touchmove once scrolling takes over, the gaze just follows until then.
const track = (x, y) => {
const r = c.getBoundingClientRect()
// Convert viewport coordinates into the canvas' CSS-pixel coordinate
// system. This remains correct with browser zoom, CSS transforms, etc.
mx = (x - r.left) * c.clientWidth / r.width
my = (y - r.top) * c.clientHeight / r.height
lastMove = performance.now()
}
addEventListener('mousemove', e => track(e.clientX, e.clientY))
const trackTouch = e => {
const t = e.touches[0]
if (t) track(t.clientX, t.clientY)
}
addEventListener('touchstart', trackTouch, { passive: true })
addEventListener('touchmove', trackTouch, { passive: true })
let gx = 0.5
let gy = 0.5
let tx = 0.5
let ty = 0.5
let yoff = 0
let vy = 0
let hidePhase = 0
let nextMove = 0
let nextHide = 4000 + Math.random() * 5000
let resurfaceAt = 0
const eyes = [
{ x: 0, y: 0, vx: 0, vy: 0, pr: 0.3 },
{ x: 0, y: 0, vx: 0, vy: 0, pr: 0.3 }
]
const ridgeY = (x, w, h) =>
h * 0.83 +
Math.sin(x * 0.012) * 10 +
Math.sin(x * 0.003 + 1.4) * 16 +
Math.sin(x * 0.02 + 0.7) * 3
const drawCloud = (x, y, s) => {
ctx.fillStyle = 'rgba(255,255,255,0.85)'
ctx.beginPath()
ctx.arc(x - s * 0.55, y + s * 0.05, s * 0.38, 0, 7)
ctx.arc(x - s * 0.12, y - s * 0.08, s * 0.48, 0, 7)
ctx.arc(x + s * 0.32, y, s * 0.42, 0, 7)
ctx.arc(x + s * 0.64, y + s * 0.1, s * 0.28, 0, 7)
ctx.fill()
}
const drawHills = (w, h) => {
ctx.fillStyle = '#b7d7a8'
ctx.beginPath()
ctx.moveTo(0, h)
ctx.lineTo(0, h * 0.63)
ctx.quadraticCurveTo(w * 0.18, h * 0.48, w * 0.35, h * 0.62)
ctx.quadraticCurveTo(w * 0.52, h * 0.78, w * 0.68, h * 0.58)
ctx.quadraticCurveTo(w * 0.82, h * 0.43, w, h * 0.57)
ctx.lineTo(w, h)
ctx.closePath()
ctx.fill()
ctx.fillStyle = '#99c685'
ctx.beginPath()
ctx.moveTo(0, h)
ctx.lineTo(0, h * 0.72)
ctx.quadraticCurveTo(w * 0.14, h * 0.6, w * 0.28, h * 0.7)
ctx.quadraticCurveTo(w * 0.46, h * 0.82, w * 0.62, h * 0.66)
ctx.quadraticCurveTo(w * 0.82, h * 0.5, w, h * 0.68)
ctx.lineTo(w, h)
ctx.closePath()
ctx.fill()
}
const drawBackground = (w, h) => {
const sky = ctx.createLinearGradient(0, 0, 0, h)
sky.addColorStop(0, '#8ed0ff')
sky.addColorStop(0.62, '#d9f1ff')
sky.addColorStop(1, '#eef9ff')
ctx.fillStyle = sky
ctx.fillRect(0, 0, w, h)
ctx.fillStyle = 'rgba(255,240,170,0.5)'
ctx.beginPath()
ctx.arc(w * 0.83, h * 0.2, h * 0.16, 0, 7)
ctx.fill()
drawCloud(w * 0.18, h * 0.2, h * 0.16)
drawCloud(w * 0.43, h * 0.14, h * 0.12)
drawCloud(w * 0.68, h * 0.24, h * 0.15)
drawHills(w, h)
for (let i = 0; i < 5; i++) {
const x = (i + 0.5) * w / 5
const y = h * 0.69 + Math.sin(i * 1.7) * 8
ctx.fillStyle = '#5f8d4e'
ctx.beginPath()
ctx.arc(x, y, 18, Math.PI, 0)
ctx.arc(x - 14, y + 2, 14, Math.PI, 0)
ctx.arc(x + 14, y + 3, 12, Math.PI, 0)
ctx.fill()
}
}
const drawCritter = (cx0, eyeY, R, now, dt) => {
const headR = R * 2
const headCx = cx0
const headCy = eyeY + R * 0.52
ctx.fillStyle = '#5fbe61'
ctx.strokeStyle = '#285838'
ctx.lineWidth = 3
ctx.beginPath()
ctx.moveTo(headCx - headR * 0.45, headCy - headR * 0.84)
ctx.quadraticCurveTo(
headCx - headR * 0.62,
headCy - headR * 1.18,
headCx - headR * 0.2,
headCy - headR * 0.94
)
ctx.fill()
ctx.stroke()
ctx.beginPath()
ctx.moveTo(headCx + headR * 0.45, headCy - headR * 0.84)
ctx.quadraticCurveTo(
headCx + headR * 0.62,
headCy - headR * 1.18,
headCx + headR * 0.2,
headCy - headR * 0.94
)
ctx.fill()
ctx.stroke()
ctx.beginPath()
ctx.arc(headCx, headCy, headR, 0, 7)
ctx.fill()
ctx.stroke()
ctx.fillStyle = 'rgba(255,255,255,0.12)'
ctx.beginPath()
ctx.arc(
headCx - headR * 0.28,
headCy - headR * 0.22,
headR * 0.4,
0,
7
)
ctx.fill()
ctx.fillStyle = '#4caa50'
for (let i = -1; i <= 1; i++) {
ctx.beginPath()
ctx.arc(
headCx + i * headR * 0.42,
headCy - headR * 0.16,
headR * 0.13,
0,
7
)
ctx.fill()
}
ctx.strokeStyle = '#285838'
ctx.lineCap = 'round'
for (let i = -1; i <= 1; i++) {
ctx.lineWidth = 4
ctx.beginPath()
ctx.moveTo(headCx + i * 10, headCy - headR * 0.94)
ctx.lineTo(
headCx + i * 16,
headCy - headR * 1.1 - Math.sin(now / 180 + i) * 3
)
ctx.stroke()
}
eyes.forEach((e, i) => {
const cx = cx0 + (i ? 1.3 : -1.3) * R
const watching = now - lastMove < 4000
let ptx
let pty
if (watching) {
const dx = mx - cx
const dy = my - eyeY
const d = Math.hypot(dx, dy) || 1
const maxReach = R * 0.43
const responseDistance = R * 4
// Direction points exactly at the cursor, while reach increases
// smoothly with cursor distance.
const reach = maxReach * Math.min(1, d / responseDistance)
ptx = dx / d * reach
pty = dy / d * reach
} else {
ptx = Math.sin(now / 900 + i * 2) * R * 0.3
pty = Math.cos(now / 1300 + i * 3) * R * 0.2
}
e.vx += (ptx - e.x) * 0.08 * dt
e.vy += (pty - e.y) * 0.08 * dt
e.vx *= Math.pow(0.82, dt)
e.vy *= Math.pow(0.82, dt)
e.x += e.vx * dt
e.y += e.vy * dt
const near = Math.hypot(mx - cx, my - eyeY) < R * 2.5
e.pr += ((near ? 0.42 : 0.3) - e.pr) * 0.1 * dt
ctx.fillStyle = '#fff'
ctx.strokeStyle = '#1f2d22'
ctx.lineWidth = 2.5
ctx.beginPath()
ctx.ellipse(cx, eyeY, R, R * 1.12, 0, 0, 7)
ctx.fill()
ctx.stroke()
ctx.save()
ctx.beginPath()
ctx.ellipse(cx, eyeY, R, R * 1.12, 0, 0, 7)
ctx.clip()
ctx.fillStyle = '#f2b84b'
ctx.beginPath()
ctx.arc(cx + e.x, eyeY + e.y, R * 0.56, 0, 7)
ctx.fill()
ctx.strokeStyle = 'rgba(140,84,8,0.45)'
ctx.lineWidth = 1
for (let a = 0; a < 12; a++) {
const ang = a / 12 * Math.PI * 2
ctx.beginPath()
ctx.moveTo(cx + e.x, eyeY + e.y)
ctx.lineTo(
cx + e.x + Math.cos(ang) * R * 0.5,
eyeY + e.y + Math.sin(ang) * R * 0.5
)
ctx.stroke()
}
ctx.fillStyle = '#191919'
ctx.beginPath()
ctx.arc(cx + e.x, eyeY + e.y, R * e.pr, 0, 7)
ctx.fill()
ctx.fillStyle = '#fff'
ctx.beginPath()
ctx.arc(
cx + e.x - R * 0.14,
eyeY + e.y - R * 0.17,
R * 0.09,
0,
7
)
ctx.fill()
ctx.restore()
ctx.strokeStyle = '#1f2d22'
ctx.lineWidth = 3
ctx.beginPath()
ctx.moveTo(cx - R * 0.7, eyeY - R * 1.2)
ctx.quadraticCurveTo(
cx,
eyeY - R * 1.48 - (i ? -1 : 1) * 2,
cx + R * 0.72,
eyeY - R * 1.12
)
ctx.stroke()
})
}
const drawForeground = (w, h) => {
ctx.fillStyle = '#69ae4b'
ctx.beginPath()
ctx.moveTo(0, h)
ctx.lineTo(0, ridgeY(0, w, h))
// Sample one step PAST the right edge (x <= w + 8): stopping at w would
// leave the path closing with a visible vertical drop at the edge
// whenever the width isn't a multiple of the 8px step.
for (let x = 0; x <= w + 8; x += 8)
ctx.lineTo(x, ridgeY(x, w, h))
ctx.lineTo(w, h)
ctx.closePath()
ctx.fill()
ctx.fillStyle = 'rgba(48,102,34,0.18)'
ctx.beginPath()
ctx.moveTo(0, h)
ctx.lineTo(0, ridgeY(0, w, h) + 10)
for (let x = 0; x <= w + 8; x += 8)
ctx.lineTo(x, ridgeY(x, w, h) + 10)
ctx.lineTo(w, h)
ctx.closePath()
ctx.fill()
ctx.strokeStyle = '#4d8d37'
ctx.lineWidth = 2
ctx.lineCap = 'round'
for (let x = 0; x <= w; x += 16) {
const y = ridgeY(x, w, h)
ctx.beginPath()
ctx.moveTo(x, y + 6)
ctx.quadraticCurveTo(x - 4, y - 10, x + 1, y - 2)
ctx.moveTo(x + 1, y + 6)
ctx.quadraticCurveTo(x + 4, y - 12, x + 3, y - 1)
ctx.stroke()
}
for (let i = 0; i < 8; i++) {
const x = (i + 0.4) * w / 8 + Math.sin(i * 2.4) * 10
const y = ridgeY(x, w, h) + 2
ctx.fillStyle = i % 2 ? '#ffdc6b' : '#ff8aa7'
ctx.beginPath()
ctx.arc(x, y, 3, 0, 7)
ctx.arc(x - 4, y + 2, 3, 0, 7)
ctx.arc(x + 4, y + 2, 3, 0, 7)
ctx.fill()
}
}
let prev = performance.now()
const frame = now => {
if (!c.isConnected) return
const dt = Math.min(now - prev, 100) / 16.7
prev = now
syncSize()
const w = c.clientWidth
// Compose the scene against the banner's layout box, not the canvas:
// a theme may extend #page-banner past #banner (e.g. summer overflows
// the artwork into the page for a masked cross-fade), and the critter
// must stay in the visible part. The overflow strip is filled with the
// flat meadow color below — a hard canvas edge would show through the
// fade, a grass extension just blends.
const h = Math.min(c.clientHeight,
c.closest('#banner')?.clientHeight || c.clientHeight)
const R = Math.min(h * 0.11, 38)
if (now > nextMove && !hidePhase) {
tx = 0.15 + Math.random() * 0.7
ty = 0.3 + Math.random() * 0.4
nextMove = now + 2500 + Math.random() * 3500
}
gx += (tx - gx) * 0.02 * dt
gy += (ty - gy) * 0.02 * dt
if (hidePhase === 0 && now > nextHide)
hidePhase = 1
if (hidePhase === 1 && yoff > h * 0.9) {
hidePhase = 2
resurfaceAt = now + 500 + Math.random() * 900
}
if (hidePhase === 2 && now > resurfaceAt) {
hidePhase = 0
nextHide = now + 5000 + Math.random() * 7000
tx = 0.15 + Math.random() * 0.7
ty = 0.3 + Math.random() * 0.4
gx = tx
gy = ty
nextMove = now + 3000 + Math.random() * 3000
}
const yTarget = hidePhase ? h : 0
vy += (yTarget - yoff) * 0.06 * dt
vy *= Math.pow(0.85, dt)
yoff += vy * dt
// Clip the scene to the composed area: the ducking critter travels
// below it, and the overflow strip is only flat meadow painted after —
// without the clip the critter would leave trails there as it sinks.
// Clipping against grass-on-grass is invisible, so the duck still
// reads as sinking into the meadow.
ctx.save()
ctx.beginPath()
ctx.rect(0, 0, w, h)
ctx.clip()
drawBackground(w, h)
const cx0 = gx * w
const ridge = ridgeY(cx0, w, h)
// Normally the full pair of eyes sits above the grass. gy gives it
// a small amount of bobbing/wandering without burying it again.
const eyeY =
ridge -
R * 1.25 +
(gy - 0.5) * R * 0.8 +
yoff
drawCritter(cx0, eyeY, R, now, dt)
drawForeground(w, h)
ctx.restore()
// Extend the meadow into any overflow below the composed scene, with
// the same two layers drawForeground leaves at the bottom (base grass
// plus the dark under-band) so the joint is invisible.
if (c.clientHeight > h) {
ctx.fillStyle = '#69ae4b'
ctx.fillRect(0, h, w, c.clientHeight - h)
ctx.fillStyle = 'rgba(48,102,34,0.18)'
ctx.fillRect(0, h, w, c.clientHeight - h)
}
requestAnimationFrame(frame)
}
requestAnimationFrame(frame)
})()
</script>
+73
View File
@@ -0,0 +1,73 @@
/* Nitro banner design: the bezier-swept artwork with wide orange stripes
(inlined by the backend into #page-banner), in neutral dark greys that
follow the page's color scheme. */
/* Banner artwork dark tones: neutral greys in light mode (retinted to the
page's violet family by the dark-scheme block below). */
.nb-base {
fill: #0b0b0d;
}
.nb-s1a {
stop-color: #242428;
}
.nb-s1b {
stop-color: #0b0b0d;
}
.nb-s2a {
stop-color: #19191d;
}
.nb-s2b {
stop-color: #060607;
}
.nb-c0 {
stop-color: #2a2a2f;
}
.nb-c1 {
stop-color: #131315;
}
.nb-c2 {
stop-color: #0b0b0d;
}
@media (prefers-color-scheme: dark) {
/* Banner dark tones tinted to the same violet family as the page. */
.nb-base {
fill: #100d18;
}
.nb-s1a {
stop-color: #292536;
}
.nb-s1b {
stop-color: #100d18;
}
.nb-s2a {
stop-color: #1e1a2b;
}
.nb-s2b {
stop-color: #090811;
}
.nb-c0 {
stop-color: #322d44;
}
.nb-c1 {
stop-color: #171422;
}
.nb-c2 {
stop-color: #100d18;
}
}
+1 -1
View File
@@ -1,4 +1,4 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1600 360" preserveAspectRatio="xMidYMid slice"> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1600 360" preserveAspectRatio="xMidYMax slice">
<defs> <defs>
<linearGradient id="flare" x1="0" y1="0" x2="1" y2="0"> <linearGradient id="flare" x1="0" y1="0" x2="1" y2="0">
<stop offset="0" stop-color="#ff6a00"/> <stop offset="0" stop-color="#ff6a00"/>

Before

Width:  |  Height:  |  Size: 2.5 KiB

After

Width:  |  Height:  |  Size: 2.5 KiB

@@ -41,6 +41,10 @@
--font-body: var(--font-montserrat); --font-body: var(--font-montserrat);
--font-heading: var(--font-literata); --font-heading: var(--font-literata);
--code-x-height: 0.517; /* Montserrat's x-height ratio */
/* Neutral grey selection instead of the accent tint: accent-colored
text (h2, links, markers) stays readable on it in both schemes. */
--selection-bg: #6664;
} }
/* Dark scheme: same identity, but the page goes deep violet (never muddy /* Dark scheme: same identity, but the page goes deep violet (never muddy
@@ -59,50 +63,19 @@
--code-bg: #12101b; --code-bg: #12101b;
/* code wells join the violet family */ /* code wells join the violet family */
} }
/* Banner dark tones tinted to the same violet family as the page. */
.nb-base {
fill: #100d18;
}
.nb-s1a {
stop-color: #292536;
}
.nb-s1b {
stop-color: #100d18;
}
.nb-s2a {
stop-color: #1e1a2b;
}
.nb-s2b {
stop-color: #090811;
}
.nb-c0 {
stop-color: #322d44;
}
.nb-c1 {
stop-color: #171422;
}
.nb-c2 {
stop-color: #100d18;
}
} }
::selection { /* Any banner used is separated from page by a thick orange line */
background: var(--accent); #banner {
color: var(--ink); border-bottom: 4px solid var(--accent);
} }
/* Oversized outlined brand, spilling off the banner edge: orange stroke, /* Oversized outlined brand, spilling off the banner edge: orange stroke,
solid black fill. */ solid black fill. */
#brand { #brand {
font-size: 10rem; /* Scales down proportionally below ~1000px: 10rem at a 62.5rem viewport,
shrinking with vmin (smaller of viewport width/height) below that. */
font-size: clamp(2.5rem, 16vmin, 10rem);
line-height: 1.2; line-height: 1.2;
font-weight: 700; font-weight: 700;
letter-spacing: 0.04em; letter-spacing: 0.04em;
@@ -112,47 +85,6 @@
text-shadow: 0 0 0.1em black; text-shadow: 0 0 0.1em black;
} }
/* Bezier-swept banner with wide orange stripes (inlined SVG), separated
from the page by a straight orange blade. */
#banner {
height: 13rem;
border-bottom: 4px solid var(--accent);
}
/* Banner artwork dark tones: neutral greys in light mode (retinted to the
page's violet family by the dark-scheme block above). */
.nb-base {
fill: #0b0b0d;
}
.nb-s1a {
stop-color: #242428;
}
.nb-s1b {
stop-color: #0b0b0d;
}
.nb-s2a {
stop-color: #19191d;
}
.nb-s2b {
stop-color: #060607;
}
.nb-c0 {
stop-color: #2a2a2f;
}
.nb-c1 {
stop-color: #131315;
}
.nb-c2 {
stop-color: #0b0b0d;
}
#nav { #nav {
font-family: var(--font-heading); font-family: var(--font-heading);
font-size: 0.95em; font-size: 0.95em;
@@ -205,21 +137,16 @@
/* Console-style headings: uppercase monospace. h1 in the page text color /* Console-style headings: uppercase monospace. h1 in the page text color
with a hazard-stripe underline, h2 deep orange, h3 cyan. */ with a hazard-stripe underline, h2 deep orange, h3 cyan. */
article h1, article h1 {
article h2,
article h3 {
text-transform: uppercase; text-transform: uppercase;
letter-spacing: 0.02em; letter-spacing: 0.02em;
}
article h1 {
color: var(--text); color: var(--text);
font-weight: 700; font-weight: 700;
padding-bottom: 0.5rem; padding-bottom: 0.5rem;
/* The hazard-stripe underline breaks out of the page box: the negative /* The hazard-stripe underline breaks out of the page box: the negative
right margin extends the h1's box (and thus its background) all the end margin extends the h1's box (and thus its background) all the
way to the viewport's right edge. */ way to the viewport's edge on that side. */
margin-right: calc((100% - 100vw) / 2); margin-inline-end: calc((100% - 100vw) / 2);
background: background:
linear-gradient(-55deg, linear-gradient(-55deg,
transparent 0 0.2rem, transparent 0 0.2rem,
@@ -255,32 +182,28 @@ article ul ul li::before {
} }
article ul ul ul li::before { article ul ul ul li::before {
content: "»"; color: var(--muted);
color: var(--accent);
} }
blockquote { blockquote {
border-left-color: var(--accent2); border-inline-start-color: var(--accent2);
background: color-mix(in oklab, var(--accent2) 6%, transparent); background: color-mix(var(--accent2) 6%, transparent);
padding: 0.25rem 0.75rem; padding: 0.25rem 0.75rem;
/* Keep the quoted text on the paragraph edge: the tinted box extends /* Keep the quoted text on the paragraph edge: the tinted box extends
past it by its own border/padding, like code blocks. */ past it by its own border/padding, like code blocks. */
margin: 0 -0.75rem 1rem -1rem; margin: 0 0 1rem;
margin-inline: -1rem -0.75rem;
} }
/* Code follows the color scheme; the dark-scheme well joins the violet /* Code follows the color scheme; the dark-scheme well joins the violet
family (--code-bg above). The orange side bar stays in both. */ family (--code-bg above). The orange side bar stays in both. */
pre { pre {
border-left: 0.25rem solid var(--accent); border-inline-start: 0.25rem solid var(--accent);
/* Text on the paragraph edge: the box extends by padding + border. */ /* Text on the paragraph edge: the box extends by padding + border. */
margin-left: calc(-0.8rem - 0.25rem); margin-inline-start: calc(-0.8rem - 0.25rem);
border-radius: 3px; border-radius: 3px;
} }
img { img {
border-radius: 3px; border-radius: 3px;
} }
::view-transition {
background: var(--bg);
}
+17
View File
@@ -0,0 +1,17 @@
/* Purple banner design: the sunrise artwork (inlined by the backend into
#page-banner) with parallax sun and a fade into the page background. */
/* Sunrise parallax: the sun and its glow rise faster than the artwork
drift (pagerite.js sets --pry on <html>), so scrolling the page makes
the sun come up. */
#page-banner .sun,
#page-banner .sun-glow {
transform-box: fill-box;
transform: translateY(calc(var(--pry, 0px) * -2));
}
/* The banner artwork fades into the page background at its bottom edge
(baked into the SVG, so a user banner replaces it cleanly). */
.banner-fade {
stop-color: var(--bg);
}
+1 -1
View File
@@ -1,4 +1,4 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1200 300" preserveAspectRatio="xMidYMid slice"> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1200 300" preserveAspectRatio="xMidYMax slice">
<defs> <defs>
<linearGradient id="sky" x1="0" y1="0" x2="0" y2="1"> <linearGradient id="sky" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#2b1b4d"/> <stop offset="0" stop-color="#2b1b4d"/>

Before

Width:  |  Height:  |  Size: 1.9 KiB

After

Width:  |  Height:  |  Size: 1.9 KiB

@@ -17,18 +17,16 @@
--line: #ffffff1c; --line: #ffffff1c;
--font-body: var(--font-literata); --font-body: var(--font-literata);
--font-heading: var(--font-fraunces); --font-heading: var(--font-fraunces);
} --code-x-height: 0.507; /* Literata's x-height ratio */
::selection {
background: var(--accent2);
color: #fff;
} }
/* Oversized tilted brand in the sky→violet gradient. */ /* Oversized tilted brand in the sky→violet gradient. */
#brand { #brand {
font-size: clamp(3.2rem, 9vw, 7.5rem); font-size: clamp(3.2rem, 9vw, 7.5rem);
line-height: 1; line-height: 1;
margin-bottom: -0.28em; /* Ensure below baseline stays visible */
padding-bottom: 0.3em;
margin-bottom: -0.3em;
transform: rotate(-2deg); transform: rotate(-2deg);
transform-origin: left bottom; transform-origin: left bottom;
background: linear-gradient(90deg, var(--accent), var(--accent2)); background: linear-gradient(90deg, var(--accent), var(--accent2));
@@ -39,25 +37,6 @@
filter: drop-shadow(0 0.15rem 0.6rem #9b6bff55); filter: drop-shadow(0 0.15rem 0.6rem #9b6bff55);
} }
/* Sunrise parallax: the sun and its glow rise faster than the artwork
drift (pagerite.js sets --pry on <html>), so scrolling the page makes
the sun come up. */
#page-banner .sun,
#page-banner .sun-glow {
transform-box: fill-box;
transform: translateY(calc(var(--pry, 0px) * -2));
}
/* The banner artwork fades into the page background at its bottom edge
(baked into the SVG, so a user banner replaces it cleanly). */
.banner-fade {
stop-color: var(--bg);
}
#banner {
min-height: 13rem;
}
/* Dark artwork: keep the nav readable with a shadow. */ /* Dark artwork: keep the nav readable with a shadow. */
#nav { #nav {
text-shadow: 0 0 0.15em black; text-shadow: 0 0 0.15em black;
@@ -81,12 +60,14 @@ article h3 {
} }
/* Theme-colored diamond markers instead of the base emoji (blue/orange /* Theme-colored diamond markers instead of the base emoji (blue/orange
clashes with this palette). */ clashes with this palette). The text-style runs heavy at full size,
so it's shrunk with font-size, the box is widened to compensate. */
article ul li::before { article ul li::before {
content: "◆"; content: "◆";
color: var(--accent); color: var(--accent);
font-size: 0.7em; font-size: 0.8em;
vertical-align: 0.15em; margin-inline-start: calc(-1 * var(--list-indent) / 0.8);
width: calc(var(--list-indent) / 0.8);
} }
article ul ul li::before { article ul ul li::before {
@@ -95,12 +76,11 @@ article ul ul li::before {
} }
article ul ul ul li::before { article ul ul ul li::before {
content: "◆";
color: var(--accent3); color: var(--accent3);
} }
blockquote { blockquote {
border-left-color: var(--accent2); border-inline-start-color: var(--accent2);
} }
/* Code panels sit slightly lighter than the page; the token colors come /* Code panels sit slightly lighter than the page; the token colors come
@@ -108,7 +88,3 @@ blockquote {
pre { pre {
--code-bg: var(--surface); --code-bg: var(--surface);
} }
::view-transition {
background: #000;
}
+52
View File
@@ -0,0 +1,52 @@
/* Wipe-reveal page transition: the old page stays put while the new one is
revealed on top of it by a clip-path wipe sweeping left to right.
Injected by the backend as #pagerite-transition when the "reveal"
transition is selected in the site settings. Mirrored on history-back
(html.nav-back: the wipe sweeps right to left); same-section navigation
crossfades (html.nav-fade). */
::view-transition-old(root),
::view-transition-new(root) {
mix-blend-mode: normal;
animation: none;
}
@keyframes reveal-right {
from {
clip-path: inset(0 100% 0 0);
}
to {
clip-path: inset(0);
}
}
::view-transition-new(root) {
clip-path: inset(0);
animation: 350ms ease-in-out reveal-right;
}
/* Reverse direction for browser back navigation. */
@keyframes reveal-left {
from {
clip-path: inset(0 0 0 100%);
}
to {
clip-path: inset(0);
}
}
html.nav-back::view-transition-new(root) {
animation-name: reveal-left;
}
/* Same-section navigation: a plain crossfade instead of the wipe. The old
snapshot stays fully opaque underneath while the new one fades in on
top — never a dip to black. */
@keyframes nav-fade-in {
from {
opacity: 0;
}
}
html.nav-fade::view-transition-new(root) {
animation: 200ms ease-in-out nav-fade-in;
}
+74
View File
@@ -0,0 +1,74 @@
/* Horizontal slide page transition: the old page slides left while the new
one follows from the right, both moving together like pages side by side.
Injected by the backend as #pagerite-transition when the "slide"
transition is selected in the site settings. Mirrored on history-back
(html.nav-back); same-section navigation crossfades (html.nav-fade). Both
snapshots cover the viewport throughout, so no background shows between
them. */
::view-transition {
background: var(--bg);
}
::view-transition-old(root),
::view-transition-new(root) {
mix-blend-mode: normal;
animation: none;
}
@keyframes slide-out-left {
to {
transform: translateX(-100%);
}
}
@keyframes slide-in-right {
from {
transform: translateX(100%);
}
}
::view-transition-old(root) {
animation: 300ms ease-in-out forwards slide-out-left;
}
::view-transition-new(root) {
animation: 300ms ease-in-out slide-in-right;
}
/* Reverse direction for browser back navigation. */
@keyframes slide-out-right {
to {
transform: translateX(100%);
}
}
@keyframes slide-in-left {
from {
transform: translateX(-100%);
}
}
html.nav-back::view-transition-old(root) {
animation-name: slide-out-right;
}
html.nav-back::view-transition-new(root) {
animation-name: slide-in-left;
}
/* Same-section navigation: a plain crossfade instead of the slide. The old
snapshot stays fully opaque underneath while the new one fades in on
top — never a dip to black. */
@keyframes nav-fade-in {
from {
opacity: 0;
}
}
html.nav-fade::view-transition-old(root) {
animation: none;
}
html.nav-fade::view-transition-new(root) {
animation: 200ms ease-in-out nav-fade-in;
}
+7
View File
@@ -0,0 +1,7 @@
/* Stars banner design: a drifting starfield (banner.html — canvas + script
inlined by the backend into #page-banner). The canvas takes the banner's
13rem layout height directly so it never renders unclipped before the
main stylesheet loads; the starfield scales to any height. The design's
opt-out of theme banner overflow/fade effects also lives in banner.html's
inline <style>, so it applies atomically with the markup (this file would
load a beat later and let the overflow flash through mid-transition). */
+73
View File
@@ -0,0 +1,73 @@
<canvas id="stars"></canvas>
<style>
#stars {
width: 100%;
/* 13rem — the banner's layout height, not 100%: a percentage only
resolves after the main stylesheet sizes #page-banner, and until
then the canvas would render at its intrinsic height, unclipped,
over the page. (This design always opts out of theme overflows —
below — so the layout height is always the right one.) */
height: 13rem;
display: block;
}
/* The night sky stays a windowed stage: undo the summer theme's banner
overflow/cross-fade — a starfield must not bleed into a daylit page.
Kept in this inlined <style> (not banner.css) so the opt-out applies
atomically with the markup; a separate stylesheet can arrive a beat
later and let the overflow flash through mid-transition. Later in
document order than theme.css, so same-specificity rules win. */
#page-banner {
inset: 0;
mask-image: none;
}
</style>
<script><!--
(() => {
const c = document.getElementById('stars')
const ctx = c.getContext('2d')
// Sync the backing store to the canvas' laid-out size. Checked every
// frame: this inline script runs before the stylesheets that size
// #page-banner, so observers/load events can still miss the transition.
// Assigning width/height also clears the canvas. DPR is read here, not
// captured: it changes with browser zoom.
const syncSize = () => {
const DPR = devicePixelRatio || 1
const w = Math.round(Math.max(1, c.clientWidth) * DPR)
const h = Math.round(Math.max(1, c.clientHeight) * DPR)
if (c.width !== w || c.height !== h) {
c.width = w
c.height = h
}
ctx.setTransform(DPR, 0, 0, DPR, 0, 0)
}
const stars = Array.from({ length: 110 }, () => ({
x: Math.random(),
y: Math.random(),
r: Math.random() * 1.4 + 0.3,
v: Math.random() * 0.05 + 0.01
}))
let prev = performance.now()
;(function frame(now) {
if (!c.isConnected) return
syncSize()
const w = c.clientWidth
const h = c.clientHeight
const dt = Math.min(now - prev, 100)
prev = now
ctx.fillStyle = '#0b0e1d'
ctx.fillRect(0, 0, w, h)
ctx.fillStyle = '#cdd6ff'
for (const s of stars) {
s.x = (s.x + (s.v * dt) / 1000) % 1
ctx.beginPath()
ctx.arc(s.x * w, s.y * h, s.r, 0, 7)
ctx.fill()
}
requestAnimationFrame(frame)
})(prev)
})()
</script>
+100
View File
@@ -0,0 +1,100 @@
/* Summer banner design: a bright illustrated meadow (inlined by the
backend into #page-banner) — rolling hills, leafy bushes, swaying
flowers, drifting clouds and a sun that rises as you scroll. */
/* The artwork fades into the page background at its bottom edge. */
.summer-fade {
stop-color: var(--bg);
}
/* Layered parallax on top of the whole-artwork drift (pagerite.js sets
--pry on <html>): the sun rises fastest, the clouds lag behind, and the
nearer a hill the less it moves — a cheap depth cue while scrolling. */
#page-banner .sun,
#page-banner .sun-glow {
transform-box: fill-box;
transform: translateY(calc(var(--pry, 0px) * -1.5));
}
#page-banner .clouds {
transform: translateY(calc(var(--pry, 0px) * -0.7));
}
#page-banner .hills-far {
transform: translateY(calc(var(--pry, 0px) * -0.3));
}
#page-banner .hills-near {
transform: translateY(calc(var(--pry, 0px) * -0.15));
}
#page-banner .meadow {
transform: translateY(calc(var(--pry, 0px) * -0.05));
}
/* Idle animations: the sun breathes, clouds float sideways, flowers sway.
Each animated group is a transform-less wrapper so the CSS transform
never clobbers the placement transforms inside. */
@media (prefers-reduced-motion: no-preference) {
#page-banner .sun-glow {
animation: summer-sun 9s ease-in-out infinite alternate;
}
#page-banner .cloud-a {
animation: summer-drift 56s ease-in-out infinite alternate;
}
#page-banner .cloud-b {
animation: summer-drift 73s ease-in-out infinite alternate-reverse;
}
#page-banner .cloud-c {
animation: summer-drift 64s ease-in-out infinite alternate;
}
#page-banner .flower {
transform-box: fill-box;
transform-origin: 50% 100%;
animation: summer-sway 5s ease-in-out infinite alternate;
}
/* Stagger the sway so the meadow doesn't move in lockstep. */
#page-banner .flower:nth-child(3n) {
animation-delay: -1.7s;
}
#page-banner .flower:nth-child(3n + 1) {
animation-delay: -3.1s;
animation-duration: 6s;
}
}
@keyframes summer-sun {
from {
opacity: 0.22;
}
to {
opacity: 0.38;
}
}
@keyframes summer-drift {
from {
transform: translateX(-1.6%);
}
to {
transform: translateX(1.6%);
}
}
@keyframes summer-sway {
from {
transform: rotate(-2.6deg);
}
to {
transform: rotate(2.6deg);
}
}
+369
View File
@@ -0,0 +1,369 @@
<svg xmlns="http://www.w3.org/2000/svg"
viewBox="0 0 1600 320"
preserveAspectRatio="xMidYMax slice">
<defs>
<linearGradient id="sky" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#8ed0ff"/>
<stop offset=".64" stop-color="#d9f1ff"/>
<stop offset="1" stop-color="#eef9ff"/>
</linearGradient>
<linearGradient id="farHill" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#b8dca7"/>
<stop offset="1" stop-color="#a4cf90"/>
</linearGradient>
<linearGradient id="nearHill" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#91c97c"/>
<stop offset="1" stop-color="#74b65e"/>
</linearGradient>
<linearGradient id="grass" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#69ae4b"/>
<stop offset="1" stop-color="#4f913b"/>
</linearGradient>
<!-- Soft fade into the page background (stop-color from banner.css) -->
<linearGradient id="meadowFade" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" class="summer-fade" stop-color="#e6f4cf" stop-opacity="0"/>
<stop offset="1" class="summer-fade" stop-color="#e6f4cf"/>
</linearGradient>
<filter id="soft">
<feGaussianBlur stdDeviation="16"/>
</filter>
<filter id="cloudShadow" x="-20%" y="-40%" width="140%" height="180%">
<feDropShadow dx="0" dy="3" stdDeviation="4"
flood-color="#4f8fb0" flood-opacity=".12"/>
</filter>
</defs>
<!-- Sky -->
<rect width="1600" height="320" fill="url(#sky)"/>
<!-- Sun (rises on scroll, glow pulses) -->
<circle class="sun-glow" cx="1330" cy="72" r="82"
fill="#fff0a2"
opacity=".28"
filter="url(#soft)"/>
<circle class="sun" cx="1330" cy="72" r="42"
fill="#fff0a2"
opacity=".82"/>
<!-- Clouds (drift on scroll, lazily float sideways) -->
<g class="clouds" fill="#fff" opacity=".82" filter="url(#cloudShadow)">
<g class="cloud cloud-a">
<g transform="translate(160 72)">
<ellipse cx="0" cy="12" rx="54" ry="20"/>
<circle cx="-27" cy="4" r="26"/>
<circle cx="7" cy="-6" r="34"/>
<circle cx="40" cy="7" r="25"/>
</g>
</g>
<g class="cloud cloud-b" opacity=".72">
<g transform="translate(650 52)">
<ellipse cx="0" cy="10" rx="45" ry="16"/>
<circle cx="-21" cy="4" r="21"/>
<circle cx="7" cy="-5" r="27"/>
<circle cx="32" cy="6" r="20"/>
</g>
</g>
<g class="cloud cloud-c" opacity=".58">
<g transform="translate(1070 110)">
<ellipse cx="0" cy="8" rx="38" ry="14"/>
<circle cx="-18" cy="3" r="18"/>
<circle cx="6" cy="-4" r="23"/>
<circle cx="28" cy="5" r="17"/>
</g>
</g>
</g>
<!-- Far rolling hills (slow parallax); drawn past the bottom edge so the
drift never exposes sky underneath -->
<path class="hills-far" d="
M0 211
C145 170 263 173 390 211
C516 249 616 237 730 194
C846 151 945 168 1065 208
C1195 251 1334 214 1600 171
L1600 340
L0 340
Z"
fill="url(#farHill)"/>
<!-- Near rolling hills -->
<path class="hills-near" d="
M0 248
C120 217 228 217 342 247
C454 278 566 269 674 228
C789 184 901 202 1024 243
C1143 282 1278 271 1391 229
C1473 199 1538 196 1600 207
L1600 340
L0 340
Z"
fill="url(#nearHill)"/>
<!-- Proper leafy bushes (ride with the near hills) -->
<g class="hills-near">
<g transform="translate(130 236)">
<path d="
M-58 28
C-61 7 -46 -5 -28 0
C-28 -20 -4 -31 11 -16
C22 -36 54 -27 53 -5
C73 -10 90 6 85 27
Z"
fill="#5e9c4b"/>
<path d="
M-47 27
C-48 10 -37 1 -23 5
C-20 -10 -4 -17 8 -7
C19 -21 40 -17 43 -1
C60 -4 70 8 68 26
Z"
fill="#70ad58"/>
</g>
<g transform="translate(1240 225) scale(.9)">
<path d="
M-58 28
C-61 7 -46 -5 -28 0
C-28 -20 -4 -31 11 -16
C22 -36 54 -27 53 -5
C73 -10 90 6 85 27
Z"
fill="#5e9c4b"/>
<path d="
M-47 27
C-48 10 -37 1 -23 5
C-20 -10 -4 -17 8 -7
C19 -21 40 -17 43 -1
C60 -4 70 8 68 26
Z"
fill="#70ad58"/>
</g>
<g transform="translate(460 250) scale(.58)">
<path d="
M-58 28
C-61 7 -46 -5 -28 0
C-28 -20 -4 -31 11 -16
C22 -36 54 -27 53 -5
C73 -10 90 6 85 27
Z"
fill="#5a9748"/>
</g>
<g transform="translate(940 245) scale(.55)">
<path d="
M-58 28
C-61 7 -46 -5 -28 0
C-28 -20 -4 -31 11 -16
C22 -36 54 -27 53 -5
C73 -10 90 6 85 27
Z"
fill="#5a9748"/>
</g>
</g>
<!-- Foreground meadow (near-static parallax) -->
<path class="meadow" d="
M0 270
C95 260 175 267 260 273
C347 279 431 265 521 262
C620 258 705 275 808 272
C914 269 1004 251 1115 258
C1229 265 1324 279 1424 270
C1491 264 1547 258 1600 262
L1600 340
L0 340
Z"
fill="url(#grass)"/>
<!-- Dense grass blades -->
<g class="meadow" fill="none"
stroke="#4b8937"
stroke-width="2.5"
stroke-linecap="round">
<path d="M20 276 q-2 -19 4 -31 M28 276 q8 -18 11 -29"/>
<path d="M65 272 q-5 -18 1 -28 M74 272 q8 -16 11 -25"/>
<path d="M110 271 q-2 -23 5 -36 M121 272 q9 -18 12 -30"/>
<path d="M165 273 q-6 -16 -1 -28 M175 273 q7 -18 11 -27"/>
<path d="M225 274 q-3 -24 4 -37 M236 274 q8 -17 12 -29"/>
<path d="M286 273 q-6 -19 -1 -31 M296 274 q9 -18 12 -28"/>
<path d="M345 270 q-3 -18 3 -31 M355 270 q8 -17 12 -28"/>
<path d="M405 266 q-4 -22 2 -35 M416 266 q9 -17 13 -28"/>
<path d="M472 263 q-3 -18 3 -30 M482 263 q8 -19 13 -29"/>
<path d="M535 263 q-6 -19 -1 -31 M546 264 q9 -19 13 -30"/>
<path d="M598 266 q-3 -24 4 -37 M609 267 q8 -19 12 -31"/>
<path d="M665 271 q-6 -17 -1 -29 M676 271 q9 -17 13 -28"/>
<path d="M735 273 q-4 -22 3 -34 M746 273 q8 -18 12 -29"/>
<path d="M805 272 q-6 -18 -1 -30 M816 271 q9 -17 13 -27"/>
<path d="M875 265 q-3 -23 4 -35 M886 265 q8 -18 12 -28"/>
<path d="M945 258 q-5 -18 1 -30 M955 258 q9 -17 12 -28"/>
<path d="M1018 256 q-3 -22 4 -34 M1029 256 q8 -18 12 -28"/>
<path d="M1090 259 q-5 -20 1 -31 M1101 259 q8 -18 12 -29"/>
<path d="M1162 264 q-3 -24 4 -36 M1173 264 q8 -18 12 -30"/>
<path d="M1234 269 q-5 -19 1 -31 M1245 269 q8 -17 12 -28"/>
<path d="M1304 273 q-3 -23 4 -35 M1315 273 q8 -18 12 -29"/>
<path d="M1374 272 q-5 -18 1 -29 M1385 272 q8 -18 12 -28"/>
<path d="M1445 268 q-3 -22 4 -34 M1456 268 q8 -17 12 -28"/>
<path d="M1515 264 q-5 -20 1 -31 M1526 264 q8 -18 12 -29"/>
<path d="M1570 263 q-3 -21 4 -33 M1581 263 q8 -17 12 -28"/>
</g>
<!-- Extra foreground tufts -->
<g class="meadow" fill="#579a40">
<path d="M80 320 l15 -45 3 45 14 -57 0 57 20 -38 -7 38z"/>
<path d="M365 320 l10 -37 5 37 15 -49 -2 49 20 -34 -7 34z"/>
<path d="M710 320 l13 -42 3 42 15 -54 -1 54 20 -38 -8 38z"/>
<path d="M1040 320 l10 -35 5 35 14 -49 0 49 18 -32 -6 32z"/>
<path d="M1360 320 l13 -43 3 43 16 -55 -2 55 20 -37 -7 37z"/>
</g>
<!-- Flowers (sway gently; each wrapped in a classed group so the CSS
transform doesn't clobber the placement transform) -->
<g class="meadow">
<g class="flower">
<g transform="translate(190 270)">
<path d="M0 8 Q-2 25 -4 42" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#ff8aa7">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ffe071"/>
</g>
</g>
<g class="flower">
<g transform="translate(310 280) scale(.8)">
<path d="M0 8 Q2 23 1 37" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#fff0a0">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ef7191"/>
</g>
</g>
<g class="flower">
<g transform="translate(540 265)">
<path d="M0 8 Q-2 25 -2 42" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#ff8aa7">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ffe071"/>
</g>
</g>
<g class="flower">
<g transform="translate(670 279) scale(.7)">
<g fill="#fff0a0">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ef7191"/>
</g>
</g>
<g class="flower">
<g transform="translate(845 272)">
<path d="M0 8 Q2 25 0 42" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#ff8aa7">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ffe071"/>
</g>
</g>
<g class="flower">
<g transform="translate(1010 265) scale(.85)">
<path d="M0 8 Q-2 24 -1 38" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#fff0a0">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ef7191"/>
</g>
</g>
<g class="flower">
<g transform="translate(1190 275)">
<g fill="#ff8aa7">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ffe071"/>
</g>
</g>
<g class="flower">
<g transform="translate(1410 269)">
<path d="M0 8 Q2 26 0 42" stroke="#42833a" stroke-width="2" fill="none"/>
<g fill="#fff0a0">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ef7191"/>
</g>
</g>
<g class="flower">
<g transform="translate(1510 282) scale(.72)">
<g fill="#ff8aa7">
<circle cx="-5" cy="0" r="5"/>
<circle cx="5" cy="0" r="5"/>
<circle cx="0" cy="-5" r="5"/>
<circle cx="0" cy="5" r="5"/>
</g>
<circle r="3" fill="#ffe071"/>
</g>
</g>
</g>
<!-- Small scattered meadow dots -->
<g class="meadow" fill="#ffe071" opacity=".9">
<circle cx="250" cy="287" r="2.5"/>
<circle cx="445" cy="281" r="2.5"/>
<circle cx="615" cy="289" r="2.5"/>
<circle cx="760" cy="283" r="2.5"/>
<circle cx="930" cy="279" r="2.5"/>
<circle cx="1125" cy="287" r="2.5"/>
<circle cx="1290" cy="281" r="2.5"/>
</g>
<g class="meadow" fill="#ff8aa7" opacity=".8">
<circle cx="120" cy="291" r="2.5"/>
<circle cx="395" cy="291" r="2.5"/>
<circle cx="730" cy="294" r="2.5"/>
<circle cx="1080" cy="294" r="2.5"/>
<circle cx="1460" cy="291" r="2.5"/>
</g>
<!-- Soft fade into the page background -->
<rect y="252" width="1600" height="68" fill="url(#meadowFade)"/>
</svg>

After

Width:  |  Height:  |  Size: 11 KiB

+143
View File
@@ -0,0 +1,143 @@
<svg xmlns="http://www.w3.org/2000/svg" width="300" height="9" viewBox="0 0 300 9">
<g fill="none" stroke-linecap="round">
<path d="M1.0 9 q2.0 -3.0 1.5 -5.5" stroke="#579a40" stroke-width="1.0"/>
<path d="M3.5 9 q1.7 -4.1 3.0 -7.5" stroke="#579a40" stroke-width="1.6"/>
<path d="M6.3 9 q-2.0 -3.8 -1.8 -6.9" stroke="#63ad4d" stroke-width="0.8"/>
<path d="M8.4 9 q-1.7 -2.3 -3.1 -4.2" stroke="#74b65e" stroke-width="1.2"/>
<path d="M11.0 9 q0.6 -3.5 1.8 -6.3" stroke="#4f913b" stroke-width="1.3"/>
<path d="M12.7 9 q-0.4 -2.6 -1.9 -4.8" stroke="#63ad4d" stroke-width="0.8"/>
<path d="M14.3 9 q-2.3 -2.9 2.5 -5.3" stroke="#4f913b" stroke-width="1.4"/>
<path d="M16.9 9 q1.8 -4.5 -2.9 -8.1" stroke="#74b65e" stroke-width="1.8"/>
<path d="M19.5 9 q0.5 -2.6 2.9 -4.7" stroke="#63ad4d" stroke-width="1.4"/>
<path d="M22.5 9 q2.3 -2.1 -1.7 -3.9" stroke="#579a40" stroke-width="1.7"/>
<path d="M25.2 9 q-2.1 -4.4 -3.8 -7.9" stroke="#42833a" stroke-width="1.0"/>
<path d="M27.7 9 q-1.9 -4.5 -3.8 -8.2" stroke="#74b65e" stroke-width="1.4"/>
<path d="M29.4 9 q-0.3 -3.7 -3.1 -6.7" stroke="#579a40" stroke-width="1.4"/>
<path d="M31.0 9 q-0.9 -3.5 0.1 -6.3" stroke="#579a40" stroke-width="0.9"/>
<path d="M32.6 9 q-0.9 -4.3 1.1 -7.8" stroke="#4f913b" stroke-width="1.4"/>
<path d="M35.7 9 q-1.4 -2.4 2.3 -4.3" stroke="#63ad4d" stroke-width="1.5"/>
<path d="M36.9 9 q-0.4 -3.2 0.0 -5.8" stroke="#63ad4d" stroke-width="1.6"/>
<path d="M38.5 9 q1.1 -4.6 0.6 -8.4" stroke="#63ad4d" stroke-width="1.2"/>
<path d="M41.3 9 q2.3 -2.9 4.0 -5.4" stroke="#63ad4d" stroke-width="1.8"/>
<path d="M43.3 9 q1.6 -2.6 4.5 -4.7" stroke="#4f913b" stroke-width="1.0"/>
<path d="M45.4 9 q-0.3 -2.6 -2.1 -4.8" stroke="#63ad4d" stroke-width="0.8"/>
<path d="M47.0 9 q0.0 -4.3 -3.8 -7.8" stroke="#42833a" stroke-width="1.6"/>
<path d="M49.6 9 q1.6 -3.1 0.3 -5.7" stroke="#4f913b" stroke-width="1.4"/>
<path d="M51.6 9 q-1.2 -2.7 -0.2 -4.9" stroke="#74b65e" stroke-width="1.2"/>
<path d="M53.7 9 q-1.0 -2.4 3.6 -4.3" stroke="#579a40" stroke-width="1.6"/>
<path d="M55.7 9 q1.4 -2.1 3.5 -3.9" stroke="#579a40" stroke-width="1.2"/>
<path d="M58.4 9 q-0.6 -2.6 0.7 -4.8" stroke="#579a40" stroke-width="1.2"/>
<path d="M60.9 9 q-2.3 -2.4 -3.7 -4.4" stroke="#579a40" stroke-width="1.0"/>
<path d="M62.7 9 q1.0 -2.3 -1.3 -4.1" stroke="#579a40" stroke-width="0.8"/>
<path d="M64.9 9 q-0.3 -4.0 -3.5 -7.2" stroke="#42833a" stroke-width="1.3"/>
<path d="M67.6 9 q1.7 -4.2 4.2 -7.6" stroke="#579a40" stroke-width="0.9"/>
<path d="M70.5 9 q-0.9 -2.3 1.3 -4.1" stroke="#42833a" stroke-width="1.7"/>
<path d="M72.8 9 q-0.2 -2.9 -2.9 -5.2" stroke="#42833a" stroke-width="1.0"/>
<path d="M74.3 9 q-0.5 -2.9 4.5 -5.2" stroke="#42833a" stroke-width="1.4"/>
<path d="M76.9 9 q-1.6 -2.3 0.5 -4.1" stroke="#42833a" stroke-width="1.7"/>
<path d="M78.4 9 q1.1 -5.1 -0.4 -9.2" stroke="#579a40" stroke-width="1.0"/>
<path d="M80.6 9 q-1.2 -3.8 2.1 -6.9" stroke="#42833a" stroke-width="1.8"/>
<path d="M81.9 9 q-1.7 -2.7 -2.8 -5.0" stroke="#42833a" stroke-width="1.0"/>
<path d="M83.8 9 q1.8 -3.1 4.5 -5.7" stroke="#74b65e" stroke-width="0.8"/>
<path d="M85.8 9 q-0.7 -4.0 -0.4 -7.3" stroke="#42833a" stroke-width="0.8"/>
<path d="M88.0 9 q-1.6 -5.0 -0.1 -9.2" stroke="#63ad4d" stroke-width="0.8"/>
<path d="M90.9 9 q1.5 -4.4 2.2 -7.9" stroke="#74b65e" stroke-width="1.5"/>
<path d="M92.3 9 q1.5 -3.3 1.5 -6.0" stroke="#579a40" stroke-width="1.2"/>
<path d="M94.7 9 q-0.4 -3.4 0.2 -6.2" stroke="#63ad4d" stroke-width="1.0"/>
<path d="M96.4 9 q-2.1 -4.8 -1.6 -8.7" stroke="#4f913b" stroke-width="1.7"/>
<path d="M99.3 9 q-0.1 -4.9 0.3 -8.9" stroke="#4f913b" stroke-width="1.7"/>
<path d="M101.1 9 q1.0 -4.6 1.4 -8.4" stroke="#4f913b" stroke-width="0.8"/>
<path d="M103.6 9 q-1.6 -4.8 -0.8 -8.8" stroke="#4f913b" stroke-width="1.7"/>
<path d="M105.1 9 q-2.5 -2.1 1.7 -3.9" stroke="#74b65e" stroke-width="1.5"/>
<path d="M108.0 9 q1.2 -2.2 -3.6 -3.9" stroke="#74b65e" stroke-width="1.8"/>
<path d="M110.5 9 q1.0 -4.2 2.3 -7.6" stroke="#579a40" stroke-width="1.6"/>
<path d="M113.6 9 q0.2 -2.4 2.7 -4.4" stroke="#42833a" stroke-width="1.5"/>
<path d="M115.8 9 q-0.7 -4.6 -1.1 -8.3" stroke="#579a40" stroke-width="1.0"/>
<path d="M117.8 9 q-2.0 -3.2 3.7 -5.8" stroke="#579a40" stroke-width="1.1"/>
<path d="M119.2 9 q-1.0 -4.8 1.2 -8.7" stroke="#4f913b" stroke-width="1.6"/>
<path d="M121.9 9 q-0.2 -4.7 -3.9 -8.5" stroke="#4f913b" stroke-width="1.7"/>
<path d="M125.0 9 q0.4 -2.2 -0.7 -3.9" stroke="#42833a" stroke-width="1.6"/>
<path d="M127.3 9 q-1.6 -5.0 -0.4 -9.0" stroke="#74b65e" stroke-width="1.4"/>
<path d="M129.3 9 q-0.2 -4.1 4.2 -7.4" stroke="#42833a" stroke-width="1.1"/>
<path d="M131.7 9 q-1.7 -2.6 -0.4 -4.7" stroke="#42833a" stroke-width="1.2"/>
<path d="M133.0 9 q2.2 -4.9 -0.0 -8.9" stroke="#579a40" stroke-width="1.2"/>
<path d="M134.8 9 q-1.2 -4.4 -0.5 -8.0" stroke="#42833a" stroke-width="1.6"/>
<path d="M137.2 9 q2.3 -4.9 1.1 -8.8" stroke="#42833a" stroke-width="1.3"/>
<path d="M139.4 9 q-2.1 -4.4 2.3 -7.9" stroke="#4f913b" stroke-width="0.9"/>
<path d="M141.6 9 q1.3 -2.4 2.5 -4.4" stroke="#74b65e" stroke-width="1.4"/>
<path d="M143.3 9 q1.8 -3.5 2.7 -6.4" stroke="#74b65e" stroke-width="1.4"/>
<path d="M146.0 9 q0.1 -4.9 0.9 -8.9" stroke="#579a40" stroke-width="0.8"/>
<path d="M147.2 9 q2.3 -3.6 -2.1 -6.5" stroke="#42833a" stroke-width="1.5"/>
<path d="M149.9 9 q0.8 -3.0 4.6 -5.4" stroke="#63ad4d" stroke-width="0.9"/>
<path d="M151.7 9 q2.5 -3.7 2.8 -6.8" stroke="#74b65e" stroke-width="1.7"/>
<path d="M154.4 9 q-1.8 -2.5 2.6 -4.5" stroke="#63ad4d" stroke-width="1.8"/>
<path d="M155.9 9 q2.4 -4.9 3.1 -8.9" stroke="#63ad4d" stroke-width="1.2"/>
<path d="M158.8 9 q2.4 -3.5 2.3 -6.3" stroke="#63ad4d" stroke-width="1.2"/>
<path d="M161.5 9 q-1.5 -4.1 2.1 -7.5" stroke="#63ad4d" stroke-width="1.7"/>
<path d="M164.2 9 q2.4 -3.4 -1.4 -6.1" stroke="#63ad4d" stroke-width="1.1"/>
<path d="M166.1 9 q1.0 -2.4 2.1 -4.3" stroke="#74b65e" stroke-width="1.1"/>
<path d="M167.7 9 q-1.2 -5.2 1.7 -9.4" stroke="#579a40" stroke-width="1.0"/>
<path d="M169.8 9 q0.9 -5.0 1.2 -9.1" stroke="#74b65e" stroke-width="1.7"/>
<path d="M171.9 9 q-2.3 -5.1 0.1 -9.2" stroke="#63ad4d" stroke-width="1.4"/>
<path d="M174.8 9 q-2.3 -4.0 -2.4 -7.4" stroke="#42833a" stroke-width="1.8"/>
<path d="M176.1 9 q2.3 -3.9 3.4 -7.1" stroke="#74b65e" stroke-width="1.6"/>
<path d="M178.2 9 q2.3 -2.2 2.4 -4.1" stroke="#579a40" stroke-width="1.7"/>
<path d="M179.6 9 q-0.5 -3.2 -3.3 -5.8" stroke="#4f913b" stroke-width="1.1"/>
<path d="M181.1 9 q-2.1 -3.9 1.2 -7.0" stroke="#4f913b" stroke-width="1.4"/>
<path d="M183.7 9 q-2.5 -3.2 0.2 -5.8" stroke="#4f913b" stroke-width="1.7"/>
<path d="M186.2 9 q0.9 -4.4 0.1 -8.1" stroke="#42833a" stroke-width="1.3"/>
<path d="M188.7 9 q-1.1 -2.8 1.0 -5.1" stroke="#579a40" stroke-width="1.1"/>
<path d="M190.3 9 q-0.9 -3.0 -3.0 -5.5" stroke="#579a40" stroke-width="1.8"/>
<path d="M193.2 9 q0.2 -5.1 0.9 -9.3" stroke="#4f913b" stroke-width="0.9"/>
<path d="M195.3 9 q-1.0 -2.3 -3.2 -4.2" stroke="#4f913b" stroke-width="1.7"/>
<path d="M196.9 9 q-1.2 -2.7 1.3 -5.0" stroke="#4f913b" stroke-width="0.9"/>
<path d="M199.2 9 q-1.5 -4.5 4.5 -8.1" stroke="#63ad4d" stroke-width="1.1"/>
<path d="M201.6 9 q1.9 -3.7 -0.1 -6.7" stroke="#63ad4d" stroke-width="1.8"/>
<path d="M202.8 9 q-0.3 -2.2 0.4 -3.9" stroke="#42833a" stroke-width="1.6"/>
<path d="M205.5 9 q0.7 -4.2 3.1 -7.6" stroke="#579a40" stroke-width="0.9"/>
<path d="M207.3 9 q0.4 -4.3 -2.1 -7.8" stroke="#63ad4d" stroke-width="0.9"/>
<path d="M209.9 9 q2.0 -3.4 -0.4 -6.3" stroke="#4f913b" stroke-width="1.6"/>
<path d="M211.1 9 q-1.7 -2.1 0.8 -3.8" stroke="#4f913b" stroke-width="1.5"/>
<path d="M214.0 9 q-2.0 -2.5 4.3 -4.6" stroke="#579a40" stroke-width="0.9"/>
<path d="M215.6 9 q-2.2 -2.2 -2.1 -4.0" stroke="#74b65e" stroke-width="1.0"/>
<path d="M217.1 9 q-1.8 -3.1 -0.9 -5.6" stroke="#63ad4d" stroke-width="1.3"/>
<path d="M220.0 9 q1.4 -4.7 -1.8 -8.5" stroke="#42833a" stroke-width="1.8"/>
<path d="M223.1 9 q0.9 -3.4 -1.5 -6.2" stroke="#42833a" stroke-width="1.7"/>
<path d="M225.3 9 q2.4 -3.5 -3.8 -6.3" stroke="#4f913b" stroke-width="0.9"/>
<path d="M227.1 9 q2.2 -3.8 -3.0 -7.0" stroke="#579a40" stroke-width="1.0"/>
<path d="M229.7 9 q2.1 -4.4 4.4 -8.0" stroke="#74b65e" stroke-width="1.3"/>
<path d="M231.0 9 q-1.2 -4.8 2.8 -8.7" stroke="#63ad4d" stroke-width="1.6"/>
<path d="M232.7 9 q-0.9 -3.1 -3.3 -5.6" stroke="#63ad4d" stroke-width="1.5"/>
<path d="M234.8 9 q2.1 -4.6 0.4 -8.4" stroke="#4f913b" stroke-width="1.0"/>
<path d="M237.9 9 q-0.9 -4.6 -0.1 -8.4" stroke="#74b65e" stroke-width="1.4"/>
<path d="M239.4 9 q1.6 -4.7 -1.0 -8.5" stroke="#579a40" stroke-width="1.6"/>
<path d="M241.3 9 q1.9 -5.1 0.9 -9.4" stroke="#579a40" stroke-width="1.4"/>
<path d="M244.3 9 q-2.2 -3.6 -1.9 -6.6" stroke="#42833a" stroke-width="1.1"/>
<path d="M245.6 9 q1.7 -4.2 -1.7 -7.6" stroke="#579a40" stroke-width="1.4"/>
<path d="M247.1 9 q0.1 -3.1 2.0 -5.7" stroke="#42833a" stroke-width="1.5"/>
<path d="M249.5 9 q-2.2 -3.8 -2.2 -6.9" stroke="#42833a" stroke-width="0.9"/>
<path d="M251.5 9 q-2.1 -4.4 3.5 -8.0" stroke="#579a40" stroke-width="0.9"/>
<path d="M253.2 9 q0.8 -4.3 0.7 -7.8" stroke="#4f913b" stroke-width="1.7"/>
<path d="M255.5 9 q1.4 -2.2 -2.9 -4.0" stroke="#4f913b" stroke-width="1.4"/>
<path d="M258.1 9 q1.9 -4.6 1.0 -8.3" stroke="#63ad4d" stroke-width="1.2"/>
<path d="M260.5 9 q-1.9 -4.8 -3.7 -8.6" stroke="#579a40" stroke-width="1.4"/>
<path d="M263.5 9 q1.4 -3.9 -1.3 -7.2" stroke="#4f913b" stroke-width="1.1"/>
<path d="M265.2 9 q-0.9 -4.8 2.7 -8.7" stroke="#42833a" stroke-width="0.9"/>
<path d="M267.4 9 q2.5 -2.2 3.0 -4.1" stroke="#4f913b" stroke-width="1.5"/>
<path d="M269.4 9 q-1.8 -1.9 -3.9 -3.5" stroke="#579a40" stroke-width="1.5"/>
<path d="M272.3 9 q-1.8 -4.5 2.9 -8.2" stroke="#74b65e" stroke-width="1.1"/>
<path d="M274.0 9 q-0.6 -2.2 1.8 -3.9" stroke="#4f913b" stroke-width="1.6"/>
<path d="M276.4 9 q1.2 -2.8 -3.9 -5.1" stroke="#63ad4d" stroke-width="1.5"/>
<path d="M277.7 9 q2.4 -4.2 3.0 -7.7" stroke="#4f913b" stroke-width="1.1"/>
<path d="M279.6 9 q0.6 -5.0 -3.3 -9.2" stroke="#4f913b" stroke-width="1.4"/>
<path d="M282.5 9 q0.6 -3.5 1.2 -6.3" stroke="#4f913b" stroke-width="1.2"/>
<path d="M284.0 9 q-2.4 -3.8 -2.5 -6.8" stroke="#74b65e" stroke-width="1.5"/>
<path d="M285.6 9 q-1.6 -5.2 -2.4 -9.4" stroke="#63ad4d" stroke-width="1.5"/>
<path d="M288.4 9 q1.6 -2.4 3.1 -4.4" stroke="#63ad4d" stroke-width="1.2"/>
<path d="M290.2 9 q-1.9 -3.0 2.9 -5.4" stroke="#63ad4d" stroke-width="0.8"/>
<path d="M291.6 9 q-0.5 -2.7 -2.1 -4.9" stroke="#579a40" stroke-width="1.8"/>
<path d="M294.2 9 q0.8 -3.7 4.4 -6.7" stroke="#63ad4d" stroke-width="0.9"/>
<path d="M296.6 9 q1.2 -3.7 4.2 -6.7" stroke="#74b65e" stroke-width="0.9"/>
<path d="M298.1 9 q1.0 -3.8 3.9 -7.0" stroke="#42833a" stroke-width="0.8"/>
</g>
</svg>

After

Width:  |  Height:  |  Size: 11 KiB

+299
View File
@@ -0,0 +1,299 @@
/* Summer theme: the banner's illustrated meadow carried through the whole
page — one palette (sky blue, grass green, sun yellow, flower pink),
soft daylight surfaces, and playful touches like a tilted brand and
flower bullets. */
:root {
color-scheme: light;
/* Every color is sampled (or text-darkened) from banner.svg. */
--bg: #e6f4cf;
--sky: #d9f1ff;
/* pale meadow — the banner fades into this */
--surface: #f8fbf0;
--text: #2c4a2f;
/* deep leaf */
--muted: #5c7a5e;
--accent: #2598d6;
/* sky, darkened for text contrast */
--accent2: #4f913b;
/* grass */
--accent3: #e95f83;
/* flower pink, darkened for text contrast */
--sun: #ffe071;
--line: #4f913b29;
--link: color-mix(var(--text) 38%, var(--accent));
--font-body: var(--font-cause);
--font-heading: var(--font-new-rocker);
--font-brand: var(--font-cause);
--code-x-height: 0.5; /* Cause's x-height ratio */
}
/* The page is the same landscape the banner paints: hazy sky light up top
(with an echo of the banner's sun in the top-right corner) settling into
a pale meadow below. Fixed, so the banner parallax reads against a
still backdrop. */
body {
background:
radial-gradient(ellipse 36rem 24rem at 85% -6rem,
#fff0a299 0,
#fff0a63d 45%,
transparent 70%),
radial-gradient(ellipse 42rem 26rem at 6% 2rem,
#8ed0ff4d 0,
transparent 70%),
radial-gradient(ellipse 30rem 22rem at 96% 62%,
#ff8aa726 0,
transparent 70%),
linear-gradient(180deg,
var(--bg) 0,
#d9eebd 24rem,
#c6e5a8 55rem,
#b8dd9c 100%);
background-attachment: fixed;
color: var(--text);
}
/* No separation between banner and page: the frame loses its background,
border and shadow, and the artwork overflows into the document below,
masked by a transparency gradient so it cross-fades into the body's fixed
meadow gradient. A plain color match is impossible — both sides are
gradients, and the parallax drift keeps the banner side moving. */
#banner {
background: none;
border-bottom: 0;
box-shadow: none;
}
#page-banner {
inset: 0 0 -6rem;
/* Paints over the page background in the overlap, but never swallows its
clicks. */
pointer-events: none;
mask-image: linear-gradient(180deg, #000 calc(100% - 6rem), transparent);
/* Canvas banner designs (eyes) default to the 13rem layout height; here
they must fill the taller overflow stage so their meadow extension
reaches through the fade strip. */
--eyes-h: 100%;
}
/* The mask above replaced the SVG's flat-color bottom fade (meadowFade) —
fading to a single --bg would reintroduce a seam against the gradient. */
#page-banner .summer-fade {
stop-opacity: 0;
}
/* Cheerful oversized tilted brand in a sky→grass→flower gradient. */
#brand {
font-size: clamp(3.2rem, 9vw, 7.5rem);
line-height: 1;
padding-bottom: 0.3em;
margin-bottom: -0.3em;
transform: rotate(-2deg);
transform-origin: left bottom;
background: linear-gradient(100deg,
var(--accent) 0 30%,
var(--accent2) 55%,
var(--accent3) 95%);
-webkit-background-clip: text;
background-clip: text;
color: transparent;
text-shadow: none;
filter:
drop-shadow(0 1px 0 #fff) drop-shadow(0 0.15rem 0.15rem #ffffff99) drop-shadow(0 0.3rem 0.4rem #4f913b33);
}
/* Navigation floats over the artwork: dark leaf text with a daylight halo. */
#nav {
color: #24452d;
text-shadow:
0 1px #ffffff,
0 0 0.2em #ffffff,
0 0 0.5em #d9f1ff;
}
#nav a {
color: #2c4a2f;
}
#nav ul ul a,
#nav span {
color: #527048;
}
#nav a:hover,
#nav .current {
color: var(--accent);
}
/* Content sits in a soft wash of sunlit meadow rather than on a white
card. The wash lives on a pseudo-element so it can carry a vertical
mask: it fades in from transparent over the strip where the banner
artwork overflows into the page, so the two never fight — no matter
which side paints on top. */
main {
position: relative;
}
main::before {
content: "";
position: absolute;
inset: 0;
z-index: -1;
background: linear-gradient(90deg,
transparent,
#f0f8dc80 15%,
#f0f8dc8c 50%,
#f0f8dc80 85%,
transparent);
mask-image: linear-gradient(180deg, transparent, #000 6rem);
}
/* The sidebar is a piece of the same meadow: glassy green with a light
top edge and a grassy shadow. */
#sidebar {
background: linear-gradient(160deg, #f4faddd9, #d9eec5cf);
border-inline-end: 1px solid #ffffff80;
border-bottom: 1px solid var(--line);
box-shadow: 0 0.3rem 1rem #4f913b1f;
border-radius: 1rem;
margin-inline: 0.5rem;
}
#sidebar a {
color: #527048;
}
#sidebar a:hover,
#sidebar .current {
color: var(--accent);
}
/* One accent per heading level: sky h1 growing a dense low strip of grass
right against its baseline (grass.svg, ~140 varied blades per tile,
served by the backend next to this stylesheet), grass h2, flower h3 in
a quieter weight. */
article h1 {
color: var(--accent);
width: fit-content;
text-shadow: 0 1px #ffffffaa;
background: url("grass.svg") bottom left / auto 0.35em repeat-x;
}
article h2 {
color: var(--accent2);
text-shadow: 0 1px #ffffff88;
}
article h3 {
color: var(--accent3);
font-weight: 550;
}
article a {
color: var(--link);
}
article a:hover {
color: var(--accent);
}
/* Fleur-de-lis on the first level, flowers below — one accent per level.
(⚜️ renders emoji-style, so it needs no shrinking like the flowers.) */
article ul li::before {
content: "⚜️";
color: var(--accent3);
}
article ul ul li::before {
content: "✿";
color: var(--accent);
}
article ul ul ul li::before {
content: "❁";
color: var(--accent2);
}
/* Quotes get a grassy edge and a wash of sunlight. */
blockquote {
color: #4d6849;
border-inline-start-color: var(--accent2);
background: linear-gradient(90deg, #fff0a238, transparent 70%);
padding-top: 0.25rem;
padding-bottom: 0.25rem;
}
/* Admonitions join the rounded, softly shaded landscape elements. */
.admonition {
border-radius: 0.6rem;
box-shadow: 0 0.2rem 0.7rem #4f913b1a;
}
/* Code rests on a shaded leaf panel, keeping the base light syntax
palette. */
pre {
--code-bg: #e7f1d6;
background: linear-gradient(160deg, #f0f7e2, #ddebc9);
border: 1px solid var(--line);
box-shadow:
inset 0 1px #ffffff99,
0 0.2rem 0.6rem #4f913b1a;
}
.copy {
background: linear-gradient(180deg, #f6faec, #e2efcf);
border-color: var(--line);
}
/* Tables: a rounded piece of the landscape. The base's header-gradient
hooks get sky→grass colors and the cell wash tints green; soft corners
and a grassy shadow replace the old hard grid. */
table {
--table-head-a: #fdfccd88;
--table-head-b: #ece9ba88;
--table-tint: var(--accent2);
border-collapse: separate;
border-spacing: 0;
border-radius: 0.7rem;
overflow: clip;
box-shadow: 0 0.1rem 0.5rem #4f913b88;
}
th {
color: #2c4a2f;
}
/* A sun-kissed sheen on hover, fading in over the cell wash. */
tbody td {
transition: background-color 0.2s;
}
tbody tr:hover td {
background-color: #fff0a238;
}
figcaption,
.dateline,
.footnote {
color: var(--muted);
}
/* Editor chrome inherits the meadow rather than the base white. */
.editor-root.overlay {
background: linear-gradient(160deg, #eef7dd, #d9eec5);
}
::view-transition {
/* Override to mid sky shade instead of the darker --bg we otherwise get */
background: var(--sky);
}
+488
View File
@@ -0,0 +1,488 @@
"""Visit analytics: collection sockets, geoip enrichment, favicon fetch.
The visitor-activity WebSocket (``/_ws``, public) and the admin analytics
stream (``/_api/ws/analytics``) plus the ``/_a`` viewer page. Client IPs are
enriched in background tasks with reverse DNS (cached PTR lookups) and the
DB-IP city MMDB (``GeoIP``, decompressed and opened once at startup);
external referrers get their favicon fetched and stored content-hashed.
Snapshot broadcasts to connected admin sockets are debounced.
"""
import asyncio
import gzip
import ipaddress
import logging
import os
import re
import shutil
import socket
from datetime import date
from functools import lru_cache
from pathlib import Path
from urllib.parse import urlparse
import httpx
import msgspec
from fastapi import APIRouter, Request, WebSocket, WebSocketDisconnect
from fastapi.responses import Response
from pagerite import analytics
from pagerite.data import resolve
from pagerite.files import _hash_name, file_store
from pagerite.state import SITE_URL, _html_response, analytics_store, data
logger = logging.getLogger(__name__)
# httpx logs every request at INFO (e.g. the favicon fetches below); our own
# one-line summary in _schedule_favicon_fetch replaces that noise.
logging.getLogger("httpx").setLevel(logging.WARNING)
router = APIRouter()
# Live WebSocket clients for the analytics stream.
_analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None
# DB-IP databases persist in the working directory (one download serves all
# sites run from it). Not the package directory: reinstalls/upgrades wipe it.
_DBIP_DIR = Path.cwd()
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
def _download_dbip() -> None:
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
today = date.today()
months = [f"{today:%Y-%m}"]
# The current month's file may not be published yet; fall back to last month.
prev = (today.replace(day=1) - date.resolution).replace(day=1)
months.append(f"{prev:%Y-%m}")
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
logger.info("DB-IP database is current (%s), skipping download", existing[-1])
return
for month in months:
url = DBIP_URL.format(month=month)
target = _DBIP_DIR / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
logger.info("Downloading %s", url)
try:
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
if r.status_code == 404:
continue
r.raise_for_status()
with open(tmp, "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk)
except httpx.HTTPError as e:
logger.warning("DB-IP download failed: %s", e)
tmp.unlink(missing_ok=True)
continue
# Verify it is actually gzip data before installing it.
try:
with gzip.open(tmp, "rb") as f:
f.read(1)
except OSError:
logger.warning("DB-IP download for %s was not valid gzip", month)
tmp.unlink(missing_ok=True)
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _DBIP_DIR.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
logger.info("DB-IP database updated to %s", target.name)
return
logger.warning("Could not download a DB-IP database")
def _geoip_db_path() -> Path | None:
"""Find a DB-IP MMDB in the working directory, preferring an already-decompressed
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present.
"""
mmdb = sorted(_DBIP_DIR.glob("dbip-*.mmdb"))
if mmdb:
return mmdb[0]
gz = sorted(_DBIP_DIR.glob("dbip-*.mmdb.gz"))
if gz:
return gz[0]
return None
class GeoIP:
"""Lazy DB-IP MMDB reader. Call ``_load()`` once at startup before
concurrent requests arrive; ``country()`` is read-only and safe to call
from ``asyncio.to_thread`` workers afterwards.
"""
def __init__(self) -> None:
self._reader: object | None = None
def _decompress(self, source: Path, target: Path) -> None:
if target.exists():
return
tmp = target.with_suffix(target.suffix + ".tmp")
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
shutil.copyfileobj(src, dst)
os.replace(tmp, target)
def _load(self) -> None:
if self._reader is not None:
return
source = _geoip_db_path()
if source is None:
return
if source.suffix == ".gz":
target = source.with_suffix("")
self._decompress(source, target)
source = target
try:
import maxminddb
self._reader = maxminddb.open_database(str(source))
except Exception:
pass
def country(self, ip: str) -> str:
"""Two-letter ISO country code for ``ip``, or "" when unavailable."""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
return (rec.get("country") or {}).get("iso_code", "")
except Exception:
pass
return ""
def city(self, ip: str) -> str:
"""City name for ``ip``, or "" when unavailable.
GeoIP sometimes appends district names in parentheses (e.g.
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
the value is stored.
"""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
city = (rec.get("city") or {}).get("names", {}).get("en", "")
if city:
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
return city
except Exception:
pass
return ""
_geoip = GeoIP()
def _client_ip(request: Request | WebSocket) -> str:
"""Client IP: first X-Forwarded-For hop (we sit behind a proxy), else
the direct peer."""
forwarded = request.headers.get("x-forwarded-for", "").split(",")[0].strip()
return forwarded or (request.client.host if request.client else "")
def _query_suffix(request: Request) -> str:
"""The request's query string as a "?..." suffix, or "" when absent."""
query = str(request.url.query)
return f"?{query}" if query else ""
@lru_cache(maxsize=4096)
def _cached_ptr(ip: str) -> str:
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
if not ip:
return ""
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ""
if (
addr.is_private
or addr.is_loopback
or addr.is_reserved
or addr.is_multicast
or addr.is_link_local
):
return ""
try:
host, _, _ = socket.gethostbyaddr(ip)
except socket.herror:
return ""
return host
async def _lookup_host(ip: str) -> str:
"""Async wrapper around ``_cached_ptr``; runs the blocking lookup in a thread."""
return await asyncio.to_thread(_cached_ptr, ip)
async def _geoip_country(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB lookup."""
return await asyncio.to_thread(_geoip.country, ip)
async def _geoip_city(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB city lookup."""
return await asyncio.to_thread(_geoip.city, ip)
async def _enrich_client(client_hash: bytes) -> None:
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
client = analytics_store.data.clients.get(client_hash)
if not client or not client.ip:
return
host = await _lookup_host(client.ip)
country = await _geoip_country(client.ip)
city = await _geoip_city(client.ip)
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
"""Start background host/geoip enrichment for the given client hashes."""
for client_hash in client_hashes:
asyncio.create_task(_enrich_client(client_hash))
#: Icon MIME -> file extension for the stored favicon name. The extension
#: reflects the actual content, not the /favicon.ico request path.
_FAVICON_EXT = {
"image/x-icon": ".ico",
"image/vnd.microsoft.icon": ".ico",
"image/png": ".png",
"image/gif": ".gif",
"image/jpeg": ".jpg",
"image/webp": ".webp",
"image/avif": ".avif",
"image/svg+xml": ".svg",
}
_FAVICON_MAX_BYTES = 65536
#: Origins with a fetch task currently in flight.
_favicon_in_flight: set[str] = set()
async def _fetch_favicon(origin: str) -> None:
"""Fetch ``{origin}/favicon.ico`` and store it content-hashed on disk.
The result (icon file name, or "" for a miss) is recorded in the
analytics store; misses are retried after analytics._FAVICON_RETRY.
Never raises: analytics must not break page serving.
"""
try:
async with httpx.AsyncClient(follow_redirects=True, timeout=8) as client:
r = await client.get(f"{origin}/favicon.ico")
body = r.content
if (
not (200 <= r.status_code < 300)
or not body
or len(body) > _FAVICON_MAX_BYTES
):
analytics_store.record_favicon(origin)
return
mime = r.headers.get("content-type", "").split(";")[0].strip().lower()
if not mime.startswith("image/"):
# Served without an image type: sniff SVG, else assume ICO.
if b"<svg" in body[:1024]:
mime = "image/svg+xml"
elif mime in ("", "application/octet-stream", "text/plain"):
mime = "image/x-icon"
else:
analytics_store.record_favicon(origin)
return
ext = _FAVICON_EXT.get(mime, ".ico")
name = _hash_name(body, f"favicon{ext}")
file_store.put(name, body)
analytics_store.record_favicon(origin, name)
except httpx.HTTPError, OSError:
analytics_store.record_favicon(origin)
finally:
_favicon_in_flight.discard(origin)
def _schedule_favicon_fetch() -> None:
"""Start background favicon fetches for origins that need one."""
origins = [
origin
for origin in analytics_store.favicon_origins_needed()
if origin not in _favicon_in_flight
]
if not origins:
return
logger.info(
"Fetching favicons: %s",
", ".join(o.removeprefix("https://") for o in origins),
)
for origin in origins:
_favicon_in_flight.add(origin)
asyncio.create_task(_fetch_favicon(origin))
async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients:
return
payload = analytics_store.display_json(_in_menu)
closed = set()
for ws in _analytics_ws_clients:
try:
await ws.send_text(payload)
except Exception:
closed.add(ws)
for ws in closed:
_analytics_ws_clients.discard(ws)
async def _debounced_analytics_broadcast() -> None:
"""Wait briefly, then broadcast the latest snapshot once."""
await asyncio.sleep(0.2)
await _broadcast_analytics()
def _schedule_analytics_broadcast() -> None:
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
global _analytics_broadcast_task
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
return
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
_debounced_analytics_broadcast()
)
def _in_menu(path: str) -> bool:
"""True when ``path`` ("/a/b" or "/") resolves to a real menu node.
Category placeholders return 404 but are real nodes: their GETs must not
count as misses in the display-time abuse classification.
"""
return resolve(data.menu, path.strip("/")) is not None
def _record_get(request: Request, *, status: int = 200) -> None:
"""Record the document GET as one raw access-log line in analytics.
Nothing is classified here — the true HTTP status, the full request path
(query included), an external referer origin and the preload flag are
stored, and visitor/crawler/abuse classification happens at display time
(see analytics.Store.display). Idle-time preloads from pagerite.js
(``x-pagerite-preload`` header) are recorded with ``pre=True``: never
counted, but a navigation later served from the in-memory page cache is
attributed this GET's status.
The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic. The root-path and
localhost checks prevent remote visitors from forging the same query.
"""
if (
request.url.path == "/"
and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1"
):
return
own_origin = SITE_URL or f"https://{urlparse(str(request.base_url)).netloc}"
referer = request.headers.get("referer", "")
if analytics._origin(referer) in (None, own_origin):
referer = ""
client_hash = analytics_store.record_get(
_client_ip(request),
request.headers.get("user-agent", ""),
f"{request.url.path}{_query_suffix(request)}",
status=status,
referer=referer,
accept_language=request.headers.get("accept-language", ""),
pre=bool(request.headers.get("x-pagerite-preload")),
)
if client_hash is not None:
_schedule_client_enrichment([client_hash])
@router.get("/_a", response_model=None)
async def analytics_page(request: Request) -> Response:
"""Render the analytics viewer as a normal site page at /_a.
The page itself is public, but the data stream (/_api/ws/analytics) stays
admin-gated like the rest of /_api, so only authorized users see the
statistics; others get the viewer with a "could not be loaded" message.
"""
return _html_response(
request,
"analytics",
"",
headers={"cache-control": "no-cache"},
etag=True,
)
@router.websocket("/_ws")
async def activity_ws(ws: WebSocket) -> None:
"""Collect visitor activity: navigations and reading-time updates.
Public, like the pages themselves (only /_api is gated); one connection
follows a browsing session. Messages are ``analytics.Ping`` structs as
JSON text frames; ``to`` set is a navigation, ``read`` alone a
reading-time update. Everything is recorded raw — known bot UAs and
abusive IPs are filtered at display time, not here. The reverse-DNS and
DB-IP geoip lookups happen in background tasks so message handling is
never delayed by slow DNS or the first MMDB decompress.
"""
ip = _client_ip(ws)
ua = ws.headers.get("user-agent", "")
accept_language = ws.headers.get("accept-language", "")
# Identify the visitor on the access-log open/close lines (the IP is
# already printed there): compact UA plus the browser's language tag.
lang, _country = analytics._parse_accept_language(accept_language)
ws.scope.setdefault("state", {})["log_extra"] = " ".join(
part for part in (analytics._compact_user_agent(ua), lang) if part
)
await ws.accept()
try:
while True:
text = await ws.receive_text()
try:
msg = msgspec.json.decode(text.encode(), type=analytics.Ping)
except msgspec.DecodeError:
continue
new_client = analytics_store.record_msg(
msg.fr,
msg.to or None,
ip,
ua,
accept_language,
hide=msg.hide,
read=msg.read,
)
if new_client is not None:
_schedule_client_enrichment([new_client])
_schedule_favicon_fetch()
except WebSocketDisconnect:
pass
@router.websocket("/_api/ws/analytics")
async def analytics_websocket(ws: WebSocket) -> None:
"""Stream the analytics snapshot, then push updates as they happen.
Admin-only via the /_api forward-auth gate, like every management
endpoint. Powers the analytics viewer rendered at /_a.
"""
await ws.accept()
await ws.send_text(analytics_store.display_json(_in_menu))
_analytics_ws_clients.add(ws)
try:
while True:
await ws.receive_text()
except Exception:
pass
finally:
_analytics_ws_clients.discard(ws)
+417
View File
@@ -0,0 +1,417 @@
"""Translator service protocol, dispatcher and its transport-independent core.
The external machine-translation service connects over WebSocket
(``/_translate/<key>``, the route itself is in api.py) and exchanges JSON
frames decoded into the tagged msgspec structs below (``bytes`` fields ride
as base64 — no manual encoding anywhere). This module holds everything
else: the message structs, the connected-client dispatcher (``Dispatcher``
— one job at a time per connection, wanted ∩ capable language matching,
requeue on disconnect), which fragments are pending for a language
(``pending_items``) and storing a result (``store_results``).
Fragments cross the wire as **prose segments**: the model only ever
receives plain text runs (Job.texts) plus per-segment context surrounds
(Job.contexts) and returns their translations (Result.texts, same order);
markup never leaves the server — reassembly is offset splicing
(``pagerite/segments.py``).
"""
import asyncio
import logging
import msgspec
from fastapi import WebSocket, WebSocketDisconnect
from kanta import Kanta
from pagerite import i18n
from pagerite.chunks import chunk_key, needs_translation
from pagerite.data import Data, Node, sorted_nodes
from pagerite.segments import Span, join, split
logger = logging.getLogger(__name__)
class Hello(msgspec.Struct, tag="hello"):
"""Client greeting on connect: the language codes its model CAN produce
(capabilities). The server offers jobs only in the intersection with
the wanted target languages (``Data.translate_langs``)."""
langs: list[str]
class TransItem(msgspec.Struct):
"""One fragment to translate: original Markdown (or a node title)."""
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
text: str
path: str #: article it came from ("" = front page), no leading slash
kind: str #: "chunk" | "title"
#: Title jobs only: the article's opening prose, so the model sees the
#: title as a heading in context, not a lone sentence.
context: str = ""
class Job(msgspec.Struct, tag="job"):
"""Server push: ONE fragment to translate.
Exactly one job is in flight per connection — the next is sent only
after this one's Result. Clients wanting parallelism open multiple
connections."""
lang: str
key: bytes #: 9-byte chunk hash (base64 in the JSON frame)
#: The fragment's prose segments (pagerite/segments.py): plain text
#: runs only — no markup, URLs, code or placeholders ever cross the
#: wire. Translate each element independently.
texts: list[str]
path: str #: article it came from ("" = front page), no leading slash
kind: str #: "chunk" | "title"
#: Per segment (parallel to texts; "" = none): the surround to
#: translate it in — a carved-out segment (link text, partial run)
#: carries its block's plain text, a title the article's opening.
#: Reference client behavior (scripts/translator.py): translate
#: segment+context together, keep the segment's part (its own line /
#: paragraph); fall back to the segment alone when the output holds no
#: separator. Contexts are not part of the result.
contexts: list[str] = msgspec.field(default_factory=list)
class TransResult(msgspec.Struct):
"""One translated fragment (storage level, see store_results)."""
key: bytes
text: str
class Result(msgspec.Struct, tag="result"):
"""Client reply: the translation of the connection's current Job
(must match its lang and key exactly)."""
lang: str
key: bytes
#: The job's segments, translated, same order and count. Each must be
#: pure prose — the server rejects the result otherwise.
texts: list[str]
#: Union of the client -> server frames (the "type" tag selects).
ClientMsg = Hello | Result
def pending_items(data: Data, lang: str) -> list[TransItem]:
"""Fragments of the site still untranslated for ``lang``, deduped by key.
Every node (published or not, pages and pure category labels alike)
contributes its title; pages also contribute each chunk
that needs translation (``needs_translation``), is not editor-flagged
no-translate (``node.no_trans``) and has no ``trans`` entry for ``lang``
yet. Content-addressed text (shared paragraphs, repeated titles) appears
once, under the first page in menu order that has it.
"""
items: list[TransItem] = []
seen: set[bytes] = set()
def emit(key: bytes, text: str, path: str, kind: str, context: str = "") -> None:
if key in seen or lang in data.trans.get(key, {}):
return
seen.add(key)
items.append(
TransItem(key=key, text=text, path=path, kind=kind, context=context)
)
def opening(node: Node) -> str:
"""The article's opening prose (first segment, capped): the title
job's context — a lone word like "About" reads as a heading on top
of an article, not as a sentence. Empty when there's no prose."""
for h in node.chunks or ():
text = data.chunks.get(h)
if text and (segs := split(text)[1]):
return segs[0][:400]
return ""
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
# An article whose primary language IS the target needs no
# translation into it — skip its title and chunks entirely.
# Category labels (chunks is None) contribute only their title:
# it is their nav-menu label.
node_lang = node.language or inherited
if node_lang != lang:
if node.title:
emit(
chunk_key(node.title),
node.title,
path,
"title",
context=opening(node),
)
for h in node.chunks or ():
text = data.chunks.get(h)
if (
text is not None
and h not in node.no_trans
and needs_translation(text)
):
emit(h, text, path, "chunk")
walk(node.children, path, node_lang)
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
return items
def store_results(data: Data, lang: str, items: list[TransResult]) -> list[str]:
"""Store machine translations for ``lang``; return the paths of the
articles that gained at least one entry.
Pure data operations: the caller wraps this in a kanta transaction and
invalidates pages. Unknown keys are stored anyway (unreferenced hashes
are never read, and the content may simply have moved on since the job
was pushed); re-storing an existing key overwrites, last wins. Every
article that gained an entry gets ``node.langs[lang]`` set (the
availability index, docs/migrate.md) — because chunks are
content-addressed, that includes pages merely sharing a fragment.
"""
stored = {item.key for item in items}
for item in items:
data.trans.setdefault(item.key, {})[lang] = item.text
pages: list[str] = []
def walk(nodes: dict[str, Node], prefix: str, inherited: str) -> None:
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
node_lang = node.language or inherited
if node_lang != lang:
keys = set(node.chunks or ())
if node.title:
keys.add(chunk_key(node.title))
if keys & stored:
node.langs[lang] = True
pages.append(path)
walk(node.children, path, node_lang)
walk(data.menu, "", i18n.ORIGINAL_LANGUAGE)
return pages
class _Connection:
"""One connected translator socket: the language codes it announced as
capabilities (Hello) and the (lang, chunk-key) job currently in flight
on it, with the segment spans to splice its Result into
(pagerite/segments.py) — one at a time, the next is sent only after its
Result.
Per-connection only: in-flight lives solely here, so on disconnect the
item simply becomes pending again and is re-offered to any free capable
connection."""
def __init__(self, capable: set[str]) -> None:
self.capable = capable
self.inflight: tuple[str, bytes] | None = None
#: Source spans of the in-flight job's segments (splice offsets
#: and link marks).
self.spans: list[Span] = []
self.original: str = "" # its full source text (for the splicing)
self.kind: str = "" # "chunk" | "title" (for the transaction action)
class Dispatcher:
"""The translator dispatcher: connected client sockets and the job
pipeline (docs/localization.md).
One single-item job at a time per connection, offered in the
intersection of the wanted languages (``Data.translate_langs``) and the
connection's announced capabilities. Pending work is derived from the
``trans`` store (``pending_items``) minus the items in flight on any
connection, so a dropped connection's in-flight item is simply
re-offered. Results are matched to content by chunk key alone. A
(lang, key) whose Result fails segment validation is skipped for the
rest of the run — generation is near-deterministic, so an immediate
retry would just re-fail.
"""
def __init__(self, data: Data, db: Kanta, invalidate) -> None:
self.data = data
self.db = db
#: Sync content-change hook (state._invalidate_pages), called inside
#: transactions; schedules the next dispatch pass.
self.invalidate = invalidate
#: Connected translator sockets and their per-connection state.
self.clients: dict[WebSocket, _Connection] = {}
#: (lang, chunk key) of fragments whose result failed validation
#: (segment count, empty or non-prose segments, segments.py) this run.
self.validation_failures: set[tuple[str, bytes]] = set()
def reset_validation_failures(self) -> None:
"""Clear the skip list of fragments rejected this run (segment
validation): a translations refresh is precisely the "another
chance" for them."""
self.validation_failures.clear()
def schedule(self) -> None:
"""Schedule a dispatch pass, if any translator is connected.
The invalidate hook is sync and called inside transactions: the
task first runs once the current coroutine awaits again, i.e. after
the transaction has committed. No-op without a running loop (CLI
use)."""
if not self.clients:
return
try:
asyncio.get_running_loop()
except RuntimeError:
return
asyncio.create_task(self._dispatch())
async def _dispatch(self) -> None:
"""Offer one pending item to every free capable connection."""
wanted = {
tag for lang in self.data.translate_langs if (tag := i18n.base_tag(lang))
}
if not wanted:
return
for ws, state in list(self.clients.items()):
if state.inflight is not None:
continue
langs = wanted & state.capable
if not langs:
continue
inflight = {s.inflight for s in self.clients.values() if s.inflight}
job = None
spans: list[Span] = []
original = ""
# Titles before articles — across languages too, so every menu
# is named before any article body is worked on (a page's name
# is its most visible string). pending_items emits in menu
# order, a page's title before its chunks; filtering by kind
# keeps that stable order within each kind.
pending = {lang: pending_items(self.data, lang) for lang in sorted(langs)}
for kind in ("title", "chunk"):
for lang in sorted(langs):
for item in pending[lang]:
if (
item.kind != kind
or (lang, item.key) in inflight
or (lang, item.key) in self.validation_failures
):
continue
spans, texts, contexts = split(item.text)
if not texts:
continue # prose that could not be located for splicing
original = item.text
if item.kind == "title" and item.context:
# A title's surround is the article's opening prose
# (TransItem.context), not its own one-word block.
contexts = [item.context] * len(texts)
job = Job(
lang=lang,
key=item.key,
texts=texts,
path=item.path,
kind=item.kind,
contexts=contexts,
)
break
if job is not None:
break
if job is not None:
break
if job is None:
continue
state.inflight = (job.lang, job.key) # before the await: no double-assign
state.spans = spans
state.original = original
state.kind = job.kind
try:
await ws.send_text(msgspec.json.encode(job).decode())
except Exception: # send failed: the receive loop cleans up
self.clients.pop(ws, None)
async def handle_ws(self, ws: WebSocket, clientkey: str) -> None:
"""The /_translate/<key> channel (docs/localization.md).
A wrong/empty key rejects the handshake (closing before accept
makes Starlette answer HTTP 403). Protocol (JSON frames): the
client opens with Hello(langs) announcing its CAPABILITIES — the
language codes its model can produce (normalized to translation
tags; "en"/empty dropped) — then answers each Job with its
Result(lang, key, texts). A Result without an in-flight job or with
a different (lang, key), a duplicate Hello, or any malformed frame
closes the socket with a protocol error.
"""
if clientkey not in self.data.translate_keys:
await ws.close(code=1008) # policy violation; pre-accept = HTTP 403
return
await ws.accept()
state: _Connection | None = None
try:
while True:
raw = await ws.receive_text()
try:
msg = msgspec.json.decode(raw.encode(), type=ClientMsg)
except msgspec.DecodeError:
await ws.close(code=1002) # protocol error
return
if isinstance(msg, Hello):
if state is not None: # one Hello per connection
await ws.close(code=1002)
return
state = _Connection(
{tag for lang in msg.langs if (tag := i18n.base_tag(lang))}
)
self.clients[ws] = state
self.schedule()
else: # Result
lang = i18n.base_tag(msg.lang)
if (
state is None # results before Hello
or state.inflight is None # no job in flight
or (lang, msg.key) != state.inflight # wrong job
):
await ws.close(code=1002)
return
texts, spans, original = msg.texts, state.spans, state.original
kind, state.kind = state.kind, ""
state.inflight = None
state.spans = []
state.original = ""
text = (
join(original, spans, texts)
if len(texts) == len(spans)
else None
)
if text is None:
# The model broke the segment contract (count
# mismatch, empty or non-prose segment): drop the
# result and skip the fragment for this run (it
# stays pending; a restart, a refresh or a model
# change gets another chance).
self.validation_failures.add((lang, msg.key))
logger.warning(
"[%s] result for chunk %s rejected: invalid segments",
lang,
msg.key.hex(),
)
self.schedule()
continue
with self.db.transaction(
f"translate:{lang}{':title' if kind == 'title' else ''}",
user=clientkey,
):
paths = store_results(
self.data, lang, [TransResult(key=msg.key, text=text)]
)
self.invalidate() # schedules the next dispatch
if paths:
logger.info(
"[%s] now available for %d page(s): %s",
lang,
len(paths),
", ".join(sorted(paths)),
)
except WebSocketDisconnect:
pass
finally:
if self.clients.pop(ws, None) is not None:
# The in-flight item (if any) is pending again; offer it around.
self.schedule()
+1082 -161
View File
File diff suppressed because it is too large Load Diff
+16 -6
View File
@@ -17,13 +17,20 @@ readme = "README.md"
requires-python = ">=3.14" requires-python = ">=3.14"
dependencies = [ dependencies = [
"blake3>=1.0.9", "blake3>=1.0.9",
"fastapi-vue>=1.3.1", "fastapi-vue~=1.4.2",
"fastapi[standard]>=0.141.1", "fastapi[standard]>=0.141.1",
"html5tagger>=2.0.0", "html5tagger>=2.0.0",
"kanta>=0.8.1", "httpx>=0.28.1",
"kanta>=0.9.0",
"markdown-it-py>=4.2.0", "markdown-it-py>=4.2.0",
"maxminddb>=3.1.1",
"mdit-py-plugins>=0.6.1", "mdit-py-plugins>=0.6.1",
"mediapreview[standard]>=0.2.3",
"platformdirs>=4.11.5",
"pygments>=2.20.0", "pygments>=2.20.0",
"python-slugify>=8.0.4",
"ua-parser>=1.0.2",
"zstandard>=0.25.0",
] ]
[project.scripts] [project.scripts]
@@ -31,11 +38,10 @@ pagerite = "pagerite.__main__:main"
[project.urls] [project.urls]
Repository = "https://git.zi.fi/LeoVasanko/pagerite" Repository = "https://git.zi.fi/LeoVasanko/pagerite"
Issues = "https://github.com/LeoVasanko/pagerite"
[dependency-groups] [dependency-groups]
dev = [ dev = []
"httpx>=0.28.1",
]
[tool.hatch.version] [tool.hatch.version]
source = "vcs" source = "vcs"
@@ -44,8 +50,12 @@ source = "vcs"
packages = ["pagerite"] packages = ["pagerite"]
[tool.hatch.build] [tool.hatch.build]
artifacts = ["pagerite/frontend-build"] # `only-packages` drops directories without an __init__.py, so the theme
# files and seed image assets must be force-included as artifacts (like
# the frontend build).
artifacts = ["pagerite/frontend-build", "pagerite/themes", "pagerite/seed-assets"]
only-packages = true only-packages = true
packages = ["pagerite"]
[tool.hatch.build.targets.sdist.hooks.custom] [tool.hatch.build.targets.sdist.hooks.custom]
path = "scripts/fastapi-vue/buildhook.py" path = "scripts/fastapi-vue/buildhook.py"
+6 -3
View File
@@ -8,6 +8,8 @@ import sys
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
import tracerite
# Import util.py from scripts/fastapi-vue (not a package, so we adjust sys.path) # Import util.py from scripts/fastapi-vue (not a package, so we adjust sys.path)
sys.path.insert(0, str(Path(__file__).with_name("fastapi-vue"))) sys.path.insert(0, str(Path(__file__).with_name("fastapi-vue")))
from devutil import ( from devutil import (
@@ -19,8 +21,8 @@ from devutil import (
setup_vite, setup_vite,
) )
DEFAULT_VITE_PORT = 3100 DEFAULT_VITE_PORT = 8200
DEFAULT_DEV_PORT = 3200 DEFAULT_DEV_PORT = 8210
HEALTH = "/?from=devserver.py" HEALTH = "/?from=devserver.py"
@@ -39,7 +41,7 @@ async def run_devserver(
viteurl, npm_install, vite = setup_vite(listen, DEFAULT_VITE_PORT) viteurl, npm_install, vite = setup_vite(listen, DEFAULT_VITE_PORT)
backurl, pagerite = setup_cli("pagerite", backend, DEFAULT_DEV_PORT) backurl, pagerite = setup_cli("pagerite", backend, DEFAULT_DEV_PORT)
# Tell the everyone by environment (vite proxy and backend devmode use these) # Tell everyone via environment (vite proxy and backend devmode use these)
os.environ["PAGERITE_VITE_URL"] = viteurl os.environ["PAGERITE_VITE_URL"] = viteurl
os.environ["PAGERITE_BACKEND_URL"] = backurl os.environ["PAGERITE_BACKEND_URL"] = backurl
os.environ["PAGERITE_DEV"] = "1" os.environ["PAGERITE_DEV"] = "1"
@@ -54,6 +56,7 @@ async def run_devserver(
def main() -> None: def main() -> None:
"""Parse CLI arguments and run the devserver.""" """Parse CLI arguments and run the devserver."""
tracerite.load()
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
description="Run Vite and FastAPI development servers", description="Run Vite and FastAPI development servers",
formatter_class=argparse.RawDescriptionHelpFormatter, formatter_class=argparse.RawDescriptionHelpFormatter,
+949
View File
@@ -0,0 +1,949 @@
#!/usr/bin/env -S uv run --script
# /// script
# requires-python = ">=3.14"
# dependencies = [
# "httpx>=0.28.1",
# "playwright>=1.45.0",
# ]
# ///
"""Generate fake browser visits, crawler hits, and abuse scans for a Pagerite site.
Browser sessions (ordinary users) come from realistic residential IPv4 and IPv6
addresses and stay mostly stable; an IPv6 host part may rotate once mid-session,
and an IPv4 session may switch to another residential address. Crawler hits come
from datacenter IPs, with each crawler profile paired to a matching provider IP
when possible. Abuse scanners fire bursts of vulnerability probes from pinned
datacenter IPs.
Run against a local dev server, e.g.:
uv run scripts/fake_traffic.py http://localhost:3200
Repeat whenever you want more traffic; each run appends new events to the
site's analytics file.
"""
from __future__ import annotations
import argparse
import logging
import random
import sys
import time
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import UTC, datetime
from typing import Any
from urllib.parse import urlencode, urljoin, urlparse
import httpx
logging.basicConfig(level=logging.INFO, format="%(message)s")
logger = logging.getLogger("fake-traffic")
@dataclass(frozen=True)
class BrowserProfile:
name: str
user_agent: str
accept_language: str
viewport: tuple[int, int]
@dataclass(frozen=True)
class CrawlerProfile:
name: str
user_agent: str
ip: str
BROWSER_PROFILES: list[BrowserProfile] = [
BrowserProfile(
"chrome-desktop",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"en-US,en;q=0.9",
(1366, 768),
),
BrowserProfile(
"safari-desktop",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"en-GB,en;q=0.9",
(1440, 900),
),
BrowserProfile(
"firefox-desktop",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"en-CA,en;q=0.8,fr;q=0.5",
(1920, 1080),
),
BrowserProfile(
"chrome-mobile",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"es-ES,es;q=0.9",
(390, 844),
),
]
CRAWLER_PROFILES: list[CrawlerProfile] = [
CrawlerProfile(
"googlebot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"66.249.64.66", # US, Google
),
CrawlerProfile(
"bingbot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"40.77.167.0", # US, Microsoft
),
CrawlerProfile(
"duckduckbot",
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
"95.217.0.1", # Germany, Hetzner VPS
),
CrawlerProfile(
"curl",
"curl/8.5.0",
"139.162.0.1", # Singapore, Linode VPS
),
]
# Residential IPv4 addresses and IPv6 /64 prefixes used for ordinary browser
# sessions. IPv6 entries keep the network part stable and randomise only the
# host part; the host may rotate once mid-session.
RESIDENTIAL_SOURCE_IPS: list[str] = [
# Residential IPv4
"91.154.140.209", # Finland, Elisa
"84.143.145.207", # Germany, Deutsche Telekom
"220.165.255.254", # China, Chinanet / China Telecom
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
# Residential IPv6 /64 prefixes
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
"240e:45d:1e60:5b0::/64", # China, China Telecom
"2409:8904:6720:4123::/64", # China, China Unicom
]
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
# the whole scan burst.
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
# and index 2 uses a fixed crawler UA.
ABUSE_SOURCE_IPS: list[str] = [
"45.63.0.12", # US, Vultr VPS
"138.197.0.89", # US, DigitalOcean / Cloudways
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
]
# Paths commonly probed by attackers looking for exposed config, admin panels,
# version control, credentials, backups, or debug endpoints.
SUSPICIOUS_PATHS: list[str] = [
"/.env",
"/env",
"/.env.local",
"/env.development",
"/config",
"/config.json",
"/config.yaml",
"/config.yml",
"/configuration.json",
"/configuration.yaml",
"/configuration.yml",
"/settings.json",
"/settings.yaml",
"/settings.yml",
"/app.config",
"/appsettings.json",
"/appsettings.Development.json",
"/credentials",
"/credentials.json",
"/secrets",
"/secrets.json",
"/.aws/credentials",
"/.ssh/id_rsa",
"/id_rsa",
"/id_rsa.pub",
"/known_hosts",
"/sftp-config.json",
"/admin",
"/administrator",
"/adminer.php",
"/login",
"/signin",
"/auth/login",
"/api/login",
"/api/.env",
"/api/config",
"/api/v1/config",
"/api/v2/config",
"/webhook",
"/webhooks",
"/callback",
"/proxy",
"/image",
"/images",
"/preview",
"/download",
"/downloads",
"/log",
"/logs",
"/debug",
"/trace",
"/phpinfo.php",
"/info.php",
"/phpmyadmin",
"/pma",
"/myadmin",
"/phpMyAdmin",
"/wp-admin",
"/wp-login.php",
"/wp-config.php",
"/xmlrpc.php",
"/wp-json/wp/v2/users",
"/.git/config",
"/.git/HEAD",
"/git/config",
"/swagger-ui.html",
"/v2/api-docs",
"/actuator/env",
"/actuator/health",
"/actuator/configprops",
"/server-status",
"/.htaccess",
"/web.config",
"/package.json",
"/composer.json",
"/vendor/autoload.php",
"/docker-compose.yml",
"/Dockerfile",
"/manage",
"/console",
"/manager",
"/manager/html",
"/metrics",
"/prometheus",
"/healthz",
"/_api",
"/api",
"/api/v1/",
"/api/v2/",
"/graphql",
"/query",
"/feed",
"/rss",
"/_debug",
"/test",
"/testing",
"/tmp",
"/temp",
"/backup",
"/backups",
"/dump",
"/dumps",
"/sql",
"/db",
"/database",
"/dump.sql",
"/backup.sql",
"/db.sql",
"/backup.zip",
"/backup.tar.gz",
"/site.zip",
"/site.tar.gz",
"/source.zip",
"/src.zip",
"/upload",
"/uploads",
"/import",
"/export",
"/token",
"/tokens",
"/oauth",
"/oauth2",
"/openid",
"/jwks",
"/keys",
"/key",
"/private",
"/public",
]
# Realistic external referers. Most sessions arrive with a generic referer;
# a subset carries matching UTM tags on the landing URL.
PLAIN_REFERRERS: list[str] = [
"https://example.com/",
"https://somedomain.com/",
"https://another-site.org/",
"https://friend-site.net/",
]
# (referer origin, utm parameter dict) pairs used for tagged traffic.
TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
("https://chatgpt.com/", {"utm_source": "chatgpt.com"}),
("https://www.google.com/", {"utm_source": "google", "utm_medium": "organic"}),
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
(
"https://news.ycombinator.com/",
{"utm_source": "hackernews", "utm_medium": "referral"},
),
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
(
"https://www.producthunt.com/",
{"utm_source": "producthunt", "utm_medium": "referral"},
),
]
# Fraction of referered sessions that also carry UTM tags.
UTM_RATE = 0.25
# Innocent-looking paths that do not exist on a Pagerite site. Hitting many of
# these from a single IP is itself a telltale of a spray-and-pray scanner.
NORMAL_404_PATHS: list[str] = [
"/about",
"/about-us",
"/services",
"/products",
"/contact",
"/contact-us",
"/team",
"/careers",
"/jobs",
"/pricing",
"/features",
"/demo",
"/trial",
"/docs",
"/documentation",
"/api-docs",
"/support",
"/help",
"/faq",
"/knowledge-base",
"/terms",
"/terms-of-service",
"/privacy",
"/privacy-policy",
"/legal",
"/blog",
"/news",
"/articles",
"/press",
"/events",
"/webinars",
"/podcast",
"/videos",
"/resources",
"/whitepapers",
"/case-studies",
"/customers",
"/clients",
"/testimonials",
"/reviews",
"/partners",
"/integrations",
"/api-reference",
"/developers",
"/status",
"/security",
"/trust",
"/compliance",
"/gdpr",
"/ccpa",
"/sitemap",
"/archive",
"/tags",
"/categories",
"/search",
"/users",
"/accounts",
"/dashboard",
"/profile",
"/settings",
"/preferences",
"/notifications",
"/messages",
"/inbox",
"/calendar",
"/reports",
"/analytics",
"/billing",
"/invoice",
"/orders",
"/cart",
"/checkout",
"/store",
"/shop",
"/home",
"/main",
"/start",
"/welcome",
"/intro",
"/overview",
"/summary",
"/portfolio",
"/projects",
"/work",
"/solutions",
]
ABUSE_USER_AGENTS: list[str] = [
# Desktop browsers
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_5 like Mac OS X) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Mobile/15E148 Safari/604.1",
# Well-known crawlers / bots
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; "
"+http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; "
"+http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html)",
"Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36 "
"(compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)",
"Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com)",
"Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html)",
"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)",
# Social / service fetchers
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
"Twitterbot/1.0",
"LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)",
"Slackbot-LinkExpanding 1.0 (+https://api.slack.com/robots)",
"WhatsApp/2.23.20.0",
# Command-line / library clients
"curl/8.5.0",
"Wget/1.21.4 (linux-gnu)",
"python-requests/2.32.3",
"Go-http-client/1.1",
"Node.js/20.5.1",
]
def _random_ipv6_host(prefix: str) -> str:
"""Return a concrete address within an IPv6 /64 prefix.
The host part is generated randomly, mimicking a fresh OS privacy address.
The input prefix must end in ``::/64`` (e.g. ``2a02:8109:ac82:6f0c::/64``).
"""
if "/" not in prefix:
return prefix
base, mask = prefix.split("/")
if mask != "64":
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
if base.endswith("::"):
base = base[:-2]
host = ":".join(f"{random.randint(0, 0xFFFF):04x}" for _ in range(4))
return f"{base}:{host}"
def _concretize_ip(entry: str) -> str:
"""Return a concrete IP address; randomise the host part for IPv6 /64 prefixes."""
if ":" in entry and "/" in entry:
return _random_ipv6_host(entry)
return entry
class _SessionIP:
"""Stable IP for a browser session, with one optional mid-session rotation.
IPv6 prefixes get a fresh random host part; IPv4 addresses are swapped for
another address from the residential pool.
"""
def __init__(self, entry: str, pool: Sequence[str]):
self.entry = entry
self.pool = pool
self._value = _concretize_ip(entry)
def current(self) -> str:
return self._value
def rotate(self) -> None:
if ":" in self.entry and "/" in self.entry:
self._value = _random_ipv6_host(self.entry)
return
# IPv4: switch to another IPv4 address from the residential pool.
for _ in range(20):
candidate_entry = random.choice(self.pool)
if ":" in candidate_entry and "/" in candidate_entry:
continue
candidate = _concretize_ip(candidate_entry)
if candidate != self._value:
self._value = candidate
return
def _sleep(base: float, jitter: float) -> None:
time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
def _normalize_url(url: str) -> str:
"""Return a usable base URL, adding missing scheme/host/port parts.
- bare ``:PORT`` becomes ``http://localhost:PORT``
- missing scheme becomes ``http://``
- otherwise returned as-is
Raises ``ValueError`` when the result is not a valid http(s) URL.
"""
raw = url.strip()
if not raw:
raise ValueError("empty URL")
if raw.startswith(":"):
raw = f"http://localhost{raw}"
elif raw.isdigit():
raw = f"http://localhost:{raw}"
elif not raw.startswith(("http://", "https://")):
raw = f"http://{raw}"
parsed = urlparse(raw)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError(f"invalid URL: {url!r}")
return raw
def _poisson_wait(rate: float) -> float:
"""Return an exponential inter-arrival time for the given Poisson rate."""
if rate <= 0:
return 0.0
return random.expovariate(rate)
def _collect_links(page: Any, include_external: bool = False) -> list[dict[str, Any]]:
"""Return links from the current page, excluding the current page.
Internal links stay on the site; external links are real https URLs found
in the page content and are marked with ``external: true``.
"""
return page.evaluate(
"""(includeExternal) => {
const loc = new URL(location.href);
const out = [];
for (const a of document.querySelectorAll('a[href]')) {
try {
const u = new URL(a.href);
const rect = a.getBoundingClientRect();
const item = {
href: a.href,
text: (a.innerText || a.title || '').trim().slice(0, 60),
visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0),
};
if (u.origin === loc.origin
&& !u.pathname.startsWith('/_')
&& !u.pathname.startsWith('/auth')
&& u.pathname !== '/favicon.ico'
&& u.pathname !== loc.pathname) {
out.push(item);
} else if (includeExternal && u.protocol === 'https:' && u.origin !== loc.origin) {
out.push({ ...item, external: true });
}
} catch { /* ignore malformed hrefs */ }
}
return out;
}""",
include_external,
)
def _click_link(page: Any, link: dict[str, Any], timeout: float = 10.0) -> bool:
"""Click an internal link and wait for the client-side URL to change."""
start_url = page.url
try:
# Prefer Playwright's native click; fall back to a JS click if the
# locator cannot be resolved or times out.
try:
page.locator(f"a[href='{link['href']}']").first.click(timeout=2000)
except Exception: # noqa: BLE001
clicked = page.evaluate(
"""(href) => {
const a = Array.from(document.querySelectorAll('a[href]'))
.find(el => el.href === href);
if (a) { a.click(); return true; }
return false;
}""",
link["href"],
)
if not clicked:
return False
# Wait for the client-side navigation to update the URL.
deadline = time.time() + timeout
while time.time() < deadline:
if page.url != start_url:
return True
page.wait_for_timeout(100)
return False
except Exception as exc: # noqa: BLE001
logger.debug("click failed on %s: %s", link.get("href"), exc)
return False
def _run_browser_session(
base: str,
paths: Sequence[str],
profile: BrowserProfile,
session_index: int,
ip_entry: str,
) -> dict[str, Any]:
from playwright.sync_api import sync_playwright
MAX_CLICKS = 6
STAY = (2.0, 6.0)
HEADLESS = True
REFERER_RATE = 0.75
INCLUDE_EXTERNAL = True
ip_provider = _SessionIP(ip_entry, RESIDENTIAL_SOURCE_IPS)
ips_used: list[str] = [ip_provider.current()]
trail: list[str] = []
start_time = datetime.now(UTC)
try:
with sync_playwright() as p:
browser = p.chromium.launch(
headless=HEADLESS,
args=["--no-sandbox", "--disable-dev-shm-usage"],
)
extra_headers = {
"X-Forwarded-For": ip_provider.current(),
"Accept-Language": profile.accept_language,
}
# Most sessions arrive from an external origin; some are direct.
# A subset of referered sessions carries realistic UTM tags on the
# landing URL; the referer origin is paired with the UTM source.
tagged: dict[str, str] = {}
if random.random() < REFERER_RATE:
if random.random() < UTM_RATE:
referer, tagged = random.choice(TAGGED_REFERRERS)
else:
referer = random.choice(PLAIN_REFERRERS)
extra_headers["Referer"] = referer
context = browser.new_context(
user_agent=profile.user_agent,
viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
extra_http_headers=extra_headers,
)
page = context.new_page()
# Update X-Forwarded-For per request; the value stays stable unless we
# explicitly rotate it once mid-session.
def _route_handler(route, request):
headers = dict(request.headers)
headers["X-Forwarded-For"] = ip_provider.current()
ips_used.append(headers["X-Forwarded-For"])
route.continue_(headers=headers)
page.route("**/*", _route_handler)
# Pick one point during the session to emulate an IP rotation.
rotate_at = random.randint(0, MAX_CLICKS - 1) if MAX_CLICKS > 0 else -1
entry = random.choice(paths) if paths else "/"
landing = urljoin(base, entry)
if tagged:
sep = "&" if "?" in landing else "?"
landing += sep + urlencode(tagged)
page.goto(landing, wait_until="networkidle")
trail.append(page.url)
for click_idx in range(MAX_CLICKS):
_sleep(random.uniform(*STAY) / 2, 0.3)
if click_idx == rotate_at:
ip_provider.rotate()
ips_used.append(ip_provider.current())
logger.debug("rotated session IP to %s", ip_provider.current())
links = _collect_links(page, INCLUDE_EXTERNAL)
visible = [item for item in links if item.get("visible")]
if not visible:
visible = links
if not visible:
break
link = random.choice(visible)
ok = _click_link(page, link)
if not ok:
# Retry once with any link (sometimes visible calc misses nav).
alt = random.choice(links) if links else None
if alt and alt is not link:
ok = _click_link(page, alt)
if not ok:
break
if link.get("external"):
# Outbound navigation: the analytics exit ping is already
# in flight. Record the external URL and end the session.
trail.append(page.url)
_sleep(0.5, 0.2)
break
page.wait_for_load_state("networkidle")
trail.append(page.url)
_sleep(random.uniform(*STAY), 0.5)
browser.close()
return {
"profile": profile.name,
"entry": entry,
"ip": ips_used[0],
"ips_seen": len(set(ips_used)),
"pages": len(trail),
"trail": [urlparse(u).path or "/" for u in trail],
"duration": (datetime.now(UTC) - start_time).total_seconds(),
}
except Exception as exc: # noqa: BLE001
logger.warning("browser session failed: %s", exc)
return {"profile": profile.name, "error": str(exc), "trail": trail}
def _run_crawler_hit(
base: str,
paths: Sequence[str],
profile: CrawlerProfile,
) -> dict[str, Any]:
path = random.choice(paths) if paths else "/"
url = urljoin(base, path)
fake_ip = profile.ip
headers = {
"User-Agent": profile.user_agent,
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.5",
}
try:
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
r = client.get(url, headers=headers)
return {
"profile": profile.name,
"path": path,
"status": r.status_code,
"ip": fake_ip,
}
except Exception as exc: # noqa: BLE001
return {"profile": profile.name, "path": path, "error": str(exc)}
def _abuse_ua() -> str:
"""Return a randomized, syntactically valid user agent for an abuse scan."""
return random.choice(ABUSE_USER_AGENTS)
def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
"""Fire a burst of vulnerability probes from a single fake IP.
Scanner 0 randomises its user agent every request, scanner 1 uses a fixed
browser UA, and scanner 2 uses a fixed crawler UA.
"""
ip_entry = ABUSE_SOURCE_IPS[ip_index % len(ABUSE_SOURCE_IPS)]
if ":" in ip_entry and "/" in ip_entry:
fake_ip = _random_ipv6_host(ip_entry)
else:
fake_ip = ip_entry
MIN_HITS = 15
MAX_HITS = 25
total_hits = random.randint(MIN_HITS, MAX_HITS)
# Ensure the burst contains both telltales: suspicious paths and more
# than ten normal-looking 404 paths.
suspicious_count = max(5, total_hits // 3)
normal_count = total_hits - suspicious_count
if normal_count < 11:
normal_count = 11
suspicious_count = max(3, total_hits - normal_count)
paths = random.choices(SUSPICIOUS_PATHS, k=suspicious_count) + random.choices(
NORMAL_404_PATHS, k=normal_count
)
random.shuffle(paths)
ua_mode = ip_index % 3
if ua_mode == 0:
get_ua = _abuse_ua
elif ua_mode == 1:
def get_ua() -> str:
return BROWSER_PROFILES[0].user_agent
else:
def get_ua() -> str:
return CRAWLER_PROFILES[0].user_agent
scan_results: list[dict[str, Any]] = []
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
for path in paths:
headers = {
"User-Agent": get_ua(),
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": random.choice(
["en-US,en;q=0.9", "en-GB,en;q=0.8", "en;q=0.7"]
),
}
try:
r = client.get(urljoin(base, path), headers=headers)
scan_results.append(
{"path": path, "status": r.status_code, "ua": headers["User-Agent"]}
)
except Exception as exc: # noqa: BLE001
scan_results.append({"path": path, "error": str(exc)})
_sleep(0.15, 0.1)
return {
"scanner": ip_index + 1,
"ip": fake_ip,
"hits": len(scan_results),
"results": scan_results,
}
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Generate fake traffic for a Pagerite site.",
formatter_class=argparse.ArgumentDefaultsHelpFormatter,
)
parser.add_argument(
"url",
nargs="?",
default="http://localhost:8200",
help="Base URL of the Pagerite site (default: http://localhost:8200). "
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
"missing scheme defaults to http://.",
)
parser.add_argument(
"-t",
"--duration",
type=float,
default=60.0,
metavar="SECONDS",
help="Rough maximum time to generate traffic (0 runs one preset batch)",
)
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
return parser.parse_args(argv)
def main(argv: Sequence[str] | None = None) -> int:
args = _parse_args(argv)
if args.verbose:
logger.setLevel(logging.DEBUG)
try:
base = _normalize_url(args.url).rstrip("/")
except ValueError as exc:
logger.error("%s", exc)
return 2
# Discover content paths from the public page tree if we can.
paths: list[str] = []
try:
r = httpx.get(urljoin(base, "/_api/pages"), timeout=10.0)
if r.status_code == 200:
paths = [page["path"] for page in r.json() if page.get("has_content")]
except Exception as exc: # noqa: BLE001
logger.debug("could not fetch page list: %s", exc)
if not paths:
paths = ["/"]
logger.info(
"Generating fake traffic against %s (%d content paths, duration=%ss)",
base,
len(paths),
args.duration,
)
results: list[dict[str, Any]] = []
arrival_rate = 1.0
def _wait() -> None:
wait = _poisson_wait(arrival_rate)
logger.debug("waiting %.2fs before next session", wait)
time.sleep(wait)
if args.duration <= 0:
# One preset batch.
for i in range(5):
if i > 0:
_wait()
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(base, paths, profile, i, ip_entry)
results.append(result)
logger.debug(" trail: %s", result.get("trail", []))
for i in range(10):
if i > 0:
_wait()
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
results.append(result)
for i in range(3):
if i > 0:
_wait()
ip_entry = ABUSE_SOURCE_IPS[i % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, i)
results.append(result)
logger.debug(
" hits: %s", [r.get("path") for r in result.get("results", [])]
)
else:
deadline = time.time() + args.duration
session_index = 0
while time.time() < deadline:
if session_index > 0:
_wait()
phase = session_index % 3
if phase == 0:
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(
base, paths, profile, session_index, ip_entry
)
logger.debug(" trail: %s", result.get("trail", []))
elif phase == 1:
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
else:
ip_entry = ABUSE_SOURCE_IPS[session_index % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, session_index // 3)
logger.debug(
" hits: %s",
[r.get("path") for r in result.get("results", [])],
)
results.append(result)
session_index += 1
ok = sum(1 for r in results if "error" not in r)
logger.info("Done: %d/%d requests succeeded.", ok, len(results))
return 0 if ok == len(results) else 1
if __name__ == "__main__":
sys.exit(main())
+48 -21
View File
@@ -7,8 +7,8 @@ import sys
from contextlib import suppress from contextlib import suppress
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, Any, Self from typing import TYPE_CHECKING, Any, Self
from urllib.parse import urlsplit
import httpx
from buildutil import find_dev_tool, find_install_tool, logger from buildutil import find_dev_tool, find_install_tool, logger
from fastapi_vue.hostutil import parse_endpoint from fastapi_vue.hostutil import parse_endpoint
@@ -105,18 +105,47 @@ class ProcessGroup:
await p.wait() await p.wait()
async def http_get_server(url: str, timeout: float) -> str | None: # noqa: ASYNC109
"""GET url with plain asyncio streams, return the response Server header.
Returns an empty string when the server responds without a Server header,
and None when the server is unreachable or doesn't answer in time.
"""
parts = urlsplit(url)
host = parts.hostname or "localhost"
port = parts.port or (443 if parts.scheme == "https" else 80)
path = parts.path or "/"
if parts.query:
path += f"?{parts.query}"
try:
async with asyncio.timeout(timeout):
reader, writer = await asyncio.open_connection(host, port)
try:
writer.write(f"GET {path} HTTP/1.0\r\nHost: {host}\r\n\r\n".encode())
await writer.drain()
data = await reader.readuntil(b"\r\n\r\n")
finally:
writer.close()
except OSError, EOFError, ValueError, TimeoutError:
return None
for line in data.decode("latin-1").split("\r\n"):
if line.lower().startswith("server:"):
return line.split(":", 1)[1].strip()
return ""
async def check_ports_free(*urls: str) -> None: async def check_ports_free(*urls: str) -> None:
"""Verify URLs are not responding (ports are free). Raise SystemExit if any respond.""" """Verify URLs are not responding (ports are free). Raise SystemExit if any respond."""
async def check(client: httpx.AsyncClient, url: str) -> None: async def check(url: str) -> None:
with suppress(httpx.RequestError): server = await http_get_server(url, timeout=0.1)
res = await client.get(url, timeout=0.1) if server is not None:
server = res.headers.get("server", "server") logger.warning(
logger.warning("Conflicting %s already running at %s", server, url) "Conflicting %s already running at %s", server or "server", url
)
raise SystemExit(1) raise SystemExit(1)
async with httpx.AsyncClient() as client: await asyncio.gather(*[check(url) for url in urls])
await asyncio.gather(*[check(client, url) for url in urls])
async def ready(url: str, path: str = "", max_attempts: int = 50) -> None: async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
@@ -128,18 +157,14 @@ async def ready(url: str, path: str = "", max_attempts: int = 50) -> None:
if not path: if not path:
return return
async with httpx.AsyncClient() as client: for attempt in range(max_attempts):
for attempt in range(max_attempts): if await http_get_server(f"{url}{path}", timeout=1.0) is not None:
try: logger.info("✓ Backend ready!")
await client.get(f"{url}{path}", timeout=1.0) return
except httpx.RequestError: if attempt == max_attempts - 1:
if attempt == max_attempts - 1: logger.warning("Backend didn't start in time")
logger.warning("Backend didn't start in time") raise SystemExit(1)
raise SystemExit(1) from None await asyncio.sleep(0.1)
await asyncio.sleep(0.1)
else:
logger.info("✓ Backend ready!")
return
def setup_vite( def setup_vite(
@@ -222,5 +247,7 @@ def setup_cli(
host = endpoints[0]["host"] host = endpoints[0]["host"]
port = endpoints[0]["port"] port = endpoints[0]["port"]
cmd = [cli, f"--listen={host}:{port}"] # Run the package as a module with the current interpreter, instead of
# relying on a PATH-installed CLI entry point.
cmd = [sys.executable, "-m", cli, f"--listen={host}:{port}"]
return f"http://{host}:{port}", cmd return f"http://{host}:{port}", cmd

Some files were not shown because too many files have changed in this diff Show More