Compare commits

...
82 Commits
Author SHA1 Message Date
LeoVasanko 1d0bc3f59d Renewed analytics format. The hide flag moves to client, and we still track but hide more robustly, to avoid noise from admins checking out their own site. 2026-08-27 16:17:02 +00:00
LeoVasanko 60d5155fd2 Don't upper case lower headings. 2026-08-27 16:14:04 +00:00
LeoVasanko 3518ac9ac7 Fixes to markdown extensions/styling. Sort crawlers most recent first. 2026-08-27 01:27:49 +00:00
LeoVasanko 5c2433b766 Correct font scaling by x size. Fix code receiving twice a smaller font size. 2026-08-27 00:15:47 +00:00
LeoVasanko 2959f971bc Implement support for {} attrs on code blocks (line right after the closing fence) and containers (::: aside {...}). 2026-08-27 00:08:51 +00:00
LeoVasanko 8ee22de060 Many more Markdown extensions and formatting improvements. 2026-08-26 23:51:05 +00:00
LeoVasanko 9be1491f0e Improved mobile layout. Banner section gets smaller and editor goes full screen. 2026-08-26 18:44:17 +00:00
LeoVasanko 29c1dc64ed Transition graph shows per article read times in minutes and seconds. 2026-08-26 18:27:09 +00:00
LeoVasanko a27958117f Maintain nitro orange line in theme.css rather than banner.css so that banner changes don't remove it. 2026-08-26 17:40:23 +00:00
LeoVasanko 56fd5f854d Selection color from base theme. 2026-08-26 17:03:01 +00:00
LeoVasanko 932aee4404 Fixes to eyes banner positioning. 2026-08-26 16:20:58 +00:00
LeoVasanko ca1c0d5a6f Summer theme gets seamless transition between banner and page. Banners updated with support for this mode. 2026-08-26 15:58:02 +00:00
LeoVasanko 9cbc4ab59d Smoother scroll effects on banners using --pry. 2026-08-26 15:09:29 +00:00
LeoVasanko 7aa2abbf8f Analytics pings with query args, cleanup. 2026-08-26 14:53:09 +00:00
LeoVasanko eddb3f22f3 Adjusted connector thickness and bead animations to better deal with highly varying rates. 2026-08-26 14:34:10 +00:00
LeoVasanko 1702281dbf Fix edge condition for gaussian smoothing. 2026-08-26 14:00:29 +00:00
LeoVasanko 38b894a789 Layout transition lags fixed with better editor panel handling. This was causing .wide sections visibly lag behind on viewport width changes. 2026-08-26 13:47:51 +00:00
LeoVasanko 06409cdaac Drop page caches on edits that may cause changes to navigation, theming etc. 2026-08-26 13:19:17 +00:00
LeoVasanko 7ba6094360 Proper cube transition background colors via base CSS mixing in theme bg. 2026-08-26 13:01:42 +00:00
LeoVasanko c59ab2f058 Fix editor-page scroll syncing issues. 2026-08-26 02:29:33 +00:00
LeoVasanko 078e80591b Maintain normal column layout while editing. 2026-08-26 02:29:14 +00:00
LeoVasanko 6563eebdba Improve list bullet positioning. 2026-08-25 22:49:02 +00:00
LeoVasanko cac74ebb6b Remove uvicorn server header. 2026-08-25 22:40:36 +00:00
LeoVasanko 8b1bb6f994 Avoid column break right after heading. 2026-08-25 22:31:52 +00:00
LeoVasanko 3173a113b6 Improved theme banner consistency across themes. 2026-08-25 22:29:28 +00:00
LeoVasanko a21f919601 More robust admin check, only adding edit pens after probe completes. 2026-08-24 21:51:35 +00:00
LeoVasanko 158b5f1961 Editor panel layout fixes, use body scroll and bidirectional sync with editor. 2026-08-24 21:41:33 +00:00
LeoVasanko f9ce0a4f3b Smarter initial analytics page when there is no visitor data yet. 2026-08-24 21:13:20 +00:00
LeoVasanko 2d595f8c15 Reduce #brand to actual size, avoid clicks on empty banner space touching it. 2026-08-24 20:58:50 +00:00
LeoVasanko b1fc8e24d7 Eyes banner follows taps not just mouse. 2026-08-24 20:54:50 +00:00
LeoVasanko c5cf799f68 Better nav layout for portrait phones. 2026-08-24 20:43:18 +00:00
LeoVasanko a5edc7b3b6 Fix crawler misclassification from /_a and orphan counts on visit scrub
Two analytics corrections verified against the production capture:

- pagerite.js suppressed pings with fr == '/_a', but fetch-navigation
  away from the analytics page had already GET-ed the target without the
  preload header; the orphaned pending hit then flushed to the crawler
  list, classifying a real user as a crawler. Navigations away from /_a
  now ping normally (the server rejects /_a as a target regardless, and
  admin noise is already handled by hide=1).

- _remove_visit only reversed the visit's creation counts, leaving
  views/transitions from later pings behind as orphans on the graph with
  no matching row in the visitor table. An in-memory per-visit count log
  now tracks every count event, so an admin hide=1 scrub reverses the
  visit completely.
2026-08-24 20:18:13 +00:00
LeoVasanko 09ebc63690 Record response status per path; mark 404 trails red in the viewer
Document GETs now stash their status (200/404) in a pending table,
consumed by the matching ping: visits gain a per-path statuses map and
crawler hits a status field. Trail links with a 404 status render in
red with the status code in the tooltip, alongside the read time.
2026-08-24 19:43:23 +00:00
LeoVasanko 2868843028 Fix inverted client filter in admin hide ping
The hide=1 branch kept the admin's own pending crawler hits (==) instead
of discarding them (!=), so an admin's document GETs flushed to the
crawler list 10s later while every other client's pending hits were
wrongly dropped. This is why ordinary admin browsers showed up as
crawlers.
2026-08-24 19:40:12 +00:00
LeoVasanko 6fcebaea3f Transition map: svg-scaled fonts, border-clipped pill text, tight crop
- Fonts scale with the svg instead of the --u constant-screen-size
  compensation (ResizeObserver machinery removed)
- Larger node text (slug 19px, count 15px)
- Pill labels no longer ellipsis-truncated: text is clipped at the pill
  border via per-node clipPaths; captions center when they fit and anchor
  left on overflow so the title's beginning survives; count lines stay
  centered
- Bounding box crops to pill half extents plus the ribbon halo instead of
  the diagonal radius, removing the large top/bottom margins
2026-08-24 18:35:36 +00:00
LeoVasanko 4a48d08a19 Analytics layout: larger charts with natural-width cap, one-line totals
- Charts grow to a larger intrinsic size (1052x174) and never upscale
  past it; centered with equal side margins above the cap, full width
  below, svg always within page bounds; overflow visible so wider fonts
  don't clip at the viewBox edge
- Totals row aligns its left edge with the charts and shrinks (gap first,
  then font) via container units to always stay on one line
2026-08-24 18:21:58 +00:00
LeoVasanko e3e29251ca Transition map: cull invisible connectors, stable beads, lane labels
- Cull connections whose thin middle would render below ~0.8px
  (MIN_WMID); drop external source/exit nodes whose connectors are
  all culled, while site page nodes always stay
- Bead simulation persists across data reloads: emitters keyed per edge
  direction, beads tracked by progress, so unrelated count changes no
  longer reshuffle bead positions
- Bead speed relative to span length: constant 1.5s traversal per edge
- Top lane labeled with a house icon; all lane labels left-aligned just
  past the source pill (half height on near-vertical branch lanes), with
  guides running to the lane end so long slugs are never truncated
2026-08-24 17:58:02 +00:00
LeoVasanko 8affc41289 Tighter analytics chart chrome
- Unified chart text at 11px system-ui; fixed size independent of theme font
- Day view y axis reads "visits / 5 min" / "views / 5 min"
- Left margin and y tick spacing tightened (MARGIN_L 56 -> 40)
- Gap between visits and views charts removed
- Legend repositioned for the larger font
2026-08-24 17:26:30 +00:00
LeoVasanko 2de4717230 Fix week overlay alignment, in-plot ISO week legend, shorter charts
- weeklySeries shifts overlaid weeks onto the current week's time axis so
  they overlay inside the plot instead of overflowing left; oldest weeks
  paint first, current week on top
- Legend moved inside the visits chart's top right: current ISO week in
  accent, past weeks as a single muted "Week M" / "Week M–N" specimen
- Past week curves use the muted color instead of faded accent
- Chart height reduced ~30% (180 -> 126)
2026-08-24 17:03:49 +00:00
LeoVasanko 563e8fcaf2 Rework analytics chart scaling; self-contained SVG charts
- Charts render as single SVGs with axis labels inside the viewBox,
  replacing the stretched plot + HTML overlay labels
- Rolling ranges end at now, t0 aligned to UTC day; bucket size follows
  the window (6h up to 31 days) so "all" at its 30-day minimum renders
  identically to "month"
- X labels always centered on their true position; no edge-align shifting
- rangeWindow simplified to rolling spans ending at now
- TransitionGraph "all" visual scale floored at the 30-day plot minimum
2026-08-24 12:31:25 +00:00
LeoVasanko 921a5484a2 Fixed-sigma smoothing of traffic history plots. 2026-08-24 05:54:43 +00:00
LeoVasanko 0e2e52fa45 Use last 24h/7d/30d/365d/all analytics data. Previously some fields were unfiltered and weekly view was based on calendar weeks. 2026-08-24 05:42:32 +00:00
LeoVasanko 075848f782 Crawlers should include all sorts of spiders along with bots and googleother. 2026-08-24 05:28:05 +00:00
LeoVasanko 7b8899af92 Transition map: bounded node scaling, concentric branch lanes, exit row at bottom 2026-08-24 04:47:06 +00:00
LeoVasanko 6d2ae104d7 Fix analytics classification: ignore bot-UA pings, skip preload GETs.
JS-running crawlers (Googlebot, GoogleOther, Applebot) execute pagerite.js
and send navigation pings, registering as visitors. Pings whose User-Agent
matches _is_bot_ua (any "bot" token plus listed exceptions) are now
ignored, so their document GETs flush to the crawler list as intended. No
source verification: a spoofed bot UA merely lands in the crawler stats,
and path-based abuse classification catches scanners regardless.

Idle-time link preloads from pagerite.js were queued as pending crawler
hits and flushed to the crawler list whenever the user navigated more than
10s later, so real visitors' subpage loads showed up as crawler hits.
Preload fetches now carry an x-pagerite-preload header and the document
GET handler skips tracking for them; the ping sent on actual navigation
does the counting.
2026-08-22 18:23:38 +00:00
LeoVasanko b7fc543a83 Transition graph layout follows navigation. 2026-08-22 18:10:44 +00:00
LeoVasanko b868033ddc Pill shaped nodes 2026-08-22 15:48:33 +00:00
LeoVasanko 8aad64cced Analytics layout update, larger, consistent text sizing. 2026-08-22 14:45:56 +00:00
LeoVasanko 319163ee7e Page caching and zstd compression. Avoid useless fetching. Mobile layouts of navigation menus improved. 2026-08-22 14:05:12 +00:00
LeoVasanko 87b16b7144 Implement /robots.txt and /sitemap.xml. Update dev proxy to all-by-default. 2026-08-22 12:13:17 +00:00
LeoVasanko 20ae6501f2 Add dynamic /sitemap.xml and /robots.txt endpoints 2026-08-22 12:01:18 +00:00
LeoVasanko a16fe88114 Auto select day if less than 24h data for new sites. 2026-08-22 01:51:36 +00:00
LeoVasanko c77598adc7 Fine tuning date formatting. 2026-08-22 00:09:57 +00:00
LeoVasanko 29f8fac013 Slightly prettier analytics URL 2026-08-21 23:58:00 +00:00
LeoVasanko 6199e5a69e Support for UTM tags in transition graph as source sites. 2026-08-21 23:50:01 +00:00
LeoVasanko 375b4b6bdb analytics: unify visitor cell across visits, crawlers and abuse tables 2026-08-21 23:32:46 +00:00
LeoVasanko fdb3e42d6f analytics: shared Client struct, grouped abuse paths, unified visitor cell 2026-08-21 23:16:15 +00:00
LeoVasanko 51a6a16221 Neater abuse table formatting. 2026-08-21 22:31:37 +00:00
LeoVasanko 0798e24d24 Desaturated house emojis 2026-08-21 22:03:55 +00:00
LeoVasanko 3be2d08ac9 SI formatting of large visitor numbers. 2026-08-21 21:36:57 +00:00
LeoVasanko b7d5b23ae6 Cleaner formatting of utm tags in visitor table. 2026-08-21 21:22:16 +00:00
LeoVasanko 6eaa1c1a8b analytics: 24h day view with bar chart for precise realtime stats. Tables redesigned with cleaner layout. Tracking article read times. Adjust connection graph visualizations by time range. Other cleanup and supporting systems. 2026-08-21 20:14:04 +00:00
LeoVasanko f341d22aa0 Improved fake traffic generation with abuse bots, utm tags etc. 2026-08-21 20:10:41 +00:00
LeoVasanko 1a479ceb24 Add --dbip CLI flag to auto-download/update the DB-IP MMDB database.
Downloads the latest dbip-city-lite-YYYY-MM.mmdb.gz before starting the
server, skipping when the local database is current, falling back to the
previous month on 404, and removing older databases after an update.
Promotes httpx to a runtime dependency.
2026-08-21 03:09:04 +00:00
LeoVasanko ff553d018a Default scheme, host and port for fake_traffic script. 2026-08-21 02:52:53 +00:00
LeoVasanko c807d48a13 Add more external content in seed data. 2026-08-21 02:50:36 +00:00
LeoVasanko 9c383c1c8b Change default port mapping to 8100/8200/8210 (prod/vite/dev). Vite gets different port to avoid caching problems when switching between it and prod. 2026-08-21 02:49:18 +00:00
LeoVasanko 462e995adc Add external link (referer/outgoing) display on connection graph. 2026-08-21 02:44:54 +00:00
LeoVasanko ea069b98da Fix analytics app not mounting on fetch-navigation to /_a
load() queried the live document for the pagerite:analytics-src meta,
but the swap never touches <head> — the meta only exists in the fetched
doc, so the app never mounted unless /_a was loaded directly. Also cache
the fetched HTML so the post-swap preload doesn't re-GET the page we
just navigated to.
2026-08-21 01:53:39 +00:00
LeoVasanko deb5419c47 analytics improvements:
- keep visitor charts y-axis minimum range at 10
- keep 'all' chart x-axis minimum span at 30 days
- group crawler hits by (ip, ua) and list top pages visited, show crawler page load counts as N× prefix
- store and display geoip city, keep geoip country overwrite
- stream live updates over WebSocket /_api/ws/analytics
- include family ring arcs in transition map crop bounds
- remove top UA summary, limit crawlers to 10 and visits to 20
- human-readable relative timestamps with UTC tooltip
2026-08-21 01:36:58 +00:00
LeoVasanko 242b62784c Add fake traffic generator script
Uses Playwright to drive Chromium through real internal link clicks,
so pagerite.js analytics pings create normal visits. Also fires HTTP
GETs with crawler user-agents to record crawler hits.

Features:
- script-local deps via uv add --script (playwright, httpx)
- rotating pool of real public IPs via X-Forwarded-For
- Poisson inter-arrival delays between sessions/hits
- configurable browsers, crawlers, clicks, and dwell time
2026-08-20 23:37:02 +00:00
LeoVasanko f78229bd30 Refactor analytics to /_a instead of under article pages. 2026-08-20 23:16:51 +00:00
LeoVasanko 7f4bc8efa4 Ignore devserver health probe in analytics tracking
The devserver polls /?from=devserver.py to check backend readiness.
Without this exclusion each poll is recorded as a crawler hit. Only
exclude the exact case: front page, that query string, and 127.0.0.1,
so remote visitors cannot hide traffic by copying the parameter.
2026-08-20 22:16:11 +00:00
LeoVasanko 3100010335 Transition map: count-scaled edges, bead flows, external links
- Edge widths grow logarithmically with the connection count (~1 px at
  a single count, uncapped); connections below 1% of total traffic are
  pruned, bounding the graph to ~100 edges.
- Beads: per-direction flows emitted at a rate linear in the count,
  each bead simulated independently in JS (no in-flight limit), offset
  onto right-hand lanes so opposing flows don't collide, running under
  the node circles with a glow.
- External links: referer origins as a node row above the map, exit
  origins fanned outwards from their source page.
- Transitions are now stored per 5-minute bucket (sparse
  from -> to -> bucket -> count) so the graph filters by time range
  like the other series; legacy analytics files are discarded.
2026-08-20 22:06:29 +00:00
LeoVasanko 7556bb7f4f analytics: add crawler tracking, pretty UA/IP display and copy-to-clipboard 2026-08-20 20:31:04 +00:00
LeoVasanko 1a1a21712e Add country flags. 2026-08-20 19:47:56 +00:00
LeoVasanko f0e6162f02 Extended analytics data collection. 2026-08-20 19:40:05 +00:00
LeoVasanko b4e8fad090 Implement analytics feature
Add server-side visit analytics collection, a public-page ping endpoint,
and a full-screen AnalyticsView for admins.

Backend:
- Add pagerite/analytics.py: Analytics/Visit model, Store, and persistence
- Wire /_a ping endpoint and GET /_api/analytics into pagerite/app.py

Frontend:
- Add full-screen AnalyticsView with visitor charts and transition map
- Add VisitorCharts and TransitionGraph subcomponents
- Add analytics JS helpers in frontend/src/analytics/
- Send navigation pings from frontend/src/pagerite.js
- Mount AnalyticsView from frontend/src/main.js
- Document the feature in docs/analytics.md and update AGENTS.md
2026-08-20 18:43:57 +00:00
LeoVasanko 11f8de2df5 Fix pretty scrollbars not appearing in production. 2026-08-19 16:33:44 +00:00
LeoVasanko 81f08e7760 Updated docs 2026-08-19 16:11:42 +00:00
LeoVasanko 1d65a57fdf SEO/social meta for content pages; full-height site editor
- views.py: description, canonical, Open Graph and twitter:card tags
  from heuristics over the rendered article — first paragraph as
  description, share image prefers a {.hero} image, then first raster,
  then first SVG; first <video> becomes og:video; published/modified
  times from the node. Absolute URLs from the request base.
- SiteEditor: panel fills the full window height; the brand-HTML and
  custom-CSS CodeMirror windows grow to share leftover space equally
  instead of fixed max-heights.
2026-08-19 15:52:24 +00:00
LeoVasanko e175b39f35 Full height site editor panel. 2026-08-19 02:19:30 +00:00
54 changed files with 7924 additions and 957 deletions
+2
View File
@@ -2,6 +2,8 @@
!.gitignore !.gitignore
*.lock *.lock
*.kantadb *.kantadb
pagerite.analytics.json
dbip-*.mmdb*
/pagerite/frontend-build /pagerite/frontend-build
package-lock.json package-lock.json
+27 -306
View File
@@ -5,297 +5,38 @@
Please instead ask the user to see from dev tools what you need, e.g. to look up something in DOM or log. Use console.log for debugging where needed (and otherwise for permanently kept useful messages in the app). Please instead ask the user to see from dev tools what you need, e.g. to look up something in DOM or log. Use console.log for debugging where needed (and otherwise for permanently kept useful messages in the app).
## What this is
Pagerite: a single-user CMS/blog. FastAPI serves HTML rendered in Python
with html5tagger; content is persisted in a kanta database and rendered on
the fly per request. Vue is used only for interactive bits (editing tools),
not for the public pages. See `docs/design-principles.md` for the design.
## Layout ## Layout
- `pagerite/` — the Python backend package (hatchling build target). Pagerite is a CMS. See `docs` for the full design and implementation details. Key files for code changes:
- Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed)
- Dev mode `scripts/devserver.py` (which the user mostly uses for auto reloads, no build needed) - `pagerite/` — Python backend package (hatchling build target).
- Avoid running the server yourself, ask the user to test - `app.py` — FastAPI app and route registration.
- `app.py`the FastAPI app. FastAPI's built-in API docs are disabled - `data.py`msgspec Structs for the kanta database.
(`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to - `markdown.py` — markdown-it-py renderer.
our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are - `views.py` — shared page layout and rendering.
registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue - `seed.py` — demo content, written only on first database creation.
inserts its file routes at the position where - `analytics.py` — visit analytics collection (see `docs/analytics.md`).
`route()` was called (during `load()` in the lifespan), so anything - `frontend/src/` — Vue editor and public-page JS entries.
defined earlier wins. The one exception is the content catch-all - `main.js` — Vue editor app entry.
`/{path:path}`, registered AFTER `frontend.route()` so that built - `analytics-main.js` — analytics page entry (mounts `AnalyticsView` at `/_a`).
frontend assets still take priority over content slugs. The `Frontend` - `pagerite.js` — public page entry.
is constructed with `spa=False` explicitly: it only serves the built - `assets/` — base CSS, Pygments styles, fonts.
files without a catch-all. The build mirrors the URL space — hashed - `scripts/devserver.py` — dev server with auto reload (the user mostly uses this; avoid running the server yourself, ask the user to test).
immutable assets under `/_assets/`, `favicon.ico` at the site root —
and an `index.html` in the build would become a `/` route, so leave it Server run by CLI entry point `uv run pagerite` (no auto reloads, build needed). Dev mode is `scripts/devserver.py` (auto reloads, no build needed).
out of the build to keep `/` ours.
- `data.py` — msgspec Structs for the kanta database. The site structure
is a tree: `Data.menu` maps top-level slugs to `Node`s, each with
`children` keyed by slug — the URL path is the slug chain. The front
page is whichever top-level node has slug "" (parallel to the other
main level pages, not their parent); it cannot have children, and
renaming its slug away leaves no front page ("/" redirects to the
first nav item). `Node.content` is
the Markdown page, or None for a pure category label whose URL renders
a placeholder page (while nav links to it point at its first child);
every label's title and slug are editable. Siblings order by the fractional `Node.order` key: a moved
item gets a fresh key relative to its new siblings, all others keep
theirs. `resolve`/`find_slot` walk the tree by path; moves are slot
detach/attach carrying the whole subtree. Legacy flat `Data.pages`
(pre-tree databases) migrates into `menu` on startup. The app owns
the `Data` object; reads are plain attribute access, writes in
`kanta.transaction(...)`.
`Data.files` is a content-addressed store (blake3[:12] + extension)
mapping file names to bytes, served at `/_f/{name}` with immutable
caching; pages reference files by absolute `/_f/` URLs so hierarchy
moves never break them. `Node.banner` is a raw trusted HTML snippet
for the header banner (img, styled div, canvas+script...); empty
inherits from the node's ancestors (front page last). It is rendered
AFTER the banner design's artwork, so author code (e.g. a `<style>`
override) always wins over the design's own styles.
`Node.banner_design` picks a banner design: a theme folder name whose
`banner.css` styles it and whose `banner.html` (arbitrary markup:
canvas + style + script) or `banner.svg` supplies the inline artwork
(wrapped in `div[data-design]`); "" = explicitly no design, None =
inherit (nearest ancestor, front page last, then the active theme's
own design if it ships banner.css/banner.svg/banner.html). The design's banner.css
is linked in `<head>` (id `pagerite-banner`) between the theme and the
custom CSS.
`Data.version` is bumped on every write
and embedded in page ETags so nav-affecting changes invalidate caches.
`Data.brand` is the site name (header link + `<title>` suffix), editable
in the site editor via `/_api/settings`; empty = no header link and
no `<title>` suffix. `Data.brand_html` is raw trusted HTML replacing the
brand link entirely (rendered in a `#brand` div on top of the banner,
next to the nav) — site-wide, not per-page like banners; edited in the
site editor with image/video upload into `Data.files`. `Data.theme` is
the active theme name (empty =
none/base only); themes are folders in `pagerite/themes/{name}`
containing `theme.css` and/or `banner.css` (+ `banner.svg` artwork and
any extra assets the CSS references, like summer's `grass.svg`),
served by the backend at `/_themes/{name}/...` — read from disk per
request (etag by mtime), never built, so on-disk edits show on the
next page load even in prod. The theme selector and banner-design
selector enumerate these folders via `GET /_api/settings`.
`Data.custom_css` is raw trusted CSS injected inline in every page
`<head>` (id `pagerite-user`) and swapped during fetch-navigation;
editable in the site editor. Font picks (heading/body/brand) in the
site editor are stored as plain `:root` rows in `custom_css`
(`--font-body: var(--font-source-sans);` format — parsed out and
rewritten on change, the `:root` block added/removed as needed),
referencing the per-family variables (`--font-source-sans` etc.) from
pagerite.css;
the base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
`Data.favicon` names a file in the content-addressed `files` store,
uploaded/cleared in the site editor via `PUT`/`DELETE
/_api/settings/favicon`; when set it is linked as `<link rel="icon">`
on every page, otherwise browsers fall back to the build's
`/favicon.ico` by convention.
- `markdown.py` — markdown-it-py renderer (html passthrough + attrs,
footnote, deflist, tasklists, admon, gfm_autolink, sub/superscript
plugins; typographer + breaks on). Custom
image rule: relative srcs resolve against the page path; an image
standing alone in its paragraph becomes a figure (captioned when
titled), while inline-with-text images and raw <img> HTML stay plain.
A `{dates}` line expands to the article's
published/updated dateline (`p.dateline`, from `Node.created`/
`modified`; left literal in previews of unsaved pages).
- `views.py` — the shared page layout as an html5tagger `Template` with
placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav
rendering straight from the `Data.menu` tree (siblings sorted by
`Node.order`; nav links to content-less labels point at their first
child via `first_leaf`, the first published descendant with content),
and page/404 rendering. If the markdown contains its own h1, the page title
is NOT rendered as an additional h1 (it still supplies <title> and nav
labels). The navbar holds
top-level items only; the current section's subitems go to a left
`#sidebar` as a nested list (the section's direct children plain,
deeper levels indented with article-list-style markers), which is
rendered when the section offers at least two
published items, or exactly one while viewing anything other than that
only page — the section index, a 404, a grandchild (so those pages can
reach the child), and also on that only page itself when it has
published children of its own; no aside element at all on the front
page, leaf
pages and the sole childless page of a one-page section. Also,
category labels are nodes without content — None *or* empty markdown —
and their nav links point at their first child page. Dynamic regions have stable ids
(`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps
(`#sidebar` may be absent on either side of a swap).
- `seed.py` — demo content written only when the database is first
created, via a `@kanta.bootstrap` handler in `app.py`.
- `frontend/src/` — the Vue editor and public-page entries.
- `main.js` — Vue editor app entry, mounts the tabbed EditorShell.
- `pagerite.js` — public page entry; runs fetch-navigation (backed by an
in-memory page cache: every visible internal link — and the current
page — is fetched once at load, clicks are then served from JS with no
fetch, and the editors' `loadPlain` keeps the cache current via a
`pagerite:page-fetched` event; articles are `cache-control: no-cache`
on the wire), scroll-reveal,
OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars
that never reserve layout space or shift the page when appearing;
native scroll APIs like `window.scrollTo` keep working; themed via the
`--os-*` variables in pagerite.css),
brand shrink-to-fit (the themed size is the maximum; JS reduces the
font-size so a long brand or narrow viewport still fits one line),
code copy buttons, and the auth check. It first probes `GET /auth/api/settings`
to detect whether Paskia SSO is available, then `GET /_api/settings` to
learn the current session's admin status. The same reverse proxy that
gates `/_api` returns 401 for anonymous users, 403 for users without
the admin permission, and 200 for admins. When Paskia is detected, a
🔑 login link (anonymous) or 🔐 profile link (logged in) is shown in
the banner corner; both are plain `<a href="/auth/">` links (Paskia
does not support being iframed, so we navigate normally), and a
`pageshow` handler re-probes auth when history navigation restores a
cached page. Admins also get the 🖊️ page/banner edit pens
and a ⚙️ site-settings pen (asset URLs from the
`pagerite:editor-src`/`-css` meta tags). If no Paskia SSO is
detected (dev/no proxy), editing is left open. Pages themselves render
identically for everyone; the real gate is the auth proxy in front of
all of `/_api`. The backend links the stylesheets in a fixed order —
base (Vite build), theme, banner design, custom CSS last — each with
a stable id so the site editor can swap them in place.
- `assets/` — shared styles and data files built by Vite and served hashed
under `/_assets/`: `pagerite.css` (base layout + conservative
variables), `pygments.css`,
and `fonts/` (self-hosted Source
Sans 3/Source Serif 4/Fraunces/Literata/Cormorant/Playfair
Display/Inter/Montserrat/Fira Code/Cause/Exo 2/New Rocker
variable woff2). The `::view-transition*` block at the end of `pagerite.css` (from
termotohtori.fi) is fragile — do not tweak. Themes are NOT built:
`pagerite/themes/{name}/theme.css` (theme overrides and font picks:
`purple` = dark dusk palette with Fraunces/Literata and a tilted
oversized gradient brand; `corporate` = light-first with automatic
`prefers-color-scheme` dark mode, Montserrat/Inter and a huge solid
brand; `nitro` = racing/HUD style following `prefers-color-scheme`
(warm light-grey page, deep violet in dark), Montserrat/Literata,
black as an accent only, a straight orange blade under the banner, and
an orange racing-tab nav clipped with a bezier `shape()`; `summer` =
light playful meadow, one palette sampled from its illustrated
`banner.svg` (sky/grass/sun/flower pink), Fraunces/Literata, a tilted
gradient brand, flower bullets, and a layered-parallax banner (sun
rises, clouds drift, nearer hills move less) with idle animations
(swaying flowers, floating clouds, breathing sun glow) wrapped in
`prefers-reduced-motion: no-preference`); standalone banner designs
(no theme.css) ship as `eyes` (a canvas critter in the grass) and
`stars` (a drifting starfield)) and the
companion `banner.css` banner designs are served by the backend.
- Vite builds ES-module `.js` outputs; the backend renders `<script
type="module">` for them (module scripts defer by default).
- The database file is `pagerite.kantadb` in the cwd (`PAGERITE_DB`
overrides); gitignored. Do not delete it without asking.
- `scripts/fastapi-vue/` — helper scripts from the fastapi-vue template
(build hook etc.), do not edit.
- `frontend/` — the Vue editor as a single tabbed `EditorShell.vue` mounted
in a host div created inside the static document. The shell hosts four
kept-alive tabs (ordered site-wide first — site, structure — then, after a
visual break, the per-page tabs — article, banner): `PageEditor.vue`
(CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`,
previewing into the visible article; editor scroll drives the article
scroll — while any editor is open the window scroll is locked
(`body.editing`), the panel exactly fills the available window height, and
only `#main` scrolls; a format bar offers Markdown helpers — bold/italic/code/link/
table/image upload, with Ctrl/Cmd-B/I/S bindings — for the
hard-to-remember syntax) — edits content and
title only, never the path — `BannerEditor.vue`
(per-page banner HTML + banner design selector, previewed into
`#page-banner`), `SiteEditor.vue` (site brand + optional custom brand
HTML with image/video upload + theme selector + font picker + favicon
upload — clicking the preview tile picks a new one — +
site-wide custom CSS, CSS injected into
`<head id="pagerite-user">`), and `StructureEditor.vue` (the
vue-draggable structure tree with
always-editable title/slug inputs per row). Media uploads everywhere use
🖼️ icon buttons (pasting into the editor works too). The article, banner and
site-settings pens are shorthands that open the shell on the matching tab;
once open, clicking a pen switches tabs (and retargets the editors to the
current page) instead of closing/remounting. The ✕ in the tab bar closes
the shell (Escape too); tabs have no close buttons of their own. Closing
only HIDES the shell — the Vue app stays mounted, so page-editor state
(unsaved text included) survives until a real page reload; saving there is
explicit (💾/Ctrl+S) and refreshes the page regions in place. Admin panels
never reload the page. In-place
page re-rendering shared by the banner/site/structure tabs lives in
`swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic
regions, replaceState). Placeholder texts are reserved for showing the
actual default in effect when a field is left empty (e.g. the pending
row's slug derived from its title); labels and help are real elements or
tooltips, never placeholders.
Everything saves immediately as you edit (brand/title/CSS debounced,
slug on commit since it renames the path), theme change swaps the
stylesheet in place, tree rows navigate in place without transitions when
focused, and the front page is a root-only row whose empty slug is
editable like any other. Every
non-empty list (and the root) ends with a non-draggable footer row
(vuedraggable `#footer` slot): clicking it starts a new pending page at
that level (its slug placeholder shows the slug derived live from the
title being typed), and while dragging it is the list's "end of list" drop
target. Committing a pending page PUTs it with empty markdown (creates
an empty page that renders with its title — saving never deletes;
deletion is the page editor's explicit choice: saving trimmed-empty
text issues a REST DELETE), then switches to the page editor tab for
the actual writing. Dropping ON the lower part of a row moves the page
under that row (the child list's container invisibly overlaps its own
row's bottom via negative margin — Sortable inserts it as the first child
natively), while a row's exposed top edge inserts a sibling before it. Row
indentation is structural (each nested list margin-indents itself), so a
dragged row previews its whole subtree at the target list's depth. The
shell is dynamic-imported onto the content page by pagerite.js when an edit
pen is clicked (the pens are injected by pagerite.js after the session
validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`).
In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`),
in prod from the hashed build assets resolved via
`frontend-build/.vite/manifest.json`. `vite.config.js` sets
`appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`,
`assetsDir: '_/assets'` (so the build mirrors the URL space;
`frontend/public/favicon.ico` lands at the build root and is served at
`/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus
`src/assets/pagerite.css` as a separate stylesheet entry; theme and
banner-design CSS are NOT built — they live in `pagerite/themes/{name}/`
and are served by the backend. There
is no `index.html` source (it would shadow `/` and turn missing dev paths
into an empty Vue shell). All outputs are ES modules. The build sets
`preserveEntrySignatures: 'exports-only'` because main.js is consumed
via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite
app builds otherwise strip unused entry exports, leaving dead edit pens.
In dev the backend links theme/banner-design stylesheets like in prod
(`/_themes/...`); only the base CSS is Vite-injected from JS, and
pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/
`#pagerite-user` elements to restore the canonical order (base < theme <
design < custom CSS). Theme switches in the site editor simply swap the
`#pagerite-theme` link href, identically in dev and prod.
vite-plugin-fastapi.js has an
auto-upgrade marker — edit `vite.config.js`, not the plugin.
- `docs/` — design documentation.
## Toolchain ## Toolchain
- Python >= 3.14, managed with **uv**. Dependencies: `fastapi[standard]`, - Python >= 3.14, managed with **uv**. Dependencies: `fastapi[standard]`, `fastapi-vue`, `html5tagger`, `kanta`, `markdown-it-py`, `mdit-py-plugins`, `pygments`, `tracerite`; dev group has `httpx`. Run anything via `uv run ...` (the venv is `.venv`).
`fastapi-vue`, `html5tagger`, `kanta`, `markdown-it-py`, `mdit-py-plugins`,
`pygments`, `tracerite`; dev group has `httpx`. Run anything via
`uv run ...` (the venv is `.venv`).
- Key libraries: - Key libraries:
- **html5tagger** — all HTML generation (`E`, `Document`, `Template`, - **html5tagger** — all HTML generation (`E`, `Document`, `Template`, `HTML` for trusted/raw HTML).
`HTML` for trusted/raw HTML).
- To create stand alone pages, begin with `doc = Document(...)` that gives a HTML5 page header - To create stand alone pages, begin with `doc = Document(...)` that gives a HTML5 page header
- Chain with `doc.p("text").br`: every attribute access creates element to doc (returning self), calls add content to current element. - Chain with `doc.p("text").br`: every attribute access creates element to doc (returning self), calls add content to current element.
- Closing tags are not used where optional, e.g. no `</p>` or `</li>` is ever included in output. Due to this proper "nesting" of content is NOT required and should be avoided. Where needed, () directly after tag define attributes and content INSIDE the element, then close the element. `with doc.ul:` and such may be used for larger chunks. - Closing tags are not used where optional, e.g. no `</p>` or `</li>` is ever included in output. Due to this proper "nesting" of content is NOT required and should be avoided. Where needed, () directly after tag define attributes and content INSIDE the element, then close the element. `with doc.ul:` and such may be used for larger chunks.
- Prefer building directly on one builder with `with` blocks (recursing - Prefer building directly on one builder with `with` blocks (recursing inside a with block for hierarchies) over preparing `E.` snippets into variables and composing them. Note `with doc.li:` alone fails (`li` has an optional end tag) — use `with doc.li.ul:` style chains, or `doc.li.a(...)` followed by a nested `with doc.ul:` block.
inside a with block for hierarchies) over preparing `E.` snippets into - `Template(builder)` freezes a builder with **Capitalized** attribute placeholders (e.g. `E.Title`, `doc.main(E.Main, id="main")`); calling it fills the slots with escaping — pass `HTML(...)` for raw HTML. Passing a list to a template slot expands it; passing a list to a normal builder call does NOT (spread it: `E.ul(*items)`).
variables and composing them. Note `with doc.li:` alone fails (`li`
has an optional end tag) — use `with doc.li.ul:` style chains, or
`doc.li.a(...)` followed by a nested `with doc.ul:` block.
- `Template(builder)` freezes a builder with **Capitalized** attribute
placeholders (e.g. `E.Title`, `doc.main(E.Main, id="main")`); calling
it fills the slots with escaping — pass `HTML(...)` for raw HTML.
Passing a list to a template slot expands it; passing a list to a
normal builder call does NOT (spread it: `E.ul(*items)`).
- To create plain HTML snippets use `E.div(E.p("content"))` etc using the `E` empty builder. - To create plain HTML snippets use `E.div(E.p("content"))` etc using the `E` empty builder.
- **kanta** — asyncio-native embedded database: `Kanta(filename, data)` - **kanta** — asyncio-native embedded database: `Kanta(filename, data)` root object, `transaction`, `flush`, snapshot/replay-log persistence.
root object, `transaction`, `flush`, snapshot/replay-log persistence.
- `async with Kanta(Data(),...) as kanta:` (or await kanta.open/close) - `async with Kanta(Data(),...) as kanta:` (or await kanta.open/close)
- `with kanta.transaction(...) as data:` - transactions only for writes - `with kanta.transaction(...) as data:` - transactions only for writes
- `data` may be referenced directly to read anywhere and to modify in transactions (`as data` is just a shorthand access) - `data` may be referenced directly to read anywhere and to modify in transactions (`as data` is just a shorthand access)
@@ -303,32 +44,12 @@ not for the public pages. See `docs/design-principles.md` for the design.
- We prefer objects rather than lists, as this works better in change diffs. E.g. `dict[str, True]` where the keys indicate presence and always have value `True`. - We prefer objects rather than lists, as this works better in change diffs. E.g. `dict[str, True]` where the keys indicate presence and always have value `True`.
- Maintaining and owning the app's own `Data` object is preferable; Kanta never copies this, only edits in place - Maintaining and owning the app's own `Data` object is preferable; Kanta never copies this, only edits in place
- Note: besides opening it every access is immediate direct variable access: no `await`, no locks, no delays - Note: besides opening it every access is immediate direct variable access: no `await`, no locks, no delays
- **fastapi-vue** — template glue for serving/building the Vue frontend; - **fastapi-vue** — template glue for serving/building the Vue frontend; keep its integration points (`Frontend`, build hook) intact.
keep its integration points (`Frontend`, build hook) intact. - **markdown-it-py** — Markdown rendering with `html=True` raw passthrough; mdit-py-plugins for footnote/deflist/tasklists/attrs; **Pygments** for server-side code highlighting (`nowrap` spans, styled by `frontend/src/assets/pygments.css` which maps token classes 1:1 onto the `--code-*` variables; light/dark palette sets live in `pagerite.css` and resolve via `light-dark()` from the theme's `color-scheme` — themes pick a set, not individual colors).
- **markdown-it-py** — Markdown rendering with `html=True` raw
passthrough; mdit-py-plugins for footnote/deflist/tasklists/attrs;
**Pygments** for server-side code highlighting (`nowrap` spans, styled
by `frontend/src/assets/pygments.css` which maps token classes 1:1 onto
the `--code-*` variables; light/dark palette sets live in
`pagerite.css` and resolve via `light-dark()` from the theme's
`color-scheme` — themes pick a set, not individual colors).
## Conventions ## Conventions
- Keep dependencies minimal; add via `uv add` and mention it. - Keep dependencies minimal; add via `uv add` and mention it.
- The public URL space belongs to content (pretty slugs at root). Reserve - The public URL space belongs to content (pretty slugs at root). Reserve only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus `/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores `[a-z0-9_-]` (the site editor filters input live via `slugify.js`, built on the `transliteration` npm package — unicode folds to ASCII, spaces become hyphens; an empty slug on a new page is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
only `/_` for the machinery (`/_api/`, `/_f/`, `/_assets/`), plus - No auth in core code; the SSO/reverse proxy gates all of `/_api` (forward-auth) and owns `/auth/` (login/logout, session validation). Pages render identically for everyone; pagerite.js adds the editing UI only after the auth server validates the session.
`/favicon.ico` from the build. Slugs are lowercase ASCII letters, digits, - Update the relevant MarkDown files when architecture, tooling, or conventions change.
hyphens and underscores `[a-z0-9_-]` (the site editor filters input live
via `slugify.js`, built on the `transliteration` npm package — unicode
folds to ASCII, spaces become hyphens; an empty slug on a new page is
derived from its title), may not begin with `_` or `.`, and such URLs are
never looked up as content.
- No auth in core code; the SSO/reverse proxy gates all of `/_api`
(forward-auth) and owns `/auth/` (login/logout, session validation).
Pages render identically for everyone; pagerite.js adds the editing UI
only after the auth server validates the session. Never add output
sanitization "for safety" against the author — embedded HTML/scripts in
Markdown are passed through deliberately.
- Update this file and `docs/design-principles.md` when architecture,
tooling, or conventions change.
+1 -3
View File
@@ -1,5 +1,3 @@
# Pagerite # Pagerite
A single-user CMS/blog. FastAPI serves HTML rendered in Python with html5tagger, A single-user CMS/blog. FastAPI serves HTML rendered in Python with html5tagger, content is persisted in a kanta database and rendered on the fly per request. Vue is used only for the interactive editing tools, not for the public pages.
content is persisted in a kanta database and rendered on the fly per request.
Vue is used only for the interactive editing tools, not for the public pages.
+324
View File
@@ -0,0 +1,324 @@
# Analytics
Server-side visit analytics. Data lives in a plain JSON file — a msgspec
Struct dumped to disk — separate from the kanta content database, path from
`PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by
`.analytics.json`, e.g. `pagerite.analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
`CrawlerHit`, `AbuseHit`) and the `Store` (in-memory data + session map,
atomic JSON persistence).
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics`
(admin-gated like every `/_api` endpoint).
- `frontend/src/pagerite.js` — client navigation pings and the 📊 pen.
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
normal site layout on the `/_a` analytics page.
- `frontend/src/analytics-main.js` — page entry that mounts `AnalyticsView`
into `#analytics-app` inside `#main`.
## What is collected
The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
`fr`, `to`, `hide` and `read` as query parameters (`fr` = source path;
falsy values are omitted):
- **Initial page load**: only `to` — the loaded path — is sent, never `fr`
(an `fr` equal to `to` would log a bogus self-transition when a session
already exists, e.g. a second tab). This ping is what starts
the visit and counts the entry page view — the document GET alone records
nothing, so bots never register (admin browsing does register, but
flagged `hide`; see **Admins** below). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent
gives them away: pings whose UA matches `_is_bot_ua` (anything calling
itself a "bot", plus known exceptions such as GoogleOther) are ignored
server-side, and their document GETs land in the crawler list instead.
No source-IP verification is done: a spoofed bot UA merely lands in the
crawler stats, and scanners that probe telltale paths are caught by the
abuse rules regardless. Reloads are not
visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a
refresh neither counts a second view nor logs a self-transition. The GET
handler stashes a cross-origin https `Referer` (origin part only —
unavailable to JS once the page has loaded) and any
`utm_*` query parameters in in-memory IP tables, consumed by the ping that
starts the visit; internal or absent referers never touch the referer table.
- **Internal fetch-navigations**: `to` is the target path, sent only after
the swap actually happened (a failed swap falls back to a full load,
whose initial ping counts the view instead — no gap, no double count).
- **External links** (`https` only): `to` is the link's full URL. This is the
exit-link record; the user may continue navigating afterwards (new tab,
back), so the exit URL is not necessarily the last trail entry. Outbound
links are stored by full URL so several links to the same domain remain
distinct.
- **Excluded**: back/forward (popstate) navigations, navigating *to* the
analytics page (`/_a` — its GET is untracked, and the server rejects it
as a ping target anyway), and everything while the user has the editor
open (`body.editing`). Admin noise, not visits. Navigating *away* from
`/_a` does ping: the fetch-navigation already GET-ed the target page
without the preload header, and without the ping that GET would flush to
the crawler list.
- **Admins**: when SSO is in use and the session is known to be an admin,
the client still pings but adds `hide=1`. The activity is recorded as
usual (navigations and all), but the `hide` flag is set on the **client
record** — so it covers everything that client ever did: visits and
crawler hits from before the login included. Hidden clients never appear
in the viewer payload: `Store.display()` drops their visits, crawler
hits, abuse hits and metadata, and computes every aggregate (site visits,
page views, transitions) from the visible visits only, so nothing needs
to be reversed or redacted. Pending crawler hits from a hidden client
are discarded when they expire, so admin browsing never lands in the
crawler list either. With no auth proxy
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
is recorded.
- The server validates `to`: internal paths must be valid slug paths
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
https origin and accepted only when the client sent exactly that.
- **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
`User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
the first 6 bytes identify a shared `Client` record. The `Client` stores
the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
`country` from the language-region subtag, and asynchronously-filled
`country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
crawler hits and abuse hits all reference this record by its hash, so
client metadata is stored once instead of repeated per event.
- The visitor IP is stored in the `Client`. A reverse-DNS lookup is
attempted for each new client and the result, when available, is stored as
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
root, it is loaded at startup and used to look up `country`/`city`. These
lookups run in background tasks after the event is stored, so the `/_a`
response is never delayed. The decompressed `dbip-*.mmdb` file is kept in
the repository root and ignored by git. The CLI flag `--dbip`
(`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
skipping the download when the local database is already current and
removing older versions after an update; without the flag only an existing
file is used.
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
hit — except idle-time link preloads from pagerite.js, which carry an
`x-pagerite-preload` header and are not tracked at all (the ping sent when
the user actually navigates to a preloaded page does the counting; forging
the header only hides a GET from the crawler stats, the path-based abuse
classification is unaffected). If a ping
from the same client arrives within 10 seconds the hit is discarded;
otherwise it is written to `crawlers` — unless the client is hidden
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as
visits or views. The `Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited; the crawler table lists
the most recent crawler first, with the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
classifies the source IP as abuse immediately, and ten plain 404s from one
IP do too. Classification reclassifies history: all earlier crawler hits
from that IP (persisted and pending) move to the `abuse` list, so a
random-UA scanner no longer pollutes the crawler stats of the legitimate
bot it impersonates. Once classified, every document GET and 404 from the
IP is recorded as an abuse hit with the full request path (query string
included), and its pings are ignored. The classified IP set (`abuse_ips`)
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
viewer, abuse hits are grouped by IP (never by client/UA — scanners
randomize theirs) in a separate "Abuse" table. Identical paths are
collapsed into one entry with their hit count; flagged paths that
triggered classification are lifted to the top, followed by other 404s and
then document GETs from the abuser. Raw User-Agent strings are shown one
per line with their occurrence counts, and the full lists are click-to-copy.
## Visits and sessions
There are no cookies. A visit is tied together by a client hash — the first
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
/64 network), the raw `User-Agent` string and the extracted
`Accept-Language` tag. The first ping from a client hash starts a new
visit; subsequent pings extend it. Pings arriving with no known session
(server restart) start a fresh visit from the first ping — treated as
missing data rather than dropped. The client-hash → visit map and the IP →
entry-referer/UTM tables are in-memory only; client metadata is stored in
`Analytics.clients` keyed by the client hash.
Each `Client` record:
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string,
- `hide` — true for admin clients (`hide=1` ping): all their visits,
crawler hits and abuse hits are recorded but excluded from every
statistic and from the viewer payload.
Each `Visit` record:
- `start` — timestamp of the first event,
- `entry` — first page (path) seen,
- `referer` — external https origin of the initial load, `""` for direct,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `trail` — the entry page and everything seen afterwards, keyed by the
timestamp of first sight (insertion order = first-seen order). Each item
holds `to` (page path or external exit URL), the accumulated active
reading time in seconds (`read`) and the most recent HTTP status seen
for the target (`status`). Re-visiting an already seen target updates
its item instead of appending.
- `navs` — every navigation ping (`fr`, `to`), keyed by its timestamp,
repeats included. The aggregates are computed from this log at display
time.
- `utm``utm_*` query parameters from the landing URL, as a dict.
Each `CrawlerHit` record:
- `start` — timestamp of the document GET,
- `entry` — page path requested,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `referer` — external https origin of the request, `""` for direct/none,
- `query` — raw query string of the request,
- `status` — HTTP status of the served response (200 for a real page, 404
for a category placeholder or missing page).
Each `AbuseHit` record:
- `start` — timestamp of the request,
- `path` — full request path including the query string (e.g. `/.env?x=1`),
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the path that triggered abuse classification (telltale
path or the 404 that crossed the threshold),
- `is_404` — true for 404 responses, false for document GETs from the
abuser.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical paths are collapsed with their counts; flagged paths
that triggered classification are lifted to the top, followed by other 404s
and then document GETs from the abuser. Within each category paths are
sorted by count descending, then by their earliest hit.
In the visitor and crawler tables, internal paths that returned a 404 status
are shown in red and the link title includes the status code, so it is easy
to tell misses from real pages at a glance.
## Aggregates
Aggregates are **not stored**; they are computed at display time by
`Store.display()` from the visit records (entry + `navs` log), skipping
hidden clients' visits. This is what allows a client to become hidden after
navigations were already logged: no counts need reversing. The computed
shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
`crawlers`, `abuse` and `clients`):
- `transitions`: time series of page transitions, sparse nested dict
`from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
referer origin or `"(direct)"` for initial loads, a page path for pings.
- `views`: time series of page loads, `path -> bucket -> count`, sparse: only
non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp).
Every load counts, including repeats within a visit; external exit origins
are not page views and are not counted here.
- `site_visits`: `bucket -> count` of new visits started, same sparse
5-minute bucketing.
Sparseness keeps quiet sites small; dropping old data is a matter of deleting
list entries (`visits` is a plain append-only list).
## Persistence
The whole `Analytics` struct is JSON-encoded and written atomically
(temp file + rename) on every recorded event. Traffic on a small CMS makes
this cheap enough; batching can be added later without changing the format.
## Viewing
The 📊 pen in the banner corner (admins only, injected by pagerite.js next to
the edit pens) links to `/_a`, the analytics page. It is a normal site page:
the standard banner, navigation and footer stay in place, and the analytics
content is rendered inside `#main`. The page itself is public, but the data
stream comes from `WebSocket /_api/ws/analytics`, which remains admin-gated
like the rest of the management API; visitors without access see the viewer
with a "could not be loaded" message.
Because it is a real page, fetch-navigation handles it like any other internal
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
range selector updates the URL hash (`#week` etc.) so links to a specific
range can be shared. When the URL has no hash, the client derives the
default from the first analytics snapshot: `day` if the recorded history
spans less than 24 hours, otherwise `week`.
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
Charts are SVG curves (Catmull-Rom over an edge-aware Gaussian — a
change-point detector splits the series at traffic-level shifts, then each
segment is smoothed independently with a fixed sigma chosen so N events in
a single bucket peak at N events per unit. The raw series is drawn faint
underneath). Values are
**per-unit rates** — per hour on the week view (5-minute bucket counts × 12,
plotted at native 5-minute resolution), per day on the month+ ranges — and
the smoothing time scale follows the unit: the month+ sigmas are 24× the
hourly ones. The y max is derived from the smoothed curves so single-bucket
spikes don't blow up the scale, and raw spikes are clamped into the plot.
Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5
labeled intervals, minor lines at fifths when integral; the minimum y-axis
range is 10 so tiny values such as a single visit are not stretched to a
fractional scale).
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
weeks in the muted color at decreasing opacity (the current week keeps the
accent color and is
truncated at the current bucket, never drawing fake zeroes for the future);
a compact legend inside the top right of the visits chart marks the current
ISO week in accent and the overlaid past weeks as "Week M" or "Week MN" on
a muted specimen. Its x labels are weekday names centered at midday UTC, without
vertical grid
lines (day boundaries would be misleading in the viewer's timezone). The
month view labels days the same lineless way — day numbers at noon UTC,
with the month name substituted for the 1st. Month, year and all are
rolling windows ending at now, aligned to UTC day boundaries at the start
so the labels span the whole range; the bucket size follows the window —
6 hours up to 31 days, daily beyond — with boundary lines at months/years
on the longer ranges. All uses the full data reach, but keeps
at least the past 30 days (identical to the month view when the site is
younger than that, bucket size included) so the chart never collapses to a
tiny sliver when the site is young. Below the charts: a **transition map** (all pages from
`/_api/pages` — top-level menu items on a large-radius circular arc whose
bottom point is the last item (each earlier item a bit higher), connected
by a top lane labeled 🏠︎ beside the home pill (50% thicker than
the branch lanes, its label font and guide offset scaled along), each item's
subtree fanning out below it in menu order along a large-radius circular
arc that leaves heading
straight down and gradually bends right, index pages without views omitted
and their children promoted in their place. The submenu structure is drawn
as wide branch lanes: one per path prefix with at least two visible
nodes, running behind the branch's node pills as circle arcs concentric
with the fan (parent levels one radius step outward, so all lanes of a
group share exactly one form), each labeled with its branch slug
left-aligned just past the first pill and allowed to run along the lane to
its end, disappearing under later pills when long — so the lanes reflect
the path
structure even where index pages are omitted — opposite transition
directions joined into organic
tapered connections whose middle width grows logarithmically (base 2)
with the daily hit rate (uncapped), connections
carrying less than 1% of the total traffic
pruned, as are those whose thin middle would render below ~0.8 px —
fainter strands are invisible and only their wide end flares would show; beads are simulated one by one in JS (requestAnimationFrame) and
flow along each edge, persisting across data reloads (emitters are keyed
per edge direction and beads tracked by progress, so an unrelated count
change never reshuffles them), emitted at a rate linearly proportional
to the directional count with no in-flight limit, opposing directions
offset onto parallel lanes. External sources and exits whose connectors are
all culled by the width threshold are dropped from their rows themselves
(the site's own page nodes always stay, connected or not). External sources show as a node row above the
map: each visit is attributed to `utm_campaign`, then `utm_source`, then the
referer origin, then any other `utm_*` tag, so UTM-tagged visits are grouped
under their campaign/source value rather than the referer domain. A UTM
source node only links to its referer when every visit carrying that tag
came from the same origin. External exits are full-size nodes in a matching
row centered below the map, so the site itself stays in the middle), per-page view
counts, the top transitions and the 50 most recent visit trails. Data is
streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest
JSON snapshot on connect and again whenever the analytics file is updated
(with a small server-side debounce to avoid flooding under high traffic).
+31
View File
@@ -0,0 +1,31 @@
# Backend
The Python backend lives in `pagerite/`.
## `app.py`
The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url`/`openapi_url=None`) because `/docs` belongs to our content. Our own routes (content pages, `/_api/...`, `/_f/...`) are registered BEFORE `frontend.route(app, "/")` is called: fastapi-vue inserts its file routes at the position where `route()` was called (during `load()` in the lifespan), so anything defined earlier wins. The one exception is the content catch-all `/{path:path}`, registered AFTER `frontend.route()` so that built frontend assets still take priority over content slugs. The `Frontend` is constructed with `spa=False` explicitly: it only serves the built files without a catch-all.
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding, and `data.version`, which bumps on every content/settings change and so transparently invalidates the whole cache. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and `data.version`; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
## `data.py`
msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model.
## `markdown.py`
markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists, admon, gfm_autolink, sub/superscript plugins; typographer + breaks on). Custom image rule: relative srcs resolve against the page path; an image standing alone in its paragraph becomes a figure (captioned when titled), while inline-with-text images and raw `<img>` HTML stay plain. A `{dates}` line expands to the article's published/updated dateline (`p.dateline`, from `Node.created`/`modified`; left literal in previews of unsaved pages).
## `views.py`
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`Data.site_url` — learned from admin browsers reporting their `location.origin` via `POST /_api/site-url`, correct even behind reverse proxies; until learned, the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. If the markdown contains its own h1, the page title is NOT rendered as an additional h1 (it still supplies `<title>` and nav labels).
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), which is rendered when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
## `seed.py`
Demo content written only when the database is first created, via a `@kanta.bootstrap` handler in `app.py`.
+35
View File
@@ -0,0 +1,35 @@
# Content model
The site structure is stored in the kanta database managed by `pagerite/data.py`.
## Site tree
`Data.menu` maps top-level slugs to `Node`s, each with `children` keyed by slug — the URL path is the slug chain. The front page is whichever top-level node has slug "" (parallel to the other main level pages, not their parent); it cannot have children, and renaming its slug away leaves no front page ("/" redirects to the first nav item).
`Node.content` is the Markdown page, or None for a pure category label whose URL renders a placeholder page (while nav links to it point at its first child); every label's title and slug are editable.
Siblings order by the fractional `Node.order` key: a moved item gets a fresh key relative to its new siblings, all others keep theirs. `resolve`/`find_slot` walk the tree by path; moves are slot detach/attach carrying the whole subtree. Legacy flat `Data.pages` (pre-tree databases) migrates into `menu` on startup. The app owns the `Data` object; reads are plain attribute access, writes in `kanta.transaction(...)`.
`Data.version` is bumped on every write and embedded in page ETags so nav-affecting changes invalidate caches.
## Files
`Data.files` is a content-addressed store (blake3[:12] + extension) mapping file names to bytes, served at `/_f/{name}` with immutable caching; pages reference files by absolute `/_f/` URLs so hierarchy moves never break them.
## Banners
`Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles.
`Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css lives in `<head>` (id `pagerite-banner`) between the theme and the custom CSS — a `<link>` in dev, an inline `<style>` in production.
## Site settings
`Data.brand` is the site name (header link + `<title>` suffix), editable in the site editor via `/_api/settings`; empty = no header link and no `<title>` suffix.
`Data.brand_html` is raw trusted HTML replacing the brand link entirely (rendered in a `#brand` div on top of the banner, next to the nav) — site-wide, not per-page like banners; edited in the site editor with image/video upload into `Data.files`.
`Data.theme` is the active theme name (empty = none/base only); themes are folders in `pagerite/themes/{name}` containing `theme.css` and/or `banner.css` (+ `banner.svg` artwork and any extra assets the CSS references, like summer's `grass.svg`), served by the backend at `/_themes/{name}/...` — read from disk per request (etag by mtime), never built, so on-disk edits show on the next page load even in prod. The theme selector and banner-design selector enumerate these folders via `GET /_api/settings`.
`Data.custom_css` is raw trusted CSS injected inline in every page `<head>` (id `pagerite-user`) and swapped during fetch-navigation; editable in the site editor. Font picks (heading/body/brand) in the site editor are stored as plain `:root` rows in `custom_css` (`--font-body: var(--font-source-sans);` format — parsed out and rewritten on change, the `:root` block added/removed as needed), referencing the per-family variables (`--font-source-sans` etc.) from `pagerite.css`; the base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
`Data.favicon` names a file in the content-addressed `files` store, uploaded/cleared in the site editor via `PUT`/`DELETE /_api/settings/favicon`; when set it is linked as `<link rel="icon">` on every page, otherwise browsers fall back to the build's `/favicon.ico` by convention.
+30 -228
View File
@@ -1,253 +1,55 @@
# Pagerite Design Principles # Pagerite Design Principles
Pagerite is a single-user CMS/blog. This document records the initial Pagerite is a single-user CMS/blog. This document records the initial high-level design decisions; it will be refined as the implementation evolves.
high-level design decisions; it will be refined as the implementation
evolves.
## Architecture ## Architecture
- **Server-side rendered.** FastAPI serves complete HTML pages, generated in - **Server-side rendered.** FastAPI serves complete HTML pages, generated in Python with **html5tagger**. There is no client-side templating or SPA for the public site.
Python with **html5tagger**. There is no client-side templating or SPA for - **Vue only where interactivity demands it.** Small interactive islands (editing tools mainly) are Vue components mounted into specific elements of the server-rendered pages. The public reading experience has no scripting requirement.
the public site. - **Persistence via kanta.** Content is stored in an asyncio-friendly kanta database. Rendering happens on the fly on each request — there are no pre-built static artifacts.
- **Vue only where interactivity demands it.** Small interactive islands
(editing tools mainly) are Vue components mounted into specific elements of
the server-rendered pages. The public reading experience has no scripting
requirement.
- **Persistence via kanta.** Content is stored in an asyncio-friendly kanta
database. Rendering happens on the fly on each request — there are no
pre-built static artifacts.
## Content model ## Content model
- Pages and blog articles are fundamentally the same kind of thing: named - Pages and blog articles are fundamentally the same kind of thing: named pieces of content. The blog/website distinction is blurred; an article is just a page (possibly with metadata such as a publication date and listing in a feed).
pieces of content. The blog/website distinction is blurred; an article is - **Pretty URLs.** Content is addressed by its name (slug), not by technical constructs — no `/cms/...` or `/blog/post1` prefixes. Slugs usually live directly at the site root; structured content may nest (`/docs/design-principles`-style). The URL space is the author's, so reserved prefixes must be kept few and deliberate: everything internal lives under `/_` (`/_api/`, `/_f/`, `/_assets/`). The only other reserved root path is `/favicon.ico`, served from the build. Slugs are lowercase ASCII letters, digits, hyphens and underscores (`[a-z0-9_-]`; input is transliterated and filtered as you type, and a new page's empty slug is derived from its title), may not begin with `_` or `.`, and such URLs are never looked up as content.
just a page (possibly with metadata such as a publication date and - **Single user, trusted author.** No auth concerns in the core design. Everything published is public; only editing tools will later sit behind access control (external SSO when that time comes). The author is trusted to create well-meaning slugs and content — no sanitization for safety, only for correctness.
listing in a feed). - **Commenting** is not planned now but the model should not preclude it later.
- **Pretty URLs.** Content is addressed by its name (slug), not by technical
constructs — no `/cms/...` or `/blog/post1` prefixes. Slugs usually live
directly at the site root; structured content may nest
(`/docs/design-principles`-style). The URL space is the author's, so
reserved prefixes must be kept few and deliberate: everything internal
lives under `/_` (`/_api/`, `/_f/`, `/_assets/`). The only
other reserved root path is `/favicon.ico`, served from the build.
Slugs are lowercase ASCII letters, digits, hyphens and underscores
(`[a-z0-9_-]`; input is transliterated and filtered as you type, and a
new page's empty slug is derived from its title), may not begin with
`_` or `.`, and such URLs are never looked up as content.
- **Single user, trusted author.** No auth concerns in the core design.
Everything published is public; only editing tools will later sit behind
access control (external SSO when that time comes). The author is trusted
to create well-meaning slugs and content — no sanitization for safety,
only for correctness.
- **Commenting** is not planned now but the model should not preclude it
later.
## Authoring format ## Authoring format
- Content is written in **Markdown** with powerful extensions (tables, - Content is written in **Markdown** with powerful extensions (tables, footnotes, code highlighting, etc.).
footnotes, code highlighting, etc.). - **Embedded HTML is passed through unfiltered**, including inline scripts and other dynamic content the author wants to post. This is safe by the single-trusted-author assumption above.
- **Embedded HTML is passed through unfiltered**, including inline scripts - Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition lists, task lists, brace-attributes, admonitions and `::: name` containers — generic `<div class="name">` wrappers (the name may be followed by brace attributes: `::: aside {.right}`), of which `::: aside` floats as a muted side box (leaning into the empty right gutter on wide single-column pages) and `::: nocols` opts its section out of column layout; tables and strikethrough from the default preset), GitHub-style alerts (`> [!NOTE]` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the admonition callout styling), with `html=True` for raw passthrough, `typographer=True` for SmartyPants-style replacements in body text (curly quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.), and `breaks=True` so single line breaks inside paragraphs become `<br>` — including inside blockquotes, where every newline is kept and a blank `>` line starts a new paragraph. Code spans/blocks and raw HTML are left untouched. Fenced code blocks are highlighted server-side with **Pygments** (`nowrap` spans styled by `/_assets/pygments-*.css`, which maps every token class onto the `--code-*` variables; the base stylesheet defines light and dark palette sets resolved via `light-dark()`, so each theme gets the set matching its `color-scheme` and may only retint `--code-bg` to keep the well in the page's color family); a JS copy button appears on hover. Should this prove limiting, we implement our own renderer on top of html5tagger, which we already use for all HTML generation.
and other dynamic content the author wants to post. This is safe by the - **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`) are stored by content hash — blake3, first 6 bytes hex + original extension — and served immutable from `/_f/{hash}.ext`. Absolute URLs that survive page renames and dedupe identical content; pages no longer own files. An image standing alone in its paragraph becomes a block `<figure>` — with `<figcaption>` when it has a title; images inline with text and raw `<img>` HTML stay plain inline images. Positioning is by attribute classes: `![alt](/_f/….avif "Caption"){.right}``{.right}`, `{.left}` float at 30% of the text column (the caption wraps within it; an explicit `width=300` makes the figure shrink-wrap the image instead), `{.wide}` goes full bleed (viewport edge to edge, or up to the docked editor; the sidebar stacks on top of it); plain attributes like `width=300` work too. The same brace syntax on a block's last line (no blank line between) applies to the whole block: a paragraph ending with `{.wide}` becomes a full-width element that breaks out of the column layout; written on the line after a block it applies to that preceding block — this is how headings, `::: containers` and code fences take classes (a wide code fence goes full bleed like a wide figure). Headings (h1/h2) clear floats, so images never overflow into the next section.
single-trusted-author assumption above.
- Renderer: **markdown-it-py** with mdit-py-plugins (footnotes, definition
lists, task lists, brace-attributes; tables and strikethrough from the
default preset), with `html=True` for raw passthrough,
`typographer=True` for SmartyPants-style replacements in body text (curly
quotes, `--` / `---` → en / em dashes, `...` → ellipsis, `(c)` → ©, etc.),
and `breaks=True` so single line breaks inside paragraphs become `<br>`.
Code spans/blocks and raw HTML are left untouched. Fenced code blocks are
highlighted server-side with
**Pygments** (`nowrap` spans styled by
`/_assets/pygments-*.css`, which maps every token class onto the `--code-*`
variables; the base stylesheet defines light and dark palette sets resolved
via `light-dark()`, so each theme gets the set matching its `color-scheme`
and may only retint `--code-bg` to keep the well in the page's color
family); a JS copy button appears on hover. Should this
prove limiting, we implement our own renderer on top of html5tagger,
which we already use for all HTML generation.
- **Files are content-addressed.** Uploads (`PUT /_api/files/{filename}`)
are stored by content hash — blake3, first 6 bytes hex + original
extension — and served immutable from `/_f/{hash}.ext`. Absolute URLs
that survive page renames and dedupe identical content; pages no longer
own files. An image standing alone in its paragraph becomes a block
`<figure>` — with `<figcaption>` when it has a title; images inline
with text and raw `<img>` HTML stay plain inline images. Positioning
is by attribute classes:
`![alt](/_f/….avif "Caption"){.right}``{.right}`, `{.left}` float at
30% of the text column (the caption wraps within it; an explicit
`width=300` makes the figure shrink-wrap the image instead),
`{.wide}` goes full bleed (viewport edge to edge, or up to the docked
editor; the sidebar stacks on top of it); plain attributes like `width=300`
work too. Headings (h1/h2) clear floats, so images never overflow into the
next section.
## Page structure and navigation ## Page structure and navigation
- All pages share one static layout, defined once as an **html5tagger - All pages share one static layout, defined once as an **html5tagger Template** with capitalized placeholders (`Title`, `Banner`, `Nav`, `Sidebar`, `Main`) filled per request. The dynamic regions carry stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`).
Template** with capitalized placeholders (`Title`, `Banner`, `Nav`, - The page top is a **full-width banner header** with the site name and the navigation bar overlaid on it — no separate chrome header. The banner combines two layers, stacked in `#page-banner` (a grid, so they overlay): first the **banner design** — a named design living in a theme folder (`pagerite/themes/{name}/banner.css` plus artwork as `banner.html` — arbitrary markup like canvas + style + script — or `banner.svg`), chosen per page via `Node.banner_design` (a design name, "" for none, None to inherit from the nearest ancestor, then the front page, then the active theme's own design). The artwork is inlined into a `div[data-design]` wrapper: SVG artwork can be recolored from the theme stylesheet (corporate's single SVG serves both light and dark mode via `var()`-driven stops). Second, **per-page author code**: `Node.banner` holds an arbitrary trusted HTML snippet (an image, a styled div, canvas + script — anything), resolved by walking up the node's ancestors to the front page and rendered **after** the design artwork, so author styles always win over the design's own. The base stylesheet falls back to a plain gradient. There is deliberately no scrim fading the banner into the page background — any such fade would ruin user-supplied designs; themes that want one bake it into their SVG (purple does).
`Sidebar`, `Main`) filled per request. The dynamic regions carry stable - **Fetch-navigation.** Links are plain `<a href>`; a small script (`frontend/src/pagerite.js`) intercepts same-origin clicks, fetches the page, and swaps the `#page-banner`, `#nav`, `#sidebar` and `#main` regions, the document title, and the site-wide custom CSS (`<style id="pagerite-user">` in `<head>`), keeping the rest of `<head>` and the layout chrome. Without JS everything works as normal page loads. Scripts inside fetched banner and content regions are re-created so they execute. Swaps run inside `document.startViewTransition` for a rotating cube page transition (CSS adapted from termotohtori.fi — the `::view-transition*` block is fragile, do not tweak; skipped under `prefers-reduced-motion`). Navigation within the same top-level section crossfades instead of rotating; browser back navigation rotates in reverse.
ids (`#page-banner`, `#nav`, `#sidebar`, `#main`). - **The site structure is a tree of labels.** `Data.menu` holds the top-level items by slug, each with `children` keyed by slug — the URL path is the slug chain. The front page is a top-level node with slug "" (an item *parallel* to the other main level pages, not their parent) and cannot have children. The header navbar holds only the top level; a top-level item is highlighted when viewing any of its subpages. When the current page is inside a main level section with children, those direct children are listed in a **left sidebar** (`#sidebar`), one level deep. The sidebar exists only when there is something to navigate — sections with fewer than two published items, leaf pages and the front page render no aside element at all. Other sections' subitems are never shown without navigating into them first.
- The page top is a **full-width banner header** with the site name and the - **Landing pages are optional.** Every label can either have content (`Node.content`, a Markdown page) or none — a content-less label renders a placeholder page (404 with a pen to create it) instead of redirecting, while nav links to it point straight at its first child, so categories need no filler content and normal navigation never sees the placeholder. Title and slug of every label are editable; renaming a slug moves the whole subtree. The sidebar never lists the section itself, avoiding title duplication with the navbar.
navigation bar overlaid on it — no separate chrome header. The banner - **Menu order is manual.** Each node has a fractional `order` key among its siblings; reordering/moving writes only the moved node (it takes a fresh value halfway between its new siblings; all other items keep theirs). New pages append at the end of their menu. Structure edits (reorder, move/rename with the whole subtree, retitle) go through `POST /_api/structure` and the editor's structure panel.
combines two layers, stacked in `#page-banner` (a grid, so they overlay):
first the **banner design** — a named design living in a theme folder
(`pagerite/themes/{name}/banner.css` plus artwork as `banner.html`
arbitrary markup like canvas + style + script — or `banner.svg`),
chosen per page via
`Node.banner_design` (a design name, "" for none, None to inherit from
the nearest ancestor, then the front page, then the active theme's own
design). The artwork is inlined into a `div[data-design]` wrapper: SVG
artwork can be recolored from the theme stylesheet (corporate's single
SVG serves both light and dark mode via `var()`-driven stops). Second,
**per-page author code**: `Node.banner` holds an arbitrary trusted HTML
snippet (an image, a styled div, canvas + script — anything), resolved by
walking up the node's ancestors to the front page and rendered **after**
the design artwork, so author styles always win over the design's own.
The base stylesheet falls back to a plain gradient. There is deliberately
no scrim fading the banner into the page background — any such fade would
ruin user-supplied designs; themes that want one bake it into their SVG
(purple does).
- **Fetch-navigation.** Links are plain `<a href>`; a small script
(`frontend/src/pagerite.js`) intercepts same-origin clicks, fetches the
page, and swaps the `#page-banner`, `#nav`, `#sidebar` and `#main` regions,
the document title, and the site-wide custom CSS (`<style id="pagerite-user">`
in `<head>`), keeping the rest of `<head>` and the layout chrome. Without JS
everything works as normal page loads. Scripts inside fetched banner and
content regions are re-created so they execute. Swaps run inside `document.startViewTransition` for a rotating
cube page transition (CSS adapted from termotohtori.fi — the
`::view-transition*` block is fragile, do not tweak; skipped under
`prefers-reduced-motion`). Navigation within the same top-level section
crossfades instead of rotating; browser back navigation rotates in
reverse.
- **The site structure is a tree of labels.** `Data.menu` holds the
top-level items by slug, each with `children` keyed by slug — the URL
path is the slug chain. The front page is a top-level node with slug ""
(an item *parallel* to the other main level pages, not their parent) and
cannot have children. The header navbar holds only the top level; a
top-level item is highlighted when viewing any of its subpages. When the
current page is inside a main level section with children, those direct
children are listed in a **left sidebar** (`#sidebar`), one level deep.
The sidebar exists only when there is something to navigate — sections
with fewer than two published items, leaf pages and the front page render
no aside element at all. Other sections' subitems
are never shown without navigating into them first.
- **Landing pages are optional.** Every label can either have content
(`Node.content`, a Markdown page) or none — a content-less label renders
a placeholder page (404 with a pen to create it) instead of redirecting,
while nav links to it point straight at its first child, so categories
need no filler content and normal navigation never sees the placeholder.
Title and slug of every label are editable; renaming a
slug moves the whole subtree. The sidebar never lists the section
itself, avoiding title duplication with the navbar.
- **Menu order is manual.** Each node has a fractional `order` key among
its siblings; reordering/moving writes only the moved node (it takes a
fresh value halfway between its new siblings; all other items keep
theirs). New pages append at the end of their menu. Structure edits
(reorder, move/rename with the whole subtree, retitle) go through
`POST /_api/structure` and the editor's structure panel.
- Unpublished pages are hidden from both nav and URL access (404). - Unpublished pages are hidden from both nav and URL access (404).
## Reading experience ## Reading experience
- The article column is sized by the **viewport, never by content**: a - The article column is sized by the **viewport, never by content**: a symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps the layout stable across navigation. The sidebar occupies the left gutter, the right gutter balances it; wide screens get columns inside long articles without changing the article's width. Columns are decided client-side (pagerite.js): the body splits into segments at h1/h2 headings and `.wide` elements (full-width separators, never inside columns), and a segment gets columns only when it holds enough text — code blocks are excluded from that measure, and a `::: nocols` container opts its whole section out.
symmetric grid (`1fr minmax(0, 78rem) 1fr`) with flexible gutters keeps - A gentle **scroll-reveal** of headings, figures and block-level elements (IntersectionObserver). It is layout-level: articles need no support for it, and `prefers-reduced-motion` disables all motion.
the layout stable across navigation. The sidebar occupies the left
gutter, the right gutter balances it; wide screens get columns inside
long articles without changing the article's width.
- A gentle **scroll-reveal** of headings, figures and block-level elements
(IntersectionObserver). It is layout-level: articles need no support
for it, and `prefers-reduced-motion` disables all motion.
## Styling ## Styling
- The base stylesheet `frontend/src/assets/pagerite.css` provides the layout, - The base stylesheet `frontend/src/assets/pagerite.css` provides the layout, typography and interaction rules with conservative CSS variables. A theme layer (`pagerite/themes/{name}/theme.css` — currently `purple`, `corporate` and `nitro`, served by the backend at `/_themes/{name}/theme.css` straight from disk, never built) overrides those variables and adds the visual styling; `Data.theme` selects the active theme (empty = none/base only) and the site editor can switch it, choosing from the theme folders found on disk. Vue may add per-component styles on top where needed. The corporate and nitro themes switch palettes automatically via `prefers-color-scheme` (corporate is light-first with a matching dark palette; nitro a warm light-grey page or, in dark mode, a deep violet one — its dark banner and orange accents carry over unchanged); purple (dusk) uses one fixed palette for everyone. Themes may restyle structural details the base leaves plain — heading colors and underlines, list markers, nav treatment, brand sizing. A theme folder may also ship a **banner design** (`banner.css` + `banner.svg`), selectable per page independently of the active theme. The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll parameter on `<html>` (event-driven, so it is still when the page is idle), the banner contents drift within their window (with scale overscan so no edge shows), and designs may key their own effects off the same parameter — purple's sun rises as you scroll.
typography and interaction rules with conservative CSS variables. A theme layer - Fonts, the shared stylesheet and pygments styles live under `frontend/src/assets/` and are emitted as hashed assets under `/_assets/` (Source Serif 4 for headings, Source Sans 3 for body, Fira Code for code by default; Fraunces, Literata, Cormorant, Playfair Display, Inter, Montserrat, Cause, Exo 2 and New Rocker kept as woff2 options with local `@font-face`, variable-weight where available). No third-party requests.
(`pagerite/themes/{name}/theme.css` — currently `purple`, `corporate`
and `nitro`, served by the backend at `/_themes/{name}/theme.css` straight
from disk, never built) overrides those variables and
adds the visual styling; `Data.theme` selects the active theme (empty = none/base
only) and the site editor can switch it, choosing from the theme folders
found on disk. Vue may add per-component styles on top
where needed. The corporate and nitro themes switch palettes automatically via
`prefers-color-scheme` (corporate is light-first with a matching dark palette;
nitro a warm light-grey page or, in dark mode, a deep violet one — its dark
banner and orange accents carry over unchanged); purple (dusk) uses one
fixed palette for everyone. Themes may restyle structural details the base
leaves plain — heading colors and underlines, list markers, nav treatment,
brand sizing. A theme folder may also ship a **banner design**
(`banner.css` + `banner.svg`), selectable per page independently of the
active theme. The banner artwork has scroll parallax: pagerite.js sets the
`--pry` scroll parameter on `<html>` (event-driven, so it is still when the
page is idle), the banner contents drift within their window (with scale
overscan so no edge shows), and designs may key their own effects off the
same parameter — purple's sun rises as you scroll.
- Fonts, the shared stylesheet and pygments styles
live under `frontend/src/assets/` and are emitted as hashed assets under
`/_assets/`
(Source Serif 4 for headings, Source Sans 3 for body, Fira Code for code
by default; Fraunces, Literata, Cormorant, Playfair Display, Inter,
Montserrat, Cause, Exo 2 and New Rocker kept as woff2 options with
local `@font-face`, variable-weight where available). No third-party
requests.
## Editing ## Editing
- Editing happens **in place**, in two modes opened by two pens: - Editing happens **in place**, in two modes opened by two pens:
- **Page mode** — the 🖊️ next to a page's heading (including 404s, which - **Page mode** — the 🖊️ next to a page's heading (including 404s, which is how new pages start) opens a CodeMirror Markdown editor docked to the left of the article: the panel is fixed to the viewport's left edge (its top tracks the banner's bottom until the banner scrolls away), the content shifts right and the sidebar hides while editing. Preview renders server-side per keystroke (no debouncing) straight into the visible article's heading and body.
is how new pages start) opens a CodeMirror Markdown editor docked to - **Site mode** — the 🖊️ on the banner opens a panel with the site **brand** (applied to the header live), a **theme** selector (swapping the theme stylesheet in place), **font** picks (heading/body/brand — stored as plain `:root` rows inside the custom CSS, referencing the base stylesheet's per-family font variables), a **site-wide custom CSS** field (injected into `<style id="pagerite-user">` in the live page head and swapped during fetch-navigation), the page's **banner design** selector (inherit / none / any design found on disk, inherited by children), the page's **banner HTML** field (supplementing the design, previewed into the real banner region, so you see exactly which banner you're editing) and the **structure tree**. Everything saves immediately as you edit — no save button, no edit mode.
the left of the article: the host sits inside `#content` (below the - Clicking a pen again closes the editor (without saving; a dirty preview reloads the page). The pens are `<button>`s wired up by `pagerite.js` — editing is an action, not a navigation. The editor's WebSocket **reconnects automatically** with local text and pending saves preserved. (All users are trusted authors for now; access control later with SSO.)
banner, never over the footer), the content shifts right and the - **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published controls. Images can be pasted straight into the editor or chosen via a file input: they upload to the content store (`PUT /_api/files/...`) and insert `![alt](/_f/hash.ext)` at the cursor.
sidebar hides while editing. Preview renders server-side per keystroke - The **structure panel** (vue-draggable tree of the whole site, in site mode) covers page management: reorder any menu level, drag across sections, add, delete (two clicks: the button arms, then deletes — no dialogs). Every node is a real label — content-less category rows offer a to give them a landing page. Deleting a category removes only its landing page (the label and its subpages stay). Every non-empty list ends with a row that starts a new page as a local-only tree row at that level; the row can be dragged into place before its title and slug are filled in and is persisted only on commit. While dragging, these rows double as "end of this list" drop targets; dropping ON the lower part of a row makes the page that row's first child (even a leaf's, creating a sublist), while a row's exposed top edge inserts a sibling before it. A dragged row's indentation previews the target list's depth. Rows are always editable: titles save while typing, slug edits commit on blur/Enter since they rename the path (moving the whole subtree). The front page is the root row with an empty slug — renaming it away leaves no front page ("/" redirects to the first nav item), and giving another top-level row the empty slug makes it the front page.
(no debouncing) straight into the visible article's heading and body. - Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a stateless JSON protocol (`open`/`render`/`save`; on save all fields are optional and absent ones keep their old values, `move_from` renames), avoiding REST polling and races. Rendering always stays server-side.
- **Site mode** — the 🖊️ on the banner opens a panel with the site - A REST API also exists for scripting, all under `/_api/`: `GET pages` (the full tree), `PUT/DELETE pages/{path}`, `GET/PUT settings` (site brand, theme and custom CSS), `POST structure` (reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`.
**brand** (applied to the header live), a **theme** selector (swapping - On startup, seed pages from `pagerite/seed.py` are added **only if missing** — existing user content is never overwritten.
the theme stylesheet in place), **font** picks (heading/body/brand —
stored as plain `:root` rows inside the custom CSS, referencing the base
stylesheet's per-family font variables), a **site-wide custom CSS** field (injected
into `<style id="pagerite-user">` in the live page head and swapped during
fetch-navigation), the page's **banner design** selector (inherit /
none / any design found on disk, inherited by children), the page's
**banner HTML** field (supplementing the design, previewed into the
real banner region, so you see exactly which banner you're editing) and
the **structure tree**. Everything saves immediately as you edit — no
save button, no edit mode.
- Clicking a pen again closes the editor (without saving; a dirty preview
reloads the page). The pens are `<button>`s wired up by `pagerite.js`
editing is an action, not a navigation. The editor's WebSocket
**reconnects automatically** with local text and pending saves preserved.
(All users are trusted authors for now; access control later with
SSO.)
- **CodeMirror 6** for Markdown editing (no WYSIWYG), title/published
controls.
Images can be pasted straight into the editor or chosen via a file
input: they upload to the content store (`PUT /_api/files/...`) and
insert `![alt](/_f/hash.ext)` at the cursor.
- The **structure panel** (vue-draggable tree of the whole site, in site
mode) covers page management: reorder any menu level, drag across
sections, add, delete (two clicks: the button arms, then deletes — no
dialogs). Every node is a real label — content-less category rows offer
a to give them a landing page.
Deleting a category removes only its landing page (the label and its
subpages stay). Every non-empty list ends with a row that starts a
new page as a local-only tree row at that level; the row can be dragged
into place before its title and slug are filled in and is persisted only
on commit. While dragging, these rows double as "end of this list"
drop targets; dropping ON the lower part of a row makes the page that
row's first child (even a leaf's, creating a sublist), while a row's
exposed top edge inserts a sibling before it. A dragged row's
indentation previews the target list's depth. Rows are always
editable: titles save while typing, slug edits commit on blur/Enter
since they rename the path (moving the whole subtree). The front page
is the root row with an empty slug — renaming it away leaves no front
page ("/" redirects to the first nav item), and giving another
top-level row the empty slug makes it the front page.
- Preview and saving go over a **WebSocket** (`/_api/ws/editor`) with a
stateless JSON protocol (`open`/`render`/`save`; on save all fields are
optional and absent ones keep their old values, `move_from` renames),
avoiding REST polling and races. Rendering always stays server-side.
- A REST API also exists for scripting, all under `/_api/`:
`GET pages` (the full tree), `PUT/DELETE pages/{path}`,
`GET/PUT settings` (site brand, theme and custom CSS), `POST structure`
(reorder/move/retitle), file upload/removal via `PUT/DELETE files/{name}`.
- On startup, seed pages from `pagerite/seed.py` are added **only if
missing** — existing user content is never overwritten.
+30
View File
@@ -0,0 +1,30 @@
# Editing interface
The Vue editor is a single tabbed `EditorShell.vue` mounted in a host div created inside the static document.
## Tabs
The shell hosts four kept-alive tabs (ordered site-wide first — site, structure — then, after a visual break, the per-page tabs — article, banner):
- `PageEditor.vue` — CodeMirror + server-rendered preview over WebSocket `/_api/ws/editor`, previewing into the visible article; editor and article scrolls are linked proportionally both ways, applied instantly (the window keeps scrolling normally while any editor is open — the panel is fixed to the viewport's left edge, its top tracking the banner's bottom edge until the banner scrolls away — and the panel scrolls internally); a format bar offers Markdown helpers — bold/italic/code/link/table/image upload, with Ctrl/Cmd-B/I/S bindings — for the hard-to-remember syntax. Edits content and title only, never the path.
- `BannerEditor.vue` — per-page banner HTML + banner design selector, previewed into `#page-banner`.
- `SiteEditor.vue` — site brand + optional custom brand HTML with image/video upload + theme selector + font picker + favicon upload — clicking the preview tile picks a new one — + site-wide custom CSS, CSS injected into `<head id="pagerite-user">`.
- `StructureEditor.vue` — the vue-draggable structure tree with always-editable title/slug inputs per row.
Media uploads everywhere use the image icon buttons (pasting into the editor works too). The article, banner and site-settings pens are shorthands that open the shell on the matching tab; once open, clicking a pen switches tabs (and retargets the editors to the current page) instead of closing/remounting. The close button in the tab bar closes the shell (Escape too); tabs have no close buttons of their own. Closing only HIDES the shell — the Vue app stays mounted, so page-editor state (unsaved text included) survives until a real page reload; saving there is explicit (Ctrl+S) and refreshes the page regions in place. Admin panels never reload the page.
In-place page re-rendering shared by the banner/site/structure tabs lives in `swapdoc.js` (`runScripts`/`loadPlain`: fetch a page, swap the dynamic regions, replaceState). It also exports `dropPageCache`, which the editor tabs call after any save that can alter the rendered HTML of other pages (theme, headings, structure, banners, site brand/CSS, favicon). Dropping the cache while editing avoids re-fetching every page immediately; the public runtime re-preloads visible links once the editor panel closes.
## Saving behavior
Everything saves immediately as you edit (brand/title/CSS debounced, slug on commit since it renames the path), theme change swaps the stylesheet in place, tree rows navigate in place without transitions when focused, and the front page is a root-only row whose empty slug is editable like any other. Saves that can affect other pages drop the prefetch cache; the cache is rebuilt when the editor panel closes so navigation stays instant.
Every non-empty list (and the root) ends with a non-draggable plus footer row (vuedraggable `#footer` slot): clicking it starts a new pending page at that level (its slug placeholder shows the slug derived live from the title being typed), and while dragging it is the list's "end of list" drop target. Committing a pending page PUTs it with empty markdown (creates an empty page that renders with its title — saving never deletes; deletion is the page editor's explicit choice: saving trimmed-empty text issues a REST DELETE), then switches to the page editor tab for the actual writing.
Dropping ON the lower part of a row moves the page under that row (the child list's container invisibly overlaps its own row's bottom via negative margin — Sortable inserts it as the first child natively), while a row's exposed top edge inserts a sibling before it. Row indentation is structural (each nested list margin-indents itself), so a dragged row previews its whole subtree at the target list's depth.
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme and banner-design CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-user` elements to restore the canonical order (base < theme < design < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
+27
View File
@@ -0,0 +1,27 @@
# Frontend runtime
The public page runtime lives in `frontend/src/`.
## `main.js`
Vue editor app entry, mounts the tabbed `EditorShell`. See `docs/editing.md` for the editor UI.
## `pagerite.js`
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire). Editors can drop the entire cache with the `pagerite:drop-page-cache` event when site-wide or page changes (theme, headings, structure, banners, etc.) invalidate the cached HTML of other pages; `main.js` triggers a fresh `pagerite:preload-pages` pass when the editor panel closes so navigation is fast again. Navigation that starts while the editor is open bypasses the cache and fetches the target page on demand. Also runs scroll-reveal, OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, and the auth check.
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`.
Asset wiring differs by mode. In dev the backend links the Vite dev-server URLs (`pagerite:editor-src`/`-css`/`pagerite:analytics-src` meta tags, `<link>` stylesheets) and Vite injects the entry CSS from JS for hot reloads. In production there are no pagerite meta tags: all page assets are inlined into the document — stylesheets as `<style>` elements in `<head>` (fixed order: base, theme, banner design, entry sheets, custom CSS last), module scripts as inline `<script>`s at the end of the body (relative chunk imports are rewritten to absolute `/_assets/` paths) — and the on-demand bundles' URLs ride in a `<script type="application/json" id="pagerite-assets">` config. The editor bundle always stays external, imported on demand when a pen is opened. Every stylesheet element carries a stable id so fetch-navigation and the site editor can sync `<head>` positionally across swaps (the analytics sheet exists on `/_a` only and is added/removed as you navigate). The analytics entry is inlined into the `/_a` page itself; pagerite.js re-creates that script element after fetch-navigating there (inline scripts don't execute on a DOM swap) and calls the module's exposed unmount before swapping away.
## `assets/`
Shared styles and data files built by Vite and served hashed under `/_assets/`: `pagerite.css` (base layout + conservative variables), `pygments.css`, and `fonts/` (self-hosted Source Sans 3/Source Serif 4/Fraunces/Literata/Cormorant/Playfair Display/Inter/Montserrat/Fira Code/Cause/Exo 2/New Rocker variable woff2).
The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak. Themes and banner designs are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details.
Vite builds ES-module `.js` outputs; in dev the backend links them as `<script type="module">` (module scripts defer by default), in production it inlines them at the end of the body.
## Database file
The database file is `pagerite.kantadb` in the cwd (`PAGERITE_DB` overrides); gitignored. Do not delete it without asking.
+5
View File
@@ -0,0 +1,5 @@
# Pagerite overview
Pagerite is a single-user CMS/blog. FastAPI serves HTML rendered in Python with html5tagger; content is persisted in a kanta database and rendered on the fly per request. Vue is used only for interactive bits (editing tools), not for the public pages.
See `docs/design-principles.md` for the high-level design and the other `docs/*.md` files for implementation details.
+37
View File
@@ -0,0 +1,37 @@
# Themes and assets
## Built assets
Files under `frontend/src/assets/` are built by Vite and served hashed under `/_assets/`:
- `pagerite.css` — base layout + conservative variables.
- `pygments.css` — Pygments token styles mapped onto the `--code-*` variables.
- `fonts/` — self-hosted variable woff2 files for Source Sans 3, Source Serif 4, Fraunces, Literata, Cormorant, Playfair Display, Inter, Montserrat, Fira Code, Cause, Exo 2 and New Rocker.
The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak.
## Themes
Themes are folders in `pagerite/themes/{name}/` containing `theme.css` and/or `banner.css` (+ `banner.svg` artwork and any extra assets the CSS references, like summer's `grass.svg`). They are served by the backend at `/_themes/{name}/...` — read from disk per request (etag by mtime), never built, so on-disk edits show on the next page load even in prod.
`Data.theme` selects the active theme (empty = none/base only) and the site editor can switch it, choosing from the theme folders found on disk. Vue may add per-component styles on top where needed.
Current themes:
- `purple` — dark dusk palette with Fraunces/Literata and a tilted oversized gradient brand.
- `corporate` — light-first with automatic `prefers-color-scheme` dark mode, Montserrat/Inter and a huge solid brand.
- `nitro` — racing/HUD style following `prefers-color-scheme` (warm light-grey page, deep violet in dark), Montserrat/Literata, black as an accent only, a straight orange blade under the banner, and an orange racing-tab nav clipped with a bezier `shape()`.
- `summer` — light playful meadow, one palette sampled from its illustrated `banner.svg` (sky/grass/sun/flower pink), Fraunces/Literata, a tilted gradient brand, flower bullets, and a layered-parallax banner (sun rises, clouds drift, nearer hills move less) with idle animations (swaying flowers, floating clouds, breathing sun glow) wrapped in `prefers-reduced-motion: no-preference`.
## Banner designs
A theme folder may also ship a banner design (`banner.css` + `banner.html` arbitrary markup or `banner.svg`), selectable per page independently of the active theme. Standalone banner designs (no theme.css) ship as:
- `eyes` — a canvas critter in the grass.
- `stars` — a drifting starfield.
The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll parameter on `<html>` (event-driven, so it is still when the page is idle), the banner contents drift within their window (with scale overscan so no edge shows), and designs may key their own effects off the same parameter.
## Stylesheet order
The backend emits the stylesheets in a fixed order — base (Vite build), theme, banner design, entry sheets, custom CSS last — each with a stable id so fetch-navigation and the site editor can sync them in place. In dev they are `<link>`s (the base is Vite-injected from JS instead); in production they are inlined as `<style>` elements. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`. Code text (Fira Code by default) is optically matched to the body font by x-height: `font-size-adjust: ex-height var(--code-x-height)` scales whatever code font is in use, so a theme that switches its body font sets `--code-x-height` to that font's x-height ratio (base: 0.478 for Source Sans 3; themes ship values for Inter, Montserrat, Literata and Cause).
+1
View File
@@ -18,6 +18,7 @@
"@codemirror/view": "^6.43.8", "@codemirror/view": "^6.43.8",
"@lezer/highlight": "^1.2.3", "@lezer/highlight": "^1.2.3",
"codemirror": "^6.0.2", "codemirror": "^6.0.2",
"country-flag-icons": "^1.6.20",
"overlayscrollbars": "^2.16.0", "overlayscrollbars": "^2.16.0",
"transliteration": "^2.6.1", "transliteration": "^2.6.1",
"vue": "^3.5.26", "vue": "^3.5.26",
+499
View File
@@ -0,0 +1,499 @@
<script setup>
// Analytics viewer rendered as a normal page inside #main. Receives live
// analytics data over /_api/ws/analytics (admin-gated by the auth proxy) and
// renders totals, smoothed visit/views curves, a transition map, and recent
// visit/crawler tables. Read-only.
// See docs/analytics.md for the data format.
import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
import {
RANGES,
rangeWindow,
filterRecordsByRange,
filterTransitionsByRange,
filterViewsByRange,
} from './analytics/time.js'
import {
calcReadStats,
calcTotalViews,
copyIp,
copyList,
formatCount,
formatAbuseRows,
formatCrawlerRows,
formatVisitRows,
} from './analytics/format.js'
import TrailLink from './TrailLink.vue'
import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue'
import { VIEW_W } from './analytics/chart.js'
// Same centering margin as the charts, so the totals row's left edge
// aligns with the chart svg above the natural width.
const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const ABUSE_MAX_LINES = 5
const data = ref(null)
const pageTree = ref(null)
const error = ref('')
const now = ref(Date.now())
let ws = null
let reconnectTimeout = null
let timeInterval = null
// The initial range comes from the URL hash (shareable links); without one,
// it is derived from the first analytics snapshot: day when the recorded
// history is shorter than 24 h, week otherwise.
const hashRange = location.hash.slice(1)
const range = ref(RANGES[hashRange] ? hashRange : 'week')
let rangePinned = Boolean(RANGES[hashRange])
function connectAnalytics() {
if (ws) return
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
ws.onopen = () => { error.value = '' }
ws.onmessage = (event) => {
try {
data.value = JSON.parse(event.data)
if (!rangePinned) {
rangePinned = true
const starts = (data.value?.visits || [])
.map((v) => Date.parse(v.start))
.filter((t) => !Number.isNaN(t))
if (starts.length && Date.now() - Math.min(...starts) < 24 * 3600 * 1000) {
range.value = 'day'
}
}
} catch {
error.value = 'analytics data could not be loaded'
}
}
ws.onerror = () => {
error.value = 'analytics data could not be loaded'
}
ws.onclose = () => {
ws = null
reconnectTimeout = setTimeout(connectAnalytics, 2000)
}
}
onMounted(async () => {
connectAnalytics()
now.value = Date.now()
timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
// The site tree for the transition map (all pages in menu order). Not
// fatal: without it the map just narrows to pages seen in transitions.
try {
const res = await fetch('/_api/pages')
if (res.ok) pageTree.value = await res.json()
} catch { /* map just narrows to pages seen in transitions */ }
})
onUnmounted(() => {
if (reconnectTimeout) clearTimeout(reconnectTimeout)
if (timeInterval) clearInterval(timeInterval)
if (ws) {
ws.onclose = null
ws.close()
ws = null
}
})
const window = computed(() => rangeWindow(range.value))
// All non-chart stats follow the selected range; the charts keep their own
// range-specific x windows (week overlays previous weeks aligned to Monday).
const rangeData = computed(() => {
if (!data.value) return null
const { t0, t1 } = window.value
return {
...data.value,
transitions: filterTransitionsByRange(data.value.transitions, t0, t1),
views: filterViewsByRange(data.value.views, t0, t1),
visits: filterRecordsByRange(data.value.visits, t0, t1),
crawlers: filterRecordsByRange(data.value.crawlers, t0, t1),
abuse: filterRecordsByRange(data.value.abuse, t0, t1),
}
})
const visits = computed(() => rangeData.value?.visits || [])
const totalViews = computed(() => calcTotalViews(rangeData.value?.views))
const readStats = computed(() => calcReadStats(visits.value))
// Keep the URL shareable when the range changes.
watch(range, (r) => {
const url = new URL(location.href)
url.hash = r
history.replaceState(null, '', url)
})
const clients = computed(() => data.value?.clients || {})
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
const crawlers = computed(() => rangeData.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, now.value))
</script>
<template>
<div class="analytics-view">
<div class="analytics-panel">
<header>
<h1>Analytics</h1>
<nav class="ranges">
<button v-for="(r, key) in RANGES" :key="key" type="button"
:class="{ active: range === key }" @click="range = key">
{{ r.label }}
</button>
</nav>
<a href="/" class="close" title="home"></a>
</header>
<p v-if="error" class="error"> {{ error }}</p>
<p v-else-if="!data" class="loading">loading</p>
<template v-else>
<section class="totals" :style="{ marginLeft: CHART_MARGIN }">
<div><strong :title="String(visits.length)">{{ formatCount(visits.length) }}</strong> visits</div>
<div><strong :title="String(totalViews)">{{ formatCount(totalViews) }}</strong> page views</div>
<div><strong>{{ readStats.avgMinPerVisit }}</strong> min/visit</div>
<div><strong>{{ readStats.avgArticleMedianMin }}</strong> min/read</div>
</section>
<VisitorCharts :data="data" :range="range" />
<TransitionGraph :data="rangeData" :window="window" :page-tree="pageTree" />
<section>
<h2>Recent visits</h2>
<div v-if="visitRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>trail</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(v, i) in visitRows" :key="i">
<td class="trail">
<TrailLink v-if="v.refererStep" :step="v.refererStep" @close="$emit('close')" />
<span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span>
<TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" @close="$emit('close')" />
</td>
<VisitorCell
:ip="v.ip"
:ip-display="v.ipDisplay"
:ua="v.ua"
:ua-raw="v.uaRaw"
:country="v.country"
:city="v.city"
:lang="v.lang"
:lang-display="v.langDisplay"
:is-host="v.isHost"
/>
<td class="last-seen muted"
:title="v.lastSeenLocal"
@click="copyList(v.lastSeenIso, $event)">{{ v.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
<p v-else class="empty">no visits recorded yet</p>
<div v-if="crawlerRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>pages crawled</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(c, i) in crawlerRows" :key="i">
<td class="trail">
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
</td>
<VisitorCell
:ip="c.ip"
:ip-display="c.ipDisplay"
:ua="c.ua"
:ua-raw="c.uaRaw"
:country="c.country"
:city="c.city"
:lang="c.lang"
:lang-display="c.langDisplay"
:is-host="c.isHost"
/>
<td class="last-seen muted"
:title="c.lastSeenLocal"
@click="copyList(c.lastSeenIso, $event)">{{ c.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
<div v-if="abuseRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>paths abused</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(a, i) in abuseRows" :key="i">
<td class="trail abuse-list clickable-list"
@click="copyList(a.allPaths, $event)">
<div class="abuse-items">
<span v-for="(p, pi) in a.paths.slice(0, ABUSE_MAX_LINES)" :key="pi"
class="inline-item">
<small v-if="p.count > 1" class="muted">{{ formatCount(p.count) }}×</small>{{ p.path }}
</span>
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
</div>
</td>
<VisitorCell
:ip="a.ip"
:ip-display="a.ipDisplay"
:ua="a.ua"
:ua-raw="a.uaRaw"
:country="a.country"
:city="a.city"
:lang="a.lang"
:lang-display="a.langDisplay"
:is-host="a.isHost"
:variant-count="a.clientCount"
/>
<td class="last-seen muted"
:title="a.lastSeenLocal"
@click="copyList(a.lastSeenIso, $event)">{{ a.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
</section>
</template>
</div>
</div>
</template>
<style scoped>
.analytics-view {
min-height: 100vh;
background: var(--bg, Canvas);
color: var(--text, CanvasText);
}
.analytics-panel {
margin: 0;
width: 100%;
/* Same 1.25rem side spacing as main's article padding. */
padding: 1.5rem 1.25rem 4rem;
/* Container for cqw-based shrink-to-fit (see .totals). */
container-type: inline-size;
}
.analytics-panel header {
display: flex;
align-items: center;
gap: 1rem;
}
.analytics-panel h1 {
margin: 0;
font-size: 1.4rem;
}
.ranges {
display: flex;
gap: 0.25rem;
margin-left: auto;
}
.ranges button {
padding: 0.2rem 0.7rem;
font: inherit;
font-size: 0.9rem;
color: var(--muted);
background: none;
border: 1px solid var(--line);
border-radius: 1rem;
cursor: pointer;
}
.ranges button:hover { color: var(--text); }
.ranges button.active {
color: var(--text);
border-color: var(--accent);
}
.close {
padding: 0 0.3rem;
background: none;
border: none;
color: var(--muted);
font-size: 1.2rem;
cursor: pointer;
}
.close:hover { color: var(--text); }
.analytics-panel h2 {
margin: 0 0 0.6rem;
font-size: 1rem;
color: var(--muted);
}
.analytics-panel section {
margin-top: 1.8rem;
}
.analytics-view a {
color: var(--text);
text-decoration: none;
}
.analytics-view a:hover { color: var(--accent); }
.analytics-view :deep(.muted) { color: var(--muted); }
.analytics-view :deep(.small) { font-size: 0.75em; }
/* One line at any width: the gap shrinks first, then the font (the number
scales along in em), both following the panel's container width. */
.totals {
display: flex;
gap: clamp(0.5rem, 3cqw, 2rem);
font-size: clamp(0.6rem, 2.2cqw, 1.1rem);
white-space: nowrap;
}
.totals strong { font-size: 1.36em; }
.visit-table-wrap {
overflow-x: auto;
}
.visit-table {
width: 100%;
border-collapse: collapse;
font-size: 0.9rem;
line-height: 1.3;
}
.visit-table th,
.visit-table td {
padding: 0.25rem 0.5rem;
border-bottom: 1px solid var(--line);
text-align: left;
vertical-align: top;
}
.visit-table th {
color: var(--muted);
font-weight: normal;
text-transform: lowercase;
position: sticky;
top: 0;
background: var(--bg, Canvas);
}
.visit-table .last-seen {
width: 6rem;
text-align: right;
white-space: nowrap;
cursor: pointer;
}
.visit-table .trail {
max-width: 20rem;
overflow-wrap: break-word;
}
.visit-table .trail a,
.visit-table .trail-link {
display: inline-block;
max-width: 8rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .trail > * + * {
margin-left: 0.5rem;
}
.analytics-view :deep(.trail-link.error),
.analytics-view :deep(.trail-link.error:hover) {
color: var(--error, #c00);
}
.visit-table .utm-tag {
display: inline-block;
max-width: 100%;
padding: 0.05rem 0.4rem;
border: 1px solid var(--line);
border-radius: 0.25rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .clickable-list {
cursor: pointer;
max-width: 22rem;
}
.visit-table .abuse-items {
display: flex;
flex-wrap: wrap;
gap: 0.15rem 0.5rem;
align-items: baseline;
}
.visit-table .inline-item {
max-width: 18rem;
min-width: 0;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
word-break: keep-all;
hyphens: none;
}
.visit-table :deep(.clickable-ip),
.visit-table .clickable-list,
.visit-table .last-seen {
cursor: pointer;
position: relative;
}
.visit-table :deep(.copy-popup) {
position: absolute;
bottom: calc(100% + 0.25rem);
left: 50%;
transform: translateX(-50%);
padding: 0.15rem 0.4rem;
background: var(--text, CanvasText);
color: var(--bg, Canvas);
border-radius: 0.25rem;
font-size: 0.75rem;
white-space: nowrap;
pointer-events: none;
z-index: 10;
}
.crawler-top-uas {
font-size: 0.9rem;
margin-bottom: 0.6rem;
}
.crawler-top-uas strong {
color: var(--muted);
}
.empty, .loading, .error { color: var(--muted); }
.error { color: var(--error, #c00); }
</style>
+3 -1
View File
@@ -6,7 +6,7 @@ import { EditorView, basicSetup } from 'codemirror'
import { EditorState } from '@codemirror/state' import { EditorState } from '@codemirror/state'
import { html } from '@codemirror/lang-html' import { html } from '@codemirror/lang-html'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
import { loadPlain, runScripts } from './swapdoc' import { dropPageCache, loadPlain, runScripts } from './swapdoc'
const props = defineProps({ const props = defineProps({
pagePath: { type: String, default: '' }, pagePath: { type: String, default: '' },
@@ -215,6 +215,8 @@ function onMessage(ev) {
} else if (msg.type === 'saved') { } else if (msg.type === 'saved') {
saveError.value = '' saveError.value = ''
pendingSave = null pendingSave = null
// Banner HTML/design changes affect the rendered page; invalidate prefetches.
dropPageCache()
refreshOnSave?.() refreshOnSave?.()
refreshOnSave = null refreshOnSave = null
} else if (msg.type === 'error') { } else if (msg.type === 'error') {
+45 -15
View File
@@ -4,9 +4,9 @@
// (/_api/ws/editor). Docked left of the article on the page itself. // (/_api/ws/editor). Docked left of the article on the page itself.
// The socket connects when the editor is opened and reconnects with // The socket connects when the editor is opened and reconnects with
// exponential backoff after a failure; unsaved text and pending saves // exponential backoff after a failure; unsaved text and pending saves
// survive a disconnect. Editor scroll drives the article scroll (while // survive a disconnect. Editor and article (window) scrolls are linked
// editing the window scroll is locked and only #main scrolls), keeping the // proportionally both ways (syncWindowToEditor / syncEditorToWindow).
// rendered article at the cursor's position. Saving (💾 / Ctrl+S) is explicit // Saving (💾 / Ctrl+S) is explicit
// and refreshes the page regions in place — never a reload — so the editor // and refreshes the page regions in place — never a reload — so the editor
// state (unsaved text included) also survives closing the shell; it is lost // state (unsaved text included) also survives closing the shell; it is lost
// only on a real page reload. // only on a real page reload.
@@ -15,7 +15,7 @@ import { EditorView, basicSetup } from 'codemirror'
import { EditorState } from '@codemirror/state' import { EditorState } from '@codemirror/state'
import { markdown } from '@codemirror/lang-markdown' import { markdown } from '@codemirror/lang-markdown'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
import { loadPlain } from './swapdoc' import { dropPageCache, loadPlain } from './swapdoc'
const props = defineProps({ const props = defineProps({
pagePath: { type: String, default: '' }, pagePath: { type: String, default: '' },
@@ -113,7 +113,9 @@ async function saveAndRefresh() {
await save() await save()
dirty.value = false dirty.value = false
// Refresh the page regions from the server so nav/sidebar changes apply // Refresh the page regions from the server so nav/sidebar changes apply
// (never a reload: the editor keeps its state). // (never a reload: the editor keeps its state). Drop the prefetch cache
// first: heading/title changes affect navigation on every page.
dropPageCache()
loadPlain(path.value) loadPlain(path.value)
} }
@@ -307,20 +309,42 @@ function onEditorShown() {
if (dirty.value) requestRender() if (dirty.value) requestRender()
} }
function syncScroll() { // Bidirectional proportional scroll sync between the CodeMirror scroller
// Editor scroll drives the article: keep the rendered page at the same // and the window (the article's scroller, also while editing). Both
// proportional position as the cursor area in the editor. While editing // directions apply instantly (never smooth — a smooth window scroll feeds
// the window scroll is locked and #main is the scrolling element. // its intermediate positions back into the editor and fights the user's
// scrolling) and coalesce to one update per frame. Loops are broken two
// ways: a driver flag held until one frame AFTER the write (the scroll
// event a programmatic write dispatches arrives asynchronously — clearing
// the flag in the writing frame would let the echo through and the two
// directions would chase each other, which showed up as random jumping
// whenever layout shifted the proportional targets mid-scroll), and a 1px
// tolerance so residual rounding is a no-op. When the panel's height
// changes mid-scroll (its top tracks the banner), the page is the driver:
// the editor is re-matched to the page's position, never vice versa.
function syncWindowToEditor() {
if (syncingScroll || !view) return if (syncingScroll || !view) return
const main = document.getElementById('main')
if (!main) return
syncingScroll = true syncingScroll = true
requestAnimationFrame(() => { requestAnimationFrame(() => {
const scroller = view.scrollDOM const scroller = view.scrollDOM
const max = scroller.scrollHeight - scroller.clientHeight const max = scroller.scrollHeight - scroller.clientHeight
const pct = max > 0 ? scroller.scrollTop / max : 0 const pct = max > 0 ? scroller.scrollTop / max : 0
main.scrollTop = pct * (main.scrollHeight - main.clientHeight) const y = pct * Math.max(0, document.documentElement.scrollHeight - innerHeight)
syncingScroll = false if (Math.abs(scrollY - y) > 1) scrollTo({ top: y, behavior: 'instant' })
requestAnimationFrame(() => { syncingScroll = false })
})
}
function syncEditorToWindow() {
if (syncingScroll || !view) return
syncingScroll = true
requestAnimationFrame(() => {
const scroller = view.scrollDOM
const pageMax = Math.max(0, document.documentElement.scrollHeight - innerHeight)
const pct = pageMax > 0 ? scrollY / pageMax : 0
const top = pct * Math.max(0, scroller.scrollHeight - scroller.clientHeight)
if (Math.abs(scroller.scrollTop - top) > 1) scroller.scrollTop = top
requestAnimationFrame(() => { syncingScroll = false })
}) })
} }
@@ -379,7 +403,11 @@ onMounted(() => {
}), }),
parent: editorEl.value, parent: editorEl.value,
}) })
view.scrollDOM.addEventListener('scroll', syncScroll) view.scrollDOM.addEventListener('scroll', syncWindowToEditor)
// Page → editor: window scroll (and resizes, e.g. the panel growing when
// the banner scrolls away) re-match the editor to the page's position.
addEventListener('scroll', syncEditorToWindow, { passive: true })
addEventListener('resize', syncEditorToWindow)
// Opening the editor means you want to write: start focused. // Opening the editor means you want to write: start focused.
view.focus() view.focus()
window.__pageritePageEditor = { window.__pageritePageEditor = {
@@ -399,6 +427,8 @@ onUnmounted(() => {
} }
view?.destroy() view?.destroy()
delete window.__pageritePageEditor delete window.__pageritePageEditor
removeEventListener('scroll', syncEditorToWindow)
removeEventListener('resize', syncEditorToWindow)
removeEventListener('keydown', onKeydown) removeEventListener('keydown', onKeydown)
removeEventListener('pagerite:editor-shown', onEditorShown) removeEventListener('pagerite:editor-shown', onEditorShown)
}) })
@@ -606,7 +636,7 @@ onUnmounted(() => {
/* CodeMirror sits inside a bordered box, like a dialog's input area, with /* CodeMirror sits inside a bordered box, like a dialog's input area, with
a slight margin to the panel edges. Wheel scroll stays in the editor and a slight margin to the panel edges. Wheel scroll stays in the editor and
drives the article (syncScroll) instead of double-scrolling. */ drives the article (syncWindowToEditor) instead of double-scrolling. */
.editor { .editor {
flex: 1; flex: 1;
min-width: 0; min-width: 0;
+51 -41
View File
@@ -8,7 +8,7 @@ import { EditorState } from '@codemirror/state'
import { css } from '@codemirror/lang-css' import { css } from '@codemirror/lang-css'
import { html } from '@codemirror/lang-html' import { html } from '@codemirror/lang-html'
import { cmHighlight, cmTheme } from './cmtheme' import { cmHighlight, cmTheme } from './cmtheme'
import { loadPlain, runScripts } from './swapdoc' import { dropPageCache, loadPlain, runScripts } from './swapdoc'
const props = defineProps({ const props = defineProps({
pagePath: { type: String, default: '' }, pagePath: { type: String, default: '' },
@@ -110,6 +110,7 @@ async function uploadFavicon(file) {
const { path: url } = await res.json() const { path: url } = await res.json()
favicon.value = url favicon.value = url
applyFavicon(url) applyFavicon(url)
dropPageCache()
} else { } else {
saveError.value = `⚠️ ${await errorDetail(res)}` saveError.value = `⚠️ ${await errorDetail(res)}`
} }
@@ -234,6 +235,7 @@ async function saveSettings(opts = {}) {
}) })
if (res.ok) { if (res.ok) {
saveError.value = '' saveError.value = ''
dropPageCache()
} else { } else {
saveError.value = '⚠️ changes could not be saved' saveError.value = '⚠️ changes could not be saved'
} }
@@ -242,30 +244,43 @@ async function saveSettings(opts = {}) {
async function onThemeChange() { async function onThemeChange() {
await saveSettings() await saveSettings()
// Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev // Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev
// and prod: swap the link in place, then re-render (the theme's default // and prod, but rendered differently: a <link> in dev, an inline <style>
// banner design and the page's stylesheet links may change with it). // in prod. Swap it in place, then re-render (the theme's default banner
let link = document.getElementById('pagerite-theme') // design and the page's stylesheets may change with it).
let el = document.getElementById('pagerite-theme')
const url = `/_themes/${theme.value}/theme.css`
if (theme.value) { if (theme.value) {
const href = `/_themes/${theme.value}/theme.css` if (el?.tagName === 'STYLE') {
if (link) { el.textContent = await (await fetch(url)).text()
link.href = href } else if (el) {
} else { el.href = url
} else if (import.meta.env.DEV) {
// Re-create after "none": keep base < theme < design < custom CSS. // Re-create after "none": keep base < theme < design < custom CSS.
// In dev there is no #pagerite-base link (the base is a // In dev there is no #pagerite-base element (the base is a
// Vite-injected <style>), so anchor to the next sheet instead of // Vite-injected <style>), so anchor to the next sheet instead of
// prepending before the base styles. // prepending before the base styles.
link = document.createElement('link') el = document.createElement('link')
link.rel = 'stylesheet' el.rel = 'stylesheet'
link.id = 'pagerite-theme' el.id = 'pagerite-theme'
link.href = href el.href = url
const before = document.getElementById('pagerite-base')?.nextSibling const before = document.getElementById('pagerite-base')?.nextSibling
?? document.getElementById('pagerite-banner') ?? document.getElementById('pagerite-banner')
?? document.getElementById('pagerite-user') ?? document.getElementById('pagerite-user')
if (before) before.before(link) if (before) before.before(el)
else document.head.append(link) else document.head.append(el)
} else {
// Prod: inline <style>, fetched from the backend-served URL.
el = document.createElement('style')
el.id = 'pagerite-theme'
el.textContent = await (await fetch(url)).text()
const before = document.getElementById('pagerite-base')?.nextSibling
?? document.getElementById('pagerite-banner')
?? document.getElementById('pagerite-user')
if (before) before.before(el)
else document.head.append(el)
} }
} else if (link) { } else if (el) {
link.remove() el.remove()
} }
loadPlain(path.value) loadPlain(path.value)
} }
@@ -561,7 +576,7 @@ onUnmounted(() => {
</div> </div>
</section> </section>
<section class="block" @paste="onBrandPaste"> <section class="block grow" @paste="onBrandPaste">
<div class="block-head"> <div class="block-head">
<span class="field-label">brand code (replaces the brand link)</span> <span class="field-label">brand code (replaces the brand link)</span>
<button <button
@@ -581,7 +596,7 @@ onUnmounted(() => {
<div ref="brandEl" class="brand-cm" /> <div ref="brandEl" class="brand-cm" />
</section> </section>
<section class="block"> <section class="block grow">
<div class="block-head"> <div class="block-head">
<span class="field-label">custom CSS (applies to every page, on top of the theme)</span> <span class="field-label">custom CSS (applies to every page, on top of the theme)</span>
</div> </div>
@@ -595,6 +610,7 @@ onUnmounted(() => {
display: flex; display: flex;
flex-direction: column; flex-direction: column;
overflow-y: auto; overflow-y: auto;
background: var(--surface);
} }
.block { .block {
@@ -606,6 +622,13 @@ onUnmounted(() => {
background: var(--surface); background: var(--surface);
} }
/* Editor blocks (brand HTML, custom CSS) share the leftover panel height
equally; their CodeMirror windows fill the block and scroll internally. */
.block.grow {
flex: 1;
min-height: 7rem;
}
.block-head { .block-head {
display: flex; display: flex;
align-items: center; align-items: center;
@@ -775,42 +798,29 @@ onUnmounted(() => {
border-color: var(--accent); border-color: var(--accent);
} }
/* Small CodeMirror window for the custom brand HTML; scrolls internally. */ /* CodeMirror windows for the brand HTML and custom CSS: fill the growing
.brand-cm { block, scroll internally. */
border: 1px solid var(--line); .brand-cm,
border-radius: 4px;
overflow: hidden;
}
.brand-cm :deep(.cm-editor) {
max-height: 7rem;
font-size: 0.85rem;
}
.brand-cm :deep(.cm-scroller) {
overflow: auto;
}
.brand-cm :deep(.cm-gutters) {
display: none;
}
/* CodeMirror window for site-wide custom CSS. */
.css-cm { .css-cm {
flex: 1;
min-height: 0;
border: 1px solid var(--line); border: 1px solid var(--line);
border-radius: 4px; border-radius: 4px;
overflow: hidden; overflow: hidden;
} }
.brand-cm :deep(.cm-editor),
.css-cm :deep(.cm-editor) { .css-cm :deep(.cm-editor) {
max-height: 12rem; height: 100%;
font-size: 0.85rem; font-size: 0.85rem;
} }
.brand-cm :deep(.cm-scroller),
.css-cm :deep(.cm-scroller) { .css-cm :deep(.cm-scroller) {
overflow: auto; overflow: auto;
} }
.brand-cm :deep(.cm-gutters),
.css-cm :deep(.cm-gutters) { .css-cm :deep(.cm-gutters) {
display: none; display: none;
} }
+5 -1
View File
@@ -10,7 +10,7 @@
import { inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue' import { inject, onActivated, onMounted, onUnmounted, provide, ref, watch } from 'vue'
import StructureTree from './StructureTree.vue' import StructureTree from './StructureTree.vue'
import { slugify } from './slugify' import { slugify } from './slugify'
import { loadPlain } from './swapdoc' import { dropPageCache, loadPlain } from './swapdoc'
const props = defineProps({ const props = defineProps({
pagePath: { type: String, default: '' }, pagePath: { type: String, default: '' },
@@ -144,6 +144,7 @@ async function commitPending() {
} }
pending.value = null pending.value = null
await refreshPages() await refreshPages()
dropPageCache()
await navigate(newPath) await navigate(newPath)
// Hand over to the page editor tab for the actual writing. // Hand over to the page editor tab for the actual writing.
shell?.switchMode('page') shell?.switchMode('page')
@@ -171,6 +172,8 @@ async function postStructure(op) {
}) })
if (res.ok) { if (res.ok) {
saveError.value = '' saveError.value = ''
// Structure changes alter navigation on every page; drop prefetches.
dropPageCache()
loadPlain(path.value) // refresh menus and content from the server loadPlain(path.value) // refresh menus and content from the server
} else { } else {
saveError.value = `⚠️ ${await errorDetail(res)}` saveError.value = `⚠️ ${await errorDetail(res)}`
@@ -234,6 +237,7 @@ async function removePage(node) {
if (res.ok) { if (res.ok) {
saveError.value = '' saveError.value = ''
refreshPages() refreshPages()
dropPageCache()
const p = node.path const p = node.path
if (p === path.value || (p && path.value.startsWith(`${p}/`))) { if (p === path.value || (p && path.value.startsWith(`${p}/`))) {
// The current page was deleted — or reduced to a category, which now // The current page was deleted — or reduced to a category, which now
+37
View File
@@ -0,0 +1,37 @@
<script setup>
import { computed } from 'vue'
import { formatCount, formatReadTime } from './analytics/format.js'
const props = defineProps({
step: { type: Object, required: true },
count: { type: Number, default: 0 },
})
defineEmits(['close'])
const hasError = computed(() => props.step.status >= 400)
const title = computed(() => {
const parts = [props.step.title]
if (props.step.readSeconds > 0) {
parts.push(formatReadTime(props.step.readSeconds))
}
if (hasError.value) {
parts.push(`${props.step.status}`)
}
return parts.filter(Boolean).join(' — ')
})
</script>
<template>
<a class="trail-link"
:class="{ error: hasError }"
:href="step.path"
:title="title"
:target="step.external ? '_blank' : undefined"
:rel="step.external ? 'noopener' : undefined"
@click="(e) => { if (!step.external) $emit('close') }">
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
<span>{{ step.slug }}</span>
</a>
</template>
+286
View File
@@ -0,0 +1,286 @@
<script setup>
/**
* Radial transition map for a pre-filtered time range.
*
* The parent filters transitions, views and visits to the selected range
* before passing them in; `window` carries the absolute [t0, t1) window
* so the visual scale can normalize against a one-week reference.
*/
import { computed, onBeforeUnmount, onMounted, shallowRef, watch } from 'vue'
import { DAY } from './analytics/time.js'
import { formatCount, formatReadTime } from './analytics/format.js'
import {
TNODE_W,
TNODE_H,
BEAD_R,
buildTransitionGraph,
} from './analytics/transitions.js'
const props = defineProps({
data: { type: Object, default: null },
window: { type: Object, required: true },
pageTree: { type: Array, default: null },
})
const dayScale = computed(() => {
const { t0, t1 } = props.window
// Convert raw counts to a daily hit rate (hits/day).
if (t0 != null && t1 != null) return DAY / (t1 - t0)
// 'all': scale by the actual data span, but never less than the 30-day
// minimum the plot enforces, so sparse young data is not over-amplified.
const times = new Set()
for (const buckets of Object.values(props.data?.views || {})) {
for (const k of Object.keys(buckets)) times.add(Date.parse(k))
}
const arr = [...times]
if (arr.length < 2) return 1
const span = Math.max(...arr) - Math.min(...arr)
return DAY / Math.max(span, 30 * DAY)
})
const graph = computed(() =>
props.data
? buildTransitionGraph(props.data, props.pageTree, props.data.visits || [], dayScale.value)
: null,
)
// Bead animation: every bead is simulated independently in JS. Each flow
// (one per edge direction) emits a bead every `interval` seconds; beads
// cross their segment in a constant TRAVERSAL_S seconds (speed relative
// to span length) and are dropped at the end.
// There is deliberately no cap on beads in flight.
// Emitters persist across data reloads, keyed by flow.key: an unchanged
// link keeps its emission phase and in-flight beads (tracked by progress,
// not absolute time), so a count change elsewhere never reshuffles them.
const beads = shallowRef([])
let rafId = 0
const emitters = new Map() // flow.key -> { flow, interval, next, alive }
const live = [] // { e, p } — beads in flight, p = progress 0..1
let lastTick = 0
const MAX_BEAD_RATE = 120 // upper bound on total beads per second
const TRAVERSAL_S = 0.4 // seconds to cross any segment, end to end
const syncBeads = (flows) => {
const reduced = matchMedia('(prefers-reduced-motion: reduce)').matches
if (!flows?.length || reduced) {
emitters.clear()
live.length = 0
beads.value = []
return
}
// Cap the total bead emission rate so a busy range cannot spawn enough
// beads to kill the page. Existing per-range time scaling is preserved;
// this is only a proportional emergency throttle when the limit is hit.
const totalRate = flows.reduce((s, f) => s + 1 / f.interval, 0)
const scale = totalRate > MAX_BEAD_RATE ? MAX_BEAD_RATE / totalRate : 1
const now = performance.now()
const seen = new Set()
for (const flow of flows) {
seen.add(flow.key)
const interval = (flow.interval / scale) * 1000
const e = emitters.get(flow.key)
if (e) {
e.flow = flow // pick up new geometry/rate, keep the phase
e.interval = interval
continue
}
// New emitter: pre-fill the traversal with evenly spaced beads (random
// phase), so the flow appears already running instead of empty.
const phase = Math.random() * interval
const dp = interval / 1000 / TRAVERSAL_S
const ne = { flow, interval, next: now + phase, alive: true }
for (let p = 1 - phase / 1000 / TRAVERSAL_S; p > 0; p -= dp) {
live.push({ e: ne, p })
}
emitters.set(flow.key, ne)
}
for (const [key, e] of emitters) {
if (!seen.has(key)) {
e.alive = false
emitters.delete(key)
}
}
for (let i = live.length - 1; i >= 0; i--) {
if (!live[i].e.alive) live.splice(i, 1)
}
}
const tick = (t) => {
const dt = lastTick ? (t - lastTick) / 1000 : 0
lastTick = t
for (const e of emitters.values()) {
while (e.next <= t) {
live.push({ e, p: 0 })
e.next += e.interval
}
}
const out = []
for (let i = live.length - 1; i >= 0; i--) {
const b = live[i]
b.p += dt / TRAVERSAL_S
if (b.p >= 1) {
live.splice(i, 1)
continue
}
const f = b.e.flow
out.push({ x: f.x1 + (f.x2 - f.x1) * b.p, y: f.y1 + (f.y2 - f.y1) * b.p })
}
beads.value = out
rafId = requestAnimationFrame(tick)
}
watch(() => graph.value?.flows, syncBeads, { immediate: true })
onMounted(() => {
if (!matchMedia('(prefers-reduced-motion: reduce)').matches) {
rafId = requestAnimationFrame(tick)
}
})
onBeforeUnmount(() => cancelAnimationFrame(rafId))
// The svg never renders larger than its natural size (1 viewBox unit = 1
// px, max-width below): the layout geometry is designed in pixel-like
// units, and upscaling would blow up the pills around their text. Narrow
// panels scale the graph down to fit (width: 100%), text along with it.
// Pill text is not truncated: text is clipped at the pill's rounded border
// (clipPath per node, inset a few units for padding). Captions center when
// they fit; overlong ones anchor left so their beginning (not their
// middle) survives the clip. Width estimate: ~0.52 em per glyph.
const fitsPill = (label, fontPx = 19) => label.length * 0.52 * fontPx <= TNODE_W - 16
const countLabel = (n) =>
n.readSec ? `${formatCount(n.views)}×${formatReadTime(n.readSec)}` : formatCount(n.views)
</script>
<template>
<section v-if="graph">
<svg class="tmap" :style="{ maxWidth: `${graph.bounds.x1 - graph.bounds.x0}px` }" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
role="img" aria-label="map of transitions between pages">
<path v-for="(a, i) in graph.arcs" :key="'a' + i"
:id="`tarc${i}`" :d="a.d" :class="['tarc', a.top && 'tarc-top']" />
<template v-for="(a, i) in graph.arcs" :key="'t' + i">
<path v-if="a.ld" :id="`tarcl${i}`" :d="a.ld" fill="none" stroke="none" />
<text v-if="a.ld" class="tarclabel" :class="{ 'tarclabel-top': a.top }"><textPath :href="`#tarcl${i}`" startOffset="0">{{ a.label }}</textPath></text>
</template>
<path v-for="(e, i) in graph.edges" :key="'e' + i"
:d="e.d" :class="['tconn', e.external && 'tconn-exit']">
<title>{{ e.title }}</title>
</path>
<circle v-for="(b, i) in beads" :key="'b' + i"
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
<g v-for="(x, i) in graph.extNodes" :key="'x' + i">
<clipPath :id="`xclip${i}`">
<rect :x="x.x - TNODE_W/2 + 6" :y="x.y - TNODE_H/2" :width="TNODE_W - 12"
:height="TNODE_H" :rx="TNODE_H/2 - 4" />
</clipPath>
<a v-if="x.href" :href="x.href" target="_blank" rel="noopener">
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<text :x="fitsPill(x.label) ? x.x : x.x - TNODE_W/2 + 8" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(x.label) ? 'middle' : 'start' }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</a>
<g v-else>
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<text :x="fitsPill(x.label) ? x.x : x.x - TNODE_W/2 + 8" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(x.label) ? 'middle' : 'start' }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</g>
</g>
<g v-for="(n, i) in graph.nodes" :key="n.path">
<clipPath :id="`nclip${i}`">
<rect :x="n.x - TNODE_W/2 + 6" :y="n.y - TNODE_H/2" :width="TNODE_W - 12"
:height="TNODE_H" :rx="TNODE_H/2 - 4" />
</clipPath>
<a :href="n.path">
<title>{{ n.title }}</title>
<rect :x="n.x - TNODE_W/2" :y="n.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2" class="tnode" />
<g :clip-path="`url(#nclip${i})`">
<text :x="fitsPill(n.label) ? n.x : n.x - TNODE_W/2 + 8" :y="n.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(n.label) ? 'middle' : 'start' }">{{ n.label }}</text>
<text :x="n.x" :y="n.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">
{{ countLabel(n) }}
</text>
</g>
</a>
</g>
</svg>
</section>
</template>
<style scoped>
/* Transition map: radial graph of internal page-to-page transitions. */
.tmap {
display: block;
width: 100%;
/* max-width is set inline to the natural content width (px = viewBox
units), so wide panels never upscale the graph beyond 1:1. */
margin: 0 auto;
}
.tmap .tconn {
fill: var(--accent);
opacity: 0.4; /* uniform, not strength-encoded: width carries that */
}
.tmap .tconn-exit {
fill: var(--text);
}
.tmap .tbead {
fill: var(--accent);
opacity: 0.85;
filter: drop-shadow(0 0 2.5px var(--accent));
}
.tmap .txnode {
fill: var(--text);
stroke: none;
}
.tmap .txnode-source { fill: var(--text); }
.tmap .txnode-exit { fill: var(--text); }
/* Branch lanes: one wide concentric arc per path prefix, running behind
the node pills around the fan's circle center; parent levels sit one
indent (radius step) outward. Each lane's label follows a short guide
arc across the first inter-node gap (the part pills never cover). */
.tmap .tarc {
fill: none;
stroke: var(--muted);
stroke-width: 16;
opacity: 0.25;
}
.tmap .tarc-top { stroke-width: 24; }
/* Lane labels are left-aligned: each guide arc starts just past the source
pill's edge, the earliest point where the text is visible. */
.tmap .tarclabel {
fill: var(--muted);
font-size: 13px;
text-anchor: start;
}
/* The top lane is 50% thicker; its 🏠︎ label scales along. */
.tmap .tarclabel-top {
font-size: 19.5px;
}
.tmap .tnode {
fill: var(--accent);
stroke: none;
}
/* Text sizes are viewBox units: they shrink along with the graph on
narrow panels. Overlong labels are clipped at the pill border. */
.tmap .tnodeslug {
fill: var(--bg, Canvas);
font-size: 19px;
text-anchor: start;
}
.tmap a { cursor: pointer; }
.tmap .tnodecount {
fill: var(--bg, Canvas);
opacity: 0.75;
font-size: 15px;
text-anchor: middle;
}
section { margin-top: 1.8rem; }
</style>
+156
View File
@@ -0,0 +1,156 @@
<script setup>
// Visitor metadata cell shared by the recent-visits, crawlers, and abuse tables.
// Displays IP/network/host, country flag/city, UA, and language when available.
// Clicking the IP copies the full address to the clipboard.
// ``variantCount`` overrides the UA line to warn when multiple client
// fingerprints share the same IP (e.g. a scanner rotating UAs).
import { computed } from 'vue'
import * as flagSvgs from 'country-flag-icons/string/3x2'
import { copyIp, formatLang } from './analytics/format.js'
const props = defineProps({
ip: { type: String, default: '' },
ipDisplay: { type: String, default: '—' },
ua: { type: String, default: '' },
uaRaw: { type: String, default: '' },
country: { type: String, default: '' },
city: { type: String, default: '' },
lang: { type: String, default: '' },
langDisplay: { type: String, default: '' },
isHost: { type: Boolean, default: false },
variantCount: { type: Number, default: 1 },
})
const hasCountry = computed(() => !!(props.country && props.country !== '—'))
const hasCity = computed(() => !!(props.city && props.city !== '—'))
const hasLocale = computed(() => hasCountry.value || hasCity.value)
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
const showLang = computed(() => langValue.value && langValue.value !== '—')
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script>
<template>
<td class="visitor-cell" :class="{ 'host-cell': isHost }">
<div class="visitor-rows">
<div class="visitor-row">
<div class="locale-line">
<span v-if="flagSvg(country)" class="flag" v-html="flagSvg(country)" :title="countryName(country) || country"></span>
<template v-if="hasCity"><small class="city-name muted">{{ city }}</small></template>
<template v-else-if="!hasLocale"></template>
</div>
<div class="ip-line">
<span class="clickable-ip small muted"
:title="ip"
@click="copyIp(ip, $event)">{{ ipDisplay }}</span>
</div>
</div>
<div class="visitor-row">
<div class="ua-line">
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
</div>
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
</div>
</div>
</td>
</template>
<style scoped>
.visitor-cell {
width: 18em;
max-width: 18em;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: top;
}
.visitor-cell.host-cell {
text-align: right;
}
.visitor-rows {
display: flex;
flex-direction: column;
gap: 0.15rem;
}
.visitor-row {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.5rem;
}
.visitor-row > * {
min-width: 0;
}
.locale-line,
.ip-line,
.ua-line {
flex: 1 1 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.locale-line {
text-align: left;
display: flex;
align-items: center;
gap: 0.3rem;
}
.ip-line {
text-align: right;
}
.ua-line {
text-align: left;
}
.locale-lang {
flex: 0 0 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
text-align: right;
}
.city-name {
display: inline-block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
vertical-align: middle;
}
.flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+201
View File
@@ -0,0 +1,201 @@
<script setup>
/**
* Visitor and page-view smoothed curves for a single shared time range.
*/
import { computed, onMounted, onUnmounted, ref } from 'vue'
import { makeSeries } from './analytics/time.js'
import {
CHART_H,
CHART_W,
MARGIN_B,
MARGIN_L,
VIEW_H,
VIEW_W,
buildChart,
} from './analytics/chart.js'
const DAY_REFRESH_MS = 15000
// Keep the whole svg within page bounds: full width below the natural
// size, centered with equal side margins above it (max() clamps the
// centering margin to 0 at the breakpoint, so the rule is continuous).
const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const props = defineProps({
data: { type: Object, default: null },
range: { type: String, required: true },
})
// Views across all pages combined into one raw bucket map.
const allViews = computed(() => {
const all = {}
for (const buckets of Object.values(props.data?.views || {})) {
for (const [k, c] of Object.entries(buckets)) all[k] = (all[k] || 0) + c
}
return all
})
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
const viewSeries = computed(() => makeSeries(allViews.value, props.range))
function freqLabel(unit) {
return unit === '5min' ? '5 min' : unit === 'hour' ? 'hourly' : 'daily'
}
/** Vertical axis caption: "visits / 5 min" on the day view, else "hourly visits" style. */
function axisLabel(unit, ylabel) {
return unit === '5min' ? `${ylabel} / 5 min` : `${freqLabel(unit)} ${ylabel}`
}
/** Legend label for the overlaid past weeks: "Week M" or "Week MN". */
function pastLabel(series) {
const oldest = series.at(-1).label.slice(5) // strip "Week "
return series.length > 2 ? `Week ${oldest}${series[1].label.slice(5)}` : `Week ${oldest}`
}
const now = ref(Date.now())
let refreshInterval = null
onMounted(() => {
refreshInterval = setInterval(() => { now.value = Date.now() }, DAY_REFRESH_MS)
})
onUnmounted(() => {
if (refreshInterval) clearInterval(refreshInterval)
})
const visitChart = computed(() => buildChart(visitSeries.value, now.value))
const viewChart = computed(() => buildChart(viewSeries.value, now.value))
</script>
<template>
<section v-for="c in [
{ ylabel: 'visits', chart: visitChart, legend: true },
{ ylabel: 'views', chart: viewChart, legend: false },
]" :key="c.ylabel">
<template v-if="c.chart">
<svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`"
:style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }"
role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)">
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
<template v-for="t in c.chart.xticks" :key="'t' + t.x">
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
class="minor vertical" />
</template>
<template v-if="c.chart.bars">
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
<path :d="c.chart.skyline" class="line" />
</template>
<template v-else>
<!-- Oldest overlay weeks first so the current week paints on top. -->
<template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
<path v-if="s.area" :d="s.area" class="area" />
<path :d="s.line" class="line" :class="{ past: s.past }"
:style="{ opacity: s.opacity }" />
</template>
</template>
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
class="axis" />
<text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y"
text-anchor="end" dominant-baseline="middle" class="ylab">{{ g.label }}</text>
<text :x="-(MARGIN_L - 10)" :y="CHART_H / 2" text-anchor="middle"
:transform="`rotate(-90 ${-(MARGIN_L - 10)} ${CHART_H / 2})`"
class="yaxis-label">{{ axisLabel(c.chart.unit, c.ylabel) }}</text>
<text v-for="t in c.chart.xticks" :key="'x' + t.x" :x="t.x" :y="CHART_H + MARGIN_B - 8"
text-anchor="middle" class="xlab">{{ t.label }}</text>
<!-- Week overlay legend, top right inside the plot: current week in
accent, one muted specimen for the whole past range. -->
<g v-if="c.legend && c.chart.series.length > 1">
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="10" y2="10" class="line" />
<text :x="CHART_W - 72" y="10" dominant-baseline="middle"
class="leglab">{{ c.chart.series[0].label }}</text>
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="25" y2="25"
class="line past" style="opacity: 0.6" />
<text :x="CHART_W - 72" y="25" dominant-baseline="middle"
class="leglab">{{ pastLabel(c.chart.series) }}</text>
</g>
</svg>
</template>
</section>
</template>
<style scoped>
/* Each chart is a self-contained SVG: the viewBox includes the axis label
margins, so nothing is positioned with HTML overlays. Never upscale past
the natural size (1 viewBox unit = 1 px, max-width set inline) — that
would blow up the constant-size text; smaller panels still scale the
chart down to fit. The margin-left (set inline) centers the chart above
its natural width; the svg always stays within page bounds.
overflow: visible lets wider fonts extend past the viewBox instead of
clipping. */
.chart {
display: block;
width: 100%;
height: auto;
overflow: visible;
}
.chart .ylab,
.chart .xlab,
.chart .yaxis-label,
.chart .leglab {
font-family: system-ui, sans-serif; /* theme fonts can be overly styled */
font-size: 11px;
fill: var(--muted);
}
.chart .ylab {
font-variant-numeric: tabular-nums;
}
.chart .minor {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
opacity: 0.35;
}
.chart .minor.vertical {
opacity: 0.25;
}
.chart .major {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
stroke-dasharray: 3 4;
opacity: 0.8;
}
.chart .axis {
stroke: var(--line);
stroke-width: 1;
vector-effect: non-scaling-stroke;
}
.chart .area {
fill: var(--accent);
opacity: 0.15;
}
.chart .bar {
fill: var(--accent);
opacity: 0.15;
}
.chart .line {
fill: none;
stroke: var(--accent);
stroke-width: 2;
vector-effect: non-scaling-stroke;
stroke-linejoin: round;
stroke-linecap: round;
}
/* Past overlay weeks contrast with the current week's accent color. */
.chart .line.past {
stroke: var(--muted);
}
.empty { color: var(--muted); }
</style>
+30
View File
@@ -0,0 +1,30 @@
// Analytics page entry: mounts AnalyticsView inside the normal page layout.
// In production the backend inlines this module into the /_a page (and
// pagerite.js re-creates the script element after fetch-navigations there);
// in dev pagerite.js imports it from the Vite dev server on demand. Either
// way it auto-mounts on #analytics-app when it evaluates, and unmounts when
// pagerite.js announces a swap away from /_a.
import { createApp } from 'vue'
import AnalyticsView from './AnalyticsView.vue'
let app = null
export function mount(container) {
if (app) return
app = createApp(AnalyticsView)
app.mount(container)
}
export function unmount() {
app?.unmount()
app = null
}
// pagerite.js calls this before swapping away from /_a; each evaluation
// (the inlined production module evaluates fresh on every visit) replaces
// the handle.
window.__pageriteAnalyticsUnmount = unmount
// Auto-mount when the page holding #analytics-app is present.
const container = document.getElementById('analytics-app')
if (container) mount(container)
+388
View File
@@ -0,0 +1,388 @@
/**
* Chart geometry, smoothing, and SVG path generation for analytics charts.
*
* Fixed 720x180 plot area inside a larger viewBox that also holds the axis
* labels, so each chart SVG is self-contained; values are per-unit rates
* (hour on the week view, day on month+).
*/
import { DAY, HOUR, MIN5, WEEK, mondayUTC } from './time.js'
import { formatCount } from './format.js'
export const CHART_W = 1000
export const CHART_H = 150
export const PAD_TOP = 14 // room above the highest point
export const MARGIN_L = 40 // y tick labels + vertical axis label
export const MARGIN_B = 24 // x tick labels
export const VIEW_W = MARGIN_L + CHART_W + 8
export const VIEW_H = CHART_H + MARGIN_B
/**
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
* most 5 intervals, so labeled ticks are always round and evenly divided.
* A minimum range of 10 keeps tiny near-zero values (e.g. a single visit)
* from being enlarged to a fractional scale; minor lines subdivide each
* major step in five when that yields integers.
*/
export function yScale(maxValue) {
let step = 1
outer: for (let exp = -3; exp < 8; exp++) {
for (const base of [1, 2, 5]) {
step = base * 10 ** exp
if (Math.ceil(maxValue / step) <= 5) break outer
}
}
let max = Math.ceil(maxValue / step) * step
if (max < 10) {
max = 10
step = 2
}
const minor = step >= 5 && step % 5 === 0 ? step / 5 : null
return { max, step, minor }
}
/**
* Edge-aware Gaussian smoothing with a fixed bandwidth. A change-point
* detector first finds traffic-level shifts (two-unit totals compared on
* both sides of each bucket; strong ratio + significance marks a candidate,
* and each run of candidates keeps only its best-scoring bucket as an
* edge). Each edge-delimited segment is then smoothed independently: every
* bucket spreads its count with a fixed Gaussian sigma chosen so N events
* in a single bucket peak at N events per unit. Mass past a detected change
* point is dropped (kernel renormalized); mass past a true series edge is
* mirrored back, so the curve doesn't fall where data simply ends. Either
* way total visitor count is preserved exactly. The unit is
* one hour on the week view and one day on the month+ views, so the
* smoothing time scale follows the range. The raw series is drawn faintly
* behind the curve for reference. Operates on raw counts.
*/
export function smooth(counts, binMinutes, unitMinutes, {
detectorWindowMinutes = 2 * unitMinutes,
// Count thresholds are defined per hour and scale with the unit, so
// "low traffic" means the same thing on hourly and daily views
// (5-20 events/hour = 120-480/day on the month+ ranges).
highTrafficEvents = 10 * unitMinutes / 60,
minRatio = 2.5,
minSignificance = 4,
} = {}) {
const n = counts.length
if (!n) return counts
const detectorWindowBins = Math.max(1, Math.round(detectorWindowMinutes / binMinutes))
const cumsum = new Float64Array(n + 1)
for (let i = 0; i < n; i++) cumsum[i + 1] = cumsum[i] + counts[i]
// Detect abrupt regime changes from aggregated traffic on both sides.
// Individual bins are deliberately ignored because even high traffic
// produces many 0-1 count bins at five-minute resolution.
const score = new Float64Array(n)
const candidate = new Uint8Array(n)
for (let i = detectorWindowBins; i < n - detectorWindowBins; i++) {
const left = cumsum[i] - cumsum[i - detectorWindowBins]
const right = cumsum[i + detectorWindowBins] - cumsum[i]
const high = Math.max(left, right)
const low = Math.min(left, right)
if (high < highTrafficEvents) continue
const ratio = (high + 1) / (low + 1)
const significance = (high - low) / Math.sqrt(high + low + 1)
if (ratio >= minRatio && significance >= minSignificance) {
candidate[i] = 1
score[i] = significance * Math.log(ratio)
}
}
// Collapse each continuous detector region to its strongest boundary.
const edges = []
for (let i = 0; i < n;) {
if (!candidate[i]) { i++; continue }
let j = i + 1
while (j < n && candidate[j]) j++
let best = i
for (let k = i + 1; k < j; k++) {
if (score[k] > score[best]) best = k
}
edges.push(best)
i = j
}
// Fixed sigma: N events in one bucket peak at N events per unit.
// sigma_bins * sqrt(2*pi) = rate = unitMinutes / binMinutes.
const sigmaBins = unitMinutes / (binMinutes * Math.sqrt(2 * Math.PI))
const radius = Math.ceil(4 * sigmaBins)
// Process each discontinuity-delimited regime independently so the
// Gaussian cannot see through a detected boundary. Each input bin spreads
// its count with the fixed sigma. Mass that would fall past a detected
// change point is dropped and the kernel renormalized; mass that would
// fall past a true series edge (first/last bin) is mirrored back into the
// segment, as if the data continued as its own reflection, so constant or
// rising data doesn't produce a spurious falling edge. Total visitor count
// is preserved apart from floating-point error.
const bounds = [0, ...edges, n]
const smoothed = new Float64Array(n)
for (let b = 0; b < bounds.length - 1; b++) {
const lo = bounds[b]
const length = bounds[b + 1] - lo
const mirrorLeft = lo === 0
const mirrorRight = lo + length === n
const segment = counts.slice(lo, lo + length)
for (let j = 0; j < length; j++) {
const count = segment[j]
if (!count) continue
// Collect (target bin, weight) pairs over the full kernel, folding
// mirrored mass at series edges and dropping mass past change points.
const spread = new Map()
let weightSum = 0
for (let i = j - radius; i <= j + radius; i++) {
let k = i
// Fold repeatedly for segments shorter than the kernel radius.
while (k < 0 || k >= length) {
if (k < 0 && mirrorLeft) k = -k - 1
else if (k >= length && mirrorRight) k = 2 * length - 1 - k
else { k = null; break }
}
if (k === null) continue
const w = Math.exp(-0.5 * ((i - j) / sigmaBins) ** 2)
spread.set(k, (spread.get(k) || 0) + w)
weightSum += w
}
for (const [k, w] of spread) {
smoothed[lo + k] += count * w / weightSum
}
}
}
return [...smoothed]
}
/**
* Catmull-Rom spline through the (smoothed) points, control points clamped
* to the plot area so the curve can never dip below zero or above the max.
*/
export function spline(pts) {
if (pts.length < 3) {
return `M${pts.map((p) => `${p.x},${p.y}`).join('L')}`
}
const clampY = (y) => Math.min(CHART_H, Math.max(PAD_TOP, y))
let d = `M${pts[0].x},${pts[0].y}`
for (let i = 0; i < pts.length - 1; i++) {
const p0 = pts[i - 1] || pts[i]
const p1 = pts[i]
const p2 = pts[i + 1]
const p3 = pts[i + 2] || p2
const c1y = clampY(p1.y + (p2.y - p0.y) / 6)
const c2y = clampY(p2.y - (p3.y - p1.y) / 6)
d += `C${p1.x + (p2.x - p0.x) / 6},${c1y} `
+ `${p2.x - (p3.x - p1.x) / 6},${c2y} ${p2.x},${p2.y}`
}
return d
}
/** Build a full chart model from a series descriptor produced by time.js. */
export function buildChart(input, now = Date.now()) {
if (!input || !input.series.length) return null
if (input.unit === '5min') return buildDayChart(input, now)
const { series, t0, t1, rate, binMinutes, unitMinutes, unit } = input
// Values are per-unit rates (hour on the week view, day on month+); the
// y max is derived from the *smoothed* curves so random single-bucket
// spikes don't blow up the scale. Smoothing works on raw counts (its edge
// detector thresholds are count-based), the result is scaled back to rates.
const smoothed = series.map((s) =>
smooth(s.points.map((p) => p.count), binMinutes, unitMinutes).map((v) => v * rate))
// Scale from the current/primary series only; older overlay weeks are drawn
// with the same scale and allowed to overflow if they are busier.
const highest = Math.max(0, ...smoothed[0])
const { max, step, minor } = yScale(highest)
const x = (t) => ((t - t0) / (t1 - t0)) * CHART_W
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const drawn = series.map((s, si) => {
const pts = s.points.map((p, i) => ({ x: x(p.t), y: y(smoothed[si][i]) }))
const line = spline(pts)
const first = pts[0]
const last = pts.at(-1)
return {
...s,
line,
area: s.area ? `${line}L${last.x},${CHART_H}L${first.x},${CHART_H}Z` : null,
}
})
// Major (labeled) and minor (hairline) y grid ticks.
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), label: fmtY(v) })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
// X ticks. Week view: weekday labels centered at midday UTC, no vertical
// lines (day boundaries would be misleading in the viewer's timezone).
// Month view: likewise lineless, day numbers at noon UTC with the month
// name substituted for the 1st (marking the month change). Longer
// ranges: boundary lines at Mondays / months / years.
const isWeek = t1 - t0 === WEEK
const isMonth = !isWeek && t1 - t0 <= 31 * DAY
let xticks
if (isWeek) {
xticks = Array.from({ length: 7 }, (_, d) => {
const t = t0 + d * DAY + 12 * HOUR
return {
x: x(t),
label: new Date(t).toLocaleDateString(undefined, {
weekday: 'short', timeZone: 'UTC',
}),
line: false,
}
})
} else if (isMonth) {
// t0 is day-aligned; label every day whose noon falls inside the range.
xticks = []
for (let day = t0; day + 12 * HOUR < t1; day += DAY) {
const date = new Date(day)
const t = day + 12 * HOUR
xticks.push({
x: x(t),
label: date.getUTCDate() === 1
? date.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' })
: String(date.getUTCDate()),
line: false,
})
}
} else {
xticks = xticksFor(t0, t1).map((t) => ({
x: x(t), label: fmtTick(t, t1 - t0), line: true,
}))
}
return { max, majors, minors, series: drawn, xticks, unit }
}
/**
* Day view: 5-minute bars for the last 24 hours. Bars are drawn at raw
* counts; the skyline uses a projected full-bucket value for the still-open
* final bucket. The y scale is derived from the projected skyline maximum.
*/
export function buildDayChart(input, now = Date.now()) {
const { series, t0, t1 } = input
const points = series[0]?.points || []
const n = points.length
if (!n) return null
const bucketMs = (t1 - t0) / n
const bucketWidth = CHART_W / n
const gap = 0.2
const barWidth = Math.max(0.2, bucketWidth - gap)
const x = (i) => i * bucketWidth + gap / 2
const prevRaw = n > 1 ? points[n - 2].count : 0
const projected = points.map((p, i) => {
if (i !== n - 1) return p.count
const bucketStart = t0 + i * bucketMs
const elapsed = Math.max(1, Math.min(bucketMs, now - bucketStart))
// Blend the observed partial bucket with the previous full bucket:
// the longer the current bucket has run, the less we borrow from it.
const share = elapsed / bucketMs
return p.count + prevRaw * (1 - share)
})
const highest = Math.max(0, ...projected)
const { max, step, minor } = yScale(highest)
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const bars = points.map((p, i) => {
const bx = x(i)
const by = y(p.count)
return {
x: bx,
y: by,
width: barWidth,
height: CHART_H - by,
raw: p.count,
projected: projected[i],
}
})
let skyline = ''
for (let i = 0; i < bars.length; i++) {
const b = bars[i]
const top = y(b.projected)
if (i === 0) {
skyline += `M${b.x},${top} H${b.x + b.width}`
} else {
skyline += ` V${top} H${b.x + b.width}`
}
}
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), label: fmtY(v) })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
const xticks = []
const tickStep = 3 * HOUR
const firstTick = Math.ceil(t0 / tickStep) * tickStep
for (let t = firstTick; t < t1; t += tickStep) {
if (t < t0) continue
const d = new Date(t)
xticks.push({
x: ((t - t0) / (t1 - t0)) * CHART_W,
label: `${String(d.getUTCHours()).padStart(2, '0')}:00`,
line: false,
})
}
return { bars, skyline: skyline.trim(), max, majors, minors, xticks, unit: '5min', series: [] }
}
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month
* boundaries up to a few years, then years. */
export function xticksFor(t0, t1) {
const span = t1 - t0
const ticks = []
if (span <= 100 * DAY) {
for (let t = mondayUTC(t0); t <= t1; t += WEEK) {
if (t >= t0) ticks.push(t)
}
return ticks
}
if (span <= 4 * 365 * DAY) {
const d = new Date(t0)
let t = Date.UTC(d.getUTCFullYear(), d.getUTCMonth() + 1, 1)
for (; t <= t1;) {
ticks.push(t)
const m = new Date(t)
t = Date.UTC(m.getUTCFullYear(), m.getUTCMonth() + 1, 1)
}
return ticks
}
const d = new Date(t0)
for (let yr = d.getUTCFullYear() + 1; Date.UTC(yr, 0, 1) <= t1; yr++) {
ticks.push(Date.UTC(yr, 0, 1))
}
return ticks
}
export function fmtTick(t, span) {
const d = new Date(t)
if (span <= 100 * DAY) {
return d.toLocaleDateString(undefined, { month: 'short', day: 'numeric', timeZone: 'UTC' })
}
if (span <= 4 * 365 * DAY) {
return d.getUTCMonth() === 0
? d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
: d.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' })
}
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
}
/** Y labels use the same compact formatter as text labels. */
export function fmtY(v) {
return formatCount(v)
}
+562
View File
@@ -0,0 +1,562 @@
/**
* Formatters and aggregators for summary sections: totals and the recent
* visit trail.
*/
/**
* IPv4 unchanged, IPv6 returns the /64 network prefix in compact form.
* Falls back to the original value when parsing fails.
*/
export const hostIP = (ip) => {
try {
if (!ip || !ip.includes(':')) return ip
const strip = (s) => s.replace(/^\[|\]$/g, '')
const norm = strip(new URL(`http://[${ip}]/`).hostname)
const [l, r] = norm.split('::').map((s) => (s ? s.split(':') : []))
const full = r
? [...l, ...Array(8 - l.length - r.length).fill('0'), ...r]
: l
return strip(
new URL(`http://[${full.slice(0, 4).join(':')}::]/`).hostname,
).replace(/::$/, '')
} catch (e) {
console.error('hostIP processing failed for:', ip, e)
return ip
}
}
function showCopiedFeedback(el) {
if (!el || typeof document === 'undefined') return
const popup = document.createElement('span')
popup.textContent = 'Copied!'
popup.className = 'copy-popup'
popup.style.cssText =
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
'pointer-events:none;z-index:10;'
el.classList.add('has-copy-popup')
el.appendChild(popup)
setTimeout(() => {
popup.remove()
el.classList.remove('has-copy-popup')
}, 1200)
}
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
export async function copyIp(ip, event) {
if (!ip) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(ip)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
export async function copyList(text, event) {
if (!text) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(text)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Total page views across every page and every bucket. */
export function calcTotalViews(views) {
let n = 0
for (const buckets of Object.values(views || {})) {
for (const c of Object.values(buckets)) n += c
}
return n
}
// Very short reads are navigation/skims, not real reading time.
export const MIN_READ_SECONDS = 10
/** path -> accumulated read seconds for a visit, derived from its trail. */
export function readMapOf(v) {
const map = {}
for (const item of Object.values(v.trail || {})) {
if (item.read) map[item.to] = (map[item.to] || 0) + item.read
}
return map
}
/** Average minutes per visit and average of per-article median read minutes. */
export function calcReadStats(visits) {
const perArticle = {}
let totalVisitSeconds = 0
let visitCount = 0
for (const v of visits || []) {
const read = readMapOf(v)
const secs = Object.values(read).filter((s) => s >= MIN_READ_SECONDS)
if (!secs.length) continue
visitCount++
totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
for (const [path, s] of Object.entries(read)) {
if (s >= MIN_READ_SECONDS) {
; (perArticle[path] || (perArticle[path] = [])).push(s)
}
}
}
const avgMinPerVisit = visitCount
? Math.max(1, Math.round(totalVisitSeconds / visitCount / 60))
: 0
let articleMedianSum = 0
const articleCount = Object.keys(perArticle).length
for (const arr of Object.values(perArticle)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median = arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
articleMedianSum += Math.max(MIN_READ_SECONDS, median)
}
const avgArticleMedianMin = articleCount
? Math.max(1, Math.round(articleMedianSum / articleCount / 60))
: 0
return { avgMinPerVisit, avgArticleMedianMin }
}
/** Build a path -> page title lookup from the site tree. */
function buildTitleMap(pageTree) {
const titles = new Map()
const walk = (items) => {
for (const item of items || []) {
titles.set(`/${item.path}`, item.title)
walk(item.children)
}
}
walk(pageTree)
return titles
}
/** Last path segment for display; front page becomes a house icon. */
function slugOf(path) {
return path === '/' ? '🏠︎' : path.split('/').pop()
}
/** Host name of an external https origin, with scheme and www. stripped. */
function externalSlug(origin) {
try {
return new URL(origin).host.replace(/^www\./, '')
} catch {
return origin.replace(/^https?:\/\//, '').replace(/^www\./, '')
}
}
/** Format one trail step: an internal page or an external https origin. */
function stepOf(path, titles) {
if (path?.startsWith('/')) {
return { path, slug: slugOf(path), title: titles.get(path) || '', external: false, home: path === '/' }
}
if (path?.startsWith('https://')) {
return {
path,
slug: externalSlug(path),
title: 'External site',
external: true,
}
}
return null
}
/**
* Human-readable relative timestamp. Adapted from cista-storage: uses
* ``Intl.RelativeTimeFormat`` for short intervals and a compact date for
* anything older than a week.
*/
export function formatWhen(ts, now = Date.now()) {
const date = new Date(ts)
const diff = date.getTime() - now
const adiff = Math.abs(diff)
const formatter = new Intl.RelativeTimeFormat('en', { numeric: 'auto' })
if (adiff <= 5000) return 'now'
if (adiff <= 60000) {
return formatter
.format(Math.round(diff / 1000), 'second')
.replace(' ago', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 3600000) {
return formatter
.format(Math.round(diff / 60000), 'minute')
.replace('utes', '')
.replace('ute', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 86400000) {
return formatter
.format(Math.round(diff / 3600000), 'hour')
.replace('hours', 'h')
.replace('hour', 'h')
.replaceAll(' ', '\u202F')
}
if (adiff <= 604800000) {
return formatter
.format(Math.round(diff / 86400000), 'day')
.replaceAll(' ', '\u202F')
}
let d = date
.toLocaleDateString('en-ie', {
weekday: 'short',
year: 'numeric',
month: 'short',
day: 'numeric',
})
.replace('Sept', 'Sep')
if (d.length === 14) d = d.replace(' ', ' \u2007')
d = d.replaceAll(' ', '\u202F').replace('\u202F', '\u00A0')
d = d.slice(0, -4) + d.slice(-2)
return d
}
/** Full UTC timestamp for tooltips, e.g. "2026-08-21 00:20:48 UTC". */
export function formatWhenTooltip(ts) {
return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC')
}
/** Full local timestamp for tooltips, e.g. "21 Aug 2026, 17:38:48". */
export function formatWhenLocal(ts) {
return new Date(ts).toLocaleString('en-ie', {
year: 'numeric',
month: 'short',
day: 'numeric',
hour: '2-digit',
minute: '2-digit',
second: '2-digit',
})
}
/** Preserve locale case with the region/country subtag upper-cased. */
export function formatLang(value) {
if (!value || value === '—') return value
const parts = value.split('-')
if (parts.length > 1) {
parts[parts.length - 1] = parts[parts.length - 1].toUpperCase()
}
return parts.join('-')
}
/** ISO 8601 UTC timestamp without subseconds, e.g. "2026-08-21T00:20:48Z". */
export function formatWhenIso(ts) {
return `${new Date(ts).toISOString().split('.')[0]}Z`
}
/**
* Compact read time for tooltips: "50s" under a minute, "1m23s" otherwise.
*/
export function formatReadTime(seconds) {
if (seconds < 60) return `${seconds}s`
return `${Math.floor(seconds / 60)}m${seconds % 60}s`
}
/**
* Compact visitor counts: plain below 1k, then 1.2k / 10k / 1.2M.
* Truncated, not rounded.
*/
export function formatCount(n) {
if (n < 1000) return String(n)
if (n < 10000) return `${Math.trunc(n / 1000)}.${Math.trunc((n % 1000) / 100)}k`
if (n < 1_000_000) return `${Math.trunc(n / 1000)}k`
return `${Math.trunc(n / 1_000_000)}.${Math.trunc((n % 1_000_000) / 100_000)}M`
}
/**
* Format recent visits for display, newest first. Each step is a linked slug
* pointing to its article; external referers/origins are shown as their
* domain name with the full origin as the link href. The link title shows the
* article heading when known, or "External site" for origins.
*/
export function formatRecentVisits(visits, pageTree, limit = 50) {
const titles = buildTitleMap(pageTree)
return [...visits]
.reverse()
.map((v) => ({
when: new Date(v.start).toLocaleString(),
steps: [v.referer, ...Object.values(v.trail || {}).map((t) => t.to)]
.map((p) => stepOf(p, titles))
.filter(Boolean),
}))
.filter((v) => v.steps.length)
.slice(0, limit)
}
/**
* Count distinct values of a visit field, sorted most-common first.
* Returns an array of [value, count] pairs.
*/
export function countByField(visits, field) {
const counts = {}
for (const v of visits || []) {
const value = v[field]
if (!value) continue
counts[value] = (counts[value] || 0) + 1
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/**
* Count UTM parameter occurrences across visits. Each distinct
* ``parameter: value`` pair is counted separately. Returns [pair, count].
*/
export function countUtmTags(visits) {
const counts = {}
for (const v of visits || []) {
for (const [key, value] of Object.entries(v.utm || {})) {
const label = `${key}: ${value}`
counts[label] = (counts[label] || 0) + 1
}
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/** Format a list of [value, count] pairs for inline display. */
export function formatCounts(entries) {
return entries.map(([value, count]) => `${value} (${count})`).join(', ')
}
/**
* Count distinct User-Agent strings among crawler hits, most common first.
* Returns an array of [ua, count] pairs. ``clients`` maps client hashes to
* client records.
*/
export function countCrawlerUas(crawlers, clients) {
const counts = {}
for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {}
const value = client.ua_pretty || client.ua || '(no UA)'
counts[value] = (counts[value] || 0) + 1
}
return Object.entries(counts).sort((a, b) => b[1] - a[1])
}
/**
* Reduce a reverse-DNS hostname to its right-most components that fit
* within ``limit`` characters. This keeps the meaningful main domain
* while avoiding absurdly long subdomains like ``xxx.yyy.zzz...provider.net``.
*/
export function mainDomain(host, limit = 24) {
if (!host) return host
const labels = host.split('.').filter(Boolean)
if (!labels.length) return host
const parts = [labels.pop()]
while (labels.length) {
const next = labels[labels.length - 1]
const candidate = `${next}.${parts.join('.')}`
if (candidate.length > limit) break
parts.unshift(labels.pop())
}
return parts.join('.')
}
/**
* Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by most recent hit
* first, with total hits as a tie-breaker.
* ``clients`` maps client hashes to client records.
*/
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
const groups = new Map()
for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {}
const g = groups.get(c.client) || {
clientHash: c.client,
client,
lastStart: 0,
pages: new Map(),
}
const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start
if (c.entry?.startsWith('/')) {
const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
existing.count += 1
if (c.status != null) existing.status = c.status
g.pages.set(c.entry, existing)
}
groups.set(c.client, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pages.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart || totalHits(b) - totalHits(a))
.slice(0, 10)
.map((g) => {
const client = g.client || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
pages: [...g.pages.entries()]
.sort((a, b) => b[1].count - a[1].count)
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical paths are collapsed into one entry with their hit count.
* Flagged paths (the ones that triggered abuse classification) are lifted to
* the top, followed by other 404s, then document GETs from the abuser. Within
* each category paths are sorted by count descending, then earliest first.
* Rows are sorted by most recent hit first. Visitor metadata comes from the
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
* how many distinct client variations the IP produced. Paths are shown
* verbatim (query string included), not resolved against the page tree.
* ``clients`` maps client hashes to client records.
*/
export function formatAbuseRows(abuse, clients, now = Date.now()) {
const groups = new Map()
for (const a of abuse || []) {
const client = (clients || {})[a.client] || {}
const ip = client.ip || ''
const g = groups.get(ip) || {
ip,
pathCounts: new Map(),
clientHashes: new Set(),
lastStart: 0,
lastClient: a.client,
}
const start = new Date(a.start).getTime()
if (start > g.lastStart) {
g.lastStart = start
g.lastClient = a.client
}
const path = a.path || ''
const existing = g.pathCounts.get(path) || {
path,
count: 0,
firstStart: start,
flag: a.flag || false,
is_404: a.is_404 || false,
}
existing.count += 1
if (start < existing.firstStart) existing.firstStart = start
if (a.flag) existing.flag = true
if (!a.is_404) existing.is_404 = false
g.pathCounts.set(path, existing)
g.clientHashes.add(a.client)
groups.set(ip, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pathCounts.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart)
.slice(0, 10)
.map((g) => {
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2)
const paths = [...g.pathCounts.values()].sort(
(a, b) =>
pathCategory(a) - pathCategory(b) ||
b.count - a.count ||
a.firstStart - b.firstStart,
)
const client = (clients || {})[g.lastClient] || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
paths: paths.map((p) => ({
path: p.path,
count: p.count,
flag: p.flag,
is_404: p.is_404,
})),
allPaths: paths
.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path))
.join('\n'),
clientCount: g.clientHashes.size,
ip: client.ip || g.ip,
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Format raw visit records as rows for a technical table. Returns objects
* with display strings; missing values become "—". ``trail`` starts with the
* external referer (when present), then the entry page and any further internal
* pages or external exit origins. Only the 20 most recent visits are shown.
* ``clients`` maps client hashes to client records.
*/
export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const client = (clients || {})[v.client] || {}
const trail = Object.values(v.trail || {})
.map((item) => {
const step = stepOf(item.to, titles)
if (step) {
if (item.read) step.readSeconds = item.read
if (item.status) step.status = item.status
}
return step
})
.filter(Boolean)
const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean)
const utm = utmValues.length ? utmValues.join(' · ') : ''
const utmTitle = Object.entries(v.utm || {})
.map(([k, value]) => `${k}=${value}`)
.join(', ')
const dash = (s) => (s || '—')
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(v.start, now),
lastSeenIso: formatWhenIso(v.start),
lastSeenLocal: formatWhenLocal(v.start),
langDisplay: formatLang(client.lang),
trail,
refererStep: stepOf(v.referer, titles),
referer: dash(v.referer),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
lang: dash(client.lang),
country: dash(client.country),
city: dash(client.city),
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
utm: utm || '—',
utmTitle,
}
})
}
+267
View File
@@ -0,0 +1,267 @@
/**
* Time ranges, week alignment and re-bucketing for analytics charts.
*
* Raw data comes as sparse 5-minute buckets; the range picks the x window
* and a coarser bucket size to keep point counts sane. The week range is
* aligned to Monday 00:00 UTC and overlays previous weeks' curves (fading
* with age), so weekly patterns compare directly.
*/
export const MIN5 = 5 * 60e3
export const HOUR = 3600e3
export const DAY = 86400e3
export const WEEK = 7 * DAY
export const RANGES = {
day: { label: 'day', span: DAY, bucket: MIN5 },
week: { label: 'week' },
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
year: { label: 'year', span: 365 * DAY, bucket: DAY },
all: { label: 'all', span: null, bucket: DAY, minSpan: 30 * DAY },
}
/** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */
export function mondayUTC(t) {
const d = Math.floor(t / DAY)
return (d - ((d + 3) % 7)) * DAY
}
/** ISO 8601 week number of the week containing t (via its Thursday). */
export function isoWeek(t) {
const d = new Date(t)
d.setUTCHours(0, 0, 0, 0)
d.setUTCDate(d.getUTCDate() + 4 - (d.getUTCDay() || 7))
const yearStart = Date.UTC(d.getUTCFullYear(), 0, 1)
return Math.ceil(((d - yearStart) / DAY + 1) / 7)
}
/** Parse sparse timestamp buckets into a { epochMs: count } map. */
export function rawTimes(buckets) {
const raw = {}
// Key by parsed timestamp: Python writes "+00:00", JS ISO uses "Z".
for (const [k, c] of Object.entries(buckets || {})) raw[Date.parse(k)] = c
return raw
}
/** Sum counts from raw 5-minute buckets between t0 (inclusive) and t1 (exclusive). */
export function sumRange(raw, t0, t1) {
let n = 0
for (let s = t0; s < t1; s += MIN5) n += raw[s] || 0
return n
}
/**
* One series per overlaid week: [this week, 1 week ago, ...], at native
* 5-minute resolution, up to 8 weeks back (and only weeks that overlap the
* recorded data at all). Each older week's timestamps are shifted forward
* onto the current week's axis so all curves overlay inside the plot.
* The current week is truncated at the current bucket
* — no fake zeroes drawn for the future. Counts are rates per hour
* (bucket count * 12): a lone visit in a 5-minute bucket reads as "12/h".
* The coarser ranges use per-day rates instead (unitMinutes = 24*60).
*/
export function weeklySeries(buckets) {
const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number)
const now = Date.now()
const thisMonday = mondayUTC(now)
if (!times.length) {
const points = []
const end = Math.min(thisMonday + WEEK, Math.floor(now / MIN5) * MIN5 + MIN5)
for (let t = thisMonday; t < end; t += MIN5) {
points.push({ t, count: 0 })
}
return {
series: [{ points, label: `Week ${isoWeek(thisMonday)}`, opacity: 1, area: true }],
t0: thisMonday,
t1: thisMonday + WEEK,
rate: HOUR / MIN5,
binMinutes: 5,
unitMinutes: 60,
unit: 'hour',
}
}
const oldest = Math.min(...times)
// Weeks back as far as the data reaches: difference in Monday indices.
const available = (thisMonday - mondayUTC(oldest)) / WEEK + 1
const count = Math.min(available, 8)
const out = []
for (let back = 0; back < count; back++) {
const start = thisMonday - back * WEEK
const end = back === 0
? Math.min(start + WEEK, Math.floor(now / MIN5) * MIN5 + MIN5)
: start + WEEK
const points = []
for (let t = start; t < end; t += MIN5) {
points.push({ t: t + back * WEEK, count: raw[t] || 0 })
}
out.push({
points,
label: `Week ${isoWeek(start)}`,
opacity: Math.max(0.15, 1 - back * 0.25),
past: back > 0,
area: back === 0,
})
}
return {
series: out,
t0: thisMonday,
t1: thisMonday + WEEK,
rate: HOUR / MIN5,
binMinutes: 5,
unitMinutes: 60,
unit: 'hour',
}
}
/**
* Rolling window for the non-week ranges (x max = now), counts converted
* to per-day rates (the unit the month+ charts are read in).
* Ranges without a fixed span use the full data reach, but never less than
* their configured minSpan so the chart keeps a readable minimum x scale.
* t0 is aligned to the UTC day so the x labels cover the whole range;
* t1 is now, so the scale never extends into the future. The bucket size
* follows the resulting window (6h up to 31 days, daily beyond), so ranges
* covering the same window — "all" at its 30-day minimum vs "month" —
* render the identical curve.
*/
export function rollingSeries(buckets, rangeKey) {
const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number)
const { span, bucket, minSpan = 0 } = RANGES[rangeKey]
const t1 = Date.now()
const t0 = Math.floor((span != null
? t1 - span
: Math.min(times.length ? Math.min(...times) : Infinity, t1 - minSpan)) / DAY) * DAY
if (!times.length) {
const bucketMs = t1 - t0 <= 31 * DAY ? Math.min(bucket, 6 * HOUR) : bucket
const points = []
for (let t = t0; t < t1; t += bucketMs) {
points.push({ t, count: 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: true }],
t0,
t1,
rate: DAY / bucketMs,
binMinutes: bucketMs / 60e3,
unitMinutes: 24 * 60,
unit: 'day',
}
}
// The bucket follows the actual window length, not the range key: when
// "all" is capped to its 30-day minimum it covers the very window "month"
// does, and daily bins would draw a different curve over the same data
// (coarser edge detection, points a day apart plotted at bin starts, the
// last point stuck at today's midnight instead of reaching now).
const bucketMs = t1 - t0 <= 31 * DAY ? Math.min(bucket, 6 * HOUR) : bucket
const points = []
for (let t = t0; t < t1; t += bucketMs) {
points.push({ t, count: sumRange(raw, t, t + bucketMs) })
}
return {
series: [{ points, label: '', opacity: 1, area: true }],
t0,
t1,
rate: DAY / bucketMs,
binMinutes: bucketMs / 60e3,
unitMinutes: 24 * 60,
unit: 'day',
}
}
/**
* Day view: raw 5-minute bucket counts for the current 24-hour window.
* No smoothing or rate conversion is applied; counts are used as-is.
*/
export function daySeries(buckets) {
const raw = rawTimes(buckets)
const now = Date.now()
const { span, bucket } = RANGES.day
const t1 = Math.floor(now / bucket) * bucket + bucket
const t0 = t1 - span
const points = []
for (let t = t0; t < t1; t += bucket) {
points.push({ t, count: raw[t] || 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: false }],
t0,
t1,
rate: 1,
binMinutes: bucket / 60e3,
unitMinutes: bucket / 60e3,
unit: '5min',
}
}
/** Dispatch to daily, weekly or rolling series based on the selected range. */
export function makeSeries(buckets, rangeKey) {
if (rangeKey === 'day') return daySeries(buckets)
if (rangeKey === 'week') return weeklySeries(buckets)
return rollingSeries(buckets, rangeKey)
}
/**
* Absolute UTC time window for a given range key. Used to filter visits,
* transitions and views for the non-chart stats on the analytics page.
* Every bounded range is a rolling span ending at now; the charts instead
* align week to Monday 00:00 UTC (overlaying previous weeks) and month+
* to UTC day boundaries, so their x windows differ from the stats range
* on purpose.
* Returns { t0, t1 } where null means unbounded.
*/
export function rangeWindow(rangeKey) {
const now = Date.now()
if (rangeKey === 'all') {
return { t0: null, t1: null }
}
const span = rangeKey === 'week' ? WEEK : RANGES[rangeKey].span
return { t0: now - span, t1: now }
}
/**
* Sum the bucketed transition matrix (from -> to -> bucket ISO -> count)
* into a plain from -> to -> count matrix for the window [t0, t1).
*/
export function filterTransitionsByRange(transitions, t0, t1) {
const out = {}
for (const [fr, tos] of Object.entries(transitions || {})) {
for (const [to, buckets] of Object.entries(tos)) {
let n = 0
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) n += c
}
if (n) {
out[fr] = out[fr] || {}
out[fr][to] = n
}
}
}
return out
}
/** Keep only the 5-minute view buckets that fall inside [t0, t1). */
export function filterViewsByRange(views, t0, t1) {
const filtered = {}
for (const [path, buckets] of Object.entries(views || {})) {
const out = {}
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out[k] = c
}
if (Object.keys(out).length) filtered[path] = out
}
return filtered
}
/** Keep only records whose start time falls inside [t0, t1). */
export function filterRecordsByRange(records, t0, t1) {
const out = []
for (const r of records || []) {
const t = Date.parse(r.start)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out.push(r)
}
return out
}
+943
View File
@@ -0,0 +1,943 @@
/**
* Radial transition map and helpers.
*
* Site map following the menu structure: top-level items in a row at the
* top (below the external source row), each item's subtree fanning out
* below it in menu order along a slightly circular downward arc. Index
* pages with no views are omitted, their children moving up in their
* place. All pages of the site are shown (from /_api/pages), plus any
* extra paths seen in transitions (deleted pages); these form their own
* top-level groups. Internal path -> path transitions join opposite
* directions into straight connections (middle width = total
* count; connectors flare into the node pills at both ends and wrap
* around their backs, surrounding them; the pills are drawn on top). Connection width grows
* logarithmically with the daily hit rate (base-2 log, one hit/day
* renders zero width, each doubling adds a fixed step, uncapped);
* connections carrying less than 1% of the total
* traffic are pruned, which naturally keeps the graph under ~100
* connections. Animated beads flow along every edge in each direction,
* emitted at time intervals inversely proportional (linear) to the
* directional count.
* External sources appear as nodes in a row above the map. Sources are
* identified from visit records in this order: utm_campaign, utm_source,
* referer, then other utm_* tags. Visits with a UTM tag are grouped under
* that tag's value, not under the referer domain. A UTM source node only
* becomes a clickable link when every visit carrying that UTM tag came
* from the same referer. External exits are full-size nodes in a row below
* the map, mirroring the source row, so the site itself stays in the
* middle. Each distinct full exit URL is its own node. Self-loops (reload
* pings) are skipped.
*/
import { MIN_READ_SECONDS, readMapOf } from './format.js'
// Nodes are constant-size pills (stadium rects) holding the slug and the
// view count on two centered lines. TNODE_BOUND is the pill's bounding
// radius, used for layout clearance and placement; connectors and flows
// use the exact outline geometry instead (pillContact below).
export const TNODE_W = 160
export const TNODE_H = 54
const TNODE_BOUND = Math.hypot(TNODE_W, TNODE_H) / 2
const PILL_R = TNODE_H / 2 // cap radius and straight-section half-height
const PILL_OFF = TNODE_W / 2 - PILL_R // x offset of the cap centers
/**
* Where the ray from a node center along (ux, uy) exits the pill outline
* (a capsule: straight top/bottom plus semicircular caps), enlarged by
* `margin`. Returns the distance `t` to the contact point and the outline
* arc position `s` of that point (see pillPointAt).
*/
const pillContact = (ux, uy, margin = 0) => {
const r = PILL_R + margin
const off = PILL_OFF + margin
const q = (Math.PI / 2) * r
// Straight top/bottom: valid when the crossing lands on the flat section.
let tf = Infinity
if (Math.abs(uy) > 1e-9) {
const t = r / Math.abs(uy)
if (Math.abs(t * ux) <= off + 1e-9) tf = t
}
// Rounded cap on the side the ray points to.
const cx = off * (ux >= 0 ? 1 : -1)
const disc = r * r - (cx * uy) ** 2
const tc = disc >= 0 ? cx * ux + Math.sqrt(disc) : Infinity
if (tf <= tc) {
const x = tf * ux
return { t: tf, s: uy > 0 ? q + off - x : q + 2 * off + Math.PI * r + x + off }
}
if (tc < Infinity) {
let th = Math.atan2(tc * uy, tc * ux - cx)
if (th < 0) th += 2 * Math.PI
const s = cx > 0
? th <= Math.PI / 2
? th * r
: q + 4 * off + Math.PI * r + (th - (3 * Math.PI) / 2) * r
: q + 2 * off + (th - Math.PI / 2) * r
return { t: tc, s }
}
return { t: TNODE_BOUND + margin, s: 0 }
}
/** Total perimeter of the (margined) pill outline. */
const pillPerimeter = (margin = 0) =>
4 * (PILL_OFF + margin) + 2 * Math.PI * (PILL_R + margin)
/**
* Point on the pill outline at arc position `s`, counterclockwise from the
* right cap tip: right cap up, top flat right-to-left, left cap down,
* bottom flat left-to-right, right cap up to the tip. Pills are never
* rotated, so the returned offset from the node center is in absolute
* coordinates.
*/
const pillPointAt = (s, margin = 0) => {
const r = PILL_R + margin
const off = PILL_OFF + margin
const P = pillPerimeter(margin)
const q = (Math.PI / 2) * r
s = ((s % P) + P) % P
if (s < q) {
const th = s / r
return [off + r * Math.cos(th), r * Math.sin(th)]
}
s -= q
if (s < 2 * off) return [off - s, r]
s -= 2 * off
if (s < Math.PI * r) {
const th = Math.PI / 2 + s / r
return [-off + r * Math.cos(th), r * Math.sin(th)]
}
s -= Math.PI * r
if (s < 2 * off) return [-off + s, -r]
s -= 2 * off
const th = (3 * Math.PI) / 2 + s / r
return [off + r * Math.cos(th), r * Math.sin(th)]
}
/** Unit tangent to the pill outline at arc position `s`, in the direction
* of increasing `s` (numeric; exact on both flats and caps). */
const pillTangent = (s, margin = 0) => {
const [x1, y1] = pillPointAt(s - 0.5, margin)
const [x2, y2] = pillPointAt(s + 0.5, margin)
const m = Math.hypot(x2 - x1, y2 - y1) || 1
return [(x2 - x1) / m, (y2 - y1) / m]
}
// Edge width: half-width = WIDTH_GROWTH * log2(daily / DAILY_REF),
// where `daily` is the connection's hit rate in hits/day (callers scale
// raw counts by DAY / range). DAILY_REF hits/day renders zero width;
// WIDTH_GROWTH is the half-width added per doubling of the rate.
// Connections whose thin middle would render below MIN_WMID are culled
// entirely (fainter strands are practically invisible), as are those
// carrying less than PRUNE_FRACTION of the total traffic (this also
// keeps the graph under ~100 connections).
const WIDTH_GROWTH = 1.4 // half-width px per doubling of the daily rate
const DAILY_REF = 0.8 // hits/day at which the width is zero
const PRUNE_FRACTION = 0.01
// ~0.8 px full width at natural size (1 viewBox unit = 1 px).
const MIN_WMID = 0.4
// Beads: each edge direction emits beads at dailyRate * BEAD_RATE beads
// per second (linear in the daily hit rate). The rate is much reduced
// from real time to keep the animation lightweight. The component
// simulates every bead independently in JS with a constant traversal
// time per edge (speed relative to span length), with no limit on beads
// in flight.
export const BEAD_R = 3.2
const BEAD_RATE = 0.0084 // beads per second per hit/day
const FLOW_OFFSET = 3 // lane offset to the right of the travel direction
const MAX_EXT_IN = 8 // referer nodes in the top row
const MAX_EXT_OUT = 12 // exit nodes in the bottom row
const EXT_GAP = 12 // vertical margin of the source/exit rows to the map
/** Flatten the site tree into navigation order via DFS. */
function buildNavigationOrder(pageTree) {
const order = new Map()
const walk = (items) => {
for (const item of items || []) {
const p = `/${item.path}`
if (!order.has(p)) order.set(p, order.size)
walk(item.children)
}
}
walk(pageTree)
return order
}
/** Map page paths to their article titles from the site tree. */
function buildTitleMap(pageTree) {
const titles = new Map()
const walk = (items) => {
for (const item of items || []) {
titles.set(`/${item.path}`, item.title)
walk(item.children)
}
}
walk(pageTree)
return titles
}
/** Extract internal page-to-page transitions, excluding self-loops. */
function collectInternalTransitions(transitions) {
const internal = []
for (const [fr, tos] of Object.entries(transitions || {})) {
if (!fr.startsWith('/')) continue
for (const [to, count] of Object.entries(tos)) {
if (to.startsWith('/') && to !== fr) internal.push({ fr, to, count })
}
}
return internal
}
/** Domain-only label for an external origin (path and www. removed).
* Pills clip the text at their border; no length cap needed. */
function extLabel(ext) {
try {
return new URL(ext).hostname.replace(/^www\./, '')
} catch {
return ext.replace(/^https?:\/\//, '').replace(/^www\./, '').split('/')[0]
}
}
/**
* Collect outgoing external transitions: page path -> full exit URL.
* Aggregated per (URL, page) pair. Incoming external links are now derived
* from visit records (which carry UTM tags), so only exits remain here.
*/
function collectExitPairs(transitions) {
const pairs = new Map() // `${ext} ${page}` -> {ext, page, out}
for (const [fr, tos] of Object.entries(transitions || {})) {
if (!fr.startsWith('/')) continue // ignore external -> anything
for (const [to, count] of Object.entries(tos)) {
if (!to.startsWith('http')) continue
const k = `${to} ${fr}`
const p = pairs.get(k) || { ext: to, page: fr, out: 0 }
p.out += count
pairs.set(k, p)
}
}
return [...pairs.values()]
}
/** Build nodes with depth and a path lookup map; children are wired to parents. */
function buildNodeTree(internal, navOrder) {
const paths = new Set(['/', ...navOrder.keys()])
for (const e of internal) { paths.add(e.fr); paths.add(e.to) }
const depth = (p) => (p === '/' ? 0 : p.split('/').length - 1)
const nodes = [...paths].map((p) => ({
path: p, depth: depth(p), angle: 0, children: [],
}))
const byPath = new Map(nodes.map((n) => [n.path, n]))
// Parent is the nearest ancestor present in the map, front page last.
const parentOf = (p) => {
let q = p
while (q !== '/') {
q = q.slice(0, q.lastIndexOf('/')) || '/'
if (byPath.has(q)) return byPath.get(q)
}
return byPath.get('/')
}
for (const n of nodes) {
if (n.path !== '/') parentOf(n.path).children.push(n)
}
return { nodes, byPath, root: byPath.get('/') }
}
/** Sort each node's children by navigation order, recursively. */
function sortByNav(root, navOrder) {
const byNav = (a, b) =>
(navOrder.get(a.path) ?? Infinity) - (navOrder.get(b.path) ?? Infinity)
|| a.path.localeCompare(b.path)
const walk = (n) => {
n.children.sort(byNav)
n.children.forEach(walk)
}
walk(root)
}
/** Compute median reading time per article in seconds. */
function buildReadSeconds(visits) {
const times = {}
for (const v of visits || []) {
for (const [path, sec] of Object.entries(readMapOf(v))) {
if (sec >= MIN_READ_SECONDS) {
; (times[path] || (times[path] = [])).push(sec)
}
}
}
const seconds = {}
for (const [path, arr] of Object.entries(times)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median =
arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
seconds[path] = Math.round(median)
}
return seconds
}
/** Compute view counts, labels and hidden flags for each node. */
function annotateNodes(nodes, viewsData, titles, readSeconds) {
const viewCount = (p) => {
let n = 0
for (const c of Object.values(viewsData?.[p] || {})) n += c
return n
}
for (const n of nodes) {
n.views = viewCount(n.path)
n.readSec = readSeconds[n.path] || 0
// Article title inside the pill (clipped at the pill border on
// render), slug as fallback for pages missing from the site tree.
n.label = titles.get(n.path) || (n.path === '/' ? '🏠︎' : n.path.split('/').pop())
n.title = titles.get(n.path) || ''
// Category (non-leaf) pages with no views in this window are omitted:
// their children move up in their place (see layoutGroups).
n.hidden = n.children.length > 0 && n.views === 0
}
}
/**
* Top-down layout following the menu structure: top-level items in an
* equally spaced row at the top (right below the external source row),
* the row following a shallow circular sag (center lowest) so connections
* between neighbors do not overlap the pills in between. Each top item's
* whole subtree fans out from it in menu (DFS preorder) order along a
* large-radius circular arc that leaves the parent heading straight down
* and gradually bends to the right — no horizontal space is reserved for fans, they
* extend under the slots to their right. Hidden index pages are omitted
* from the fan; when the top item itself is hidden, the fan shifts one
* slot up, the first visible child taking the top position. Branch lanes
* labeled with the branch slug (see the branch-lane pass at the end)
* keep the omitted menu levels visible.
*/
function layoutGroups(root) {
// Top slots are spaced well over one pill width apart regardless of
// fan sizes.
const SLOT = TNODE_W + 100
const CLEAR = TNODE_W * 0.8 // fan spacing per member along the curve
// First pass: visible members per group, in menu order. Hidden index
// pages are skipped, but their children still appear. The front page
// forms its own group. groupRoots keeps each group's subtree root for
// the branch-curve pass below.
const groups = []
const groupRoots = []
for (const g of [root, ...root.children]) {
const members = []
if (g === root) {
if (!g.hidden) members.push(g)
} else {
const walk = (n) => {
if (!n.hidden) members.push(n)
n.children.forEach(walk)
}
walk(g)
}
if (members.length) {
groups.push(members)
groupRoots.push(g)
}
}
// Top row on a large-radius circular arc whose bottom point is the
// LAST top item: each earlier item sits a bit higher (drop = 15% of
// the row span). Flat row when there is a single group.
const half = ((groups.length - 1) * SLOT) / 2 || 1
const span = (groups.length - 1) * SLOT
const topD = span * 0.15
const R_T = span ? (span * span + topD * topD) / (2 * topD) : 0
const topY = span
? (x) => topD - R_T + Math.sqrt(R_T * R_T - (x - half) * (x - half))
: () => 0
// Second pass: place groups. Fan members follow a circular arc of
// large radius FAN_R centered at (gx + FAN_R, y0): the trail leaves
// the top node heading straight down (vertical tangent) and bends
// right gently, member i at arc angle π i·CLEAR/FAN_R (spaced by
// arc length CLEAR). A circle — not a spline — so the branch lanes
// below can be concentric arcs: identical forms, only radii differ.
const FAN_R = 1000
groups.forEach((members, gi) => {
const gx = gi * SLOT - half
const y0 = topY(gx)
members[0].x = gx
members[0].y = y0
for (let i = 1; i < members.length; i++) {
const th = Math.PI - (i * CLEAR) / FAN_R
members[i].x = gx + FAN_R * (1 + Math.cos(th))
members[i].y = y0 + FAN_R * Math.sin(th)
}
})
// Branch lanes: one wide arc per path prefix (slug depth ≥ 1) whose
// subtree holds at least two visible nodes (a branch's visible nodes
// form one contiguous run in the fan's DFS preorder). Every lane of a
// group is an arc around the group's fan center with a radius one
// INDENT larger per parent level — concentric circles, so all lanes
// share exactly one form. Lanes span their branch's nodes plus a
// little extra tucked under the first/last pill (so the line caps are
// never visible) and run behind the pills. A label arc carries the
// branch slug, left-aligned just past the first pill and free to run
// to the lane's end — longer text simply passes under later pills,
// which are drawn on top. Hidden (unplaced) index pages still
// define a lane: it follows their promoted children, so lanes reflect
// the path structure rather than page existence.
const INDENT = 20 // lane spacing (radius) per nesting level (> lane width)
const END_TUCK = 22 // arc units tucked under the first/last pill
const GAP_TRIM = 32 // label arc clearance from the pills
// Labels are left-aligned on their guide: the guide starts just past the
// source pill's edge, the earliest point where the text is visible.
const LABEL_PAD = 6
// The label guide rides GUIDE_OFF outward of the lane centerline: the
// text's alphabetic baseline sits on the guide, so this puts the
// glyph middle (not the baseline) on the lane center at any zoom —
// dominant-baseline tricks are em-based and break under downscale.
const GUIDE_OFF = 3.5
const branches = []
groups.forEach((members, gi) => {
const g = groupRoots[gi]
if (g === root || members.length < 2) return
const idx = new Map(members.map((n, i) => [n, i]))
const C = [gi * SLOT - half + FAN_R, topY(gi * SLOT - half)]
const walk = (n) => {
let first = Infinity
let last = -1
const span = (m) => {
const k = idx.get(m)
if (k !== undefined) {
first = Math.min(first, k)
last = Math.max(last, k)
}
m.children.forEach(span)
}
span(n)
if (n.depth >= 1 && last > first) {
branches.push({ depth: n.depth, name: n.path.split('/').pop(), C, first, last })
}
n.children.forEach(walk)
}
walk(g)
})
const depthMax = branches.reduce((d, b) => Math.max(d, b.depth), 1)
let arcLeft = Infinity // leftmost lane point, for the bounding box
const arcs = branches.map(({ depth, name, C, first, last }) => {
const R = FAN_R + (depthMax - depth) * INDENT
const th = (i) => Math.PI - (i * CLEAR) / FAN_R
const pt = (a, r) => [C[0] + r * Math.cos(a), C[1] + r * Math.sin(a)]
// Arc from angle a down to angle b (a > b; visually counterclockwise
// from the west point downward, hence sweep flag 0).
const arc = (a, b, r) => {
const [x0, y0] = pt(a, r)
const [x1, y1] = pt(b, r)
return `M ${x0.toFixed(2)} ${y0.toFixed(2)} A ${r.toFixed(2)} ${r.toFixed(2)} 0 0 0 ${x1.toFixed(2)} ${y1.toFixed(2)}`
}
const d = arc(th(first) + END_TUCK / R, th(last) - END_TUCK / R, R)
// Start just past the first pill: lanes leave the source node nearly
// vertically, so the pill's extent along the arc is its half height.
// The guide runs to the lane's end so long slugs are never cut off.
const ld = arc(th(first) - (TNODE_H / 2 + LABEL_PAD) / R,
th(last) - END_TUCK / R, R + GUIDE_OFF)
arcLeft = Math.min(arcLeft, pt(th(first) + END_TUCK / R, R)[0])
return { d, ld, label: name }
})
// Top lane: an arc along the top row's own circle, connecting
// the top nodes of all groups and tucked under the first and last of
// them (the arc bottoms at the last item, so it continues rightward
// under its pill). Drawn 50% thicker than branch lanes. A 🏠︎ label
// marks the lane right after the home pill, on a guide arc like
// the branch labels but with the offset and clearance scaled up by the
// same 50% to keep the glyph centered on the wider lane.
if (span) {
const d = `M ${(-half - END_TUCK).toFixed(2)} ${topY(-half - END_TUCK).toFixed(2)} `
+ `A ${R_T.toFixed(2)} ${R_T.toFixed(2)} 0 0 0 ${(half + END_TUCK).toFixed(2)} ${topY(half + END_TUCK).toFixed(2)}`
const rG = R_T + GUIDE_OFF * 1.5
const ptG = (x) => [x, topD - R_T + Math.sqrt(rG * rG - (x - half) ** 2)]
// Left-aligned like the branch labels: the guide starts just past the
// home pill's edge (scaled with the lane thickness).
const g0 = -half + TNODE_W / 2 + LABEL_PAD * 1.5
const g1 = SLOT - half - TNODE_W / 2 - GAP_TRIM * 1.5
const [gx0, gy0] = ptG(g0)
const [gx1, gy1] = ptG(g1)
arcs.unshift({
d,
ld: `M ${gx0.toFixed(2)} ${gy0.toFixed(2)} A ${rG.toFixed(2)} ${rG.toFixed(2)} 0 0 0 ${gx1.toFixed(2)} ${gy1.toFixed(2)}`,
label: '🏠︎',
top: true,
})
}
return { arcs, arcLeft }
}
/** Collapse opposite transition directions into one unordered pair per page pair. */
function aggregatePairs(internal) {
const pairs = new Map() // unordered pair key -> [countAB, countBA]
for (const e of internal) {
const forward = e.fr < e.to
const k = forward ? `${e.fr} ${e.to}` : `${e.to} ${e.fr}`
const c = pairs.get(k) || [0, 0]
c[forward ? 0 : 1] += e.count
pairs.set(k, c)
}
return pairs
}
const fmtPt = (p) => `${p[0].toFixed(2)} ${p[1].toFixed(2)}`
/**
* Build one ribbon edge between two nodes with counts ab and ba.
* `wMid` is the half-width of the thin middle (already strength-scaled by
* the caller). Each end flares into the node's pill surround (the outline
* enlarged by margin S): the flare contact points follow the pill outline
* a constant arc distance to each side of the direct contact point, and
* the back of the ribbon wraps all the way around the pill between them,
* surrounding the node. The pills themselves are drawn on top.
*/
function buildRibbon(a, b, ab, ba, wMid, external = false) {
const count = ab + ba
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len
const uy = (b.y - a.y) / len
const nx = -uy
const ny = ux
// Direct contact: where the centerline exits each pill's surround.
const S = 4
const cA = pillContact(ux, uy, S)
const cB = pillContact(-ux, -uy, S)
// Flares take a fair share of the free span while leaving the
// count-scaled thin middle a visible share of the connection length.
// The maximum flare length scales with the contact distance so wide
// approach angles still show a wide connector end.
const free = Math.max(0, len - cA.t - cB.t)
const FLARE = Math.min(Math.max(cA.t, cB.t) * 1.2, free * 0.4)
// Flare endpoints: walk the outline a constant arc distance to each
// side of the direct contact point (spanning flats and caps alike).
const D = (Math.PI / 4) * (PILL_R + S)
// Per node: endpoints for the +n (left) and -n (right) flare sides,
// each with its arc position, absolute point, and an outline tangent
// oriented back toward the direct contact point.
const ends = (cx, cy, contact) => {
const pick = (s) => {
const [px, py] = pillPointAt(s, S)
// Outline tangent oriented back toward the direct contact point
// (the flare side sweeps from the contact point around to its
// endpoint and into the connection), so it can never fork outward.
const tan = pillTangent(s, S)
if (s > contact.s) { tan[0] = -tan[0]; tan[1] = -tan[1] }
return { s, p: [cx + px, cy + py], tan, side: px * nx + py * ny }
}
const plus = pick(contact.s + D)
const minus = pick(contact.s - D)
return plus.side >= 0 ? [plus, minus] : [minus, plus]
}
const [aLeftEnd, aRightEnd] = ends(a.x, a.y, cA)
const [bLeftEnd, bRightEnd] = ends(b.x, b.y, cB)
// Point on the connection centerline at distance t from A, offset s
// perpendicular to it.
const P = (t, s) => [
a.x + t * ux + s * nx,
a.y + t * uy + s * ny,
]
// One side of a flare: from the outline endpoint, leaving tangent to
// the pill outline, to the connection middle arriving parallel with
// the centerline. The tangent pull is clamped so the control point
// stays well on its own side of the centerline — otherwise a long
// flare on a rounded cap crosses the opposite side.
const flarePoints = (end, midT, s, dir) => {
let hEnd = FLARE * 0.65
const hMid = FLARE * 0.4
const tanS = end.tan[0] * nx + end.tan[1] * ny // inward rate
if (tanS * end.side < 0) {
hEnd = Math.min(hEnd, (Math.abs(end.side) * 0.6) / Math.abs(tanS))
}
return {
pEnd: end.p,
cEnd: [end.p[0] + end.tan[0] * hEnd, end.p[1] + end.tan[1] * hEnd],
cMid: P(midT - dir * hMid, s * wMid),
pMid: P(midT, s * wMid),
}
}
// Emit a cubic in either traversal direction. Reversing a cubic requires
// swapping its control points, rather than recalculating the geometry.
const curve = (f, reverse = false) => {
if (!reverse) {
return `C ${fmtPt(f.cEnd)} ${fmtPt(f.cMid)} ${fmtPt(f.pMid)} `
}
return `C ${fmtPt(f.cMid)} ${fmtPt(f.cEnd)} ${fmtPt(f.pEnd)} `
}
// Trace the surround outline the long way around (behind the node) from
// arc s1 to arc s2. Sampled as a polyline: the visible result is a thin
// halo hugging the pill, so exact arc segments are unnecessary.
const outlineWrap = (cx, cy, s1, s2) => {
const per = pillPerimeter(S)
const dPlus = ((s2 - s1) % per + per) % per
const total = dPlus > per / 2 ? dPlus : per - dPlus
const dir = dPlus > per / 2 ? 1 : -1
const n = Math.max(4, Math.ceil(total / 6))
let out = ''
for (let i = 1; i <= n; i++) {
const [x, y] = pillPointAt(s1 + (dir * total * i) / n, S)
out += `L ${(cx + x).toFixed(2)} ${(cy + y).toFixed(2)} `
}
return out
}
const aLeft = flarePoints(aLeftEnd, cA.t + FLARE, 1, 1)
const bLeft = flarePoints(bLeftEnd, len - cB.t - FLARE, 1, -1)
const bRight = flarePoints(bRightEnd, len - cB.t - FLARE, -1, -1)
const aRight = flarePoints(aRightEnd, cA.t + FLARE, -1, 1)
// Each end wraps the full back of the node pill between its two flare
// contact points (bLeft -> bRight around B, aRight -> aLeft around A).
const d = `M ${fmtPt(aLeft.pEnd)} `
+ curve(aLeft)
+ `L ${fmtPt(bLeft.pMid)} `
+ curve(bLeft, true)
+ outlineWrap(b.x, b.y, bLeftEnd.s, bRightEnd.s)
+ curve(bRight)
+ `L ${fmtPt(aRight.pMid)} `
+ curve(aRight, true)
+ outlineWrap(a.x, a.y, aRightEnd.s, aLeftEnd.s)
+ 'Z'
return {
d,
title: `${a.path}${b.path}: ${count} (${ab} / ${ba})`,
external,
}
}
/**
* Flow descriptors for the bead animation, one per edge direction with a
* nonzero count: a straight segment running from inside the source node
* to inside the target node (beads render under the node pills, so
* they emerge from and vanish beneath the nodes rather than popping in
* at the surround), plus the emission interval (seconds between beads,
* inverse of the daily hit rate * BEAD_RATE). Each segment is offset to the
* right-hand side of its travel direction, so opposing flows on the same
* edge run on parallel lanes instead of colliding. The component turns
* these into independently simulated beads.
*/
function buildFlows(a, b, ab, ba, dayScale = 1) {
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len
const uy = (b.y - a.y) / len
const rA = pillContact(ux, uy).t
const rB = pillContact(-ux, -uy).t
const t0 = rA / 3
const t1 = len - rB / 3
if (t1 - t0 < 12) return []
// Unit normal pointing to the visual right of the A -> B direction.
const rx = -uy
const ry = ux
const span = t1 - t0
const flow = (count, fromT, toT) => {
// Each direction shifts to its own right, away from the opposing lane.
const s = fromT < toT ? FLOW_OFFSET : -FLOW_OFFSET
return {
x1: a.x + fromT * ux + s * rx,
y1: a.y + fromT * uy + s * ry,
x2: a.x + toT * ux + s * rx,
y2: a.y + toT * uy + s * ry,
len: span,
interval: 1 / (count * BEAD_RATE * dayScale),
}
}
const flows = []
// Stable key per edge direction so the component's bead simulation can
// match flows across data reloads and keep bead phases/positions.
if (ab) flows.push({ ...flow(ab, t0, t1), key: `${a.path} ${b.path}` })
if (ba) flows.push({ ...flow(ba, t1, t0), key: `${b.path} ${a.path}` })
return flows
}
/**
* Half-width for a connection middle: base-2 logarithmic in the daily
* hit rate, zero at DAILY_REF hits/day, uncapped. Absolute on purpose —
* cool routes stay visible regardless of how hot the hottest connection
* is. Callers cull results below MIN_WMID.
*/
const scaledWidth = (daily) => {
if (daily <= 0) return 0
return WIDTH_GROWTH * Math.log2(daily / DAILY_REF)
}
/**
* Build ribbon edges and bead flows for every aggregated page-to-page
* pair. Pairs carrying less than PRUNE_FRACTION of the total internal
* traffic are pruned (this naturally bounds the graph to ~100 edges).
*/
function buildInternalEdges(pairs, byPath, dayScale = 1) {
let total = 0
for (const [, [ab, ba]] of pairs) total += ab + ba
const minCount = total * PRUNE_FRACTION
const edges = []
const flows = []
for (const [k, [ab, ba]] of pairs) {
if (ab + ba < minCount) continue
const [pf, pt] = k.split(' ')
const a = byPath.get(pf)
const b = byPath.get(pt)
if (a.hidden || b.hidden) continue // unplaced index pages are omitted
const wMid = scaledWidth((ab + ba) * dayScale)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(a, b, ab, ba, wMid))
flows.push(...buildFlows(a, b, ab, ba, dayScale))
}
return { edges, flows }
}
const UTM_PRIORITY = ['utm_campaign', 'utm_source']
const UTM_FALLBACK = ['utm_medium', 'utm_content', 'utm_term', 'utm_id']
/** Identify the source of a visit according to the requested priority. */
function identifySource(visit) {
const utm = visit.utm || {}
for (const k of UTM_PRIORITY) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
if (visit.referer?.startsWith('http')) {
return { value: visit.referer, isUtm: false }
}
for (const k of UTM_FALLBACK) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
return null
}
/**
* Collect source -> entry page pairs from visit records. Sources are
* identified by UTM campaign/source (then referer, then other UTM tags).
* A UTM source only gets a link href when every visit using that source
* came from the same referer; referer sources always link to their origin.
*/
function collectSourcePairs(visits) {
const groups = new Map() // `${source}\0${page}` -> pair
for (const v of visits || []) {
const src = identifySource(v)
if (!src) continue
const k = `${src.value}\0${v.entry}`
const p = groups.get(k) || {
source: src.value,
page: v.entry,
in: 0,
refs: new Set(),
missingRef: false,
href: null,
isUtm: src.isUtm,
}
p.in += 1
if (v.referer?.startsWith('http')) {
p.refs.add(v.referer)
} else {
p.missingRef = true
}
groups.set(k, p)
}
for (const p of groups.values()) {
if (p.isUtm && !p.missingRef && p.refs.size === 1) {
const ref = [...p.refs][0]
if (ref.startsWith('http')) p.href = ref
} else if (!p.isUtm && p.source.startsWith('http')) {
p.href = p.source
}
}
return [...groups.values()]
}
/**
* Place external source and exit nodes and build their edges and bead
* flows.
* Sources (incoming links) are derived from visit UTM/referer data and form
* a row centered above the map, hottest first; exits come from the
* transition matrix and form a matching row centered below the map, so
* the site itself stays in the middle. Both rows sit EXT_GAP beyond the
* map's bounds.
* Widths and pruning use the same log scale and traffic-share rule as
* internal connections.
*/
function buildExternal({ sources, exits }, byPath, innerBounds, dayScale = 1) {
const extNodes = []
const edges = []
const flows = []
let extTotal = 0
for (const p of sources) extTotal += p.in
for (const p of exits) extTotal += p.out
const minCount = extTotal * PRUNE_FRACTION
const liveSources = sources.filter((p) => byPath.has(p.page))
const liveExits = exits.filter((p) => byPath.has(p.page))
if (!liveSources.length && !liveExits.length) return { extNodes, edges, flows }
const width = (count) => scaledWidth(count * dayScale)
// Incoming: one source node per identified source, in a row centered
// above the map, with an edge to each page that source led to. A source
// whose connectors are all culled (below MIN_WMID) is dropped itself.
const bySource = new Map() // source -> pairs, sorted by total incoming count
for (const p of liveSources.filter((p) => p.in >= minCount)) {
const g = bySource.get(p.source) || []
g.push(p)
bySource.set(p.source, g)
}
const origins = [...bySource]
.map(([source, ps]) => ({
source,
ps,
total: ps.reduce((s, p) => s + p.in, 0),
href: ps[0].href,
isUtm: ps[0].isUtm,
}))
.sort((a, b) => b.total - a.total)
.slice(0, MAX_EXT_IN)
.filter(({ ps }) =>
ps.some((p) => !byPath.get(p.page).hidden && width(p.in) >= MIN_WMID))
if (origins.length) {
const cx = (innerBounds.x0 + innerBounds.x1) / 2
const y = innerBounds.y0 - TNODE_BOUND - EXT_GAP
const spacing = TNODE_W + 44
const x0 = cx - ((origins.length - 1) * spacing) / 2
origins.forEach(({ source, ps, total, href, isUtm }, i) => {
const label = isUtm ? source : extLabel(source)
const xn = {
path: source,
href,
label, // clipped at the pill border on render
x: x0 + i * spacing,
y,
count: total,
kind: 'source',
}
extNodes.push(xn)
for (const p of ps) {
const page = byPath.get(p.page)
if (page.hidden) continue
const wMid = width(p.in)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(xn, page, p.in, 0, wMid, true))
flows.push(...buildFlows(xn, page, p.in, 0, dayScale))
}
})
}
// Outgoing: one exit node per distinct full URL (so several links to
// the same domain stay distinct), showing the total count across all
// pages linking to it, in a row centered below the map (hottest
// first), mirroring the source row above. Each (URL, page) pair
// contributes an edge from that page. An exit whose connectors are all
// culled (below MIN_WMID) is dropped itself.
const byExt = new Map() // full URL -> { ext, out, pairs }
for (const p of liveExits.filter((p) => p.out >= minCount)) {
const g = byExt.get(p.ext) || { ext: p.ext, out: 0, pairs: [] }
g.out += p.out
g.pairs.push(p)
byExt.set(p.ext, g)
}
const targets = [...byExt.values()]
.sort((a, b) => b.out - a.out)
.slice(0, MAX_EXT_OUT)
.filter(({ pairs }) =>
pairs.some((p) => !byPath.get(p.page).hidden && width(p.out) >= MIN_WMID))
if (targets.length) {
const cx = (innerBounds.x0 + innerBounds.x1) / 2
const y = innerBounds.y1 + TNODE_BOUND + EXT_GAP
const spacing = TNODE_W + 44
const x0 = cx - ((targets.length - 1) * spacing) / 2
targets.forEach(({ ext, out, pairs }, i) => {
const xn = {
path: ext,
href: ext,
label: extLabel(ext),
x: x0 + i * spacing,
y,
count: out,
kind: 'exit',
}
extNodes.push(xn)
for (const p of pairs) {
const page = byPath.get(p.page)
if (page.hidden) continue
const wMid = width(p.out)
if (wMid < MIN_WMID) continue
edges.push(buildRibbon(page, xn, p.out, 0, wMid, true))
flows.push(...buildFlows(page, xn, p.out, 0, dayScale))
}
})
}
return { extNodes, edges, flows }
}
/**
* Build the transition map model.
* Returns { nodes, edges, flows, extNodes, arcs, bounds } or null when
* there is nothing to show. `arcs` holds the branch curves; `nodes` only
* contains placed (visible) nodes. `dayScale` converts raw counts to a
* daily hit rate (DAY / range ms) for edge widths and bead rates.
*/
export function buildTransitionGraph(data, pageTree, visits = [], dayScale = 1) {
const internal = collectInternalTransitions(data?.transitions)
const sources = collectSourcePairs(visits)
const exits = collectExitPairs(data?.transitions)
const navOrder = buildNavigationOrder(pageTree)
const titles = buildTitleMap(pageTree)
const readSeconds = buildReadSeconds(visits)
if (!internal.length && !navOrder.size) return null
const { nodes, byPath, root } = buildNodeTree(internal, navOrder)
sortByNav(root, navOrder)
annotateNodes(nodes, data?.views, titles, readSeconds)
const { arcs, arcLeft } = layoutGroups(root)
const placed = nodes.filter((n) => !n.hidden)
const pairs = aggregatePairs(internal)
const { edges, flows } = buildInternalEdges(pairs, byPath, dayScale)
// Tight bounding box of the placed page nodes, extended to cover the
// branch curves running left of the pills; external nodes extend it.
// Margins cover just the ribbon surround (pill outline + flare margin
// S=4) plus a small pad: the pill's half width/height, not its diagonal
// radius, so the map crops tight especially at top and bottom.
const pad = 8
const MX = TNODE_W / 2 + 4 + pad
const MY = TNODE_H / 2 + 4 + pad
const xs = placed.map((n) => n.x)
const ys = placed.map((n) => n.y)
const bounds = {
x0: Math.min(Math.min(...xs) - MX, arcLeft - pad),
y0: Math.min(...ys) - MY,
x1: Math.max(...xs) + MX,
y1: Math.max(...ys) + MY,
}
const ext = buildExternal({ sources, exits }, byPath, bounds, dayScale)
for (const xn of ext.extNodes) {
bounds.x0 = Math.min(bounds.x0, xn.x - MX)
bounds.y0 = Math.min(bounds.y0, xn.y - MY)
bounds.x1 = Math.max(bounds.x1, xn.x + MX)
bounds.y1 = Math.max(bounds.y1, xn.y + MY)
}
return {
nodes: placed,
edges: [...edges, ...ext.edges],
flows: [...flows, ...ext.flows],
extNodes: ext.extNodes,
arcs,
bounds,
}
}
+266 -88
View File
@@ -56,6 +56,11 @@
/* The brand follows the heading font unless overridden separately. */ /* The brand follows the heading font unless overridden separately. */
--font-brand: var(--font-heading); --font-brand: var(--font-heading);
--font-code: var(--font-fira-code); --font-code: var(--font-fira-code);
/* Target x-height ratio for code text: set to the body font's ratio
(here Source Sans 3's 0.478) so font-size-adjust can scale the code
font to the same optical height. Themes retune it to their body font
(measured: Literata 0.507, Inter 0.546, Montserrat 0.517, Cause 0.5). */
--code-x-height: 0.478;
/* Width of the docked editor panel (used both here for shifting the page /* Width of the docked editor panel (used both here for shifting the page
and in the Vue editor's own styles). */ and in the Vue editor's own styles). */
--editor-w: min(46rem, 50vw); --editor-w: min(46rem, 50vw);
@@ -65,8 +70,23 @@
box-sizing: border-box; box-sizing: border-box;
} }
::selection {
background: color-mix(var(--accent) 30%, transparent);
color: inherit;
}
/* Links never underline — including SVG link text, which the UA stylesheet
underlines by default. */
a {
text-decoration: none;
}
html { html {
scroll-behavior: smooth; scroll-behavior: smooth;
/* No rubber-band bounce past the page ends (macOS trackpads): the banner
parallax in --pry is driven by scrollY and overscroll would let the
artwork drift beyond its designed range. */
overscroll-behavior: none;
/* Native-scrollbar fallback styling (JS off or before pagerite.js runs): /* Native-scrollbar fallback styling (JS off or before pagerite.js runs):
thin, theme-muted thumb on a transparent track. With JS the scrollbars thin, theme-muted thumb on a transparent track. With JS the scrollbars
are replaced by OverlayScrollbars (see pagerite.js) — floating, are replaced by OverlayScrollbars (see pagerite.js) — floating,
@@ -116,7 +136,11 @@ body {
display: flex; display: flex;
flex-direction: column; flex-direction: column;
justify-content: flex-end; justify-content: flex-end;
min-height: 11rem; /* Height scales down proportionally on small screens: 13rem at 800px
(50rem) viewport, shrinking with the smaller of viewport width/height
(vmin) below that, floored at 8rem. */
height: clamp(8rem, 26vmin, 13rem);
box-sizing: content-box;
background: linear-gradient(135deg, var(--surface), var(--bg)); background: linear-gradient(135deg, var(--surface), var(--bg));
border-bottom: 1px solid var(--line); border-bottom: 1px solid var(--line);
} }
@@ -181,11 +205,17 @@ body {
#brand { #brand {
font-family: var(--font-brand); font-family: var(--font-brand);
font-weight: 700; font-weight: 700;
font-size: 2.4rem; /* Scales down proportionally on small screens, same curve as the banner
height: 2.4rem at 800px, shrinking with vmin below that. */
font-size: clamp(1.4rem, 4.8vmin, 2.4rem);
text-decoration: none; text-decoration: none;
/* One line always: pagerite.js shrinks the font size to fit instead of /* One line always: pagerite.js shrinks the font size to fit instead of
wrapping (the themed size is the maximum). */ wrapping (the themed size is the maximum). */
white-space: nowrap; white-space: nowrap;
/* Shrink-wrap to the text: as a flex child of the column-direction
#banner it would otherwise stretch full-width, making the empty banner
area beside the text a link to the front page. */
align-self: flex-start;
margin: auto 1.25rem 0; margin: auto 1.25rem 0;
padding-top: 1.5rem; padding-top: 1.5rem;
color: var(--text); color: var(--text);
@@ -243,8 +273,6 @@ body {
position: relative; position: relative;
display: grid; display: grid;
grid-template-columns: minmax(0, 1fr) minmax(0, 78rem) minmax(0, 1fr); grid-template-columns: minmax(0, 1fr) minmax(0, 78rem) minmax(0, 1fr);
/* The docked editor pushes the content (not the header) right. */
transition: margin-left 0.25s ease;
} }
/* Long articles (.multicol is added by pagerite.js based on content length) /* Long articles (.multicol is added by pagerite.js based on content length)
@@ -259,10 +287,9 @@ body:has(.multicol) #content {
body.editing #content { body.editing #content {
margin-left: var(--editor-w); margin-left: var(--editor-w);
padding-left: 1rem; /* No padding/gap here: the article area starts flush at the panel's right
/* gap between the docked editor and the content */ edge so that full-bleed .wide images (anchored to that edge below) line
/* No overflow clipping here: .editor-host lives outside this box up with it exactly. */
(negative left), and .wide shrink-wraps to the remaining space. */
} }
/* The sidebar's gutter space is needed by the editor instead. */ /* The sidebar's gutter space is needed by the editor instead. */
@@ -270,34 +297,23 @@ body.editing #sidebar {
display: none; display: none;
} }
/* While editing, the window itself does not scroll: the editor panel is /* While editing, the window keeps scrolling normally (the overlay
exactly the remaining window height, and only the article area (#main) scrollbars take no layout space, so the vw-based .wide bleed stays exact
scrolls. Without the editor, normal full-page scroll applies. */ and no horizontal scrollbar appears). The editor panel is fixed to the
body.editing { viewport's left edge; main.js sets its top each scroll frame — the
height: 100vh; banner's bottom edge while the banner is visible, else the viewport top.
overflow: hidden; The panel scrolls internally. */
}
body.editing #content {
min-height: 0;
grid-template-rows: 100%;
}
body.editing #main {
height: 100%;
min-height: 0;
overflow-y: auto;
}
/* The editor host lives inside #content: it starts below the banner and
ends above the footer. With the page scroll locked while editing, the
panel exactly fills that space — the full available window height. */
.editor-host { .editor-host {
position: absolute; position: fixed;
top: 0; top: 0;
/* main.js: banner bottom while visible, else 0 */
bottom: 0; bottom: 0;
left: calc(0px - var(--editor-w)); left: 0;
width: var(--editor-w); width: var(--editor-w);
/* Above the sidebar, .edit-link and the banner's top-right pens
(z-index 10) while sliding in/out (the host now lives at the end of
<body>, so it needs its own stacking level). */
z-index: 10;
} }
.editor-root.overlay { .editor-root.overlay {
@@ -375,7 +391,7 @@ body.editing #main {
#sidebar ul ul { #sidebar ul ul {
gap: 0.5em; gap: 0.5em;
margin-top: 0.5em; margin-top: 0.5em;
padding-inline-start: 1em; padding-inline-start: 1.8em;
} }
#sidebar ul ul li::before { #sidebar ul ul li::before {
@@ -408,7 +424,10 @@ main {
article h1, article h1,
article h2, article h2,
article h3 { article h3,
article h4,
article h5,
article h6 {
font-family: var(--font-heading); font-family: var(--font-heading);
font-weight: 600; font-weight: 600;
line-height: 1.25; line-height: 1.25;
@@ -428,13 +447,31 @@ article dl,
article blockquote, article blockquote,
article pre, article pre,
article figure, article figure,
article table { article table,
article h3,
article h4,
article h5,
article h6 {
margin-top: 0; margin-top: 0;
margin-bottom: 1rem; margin-bottom: 1rem;
} }
article h3 { /* Headings separate from the text above via a top margin on the sibling
margin: 1.4rem 0 0.4rem; combinator: a heading that is the first child of a container (e.g. the
top of a .colseg column segment) gets no gap, and browsers truncate the
margin at column breaks, so column tops stay aligned. */
article h3,
article h4,
article h5,
article h6 {
margin-bottom: 0.4rem;
}
article * + h3,
article * + h4,
article * + h5,
article * + h6 {
margin-top: 1.4rem;
} }
/* Lists: small diamond emoji markers — blue 🔹 on odd nesting levels, /* Lists: small diamond emoji markers — blue 🔹 on odd nesting levels,
@@ -442,7 +479,7 @@ article h3 {
lines align. */ lines align. */
article ul { article ul {
list-style: none; list-style: none;
padding-inline-start: 1em; padding-inline-start: 1.8em;
} }
article ul li::before { article ul li::before {
@@ -525,9 +562,10 @@ article dd {
} }
/* Multi-column reading, but only for long articles (pagerite.js adds /* Multi-column reading, but only for long articles (pagerite.js adds
.multicol based on content length and splits the body into .colseg segments .multicol based on content length — code blocks excluded — and splits
separated by full-width h2s and wide figures; only segments with enough the body into .colseg segments separated by full-width h2s and .wide
text get .cols). No fixed breakpoint: `columns: 30rem` lets CSS fit as elements; only segments with enough text get .cols, and a ::: nocols
container opts its section out). No fixed breakpoint: `columns: 30rem` lets CSS fit as
many columns of at least 30rem as the article's current width allows — many columns of at least 30rem as the article's current width allows —
since .multicol also uncaps the article width (see #content above), a since .multicol also uncaps the article width (see #content above), a
wider window simply yields more columns. */ wider window simply yields more columns. */
@@ -540,6 +578,13 @@ article dd {
.multicol .colseg { .multicol .colseg {
margin-bottom: 1rem; margin-bottom: 1rem;
h3,
h4,
h5,
h6 {
break-after: avoid-column;
}
p, p,
li { li {
break-inside: avoid-column; break-inside: avoid-column;
@@ -549,7 +594,9 @@ article dd {
pre, pre,
blockquote, blockquote,
table, table,
dl { dl,
.admonition,
.markdown-alert {
break-inside: avoid; break-inside: avoid;
} }
} }
@@ -569,10 +616,11 @@ article a:hover {
color: var(--accent); color: var(--accent);
} }
/* Blockquotes: inner paragraphs carry no margins (spacing comes from the /* Blockquotes: spacing comes from the blockquote itself (bottom-only like
blockquote itself, bottom-only like everything else in articles). The everything else in articles); inner paragraphs keep only the gap between
negative left margin pushes the bar out past the text edge, so quoted them. The negative left margin pushes the bar out past the text edge, so
text aligns with the surrounding paragraphs — same trick as code blocks. */ quoted text aligns with the surrounding paragraphs — same trick as code
blocks. */
blockquote { blockquote {
margin: 0 0 1rem -0.5rem; margin: 0 0 1rem -0.5rem;
padding: 0 0 0 0.25rem; padding: 0 0 0 0.25rem;
@@ -584,42 +632,128 @@ blockquote p {
margin: 0; margin: 0;
} }
/* Admonitions (markdown !!! note/warning/...): a lightweight callout in blockquote p + p {
the blockquote idiom — accent bar and a faint wash, recolored per type. margin-top: 0.6rem;
}
/* Admonitions (markdown !!! note/warning/...) and GitHub-style alerts
(> [!NOTE] ...): a lightweight callout in the blockquote idiom — accent
bar and a faint wash, recolored per type, with a type emoji on the
title. The negative left margin pushes bar and wash out past the text
edge so the inner text aligns with surrounding paragraphs — same trick
as blockquotes and code blocks (margin-left = border + padding-left).
Bottom-only margins like everything else in articles; inner paragraphs Bottom-only margins like everything else in articles; inner paragraphs
carry no margins of their own. */ carry no margins of their own. */
.admonition { .admonition,
margin: 0 0 1rem; .markdown-alert {
margin: 0 0 1rem -1.15rem;
padding: 0.4rem 0.9rem; padding: 0.4rem 0.9rem;
border-left: 0.25rem solid var(--admonition-color, var(--accent)); border-left: 0.25rem solid var(--admonition-color, var(--accent));
border-radius: 0 0.3rem 0.3rem 0; border-radius: 0 0.3rem 0.3rem 0;
background: color-mix(in srgb, var(--admonition-color, var(--accent)) 7%, transparent); background: color-mix(in srgb, var(--admonition-color, var(--accent)) 7%, transparent);
} }
.admonition > :last-child { .admonition> :last-child,
.markdown-alert> :last-child {
margin-bottom: 0; margin-bottom: 0;
} }
.admonition-title { .admonition-title,
.markdown-alert-title {
margin: 0 0 0.2rem; margin: 0 0 0.2rem;
font-weight: 600; font-weight: 600;
color: var(--admonition-color, var(--accent)); color: var(--admonition-color, var(--accent));
} }
.admonition-title::before,
.markdown-alert-title::before {
padding-right: 0.35em;
}
.admonition.note .admonition-title::before,
.markdown-alert-note .markdown-alert-title::before {
content: "️";
}
.admonition.tip .admonition-title::before,
.admonition.hint .admonition-title::before,
.markdown-alert-tip .markdown-alert-title::before {
content: "✨";
}
.admonition.important .admonition-title::before,
.markdown-alert-important .markdown-alert-title::before {
content: "❗";
}
.admonition.success .admonition-title::before {
content: "✅";
}
.admonition.warning .admonition-title::before,
.markdown-alert-warning .markdown-alert-title::before {
content: "⚠️";
}
.admonition.caution .admonition-title::before,
.markdown-alert-caution .markdown-alert-title::before {
content: "🔥";
}
.admonition.danger .admonition-title::before,
.admonition.failure .admonition-title::before {
content: "⛔";
}
.admonition.tip, .admonition.tip,
.admonition.important, .admonition.important,
.admonition.hint, .admonition.hint,
.admonition.success { .admonition.success,
.markdown-alert-tip,
.markdown-alert-important {
--admonition-color: var(--accent2); --admonition-color: var(--accent2);
} }
.admonition.warning, .admonition.warning,
.admonition.caution, .admonition.caution,
.admonition.danger, .admonition.danger,
.admonition.failure { .admonition.failure,
.markdown-alert-warning,
.markdown-alert-caution {
--admonition-color: var(--accent3); --admonition-color: var(--accent3);
} }
/* Asides (::: aside): a floated side box in the floated-figure idiom;
consecutive asides stack (clear: right). On wide single-column pages it
leans into the empty right gutter (below 104rem the gutter cannot hold
the box; multicol pages have no right gutter at all, and while editing
the docked panel reshapes the gutters — in all these it stays a plain
float). Headings already clear floats, so asides never bleed into the
next section. */
.aside {
float: right;
clear: right;
width: 30%;
max-width: 20rem;
margin: 0.3rem 0 1rem 1.2rem;
padding: 0.6rem 0.9rem;
font-size: 0.9rem;
color: var(--muted);
background: color-mix(in srgb, var(--accent) 6%, transparent);
border-radius: 0.3rem;
}
.aside> :last-child {
margin-bottom: 0;
}
@media (min-width: 104rem) {
body:not(.editing):not(:has(.multicol)) .aside {
width: 12rem;
margin-right: -13rem;
}
}
pre { pre {
overflow-x: auto; overflow-x: auto;
padding: 0.5rem 0.8rem; padding: 0.5rem 0.8rem;
@@ -633,10 +767,15 @@ pre {
position: relative; position: relative;
} }
/* Inline code integrates with the text, no box of its own */ /* Inline code integrates with the text, no box of its own. Instead of a
p code, fixed em shrink (which can't fit every body/code font pairing — Fira
li code { Code's x-height ratio 0.525 is taller than Source Sans 3's 0.478 yet
font-size: 0.85em; shorter than Inter's 0.546), font-size-adjust scales whatever code font
is in use so its x-height matches the body font's ratio. Browsers
without font-size-adjust get unadjusted 1em code, which is fine. */
code {
font-family: var(--font-code);
font-size-adjust: ex-height var(--code-x-height);
} }
/* Click-to-copy button (added by pagerite.js) */ /* Click-to-copy button (added by pagerite.js) */
@@ -665,11 +804,6 @@ pre:hover .copy,
color: var(--accent); color: var(--accent);
} }
code {
font-family: var(--font-code);
font-size: 0.88em;
}
/* Tables separate by color, not lines: the header is a soft vertical /* Tables separate by color, not lines: the header is a soft vertical
gradient tinted with the theme's accent (themes can override the gradient tinted with the theme's accent (themes can override the
--table-head-* stops outright), body cells carry a very faint diagonal --table-head-* stops outright), body cells carry a very faint diagonal
@@ -698,12 +832,6 @@ td {
transparent 75%); transparent 75%);
} }
tbody tr:nth-child(even) td {
background: linear-gradient(160deg,
color-mix(in srgb, var(--table-tint, var(--accent)) 9%, transparent),
color-mix(in srgb, var(--table-tint, var(--accent)) 3%, transparent) 75%);
}
tbody tr+tr td { tbody tr+tr td {
border-top: 1px solid color-mix(in srgb, var(--table-tint, var(--accent)) 12%, transparent); border-top: 1px solid color-mix(in srgb, var(--table-tint, var(--accent)) 12%, transparent);
} }
@@ -797,17 +925,21 @@ figure:has(img[width]) {
The rules below re-anchor the bleed for the layouts where the article The rules below re-anchor the bleed for the layouts where the article
is not viewport-centered; each just overrides width/margin-inline, and is not viewport-centered; each just overrides width/margin-inline, and
later rules win at equal specificity. */ later rules win at equal specificity. The analytics dashboard uses the
figure:has(.wide) { same breakout directly on its container (div.wide — it is the page's
whole content, not a figure), and code blocks via a trailing {.wide}
line (fence block attrs land on <pre> itself). */
figure:has(.wide),
div.wide,
pre.wide {
width: 100vw; width: 100vw;
max-width: none; max-width: none;
margin-inline: calc(50% - 50vw); margin-inline: calc(50% - 50vw);
/* Follow the docked editor's margin-left transition smoothly. */
transition: margin-left 0.25s ease;
} }
/* Full bleed means edge to edge — no rounded corners. */ /* Full bleed means edge to edge — no rounded corners. */
figure:has(.wide) img { figure:has(.wide) img,
pre.wide {
border-radius: 0; border-radius: 0;
} }
@@ -815,22 +947,25 @@ figure:has(.wide) img {
gutter), so the bleed anchors at the left gutter — the 1fr share of the gutter), so the bleed anchors at the left gutter — the 1fr share of the
1fr + 4fr grid, i.e. 20vw — plus main's padding, and spans on to the 1fr + 4fr grid, i.e. 20vw — plus main's padding, and spans on to the
right viewport edge. */ right viewport edge. */
body:has(.multicol) figure:has(.wide) { body:has(.multicol) figure:has(.wide),
body:has(.multicol) pre.wide {
margin-inline: calc(-20vw - 1.25rem) 0; margin-inline: calc(-20vw - 1.25rem) 0;
} }
/* Editing: shrink the bleed to the space right of the docked editor. */ /* Editing: shrink the bleed to the space right of the docked editor. The
body.editing figure:has(.wide) { window keeps its overlay scrollbars while editing, so — unlike a classic
scrollbar — they take no layout space and the vw math stays exact. */
body.editing figure:has(.wide),
body.editing pre.wide {
width: calc(100vw - var(--editor-w)); width: calc(100vw - var(--editor-w));
margin-inline: calc(50% - (100vw - var(--editor-w)) / 2); margin-inline: calc(50% - (100vw - var(--editor-w)) / 2);
} }
/* Editing + multicol: the left gutter is 1/5 of the space right of the /* Editing + multicol: the left gutter is 1/5 of the space right of the
editor, and the bleed also crosses #content's 1rem editing gap plus editor, and the bleed also crosses main's 1.25rem left padding. */
main's 1.25rem left padding (the gutter shrink from the padding roughly body.editing:has(.multicol) figure:has(.wide),
cancels the rounding): 2rem in total. */ body.editing:has(.multicol) pre.wide {
body.editing:has(.multicol) figure:has(.wide) { margin-inline: calc((100vw - var(--editor-w)) / -5 - 1.25rem) 0;
margin-inline: calc((100vw - var(--editor-w)) / -5 - 2rem) 0;
} }
/* Narrow windows with a sidebar: below 102rem the symmetric gutters can no /* Narrow windows with a sidebar: below 102rem the symmetric gutters can no
@@ -841,7 +976,8 @@ body.editing:has(.multicol) figure:has(.wide) {
entirely on pages without sub-navigation, and excluded while editing, entirely on pages without sub-navigation, and excluded while editing,
where the editing rules above apply instead. */ where the editing rules above apply instead. */
@media (max-width: 102rem) { @media (max-width: 102rem) {
body:has(#sidebar):not(.editing) figure:has(.wide) { body:has(#sidebar):not(.editing) figure:has(.wide),
body:has(#sidebar):not(.editing) pre.wide {
margin-inline: -13.25rem 0; margin-inline: -13.25rem 0;
} }
} }
@@ -904,6 +1040,42 @@ article h2 {
(explicit img widths still shrink-wrap), while .wide keeps its full (explicit img widths still shrink-wrap), while .wide keeps its full
viewport bleed. */ viewport bleed. */
@media (max-width: 48rem) { @media (max-width: 48rem) {
/* Nav type shrinks fluidly as space runs out. The nav font-size is
em-based both in base and in every theme override, so scaling the
banner's font-size (nothing else in the banner is em-sized — brand and
gaps use rem) reaches the nav through all themes with a single rule.
2.6vw crosses 1rem at ≈38.5rem, so only genuinely narrow viewports
shrink. */
#banner {
font-size: clamp(0.65rem, 2.6vw, 1rem);
}
/* Tighter margins/padding/gaps: the 1.25rem side gutter is wasted space
on a phone. */
#brand {
margin-inline: 0.6rem;
}
#nav {
padding: 0.25rem 0.6rem;
}
#nav ul {
gap: 0.15rem 0.9rem;
}
/* The editor panel takes over the entire viewport: no space left for
the banner or the page content (main.js pins its top to 0 at these
widths). */
body.editing #content {
margin-left: 0;
}
.editor-host {
width: 100vw;
}
#content { #content {
display: flex; display: flex;
flex-direction: column; flex-direction: column;
@@ -915,13 +1087,19 @@ article h2 {
max-height: none; max-height: none;
overflow-y: visible; overflow-y: visible;
border-radius: 0; border-radius: 0;
padding: 0.5rem 1rem; padding: 0.4rem 0.8rem;
/* Smaller type: the horizontal link strip fits roughly a third more
items per line. The nested-list gaps below are em-based and shrink
along. */
font-size: 0.8rem;
} }
#sidebar ul { /* Only the main level becomes a horizontal wrapping strip; submenus stay
vertical blocks attached under their parent item. */
#sidebar>ul {
flex-direction: row; flex-direction: row;
flex-wrap: wrap; flex-wrap: wrap;
gap: 0.5rem 1.2rem; gap: 0.3rem 0.75rem;
} }
figure:has(.right), figure:has(.right),
@@ -996,11 +1174,11 @@ footer {
padding: 0; padding: 0;
} }
/* Rotating-cube page transition, adapted from termotohtori.fi. /* Rotating-cube page transition. */
FRAGILE: do not tweak; the view-transition pseudo-tree is picky. */
::view-transition { ::view-transition {
perspective: 1000px; perspective: 1000px;
inset: 0; inset: 0;
background: color-mix(var(--bg) 50%, black 50%);
} }
::view-transition-group(root), ::view-transition-group(root),
+67 -5
View File
@@ -22,6 +22,57 @@ let host = null
let app = null let app = null
let savedTitle = null let savedTitle = null
let visible = false let visible = false
let slideAnimation = null
const SLIDE_MS = 250 // keep in sync with the panel slide in pagerite.css
// The layout switches instantly when .editing toggles — no margin/width
// transitions anywhere, so viewport resizes (and the vw-based .wide bleed)
// always stay instant. The visible slide is a compositor-only FLIP
// transform on #content, running in sync with the panel's own slide
// (editor-slide-in / .closing in pagerite.css): both move by --editor-w
// over the same duration and easing, so .wide's left edge tracks the
// panel's right edge exactly throughout.
function setEditingClass(enable) {
const content = document.getElementById('content')
const before = content.getBoundingClientRect().left
document.body.classList.toggle('editing', enable)
const delta = before - content.getBoundingClientRect().left
slideAnimation?.cancel()
if (delta) {
slideAnimation = content.animate(
{ transform: [`translateX(${delta}px)`, 'translateX(0)'] },
{ duration: SLIDE_MS, easing: 'ease' }
)
}
}
// The panel is fixed to the viewport's left edge (pagerite.css) but tracks
// the page: its top is the banner's bottom edge while the banner is visible
// (= #content's top edge), and the viewport top once the banner has
// scrolled away. The window keeps scrolling normally while editing.
// Below 48rem the panel covers the entire viewport (pagerite.css), so its
// top stays 0 regardless of the banner.
const narrow = matchMedia('(max-width: 48rem)')
function trackPanelTop() {
const content = document.getElementById('content')
if (host && content) {
host.style.top = narrow.matches
? '0px'
: `${Math.max(0, content.getBoundingClientRect().top)}px`
}
}
function startTrackingPanel() {
trackPanelTop()
addEventListener('scroll', trackPanelTop, { passive: true })
addEventListener('resize', trackPanelTop)
}
function stopTrackingPanel() {
removeEventListener('scroll', trackPanelTop)
removeEventListener('resize', trackPanelTop)
}
export function openEditor(path, { mode = 'page' } = {}) { export function openEditor(path, { mode = 'page' } = {}) {
if (app) { if (app) {
@@ -36,9 +87,13 @@ export function openEditor(path, { mode = 'page' } = {}) {
savedTitle = document.title savedTitle = document.title
host = document.createElement('div') host = document.createElement('div')
host.className = 'editor-host' host.className = 'editor-host'
// Docked inside #content: below the banner, next to the article only. // Appended to <body>, not #content: the open/close slide transforms
document.getElementById('content').prepend(host) // #content (setEditingClass), and a transformed element becomes the
document.body.classList.add('editing') // containing block for fixed-position descendants — the panel would be
// dragged along with the content instead of sliding on its own.
document.body.append(host)
setEditingClass(true)
startTrackingPanel()
// Which tab is active; pagerite.js uses this to decide whether a pen click // Which tab is active; pagerite.js uses this to decide whether a pen click
// closes the panel or switches tabs. // closes the panel or switches tabs.
document.body.dataset.editorMode = mode document.body.dataset.editorMode = mode
@@ -64,7 +119,8 @@ function showEditor() {
savedTitle = document.title savedTitle = document.title
host.style.display = '' host.style.display = ''
host.firstElementChild?.classList.remove('closing') host.firstElementChild?.classList.remove('closing')
document.body.classList.add('editing') setEditingClass(true)
startTrackingPanel()
visible = true visible = true
dispatchEvent(new CustomEvent('pagerite:editor-shown')) dispatchEvent(new CustomEvent('pagerite:editor-shown'))
} }
@@ -72,7 +128,8 @@ function showEditor() {
export function closeEditor() { export function closeEditor() {
if (!visible) return if (!visible) return
visible = false visible = false
document.body.classList.remove('editing') stopTrackingPanel()
setEditingClass(false)
// dataset.editorMode is kept while hidden: the tabs use it to tell whether // dataset.editorMode is kept while hidden: the tabs use it to tell whether
// a pagerite:editor-shown event targets them. // a pagerite:editor-shown event targets them.
// Slide the panel out in sync with the page shifting back, then hide it. // Slide the panel out in sync with the page shifting back, then hide it.
@@ -80,6 +137,9 @@ export function closeEditor() {
const h = host const h = host
setTimeout(() => { h.style.display = 'none' }, 250) setTimeout(() => { h.style.display = 'none' }, 250)
dispatchEvent(new CustomEvent('pagerite:editor-hidden')) dispatchEvent(new CustomEvent('pagerite:editor-hidden'))
// The editor may have dropped the prefetch cache; warm it again for the
// now-final page so navigation stays instant.
dispatchEvent(new CustomEvent('pagerite:preload-pages'))
// Restore the server-rendered title for the current URL. Re-fetching makes // Restore the server-rendered title for the current URL. Re-fetching makes
// sure a brand change in the site editor or an in-place navigation leaves // sure a brand change in the site editor or an in-place navigation leaves
// the correct public title behind. // the correct public title behind.
@@ -96,3 +156,5 @@ export function closeEditor() {
if (!visible && restoreTitle != null) document.title = restoreTitle if (!visible && restoreTitle != null) document.title = restoreTitle
}) })
} }
+461 -82
View File
@@ -53,8 +53,22 @@ import "overlayscrollbars/overlayscrollbars.css";
// pageshow handler below re-probes auth to refresh the pens). // pageshow handler below re-probes auth to refresh the pens).
let ssoAvailable = false; let ssoAvailable = false;
let isAdmin = false; let isAdmin = false;
let authReady = false;
let editorMeta = null; let editorMeta = null;
// Asset URLs for the on-demand bundles. Dev renders them as
// pagerite:* meta tags (Vite dev-server URLs); production inlines all
// page assets and carries the on-demand URLs in a JSON script instead.
const assets = (() => {
const el = document.getElementById("pagerite-assets");
if (el) return JSON.parse(el.textContent);
const map = {};
for (const m of document.querySelectorAll('meta[name^="pagerite:"]')) {
map[m.name] = m.content;
}
return map;
})();
function makePen(mode) { function makePen(mode) {
const btn = document.createElement("button"); const btn = document.createElement("button");
btn.type = "button"; btn.type = "button";
@@ -93,32 +107,55 @@ import "overlayscrollbars/overlayscrollbars.css";
return a; return a;
} }
function removePens() {
document.querySelectorAll(".editor-pens, #main article button.edit-link")
.forEach((el) => el.remove());
}
function renderAuthUi() { function renderAuthUi() {
// Always start from a clean slate: if the auth probe is still running we
// must not show any admin UI, and if it came back negative we must drop
// pens that may have been injected while the browser cache made us look
// authenticated.
removePens();
if (!authReady) return;
// Editing is open for admins and, as a dev/no-proxy fallback, when no // Editing is open for admins and, as a dev/no-proxy fallback, when no
// Paskia SSO is detected at all. // Paskia SSO is detected at all.
const canEdit = isAdmin || !ssoAvailable; const canEdit = isAdmin || !ssoAvailable;
// The analytics page is a read-only dashboard: editing pens and the side
// panel do not apply there. Login/logout links are still useful.
const onAnalytics = currentPath === "/_a";
const banner = document.getElementById("page-banner"); const banner = document.getElementById("page-banner");
if (banner) { if (banner) {
const old = banner.parentElement.querySelector(".editor-pens");
if (old) old.remove();
const pens = document.createElement("div"); const pens = document.createElement("div");
pens.className = "editor-pens"; pens.className = "editor-pens";
if (canEdit) { if (canEdit && !onAnalytics) {
pens.append(makePen("banner")); pens.append(makePen("banner"));
pens.append(makePen("site")); pens.append(makePen("site"));
// Analytics viewer is now a normal page at /_a.
const a = document.createElement("a");
a.className = "edit-link analytics-link";
a.href = "/_a";
a.title = "analytics";
a.textContent = "📊";
pens.append(a);
} }
if (ssoAvailable) pens.append(makeAuthLink(isAdmin)); if (ssoAvailable) pens.append(makeAuthLink(isAdmin));
banner.after(pens); banner.after(pens);
} }
if (canEdit) injectPagePen(); if (canEdit && !onAnalytics) injectPagePen();
} }
async function setupAuth() { async function setupAuth() {
const src = document.querySelector('meta[name="pagerite:editor-src"]')?.content; authReady = false;
if (!src) return; renderAuthUi();
const src = assets["pagerite:editor-src"];
if (!src) { authReady = true; renderAuthUi(); pingEntryOnce(); return; }
editorMeta = { editorMeta = {
src, src,
css: document.querySelector('meta[name="pagerite:editor-css"]')?.content, css: assets["pagerite:editor-css"],
}; };
// Detect whether Paskia SSO is available on this site. // Detect whether Paskia SSO is available on this site.
@@ -137,7 +174,30 @@ import "overlayscrollbars/overlayscrollbars.css";
// No auth proxy / dev. // No auth proxy / dev.
} }
if (isAdmin) {
// Teach the backend the site's public origin (used for absolute
// social/canonical URLs): unlike request headers, location.origin
// reflects the real scheme and host even behind reverse proxies.
fetch("/_api/site-url", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ url: location.origin }),
}).catch(() => {});
// Warm the cache with the editor bundle: the hashed asset is
// immutable, so preloading costs nothing and the pens then open
// instantly. The analytics page has no editor.
if (currentPath !== "/_a" && !import.meta.env.DEV) {
const preload = document.createElement("link");
preload.rel = "modulepreload";
preload.href = src;
document.head.append(preload);
}
}
authReady = true;
renderAuthUi(); renderAuthUi();
placeEditPen();
pingEntryOnce();
} }
// Returning to the page via history back/forward may restore a cached // Returning to the page via history back/forward may restore a cached
@@ -175,9 +235,26 @@ import "overlayscrollbars/overlayscrollbars.css";
btn.textContent = "copy"; btn.textContent = "copy";
btn.addEventListener("click", async () => { btn.addEventListener("click", async () => {
const code = pre.querySelector("code"); const code = pre.querySelector("code");
await navigator.clipboard.writeText( const text = (code || pre).textContent.replace(/\n$/, "");
(code || pre).textContent.replace(/\n$/, ""), // navigator.clipboard exists only in secure contexts (https or
); // localhost); viewing over plain http needs the textarea fallback.
try {
if (navigator.clipboard) {
await navigator.clipboard.writeText(text);
} else {
const ta = document.createElement("textarea");
ta.value = text;
ta.style.cssText = "position:fixed;opacity:0";
document.body.append(ta);
ta.select();
document.execCommand("copy");
ta.remove();
}
} catch {
btn.textContent = "failed";
setTimeout(() => (btn.textContent = "copy"), 1500);
return;
}
btn.textContent = "copied"; btn.textContent = "copied";
btn.classList.add("copied"); btn.classList.add("copied");
setTimeout(() => { setTimeout(() => {
@@ -209,8 +286,66 @@ import "overlayscrollbars/overlayscrollbars.css";
// created page has no pen for commitPending's handover click. // created page has no pen for commitPending's handover click.
renderAuthUi(); renderAuthUi();
placeEditPen(); placeEditPen();
// Preview swaps also wipe the .colseg wrappers (the server render has
// none), which would drop the multi-column layout until a full reload;
// re-split so columns survive both live editing and closing the editor.
const main = document.getElementById("main");
if (main) applyMulticol(main);
}); });
// Multi-column layout only when there is enough text to justify it.
// Split the body into columned segments: h1s, h2s and wide elements are
// full-width separators and never go inside columns.
function applyMulticol(main) {
const article = main.querySelector("article");
if (!article) return;
const body = article.querySelector(".body");
// Code blocks don't read as flowing text and are often generated
// filler; exclude them when measuring whether the text justifies
// columns.
const textLen = (el) => {
let n = el.textContent.trim().length;
for (const pre of el.querySelectorAll("pre")) n -= pre.textContent.length;
return n;
};
article.classList.toggle(
"multicol",
!!body && textLen(body) > 1800,
);
if (body && article.classList.contains("multicol")
&& !body.querySelector(".colseg")) {
// h1s, h2s and wide elements (a {.wide} block or anything holding
// one, e.g. a figure with a wide image) are full-width separators
const isSeparator = (el) =>
el.tagName === "H1" || el.tagName === "H2"
|| el.classList.contains("wide")
|| el.querySelector(".wide") !== null;
let seg = null;
for (const el of [...body.children]) {
if (isSeparator(el)) {
seg = null;
body.append(el);
} else {
if (!seg) {
seg = document.createElement("div");
seg.className = "colseg";
body.append(seg);
}
seg.append(el);
}
}
// Columns are per section: only segments with enough text get them,
// so a short ingress or a brief section stays single-column. A
// .nocols container (::: nocols) opts its whole section out.
for (const s of body.querySelectorAll(".colseg")) {
s.classList.toggle(
"cols",
s.querySelector(".nocols") === null && textLen(s) > 600,
);
}
}
}
function applyEffects() { function applyEffects() {
(window.requestIdleCallback || setTimeout)(preload); (window.requestIdleCallback || setTimeout)(preload);
const main = document.getElementById("main"); const main = document.getElementById("main");
@@ -219,44 +354,8 @@ import "overlayscrollbars/overlayscrollbars.css";
// buttons; re-add whichever auth UI is appropriate for this session. // buttons; re-add whichever auth UI is appropriate for this session.
renderAuthUi(); renderAuthUi();
placeEditPen(); placeEditPen();
// Multi-column layout only when there is enough text to justify it. fitNav();
// Split the body into columned segments: h1s, h2s and wide figures are applyMulticol(main);
// full-width separators and never go inside columns.
const article = main.querySelector("article");
if (article) {
const body = article.querySelector(".body");
article.classList.toggle(
"multicol",
!!body && body.textContent.trim().length > 1800,
);
if (body && article.classList.contains("multicol")
&& !body.querySelector(".colseg")) {
// h1s, h2s and anything holding a wide image are full-width
// separators
const isSeparator = (el) =>
el.tagName === "H1" || el.tagName === "H2"
|| el.querySelector("img.wide") !== null;
let seg = null;
for (const el of [...body.children]) {
if (isSeparator(el)) {
seg = null;
body.append(el);
} else {
if (!seg) {
seg = document.createElement("div");
seg.className = "colseg";
body.append(seg);
}
seg.append(el);
}
}
// Columns are per section: only segments with enough text get them,
// so a short ingress or a brief section stays single-column.
for (const s of body.querySelectorAll(".colseg")) {
s.classList.toggle("cols", s.textContent.trim().length > 600);
}
}
}
if (reduceMotion.matches) return; if (reduceMotion.matches) return;
for (const el of main.querySelectorAll( for (const el of main.querySelectorAll(
"h2, h3, figure, img, pre, blockquote, table, dl, .task-list-item", "h2, h3, figure, img, pre, blockquote, table, dl, .task-list-item",
@@ -274,14 +373,32 @@ import "overlayscrollbars/overlayscrollbars.css";
// internal link is fetched exactly once, and navigation is served from // internal link is fetched exactly once, and navigation is served from
// memory with no fetch at all. Editor re-renders (swapdoc.loadPlain) // memory with no fetch at all. Editor re-renders (swapdoc.loadPlain)
// announce their fresh copies via pagerite:page-fetched, keeping the // announce their fresh copies via pagerite:page-fetched, keeping the
// cache in sync after edits. // cache in sync after edits. The current page is NOT preloaded: we just
// received it as the document (re-fetching would be redundant, and
// browser heuristics may send it without if-none-match, defeating the
// conditional request); it enters the cache when navigated to.
const pageCache = new Map(); // pathname -> HTML text const pageCache = new Map(); // pathname -> HTML text
addEventListener("pagerite:page-fetched", (ev) => { addEventListener("pagerite:page-fetched", (ev) => {
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html); pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html);
}); });
// Editors mutate site-wide state (theme, structure, headings, banners),
// which can change the rendered HTML of every cached page. Drop the whole
// cache so stale prefetches are never served; the current page is re-fetched
// by the editor's own loadPlain and re-cached afterwards. Re-preloading is
// deferred until the editor panel closes to avoid hammering the server.
addEventListener("pagerite:drop-page-cache", () => {
pageCache.clear();
});
// When the editor panel closes, warm the cache again for the visible links
// on the (now final) page so subsequent navigation stays instant.
addEventListener("pagerite:preload-pages", () => {
preload();
});
function preload() { function preload() {
const urls = new Set([location.pathname]); const urls = new Set();
for (const a of document.querySelectorAll( for (const a of document.querySelectorAll(
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]', '#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
)) { )) {
@@ -289,7 +406,10 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
for (const url of urls) { for (const url of urls) {
if (pageCache.has(url)) continue; if (pageCache.has(url)) continue;
fetch(url) // x-pagerite-preload: idle cache warm-up, not a page view — the
// server excludes these GETs from analytics (the ping sent on actual
// navigation does the counting).
fetch(url, { headers: { "x-pagerite-preload": "1" } })
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html") .then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
? r.text() : "")) ? r.text() : ""))
.then((html) => { if (html) pageCache.set(url, html); }) .then((html) => { if (html) pageCache.set(url, html); })
@@ -311,31 +431,202 @@ import "overlayscrollbars/overlayscrollbars.css";
// themes for their own effects (e.g. the purple sun rising faster than // themes for their own effects (e.g. the purple sun rising faster than
// the drift). Event-driven only: perfectly still when the page is idle. // the drift). Event-driven only: perfectly still when the page is idle.
if (!reduceMotion.matches) { if (!reduceMotion.matches) {
let ticking = false; // rAF loop that eases the value toward the live scroll position. Reading
// scrollY every frame (rather than only on scroll events) also picks up
// the in-between positions of Chrome/macOS momentum scrolling, whose
// scroll events fire late and coarsely. The loop idles once settled.
let value = Math.min(scrollY * 0.1, 30);
let running = false;
const drift = () => { const drift = () => {
ticking = false; const target = Math.min(scrollY * 0.1, 30);
document.documentElement.style.setProperty( value += (target - value) * 0.12;
"--pry", `${Math.min(scrollY * 0.1, 30)}px`, if (Math.abs(target - value) < 0.05) {
); value = target;
running = false;
}
document.documentElement.style.setProperty("--pry", `${value}px`);
if (running) requestAnimationFrame(drift);
}; };
addEventListener("scroll", () => { addEventListener("scroll", () => {
if (!ticking) { if (!running) {
ticking = true; running = true;
requestAnimationFrame(drift); requestAnimationFrame(drift);
} }
}, { passive: true }); }, { passive: true });
} }
// --- Analytics pings ---------------------------------------------------
// Fire-and-forget POSTs to /_a with the fields as query parameters (a
// beacon can carry no body, and query args show in server logs next to
// the document GET they refer to): on the initial page load (starts the
// visit — the server counts nothing from the document GET alone), for
// internal fetch-navigations, for external https exits, and on window
// close. ``read`` is the active time (ms) spent on ``fr``.
// Reading time pauses after 1 minute of inactivity and resumes on the
// next mouse/touch/scroll/keyboard event.
// Excluded: back/forward (popstate never pings), everything while the
// editor is open (body.editing — admin noise, not visits), and
// navigations TO the analytics page (/_a — admin machinery, and the
// server rejects it as a ping target anyway). Navigations AWAY from /_a
// must ping: load() already fetched the target page without the preload
// header, and without the ping that GET would flush to the crawler list.
// Admins (when SSO is actually in use — with no auth proxy "admin" is
// everyone's state) ping normally but with hide=1: the server then
// records nothing and scrubs any session the same browser accumulated
// before logging in, so admins never show up as visits or crawlers.
// See docs/analytics.md.
// fetch wrapper: every key of ``params`` becomes a query arg on /_a
// (falsy values are omitted). Admins get hide=1. ``beacon`` uses
// sendBeacon when available, for unload-time pings.
function pingFetch(params, { beacon = false } = {}) {
const query = new URLSearchParams();
if (ssoAvailable && isAdmin) params = { ...params, hide: 1 };
for (const [key, value] of Object.entries(params)) {
if (value) query.set(key, value);
}
const url = `/_a?${query}`;
try {
if (beacon && navigator.sendBeacon) {
navigator.sendBeacon(url);
} else {
fetch(url, { method: "POST", keepalive: true });
}
} catch { /* analytics must never break navigation */ }
}
function ping({ to, fr = currentPath, read = 0, beacon = false } = {}) {
if (document.body.classList.contains("editing")) return;
if (to === "/_a") return;
pingFetch({ fr, to, read: Math.round(read / 1000) }, { beacon });
}
// Active reading time for the current page. The clock stops after 1 minute
// without activity and restarts on the next mouse/touch/scroll/keyboard
// event.
const INACTIVE_MS = 60_000;
let readStart = performance.now();
let readElapsed = 0;
let reading = true;
let readInactivityTimer = null;
let closePingedFor = null;
function markReadActivity() {
if (!reading) {
reading = true;
readStart = performance.now();
}
clearTimeout(readInactivityTimer);
readInactivityTimer = setTimeout(() => {
if (reading) {
readElapsed += performance.now() - readStart;
reading = false;
}
}, INACTIVE_MS);
}
function takeReadTime() {
if (reading) {
readElapsed += performance.now() - readStart;
readStart = performance.now();
}
const ms = Math.max(0, Math.round(readElapsed));
readElapsed = 0;
return ms;
}
function resetReadTime() {
readElapsed = 0;
reading = true;
readStart = performance.now();
clearTimeout(readInactivityTimer);
}
function sendClosePing() {
if (closePingedFor === currentPath) return;
closePingedFor = currentPath;
const read = takeReadTime();
if (Math.round(read / 1000) <= 0) return;
ping({ read, beacon: true });
}
for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) {
addEventListener(ev, markReadActivity, { passive: true });
}
addEventListener("pagehide", sendClosePing);
// The initial page load pings too — it is what starts the visit and
// counts the entry page view (the document GET alone records nothing).
// It carries only ``to``: the server attributes the entry to the referer
// it saw on the document GET (unavailable to JS once loaded), and an
// ``fr`` equal to ``to`` would log a bogus self-transition when a
// session already exists (e.g. a second tab).
// Sent once per load, after the auth probes so the admin gate applies;
// the pageshow re-probe must not ping again. Reloads are not visits:
// pinging them would double-count the view and log a self-transition.
let entryPinged = false;
function pingEntryOnce() {
if (entryPinged) return;
entryPinged = true;
const nav = performance.getEntriesByType?.("navigation")[0];
if (nav ? nav.type === "reload" : performance.navigation?.type === 1) return;
ping({ to: currentPath, fr: "" });
}
// --- Analytics page mount/unmount --------------------------------------
// The analytics page is a normal page whose body is rendered by the server
// but whose content is a Vue app. In dev the entry module is imported from
// the Vite dev server on demand; in production it is inlined into the /_a
// page as script#pagerite-js-analytics, which a fetch-navigation swap does
// not execute — re-create the element so the fresh module auto-mounts on
// #analytics-app (see analytics-main.js). The module exposes its unmount
// as window.__pageriteAnalyticsUnmount.
function teardownAnalytics() {
// Remove even the server-rendered script element so a later return to
// /_a re-mounts from a fresh copy (the module has torn itself down).
document.getElementById("pagerite-js-analytics")?.remove();
window.__pageriteAnalyticsUnmount?.();
window.__pageriteAnalyticsUnmount = null;
}
async function mountAnalytics(doc) {
if (!doc.getElementById("analytics-app")) return;
// Already mounted: on a full /_a load the inline script has run.
if (document.getElementById("pagerite-js-analytics")) return;
const inline = doc.getElementById("pagerite-js-analytics");
if (inline) {
const s = document.createElement("script");
for (const a of inline.attributes) s.setAttribute(a.name, a.value);
s.textContent = inline.textContent;
document.body.append(s);
return;
}
try {
// Dev: the cached module auto-mounts only on its first evaluation,
// so call mount() explicitly for repeat visits (it no-ops when the
// app is already up).
const mod = await import(/* @vite-ignore */ assets["pagerite:analytics-src"]);
const container = document.getElementById("analytics-app");
if (container) mod.mount(container);
} catch (e) {
console.error("analytics mount failed:", e);
}
}
// --- Fetch navigation ------------------------------------------------ // --- Fetch navigation ------------------------------------------------
async function load(url, push = true, back = false) { async function load(url, push = true, back = false) {
// Navigating with the editor open closes it; unsaved edits are lost // Navigating with the editor open closes it; unsaved edits are lost
// (the region swap discards the previewed changes anyway). // (the region swap discards the previewed changes anyway). Cache must be
if (document.body.classList.contains("editing")) { // bypassed for this navigation because the editor may have invalidated
// the prefetched copies of other pages.
const editing = document.body.classList.contains("editing");
if (editing) {
editorModule?.then((m) => m.closeEditor()); editorModule?.then((m) => m.closeEditor());
} }
teardownAnalytics();
let doc; let doc;
let finalUrl = url; let finalUrl = url;
const cached = pageCache.get(new URL(url, location.href).pathname); const cached = !editing && pageCache.get(new URL(url, location.href).pathname);
if (cached) { if (cached) {
doc = new DOMParser().parseFromString(cached, "text/html"); doc = new DOMParser().parseFromString(cached, "text/html");
} else { } else {
@@ -345,15 +636,19 @@ import "overlayscrollbars/overlayscrollbars.css";
if (!res.ok || !type.includes("text/html")) throw new Error("not a page"); if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
// Reflect any redirect the server issued. // Reflect any redirect the server issued.
if (res.redirected) finalUrl = res.url; if (res.redirected) finalUrl = res.url;
doc = new DOMParser().parseFromString(await res.text(), "text/html"); const html = await res.text();
// Populate the cache too, so returning here (back/forward, or a
// self-link in the nav) is served from memory.
pageCache.set(new URL(finalUrl, location.href).pathname, html);
doc = new DOMParser().parseFromString(html, "text/html");
} catch { } catch {
location.href = url; // fall back to a normal navigation location.href = url; // fall back to a normal navigation
return; return false;
} }
} }
if (REGIONS.some((id) => !doc.getElementById(id))) { if (REGIONS.some((id) => !doc.getElementById(id))) {
location.href = url; location.href = url;
return; return false;
} }
const doit = () => { const doit = () => {
for (const id of REGIONS) { for (const id of REGIONS) {
@@ -372,25 +667,48 @@ import "overlayscrollbars/overlayscrollbars.css";
} else if (oldSidebar) { } else if (oldSidebar) {
oldSidebar.remove(); oldSidebar.remove();
} }
// Site-wide custom CSS lives in <head id="pagerite-user"> and must be // Stylesheets live in <head> with stable ids — links in dev, inline
// kept in sync across fetch-navigations. It is kept last in <head>: // <style> elements in production — and must follow the swap: the
// in dev Vite injects the base stylesheet after the server-rendered // analytics sheet exists on /_a only, and theme/banner/custom CSS
// tag, and equal-specificity :root rules are decided by order. // may have changed since this page was loaded. Diff by id, keeping
const oldUserStyle = document.getElementById("pagerite-user"); // the fresh document's order; unchanged sheets keep their elements
const newUserStyle = doc.getElementById("pagerite-user"); // so their @keyframes are never torn down. Editor-injected sheets
if (oldUserStyle && newUserStyle) { // (data-pagerite, no id) and Vite's dev styles (no id) are left
oldUserStyle.textContent = newUserStyle.textContent; // alone. Mirrors the head sync in swapdoc.js.
document.head.appendChild(oldUserStyle); const sel = 'link[rel="stylesheet"][id], style[id]';
} else if (newUserStyle) { const fresh = [...doc.head.querySelectorAll(sel)];
document.head.appendChild(document.importNode(newUserStyle, true)); const freshIds = new Set(fresh.map((el) => el.id));
} else if (oldUserStyle) { for (const el of [...document.head.querySelectorAll(sel)]) {
oldUserStyle.remove(); if (!freshIds.has(el.id)) el.remove();
} }
let anchor = null;
for (const el of fresh) {
const cur = document.getElementById(el.id);
if (cur && cur.outerHTML === el.outerHTML) {
anchor = cur;
continue;
}
const imported = document.importNode(el, true);
if (cur) cur.replaceWith(imported);
else if (anchor) anchor.after(imported);
else {
const base = document.getElementById("pagerite-base");
if (base) base.after(imported);
else document.head.append(imported);
}
anchor = imported;
}
// Custom CSS must stay last: equal-specificity :root rules (font
// variables) are decided by order, and in dev Vite injects the base
// stylesheet after the server-rendered tag.
const userStyle = document.getElementById("pagerite-user");
if (userStyle) document.head.appendChild(userStyle);
document.title = doc.title; document.title = doc.title;
// Banners may contain scripts (canvas etc.), content pages may too. // Banners may contain scripts (canvas etc.), content pages may too.
runScripts(document.getElementById("page-banner")); runScripts(document.getElementById("page-banner"));
runScripts(document.getElementById("main")); runScripts(document.getElementById("main"));
applyEffects(); applyEffects();
mountAnalytics(doc);
}; };
// Rotating cube page transition (see the FRAGILE block in pagerite.css); // Rotating cube page transition (see the FRAGILE block in pagerite.css);
// mirrored when navigating back through history. Navigation within the // mirrored when navigating back through history. Navigation within the
@@ -410,6 +728,7 @@ import "overlayscrollbars/overlayscrollbars.css";
currentPath = new URL(finalUrl, location.href).pathname; currentPath = new URL(finalUrl, location.href).pathname;
if (push) history.pushState(null, "", finalUrl); if (push) history.pushState(null, "", finalUrl);
scrollTo(0, 0); scrollTo(0, 0);
return true;
} }
addEventListener("click", (ev) => { addEventListener("click", (ev) => {
@@ -447,16 +766,37 @@ import "overlayscrollbars/overlayscrollbars.css";
const a = ev.target.closest("a[href]"); const a = ev.target.closest("a[href]");
if (!a || a.target || a.hasAttribute("download")) return; if (!a || a.target || a.hasAttribute("download")) return;
const url = new URL(a.href, location.href); const url = new URL(a.href, location.href);
if (url.origin !== location.origin) return; if (url.origin !== location.origin) {
// External link: the browser navigates; record the full https URL so
// different links to the same domain stay distinct in analytics.
if (url.protocol === "https:") {
closePingedFor = currentPath;
ping({ to: url.href, read: takeReadTime() });
}
return;
}
// Same-page anchor links (footnotes etc.): let the browser handle them // Same-page anchor links (footnotes etc.): let the browser handle them
if (url.pathname === location.pathname && url.hash) return; if (url.pathname === location.pathname && url.hash) return;
// Machinery and auth endpoints are never fetch-navigated. // Machinery and auth endpoints are never fetch-navigated, except the
if (url.pathname.startsWith("/_") || url.pathname.startsWith("/auth")) return; // public analytics viewer page at /_a.
if ((url.pathname.startsWith("/_") && url.pathname !== "/_a")
|| url.pathname.startsWith("/auth")) return;
ev.preventDefault(); ev.preventDefault();
load(url); // Capture the source now: load() updates currentPath before pinging.
const from = currentPath;
load(url).then((ok) => {
if (!ok) return;
closePingedFor = null;
ping({ to: url.pathname, fr: from, read: takeReadTime() });
resetReadTime();
});
}); });
addEventListener("popstate", () => load(location.href, false, true)); addEventListener("popstate", () => {
// Hash-only history entries are not navigations.
if (location.pathname === currentPath) return;
load(location.href, false, true);
});
// --- Task-list checkboxes ------------------------------------------------ // --- Task-list checkboxes ------------------------------------------------
// Checkboxes in the rendered article are live: toggling them edits the // Checkboxes in the rendered article are live: toggling them edits the
@@ -520,6 +860,45 @@ import "overlayscrollbars/overlayscrollbars.css";
fit(); fit();
} }
// --- Nav condense-to-fit -------------------------------------------------
// The top nav stays on one row even on too-narrow screens: first the link
// gaps shrink, then the nav's side padding, and only in extreme cases the
// font size. #nav is replaced on fetch-navigation swaps, so this re-runs
// from applyEffects (fresh elements each time); CSS keeps flex-wrap: wrap
// as the no-JS fallback.
function fitNav() {
const nav = document.getElementById("nav");
const ul = nav?.querySelector("ul");
if (!ul) return;
// Restore the themed defaults before measuring.
nav.style.fontSize = "";
nav.style.paddingInline = "";
ul.style.columnGap = "";
ul.style.flexWrap = "nowrap";
const overflow = () => ul.scrollWidth - ul.clientWidth;
if (overflow() <= 0) return;
// 1) shrink the gaps between items (down to a fifth of the themed gap)
const gap = parseFloat(getComputedStyle(ul).columnGap) || 0;
const joints = Math.max(ul.children.length - 1, 1);
if (gap > 0) {
ul.style.columnGap = `${Math.max(0.2 * gap, gap - overflow() / joints)}px`;
}
// 2) shrink the nav's side padding (down to 0.4x)
if (overflow() > 0) {
const pad = parseFloat(getComputedStyle(nav).paddingInlineStart) || 0;
nav.style.paddingInline = `${Math.max(0.4 * pad, pad - overflow() / 2)}px`;
}
// 3) shrink the font to fit what remains
if (overflow() > 0) {
const fs = parseFloat(getComputedStyle(nav).fontSize);
nav.style.fontSize = `${fs * ul.clientWidth / ul.scrollWidth}px`;
}
}
addEventListener("resize", fitNav);
document.fonts?.ready.then(fitNav);
setupAuth(); setupAuth();
applyEffects(); applyEffects();
mountAnalytics(document);
})(); })();
+34 -19
View File
@@ -3,6 +3,14 @@
// Used by BannerEditor (banner design changes), SiteEditor (theme changes) // Used by BannerEditor (banner design changes), SiteEditor (theme changes)
// and StructureEditor (tree navigation). // and StructureEditor (tree navigation).
// Drop the public page runtime's in-memory prefetch cache. Editors call this
// whenever a site-wide or page change invalidates the cached HTML of other
// pages (theme, headings, structure, banner, etc.). The cache is rebuilt by
// re-preloading visible links once the editor panel closes.
export function dropPageCache() {
dispatchEvent(new CustomEvent('pagerite:drop-page-cache'))
}
export function runScripts(root) { export function runScripts(root) {
// Scripts injected via innerHTML do not execute; re-create them. // Scripts injected via innerHTML do not execute; re-create them.
if (!root) return if (!root) return
@@ -59,32 +67,39 @@ function swapRegions(doc) {
curUserStyle.remove() curUserStyle.remove()
} }
// Theme and other public stylesheets live in <head>, rendered with stable // Theme and other public stylesheets live in <head>, rendered with stable
// ids by the backend; sync them positionally so the custom CSS (rendered // ids by the backend (links in dev, inline <style> elements in prod);
// last) always keeps winning by order. Diff-based: unchanged sheets keep // sync them positionally so the custom CSS (rendered last) always keeps
// their elements, so their @keyframes are never torn down (re-creating // winning by order. Diff-based: unchanged sheets keep their elements, so
// keyframes would replay the editor's slide-in animation). // their @keyframes are never torn down (re-creating keyframes would
const freshLinks = [...doc.head.querySelectorAll('link[rel="stylesheet"]')] // replay the editor's slide-in animation).
const freshIds = new Set(freshLinks.map((l) => l.id)) const sel = 'link[rel="stylesheet"][id], style[id]'
for (const link of [...document.head.querySelectorAll('link[rel="stylesheet"]')]) { const freshEls = [...doc.head.querySelectorAll(sel)]
if (!link.dataset.pagerite && !freshIds.has(link.id)) link.remove() const freshIds = new Set(freshEls.map((el) => el.id))
for (const el of [...document.head.querySelectorAll(sel)]) {
if (!freshIds.has(el.id)) el.remove()
} }
// Insert missing sheets in the fresh document's order, each right after // Insert missing sheets in the fresh document's order, each right after
// its predecessor's element. The first sheet rendered is always the base // its predecessor's element. The first sheet rendered is always the base
// CSS, so its link doubles as the fallback anchor when nothing matched yet // CSS, so its element doubles as the fallback anchor when nothing matched
// (e.g. no theme was selected before and the position is otherwise lost). // yet (e.g. no theme was selected before and the position is otherwise
// lost).
let anchor = null let anchor = null
for (const link of freshLinks) { for (const el of freshEls) {
const cur = link.id && document.getElementById(link.id) const cur = el.id && document.getElementById(el.id)
if (cur && cur.href === link.href) { if (cur && cur.outerHTML === el.outerHTML) {
anchor = cur anchor = cur
continue continue
} }
const el = document.importNode(link, true) const imported = document.importNode(el, true)
// Same id, new URL (theme switch): replace in place, keeping position. // Same id, new content (theme switch): replace in place, keeping position.
if (cur) cur.replaceWith(el) if (cur) cur.replaceWith(imported)
else if (anchor) anchor.after(el) else if (anchor) anchor.after(imported)
else document.getElementById('pagerite-base')?.after(el) ?? document.head.append(el) else {
anchor = el const base = document.getElementById('pagerite-base')
if (base) base.after(imported)
else document.head.append(imported)
}
anchor = imported
} }
// The editor keeps its own title while open; only inherit the server title // The editor keeps its own title while open; only inherit the server title
// when navigating outside the editor (e.g. fetch-navigation swaps). // when navigating outside the editor (e.g. fetch-navigation swaps).
+1 -1
View File
@@ -11,7 +11,7 @@
*/ */
export default function fastapiVue({ paths = ["/api"] } = {}) { export default function fastapiVue({ paths = ["/api"] } = {}) {
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:3200" const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:8210"
// Build proxy configuration for each path // Build proxy configuration for each path
const proxy = {} const proxy = {}
+6 -5
View File
@@ -7,15 +7,15 @@ import vueDevTools from 'vite-plugin-vue-devtools'
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200' const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200'
// Proxy content pages (/slug, /path/to/slug) to the FastAPI backend in dev. // Proxy everything except Vite's own dev-time paths and the backend machinery
// Excludes Vite internals (/@..., /src, /node_modules, /__...) and the // to the FastAPI backend in dev. /_api, /_f, /_themes and /_a are handled by
// backend's /_ prefix. /_api and /_f are handled by the fastapi-vue plugin. // the fastapi-vue plugin, and /@..., /src, /node_modules, /__... stay with Vite.
const CONTENT_PROXY = '^\\/(?!_|@|src|node_modules|__)(?:[^./?]+(?:\\/[^./?]+)*)?(?:\\?.*)?$' const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
// https://vite.dev/config/ // https://vite.dev/config/
export default defineConfig({ export default defineConfig({
plugins: [ plugins: [
fastapiVue({ paths: ["/_api", "/_f", "/_themes"] }), fastapiVue({ paths: ["/_api", "/_f", "/_themes", "/_a"] }),
vue(), vue(),
vueDevTools(), vueDevTools(),
], ],
@@ -39,6 +39,7 @@ export default defineConfig({
input: { input: {
main: fileURLToPath(new URL('./src/main.js', import.meta.url)), main: fileURLToPath(new URL('./src/main.js', import.meta.url)),
pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)), pagerite: fileURLToPath(new URL('./src/pagerite.js', import.meta.url)),
analytics: fileURLToPath(new URL('./src/analytics-main.js', import.meta.url)),
// Only the base CSS is built; theme/banner-design stylesheets live // Only the base CSS is built; theme/banner-design stylesheets live
// in pagerite/themes/{name}/ and are served by the backend as-is. // in pagerite/themes/{name}/ and are served by the backend as-is.
pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)), pagerite_base: fileURLToPath(new URL('./src/assets/pagerite.css', import.meta.url)),
+70 -2
View File
@@ -1,14 +1,74 @@
# auto-upgrade@fastapi-vue-setup - remove this if you modify this file
"""Command-line entry point for running the backend server.""" """Command-line entry point for running the backend server."""
import argparse import argparse
import gzip
import os import os
import sys
from datetime import date
from pathlib import Path
import httpx
from fastapi_vue import server from fastapi_vue import server
DEFAULT_PORT = 3100 DEFAULT_PORT = 8100
DEVMODE = os.getenv("PAGERITE_DEV") == "1" DEVMODE = os.getenv("PAGERITE_DEV") == "1"
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
_REPO_ROOT = Path(__file__).resolve().parent.parent
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
def _download_dbip() -> None:
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
today = date.today()
months = [f"{today:%Y-%m}"]
# The current month's file may not be published yet; fall back to last month.
prev = (today.replace(day=1) - date.resolution).replace(day=1)
months.append(f"{prev:%Y-%m}")
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download")
return
for month in months:
url = DBIP_URL.format(month=month)
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
print(f"pagerite: downloading {url}")
try:
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
if r.status_code == 404:
continue
r.raise_for_status()
with open(tmp, "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk)
except httpx.HTTPError as e:
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
tmp.unlink(missing_ok=True)
continue
# Verify it is actually gzip data before installing it.
try:
with gzip.open(tmp, "rb") as f:
f.read(1)
except OSError:
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr)
tmp.unlink(missing_ok=True)
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
print(f"pagerite: DB-IP database updated to {target.name}")
return
print("pagerite: could not download a DB-IP database", file=sys.stderr)
def main() -> None: def main() -> None:
"""Run the backend server with optional arguments.""" """Run the backend server with optional arguments."""
@@ -19,12 +79,20 @@ def main() -> None:
action="append", action="append",
help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."), help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."),
) )
parser.add_argument(
"--dbip",
action="store_true",
help="Download/update the DB-IP city lite database before starting.",
)
args = parser.parse_args() args = parser.parse_args()
if args.dbip:
_download_dbip()
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {} dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
server.run( server.run(
"pagerite.app:app", "pagerite.app:app",
listen=args.listen, listen=args.listen,
default_port=DEFAULT_PORT, default_port=DEFAULT_PORT,
server_header=False,
**dev, **dev,
) )
+839
View File
@@ -0,0 +1,839 @@
"""Server-side visit analytics (collection only; see docs/analytics.md).
Events come from navigation pings POSTed to /_a by pagerite.js: the first
ping on page load starts a visit, later pings extend it, and pings with no
known session start a fresh one (missing data, not dropped). The document
GET handler stashes the entry referer (external https origin) and any
utm_* query parameters in in-memory IP tables, consumed when the ping
starts the visit; nothing is counted without a ping (plain bots that only
fetch documents end up in the crawler list). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their UA gives them
away (``_is_bot_ua``) and their pings are ignored, so they land in the
crawler list too. Idle-time link preloads from pagerite.js carry an
``x-pagerite-preload`` header and are not tracked at all — the ping sent
when the user actually navigates does the counting.
Admin clients ping with ``hide=1``: the client record is flagged ``hide``,
which covers everything that client ever did — visits and crawler hits
from before the login included. Aggregates (site visits, page views,
transitions) are not stored; they are computed at display time from the
visit records, excluding hidden clients, and hidden clients' visits,
crawler hits, abuse hits and metadata are left out of the viewer payload
entirely. Scanner telltale 404s
(dotpaths, *.php) classify the source IP as abuse; its hits — including
earlier crawler hits — are moved to the abuse list, which the viewer
groups by IP with full request paths. Client metadata (IP, UA, language,
country/city, host) is stored once per unique client hash and referenced
from visits, crawler hits and abuse hits. The session map is in-memory
only.
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
rewritten atomically on every recorded event.
"""
import ipaddress
import os
import re
import tempfile
from collections.abc import Callable
from contextlib import suppress
from datetime import UTC, datetime, timedelta
from pathlib import Path
from urllib.parse import parse_qs, urlparse
import blake3
import msgspec
from ua_parser import parse
def _compact_user_agent(ua: str) -> str:
"""Format a User-Agent string into a compact display form.
Returns the original UA when the parser cannot identify the browser/OS.
"""
if not ua or not ua.strip() or ua == "-":
return ""
r = parse(ua)
browser = r.user_agent.family if r.user_agent else None
ver = r.user_agent.major if r.user_agent else ""
os_name = r.os.family if r.os else None
dev = r.device.family if r.device else None
if browser in (None, "Other") and os_name in (None, "Other"):
return ua
if browser and browser != "Other":
browser = browser.split()[0]
else:
browser = ""
os_name = os_name if os_name and os_name != "Other" else ""
if dev in (None, "Other") or dev == browser:
dev = ""
parts = [f"{browser}/{ver}" if browser else "", os_name, dev]
return " ".join(p for p in parts if p).strip()
class Client(msgspec.Struct, omit_defaults=True):
"""Client metadata shared by visits, crawler hits and abuse hits.
Identified by a 6-byte blake3 hash of the IPv4 address or IPv6 /64
network, the full User-Agent string and the extracted language tag.
Country/city/host are filled in asynchronously after the first event.
"""
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
ip: str = ""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter country code from the DB-IP geoip lookup, or "".
country: str = ""
#: City name from the DB-IP geoip lookup, or "".
city: str = ""
#: Raw User-Agent header.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
#: True for admin clients (hide=1 ping): their visits, crawler hits and
#: abuse hits are recorded but excluded from all statistics and from
#: the viewer payload.
hide: bool = False
class Nav(msgspec.Struct, omit_defaults=True):
"""One navigation inside a visit: from ``fr`` to ``to``.
``to`` is an internal page path or an external https exit URL. Every
navigation is logged (repeats included), keyed by its timestamp in
``Visit.navs``, so display-time aggregates can count views and
transitions; ``Visit.trail`` keeps the first-seen order.
"""
fr: str
to: str
class TrailItem(msgspec.Struct, omit_defaults=True):
"""One first-seen target in a visit trail: a page or external exit URL.
``read`` accumulates active reading time (seconds) across the whole
visit; ``status`` is the most recent HTTP status seen for the target.
"""
to: str
#: Accumulated active reading time in seconds.
read: int = 0
#: Most recent HTTP status of the response (200 or 404).
status: int = 200
class Visit(msgspec.Struct, omit_defaults=True):
"""One visit: the initial-load data plus everything seen afterwards.
``trail`` holds the entry page and everything seen afterwards, keyed by
the timestamp of first sight (insertion order = first-seen order);
re-visiting an already seen target updates its item instead of
appending. Client metadata is held in ``Analytics.clients`` keyed by
``client``.
"""
start: datetime
entry: str
#: External https origin of the initial load, "" for direct visits.
referer: str = ""
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: First-seen targets keyed by their timestamp (entry included).
trail: dict[datetime, TrailItem] = {}
#: Every navigation ping (repeats included) keyed by its timestamp; the
#: aggregates are computed from this log at display time.
navs: dict[datetime, Nav] = {}
#: UTM query parameters from the landing URL, keyed by parameter name.
utm: dict[str, str] = {}
class CrawlerHit(msgspec.Struct, omit_defaults=True):
"""A document GET that was never followed by an analytics ping.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime
entry: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: External https origin of the initial load, "" for direct/none.
referer: str = ""
#: Raw query string of the landing URL (UTM tags can be parsed from it).
query: str = ""
#: HTTP status of the served response (200 or 404 for content pages).
status: int = 200
class AbuseHit(msgspec.Struct, omit_defaults=True):
"""A request from an IP classified as a scanner/abuser.
Unlike crawler hits the full request path (query string included) is
kept: the interesting part is exactly which paths were probed.
``flag`` marks the path that triggered classification; ``is_404``
distinguishes 404 responses from document GETs made by the abuser.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime
#: Full request path including the query string (e.g. "/.env?x=1").
path: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: True when this path triggered abuse classification (telltale path
#: or the 404 that crossed the threshold).
flag: bool = False
#: True for 404 responses; false for document GETs from the abuser.
is_404: bool = False
class Analytics(msgspec.Struct, omit_defaults=True):
"""Root of the analytics JSON file. Append-only by design: old data is
dropped by deleting list entries / bucket keys."""
visits: list[Visit] = []
#: Document GETs that never produced a ping, treated as crawler/bot hits.
crawlers: list[CrawlerHit] = []
#: Requests from abusive IPs (see AbuseHit), grouped by IP in the viewer.
abuse: list[AbuseHit] = []
#: Client metadata keyed by 6-byte blake3 hash.
clients: dict[bytes, Client] = {}
#: IPs classified as scanners/abusers (keys; values always True).
abuse_ips: dict[str, bool] = {}
class Display(msgspec.Struct, omit_defaults=True):
"""The viewer payload: visible data plus display-time aggregates.
Hidden clients are excluded everywhere: their visits, crawler hits,
abuse hits and metadata are dropped, and the aggregates are computed
from the visible visits only.
The aggregate shapes match what the viewer consumes: sparse 5-minute
buckets keyed by their floored ISO timestamp.
"""
visits: list[Visit] = []
crawlers: list[CrawlerHit] = []
abuse: list[AbuseHit] = []
clients: dict[bytes, Client] = {}
#: Page transitions per 5-minute bucket (sparse):
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
#: "(direct)" for initial loads, a page path for pings.
transitions: dict[str, dict[str, dict[str, int]]] = {}
#: Page views per 5-minute bucket: path -> bucket ISO -> count (sparse).
views: dict[str, dict[str, int]] = {}
#: New visits per 5-minute bucket: bucket ISO -> count (sparse).
site_visits: dict[str, int] = {}
def _bucket(now: datetime) -> str:
"""Start of the 5-minute interval containing ``now``, as ISO string."""
return now.replace(minute=now.minute // 5 * 5, second=0, microsecond=0).isoformat()
def _origin(url: str) -> str | None:
"""The origin part of an https URL (scheme://host[:port]), else None."""
try:
parsed = urlparse(url)
except ValueError:
return None
if parsed.scheme != "https" or not parsed.netloc:
return None
return f"https://{parsed.netloc}"
def _external_target(url: str) -> str | None:
"""A valid https URL (origin or full page), else None."""
try:
parsed = urlparse(url)
except ValueError:
return None
if parsed.scheme != "https" or not parsed.netloc:
return None
return url
_SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*")
def _internal_path(to: str) -> str | None:
"""A valid internal page path ("/" or slug segments), else None."""
path = to.split("?")[0].split("#")[0].strip("/")
if not path:
return "/"
if all(_SEGMENT.fullmatch(seg) for seg in path.split("/")):
return f"/{path}"
return None
def _parse_accept_language(value: str) -> tuple[str, str]:
"""First Accept-Language tag and the region/country subtag if present.
``en-US, fr;q=0.9`` -> ("en-us", "US"). Wildcards and missing regions
produce an empty country. The region is intentionally approximate:
it reflects the browser's language preference, not geo-location.
"""
if not value:
return "", ""
tag = value.split(",")[0].split(";")[0].strip()
if not tag or tag == "*":
return "", ""
lang = tag.lower()
country = ""
# Region subtags follow the initial language tag (en-US, zh-Hans-CN).
# A bare two-letter tag such as "fr" is a language code, not a region.
for part in reversed(tag.split("-")[1:]):
if len(part) == 2 and part.isalpha():
country = part.upper()
break
return lang, country
def _utm_tags(query: str) -> dict[str, str]:
"""UTM parameters from a query string, keeping only the first value."""
if not query:
return {}
parsed = parse_qs(query, keep_blank_values=True)
return {k: v[0] for k, v in parsed.items() if k.startswith("utm_")}
_CRAWLER_TIMEOUT = timedelta(seconds=10)
#: UAs of JS-running crawlers, which would register as visitors on their
#: ping. Anything calling itself a "bot" or "spider" matches; known crawlers
#: without those tokens (GoogleOther) are listed as extra alternates. No
#: source verification: a spoofed bot UA just lands in the crawler list, and
#: scanners that probe telltale paths are caught by the abuse rules anyway.
_BOT_UA = re.compile(r"bot|spider|googleother", re.IGNORECASE)
def _is_bot_ua(ua: str) -> bool:
"""True when the UA claims a crawler identity (bot or spider)."""
return bool(_BOT_UA.search(ua))
#: Plain-404 count per IP that classifies it as abuse even without a
#: telltale path hit.
_ABUSE_404_THRESHOLD = 10
#: Paths that instantly classify an IP as abuse when they 404: any segment
#: starting with a dot ("/.env", "/.git/config") or ending in ".php".
_ABUSE_PATH = re.compile(r"(^|/)\.|\.php$", re.IGNORECASE)
def _is_abuse_path(path: str) -> bool:
"""Telltale scanner path: dot segment or *.php."""
return bool(_ABUSE_PATH.search(path.split("?")[0]))
def _network_ip(ip: str) -> str:
"""IPv4 address unchanged, IPv6 collapsed to its /64 network address.
We hash the network rather than the full address so that clients in the
same /64 (a typical end-user allocation) are treated as one visitor.
"""
if not ip:
return ip
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ip
if isinstance(addr, ipaddress.IPv6Address):
return str(ipaddress.IPv6Network(f"{ip}/64", strict=False).network_address)
return ip
def _client_hash(ip: str, ua: str, lang: str) -> bytes:
"""6-byte blake3 digest identifying a visitor/client tuple.
The key is the prettified IP (IPv6 /64), the raw UA string and the
extracted language tag, separated by null bytes.
"""
return blake3.blake3(
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
).digest()[:6]
class Store:
"""In-memory analytics data plus the client-hash -> visit session map."""
def __init__(self, path: Path) -> None:
self.path = path
self.data = Analytics()
if path.exists():
try:
self.data = msgspec.json.decode(path.read_bytes(), type=Analytics)
except msgspec.DecodeError, OSError:
pass # legacy schema / corrupt or unreadable file: start fresh
#: client hash -> index of the current visit in data.visits
self.sessions: dict[bytes, int] = {}
#: ip -> external https origin of the latest document GET carrying
#: one, stashed for the visit the client's initial ping starts.
#: Internal or absent referers never touch the table.
self.pending_referers: dict[str, str] = {}
#: ip -> utm_* query parameters from the latest document GET that
#: carried any, stashed for the visit the client's initial ping starts.
#: Only non-empty sets are stored, so a later parameter-less page
#: does not overwrite an earlier tagged landing URL.
self.pending_utms: dict[str, dict[str, str]] = {}
#: Document GETs that have not yet been matched by a ping. Kept
#: in RAM only; expired entries are written to ``data.crawlers``.
self.pending_crawlers: list[CrawlerHit] = []
#: client hash -> {path: status} for recent document GETs, consumed
#: by the matching ping to record the status of each visited path.
self.pending_statuses: dict[bytes, dict[str, int]] = {}
#: ip -> number of plain (non-telltale) 404s seen, in RAM only;
#: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse.
self.not_found_counts: dict[str, int] = {}
#: Callables to notify when persisted data changes. Registered by the
#: analytics WebSocket broadcaster.
self._on_change: list[Callable[[], None]] = []
def subscribe(self, callback: Callable[[], None]) -> None:
"""Register a callback to be called after every persisted change."""
if callback not in self._on_change:
self._on_change.append(callback)
def unsubscribe(self, callback: Callable[[], None]) -> None:
"""Remove a previously registered change callback."""
with suppress(ValueError):
self._on_change.remove(callback)
def _notify(self) -> None:
for callback in self._on_change:
callback()
def _save(self) -> None:
"""Rewrite the JSON file atomically (temp file + rename)."""
try:
fd, tmp = tempfile.mkstemp(
dir=self.path.parent, prefix=self.path.name, suffix=".tmp"
)
with os.fdopen(fd, "wb") as f:
f.write(msgspec.json.encode(self.data))
os.replace(tmp, self.path)
except OSError:
pass # analytics must never break page serving
else:
self._notify()
def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]:
"""Move expired pending crawler hits into persistent ``data.crawlers``.
Hits from a hidden client (admin) are discarded instead of
persisted — admin browsing must not land in the crawler list.
Returns the client hashes of the newly flushed hits so callers can
schedule async enrichment.
"""
if not self.pending_crawlers:
return []
now = now or datetime.now(UTC)
cutoff = now - _CRAWLER_TIMEOUT
expired: list[CrawlerHit] = []
remaining: list[CrawlerHit] = []
for hit in self.pending_crawlers:
if hit.start > cutoff:
remaining.append(hit)
continue
client = self.data.clients.get(hit.client)
if client is not None and client.hide:
continue # hidden admin client: not a crawler
expired.append(hit)
if not expired:
self.pending_crawlers = remaining
return []
self.pending_crawlers = remaining
self.data.crawlers.extend(expired)
self._save()
return [hit.client for hit in expired]
def _hidden(self, client_hash: bytes) -> bool:
"""True when the client record is flagged hidden (admin)."""
client = self.data.clients.get(client_hash)
return client is not None and client.hide
def display(self) -> Display:
"""Build the viewer payload, excluding hidden clients.
The aggregates (site visits, page views, transitions) are computed
here from the visit records rather than stored, so a client that
becomes hidden after navigations were already logged disappears
from every statistic. Internal-path navigations count as page
views; external https targets are transitions only.
"""
visits = [v for v in self.data.visits if not self._hidden(v.client)]
display = Display(
visits=visits,
crawlers=[h for h in self.data.crawlers if not self._hidden(h.client)],
abuse=[h for h in self.data.abuse if not self._hidden(h.client)],
clients={h: c for h, c in self.data.clients.items() if not c.hide},
)
for visit in visits:
bucket = _bucket(visit.start)
site = display.site_visits
site[bucket] = site.get(bucket, 0) + 1
entry_views = display.views.setdefault(visit.entry, {})
entry_views[bucket] = entry_views.get(bucket, 0) + 1
fr = visit.referer or "(direct)"
buckets = display.transitions.setdefault(fr, {}).setdefault(visit.entry, {})
buckets[bucket] = buckets.get(bucket, 0) + 1
for t, nav in visit.navs.items():
nb = _bucket(t)
if nav.to.startswith("/"):
nav_views = display.views.setdefault(nav.to, {})
nav_views[nb] = nav_views.get(nb, 0) + 1
nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(nav.to, {})
nbuckets[nb] = nbuckets.get(nb, 0) + 1
return display
def display_json(self) -> str:
"""The ``display()`` payload as a JSON string for the WebSocket."""
return msgspec.json.encode(self.display()).decode()
def _client_ip(self, client_hash: bytes) -> str:
"""Return the IP stored for ``client_hash``, or "" if missing."""
client = self.data.clients.get(client_hash)
return client.ip if client else ""
def _ensure_client(
self,
ip: str,
ua: str,
lang: str,
*,
country: str = "",
) -> bytes:
"""Get or create a ``Client`` record; return its 6-byte hash."""
h = _client_hash(ip, ua, lang)
if h not in self.data.clients:
self.data.clients[h] = Client(
ip=ip,
ua=ua,
ua_pretty=_compact_user_agent(ua),
lang=lang,
country=country,
)
self._save()
return h
def enrich_client(
self,
client_hash: bytes,
*,
host: str = "",
country: str = "",
city: str = "",
) -> None:
"""Fill in host/geoip fields on a client record after async lookups."""
client = self.data.clients.get(client_hash)
if client is None:
return
changed = False
if host and not client.host:
client.host = host
changed = True
if country:
client.country = country
changed = True
if city:
client.city = city
changed = True
if changed:
self._save()
def _abuse_hit(
self,
client_hash: bytes,
path: str,
start: datetime | None = None,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Append one abuse hit referencing a client by hash."""
self.data.abuse.append(
AbuseHit(
start=start or datetime.now(UTC),
path=path,
client=client_hash,
flag=flag,
is_404=is_404,
)
)
def classify_abuse(
self,
ip: str,
client_hash: bytes,
path: str,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Classify an IP as a scanner/abuser and record the triggering hit.
All earlier crawler hits from the same IP (persisted and pending)
are moved to the abuse list — a random-UA scanner must not pollute
the crawler stats of the legitimate bots it impersonates.
"""
if ip not in self.data.abuse_ips:
self.data.abuse_ips[ip] = True
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
if moved:
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip]
for h in moved:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip]
if pending:
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip]
for h in pending:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
self._abuse_hit(client_hash, path, flag=flag, is_404=is_404)
self._save()
def track_404(
self,
ip: str,
ua: str,
path: str,
accept_language: str = "",
) -> bytes:
"""Record a 404 response for ``path`` (full path, query included).
A telltale path (dot segment or *.php) classifies the IP as abuse
immediately; enough plain 404s from one IP do too. Hits from
already-classified IPs go straight to the abuse list.
Returns the client hash so callers can schedule async enrichment.
"""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
self._abuse_hit(client_hash, path, flag=_is_abuse_path(path), is_404=True)
self._save()
return client_hash
if _is_abuse_path(path):
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
self.not_found_counts[ip] = self.not_found_counts.get(ip, 0) + 1
if self.not_found_counts[ip] >= _ABUSE_404_THRESHOLD:
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
return client_hash
def _new_visit(
self,
entry: str,
referer: str,
client_hash: bytes,
utm: dict[str, str] | None = None,
status: int = 200,
) -> Visit:
now = datetime.now(UTC)
visit = Visit(
start=now,
entry=entry,
referer=referer,
client=client_hash,
utm=utm or {},
)
visit.trail[now] = TrailItem(to=entry, status=status)
self.data.visits.append(visit)
self.sessions[client_hash] = len(self.data.visits) - 1
return visit
def track_entry(
self,
referer: str,
own_origin: str,
ip: str,
ua: str,
full_path: str,
accept_language: str = "",
*,
status: int = 200,
) -> list[bytes]:
"""Stash the entry referer/UTM tags and queue a pending crawler hit.
Nothing is counted here — the client's initial /_a ping starts the
visit (only non-admin clients ping). Only a cross-origin https
referer updates the table; an internal or absent referer leaves any
stashed origin untouched. UTM parameters are kept only when the
landing URL actually carries them, so a subsequent parameter-less page
does not erase an earlier tagged landing.
Every document GET is also queued as a pending crawler hit. If a ping
from the same client arrives within ``_CRAWLER_TIMEOUT``, the hit is
discarded; otherwise it is flushed to ``data.crawlers``. The
Accept-Language header is stored on the client record immediately;
host/geoip are filled in later by async enrichment.
GETs from IPs already classified as abuse are recorded as abuse hits
with the full request path (query string included).
Returns the client hashes of any hits flushed to persistent storage,
so callers can schedule async enrichment.
"""
entry = full_path.split("?")[0]
query = full_path.split("?", 1)[1] if "?" in full_path else ""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
flushed = self._flush_crawlers()
self._abuse_hit(client_hash, full_path, is_404=False, flag=False)
self._save()
return flushed
now = datetime.now(UTC)
flushed = self._flush_crawlers(now)
if referer:
origin = _origin(referer)
if origin is not None and origin != own_origin:
self.pending_referers[ip] = origin
utms = _utm_tags(query)
if utms:
self.pending_utms[ip] = utms
self.pending_crawlers.append(
CrawlerHit(
start=now,
entry=entry,
client=client_hash,
referer=self.pending_referers.get(ip, ""),
query=query,
status=status,
)
)
self.pending_statuses.setdefault(client_hash, {})[entry] = status
return flushed
def _add_read(self, client_hash: bytes, path: str, seconds: int) -> None:
"""Add ``seconds`` of reading time for ``path`` to the current visit."""
if seconds <= 0:
return
index = self.sessions.get(client_hash)
if index is None or index >= len(self.data.visits):
return
visit = self.data.visits[index]
for item in visit.trail.values():
if item.to == path:
item.read += seconds
return
def ping(
self,
from_: str,
to: str | None,
ip: str,
ua: str,
accept_language: str = "",
hide: bool = False,
read: int = 0,
) -> tuple[int | None, list[bytes]]:
"""Record a client navigation ping ({from, to, read} from pagerite.js).
``to`` is an internal path ("/...") or an https URL for exit links; a
missing/empty ``to`` means the page is being closed and only the
``read`` time should be recorded. The transition is always counted when
``to`` is present; the trail only grows on first sight of a page within
the visit. ``read`` is the active time (seconds) spent on ``from_``.
A ping with no known session starts a fresh visit, consuming the
referer and UTM tags stashed by the document GET if there are any.
``hide`` is set by admin clients: the client record is flagged
``hide`` — which covers everything it ever did, including visits and
crawler hits from before the login — and the navigation is recorded
normally. Hidden clients are excluded from every statistic and list
at display time, and their pending crawler hits are discarded.
Pings from IPs classified as abuse, and pings whose User-Agent
claims a JS-running crawler identity (``_is_bot_ua``), are ignored
entirely — the crawler's pending hits stay queued and flush to
``data.crawlers`` normally.
Returns the index of the new visit when one is created (or None) and
the client hashes of any crawler hits flushed by this call, so callers
can schedule async enrichment (host, geoip country/city).
"""
flushed = self._flush_crawlers()
lang, country = _parse_accept_language(accept_language)
if hide:
# Admin ping: flag the client hidden and never a crawler hit.
# The flag lives on the client record, so it covers visits and
# crawler hits from before the login too; display-time
# aggregation excludes hidden clients from every statistic.
client_hash = self._ensure_client(ip, ua, lang, country=country)
self.data.clients[client_hash].hide = True
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash
]
else:
client_hash = _client_hash(ip, ua, lang)
if ip in self.data.abuse_ips:
return None, flushed
if _is_bot_ua(ua):
# A JS-running crawler (Googlebot, GoogleOther, Applebot
# execute JS and ping): never a visit. Its pending crawler
# hits are kept and flush to ``data.crawlers`` normally.
return None, flushed
# A real visitor ping cancels any pending crawler hits from
# this client.
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash
]
fr_path = _internal_path(from_) if from_ else ""
if fr_path and read > 0:
self._add_read(client_hash, fr_path, read)
if not to:
if read > 0 or hide:
self._save()
return None, flushed
if to.startswith("/") and not to.startswith("//"):
target = _internal_path(to) or ""
else:
target = _external_target(to) or ""
if not target:
return None, flushed
index = self.sessions.get(client_hash)
fr = fr_path or "(direct)"
statuses = self.pending_statuses.setdefault(client_hash, {})
target_status = statuses.pop(target, None) or 200
if not statuses:
self.pending_statuses.pop(client_hash, None)
if index is None or index >= len(self.data.visits):
# No known session: the initial ping of a fresh page load (or
# missing data after a server restart) — start a visit.
index = len(self.data.visits)
self._ensure_client(ip, ua, lang, country=country)
self._new_visit(
target,
self.pending_referers.pop(ip, ""),
client_hash,
utm=self.pending_utms.pop(ip, {}),
status=target_status,
)
else:
visit = self.data.visits[index]
now = datetime.now(UTC)
visit.navs[now] = Nav(fr=fr, to=target)
# First-seen only: repeat pages and repeated exits update the
# existing trail item (most recent status) instead of appending.
for item in visit.trail.values():
if item.to == target:
item.status = target_status
break
else:
visit.trail[now] = TrailItem(to=target, status=target_status)
self._save()
visit_index = index if index is not None and index < len(self.data.visits) else None
return visit_index, flushed
+550 -10
View File
@@ -12,23 +12,40 @@ walking the tree (``resolve``), moves are slot detach/attach
(``find_slot``) with a fresh order key from the new siblings. (``find_slot``) with a fresh order key from the new siblings.
""" """
import asyncio
import gzip
import ipaddress
import mimetypes import mimetypes
import os import os
import re import re
import shutil
import socket
from collections.abc import AsyncIterator from collections.abc import AsyncIterator
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from datetime import UTC, datetime from datetime import UTC, datetime
from email.utils import format_datetime from email.utils import format_datetime
from functools import lru_cache
from pathlib import Path from pathlib import Path
from urllib.parse import urlparse
from xml.sax.saxutils import escape as xml_escape
import blake3 import blake3
from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect import msgspec
from fastapi.responses import HTMLResponse, RedirectResponse, Response from fastapi import (
FastAPI,
HTTPException,
Query,
Request,
WebSocket,
WebSocketDisconnect,
)
from fastapi.responses import RedirectResponse, Response
from fastapi_vue import Frontend from fastapi_vue import Frontend
from kanta import Kanta from kanta import Kanta
from pydantic import BaseModel from pydantic import BaseModel
from zstandard import ZstdCompressor
from pagerite import seed, views from pagerite import analytics, seed, views
from pagerite.__main__ import DEVMODE from pagerite.__main__ import DEVMODE
from pagerite.data import ( from pagerite.data import (
Data, Data,
@@ -43,6 +60,104 @@ from pagerite.markdown import has_h1, render, toggle_task
DB_PATH = os.getenv("PAGERITE_DB", "pagerite.kantadb") DB_PATH = os.getenv("PAGERITE_DB", "pagerite.kantadb")
# Visit analytics go to their own JSON file, not the kanta database.
ANALYTICS_PATH = Path(
os.getenv("PAGERITE_ANALYTICS", DB_PATH.replace(".kantadb", "") + ".analytics.json")
)
analytics_store = analytics.Store(ANALYTICS_PATH)
# Live WebSocket clients for the analytics stream.
_analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None
# Repository root from this file's location (pagerite/app.py -> ..).
_REPO_ROOT = Path(__file__).resolve().parent.parent
def _geoip_db_path() -> Path | None:
"""Find a DB-IP MMDB in the repo root, preferring an already-decompressed
``.mmdb`` over the matching ``.mmdb.gz``. Returns None if none is present.
"""
mmdb = sorted(_REPO_ROOT.glob("dbip-*.mmdb"))
if mmdb:
return mmdb[0]
gz = sorted(_REPO_ROOT.glob("dbip-*.mmdb.gz"))
if gz:
return gz[0]
return None
class GeoIP:
"""Lazy DB-IP MMDB reader. Call ``_load()`` once at startup before
concurrent requests arrive; ``country()`` is read-only and safe to call
from ``asyncio.to_thread`` workers afterwards.
"""
def __init__(self) -> None:
self._reader: object | None = None
def _decompress(self, source: Path, target: Path) -> None:
if target.exists():
return
tmp = target.with_suffix(target.suffix + ".tmp")
with gzip.open(source, "rb") as src, open(tmp, "wb") as dst:
shutil.copyfileobj(src, dst)
os.replace(tmp, target)
def _load(self) -> None:
if self._reader is not None:
return
source = _geoip_db_path()
if source is None:
return
if source.suffix == ".gz":
target = source.with_suffix("")
self._decompress(source, target)
source = target
try:
import maxminddb
self._reader = maxminddb.open_database(str(source))
except Exception:
pass
def country(self, ip: str) -> str:
"""Two-letter ISO country code for ``ip``, or "" when unavailable."""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
return (rec.get("country") or {}).get("iso_code", "")
except Exception:
pass
return ""
def city(self, ip: str) -> str:
"""City name for ``ip``, or "" when unavailable.
GeoIP sometimes appends district names in parentheses (e.g.
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
the value is stored.
"""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
city = (rec.get("city") or {}).get("names", {}).get("en", "")
if city:
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
return city
except Exception:
pass
return ""
_geoip = GeoIP()
# Our own data root; kanta edits it in place, reads are plain attribute access. # Our own data root; kanta edits it in place, reads are plain attribute access.
data = Data() data = Data()
kanta = Kanta(DB_PATH, data) kanta = Kanta(DB_PATH, data)
@@ -142,11 +257,16 @@ def _seed(data: Data) -> None:
@asynccontextmanager @asynccontextmanager
async def lifespan(_app: FastAPI) -> AsyncIterator[None]: async def lifespan(_app: FastAPI) -> AsyncIterator[None]:
"""Open the database, migrate legacy content, load assets.""" """Open the database, migrate legacy content, load assets, load GeoIP."""
await kanta.open() await kanta.open()
_migrate_legacy() _migrate_legacy()
await frontend.load() await frontend.load()
# Decompress/open the DB-IP MMDB once at startup. Lookups are then
# read-only and safe to run in background ``to_thread`` workers.
await asyncio.to_thread(_geoip._load)
analytics_store.subscribe(_schedule_analytics_broadcast)
yield yield
analytics_store.unsubscribe(_schedule_analytics_broadcast)
await kanta.close() await kanta.close()
@@ -170,6 +290,79 @@ async def _headers(request: Request, call_next) -> Response:
return response return response
# Dynamic HTML is compressed per request at level 9 (static assets are
# already pre-compressed by fastapi-vue's Frontend).
_zstd = ZstdCompressor(9)
def _render_html(kind: str, path: str, base_url: str) -> str:
"""Render one of the generated pages (see _html_response)."""
if kind == "page":
return views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, base_url)
if kind == "category":
return views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
if kind == "not-found":
return views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
return views.render_analytics(data.menu, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
@lru_cache(maxsize=128)
def _cached_body(kind: str, path: str, base_url: str, version: int, zstd: bool) -> bytes:
"""Rendered page body. Every input the output depends on is in the key:
data.version bumps on any content/settings change, base_url feeds the
social meta URLs, and zstd selects the stored encoding (both variants
are cached rather than re-compressed).
"""
body = _render_html(kind, path, base_url).encode()
return _zstd.compress(body) if zstd else body
def _html_response(
request: Request,
kind: str,
path: str,
status_code: int = 200,
headers: dict | None = None,
etag: bool = False,
) -> Response:
"""Response for a generated page, zstd-compressed when the client
accepts it (no gzip fallback).
Done per handler rather than in middleware so that Frontend's
already-compressed asset responses are never touched. The ETag stays
identical across encodings (revalidation compares it before
compression); ``vary: accept-encoding`` keeps caches from mixing the
representations. In dev the cache is bypassed so theme/design edits on
disk apply immediately.
``etag=True`` derives the validator from a blake3 hash of the
(uncompressed) body — for pages like /_a that have no Node whose
modified timestamp could serve as one — and answers matching
if-none-match revalidations with a 304.
"""
zstd = "zstd" in request.headers.get("accept-encoding", "")
# Absolute social/canonical URLs use the learned public origin; until
# an admin visit teaches it, fall back to the request's own base URL.
base_url = data.site_url or str(request.base_url).rstrip("/")
if DEVMODE:
identity = _render_html(kind, path, base_url).encode()
body = _zstd.compress(identity) if zstd else identity
else:
identity = _cached_body(kind, path, base_url, data.version, False)
body = _cached_body(kind, path, base_url, data.version, True) if zstd else identity
h = dict(headers or {})
if zstd:
h["vary"] = "accept-encoding"
if etag:
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
h["etag"] = tag
if request.headers.get("if-none-match") == tag:
return Response(status_code=304, headers=h)
if zstd:
h["content-encoding"] = "zstd"
return Response(body, status_code, h, media_type="text/html")
class PageIn(BaseModel): class PageIn(BaseModel):
"""Payload for creating or replacing a page.""" """Payload for creating or replacing a page."""
@@ -323,6 +516,33 @@ async def put_settings(settings: SettingsIn) -> None:
data.version += 1 data.version += 1
class SiteUrlIn(BaseModel):
"""Payload for learning the site's public origin."""
url: str
@app.post("/_api/site-url", status_code=204)
async def learn_site_url(payload: SiteUrlIn) -> None:
"""Learn the site's public origin (scheme + host) from an admin browser.
pagerite.js reports location.origin once an admin session is detected:
unlike request Host headers it reflects the real public scheme and host
even behind reverse proxies, with zero manual configuration. Stored in
the database with a version bump so cached pages re-render with correct
absolute social/canonical URLs.
"""
url = payload.url.rstrip("/")
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or not parsed.netloc or parsed.path:
raise HTTPException(400, "not an origin")
if url == data.site_url:
return
with kanta.transaction("learn site url"):
data.site_url = url
data.version += 1
@app.put("/_api/settings/favicon") @app.put("/_api/settings/favicon")
async def put_favicon(request: Request) -> dict[str, str]: async def put_favicon(request: Request) -> dict[str, str]:
"""Upload a favicon into the content-addressed store and activate it. """Upload a favicon into the content-addressed store and activate it.
@@ -495,6 +715,192 @@ async def delete_page(path: str) -> None:
_SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$") _SLUG_RE = re.compile(r"^[a-z0-9][a-z0-9_-]*$")
def _client_ip(request: Request) -> str:
"""Client IP: first X-Forwarded-For hop (we sit behind a proxy), else
the direct peer."""
forwarded = request.headers.get("x-forwarded-for", "").split(",")[0].strip()
return forwarded or (request.client.host if request.client else "")
def _query_suffix(request: Request) -> str:
"""The request's query string as a "?..." suffix, or "" when absent."""
query = str(request.url.query)
return f"?{query}" if query else ""
@lru_cache(maxsize=4096)
def _cached_ptr(ip: str) -> str:
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
if not ip:
return ""
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ""
if addr.is_private or addr.is_loopback or addr.is_reserved or addr.is_multicast or addr.is_link_local:
return ""
try:
host, _, _ = socket.gethostbyaddr(ip)
except socket.herror:
return ""
return host
async def _lookup_host(ip: str) -> str:
"""Async wrapper around ``_cached_ptr``; runs the blocking lookup in a thread."""
return await asyncio.to_thread(_cached_ptr, ip)
async def _geoip_country(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB lookup."""
return await asyncio.to_thread(_geoip.country, ip)
async def _geoip_city(ip: str) -> str:
"""Async wrapper around the DB-IP MMDB city lookup."""
return await asyncio.to_thread(_geoip.city, ip)
async def _enrich_client(client_hash: bytes) -> None:
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
client = analytics_store.data.clients.get(client_hash)
if not client or not client.ip:
return
host = await _lookup_host(client.ip)
country = await _geoip_country(client.ip)
city = await _geoip_city(client.ip)
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
"""Start background host/geoip enrichment for the given client hashes."""
for client_hash in client_hashes:
asyncio.create_task(_enrich_client(client_hash))
async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients:
return
payload = analytics_store.display_json()
closed = set()
for ws in _analytics_ws_clients:
try:
await ws.send_text(payload)
except Exception:
closed.add(ws)
for ws in closed:
_analytics_ws_clients.discard(ws)
async def _debounced_analytics_broadcast() -> None:
"""Wait briefly, then broadcast the latest snapshot once."""
await asyncio.sleep(0.2)
await _broadcast_analytics()
def _schedule_analytics_broadcast() -> None:
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
global _analytics_broadcast_task
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
return
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
_debounced_analytics_broadcast()
)
@app.get("/_a", response_model=None)
async def analytics_page(request: Request) -> Response:
"""Render the analytics viewer as a normal site page at /_a.
The page itself is public, but the data stream (/_api/ws/analytics) stays
admin-gated like the rest of /_api, so only authorized users see the
statistics; others get the viewer with a "could not be loaded" message.
"""
return _html_response(
request,
"analytics",
"",
headers={"cache-control": "no-cache"},
etag=True,
)
@app.post("/_a", status_code=204)
async def analytics_ping(
request: Request,
fr: str = Query(""),
to: str | None = Query(None),
hide: int = Query(0),
read: int = Query(0),
) -> None:
"""Record a navigation ping (?fr=&to=&hide=&read=); fire-and-forget.
The initial page-load ping carries only ``to``: the entry is attributed
to the referer/UTM tags stashed by the document GET (see _track_entry),
which JS cannot see once the page has loaded.
The reverse-DNS and DB-IP geoip lookups happen in a background task so
the response is never delayed by slow DNS or the first MMDB decompress.
"""
ip = _client_ip(request)
visit_index, flushed_clients = analytics_store.ping(
fr,
to,
ip,
request.headers.get("user-agent", ""),
request.headers.get("accept-language", ""),
hide=bool(hide),
read=read,
)
if visit_index is not None:
visit = analytics_store.data.visits[visit_index]
asyncio.create_task(_enrich_client(visit.client))
_schedule_client_enrichment(flushed_clients)
def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]:
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
Nothing is counted on the GET itself — the client's /_a ping starts the
visit, so bots never register as visits (JS-running crawlers ping too,
but the ping handler ignores known bot UAs). (Admin clients ping too,
but with hide=1, which flags their visit hidden: it is recorded but
excluded from all statistics and from the crawler list.)
The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
logged as a crawler hit. The root-path and localhost checks prevent
remote visitors from hiding traffic with the same query string.
Returns the client hashes of any pending crawler hits flushed to persistent
storage, so callers can schedule async geoip and reverse-DNS enrichment.
"""
if request.headers.get("x-pagerite-preload"):
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
# ping sent when the user actually navigates does the counting.
# (Forging the header only hides a GET from the crawler stats; the
# path-based abuse classification is unaffected.)
return []
if (
path == ""
and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1"
):
return []
own_origin = f"https://{urlparse(str(request.base_url)).netloc}"
full_path = f"{request.url.path}{_query_suffix(request)}"
return analytics_store.track_entry(
request.headers.get("referer", ""),
own_origin,
_client_ip(request),
request.headers.get("user-agent", ""),
full_path,
request.headers.get("accept-language", ""),
status=status,
)
def _http_date(dt: datetime) -> str: def _http_date(dt: datetime) -> str:
"""RFC 7231 date for the Last-Modified header.""" """RFC 7231 date for the Last-Modified header."""
return format_datetime(dt.astimezone(UTC), usegmt=True) return format_datetime(dt.astimezone(UTC), usegmt=True)
@@ -510,6 +916,15 @@ def _is_reserved(path: str) -> bool:
return any(not _SLUG_RE.match(seg) for seg in path.split("/")) return any(not _SLUG_RE.match(seg) for seg in path.split("/"))
def _is_trackable_path(path: str) -> bool:
"""Content URLs only: skip auth endpoints and reserved/machinery paths."""
if not path:
return True
if path == "auth" or path.startswith("auth/"):
return False
return not _is_reserved(path)
def _check_reserved(path: str) -> None: def _check_reserved(path: str) -> None:
"""Reject paths that do not follow the slug charset.""" """Reject paths that do not follow the slug charset."""
if _is_reserved(path): if _is_reserved(path):
@@ -519,6 +934,25 @@ def _check_reserved(path: str) -> None:
) )
@app.websocket("/_api/ws/analytics")
async def analytics_websocket(ws: WebSocket) -> None:
"""Stream the analytics snapshot, then push updates as they happen.
Admin-only via the /_api forward-auth gate, like every management
endpoint. Powers the analytics viewer rendered at /_a.
"""
await ws.accept()
await ws.send_text(analytics_store.display_json())
_analytics_ws_clients.add(ws)
try:
while True:
await ws.receive_text()
except Exception:
pass
finally:
_analytics_ws_clients.discard(ws)
@app.websocket("/_api/ws/editor") @app.websocket("/_api/ws/editor")
async def editor_ws(ws: WebSocket) -> None: async def editor_ws(ws: WebSocket) -> None:
"""Editor session: open pages, render previews, save — over one socket. """Editor session: open pages, render previews, save — over one socket.
@@ -670,22 +1104,108 @@ async def front_page(request: Request) -> Response:
return await show_page(request, "") return await show_page(request, "")
@app.get("/sitemap.xml")
async def sitemap(request: Request) -> Response:
"""Dynamically generate a sitemap of all published article pages."""
base = str(request.base_url).rstrip("/")
entries: list[tuple[str, datetime, int]] = []
def walk(
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
) -> None:
first_content_slug = next(
(
slug
for slug, node in sorted_nodes(nodes)
if node.published and node.content is not None
),
None,
)
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
depth = path.count("/") if path else 0
if (
not parent_has_content
and slug == first_content_slug
and node.published
and node.content is not None
and depth > 0
):
depth -= 1
if node.published and node.content is not None:
entries.append((path, node.modified, depth))
if node.children:
walk(node.children, path, node.content is not None)
walk(data.menu, "")
def priority(depth: int) -> float:
return max(0.1, 1.0 - depth * 0.2)
lines = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
]
for path, modified, depth in entries:
loc = xml_escape(f"{base}/{path}" if path else base)
lastmod = (
modified.astimezone(UTC).replace(microsecond=0).isoformat().replace("+00:00", "Z")
)
lines.append(
f" <url>"
f"<loc>{loc}</loc>"
f"<lastmod>{lastmod}</lastmod>"
f"<priority>{priority(depth):.1f}</priority>"
f"</url>"
)
lines.append("</urlset>")
return Response(
"\n".join(lines),
media_type="application/xml",
headers={"cache-control": "no-cache"},
)
@app.get("/robots.txt")
async def robots_txt(request: Request) -> Response:
"""Allow all crawling and point crawlers at the sitemap."""
base = str(request.base_url).rstrip("/")
body = f"User-agent: *\nAllow: /\nSitemap: {base}/sitemap.xml\n"
return Response(
body,
media_type="text/plain",
headers={"cache-control": "no-cache"},
)
# Vue build asset routes are inserted at this position during load(): the # Vue build asset routes are inserted at this position during load(): the
# build mirrors the URL space (/_assets/*, /favicon.ico at the root). # build mirrors the URL space (/_assets/*, /favicon.ico at the root).
frontend.route(app, "/") frontend.route(app, "/")
@app.get("/{path:path}", response_model=None) @app.get("/{path:path}", response_model=None)
async def show_page(request: Request, path: str) -> HTMLResponse | Response: async def show_page(request: Request, path: str) -> Response:
"""Render the content page at a slug path, or 404. """Render the content page at a slug path, or 404.
A node without content is a category label: its URL renders a A node without content is a category label: its URL renders a
placeholder page (nav links point straight at its first child). placeholder page (nav links point straight at its first child).
""" """
path = path.strip("/") path = path.strip("/")
ua = request.headers.get("user-agent", "")
accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path): if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its # Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page. # built-in 404 instead of rendering an editable article page.
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
# as abuse in analytics.
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
raise HTTPException(404) raise HTTPException(404)
chain = resolve(data.menu, path) chain = resolve(data.menu, path)
node = chain[-1] if chain else None node = chain[-1] if chain else None
@@ -699,8 +1219,13 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
etag = f'"{path}@{node.modified.timestamp()}v{data.version}"' etag = f'"{path}@{node.modified.timestamp()}v{data.version}"'
if request.headers.get("if-none-match") == etag: if request.headers.get("if-none-match") == etag:
return Response(status_code=304) return Response(status_code=304)
return HTMLResponse( if _is_trackable_path(path):
views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), flushed = _track_entry(path, request)
_schedule_client_enrichment(flushed)
return _html_response(
request,
"page",
path,
headers={ headers={
"etag": etag, "etag": etag,
"last-modified": _http_date(node.modified), "last-modified": _http_date(node.modified),
@@ -710,8 +1235,13 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
if node is not None and node.published and node.content is None: if node is not None and node.published and node.content is None:
# Category label without a landing page: placeholder with the pen # Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real). # to create it (404 — no page here, but the node is real).
return HTMLResponse( if _is_trackable_path(path):
views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(
request,
"category",
path,
404, 404,
headers={ headers={
"last-modified": _http_date(node.modified), "last-modified": _http_date(node.modified),
@@ -724,4 +1254,14 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
for slug, item in sorted_nodes(data.menu): for slug, item in sorted_nodes(data.menu):
if item.published: if item.published:
return RedirectResponse(f"/{slug}") return RedirectResponse(f"/{slug}")
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), 404) if _is_trackable_path(path):
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(request, "not-found", path, 404)
+6
View File
@@ -100,6 +100,12 @@ class Data(msgspec.Struct):
#: Favicon: name of a file in `files` (content-addressed), linked as #: Favicon: name of a file in `files` (content-addressed), linked as
#: <link rel="icon"> on every page. Empty = the build's /favicon.ico. #: <link rel="icon"> on every page. Empty = the build's /favicon.ico.
favicon: str = "" favicon: str = ""
#: Public origin (scheme + host) of the site, learned from admin
#: browsers (POST /_api/site-url — location.origin is correct even
#: behind reverse proxies, unlike request Host headers). Used for
#: absolute social/canonical URLs; empty = fall back to the request's
#: own base URL.
site_url: str = ""
#: Legacy flat page store (pre-tree databases); migrated into `menu` #: Legacy flat page store (pre-tree databases); migrated into `menu`
#: on startup, then cleared. Never written otherwise. #: on startup, then cleared. Never written otherwise.
pages: dict[str, Page] = {} pages: dict[str, Page] = {}
+149 -2
View File
@@ -4,8 +4,19 @@ Raw HTML (including inline scripts) is passed through unfiltered: the
single author is trusted. Extensions: tables and strikethrough (from the single author is trusted. Extensions: tables and strikethrough (from the
"default" preset), footnotes, definition lists, task lists, "default" preset), footnotes, definition lists, task lists,
brace-attributes (`{.class width=300}` on any element, images in brace-attributes (`{.class width=300}` on any element, images in
particular) and admonitions (``!!! note Title`` with an indented body — particular), admonitions (``!!! note Title`` with an indented body —
note/tip/warning/etc., the title optional). Bare URLs autolink (GFM), with note/tip/warning/etc., the title optional) and GitHub-style alerts
(``> [!NOTE]`` / TIP / IMPORTANT / WARNING / CAUTION, rendered in the
same callout styling). ``::: name`` opens a generic container rendered
as ``<div class="name">`` and closed by a matching ``:::`` (nest by
giving the outer container more colons, e.g. `::::`); the name may be
followed by brace attributes (``::: aside {.right}``). ``::: aside``
floats as a side box beside the text and ``::: nocols`` opts its
section out of the column layout. A brace-attribute
line as a block's last line (no blank line between) applies to the whole
block, e.g. a paragraph ending with ``{.wide}`` breaks out of the column
layout as a full-width element; written after a block (code fence,
heading, container, ...) it applies to that preceding block. Bare URLs autolink (GFM), with
the ``https://`` scheme hidden in the link text (``http://`` and other the ``https://`` scheme hidden in the link text (``http://`` and other
schemes stay visible; manually labelled links are untouched), and schemes stay visible; manually labelled links are untouched), and
``H~2~O`` / ``x^2^`` give sub/superscripts. ``H~2~O`` / ``x^2^`` give sub/superscripts.
@@ -33,6 +44,8 @@ from markdown_it.common.utils import escapeHtml
from markdown_it.renderer import RendererHTML from markdown_it.renderer import RendererHTML
from mdit_py_plugins.admon import admon_plugin from mdit_py_plugins.admon import admon_plugin
from mdit_py_plugins.attrs import attrs_plugin from mdit_py_plugins.attrs import attrs_plugin
from mdit_py_plugins.attrs.parse import ParseError, parse as parse_attrs
from mdit_py_plugins.container import container_plugin
from mdit_py_plugins.deflist import deflist_plugin from mdit_py_plugins.deflist import deflist_plugin
from mdit_py_plugins.footnote import footnote_plugin from mdit_py_plugins.footnote import footnote_plugin
from mdit_py_plugins.gfm_autolink import gfm_autolink_plugin from mdit_py_plugins.gfm_autolink import gfm_autolink_plugin
@@ -65,6 +78,31 @@ def _highlight(text: str, lang: str, _attrs: str) -> str:
return highlight(text, lexer, _formatter) return highlight(text, lexer, _formatter)
def _fence_rule(
self: RendererHTML,
tokens,
idx: int,
options,
env: dict,
) -> str:
"""Render a fenced code block.
Like the default fence rule, but block attributes (a trailing `{...}`
line, applied to the fence token by _block_attrs) go on the <pre> — the
block element — instead of the <code>, which keeps only the language
class. This is what makes e.g. `{.wide}` or `{style="..."}` after a
code fence style the block itself.
"""
token = tokens[idx]
info = token.info.strip() if token.info else ""
lang = info.split(maxsplit=1)[0] if info else ""
highlighted = (_highlight(token.content, lang, "")
or escapeHtml(token.content))
code_class = f' class="{options.langPrefix}{lang}"' if lang else ""
return (f"<pre{self.renderAttrs(token)}><code{code_class}>"
f"{highlighted}</code></pre>\n")
def _image_rule( def _image_rule(
self: RendererHTML, self: RendererHTML,
tokens, tokens,
@@ -107,6 +145,10 @@ def _unwrap_lone_figures(state) -> None:
if child and child.type == "image": if child and child.type == "image":
if (tokens[i - 1].type == "paragraph_open" if (tokens[i - 1].type == "paragraph_open"
and tokens[i + 1].type == "paragraph_close"): and tokens[i + 1].type == "paragraph_close"):
# A lone image becomes a <figure> (see _image_rule); block
# attrs on the paragraph (e.g. a trailing {.wide} line) move
# onto the image so they survive the unwrap.
_apply_attrs(child, tokens[i - 1].attrs or {})
tokens[i - 1].hidden = True tokens[i - 1].hidden = True
tokens[i + 1].hidden = True tokens[i + 1].hidden = True
@@ -149,6 +191,104 @@ def _shorten_autolinks(state) -> None:
text.content = text.content.removeprefix("https://") text.content = text.content.removeprefix("https://")
_CONTAINER_NAME_RE = re.compile(r"[a-zA-Z][\w-]*")
def _apply_attrs(token, attrs: dict) -> None:
"""Join/set parsed brace attributes (`{.class key=value}`) on a token."""
for key, value in attrs.items():
if key == "class":
token.attrJoin("class", value)
else:
token.attrSet(key, value)
def _container_validate(params: str, _markup: str) -> bool:
"""`::: name`, optionally followed by brace attrs (`::: aside {.right}`)."""
name, _, rest = params.strip().partition(" ")
if not _CONTAINER_NAME_RE.fullmatch(name):
return False
rest = rest.strip()
if not rest:
return True
try:
pos, _ = parse_attrs(rest)
except ParseError:
return False
# parse() stops at (returns the index of) the closing brace.
return pos == len(rest) - 1
def _container_render(self, tokens, idx, options, env):
"""Render `::: name {attrs}` containers as `<div class="name">`."""
token = tokens[idx]
if token.nesting == 1:
name, _, rest = token.info.strip().partition(" ")
token.attrJoin("class", name)
if rest.strip():
_, attrs = parse_attrs(rest.strip())
_apply_attrs(token, attrs)
return self.renderToken(tokens, idx, options, env)
def _block_attrs(state) -> None:
"""Apply `{.class key=value}` on a block's last line to the block.
The inline attrs plugin only covers attributes right after an image,
code span or link; this extends the same brace syntax to whole blocks,
e.g. a paragraph ending with a `{.wide}` line (no blank line between)
gets the `wide` class and thereby breaks out of the column layout.
A lone `{...}` paragraph applies to the previous block instead (this
is how headings take attributes, since a heading's next line always
starts a new paragraph). Runs before the typographer so quotes inside
attributes stay straight.
"""
tokens = state.tokens
for i, token in enumerate(tokens):
if token.type != "inline" or not token.children:
continue
text = token.children[-1]
if (text.type != "text" or not text.content.startswith("{")
or not text.content.endswith("}")):
continue
try:
_, attrs = parse_attrs(text.content.strip())
except ParseError:
continue
standalone = len(token.children) == 1
if not standalone and token.children[-2].type != "softbreak":
continue
# The target: the enclosing block for a trailing attrs line, or the
# previous same-level block for a standalone attrs paragraph —
# including self-contained blocks like code fences and <hr>. Never
# a hidden token (tight-list paragraphs render no tag to hold the
# attributes) — in that case leave the text untouched instead of
# silently swallowing it.
own = i - 1 # standalone: the attrs paragraph's own opening token
j = i - 1
while j >= 0:
target = tokens[j]
if target.hidden:
pass
elif standalone:
if (j != own and target.level == tokens[own].level
and (target.nesting == 1
or target.type in ("fence", "code_block", "hr"))):
break
elif target.nesting == 1:
break
j -= 1
if j < 0:
continue
_apply_attrs(tokens[j], attrs)
if standalone:
tokens[own].hidden = True
token.children = []
tokens[i + 1].hidden = True
else:
del token.children[-2:]
md = ( md = (
MarkdownIt( MarkdownIt(
"default", "default",
@@ -161,6 +301,8 @@ md = (
) )
.use(attrs_plugin) .use(attrs_plugin)
.use(admon_plugin) .use(admon_plugin)
.use(container_plugin, "block", validate=_container_validate,
render=_container_render)
.use(footnote_plugin) .use(footnote_plugin)
.use(deflist_plugin) .use(deflist_plugin)
.use(tasklists_plugin, enabled=True) .use(tasklists_plugin, enabled=True)
@@ -169,6 +311,11 @@ md = (
.use(superscript_plugin) .use(superscript_plugin)
) )
md.add_render_rule("image", _image_rule) md.add_render_rule("image", _image_rule)
md.add_render_rule("fence", _fence_rule)
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
md.options["alerts"] = True
# Block attrs must be stripped before the typographer curlifies their quotes.
md.core.ruler.before("replacements", "block_attrs", _block_attrs)
md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures) md.core.ruler.push("unwrap_lone_figures", _unwrap_lone_figures)
md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes) md.core.ruler.push("tag_task_checkboxes", _tag_task_checkboxes)
md.core.ruler.push("shorten_autolinks", _shorten_autolinks) md.core.ruler.push("shorten_autolinks", _shorten_autolinks)
+5 -3
View File
@@ -29,6 +29,7 @@ Where to go next:
- The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side. - The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side.
- The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read. - The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read.
- Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree. - Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree.
- Elsewhere on the web: [![xkcd 927: Standards](https://imgs.xkcd.com/comics/standards.png "xkcd 927: Standards"){width=240}](https://xkcd.com/927/) — a cautionary tale about adding one more standard.
![Abstract waves](waves.svg "Generated SVG artwork, attached to this page"){width=420} ![Abstract waves](waves.svg "Generated SVG artwork, attached to this page"){width=420}
@@ -68,15 +69,15 @@ Every feature below is shown twice: first the Markdown source, then how it rende
### A subsection ### A subsection
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a *Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a
[link to the front page](/). Plain URLs become links automatically: [link to the front page](/). An image that links to its page:
https://example.com — and a hard line break [![xkcd 1179: ISO 8601](https://imgs.xkcd.com/comics/iso_8601.png "xkcd 1179: ISO 8601"){width=240}](https://xkcd.com/1179/) — and a hard line break
is just a newline. is just a newline.
``` ```
## A section heading ## A section heading
### A subsection ### A subsection
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). Plain URLs become links automatically: https://example.com — and a hard line break *Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). An image that links to its page: [![xkcd 1179: ISO 8601](https://imgs.xkcd.com/comics/iso_8601.png "xkcd 1179: ISO 8601"){width=240}](https://xkcd.com/1179/) — and a hard line break
is just a newline. is just a newline.
## Lists and quotes ## Lists and quotes
@@ -354,6 +355,7 @@ This site runs on **Pagerite**: FastAPI + html5tagger + kanta, with content writ
- [How to edit this site](/docs/editing) - [How to edit this site](/docs/editing)
- [Markdown features](/docs/markdown/basics) - [Markdown features](/docs/markdown/basics)
- [The showcase](/showcase/gallery) - [The showcase](/showcase/gallery)
- [![xkcd 2347: Dependency](https://imgs.xkcd.com/comics/dependency.png "xkcd 2347: Dependency"){width=240}](https://xkcd.com/2347/) — a small comic about small dependencies
*Replace this page with whatever your site is about.* *Replace this page with whatever your site is about.*
""" """
-5
View File
@@ -3,11 +3,6 @@
SVG from the active palette (var(--accent)), so one SVG serves light SVG from the active palette (var(--accent)), so one SVG serves light
and dark — and other themes too. */ and dark — and other themes too. */
#banner {
min-height: 15rem;
border-bottom: none;
}
/* Artwork colors, light mode */ /* Artwork colors, light mode */
.cb-bg0 { .cb-bg0 {
stop-color: #ffffff; stop-color: #ffffff;
+2 -11
View File
@@ -16,6 +16,7 @@
--line: #12203f14; --line: #12203f14;
--font-body: var(--font-inter); --font-body: var(--font-inter);
--font-heading: var(--font-montserrat); --font-heading: var(--font-montserrat);
--code-x-height: 0.546; /* Inter's x-height ratio */
} }
@media (prefers-color-scheme: dark) { @media (prefers-color-scheme: dark) {
@@ -32,18 +33,13 @@
} }
} }
::selection {
background: var(--accent);
color: #fff;
}
/* Genuinely large solid brand with a soft blue shadow overlapping the /* Genuinely large solid brand with a soft blue shadow overlapping the
artwork — conservative, but unmissable. */ artwork — conservative, but unmissable. */
#brand { #brand {
font-size: clamp(4rem, 11vw, 8.5rem); font-size: clamp(4rem, 11vw, 8.5rem);
font-weight: 800; font-weight: 800;
letter-spacing: -0.04em; letter-spacing: -0.04em;
line-height: 1; line-height: 1.1;
white-space: nowrap; white-space: nowrap;
color: var(--accent2); color: var(--accent2);
filter: drop-shadow(0 0.4rem 1.4rem rgb(10 92 255 / 0.3)); filter: drop-shadow(0 0.4rem 1.4rem rgb(10 92 255 / 0.3));
@@ -112,7 +108,6 @@ article h3 {
font-weight: 700; font-weight: 700;
font-size: 0.95rem; font-size: 0.95rem;
letter-spacing: 0.08em; letter-spacing: 0.08em;
text-transform: uppercase;
color: var(--muted); color: var(--muted);
} }
@@ -140,7 +135,3 @@ pre {
img { img {
border-radius: 4px; border-radius: 4px;
} }
::view-transition {
background: var(--bg);
}
+9 -3
View File
@@ -1,7 +1,13 @@
/* Eyes banner design: a canvas critter watching the cursor from the /* Eyes banner design: a canvas critter watching the cursor from the
grass (banner.html — markup + styles + script inlined by the backend grass (banner.html — markup + styles + script inlined by the backend
into #page-banner). Fixed-height stage matching the canvas. */ into #page-banner). The canvas fills the banner; the scene composes
against the 13rem layout box and extends flat grass into any overflow
a theme adds below (summer's cross-fade strip). */
#banner { /* Opt out of the base parallax (scale overscan + --pry drift): it lands on
height: 240px; the design wrapper the backend puts around banner.html, and scaling from
the bottom edge would crop the top of the composed scene — pushing the
critter out of view — while this canvas animates on its own. */
#page-banner>[data-design="eyes"] {
transform: none;
} }
+78 -19
View File
@@ -2,7 +2,13 @@
<style> <style>
#eyes { #eyes {
width: 100%; width: 100%;
height: 240px; /* 13rem — the banner's layout height, correct from the first frame,
before any external stylesheet has sized #page-banner. Themes that
extend the banner past the layout box (summer's overflow fade) raise
--eyes-h to 100% so the canvas follows the taller stage; the script
still composes the scene against the 13rem box and only extends the
meadow, so the scene itself never shifts. */
height: var(--eyes-h, 13rem);
display: block; display: block;
} }
</style> </style>
@@ -10,32 +16,51 @@
(() => { (() => {
const c = document.getElementById('eyes') const c = document.getElementById('eyes')
const ctx = c.getContext('2d') const ctx = c.getContext('2d')
const DPR = devicePixelRatio || 1
const fit = () => { // Sync the backing store to the canvas' laid-out size. Checked every
const w = Math.max(1, c.clientWidth) // frame: this inline script runs before the stylesheets that size
const h = Math.max(1, c.clientHeight) // #page-banner, so observers/load events can still miss the transition.
c.width = Math.round(w * DPR) // Assigning width/height also clears the canvas. DPR is read here, not
c.height = Math.round(h * DPR) // captured: it changes with browser zoom.
const syncSize = () => {
const DPR = devicePixelRatio || 1
const w = Math.round(Math.max(1, c.clientWidth) * DPR)
const h = Math.round(Math.max(1, c.clientHeight) * DPR)
if (c.width !== w || c.height !== h) {
c.width = w
c.height = h
}
ctx.setTransform(DPR, 0, 0, DPR, 0, 0) ctx.setTransform(DPR, 0, 0, DPR, 0, 0)
} }
fit()
addEventListener('resize', fit)
let mx = 0 let mx = 0
let my = 0 let my = 0
let lastMove = 0 let lastMove = 0
addEventListener('mousemove', e => { // Mouse and touch tracked with separate listeners (pointer events arrive
// too late on some mobile browsers). Passive listeners: a drag on the
// banner still scrolls the page — on browsers that stop delivering
// touchmove once scrolling takes over, the gaze just follows until then.
const track = (x, y) => {
const r = c.getBoundingClientRect() const r = c.getBoundingClientRect()
// Convert viewport coordinates into the canvas' CSS-pixel coordinate // Convert viewport coordinates into the canvas' CSS-pixel coordinate
// system. This remains correct with browser zoom, CSS transforms, etc. // system. This remains correct with browser zoom, CSS transforms, etc.
mx = (e.clientX - r.left) * c.clientWidth / r.width mx = (x - r.left) * c.clientWidth / r.width
my = (e.clientY - r.top) * c.clientHeight / r.height my = (y - r.top) * c.clientHeight / r.height
lastMove = performance.now() lastMove = performance.now()
}) }
addEventListener('mousemove', e => track(e.clientX, e.clientY))
const trackTouch = e => {
const t = e.touches[0]
if (t) track(t.clientX, t.clientY)
}
addEventListener('touchstart', trackTouch, { passive: true })
addEventListener('touchmove', trackTouch, { passive: true })
let gx = 0.5 let gx = 0.5
let gy = 0.5 let gy = 0.5
@@ -54,7 +79,7 @@
] ]
const ridgeY = (x, w, h) => const ridgeY = (x, w, h) =>
h * 0.72 + h * 0.83 +
Math.sin(x * 0.012) * 10 + Math.sin(x * 0.012) * 10 +
Math.sin(x * 0.003 + 1.4) * 16 + Math.sin(x * 0.003 + 1.4) * 16 +
Math.sin(x * 0.02 + 0.7) * 3 Math.sin(x * 0.02 + 0.7) * 3
@@ -114,7 +139,7 @@
for (let i = 0; i < 5; i++) { for (let i = 0; i < 5; i++) {
const x = (i + 0.5) * w / 5 const x = (i + 0.5) * w / 5
const y = h * 0.58 + Math.sin(i * 1.7) * 8 const y = h * 0.69 + Math.sin(i * 1.7) * 8
ctx.fillStyle = '#5f8d4e' ctx.fillStyle = '#5f8d4e'
ctx.beginPath() ctx.beginPath()
@@ -307,7 +332,10 @@
ctx.moveTo(0, h) ctx.moveTo(0, h)
ctx.lineTo(0, ridgeY(0, w, h)) ctx.lineTo(0, ridgeY(0, w, h))
for (let x = 0; x <= w; x += 8) // Sample one step PAST the right edge (x <= w + 8): stopping at w would
// leave the path closing with a visible vertical drop at the edge
// whenever the width isn't a multiple of the 8px step.
for (let x = 0; x <= w + 8; x += 8)
ctx.lineTo(x, ridgeY(x, w, h)) ctx.lineTo(x, ridgeY(x, w, h))
ctx.lineTo(w, h) ctx.lineTo(w, h)
@@ -319,7 +347,7 @@
ctx.moveTo(0, h) ctx.moveTo(0, h)
ctx.lineTo(0, ridgeY(0, w, h) + 10) ctx.lineTo(0, ridgeY(0, w, h) + 10)
for (let x = 0; x <= w; x += 8) for (let x = 0; x <= w + 8; x += 8)
ctx.lineTo(x, ridgeY(x, w, h) + 10) ctx.lineTo(x, ridgeY(x, w, h) + 10)
ctx.lineTo(w, h) ctx.lineTo(w, h)
@@ -362,8 +390,17 @@
const dt = Math.min(now - prev, 100) / 16.7 const dt = Math.min(now - prev, 100) / 16.7
prev = now prev = now
syncSize()
const w = c.clientWidth const w = c.clientWidth
const h = c.clientHeight // Compose the scene against the banner's layout box, not the canvas:
// a theme may extend #page-banner past #banner (e.g. summer overflows
// the artwork into the page for a masked cross-fade), and the critter
// must stay in the visible part. The overflow strip is filled with the
// flat meadow color below — a hard canvas edge would show through the
// fade, a grass extension just blends.
const h = Math.min(c.clientHeight,
c.closest('#banner')?.clientHeight || c.clientHeight)
const R = Math.min(h * 0.11, 38) const R = Math.min(h * 0.11, 38)
if (now > nextMove && !hidePhase) { if (now > nextMove && !hidePhase) {
@@ -399,6 +436,16 @@
vy *= Math.pow(0.85, dt) vy *= Math.pow(0.85, dt)
yoff += vy * dt yoff += vy * dt
// Clip the scene to the composed area: the ducking critter travels
// below it, and the overflow strip is only flat meadow painted after —
// without the clip the critter would leave trails there as it sinks.
// Clipping against grass-on-grass is invisible, so the duck still
// reads as sinking into the meadow.
ctx.save()
ctx.beginPath()
ctx.rect(0, 0, w, h)
ctx.clip()
drawBackground(w, h) drawBackground(w, h)
const cx0 = gx * w const cx0 = gx * w
@@ -415,6 +462,18 @@
drawCritter(cx0, eyeY, R, now, dt) drawCritter(cx0, eyeY, R, now, dt)
drawForeground(w, h) drawForeground(w, h)
ctx.restore()
// Extend the meadow into any overflow below the composed scene, with
// the same two layers drawForeground leaves at the bottom (base grass
// plus the dark under-band) so the joint is invisible.
if (c.clientHeight > h) {
ctx.fillStyle = '#69ae4b'
ctx.fillRect(0, h, w, c.clientHeight - h)
ctx.fillStyle = 'rgba(48,102,34,0.18)'
ctx.fillRect(0, h, w, c.clientHeight - h)
}
requestAnimationFrame(frame) requestAnimationFrame(frame)
} }
+1 -7
View File
@@ -2,13 +2,6 @@
(inlined by the backend into #page-banner), in neutral dark greys that (inlined by the backend into #page-banner), in neutral dark greys that
follow the page's color scheme. */ follow the page's color scheme. */
/* Bezier-swept banner with wide orange stripes (inlined SVG), separated
from the page by a straight orange blade. */
#banner {
height: 13rem;
border-bottom: 4px solid var(--accent);
}
/* Banner artwork dark tones: neutral greys in light mode (retinted to the /* Banner artwork dark tones: neutral greys in light mode (retinted to the
page's violet family by the dark-scheme block below). */ page's violet family by the dark-scheme block below). */
.nb-base { .nb-base {
@@ -44,6 +37,7 @@
} }
@media (prefers-color-scheme: dark) { @media (prefers-color-scheme: dark) {
/* Banner dark tones tinted to the same violet family as the page. */ /* Banner dark tones tinted to the same violet family as the page. */
.nb-base { .nb-base {
fill: #100d18; fill: #100d18;
+10 -12
View File
@@ -41,6 +41,7 @@
--font-body: var(--font-montserrat); --font-body: var(--font-montserrat);
--font-heading: var(--font-literata); --font-heading: var(--font-literata);
--code-x-height: 0.517; /* Montserrat's x-height ratio */
} }
/* Dark scheme: same identity, but the page goes deep violet (never muddy /* Dark scheme: same identity, but the page goes deep violet (never muddy
@@ -63,13 +64,19 @@
::selection { ::selection {
background: var(--accent); background: var(--accent);
color: var(--ink); }
/* Any banner used is separated from page by a thick orange line */
#banner {
border-bottom: 4px solid var(--accent);
} }
/* Oversized outlined brand, spilling off the banner edge: orange stroke, /* Oversized outlined brand, spilling off the banner edge: orange stroke,
solid black fill. */ solid black fill. */
#brand { #brand {
font-size: 10rem; /* Scales down proportionally below ~1000px: 10rem at a 62.5rem viewport,
shrinking with vmin (smaller of viewport width/height) below that. */
font-size: clamp(2.5rem, 16vmin, 10rem);
line-height: 1.2; line-height: 1.2;
font-weight: 700; font-weight: 700;
letter-spacing: 0.04em; letter-spacing: 0.04em;
@@ -131,14 +138,9 @@
/* Console-style headings: uppercase monospace. h1 in the page text color /* Console-style headings: uppercase monospace. h1 in the page text color
with a hazard-stripe underline, h2 deep orange, h3 cyan. */ with a hazard-stripe underline, h2 deep orange, h3 cyan. */
article h1, article h1 {
article h2,
article h3 {
text-transform: uppercase; text-transform: uppercase;
letter-spacing: 0.02em; letter-spacing: 0.02em;
}
article h1 {
color: var(--text); color: var(--text);
font-weight: 700; font-weight: 700;
padding-bottom: 0.5rem; padding-bottom: 0.5rem;
@@ -206,7 +208,3 @@ pre {
img { img {
border-radius: 3px; border-radius: 3px;
} }
::view-transition {
background: var(--bg);
}
-4
View File
@@ -15,7 +15,3 @@
.banner-fade { .banner-fade {
stop-color: var(--bg); stop-color: var(--bg);
} }
#banner {
min-height: 13rem;
}
+1 -9
View File
@@ -17,11 +17,7 @@
--line: #ffffff1c; --line: #ffffff1c;
--font-body: var(--font-literata); --font-body: var(--font-literata);
--font-heading: var(--font-fraunces); --font-heading: var(--font-fraunces);
} --code-x-height: 0.507; /* Literata's x-height ratio */
::selection {
background: var(--accent2);
color: #fff;
} }
/* Oversized tilted brand in the sky→violet gradient. */ /* Oversized tilted brand in the sky→violet gradient. */
@@ -91,7 +87,3 @@ blockquote {
pre { pre {
--code-bg: var(--surface); --code-bg: var(--surface);
} }
::view-transition {
background: #000;
}
+6 -6
View File
@@ -1,7 +1,7 @@
/* Stars banner design: a drifting starfield (banner.html — canvas + script /* Stars banner design: a drifting starfield (banner.html — canvas + script
inlined by the backend into #page-banner). Fixed-height stage matching inlined by the backend into #page-banner). The canvas takes the banner's
the canvas. */ 13rem layout height directly so it never renders unclipped before the
main stylesheet loads; the starfield scales to any height. The design's
#banner { opt-out of theme banner overflow/fade effects also lives in banner.html's
height: 240px; inline <style>, so it applies atomically with the markup (this file would
} load a beat later and let the overflow flash through mid-transition). */
+31 -10
View File
@@ -2,27 +2,47 @@
<style> <style>
#stars { #stars {
width: 100%; width: 100%;
height: 240px; /* 13rem — the banner's layout height, not 100%: a percentage only
resolves after the main stylesheet sizes #page-banner, and until
then the canvas would render at its intrinsic height, unclipped,
over the page. (This design always opts out of theme overflows —
below — so the layout height is always the right one.) */
height: 13rem;
display: block; display: block;
} }
/* The night sky stays a windowed stage: undo the summer theme's banner
overflow/cross-fade — a starfield must not bleed into a daylit page.
Kept in this inlined <style> (not banner.css) so the opt-out applies
atomically with the markup; a separate stylesheet can arrive a beat
later and let the overflow flash through mid-transition. Later in
document order than theme.css, so same-specificity rules win. */
#page-banner {
inset: 0;
mask-image: none;
}
</style> </style>
<script><!-- <script><!--
(() => { (() => {
const c = document.getElementById('stars') const c = document.getElementById('stars')
const ctx = c.getContext('2d') const ctx = c.getContext('2d')
const DPR = devicePixelRatio || 1
const fit = () => { // Sync the backing store to the canvas' laid-out size. Checked every
const w = Math.max(1, c.clientWidth) // frame: this inline script runs before the stylesheets that size
const h = Math.max(1, c.clientHeight) // #page-banner, so observers/load events can still miss the transition.
c.width = Math.round(w * DPR) // Assigning width/height also clears the canvas. DPR is read here, not
c.height = Math.round(h * DPR) // captured: it changes with browser zoom.
const syncSize = () => {
const DPR = devicePixelRatio || 1
const w = Math.round(Math.max(1, c.clientWidth) * DPR)
const h = Math.round(Math.max(1, c.clientHeight) * DPR)
if (c.width !== w || c.height !== h) {
c.width = w
c.height = h
}
ctx.setTransform(DPR, 0, 0, DPR, 0, 0) ctx.setTransform(DPR, 0, 0, DPR, 0, 0)
} }
fit()
addEventListener('resize', fit)
const stars = Array.from({ length: 110 }, () => ({ const stars = Array.from({ length: 110 }, () => ({
x: Math.random(), x: Math.random(),
y: Math.random(), y: Math.random(),
@@ -33,6 +53,7 @@
let prev = performance.now() let prev = performance.now()
;(function frame(now) { ;(function frame(now) {
if (!c.isConnected) return if (!c.isConnected) return
syncSize()
const w = c.clientWidth const w = c.clientWidth
const h = c.clientHeight const h = c.clientHeight
const dt = Math.min(now - prev, 100) const dt = Math.min(now - prev, 100)
+40 -15
View File
@@ -2,10 +2,6 @@
backend into #page-banner) — rolling hills, leafy bushes, swaying backend into #page-banner) — rolling hills, leafy bushes, swaying
flowers, drifting clouds and a sun that rises as you scroll. */ flowers, drifting clouds and a sun that rises as you scroll. */
#banner {
min-height: 15rem;
}
/* The artwork fades into the page background at its bottom edge. */ /* The artwork fades into the page background at its bottom edge. */
.summer-fade { .summer-fade {
stop-color: var(--bg); stop-color: var(--bg);
@@ -44,9 +40,17 @@
animation: summer-sun 9s ease-in-out infinite alternate; animation: summer-sun 9s ease-in-out infinite alternate;
} }
#page-banner .cloud-a { animation: summer-drift 56s ease-in-out infinite alternate; } #page-banner .cloud-a {
#page-banner .cloud-b { animation: summer-drift 73s ease-in-out infinite alternate-reverse; } animation: summer-drift 56s ease-in-out infinite alternate;
#page-banner .cloud-c { animation: summer-drift 64s ease-in-out infinite alternate; } }
#page-banner .cloud-b {
animation: summer-drift 73s ease-in-out infinite alternate-reverse;
}
#page-banner .cloud-c {
animation: summer-drift 64s ease-in-out infinite alternate;
}
#page-banner .flower { #page-banner .flower {
transform-box: fill-box; transform-box: fill-box;
@@ -55,21 +59,42 @@
} }
/* Stagger the sway so the meadow doesn't move in lockstep. */ /* Stagger the sway so the meadow doesn't move in lockstep. */
#page-banner .flower:nth-child(3n) { animation-delay: -1.7s; } #page-banner .flower:nth-child(3n) {
#page-banner .flower:nth-child(3n + 1) { animation-delay: -3.1s; animation-duration: 6s; } animation-delay: -1.7s;
}
#page-banner .flower:nth-child(3n + 1) {
animation-delay: -3.1s;
animation-duration: 6s;
}
} }
@keyframes summer-sun { @keyframes summer-sun {
from { opacity: 0.22; } from {
to { opacity: 0.38; } opacity: 0.22;
}
to {
opacity: 0.38;
}
} }
@keyframes summer-drift { @keyframes summer-drift {
from { transform: translateX(-1.6%); } from {
to { transform: translateX(1.6%); } transform: translateX(-1.6%);
}
to {
transform: translateX(1.6%);
}
} }
@keyframes summer-sway { @keyframes summer-sway {
from { transform: rotate(-2.6deg); } from {
to { transform: rotate(2.6deg); } transform: rotate(-2.6deg);
}
to {
transform: rotate(2.6deg);
}
} }
+44 -11
View File
@@ -8,6 +8,7 @@
/* Every color is sampled (or text-darkened) from banner.svg. */ /* Every color is sampled (or text-darkened) from banner.svg. */
--bg: #e6f4cf; --bg: #e6f4cf;
--sky: #d9f1ff;
/* pale meadow — the banner fades into this */ /* pale meadow — the banner fades into this */
--surface: #f8fbf0; --surface: #f8fbf0;
--text: #2c4a2f; --text: #2c4a2f;
@@ -28,6 +29,7 @@
--font-body: var(--font-cause); --font-body: var(--font-cause);
--font-heading: var(--font-new-rocker); --font-heading: var(--font-new-rocker);
--font-brand: var(--font-cause); --font-brand: var(--font-cause);
--code-x-height: 0.5; /* Cause's x-height ratio */
} }
/* The page is the same landscape the banner paints: hazy sky light up top /* The page is the same landscape the banner paints: hazy sky light up top
@@ -56,15 +58,33 @@ body {
color: var(--text); color: var(--text);
} }
::selection { /* No separation between banner and page: the frame loses its background,
background: var(--sun); border and shadow, and the artwork overflows into the document below,
color: #3d5223; masked by a transparency gradient so it cross-fades into the body's fixed
meadow gradient. A plain color match is impossible — both sides are
gradients, and the parallax drift keeps the banner side moving. */
#banner {
background: none;
border-bottom: 0;
box-shadow: none;
} }
/* The banner frame ties into the meadow beneath it. */ #page-banner {
#banner { inset: 0 0 -6rem;
border-bottom-color: #4f913b30; /* Paints over the page background in the overlap, but never swallows its
box-shadow: 0 0.3rem 1rem #4f913b22; clicks. */
pointer-events: none;
mask-image: linear-gradient(180deg, #000 calc(100% - 6rem), transparent);
/* Canvas banner designs (eyes) default to the 13rem layout height; here
they must fill the taller overflow stage so their meadow extension
reaches through the fade strip. */
--eyes-h: 100%;
}
/* The mask above replaced the SVG's flat-color bottom fade (meadowFade) —
fading to a single --bg would reintroduce a seam against the gradient. */
#page-banner .summer-fade {
stop-opacity: 0;
} }
/* Cheerful oversized tilted brand in a sky→grass→flower gradient. */ /* Cheerful oversized tilted brand in a sky→grass→flower gradient. */
@@ -114,14 +134,26 @@ body {
} }
/* Content sits in a soft wash of sunlit meadow rather than on a white /* Content sits in a soft wash of sunlit meadow rather than on a white
card. */ card. The wash lives on a pseudo-element so it can carry a vertical
mask: it fades in from transparent over the strip where the banner
artwork overflows into the page, so the two never fight — no matter
which side paints on top. */
main { main {
position: relative;
}
main::before {
content: "";
position: absolute;
inset: 0;
z-index: -1;
background: linear-gradient(90deg, background: linear-gradient(90deg,
transparent, transparent,
#f0f8dc80 15%, #f0f8dc80 15%,
#f0f8dc8c 50%, #f0f8dc8c 50%,
#f0f8dc80 85%, #f0f8dc80 85%,
transparent); transparent);
mask-image: linear-gradient(180deg, transparent, #000 6rem);
} }
/* The sidebar is a piece of the same meadow: glassy green with a light /* The sidebar is a piece of the same meadow: glassy green with a light
@@ -131,6 +163,8 @@ main {
border-right: 1px solid #ffffff80; border-right: 1px solid #ffffff80;
border-bottom: 1px solid var(--line); border-bottom: 1px solid var(--line);
box-shadow: 0 0.3rem 1rem #4f913b1f; box-shadow: 0 0.3rem 1rem #4f913b1f;
border-radius: 1rem;
margin-inline: 0.5rem;
} }
#sidebar a { #sidebar a {
@@ -260,8 +294,7 @@ figcaption,
background: linear-gradient(160deg, #eef7dd, #d9eec5); background: linear-gradient(160deg, #eef7dd, #d9eec5);
} }
/* The page transition exposes summer color around the rotating
snapshots. */
::view-transition { ::view-transition {
background: var(--bg); /* Override to mid sky shade instead of the darker --bg we otherwise get */
background: var(--sky);
} }
+277 -28
View File
@@ -15,8 +15,10 @@ own URL renders a placeholder page (render_category).
""" """
from pathlib import Path from pathlib import Path
from html import unescape
import json import json
import os import os
import re
from html5tagger import HTML, Document, E, Template from html5tagger import HTML, Document, E, Template
@@ -108,56 +110,139 @@ def _editor_css_url(vite_url: str | None) -> str | None:
return None return None
def _inline_asset(url: str) -> str:
"""Read a served asset's content for inlining into the page (prod only).
Handles build assets (``/_assets/...`` from the Vite build) and theme
files (``/_themes/{name}/...`` from pagerite/themes/).
"""
if url.startswith("/_themes/"):
name, _, file = url.removeprefix("/_themes/").partition("/")
if _valid_name(name) and _valid_name(file):
return (THEMES / name / file).read_text()
raise ValueError(f"not a theme asset: {url}")
return (BUILD / url.lstrip("/")).read_text()
def _inline_script(url: str) -> str:
"""Read a built JS bundle for inlining (prod only).
Inline modules resolve relative imports against the document URL, not
the bundle's directory, so rewrite the build's relative chunk
specifiers ("./chunk.js") to absolute /_assets/ paths.
"""
js = _inline_asset(url)
for chunk in _manifest().values():
file = chunk.get("file", "")
if file.endswith(".js"):
js = js.replace(f'"./{file.rsplit("/", 1)[-1]}"', f'"/{file}"')
return js
def _layout( def _layout(
modules: list[str] = (), modules: list[str] = (),
stylesheets: list[str] = (),
custom_css: str = "", custom_css: str = "",
theme: str = "", theme: str = "",
banner_design: str = "", banner_design: str = "",
favicon: str = "", favicon: str = "",
social: dict[str, str] | None = None,
) -> Template: ) -> Template:
"""Page layout template with standard asset URLs and ES-module scripts. """Page layout template with standard assets and ES-module scripts.
Stylesheets use ``blocking="render"`` so the browser waits for them before In dev (PAGERITE_VITE_URL set) assets are linked from the Vite dev
showing the page, avoiding a flash of unstyled content. Order matters and server and stylesheets use ``blocking="render"`` so the browser waits
is fixed: base (Vite build, absent in dev where Vite injects it from JS), for them before showing the page, avoiding a flash of unstyled content.
theme and banner design (backend-served from pagerite/themes/), then the In production all page assets are inlined into the document: stylesheets
user's custom CSS last so it always wins. become ``<style>`` elements and module scripts inline ``<script>``s, so
a page loads with no asset round trips. The on-demand bundles (editor,
analytics) stay external in both modes.
Order matters and is fixed: base (Vite build, absent in dev where Vite
injects it from JS), theme and banner design (from pagerite/themes/),
entry-specific stylesheets (e.g. overlayscrollbars.css), then the user's
custom CSS last so it always wins.
In dev, pagerite.js re-appends the backend-rendered theme/design links In dev, pagerite.js re-appends the backend-rendered theme/design links
(and the custom CSS) after the Vite-injected base styles, keeping this (and the custom CSS) after the Vite-injected base styles, keeping this
order intact. order intact.
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
property attributes, everything else (description, twitter:*) as name.
""" """
doc = Document(E.Title, lang="en") doc = Document(E.Title, lang="en")
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs # Responsive layout (see the 48rem breakpoint in pagerite.css) needs
# the real device width, not the default 980px layout viewport. # the real device width, not the default 980px layout viewport.
doc.meta(name="viewport", content="width=device-width, initial-scale=1") doc.meta(name="viewport", content="width=device-width, initial-scale=1")
for key, value in (social or {}).items():
if value:
if key.startswith(("og:", "article:")):
doc.meta(property=key, content=value)
elif key == "canonical":
doc.link(rel="canonical", href=value)
else:
doc.meta(name=key, content=value)
# A custom favicon (from the site editor) is linked explicitly; without # A custom favicon (from the site editor) is linked explicitly; without
# one, browsers fall back to the build's /favicon.ico by convention. # one, browsers fall back to the build's /favicon.ico by convention.
if favicon: if favicon:
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon") doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
# Editor asset URLs for pagerite.js, which injects the 🖊️ edit pens # Asset URLs for the on-demand bundles (editor, analytics) for
# itself once it has validated the session (pages render identically # pagerite.js, which injects the 🖊️ edit pens itself once it has
# for everyone; editing is gated by the auth proxy in front of /_api). # validated the session (pages render identically for everyone; editing
script, editor_css = _editor_assets() # is gated by the auth proxy in front of /_api). Dev passes the Vite
doc.meta(name="pagerite:editor-src", content=script[-1]) # dev-server URLs as meta tags (Vite serves the modules and injects
if editor_css: # their CSS for hot reloads); production inlines all page assets and
doc.meta(name="pagerite:editor-css", content=editor_css) # carries the on-demand URLs in one JSON script instead.
# Stylesheet links carry stable ids so the site editor's hot swap can
# keep each sheet at its rendered position (see swapRegions).
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = os.environ.get("PAGERITE_VITE_URL")
editor_scripts, editor_css = _editor_assets()
config = {
"pagerite:editor-src": editor_scripts[-1],
"pagerite:analytics-src": _analytics_assets()[0][0],
}
if editor_css:
config["pagerite:editor-css"] = editor_css
if vite_url:
for key, value in config.items():
doc.meta(name=key, content=value)
else:
# Inert JSON script; URLs never contain "</", but stay safe.
doc.script(
HTML(json.dumps(config).replace("</", "<\\/")),
type="application/json",
id="pagerite-assets",
)
# Stylesheets carry stable ids so the fetch-navigation and the site
# editor's hot swap can sync <head> positionally (see swapdoc.js).
# Production inlines the CSS as <style> elements: one less round trip
# per sheet, and fetch-navigation can carry them across swaps whole.
sheets = [ sheets = [
("pagerite-base", _base_css_url(vite_url)), ("pagerite-base", _base_css_url(vite_url)),
("pagerite-theme", _theme_css_url(theme)), ("pagerite-theme", _theme_css_url(theme)),
("pagerite-banner", _banner_css_url(banner_design)), ("pagerite-banner", _banner_css_url(banner_design)),
] ]
for id_, url in sheets: for id_, url in sheets:
if url: if not url:
continue
if vite_url:
doc.link(rel="stylesheet", href=url, blocking="render", id=id_) doc.link(rel="stylesheet", href=url, blocking="render", id=id_)
else:
doc.style(HTML(_inline_asset(url)), id=id_)
for url in stylesheets:
if vite_url:
doc.link(rel="stylesheet", href=url, blocking="render")
else:
# Id from the file stem minus the content hash, so the head
# sync can match sheets across pages (e.g. the analytics sheet
# exists on /_a only and is added/removed on swaps).
stem = url.rsplit("/", 1)[-1].removesuffix(".css")
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
doc.style(HTML(_inline_asset(url)), id=f"pagerite-css-{name}")
for src in modules: for src in modules:
doc.script(src=src, type="module") if vite_url:
doc.script(src=src, type="module")
if custom_css.strip(): if custom_css.strip():
doc.style(custom_css, id="pagerite-user") doc.style(custom_css, id="pagerite-user")
return Template( body = (
doc doc
.header( .header(
E.div(E.Banner, id="page-banner"), E.div(E.Banner, id="page-banner"),
@@ -170,8 +255,23 @@ def _layout(
E.main(E.Main, id="main"), E.main(E.Main, id="main"),
id="content", id="content",
) )
.footer(None), # kept empty for now; zero-height (see pagerite.css) .footer(None) # kept empty for now; zero-height (see pagerite.css)
) )
if not vite_url:
# Inline the bundles at the end of the body: module scripts are
# deferred anyway, and the page can render before they execute.
# Escape "</script" so it cannot terminate the element early (only
# ever occurs inside string literals, where the backslash escape is
# a no-op).
for src in modules:
js = re.sub(r"</script", r"<\\/script", _inline_script(src), flags=re.I)
# Stable id from the file stem minus the content hash; the
# analytics page's script (pagerite-js-analytics) is found and
# re-created by pagerite.js on fetch-navigations to /_a.
stem = src.rsplit("/", 1)[-1].removesuffix(".js")
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
body.script(HTML(js), type="module", id=f"pagerite-js-{name}")
return Template(body)
def _brand_link(brand: str, brand_html: str = "") -> HTML: def _brand_link(brand: str, brand_html: str = "") -> HTML:
@@ -426,6 +526,87 @@ def page_content(menu: dict[str, Node], path: str) -> HTML:
return HTML(str(doc)) return HTML(str(doc))
_FIRST_P = re.compile(r"<p[^>]*>(.*?)</p>", re.S)
_TAG = re.compile(r"<[^>]+>")
_IMG_TAG = re.compile(r"<img\b[^>]*>")
_VIDEO_TAG = re.compile(r"<video\b[^>]*>")
_ATTR_SRC = re.compile(r'src="([^"]+)"')
_ATTR_CLASS = re.compile(r'class="([^"]*)"')
def _share_media(html: str, base_url: str) -> tuple[str, str]:
"""(image, video) share URLs from the rendered article.
Image preference: an image with class "hero" (author override, may
appear anywhere in the article), then the first raster image (SVGs
rasterize poorly or not at all on many social scrapers), then the
first SVG. Video: the first <video> — og:video is in the OGP spec and
honored mainly by Facebook; X/Twitter ignores it. Absolute URLs are
built from the request base, scrapers cannot use relative ones.
"""
if not base_url:
return "", ""
def absolute(src: str) -> str:
src = unescape(src)
return src if src.startswith(("http://", "https://")) else f"{base_url}{src}"
hero = raster = svg = video = ""
for tag in _IMG_TAG.findall(html):
if not (src := _ATTR_SRC.search(tag)):
continue
src = src.group(1)
cls = _ATTR_CLASS.search(tag)
if cls and "hero" in cls.group(1).split():
hero = src
break
if src.lower().split("?")[0].endswith(".svg"):
svg = svg or src
else:
raster = raster or src
# Keep scanning: a later hero still wins.
for tag in _VIDEO_TAG.findall(html):
if m := _ATTR_SRC.search(tag):
video = m.group(1)
break
image = hero or raster or svg
return (absolute(image) if image else "", absolute(video) if video else "")
def _social_meta(
node: Node, path: str, title: str, html: str, brand: str, base_url: str,
) -> dict[str, str]:
"""Open Graph/Twitter/SEO meta tags for a content page.
Heuristics over the rendered article: the description is the first
paragraph's text (truncated at ~200 chars on a word boundary), the
share image the article's first <img> — authors lead with their most
representative figure. Absolute URLs are built from the request's base
(social scrapers cannot use relative ones).
"""
url = f"{base_url}/{path}" if base_url else ""
m = _FIRST_P.search(html)
text = unescape(_TAG.sub("", m.group(1) if m else ""))
text = " ".join(text.split())
if len(text) > 200:
text = text[:200].rsplit(" ", 1)[0] + ""
image, video = _share_media(html, base_url)
return {
"description": text,
"canonical": url,
"og:type": "article",
"og:title": title,
"og:description": text,
"og:url": url,
"og:site_name": brand,
"og:image": image,
"og:video": video,
"article:published_time": node.created.isoformat(),
"article:modified_time": node.modified.isoformat(),
"twitter:card": "summary_large_image" if image else "summary",
}
def render_page( def render_page(
menu: dict[str, Node], menu: dict[str, Node],
path: str, path: str,
@@ -434,18 +615,24 @@ def render_page(
theme: str = "", theme: str = "",
favicon: str = "", favicon: str = "",
brand_html: str = "", brand_html: str = "",
base_url: str = "",
) -> str: ) -> str:
"""Render a full HTML page for the slug path.""" """Render a full HTML page for the slug path."""
node = resolve(menu, path)[-1] node = resolve(menu, path)[-1]
title = _title(path.rpartition("/")[2], node) title = _title(path.rpartition("/")[2], node)
main = page_content(menu, path)
social = _social_meta(node, path, title, str(main), brand, base_url)
return str( return str(
_layout(_page_assets(), custom_css, theme, banner_design(menu, path, theme), favicon)( _layout(
*_page_assets(), custom_css, theme, banner_design(menu, path, theme),
favicon, social,
)(
Title=f"{title} {brand}" if brand else title, Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html),
Nav=nav_html(menu, path), Nav=nav_html(menu, path),
Sidebar=sidebar_html(menu, path), Sidebar=sidebar_html(menu, path),
Banner=banner_html(menu, path, theme), Banner=banner_html(menu, path, theme),
Main=page_content(menu, path), Main=main,
), ),
) )
@@ -476,7 +663,7 @@ def render_category(
else: else:
doc.p("This section has no page of its own yet.") doc.p("This section has no page of its own yet.")
return str( return str(
_layout(_page_assets(), custom_css, theme, banner_design(menu, path, theme), favicon)( _layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), favicon)(
Title=f"{title} {brand}" if brand else title, Title=f"{title} {brand}" if brand else title,
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html),
Nav=nav_html(menu, path), Nav=nav_html(menu, path),
@@ -502,7 +689,7 @@ def render_not_found(
doc.h1("Not Found") doc.h1("Not Found")
doc.p(f"No article at /{path}. If there was before, it may have been deleted.") doc.p(f"No article at /{path}. If there was before, it may have been deleted.")
return str( return str(
_layout(_page_assets(), custom_css, theme, banner_design(menu, path, theme), favicon)( _layout(*_page_assets(), custom_css, theme, banner_design(menu, path, theme), favicon)(
Title=f"Not Found {brand}" if brand else "Not Found", Title=f"Not Found {brand}" if brand else "Not Found",
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html),
Nav=nav_html(menu, path), Nav=nav_html(menu, path),
@@ -513,19 +700,23 @@ def render_not_found(
) )
def _page_assets() -> list[str]: def _page_assets() -> tuple[list[str], list[str]]:
"""Script URLs for public pages (pagerite entry). """Script and stylesheet URLs for public pages (pagerite entry).
Dev mode loads the entry from the Vite dev server; production uses Dev mode loads the entry from the Vite dev server; production uses
the Vite build manifest to resolve the hashed asset names. the Vite build manifest to resolve the hashed asset names. CSS imported
by the entry (e.g. overlayscrollbars.css) is extracted by Vite and must
be linked separately.
""" """
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = os.environ.get("PAGERITE_VITE_URL")
if vite_url: if vite_url:
return [f"{vite_url}/src/pagerite.js"] return [f"{vite_url}/src/pagerite.js"], []
if "page" not in _asset_cache: if "page" not in _asset_cache:
manifest = _manifest() manifest = _manifest()
entry = manifest["src/pagerite.js"] entry = manifest["src/pagerite.js"]
_asset_cache["page"] = [f"/{entry['file']}"] scripts = [f"/{entry['file']}"]
stylesheets = [f"/{css}" for css in entry.get("css", [])]
_asset_cache["page"] = scripts, stylesheets
return _asset_cache["page"] return _asset_cache["page"]
@@ -543,3 +734,61 @@ def _editor_assets() -> tuple[list[str], str | None]:
entry = manifest["src/main.js"] entry = manifest["src/main.js"]
_asset_cache["editor"] = [f"/{entry['file']}"], _editor_css_url(None) _asset_cache["editor"] = [f"/{entry['file']}"], _editor_css_url(None)
return _asset_cache["editor"] return _asset_cache["editor"]
def _analytics_assets() -> tuple[list[str], list[str]]:
"""Script and stylesheet URLs for the analytics page entry."""
vite_url = os.environ.get("PAGERITE_VITE_URL")
if vite_url:
return [f"{vite_url}/src/analytics-main.js"], []
if "analytics" not in _asset_cache:
manifest = _manifest()
entry = manifest["src/analytics-main.js"]
scripts = [f"/{entry['file']}"]
stylesheets = [f"/{css}" for css in entry.get("css", [])]
_asset_cache["analytics"] = scripts, stylesheets
return _asset_cache["analytics"]
def render_analytics(
menu: dict[str, Node],
brand: str = SITE_NAME,
custom_css: str = "",
theme: str = "",
favicon: str = "",
brand_html: str = "",
) -> str:
"""Render the analytics viewer as a normal page at /_a.
The analytics entry is inlined into this page only (prod) or loaded
from the Vite dev server (dev); its stylesheet rides along in <head>
so fetch-navigations can sync it into the live document. The initial
range is not rendered in: the client takes it from the URL hash or
derives it from the analytics data itself.
"""
page_scripts, page_stylesheets = _page_assets()
analytics_scripts, analytics_stylesheets = _analytics_assets()
scripts = page_scripts + analytics_scripts
stylesheets = page_stylesheets + analytics_stylesheets
doc = E.article
with doc:
# .wide: the dashboard breaks out of the article column to the full
# viewport width, like wide figures (see the .wide rules).
doc.div(id="analytics-app", class_="wide")
return str(
_layout(
scripts,
stylesheets,
custom_css,
theme,
banner_design(menu, "_a", theme),
favicon,
)(
Title=f"Analytics {brand}" if brand else "Analytics",
Brand=_brand_link(brand, brand_html),
Nav=nav_html(menu, "_a"),
Sidebar=sidebar_html(menu, "_a"),
Banner=banner_html(menu, "_a", theme),
Main=HTML(str(doc)),
),
)
+5 -3
View File
@@ -20,10 +20,14 @@ dependencies = [
"fastapi-vue>=1.3.1", "fastapi-vue>=1.3.1",
"fastapi[standard]>=0.141.1", "fastapi[standard]>=0.141.1",
"html5tagger>=2.0.0", "html5tagger>=2.0.0",
"httpx>=0.28.1",
"kanta>=0.8.1", "kanta>=0.8.1",
"markdown-it-py>=4.2.0", "markdown-it-py>=4.2.0",
"maxminddb>=3.1.1",
"mdit-py-plugins>=0.6.1", "mdit-py-plugins>=0.6.1",
"pygments>=2.20.0", "pygments>=2.20.0",
"ua-parser>=1.0.2",
"zstandard>=0.25.0",
] ]
[project.scripts] [project.scripts]
@@ -33,9 +37,7 @@ pagerite = "pagerite.__main__:main"
Repository = "https://git.zi.fi/LeoVasanko/pagerite" Repository = "https://git.zi.fi/LeoVasanko/pagerite"
[dependency-groups] [dependency-groups]
dev = [ dev = []
"httpx>=0.28.1",
]
[tool.hatch.version] [tool.hatch.version]
source = "vcs" source = "vcs"
+2 -2
View File
@@ -19,8 +19,8 @@ from devutil import (
setup_vite, setup_vite,
) )
DEFAULT_VITE_PORT = 3100 DEFAULT_VITE_PORT = 8200
DEFAULT_DEV_PORT = 3200 DEFAULT_DEV_PORT = 8210
HEALTH = "/?from=devserver.py" HEALTH = "/?from=devserver.py"
+941
View File
@@ -0,0 +1,941 @@
#!/usr/bin/env -S uv run --script
# /// script
# requires-python = ">=3.14"
# dependencies = [
# "httpx>=0.28.1",
# "playwright>=1.45.0",
# ]
# ///
"""Generate fake browser visits, crawler hits, and abuse scans for a Pagerite site.
Browser sessions (ordinary users) come from realistic residential IPv4 and IPv6
addresses and stay mostly stable; an IPv6 host part may rotate once mid-session,
and an IPv4 session may switch to another residential address. Crawler hits come
from datacenter IPs, with each crawler profile paired to a matching provider IP
when possible. Abuse scanners fire bursts of vulnerability probes from pinned
datacenter IPs.
Run against a local dev server, e.g.:
uv run scripts/fake_traffic.py http://localhost:3200
Repeat whenever you want more traffic; each run appends new events to the
site's analytics file.
"""
from __future__ import annotations
import argparse
import logging
import random
import sys
import time
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import UTC, datetime
from typing import Any
from urllib.parse import urlencode, urljoin, urlparse
import httpx
logging.basicConfig(level=logging.INFO, format="%(message)s")
logger = logging.getLogger("fake-traffic")
@dataclass(frozen=True)
class BrowserProfile:
name: str
user_agent: str
accept_language: str
viewport: tuple[int, int]
@dataclass(frozen=True)
class CrawlerProfile:
name: str
user_agent: str
ip: str
BROWSER_PROFILES: list[BrowserProfile] = [
BrowserProfile(
"chrome-desktop",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"en-US,en;q=0.9",
(1366, 768),
),
BrowserProfile(
"safari-desktop",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"en-GB,en;q=0.9",
(1440, 900),
),
BrowserProfile(
"firefox-desktop",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"en-CA,en;q=0.8,fr;q=0.5",
(1920, 1080),
),
BrowserProfile(
"chrome-mobile",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"es-ES,es;q=0.9",
(390, 844),
),
]
CRAWLER_PROFILES: list[CrawlerProfile] = [
CrawlerProfile(
"googlebot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"66.249.64.66", # US, Google
),
CrawlerProfile(
"bingbot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"40.77.167.0", # US, Microsoft
),
CrawlerProfile(
"duckduckbot",
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
"95.217.0.1", # Germany, Hetzner VPS
),
CrawlerProfile(
"curl",
"curl/8.5.0",
"139.162.0.1", # Singapore, Linode VPS
),
]
# Residential IPv4 addresses and IPv6 /64 prefixes used for ordinary browser
# sessions. IPv6 entries keep the network part stable and randomise only the
# host part; the host may rotate once mid-session.
RESIDENTIAL_SOURCE_IPS: list[str] = [
# Residential IPv4
"91.154.140.209", # Finland, Elisa
"84.143.145.207", # Germany, Deutsche Telekom
"220.165.255.254", # China, Chinanet / China Telecom
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
# Residential IPv6 /64 prefixes
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
"240e:45d:1e60:5b0::/64", # China, China Telecom
"2409:8904:6720:4123::/64", # China, China Unicom
]
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
# the whole scan burst.
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
# and index 2 uses a fixed crawler UA.
ABUSE_SOURCE_IPS: list[str] = [
"45.63.0.12", # US, Vultr VPS
"138.197.0.89", # US, DigitalOcean / Cloudways
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
]
# Paths commonly probed by attackers looking for exposed config, admin panels,
# version control, credentials, backups, or debug endpoints.
SUSPICIOUS_PATHS: list[str] = [
"/.env",
"/env",
"/.env.local",
"/env.development",
"/config",
"/config.json",
"/config.yaml",
"/config.yml",
"/configuration.json",
"/configuration.yaml",
"/configuration.yml",
"/settings.json",
"/settings.yaml",
"/settings.yml",
"/app.config",
"/appsettings.json",
"/appsettings.Development.json",
"/credentials",
"/credentials.json",
"/secrets",
"/secrets.json",
"/.aws/credentials",
"/.ssh/id_rsa",
"/id_rsa",
"/id_rsa.pub",
"/known_hosts",
"/sftp-config.json",
"/admin",
"/administrator",
"/adminer.php",
"/login",
"/signin",
"/auth/login",
"/api/login",
"/api/.env",
"/api/config",
"/api/v1/config",
"/api/v2/config",
"/webhook",
"/webhooks",
"/callback",
"/proxy",
"/image",
"/images",
"/preview",
"/download",
"/downloads",
"/log",
"/logs",
"/debug",
"/trace",
"/phpinfo.php",
"/info.php",
"/phpmyadmin",
"/pma",
"/myadmin",
"/phpMyAdmin",
"/wp-admin",
"/wp-login.php",
"/wp-config.php",
"/xmlrpc.php",
"/wp-json/wp/v2/users",
"/.git/config",
"/.git/HEAD",
"/git/config",
"/swagger-ui.html",
"/v2/api-docs",
"/actuator/env",
"/actuator/health",
"/actuator/configprops",
"/server-status",
"/.htaccess",
"/web.config",
"/package.json",
"/composer.json",
"/vendor/autoload.php",
"/docker-compose.yml",
"/Dockerfile",
"/manage",
"/console",
"/manager",
"/manager/html",
"/metrics",
"/prometheus",
"/healthz",
"/_api",
"/api",
"/api/v1/",
"/api/v2/",
"/graphql",
"/query",
"/feed",
"/rss",
"/_debug",
"/test",
"/testing",
"/tmp",
"/temp",
"/backup",
"/backups",
"/dump",
"/dumps",
"/sql",
"/db",
"/database",
"/dump.sql",
"/backup.sql",
"/db.sql",
"/backup.zip",
"/backup.tar.gz",
"/site.zip",
"/site.tar.gz",
"/source.zip",
"/src.zip",
"/upload",
"/uploads",
"/import",
"/export",
"/token",
"/tokens",
"/oauth",
"/oauth2",
"/openid",
"/jwks",
"/keys",
"/key",
"/private",
"/public",
]
# Realistic external referers. Most sessions arrive with a generic referer;
# a subset carries matching UTM tags on the landing URL.
PLAIN_REFERRERS: list[str] = [
"https://example.com/",
"https://somedomain.com/",
"https://another-site.org/",
"https://friend-site.net/",
]
# (referer origin, utm parameter dict) pairs used for tagged traffic.
TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
("https://chatgpt.com/", {"utm_source": "chatgpt.com"}),
("https://www.google.com/", {"utm_source": "google", "utm_medium": "organic"}),
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
]
# Fraction of referered sessions that also carry UTM tags.
UTM_RATE = 0.25
# Innocent-looking paths that do not exist on a Pagerite site. Hitting many of
# these from a single IP is itself a telltale of a spray-and-pray scanner.
NORMAL_404_PATHS: list[str] = [
"/about",
"/about-us",
"/services",
"/products",
"/contact",
"/contact-us",
"/team",
"/careers",
"/jobs",
"/pricing",
"/features",
"/demo",
"/trial",
"/docs",
"/documentation",
"/api-docs",
"/support",
"/help",
"/faq",
"/knowledge-base",
"/terms",
"/terms-of-service",
"/privacy",
"/privacy-policy",
"/legal",
"/blog",
"/news",
"/articles",
"/press",
"/events",
"/webinars",
"/podcast",
"/videos",
"/resources",
"/whitepapers",
"/case-studies",
"/customers",
"/clients",
"/testimonials",
"/reviews",
"/partners",
"/integrations",
"/api-reference",
"/developers",
"/status",
"/security",
"/trust",
"/compliance",
"/gdpr",
"/ccpa",
"/sitemap",
"/archive",
"/tags",
"/categories",
"/search",
"/users",
"/accounts",
"/dashboard",
"/profile",
"/settings",
"/preferences",
"/notifications",
"/messages",
"/inbox",
"/calendar",
"/reports",
"/analytics",
"/billing",
"/invoice",
"/orders",
"/cart",
"/checkout",
"/store",
"/shop",
"/home",
"/main",
"/start",
"/welcome",
"/intro",
"/overview",
"/summary",
"/portfolio",
"/projects",
"/work",
"/solutions",
]
ABUSE_USER_AGENTS: list[str] = [
# Desktop browsers
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_5 like Mac OS X) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Mobile/15E148 Safari/604.1",
# Well-known crawlers / bots
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; "
"+http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; "
"+http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html)",
"Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36 "
"(compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)",
"Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com)",
"Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html)",
"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)",
# Social / service fetchers
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
"Twitterbot/1.0",
"LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)",
"Slackbot-LinkExpanding 1.0 (+https://api.slack.com/robots)",
"WhatsApp/2.23.20.0",
# Command-line / library clients
"curl/8.5.0",
"Wget/1.21.4 (linux-gnu)",
"python-requests/2.32.3",
"Go-http-client/1.1",
"Node.js/20.5.1",
]
def _random_ipv6_host(prefix: str) -> str:
"""Return a concrete address within an IPv6 /64 prefix.
The host part is generated randomly, mimicking a fresh OS privacy address.
The input prefix must end in ``::/64`` (e.g. ``2a02:8109:ac82:6f0c::/64``).
"""
if "/" not in prefix:
return prefix
base, mask = prefix.split("/")
if mask != "64":
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
if base.endswith("::"):
base = base[:-2]
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
return f"{base}:{host}"
def _concretize_ip(entry: str) -> str:
"""Return a concrete IP address; randomise the host part for IPv6 /64 prefixes."""
if ":" in entry and "/" in entry:
return _random_ipv6_host(entry)
return entry
class _SessionIP:
"""Stable IP for a browser session, with one optional mid-session rotation.
IPv6 prefixes get a fresh random host part; IPv4 addresses are swapped for
another address from the residential pool.
"""
def __init__(self, entry: str, pool: Sequence[str]):
self.entry = entry
self.pool = pool
self._value = _concretize_ip(entry)
def current(self) -> str:
return self._value
def rotate(self) -> None:
if ":" in self.entry and "/" in self.entry:
self._value = _random_ipv6_host(self.entry)
return
# IPv4: switch to another IPv4 address from the residential pool.
for _ in range(20):
candidate_entry = random.choice(self.pool)
if ":" in candidate_entry and "/" in candidate_entry:
continue
candidate = _concretize_ip(candidate_entry)
if candidate != self._value:
self._value = candidate
return
def _sleep(base: float, jitter: float) -> None:
time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
def _normalize_url(url: str) -> str:
"""Return a usable base URL, adding missing scheme/host/port parts.
- bare ``:PORT`` becomes ``http://localhost:PORT``
- missing scheme becomes ``http://``
- otherwise returned as-is
Raises ``ValueError`` when the result is not a valid http(s) URL.
"""
raw = url.strip()
if not raw:
raise ValueError("empty URL")
if raw.startswith(":"):
raw = f"http://localhost{raw}"
elif raw.isdigit():
raw = f"http://localhost:{raw}"
elif not raw.startswith(("http://", "https://")):
raw = f"http://{raw}"
parsed = urlparse(raw)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError(f"invalid URL: {url!r}")
return raw
def _poisson_wait(rate: float) -> float:
"""Return an exponential inter-arrival time for the given Poisson rate."""
if rate <= 0:
return 0.0
return random.expovariate(rate)
def _collect_links(page: Any, include_external: bool = False) -> list[dict[str, Any]]:
"""Return links from the current page, excluding the current page.
Internal links stay on the site; external links are real https URLs found
in the page content and are marked with ``external: true``.
"""
return page.evaluate(
"""(includeExternal) => {
const loc = new URL(location.href);
const out = [];
for (const a of document.querySelectorAll('a[href]')) {
try {
const u = new URL(a.href);
const rect = a.getBoundingClientRect();
const item = {
href: a.href,
text: (a.innerText || a.title || '').trim().slice(0, 60),
visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0),
};
if (u.origin === loc.origin
&& !u.pathname.startsWith('/_')
&& !u.pathname.startsWith('/auth')
&& u.pathname !== '/favicon.ico'
&& u.pathname !== loc.pathname) {
out.push(item);
} else if (includeExternal && u.protocol === 'https:' && u.origin !== loc.origin) {
out.push({ ...item, external: true });
}
} catch { /* ignore malformed hrefs */ }
}
return out;
}""",
include_external,
)
def _click_link(page: Any, link: dict[str, Any], timeout: float = 10.0) -> bool:
"""Click an internal link and wait for the client-side URL to change."""
start_url = page.url
try:
# Prefer Playwright's native click; fall back to a JS click if the
# locator cannot be resolved or times out.
try:
page.locator(f"a[href='{link['href']}']").first.click(timeout=2000)
except Exception: # noqa: BLE001
clicked = page.evaluate(
"""(href) => {
const a = Array.from(document.querySelectorAll('a[href]'))
.find(el => el.href === href);
if (a) { a.click(); return true; }
return false;
}""",
link["href"],
)
if not clicked:
return False
# Wait for the client-side navigation to update the URL.
deadline = time.time() + timeout
while time.time() < deadline:
if page.url != start_url:
return True
page.wait_for_timeout(100)
return False
except Exception as exc: # noqa: BLE001
logger.debug("click failed on %s: %s", link.get("href"), exc)
return False
def _run_browser_session(
base: str,
paths: Sequence[str],
profile: BrowserProfile,
session_index: int,
ip_entry: str,
) -> dict[str, Any]:
from playwright.sync_api import sync_playwright
MAX_CLICKS = 6
STAY = (2.0, 6.0)
HEADLESS = True
REFERER_RATE = 0.75
INCLUDE_EXTERNAL = True
ip_provider = _SessionIP(ip_entry, RESIDENTIAL_SOURCE_IPS)
ips_used: list[str] = [ip_provider.current()]
trail: list[str] = []
start_time = datetime.now(UTC)
try:
with sync_playwright() as p:
browser = p.chromium.launch(
headless=HEADLESS,
args=["--no-sandbox", "--disable-dev-shm-usage"],
)
extra_headers = {
"X-Forwarded-For": ip_provider.current(),
"Accept-Language": profile.accept_language,
}
# Most sessions arrive from an external origin; some are direct.
# A subset of referered sessions carries realistic UTM tags on the
# landing URL; the referer origin is paired with the UTM source.
tagged: dict[str, str] = {}
if random.random() < REFERER_RATE:
if random.random() < UTM_RATE:
referer, tagged = random.choice(TAGGED_REFERRERS)
else:
referer = random.choice(PLAIN_REFERRERS)
extra_headers["Referer"] = referer
context = browser.new_context(
user_agent=profile.user_agent,
viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
extra_http_headers=extra_headers,
)
page = context.new_page()
# Update X-Forwarded-For per request; the value stays stable unless we
# explicitly rotate it once mid-session.
def _route_handler(route, request):
headers = dict(request.headers)
headers["X-Forwarded-For"] = ip_provider.current()
ips_used.append(headers["X-Forwarded-For"])
route.continue_(headers=headers)
page.route("**/*", _route_handler)
# Pick one point during the session to emulate an IP rotation.
rotate_at = random.randint(0, MAX_CLICKS - 1) if MAX_CLICKS > 0 else -1
entry = random.choice(paths) if paths else "/"
landing = urljoin(base, entry)
if tagged:
sep = "&" if "?" in landing else "?"
landing += sep + urlencode(tagged)
page.goto(landing, wait_until="networkidle")
trail.append(page.url)
for click_idx in range(MAX_CLICKS):
_sleep(random.uniform(*STAY) / 2, 0.3)
if click_idx == rotate_at:
ip_provider.rotate()
ips_used.append(ip_provider.current())
logger.debug("rotated session IP to %s", ip_provider.current())
links = _collect_links(page, INCLUDE_EXTERNAL)
visible = [item for item in links if item.get("visible")]
if not visible:
visible = links
if not visible:
break
link = random.choice(visible)
ok = _click_link(page, link)
if not ok:
# Retry once with any link (sometimes visible calc misses nav).
alt = random.choice(links) if links else None
if alt and alt is not link:
ok = _click_link(page, alt)
if not ok:
break
if link.get("external"):
# Outbound navigation: the analytics exit ping is already
# in flight. Record the external URL and end the session.
trail.append(page.url)
_sleep(0.5, 0.2)
break
page.wait_for_load_state("networkidle")
trail.append(page.url)
_sleep(random.uniform(*STAY), 0.5)
browser.close()
return {
"profile": profile.name,
"entry": entry,
"ip": ips_used[0],
"ips_seen": len(set(ips_used)),
"pages": len(trail),
"trail": [urlparse(u).path or "/" for u in trail],
"duration": (datetime.now(UTC) - start_time).total_seconds(),
}
except Exception as exc: # noqa: BLE001
logger.warning("browser session failed: %s", exc)
return {"profile": profile.name, "error": str(exc), "trail": trail}
def _run_crawler_hit(
base: str,
paths: Sequence[str],
profile: CrawlerProfile,
) -> dict[str, Any]:
path = random.choice(paths) if paths else "/"
url = urljoin(base, path)
fake_ip = profile.ip
headers = {
"User-Agent": profile.user_agent,
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.5",
}
try:
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
r = client.get(url, headers=headers)
return {
"profile": profile.name,
"path": path,
"status": r.status_code,
"ip": fake_ip,
}
except Exception as exc: # noqa: BLE001
return {"profile": profile.name, "path": path, "error": str(exc)}
def _abuse_ua() -> str:
"""Return a randomized, syntactically valid user agent for an abuse scan."""
return random.choice(ABUSE_USER_AGENTS)
def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
"""Fire a burst of vulnerability probes from a single fake IP.
Scanner 0 randomises its user agent every request, scanner 1 uses a fixed
browser UA, and scanner 2 uses a fixed crawler UA.
"""
ip_entry = ABUSE_SOURCE_IPS[ip_index % len(ABUSE_SOURCE_IPS)]
if ":" in ip_entry and "/" in ip_entry:
fake_ip = _random_ipv6_host(ip_entry)
else:
fake_ip = ip_entry
MIN_HITS = 15
MAX_HITS = 25
total_hits = random.randint(MIN_HITS, MAX_HITS)
# Ensure the burst contains both telltales: suspicious paths and more
# than ten normal-looking 404 paths.
suspicious_count = max(5, total_hits // 3)
normal_count = total_hits - suspicious_count
if normal_count < 11:
normal_count = 11
suspicious_count = max(3, total_hits - normal_count)
paths = random.choices(SUSPICIOUS_PATHS, k=suspicious_count) + random.choices(
NORMAL_404_PATHS, k=normal_count
)
random.shuffle(paths)
ua_mode = ip_index % 3
if ua_mode == 0:
get_ua = _abuse_ua
elif ua_mode == 1:
def get_ua() -> str:
return BROWSER_PROFILES[0].user_agent
else:
def get_ua() -> str:
return CRAWLER_PROFILES[0].user_agent
scan_results: list[dict[str, Any]] = []
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
for path in paths:
headers = {
"User-Agent": get_ua(),
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": random.choice(
["en-US,en;q=0.9", "en-GB,en;q=0.8", "en;q=0.7"]
),
}
try:
r = client.get(urljoin(base, path), headers=headers)
scan_results.append(
{"path": path, "status": r.status_code, "ua": headers["User-Agent"]}
)
except Exception as exc: # noqa: BLE001
scan_results.append({"path": path, "error": str(exc)})
_sleep(0.15, 0.1)
return {
"scanner": ip_index + 1,
"ip": fake_ip,
"hits": len(scan_results),
"results": scan_results,
}
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Generate fake traffic for a Pagerite site.",
formatter_class=argparse.ArgumentDefaultsHelpFormatter,
)
parser.add_argument(
"url",
nargs="?",
default="http://localhost:8200",
help="Base URL of the Pagerite site (default: http://localhost:8200). "
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
"missing scheme defaults to http://.",
)
parser.add_argument(
"-t",
"--duration",
type=float,
default=60.0,
metavar="SECONDS",
help="Rough maximum time to generate traffic (0 runs one preset batch)",
)
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
return parser.parse_args(argv)
def main(argv: Sequence[str] | None = None) -> int:
args = _parse_args(argv)
if args.verbose:
logger.setLevel(logging.DEBUG)
try:
base = _normalize_url(args.url).rstrip("/")
except ValueError as exc:
logger.error("%s", exc)
return 2
# Discover content paths from the public page tree if we can.
paths: list[str] = []
try:
r = httpx.get(urljoin(base, "/_api/pages"), timeout=10.0)
if r.status_code == 200:
paths = [page["path"] for page in r.json() if page.get("has_content")]
except Exception as exc: # noqa: BLE001
logger.debug("could not fetch page list: %s", exc)
if not paths:
paths = ["/"]
logger.info(
"Generating fake traffic against %s (%d content paths, duration=%ss)",
base,
len(paths),
args.duration,
)
results: list[dict[str, Any]] = []
arrival_rate = 1.0
def _wait() -> None:
wait = _poisson_wait(arrival_rate)
logger.debug("waiting %.2fs before next session", wait)
time.sleep(wait)
if args.duration <= 0:
# One preset batch.
for i in range(5):
if i > 0:
_wait()
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(base, paths, profile, i, ip_entry)
results.append(result)
logger.debug(" trail: %s", result.get("trail", []))
for i in range(10):
if i > 0:
_wait()
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
results.append(result)
for i in range(3):
if i > 0:
_wait()
ip_entry = ABUSE_SOURCE_IPS[i % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, i)
results.append(result)
logger.debug(
" hits: %s", [r.get("path") for r in result.get("results", [])]
)
else:
deadline = time.time() + args.duration
session_index = 0
while time.time() < deadline:
if session_index > 0:
_wait()
phase = session_index % 3
if phase == 0:
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(
base, paths, profile, session_index, ip_entry
)
logger.debug(" trail: %s", result.get("trail", []))
elif phase == 1:
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
else:
ip_entry = ABUSE_SOURCE_IPS[session_index % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, session_index // 3)
logger.debug(
" hits: %s",
[r.get("path") for r in result.get("results", [])],
)
results.append(result)
session_index += 1
ok = sum(1 for r in results if "error" not in r)
logger.info("Done: %d/%d requests succeeded.", ok, len(results))
return 0 if ok == len(results) else 1
if __name__ == "__main__":
sys.exit(main())