Compare commits

..
35 Commits
Author SHA1 Message Date
LeoVasanko 2d595f8c15 Reduce #brand to actual size, avoid clicks on empty banner space touching it. 2026-08-24 20:58:50 +00:00
LeoVasanko b1fc8e24d7 Eyes banner follows taps not just mouse. 2026-08-24 20:54:50 +00:00
LeoVasanko c5cf799f68 Better nav layout for portrait phones. 2026-08-24 20:43:18 +00:00
LeoVasanko a5edc7b3b6 Fix crawler misclassification from /_a and orphan counts on visit scrub
Two analytics corrections verified against the production capture:

- pagerite.js suppressed pings with fr == '/_a', but fetch-navigation
  away from the analytics page had already GET-ed the target without the
  preload header; the orphaned pending hit then flushed to the crawler
  list, classifying a real user as a crawler. Navigations away from /_a
  now ping normally (the server rejects /_a as a target regardless, and
  admin noise is already handled by hide=1).

- _remove_visit only reversed the visit's creation counts, leaving
  views/transitions from later pings behind as orphans on the graph with
  no matching row in the visitor table. An in-memory per-visit count log
  now tracks every count event, so an admin hide=1 scrub reverses the
  visit completely.
2026-08-24 20:18:13 +00:00
LeoVasanko 09ebc63690 Record response status per path; mark 404 trails red in the viewer
Document GETs now stash their status (200/404) in a pending table,
consumed by the matching ping: visits gain a per-path statuses map and
crawler hits a status field. Trail links with a 404 status render in
red with the status code in the tooltip, alongside the read time.
2026-08-24 19:43:23 +00:00
LeoVasanko 2868843028 Fix inverted client filter in admin hide ping
The hide=1 branch kept the admin's own pending crawler hits (==) instead
of discarding them (!=), so an admin's document GETs flushed to the
crawler list 10s later while every other client's pending hits were
wrongly dropped. This is why ordinary admin browsers showed up as
crawlers.
2026-08-24 19:40:12 +00:00
LeoVasanko 6fcebaea3f Transition map: svg-scaled fonts, border-clipped pill text, tight crop
- Fonts scale with the svg instead of the --u constant-screen-size
  compensation (ResizeObserver machinery removed)
- Larger node text (slug 19px, count 15px)
- Pill labels no longer ellipsis-truncated: text is clipped at the pill
  border via per-node clipPaths; captions center when they fit and anchor
  left on overflow so the title's beginning survives; count lines stay
  centered
- Bounding box crops to pill half extents plus the ribbon halo instead of
  the diagonal radius, removing the large top/bottom margins
2026-08-24 18:35:36 +00:00
LeoVasanko 4a48d08a19 Analytics layout: larger charts with natural-width cap, one-line totals
- Charts grow to a larger intrinsic size (1052x174) and never upscale
  past it; centered with equal side margins above the cap, full width
  below, svg always within page bounds; overflow visible so wider fonts
  don't clip at the viewBox edge
- Totals row aligns its left edge with the charts and shrinks (gap first,
  then font) via container units to always stay on one line
2026-08-24 18:21:58 +00:00
LeoVasanko e3e29251ca Transition map: cull invisible connectors, stable beads, lane labels
- Cull connections whose thin middle would render below ~0.8px
  (MIN_WMID); drop external source/exit nodes whose connectors are
  all culled, while site page nodes always stay
- Bead simulation persists across data reloads: emitters keyed per edge
  direction, beads tracked by progress, so unrelated count changes no
  longer reshuffle bead positions
- Bead speed relative to span length: constant 1.5s traversal per edge
- Top lane labeled with a house icon; all lane labels left-aligned just
  past the source pill (half height on near-vertical branch lanes), with
  guides running to the lane end so long slugs are never truncated
2026-08-24 17:58:02 +00:00
LeoVasanko 8affc41289 Tighter analytics chart chrome
- Unified chart text at 11px system-ui; fixed size independent of theme font
- Day view y axis reads "visits / 5 min" / "views / 5 min"
- Left margin and y tick spacing tightened (MARGIN_L 56 -> 40)
- Gap between visits and views charts removed
- Legend repositioned for the larger font
2026-08-24 17:26:30 +00:00
LeoVasanko 2de4717230 Fix week overlay alignment, in-plot ISO week legend, shorter charts
- weeklySeries shifts overlaid weeks onto the current week's time axis so
  they overlay inside the plot instead of overflowing left; oldest weeks
  paint first, current week on top
- Legend moved inside the visits chart's top right: current ISO week in
  accent, past weeks as a single muted "Week M" / "Week M–N" specimen
- Past week curves use the muted color instead of faded accent
- Chart height reduced ~30% (180 -> 126)
2026-08-24 17:03:49 +00:00
LeoVasanko 563e8fcaf2 Rework analytics chart scaling; self-contained SVG charts
- Charts render as single SVGs with axis labels inside the viewBox,
  replacing the stretched plot + HTML overlay labels
- Rolling ranges end at now, t0 aligned to UTC day; bucket size follows
  the window (6h up to 31 days) so "all" at its 30-day minimum renders
  identically to "month"
- X labels always centered on their true position; no edge-align shifting
- rangeWindow simplified to rolling spans ending at now
- TransitionGraph "all" visual scale floored at the 30-day plot minimum
2026-08-24 12:31:25 +00:00
LeoVasanko 921a5484a2 Fixed-sigma smoothing of traffic history plots. 2026-08-24 05:54:43 +00:00
LeoVasanko 0e2e52fa45 Use last 24h/7d/30d/365d/all analytics data. Previously some fields were unfiltered and weekly view was based on calendar weeks. 2026-08-24 05:42:32 +00:00
LeoVasanko 075848f782 Crawlers should include all sorts of spiders along with bots and googleother. 2026-08-24 05:28:05 +00:00
LeoVasanko 7b8899af92 Transition map: bounded node scaling, concentric branch lanes, exit row at bottom 2026-08-24 04:47:06 +00:00
LeoVasanko 6d2ae104d7 Fix analytics classification: ignore bot-UA pings, skip preload GETs.
JS-running crawlers (Googlebot, GoogleOther, Applebot) execute pagerite.js
and send navigation pings, registering as visitors. Pings whose User-Agent
matches _is_bot_ua (any "bot" token plus listed exceptions) are now
ignored, so their document GETs flush to the crawler list as intended. No
source verification: a spoofed bot UA merely lands in the crawler stats,
and path-based abuse classification catches scanners regardless.

Idle-time link preloads from pagerite.js were queued as pending crawler
hits and flushed to the crawler list whenever the user navigated more than
10s later, so real visitors' subpage loads showed up as crawler hits.
Preload fetches now carry an x-pagerite-preload header and the document
GET handler skips tracking for them; the ping sent on actual navigation
does the counting.
2026-08-22 18:23:38 +00:00
LeoVasanko b7fc543a83 Transition graph layout follows navigation. 2026-08-22 18:10:44 +00:00
LeoVasanko b868033ddc Pill shaped nodes 2026-08-22 15:48:33 +00:00
LeoVasanko 8aad64cced Analytics layout update, larger, consistent text sizing. 2026-08-22 14:45:56 +00:00
LeoVasanko 319163ee7e Page caching and zstd compression. Avoid useless fetching. Mobile layouts of navigation menus improved. 2026-08-22 14:05:12 +00:00
LeoVasanko 87b16b7144 Implement /robots.txt and /sitemap.xml. Update dev proxy to all-by-default. 2026-08-22 12:13:17 +00:00
LeoVasanko 20ae6501f2 Add dynamic /sitemap.xml and /robots.txt endpoints 2026-08-22 12:01:18 +00:00
LeoVasanko a16fe88114 Auto select day if less than 24h data for new sites. 2026-08-22 01:51:36 +00:00
LeoVasanko c77598adc7 Fine tuning date formatting. 2026-08-22 00:09:57 +00:00
LeoVasanko 29f8fac013 Slightly prettier analytics URL 2026-08-21 23:58:00 +00:00
LeoVasanko 6199e5a69e Support for UTM tags in transition graph as source sites. 2026-08-21 23:50:01 +00:00
LeoVasanko 375b4b6bdb analytics: unify visitor cell across visits, crawlers and abuse tables 2026-08-21 23:32:46 +00:00
LeoVasanko fdb3e42d6f analytics: shared Client struct, grouped abuse paths, unified visitor cell 2026-08-21 23:16:15 +00:00
LeoVasanko 51a6a16221 Neater abuse table formatting. 2026-08-21 22:31:37 +00:00
LeoVasanko 0798e24d24 Desaturated house emojis 2026-08-21 22:03:55 +00:00
LeoVasanko 3be2d08ac9 SI formatting of large visitor numbers. 2026-08-21 21:36:57 +00:00
LeoVasanko b7d5b23ae6 Cleaner formatting of utm tags in visitor table. 2026-08-21 21:22:16 +00:00
LeoVasanko 6eaa1c1a8b analytics: 24h day view with bar chart for precise realtime stats. Tables redesigned with cleaner layout. Tracking article read times. Adjust connection graph visualizations by time range. Other cleanup and supporting systems. 2026-08-21 20:14:04 +00:00
LeoVasanko f341d22aa0 Improved fake traffic generation with abuse bots, utm tags etc. 2026-08-21 20:10:41 +00:00
28 changed files with 3994 additions and 1362 deletions
+178 -75
View File
@@ -5,8 +5,9 @@ Struct dumped to disk — separate from the kanta content database, path from
`PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by `PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by
`.analytics.json`, e.g. `pagerite.analytics.json`). `.analytics.json`, e.g. `pagerite.analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Visit`) and the `Store` - `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
(in-memory data + session map, atomic JSON persistence). `CrawlerHit`, `AbuseHit`) and the `Store` (in-memory data + session map,
atomic JSON persistence).
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`), - `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics` the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics`
(admin-gated like every `/_api` endpoint). (admin-gated like every `/_api` endpoint).
@@ -23,7 +24,14 @@ The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
- **Initial page load**: `to` is the loaded path. This ping is what starts - **Initial page load**: `to` is the loaded path. This ping is what starts
the visit and counts the entry page view — the document GET alone records the visit and counts the entry page view — the document GET alone records
nothing, so bots and admin browsing never register. Reloads are not nothing, so bots and admin browsing never register. JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent
gives them away: pings whose UA matches `_is_bot_ua` (anything calling
itself a "bot", plus known exceptions such as GoogleOther) are ignored
server-side, and their document GETs land in the crawler list instead.
No source-IP verification is done: a spoofed bot UA merely lands in the
crawler stats, and scanners that probe telltale paths are caught by the
abuse rules regardless. Reloads are not
visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a
refresh neither counts a second view nor logs a self-transition. The GET refresh neither counts a second view nor logs a self-transition. The GET
handler stashes a cross-origin https `Referer` (origin part only) and any handler stashes a cross-origin https `Referer` (origin part only) and any
@@ -37,81 +45,144 @@ The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
back), so the exit URL is not necessarily the last trail entry. Outbound back), so the exit URL is not necessarily the last trail entry. Outbound
links are stored by full URL so several links to the same domain remain links are stored by full URL so several links to the same domain remain
distinct. distinct.
- **Excluded**: back/forward (popstate) navigations, navigation involving - **Excluded**: back/forward (popstate) navigations, navigating *to* the
the analytics page itself (`/_a`), and everything while the user is known to analytics page (`/_a` — its GET is untracked, and the server rejects it
be an admin *and SSO is actually in use* — with no auth proxy (dev/test) as a ping target anyway), and everything while the user has the editor
"admin" is everyone's state, so the gate is off and everything is recorded — open (`body.editing`). Admin noise, not visits. Navigating *away* from
or has the editor open (`body.editing`). Admin noise, not visits. `/_a` does ping: the fetch-navigation already GET-ed the target page
without the preload header, and without the ping that GET would flush to
the crawler list.
- **Admins**: when SSO is in use and the session is known to be an admin,
the client still pings but adds `hide=1`. The server then records
nothing — and if the same client session already had a visit from before
logging in, that visit is removed from the JSON along with every count
it recorded — an in-memory per-visit log of count events makes full
reversal possible. With no auth proxy (dev/test)
"admin" is everyone's state, so `hide` stays 0 and everything is recorded.
- The server validates `to`: internal paths must be valid slug paths - The server validates `to`: internal paths must be valid slug paths
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the ("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
https origin and accepted only when the client sent exactly that. https origin and accepted only when the client sent exactly that.
- The initial ping also records the visitor's `User-Agent` and - **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
`Accept-Language` headers. The first `Accept-Language` tag is stored as `User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
`lang` (e.g. `en-us`) and its region subtag, if present, is stored as the first 6 bytes identify a shared `Client` record. The `Client` stores
an initial `country` (e.g. `US`). the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
- The visitor IP is stored. A reverse-DNS lookup is attempted for each new `country` from the language-region subtag, and asynchronously-filled
visit and the result, when available, is cached in RAM and stored as `country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
`host`; local/reserved/multicast addresses are skipped. crawler hits and abuse hits all reference this record by its hash, so
- If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the client metadata is stored once instead of repeated per event.
repository root, it is loaded at startup and used to look up a more accurate - The visitor IP is stored in the `Client`. A reverse-DNS lookup is
`country`. The MMDB lookup and the reverse-DNS lookup run in background attempted for each new client and the result, when available, is stored as
tasks after the visit is stored, so the `/ _a` response is never delayed. `host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
The decompressed `dbip-*.mmdb` file is kept in the repository root and file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
ignored by git. The CLI flag `--dbip` (`uv run pagerite --dbip`) downloads root, it is loaded at startup and used to look up `country`/`city`. These
the latest `dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server lookups run in background tasks after the event is stored, so the `/_a`
starts, skipping the download when the local database is already current and response is never delayed. The decompressed `dbip-*.mmdb` file is kept in
the repository root and ignored by git. The CLI flag `--dbip`
(`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
skipping the download when the local database is already current and
removing older versions after an update; without the flag only an existing removing older versions after an update; without the flag only an existing
file is used. file is used.
- **Crawler hits**: every document GET is queued in RAM as a pending crawler - **Crawler hits**: every document GET is queued in RAM as a pending crawler
hit. If a ping from the same (IP, User-Agent) pair arrives within 10 hit — except idle-time link preloads from pagerite.js, which carry an
seconds the hit is discarded; otherwise it is written to `crawlers`. `x-pagerite-preload` header and are not tracked at all (the ping sent when
Crawlers do not count as visits or views. In the analytics viewer, crawler the user actually navigates to a preloaded page does the counting; forging
hits are grouped by the same (IP, User-Agent) pair and shown as a trail of the header only hides a GET from the crawler stats, the path-based abuse
internal pages that crawler visited; the crawler table lists the most active classification is unaffected). If a ping
crawlers first rather than the most recent hits. from the same client arrives within 10 seconds the hit is discarded;
otherwise it is written to `crawlers`. Crawlers do not count as
visits or views. The `Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited; the crawler table lists
the most active crawlers first rather than the most recent hits.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
classifies the source IP as abuse immediately, and ten plain 404s from one
IP do too. Classification reclassifies history: all earlier crawler hits
from that IP (persisted and pending) move to the `abuse` list, so a
random-UA scanner no longer pollutes the crawler stats of the legitimate
bot it impersonates. Once classified, every document GET and 404 from the
IP is recorded as an abuse hit with the full request path (query string
included), and its pings are ignored. The classified IP set (`abuse_ips`)
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
viewer, abuse hits are grouped by IP (never by client/UA — scanners
randomize theirs) in a separate "Abuse" table. Identical paths are
collapsed into one entry with their hit count; flagged paths that
triggered classification are lifted to the top, followed by other 404s and
then document GETs from the abuser. Raw User-Agent strings are shown one
per line with their occurrence counts, and the full lists are click-to-copy.
## Visits and sessions ## Visits and sessions
There are no cookies. A visit is tied together by the (IP, User-Agent) pair There are no cookies. A visit is tied together by a client hash — the first
(IP from the first `X-Forwarded-For` hop — we sit behind a proxy — else the 6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
direct peer): the first ping from a pair starts a new visit, subsequent /64 network), the raw `User-Agent` string and the extracted
pings extend it. Pings arriving with no known session (server restart) `Accept-Language` tag. The first ping from a client hash starts a new
start a fresh visit from the first ping — treated as missing data rather visit; subsequent pings extend it. Pings arriving with no known session
than dropped. The (IP, UA) → visit map and the IP → entry-referer/UTM (server restart) start a fresh visit from the first ping — treated as
tables are in-memory only, but the IP and any resolvable reverse-DNS host missing data rather than dropped. The client-hash → visit map and the IP →
name are stored on the `Visit` record itself. entry-referer/UTM tables are in-memory only; client metadata is stored in
`Analytics.clients` keyed by the client hash.
Each `Client` record:
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string.
Each `Visit` record: Each `Visit` record:
- `start` — timestamp of the first event, - `start` — timestamp of the first event,
- `entry` — first page (path) seen, - `entry` — first page (path) seen,
- `referer` — external https origin of the initial load, `""` for direct, - `referer` — external https origin of the initial load, `""` for direct,
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer), - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `trail` — everything seen afterwards in first-seen order: page paths and - `trail` — everything seen afterwards in first-seen order: page paths and
external exit URLs. Re-visiting an already seen page (incl. the entry) external exit URLs. Re-visiting an already seen page (incl. the entry)
does not append. does not append.
- `lang` — first `Accept-Language` tag, lowercased (e.g. `en-us`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string from the initial ping,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string,
- `utm``utm_*` query parameters from the landing URL, as a dict. - `utm``utm_*` query parameters from the landing URL, as a dict.
- `read` — active reading time per path (seconds), keyed by path.
- `statuses` — HTTP status of the response when each path was first seen
(200 or 404), keyed by path.
Each `CrawlerHit` record: Each `CrawlerHit` record:
- `start` — timestamp of the document GET, - `start` — timestamp of the document GET,
- `entry` — page path requested, - `entry` — page path requested,
- `ip` — IP address, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `ua` — raw `User-Agent` header,
- `ua_pretty` — compact display form of the UA when parsable,
- `referer` — external https origin of the request, `""` for direct/none, - `referer` — external https origin of the request, `""` for direct/none,
- `query` — raw query string of the request. - `query` — raw query string of the request,
- `status` — HTTP status of the served response (200 for a real page, 404
for a category placeholder or missing page).
Crawler hits are grouped by User-Agent in the analytics viewer. Each `AbuseHit` record:
- `start` — timestamp of the request,
- `path` — full request path including the query string (e.g. `/.env?x=1`),
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the path that triggered abuse classification (telltale
path or the 404 that crossed the threshold),
- `is_404` — true for 404 responses, false for document GETs from the
abuser.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical paths are collapsed with their counts; flagged paths
that triggered classification are lifted to the top, followed by other 404s
and then document GETs from the abuser. Within each category paths are
sorted by count descending, then by their earliest hit.
In the visitor and crawler tables, internal paths that returned a 404 status
are shown in red and the link title includes the status code, so it is easy
to tell misses from real pages at a glance.
## Aggregates ## Aggregates
@@ -148,18 +219,19 @@ with a "could not be loaded" message.
Because it is a real page, fetch-navigation handles it like any other internal Because it is a real page, fetch-navigation handles it like any other internal
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
range selector updates the URL query string (`?range=week` etc.) so links to range selector updates the URL hash (`#week` etc.) so links to a specific
a specific range can be shared. range can be shared. When the URL has no hash, the client derives the
default from the first analytics snapshot: `day` if the recorded history
spans less than 24 hours, otherwise `week`.
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open` `AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
page-chrome hiding and `#/analytics/<range>` hash routing have been removed. page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
Charts are SVG curves (Catmull-Rom over an edge-aware adaptive Gaussian — Charts are SVG curves (Catmull-Rom over an edge-aware Gaussian — a
a change-point detector splits the series at traffic-level shifts, then change-point detector splits the series at traffic-level shifts, then each
each segment is smoothed with a bandwidth that ramps with a broad pilot segment is smoothed independently with a fixed sigma chosen so N events in
estimate of the local rate: isolated events stay narrow (~0.4-unit sigma, a single bucket peak at N events per unit. The raw series is drawn faint
peaking at ~1 event/unit), busy traffic widens to a 1-unit sigma. The raw underneath). Values are
series is drawn faint underneath). Values are
**per-unit rates** — per hour on the week view (5-minute bucket counts × 12, **per-unit rates** — per hour on the week view (5-minute bucket counts × 12,
plotted at native 5-minute resolution), per day on the month+ ranges — and plotted at native 5-minute resolution), per day on the month+ ranges — and
the smoothing time scale follows the unit: the month+ sigmas are 24× the the smoothing time scale follows the unit: the month+ sigmas are 24× the
@@ -170,27 +242,58 @@ labeled intervals, minor lines at fifths when integral; the minimum y-axis
range is 10 so tiny values such as a single visit are not stretched to a range is 10 so tiny values such as a single visit are not stretched to a
fractional scale). fractional scale).
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
weeks in the same accent color at decreasing opacity (the current week is weeks in the muted color at decreasing opacity (the current week keeps the
accent color and is
truncated at the current bucket, never drawing fake zeroes for the future); truncated at the current bucket, never drawing fake zeroes for the future);
its x labels are weekday names centered at midday UTC, without vertical grid a compact legend inside the top right of the visits chart marks the current
ISO week in accent and the overlaid past weeks as "Week M" or "Week MN" on
a muted specimen. Its x labels are weekday names centered at midday UTC, without
vertical grid
lines (day boundaries would be misleading in the viewer's timezone). The lines (day boundaries would be misleading in the viewer's timezone). The
month view labels days the same lineless way — day numbers at noon UTC, month view labels days the same lineless way — day numbers at noon UTC,
with the month name substituted for the 1st. Year is a rolling 365-day window ending at now, re-bucketed to daily points, with the month name substituted for the 1st. Month, year and all are
with boundary lines at months/years. All uses the full data reach, but keeps rolling windows ending at now, aligned to UTC day boundaries at the start
at least the past 30 days so the chart never collapses to a tiny sliver when so the labels span the whole range; the bucket size follows the window —
the site is young. Below the charts: a radial **transition map** (all pages from 6 hours up to 31 days, daily beyond — with boundary lines at months/years
`/_api/pages` — front page at the center, each slug level on its own ring, on the longer ranges. All uses the full data reach, but keeps
siblings clockwise in navigation order from the top, radial gap equal to at least the past 30 days (identical to the month view when the site is
the arc spacing — opposite transition directions joined into organic younger than that, bucket size included) so the chart never collapses to a
tiny sliver when the site is young. Below the charts: a **transition map** (all pages from
`/_api/pages` — top-level menu items on a large-radius circular arc whose
bottom point is the last item (each earlier item a bit higher), connected
by a top lane labeled 🏠︎ beside the home pill (50% thicker than
the branch lanes, its label font and guide offset scaled along), each item's
subtree fanning out below it in menu order along a large-radius circular
arc that leaves heading
straight down and gradually bends right, index pages without views omitted
and their children promoted in their place. The submenu structure is drawn
as wide branch lanes: one per path prefix with at least two visible
nodes, running behind the branch's node pills as circle arcs concentric
with the fan (parent levels one radius step outward, so all lanes of a
group share exactly one form), each labeled with its branch slug
left-aligned just past the first pill and allowed to run along the lane to
its end, disappearing under later pills when long — so the lanes reflect
the path
structure even where index pages are omitted — opposite transition
directions joined into organic
tapered connections whose middle width grows logarithmically with the tapered connections whose middle width grows logarithmically with the
count (a single count renders as a ~1 px line, uncapped), connections count (uncapped), connections
carrying less than 1% of the total traffic carrying less than 1% of the total traffic
pruned; beads are simulated one by one in JS (requestAnimationFrame) and pruned, as are those whose thin middle would render below ~0.8 px —
flow along each edge, emitted at a rate linearly proportional fainter strands are invisible and only their wide end flares would show; beads are simulated one by one in JS (requestAnimationFrame) and
flow along each edge, persisting across data reloads (emitters are keyed
per edge direction and beads tracked by progress, so an unrelated count
change never reshuffles them), emitted at a rate linearly proportional
to the directional count with no in-flight limit, opposing directions to the directional count with no in-flight limit, opposing directions
offset onto parallel lanes. External referers show as a node row above the offset onto parallel lanes. External sources and exits whose connectors are
map, external exits as small nodes fanned outwards from their source all culled by the width threshold are dropped from their rows themselves
page), per-page view (the site's own page nodes always stay, connected or not). External sources show as a node row above the
map: each visit is attributed to `utm_campaign`, then `utm_source`, then the
referer origin, then any other `utm_*` tag, so UTM-tagged visits are grouped
under their campaign/source value rather than the referer domain. A UTM
source node only links to its referer when every visit carrying that tag
came from the same origin. External exits are full-size nodes in a matching
row centered below the map, so the site itself stays in the middle), per-page view
counts, the top transitions and the 50 most recent visit trails. Data is counts, the top transitions and the 50 most recent visit trails. Data is
streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest
JSON snapshot on connect and again whenever the analytics file is updated JSON snapshot on connect and again whenever the analytics file is updated
+3 -1
View File
@@ -8,6 +8,8 @@ The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours. The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding, and `data.version`, which bumps on every content/settings change and so transparently invalidates the whole cache. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and `data.version`; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
## `data.py` ## `data.py`
msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model. msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model.
@@ -20,7 +22,7 @@ markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists,
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering. The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the request base URL; `article:published/modified_time` come from `Node.created`/`modified`. If the markdown contains its own h1, the page title is NOT rendered as an additional h1 (it still supplies `<title>` and nav labels). Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`Data.site_url` — learned from admin browsers reporting their `location.origin` via `POST /_api/site-url`, correct even behind reverse proxies; until learned, the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. If the markdown contains its own h1, the page title is NOT rendered as an additional h1 (it still supplies `<title>` and nav labels).
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), which is rendered when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap). The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), which is rendered when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
+1 -1
View File
@@ -20,7 +20,7 @@ Siblings order by the fractional `Node.order` key: a moved item gets a fresh key
`Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles. `Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles.
`Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css is linked in `<head>` (id `pagerite-banner`) between the theme and the custom CSS. `Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css lives in `<head>` (id `pagerite-banner`) between the theme and the custom CSS — a `<link>` in dev, an inline `<style>` in production.
## Site settings ## Site settings
+1 -1
View File
@@ -25,6 +25,6 @@ Dropping ON the lower part of a row moves the page under that row (the child lis
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`. The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme and banner-design CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-user` elements to restore the canonical order (base < theme < design < custom CSS). Theme switches in the site editor simply swap the `#pagerite-theme` link href, identically in dev and prod. `vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme and banner-design CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-user` elements to restore the canonical order (base < theme < design < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin. `vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
+5 -3
View File
@@ -8,9 +8,11 @@ Vue editor app entry, mounts the tabbed `EditorShell`. See `docs/editing.md` for
## `pagerite.js` ## `pagerite.js`
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link — and the current page — is fetched once at load, clicks are then served from JS with no fetch, and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire), scroll-reveal, OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), code copy buttons, and the auth check. Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire), scroll-reveal, OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, and the auth check.
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen (asset URLs from the `pagerite:editor-src`/`-css` meta tags). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`. The backend links the stylesheets in a fixed order — base (Vite build), theme, banner design, custom CSS last — each with a stable id so the site editor can swap them in place. It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`.
Asset wiring differs by mode. In dev the backend links the Vite dev-server URLs (`pagerite:editor-src`/`-css`/`pagerite:analytics-src` meta tags, `<link>` stylesheets) and Vite injects the entry CSS from JS for hot reloads. In production there are no pagerite meta tags: all page assets are inlined into the document — stylesheets as `<style>` elements in `<head>` (fixed order: base, theme, banner design, entry sheets, custom CSS last), module scripts as inline `<script>`s at the end of the body (relative chunk imports are rewritten to absolute `/_assets/` paths) — and the on-demand bundles' URLs ride in a `<script type="application/json" id="pagerite-assets">` config. The editor bundle always stays external, imported on demand when a pen is opened. Every stylesheet element carries a stable id so fetch-navigation and the site editor can sync `<head>` positionally across swaps (the analytics sheet exists on `/_a` only and is added/removed as you navigate). The analytics entry is inlined into the `/_a` page itself; pagerite.js re-creates that script element after fetch-navigating there (inline scripts don't execute on a DOM swap) and calls the module's exposed unmount before swapping away.
## `assets/` ## `assets/`
@@ -18,7 +20,7 @@ Shared styles and data files built by Vite and served hashed under `/_assets/`:
The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak. Themes and banner designs are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details. The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak. Themes and banner designs are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details.
Vite builds ES-module `.js` outputs; the backend renders `<script type="module">` for them (module scripts defer by default). Vite builds ES-module `.js` outputs; in dev the backend links them as `<script type="module">` (module scripts defer by default), in production it inlines them at the end of the body.
## Database file ## Database file
+1 -1
View File
@@ -34,4 +34,4 @@ The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll para
## Stylesheet order ## Stylesheet order
The backend links the stylesheets in a fixed order — base (Vite build), theme, banner design, custom CSS last — each with a stable id so the site editor can swap them in place. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`. The backend emits the stylesheets in a fixed order — base (Vite build), theme, banner design, entry sheets, custom CSS last — each with a stable id so fetch-navigation and the site editor can sync them in place. In dev they are `<link>`s (the base is Vite-injected from JS instead); in production they are inlined as `<style>` elements. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
+232 -124
View File
@@ -5,20 +5,34 @@
// visit/crawler tables. Read-only. // visit/crawler tables. Read-only.
// See docs/analytics.md for the data format. // See docs/analytics.md for the data format.
import { computed, onMounted, onUnmounted, ref, watch } from 'vue' import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
import { RANGES } from './analytics/time.js'
import { import {
RANGES,
rangeWindow,
filterRecordsByRange,
filterTransitionsByRange,
filterViewsByRange,
} from './analytics/time.js'
import {
calcReadStats,
calcTotalViews, calcTotalViews,
copyIp, copyIp,
copyList,
formatCount,
formatAbuseRows,
formatCrawlerRows, formatCrawlerRows,
formatVisitRows, formatVisitRows,
} from './analytics/format.js' } from './analytics/format.js'
import * as flagSvgs from 'country-flag-icons/string/3x2' import TrailLink from './TrailLink.vue'
import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue' import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue' import VisitorCharts from './VisitorCharts.vue'
import { VIEW_W } from './analytics/chart.js'
const props = defineProps({ // Same centering margin as the charts, so the totals row's left edge
initialRange: { type: String, default: 'week' }, // aligns with the chart svg above the natural width.
}) const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const ABUSE_MAX_LINES = 5
const data = ref(null) const data = ref(null)
const pageTree = ref(null) const pageTree = ref(null)
@@ -28,6 +42,13 @@ let ws = null
let reconnectTimeout = null let reconnectTimeout = null
let timeInterval = null let timeInterval = null
// The initial range comes from the URL hash (shareable links); without one,
// it is derived from the first analytics snapshot: day when the recorded
// history is shorter than 24 h, week otherwise.
const hashRange = location.hash.slice(1)
const range = ref(RANGES[hashRange] ? hashRange : 'week')
let rangePinned = Boolean(RANGES[hashRange])
function connectAnalytics() { function connectAnalytics() {
if (ws) return if (ws) return
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:' const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
@@ -36,6 +57,15 @@ function connectAnalytics() {
ws.onmessage = (event) => { ws.onmessage = (event) => {
try { try {
data.value = JSON.parse(event.data) data.value = JSON.parse(event.data)
if (!rangePinned) {
rangePinned = true
const starts = (data.value?.visits || [])
.map((v) => Date.parse(v.start))
.filter((t) => !Number.isNaN(t))
if (starts.length && Date.now() - Math.min(...starts) < 24 * 3600 * 1000) {
range.value = 'day'
}
}
} catch { } catch {
error.value = 'analytics data could not be loaded' error.value = 'analytics data could not be loaded'
} }
@@ -52,7 +82,7 @@ function connectAnalytics() {
onMounted(async () => { onMounted(async () => {
connectAnalytics() connectAnalytics()
now.value = Date.now() now.value = Date.now()
timeInterval = setInterval(() => { now.value = Date.now() }, 30000) timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
// The site tree for the transition map (all pages in menu order). Not // The site tree for the transition map (all pages in menu order). Not
// fatal: without it the map just narrows to pages seen in transitions. // fatal: without it the map just narrows to pages seen in transitions.
try { try {
@@ -71,34 +101,40 @@ onUnmounted(() => {
} }
}) })
const visits = computed(() => data.value?.visits || []) const window = computed(() => rangeWindow(range.value))
const totalViews = computed(() => calcTotalViews(data.value?.views))
const range = ref(RANGES[props.initialRange] ? props.initialRange : 'week') // All non-chart stats follow the selected range; the charts keep their own
// range-specific x windows (week overlays previous weeks aligned to Monday).
const rangeData = computed(() => {
if (!data.value) return null
const { t0, t1 } = window.value
return {
...data.value,
transitions: filterTransitionsByRange(data.value.transitions, t0, t1),
views: filterViewsByRange(data.value.views, t0, t1),
visits: filterRecordsByRange(data.value.visits, t0, t1),
crawlers: filterRecordsByRange(data.value.crawlers, t0, t1),
abuse: filterRecordsByRange(data.value.abuse, t0, t1),
}
})
const visits = computed(() => rangeData.value?.visits || [])
const totalViews = computed(() => calcTotalViews(rangeData.value?.views))
const readStats = computed(() => calcReadStats(visits.value))
// Keep the URL shareable when the range changes. // Keep the URL shareable when the range changes.
watch(range, (r) => { watch(range, (r) => {
const url = new URL(location.href) const url = new URL(location.href)
url.searchParams.set('range', r) url.hash = r
history.replaceState(null, '', url) history.replaceState(null, '', url)
}) })
const visitRows = computed(() => formatVisitRows(visits.value, pageTree.value, now.value)) const clients = computed(() => data.value?.clients || {})
const crawlers = computed(() => data.value?.crawlers || []) const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, pageTree.value, now.value)) const crawlers = computed(() => rangeData.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
const abuseRows = computed(() => formatAbuseRows(rangeData.value?.abuse || [], clients.value, now.value))
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script> </script>
<template> <template>
@@ -117,13 +153,15 @@ function countryName(code) {
<p v-if="error" class="error"> {{ error }}</p> <p v-if="error" class="error"> {{ error }}</p>
<p v-else-if="!data" class="loading">loading</p> <p v-else-if="!data" class="loading">loading</p>
<template v-else> <template v-else>
<section class="totals"> <section class="totals" :style="{ marginLeft: CHART_MARGIN }">
<div><strong>{{ visits.length }}</strong> visits</div> <div><strong :title="String(visits.length)">{{ formatCount(visits.length) }}</strong> visits</div>
<div><strong>{{ totalViews }}</strong> page views</div> <div><strong :title="String(totalViews)">{{ formatCount(totalViews) }}</strong> page views</div>
<div><strong>{{ readStats.avgMinPerVisit }}</strong> min/visit</div>
<div><strong>{{ readStats.avgArticleMedianMin }}</strong> min/read</div>
</section> </section>
<VisitorCharts :data="data" :range="range" /> <VisitorCharts :data="data" :range="range" />
<TransitionGraph :data="data" :range="range" :page-tree="pageTree" /> <TransitionGraph :data="rangeData" :window="window" :page-tree="pageTree" />
<section> <section>
<h2>Recent visits</h2> <h2>Recent visits</h2>
@@ -131,83 +169,112 @@ function countryName(code) {
<table class="visit-table"> <table class="visit-table">
<thead> <thead>
<tr> <tr>
<th>when</th>
<th>trail</th> <th>trail</th>
<th>referer</th> <th>visitor</th>
<th>ip</th> <th class="last-seen">last seen</th>
<th>lang</th>
<th>country</th>
<th>ua</th>
<th>utm</th>
</tr> </tr>
</thead> </thead>
<tbody> <tbody>
<tr v-for="(v, i) in visitRows" :key="i"> <tr v-for="(v, i) in visitRows" :key="i">
<td class="when" :title="v.whenTooltip">{{ v.when }}</td>
<td class="trail"> <td class="trail">
<a v-for="(s, si) in v.trail" :key="si" <TrailLink v-if="v.refererStep" :step="v.refererStep" @close="$emit('close')" />
:href="s.path" :title="s.title" <span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span>
:target="s.external ? '_blank' : undefined" <TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" @close="$emit('close')" />
:rel="s.external ? 'noopener' : undefined"
@click="(e) => { if (!s.external) $emit('close') }">
{{ s.slug }}
</a>
</td> </td>
<td>{{ v.referer }}</td> <VisitorCell
<td> :ip="v.ip"
<span class="clickable-ip" :ip-display="v.ipDisplay"
:title="`Click to copy full IP: ${v.ip}`" :ua="v.ua"
@click="copyIp(v.ip)">{{ v.ipDisplay }}</span> :ua-raw="v.uaRaw"
</td> :country="v.country"
<td>{{ v.lang }}</td> :city="v.city"
<td class="country"> :lang="v.lang"
<span v-if="flagSvg(v.country)" class="flag" v-html="flagSvg(v.country)" :title="countryName(v.country) || v.country"></span> :lang-display="v.langDisplay"
<template v-if="v.city !== '—'">{{ countryName(v.country) || v.country }}<br><small class="muted">{{ v.city }}</small></template> :is-host="v.isHost"
<template v-else-if="v.country !== '—'">{{ countryName(v.country) || v.country }}</template> />
<template v-else></template> <td class="last-seen muted"
</td> :title="v.lastSeenLocal"
<td class="ua" :title="v.uaRaw">{{ v.ua }}</td> @click="copyList(v.lastSeenIso, $event)">{{ v.lastSeen }}</td>
<td>{{ v.utm }}</td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
</div> </div>
<p v-else class="empty">no visits recorded yet</p> <p v-else class="empty">no visits recorded yet</p>
</section>
<section>
<h2>Crawlers</h2>
<div v-if="crawlerRows.length" class="visit-table-wrap"> <div v-if="crawlerRows.length" class="visit-table-wrap">
<table class="visit-table"> <table class="visit-table">
<thead> <thead>
<tr> <tr>
<th>when</th> <th>pages crawled</th>
<th>pages</th> <th>visitor</th>
<th>ip</th> <th class="last-seen">last seen</th>
<th>ua</th>
</tr> </tr>
</thead> </thead>
<tbody> <tbody>
<tr v-for="(c, i) in crawlerRows" :key="i"> <tr v-for="(c, i) in crawlerRows" :key="i">
<td class="when" :title="c.whenTooltip">{{ c.when }}</td>
<td class="trail"> <td class="trail">
<a v-for="(s, si) in c.pages" :key="si" <TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
:href="s.path" :title="`${s.title}${s.count > 1 ? ` (${s.count} hits)` : ''}`"
@click="$emit('close')">
<small v-if="s.count > 1" class="muted">{{ s.count }}×</small>{{ s.slug }}
</a>
</td> </td>
<td> <VisitorCell
<span class="clickable-ip" :ip="c.ip"
:title="`Click to copy full IP: ${c.ip}`" :ip-display="c.ipDisplay"
@click="copyIp(c.ip)">{{ c.ipDisplay }}</span> :ua="c.ua"
</td> :ua-raw="c.uaRaw"
<td class="ua" :title="c.uaRaw">{{ c.ua }}</td> :country="c.country"
:city="c.city"
:lang="c.lang"
:lang-display="c.langDisplay"
:is-host="c.isHost"
/>
<td class="last-seen muted"
:title="c.lastSeenLocal"
@click="copyList(c.lastSeenIso, $event)">{{ c.lastSeen }}</td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
</div> </div>
<p v-else class="empty">no crawler hits recorded yet</p> <p v-else class="empty">no crawler hits recorded yet</p>
<div v-if="abuseRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>paths abused</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(a, i) in abuseRows" :key="i">
<td class="trail abuse-list clickable-list"
@click="copyList(a.allPaths, $event)">
<div class="abuse-items">
<span v-for="(p, pi) in a.paths.slice(0, ABUSE_MAX_LINES)" :key="pi"
class="inline-item">
<small v-if="p.count > 1" class="muted">{{ formatCount(p.count) }}×</small>{{ p.path }}
</span>
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
</div>
</td>
<VisitorCell
:ip="a.ip"
:ip-display="a.ipDisplay"
:ua="a.ua"
:ua-raw="a.uaRaw"
:country="a.country"
:city="a.city"
:lang="a.lang"
:lang-display="a.langDisplay"
:is-host="a.isHost"
:variant-count="a.clientCount"
/>
<td class="last-seen muted"
:title="a.lastSeenLocal"
@click="copyList(a.lastSeenIso, $event)">{{ a.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
</section> </section>
</template> </template>
</div> </div>
@@ -222,9 +289,12 @@ function countryName(code) {
} }
.analytics-panel { .analytics-panel {
margin: 0 auto; margin: 0;
width: min(60rem, 96vw); width: 100%;
padding: 1.5rem 2rem 4rem; /* Same 1.25rem side spacing as main's article padding. */
padding: 1.5rem 1.25rem 4rem;
/* Container for cqw-based shrink-to-fit (see .totals). */
container-type: inline-size;
} }
.analytics-panel header { .analytics-panel header {
@@ -247,7 +317,7 @@ function countryName(code) {
.ranges button { .ranges button {
padding: 0.2rem 0.7rem; padding: 0.2rem 0.7rem;
font: inherit; font: inherit;
font-size: 0.85rem; font-size: 0.9rem;
color: var(--muted); color: var(--muted);
background: none; background: none;
border: 1px solid var(--line); border: 1px solid var(--line);
@@ -282,12 +352,24 @@ function countryName(code) {
margin-top: 1.8rem; margin-top: 1.8rem;
} }
.analytics-view a {
color: var(--text);
text-decoration: none;
}
.analytics-view a:hover { color: var(--accent); }
.analytics-view :deep(.muted) { color: var(--muted); }
.analytics-view :deep(.small) { font-size: 0.75em; }
/* One line at any width: the gap shrinks first, then the font (the number
scales along in em), both following the panel's container width. */
.totals { .totals {
display: flex; display: flex;
gap: 2rem; gap: clamp(0.5rem, 3cqw, 2rem);
font-size: 1.1rem; font-size: clamp(0.6rem, 2.2cqw, 1.1rem);
white-space: nowrap;
} }
.totals strong { font-size: 1.5rem; } .totals strong { font-size: 1.36em; }
.visit-table-wrap { .visit-table-wrap {
overflow-x: auto; overflow-x: auto;
@@ -296,8 +378,7 @@ function countryName(code) {
.visit-table { .visit-table {
width: 100%; width: 100%;
border-collapse: collapse; border-collapse: collapse;
font-family: monospace; font-size: 0.9rem;
font-size: 0.82rem;
line-height: 1.3; line-height: 1.3;
} }
@@ -318,9 +399,11 @@ function countryName(code) {
background: var(--bg, Canvas); background: var(--bg, Canvas);
} }
.visit-table .when { .visit-table .last-seen {
width: 6rem;
text-align: right;
white-space: nowrap; white-space: nowrap;
color: var(--muted); cursor: pointer;
} }
.visit-table .trail { .visit-table .trail {
@@ -328,54 +411,79 @@ function countryName(code) {
overflow-wrap: break-word; overflow-wrap: break-word;
} }
.visit-table .trail a { .visit-table .trail a,
color: var(--text); .visit-table .trail-link {
text-decoration: none; display: inline-block;
max-width: 8rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
} }
.visit-table .trail a:hover { color: var(--accent); } .visit-table .trail > * + * {
.visit-table .trail a + a {
margin-left: 0.5rem; margin-left: 0.5rem;
} }
.visit-table .trail small, .analytics-view :deep(.trail-link.error),
.visit-table small.muted { .analytics-view :deep(.trail-link.error:hover) {
color: var(--muted); color: var(--error, #c00);
font-size: 0.75em;
} }
.visit-table .clickable-ip { .visit-table .utm-tag {
cursor: pointer; display: inline-block;
text-decoration: underline; max-width: 100%;
text-decoration-style: dotted; padding: 0.05rem 0.4rem;
} border: 1px solid var(--line);
border-radius: 0.25rem;
.visit-table .clickable-ip:hover { white-space: nowrap;
color: var(--accent);
}
.visit-table .ua {
max-width: 18rem;
overflow: hidden; overflow: hidden;
text-overflow: ellipsis; text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .clickable-list {
cursor: pointer;
max-width: 22rem;
}
.visit-table .abuse-items {
display: flex;
flex-wrap: wrap;
gap: 0.15rem 0.5rem;
align-items: baseline;
}
.visit-table .inline-item {
max-width: 18rem;
min-width: 0;
white-space: nowrap; white-space: nowrap;
}
.visit-table .country .flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden; overflow: hidden;
border: 1px solid var(--line); text-overflow: ellipsis;
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset; word-break: keep-all;
hyphens: none;
} }
.visit-table .country .flag :deep(svg) { .visit-table :deep(.clickable-ip),
width: 100%; .visit-table .clickable-list,
height: 100%; .visit-table .last-seen {
display: block; cursor: pointer;
position: relative;
}
.visit-table :deep(.copy-popup) {
position: absolute;
bottom: calc(100% + 0.25rem);
left: 50%;
transform: translateX(-50%);
padding: 0.15rem 0.4rem;
background: var(--text, CanvasText);
color: var(--bg, Canvas);
border-radius: 0.25rem;
font-size: 0.75rem;
white-space: nowrap;
pointer-events: none;
z-index: 10;
} }
.crawler-top-uas { .crawler-top-uas {
+29 -16
View File
@@ -242,30 +242,43 @@ async function saveSettings(opts = {}) {
async function onThemeChange() { async function onThemeChange() {
await saveSettings() await saveSettings()
// Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev // Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev
// and prod: swap the link in place, then re-render (the theme's default // and prod, but rendered differently: a <link> in dev, an inline <style>
// banner design and the page's stylesheet links may change with it). // in prod. Swap it in place, then re-render (the theme's default banner
let link = document.getElementById('pagerite-theme') // design and the page's stylesheets may change with it).
let el = document.getElementById('pagerite-theme')
const url = `/_themes/${theme.value}/theme.css`
if (theme.value) { if (theme.value) {
const href = `/_themes/${theme.value}/theme.css` if (el?.tagName === 'STYLE') {
if (link) { el.textContent = await (await fetch(url)).text()
link.href = href } else if (el) {
} else { el.href = url
} else if (import.meta.env.DEV) {
// Re-create after "none": keep base < theme < design < custom CSS. // Re-create after "none": keep base < theme < design < custom CSS.
// In dev there is no #pagerite-base link (the base is a // In dev there is no #pagerite-base element (the base is a
// Vite-injected <style>), so anchor to the next sheet instead of // Vite-injected <style>), so anchor to the next sheet instead of
// prepending before the base styles. // prepending before the base styles.
link = document.createElement('link') el = document.createElement('link')
link.rel = 'stylesheet' el.rel = 'stylesheet'
link.id = 'pagerite-theme' el.id = 'pagerite-theme'
link.href = href el.href = url
const before = document.getElementById('pagerite-base')?.nextSibling const before = document.getElementById('pagerite-base')?.nextSibling
?? document.getElementById('pagerite-banner') ?? document.getElementById('pagerite-banner')
?? document.getElementById('pagerite-user') ?? document.getElementById('pagerite-user')
if (before) before.before(link) if (before) before.before(el)
else document.head.append(link) else document.head.append(el)
} else {
// Prod: inline <style>, fetched from the backend-served URL.
el = document.createElement('style')
el.id = 'pagerite-theme'
el.textContent = await (await fetch(url)).text()
const before = document.getElementById('pagerite-base')?.nextSibling
?? document.getElementById('pagerite-banner')
?? document.getElementById('pagerite-user')
if (before) before.before(el)
else document.head.append(el)
} }
} else if (link) { } else if (el) {
link.remove() el.remove()
} }
loadPlain(path.value) loadPlain(path.value)
} }
+37
View File
@@ -0,0 +1,37 @@
<script setup>
import { computed } from 'vue'
import { formatCount, formatReadTime } from './analytics/format.js'
const props = defineProps({
step: { type: Object, required: true },
count: { type: Number, default: 0 },
})
defineEmits(['close'])
const hasError = computed(() => props.step.status >= 400)
const title = computed(() => {
const parts = [props.step.title]
if (props.step.readSeconds > 0) {
parts.push(formatReadTime(props.step.readSeconds))
}
if (hasError.value) {
parts.push(`${props.step.status}`)
}
return parts.filter(Boolean).join(' — ')
})
</script>
<template>
<a class="trail-link"
:class="{ error: hasError }"
:href="step.path"
:title="title"
:target="step.external ? '_blank' : undefined"
:rel="step.external ? 'noopener' : undefined"
@click="(e) => { if (!step.external) $emit('close') }">
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
<span>{{ step.slug }}</span>
</a>
</template>
+200 -92
View File
@@ -1,126 +1,212 @@
<script setup> <script setup>
/** /**
* Radial transition map filtered to the selected time range. * Radial transition map for a pre-filtered time range.
* *
* Transitions are stored per 5-minute bucket (from -> to -> bucket -> * The parent filters transitions, views and visits to the selected range
* count), so the graph sums the buckets falling inside the selected * before passing them in; `window` carries the absolute [t0, t1) window
* range, exactly like the charts and per-page views do. * so the visual scale can normalize against a one-week reference.
*/ */
import { computed, onBeforeUnmount, shallowRef, watch } from 'vue' import { computed, onBeforeUnmount, onMounted, shallowRef, watch } from 'vue'
import { rangeWindow } from './analytics/time.js' import { DAY, WEEK } from './analytics/time.js'
import { formatCount } from './analytics/format.js'
import { import {
TNODE_R, TNODE_W,
TNODE_H,
BEAD_R, BEAD_R,
BEAD_SPEED,
buildTransitionGraph, buildTransitionGraph,
filterTransitionsByRange,
filterViewsByRange,
} from './analytics/transitions.js' } from './analytics/transitions.js'
const props = defineProps({ const props = defineProps({
data: { type: Object, default: null }, data: { type: Object, default: null },
range: { type: String, required: true }, window: { type: Object, required: true },
pageTree: { type: Array, default: null }, pageTree: { type: Array, default: null },
}) })
const window = computed(() => rangeWindow(props.range)) const visualScale = computed(() => {
const { t0, t1 } = props.window
const filteredData = computed(() => { if (t0 != null && t1 != null) return WEEK / (t1 - t0)
if (!props.data) return null // 'all': scale by the actual data span, but never less than the 30-day
const { t0, t1 } = window.value // minimum the plot enforces, so sparse young data is not over-amplified.
return { const times = new Set()
transitions: filterTransitionsByRange(props.data.transitions, t0, t1), for (const buckets of Object.values(props.data?.views || {})) {
views: filterViewsByRange(props.data.views, t0, t1), for (const k of Object.keys(buckets)) times.add(Date.parse(k))
} }
const arr = [...times]
if (arr.length < 2) return 1
const span = Math.max(...arr) - Math.min(...arr)
return WEEK / Math.max(span, 30 * DAY)
}) })
const graph = computed(() => const graph = computed(() =>
filteredData.value props.data
? buildTransitionGraph(filteredData.value, props.pageTree) ? buildTransitionGraph(props.data, props.pageTree, props.data.visits || [], visualScale.value)
: null, : null,
) )
// Bead animation: every bead is simulated independently in JS. Each flow // Bead animation: every bead is simulated independently in JS. Each flow
// (one per edge direction) emits a bead every `interval` seconds; beads // (one per edge direction) emits a bead every `interval` seconds; beads
// travel at BEAD_SPEED along the segment and are dropped at the end. // cross their segment in a constant TRAVERSAL_S seconds (speed relative
// to span length) and are dropped at the end.
// There is deliberately no cap on beads in flight. // There is deliberately no cap on beads in flight.
// Emitters persist across data reloads, keyed by flow.key: an unchanged
// link keeps its emission phase and in-flight beads (tracked by progress,
// not absolute time), so a count change elsewhere never reshuffles them.
const beads = shallowRef([]) const beads = shallowRef([])
let rafId = 0 let rafId = 0
const emitters = new Map() // flow.key -> { flow, interval, next, alive }
const live = [] // { e, p } — beads in flight, p = progress 0..1
let lastTick = 0
const startBeads = (flows) => { const MAX_BEAD_RATE = 120 // upper bound on total beads per second
cancelAnimationFrame(rafId) const TRAVERSAL_S = 1.5 // seconds to cross any segment, end to end
beads.value = []
if (!flows?.length) return
if (matchMedia('(prefers-reduced-motion: reduce)').matches) return
const live = [] // { flow, t0 } — one entry per bead in flight const syncBeads = (flows) => {
const now = performance.now() const reduced = matchMedia('(prefers-reduced-motion: reduce)').matches
const emitters = flows.map((flow) => { if (!flows?.length || reduced) {
const interval = flow.interval * 1000 emitters.clear()
// Pre-fill the traversal with evenly spaced beads (random phase), so live.length = 0
// the flow appears already running instead of starting empty. beads.value = []
const phase = Math.random() * interval return
for (let t = now - (flow.len / BEAD_SPEED) * 1000 + phase; t <= now; t += interval) {
live.push({ flow, t0: t })
}
return { flow, interval, next: now + phase }
})
const tick = (t) => {
for (const e of emitters) {
while (e.next <= t) {
live.push({ flow: e.flow, t0: e.next })
e.next += e.interval
}
}
const out = []
for (let i = live.length - 1; i >= 0; i--) {
const b = live[i]
const p = ((t - b.t0) / 1000) * BEAD_SPEED / b.flow.len
if (p >= 1) {
live.splice(i, 1)
continue
}
out.push({
x: b.flow.x1 + (b.flow.x2 - b.flow.x1) * p,
y: b.flow.y1 + (b.flow.y2 - b.flow.y1) * p,
})
}
beads.value = out
rafId = requestAnimationFrame(tick)
} }
// Cap the total bead emission rate so a busy range cannot spawn enough
// beads to kill the page. Existing per-range time scaling is preserved;
// this is only a proportional emergency throttle when the limit is hit.
const totalRate = flows.reduce((s, f) => s + 1 / f.interval, 0)
const scale = totalRate > MAX_BEAD_RATE ? MAX_BEAD_RATE / totalRate : 1
const now = performance.now()
const seen = new Set()
for (const flow of flows) {
seen.add(flow.key)
const interval = (flow.interval / scale) * 1000
const e = emitters.get(flow.key)
if (e) {
e.flow = flow // pick up new geometry/rate, keep the phase
e.interval = interval
continue
}
// New emitter: pre-fill the traversal with evenly spaced beads (random
// phase), so the flow appears already running instead of empty.
const phase = Math.random() * interval
const dp = interval / 1000 / TRAVERSAL_S
const ne = { flow, interval, next: now + phase, alive: true }
for (let p = 1 - phase / 1000 / TRAVERSAL_S; p > 0; p -= dp) {
live.push({ e: ne, p })
}
emitters.set(flow.key, ne)
}
for (const [key, e] of emitters) {
if (!seen.has(key)) {
e.alive = false
emitters.delete(key)
}
}
for (let i = live.length - 1; i >= 0; i--) {
if (!live[i].e.alive) live.splice(i, 1)
}
}
const tick = (t) => {
const dt = lastTick ? (t - lastTick) / 1000 : 0
lastTick = t
for (const e of emitters.values()) {
while (e.next <= t) {
live.push({ e, p: 0 })
e.next += e.interval
}
}
const out = []
for (let i = live.length - 1; i >= 0; i--) {
const b = live[i]
b.p += dt / TRAVERSAL_S
if (b.p >= 1) {
live.splice(i, 1)
continue
}
const f = b.e.flow
out.push({ x: f.x1 + (f.x2 - f.x1) * b.p, y: f.y1 + (f.y2 - f.y1) * b.p })
}
beads.value = out
rafId = requestAnimationFrame(tick) rafId = requestAnimationFrame(tick)
} }
watch(() => graph.value?.flows, startBeads, { immediate: true }) watch(() => graph.value?.flows, syncBeads, { immediate: true })
onMounted(() => {
if (!matchMedia('(prefers-reduced-motion: reduce)').matches) {
rafId = requestAnimationFrame(tick)
}
})
onBeforeUnmount(() => cancelAnimationFrame(rafId)) onBeforeUnmount(() => cancelAnimationFrame(rafId))
// The svg never renders larger than its natural size (1 viewBox unit = 1
// px, max-width below): the layout geometry is designed in pixel-like
// units, and upscaling would blow up the pills around their text. Narrow
// panels scale the graph down to fit (width: 100%), text along with it.
// Pill text is not truncated: text is clipped at the pill's rounded border
// (clipPath per node, inset a few units for padding). Captions center when
// they fit; overlong ones anchor left so their beginning (not their
// middle) survives the clip. Width estimate: ~0.52 em per glyph.
const fitsPill = (label, fontPx = 19) => label.length * 0.52 * fontPx <= TNODE_W - 16
const countLabel = (n) =>
n.readMin ? `${formatCount(n.views)}×${n.readMin}m` : formatCount(n.views)
</script> </script>
<template> <template>
<section v-if="graph"> <section v-if="graph">
<svg class="tmap" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`" <svg class="tmap" :style="{ maxWidth: `${graph.bounds.x1 - graph.bounds.x0}px` }" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
role="img" aria-label="map of transitions between pages"> role="img" aria-label="map of transitions between pages">
<path v-for="(a, i) in graph.arcs" :key="'a' + i" <path v-for="(a, i) in graph.arcs" :key="'a' + i"
:d="a.d" class="tarc" /> :id="`tarc${i}`" :d="a.d" :class="['tarc', a.top && 'tarc-top']" />
<template v-for="(a, i) in graph.arcs" :key="'t' + i">
<path v-if="a.ld" :id="`tarcl${i}`" :d="a.ld" fill="none" stroke="none" />
<text v-if="a.ld" class="tarclabel" :class="{ 'tarclabel-top': a.top }"><textPath :href="`#tarcl${i}`" startOffset="0">{{ a.label }}</textPath></text>
</template>
<path v-for="(e, i) in graph.edges" :key="'e' + i" <path v-for="(e, i) in graph.edges" :key="'e' + i"
:d="e.d" class="tconn"> :d="e.d" :class="['tconn', e.external && 'tconn-exit']">
<title>{{ e.title }}</title> <title>{{ e.title }}</title>
</path> </path>
<circle v-for="(b, i) in beads" :key="'b' + i" <circle v-for="(b, i) in beads" :key="'b' + i"
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" /> :cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
<g v-for="(x, i) in graph.extNodes" :key="'x' + i"> <g v-for="(x, i) in graph.extNodes" :key="'x' + i">
<a :href="x.path" target="_blank" rel="noopener" :title="x.path"> <clipPath :id="`xclip${i}`">
<circle :cx="x.x" :cy="x.y" :r="x.r" <rect :x="x.x - TNODE_W/2 + 6" :y="x.y - TNODE_H/2" :width="TNODE_W - 12"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" /> :height="TNODE_H" :rx="TNODE_H/2 - 4" />
<text :x="x.x" :y="x.y - 2" class="tnodeslug">{{ x.label }}</text> </clipPath>
<text :x="x.x" :y="x.y + 12" class="tnodecount">{{ x.count }}</text> <a v-if="x.href" :href="x.href" target="_blank" rel="noopener">
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<text :x="fitsPill(x.label) ? x.x : x.x - TNODE_W/2 + 8" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(x.label) ? 'middle' : 'start' }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</a> </a>
<g v-else>
<title>{{ x.path }}</title>
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<g :clip-path="`url(#xclip${i})`">
<text :x="fitsPill(x.label) ? x.x : x.x - TNODE_W/2 + 8" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(x.label) ? 'middle' : 'start' }">{{ x.label }}</text>
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
</g>
</g>
</g> </g>
<g v-for="n in graph.nodes" :key="n.path"> <g v-for="(n, i) in graph.nodes" :key="n.path">
<a :href="n.path" :title="n.title"> <clipPath :id="`nclip${i}`">
<circle :cx="n.x" :cy="n.y" :r="TNODE_R" class="tnode" /> <rect :x="n.x - TNODE_W/2 + 6" :y="n.y - TNODE_H/2" :width="TNODE_W - 12"
<text :x="n.x" :y="n.y - 2" class="tnodeslug">{{ n.label }}</text> :height="TNODE_H" :rx="TNODE_H/2 - 4" />
<text :x="n.x" :y="n.y + 12" class="tnodecount">{{ n.views }}</text> </clipPath>
<a :href="n.path">
<title>{{ n.title }}</title>
<rect :x="n.x - TNODE_W/2" :y="n.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2" class="tnode" />
<g :clip-path="`url(#nclip${i})`">
<text :x="fitsPill(n.label) ? n.x : n.x - TNODE_W/2 + 8" :y="n.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle" :style="{ textAnchor: fitsPill(n.label) ? 'middle' : 'start' }">{{ n.label }}</text>
<text :x="n.x" :y="n.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">
{{ countLabel(n) }}
</text>
</g>
</a> </a>
</g> </g>
</svg> </svg>
@@ -132,44 +218,66 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
.tmap { .tmap {
display: block; display: block;
width: 100%; width: 100%;
max-width: 36rem; /* max-width is set inline to the natural content width (px = viewBox
units), so wide panels never upscale the graph beyond 1:1. */
margin: 0 auto; margin: 0 auto;
} }
.tmap .tconn { .tmap .tconn {
fill: var(--accent); fill: var(--accent);
opacity: 0.4; /* uniform, not strength-encoded: width carries that */ opacity: 0.4; /* uniform, not strength-encoded: width carries that */
} }
.tmap .tconn-exit {
fill: var(--text);
}
.tmap .tbead { .tmap .tbead {
fill: var(--accent); fill: var(--accent);
opacity: 0.85; opacity: 0.85;
filter: drop-shadow(0 0 2.5px var(--accent)); filter: drop-shadow(0 0 2.5px var(--accent));
} }
.tmap .txnode { .tmap .txnode {
fill: var(--bg, Canvas); fill: var(--text);
stroke-width: 1.5; stroke: none;
} }
.tmap .txnode-source { stroke: var(--text); } .tmap .txnode-source { fill: var(--text); }
.tmap .txnode-exit { stroke: var(--muted); } .tmap .txnode-exit { fill: var(--text); }
/* Branch lanes: one wide concentric arc per path prefix, running behind
the node pills around the fan's circle center; parent levels sit one
indent (radius step) outward. Each lane's label follows a short guide
arc across the first inter-node gap (the part pills never cover). */
.tmap .tarc { .tmap .tarc {
fill: none; fill: none;
stroke: var(--line); stroke: var(--muted);
stroke-width: 1; stroke-width: 16;
opacity: 0.25;
}
.tmap .tarc-top { stroke-width: 24; }
/* Lane labels are left-aligned: each guide arc starts just past the source
pill's edge, the earliest point where the text is visible. */
.tmap .tarclabel {
fill: var(--muted);
font-size: 13px;
text-anchor: start;
}
/* The top lane is 50% thicker; its 🏠︎ label scales along. */
.tmap .tarclabel-top {
font-size: 19.5px;
} }
.tmap .tnode { .tmap .tnode {
fill: var(--bg, Canvas); fill: var(--accent);
stroke: var(--accent); stroke: none;
stroke-width: 1.5;
} }
/* Text sizes are viewBox units: they shrink along with the graph on
narrow panels. Overlong labels are clipped at the pill border. */
.tmap .tnodeslug { .tmap .tnodeslug {
fill: var(--text); fill: var(--bg, Canvas);
font-size: 11px; font-size: 19px;
text-anchor: middle; text-anchor: start;
} }
.tmap a { cursor: pointer; } .tmap a { cursor: pointer; }
.tmap a:hover .tnodeslug { fill: var(--accent); }
.tmap .tnodecount { .tmap .tnodecount {
fill: var(--muted); fill: var(--bg, Canvas);
font-size: 10px; opacity: 0.75;
font-size: 15px;
text-anchor: middle; text-anchor: middle;
} }
+156
View File
@@ -0,0 +1,156 @@
<script setup>
// Visitor metadata cell shared by the recent-visits, crawlers, and abuse tables.
// Displays IP/network/host, country flag/city, UA, and language when available.
// Clicking the IP copies the full address to the clipboard.
// ``variantCount`` overrides the UA line to warn when multiple client
// fingerprints share the same IP (e.g. a scanner rotating UAs).
import { computed } from 'vue'
import * as flagSvgs from 'country-flag-icons/string/3x2'
import { copyIp, formatLang } from './analytics/format.js'
const props = defineProps({
ip: { type: String, default: '' },
ipDisplay: { type: String, default: '—' },
ua: { type: String, default: '' },
uaRaw: { type: String, default: '' },
country: { type: String, default: '' },
city: { type: String, default: '' },
lang: { type: String, default: '' },
langDisplay: { type: String, default: '' },
isHost: { type: Boolean, default: false },
variantCount: { type: Number, default: 1 },
})
const hasCountry = computed(() => !!(props.country && props.country !== '—'))
const hasCity = computed(() => !!(props.city && props.city !== '—'))
const hasLocale = computed(() => hasCountry.value || hasCity.value)
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
const showLang = computed(() => langValue.value && langValue.value !== '—')
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script>
<template>
<td class="visitor-cell" :class="{ 'host-cell': isHost }">
<div class="visitor-rows">
<div class="visitor-row">
<div class="locale-line">
<span v-if="flagSvg(country)" class="flag" v-html="flagSvg(country)" :title="countryName(country) || country"></span>
<template v-if="hasCity"><small class="city-name muted">{{ city }}</small></template>
<template v-else-if="!hasLocale"></template>
</div>
<div class="ip-line">
<span class="clickable-ip small muted"
:title="ip"
@click="copyIp(ip, $event)">{{ ipDisplay }}</span>
</div>
</div>
<div class="visitor-row">
<div class="ua-line">
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
</div>
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
</div>
</div>
</td>
</template>
<style scoped>
.visitor-cell {
width: 18em;
max-width: 18em;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: top;
}
.visitor-cell.host-cell {
text-align: right;
}
.visitor-rows {
display: flex;
flex-direction: column;
gap: 0.15rem;
}
.visitor-row {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.5rem;
}
.visitor-row > * {
min-width: 0;
}
.locale-line,
.ip-line,
.ua-line {
flex: 1 1 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.locale-line {
text-align: left;
display: flex;
align-items: center;
gap: 0.3rem;
}
.ip-line {
text-align: right;
}
.ua-line {
text-align: left;
}
.locale-lang {
flex: 0 0 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
text-align: right;
}
.city-name {
display: inline-block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
vertical-align: middle;
}
.flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+116 -103
View File
@@ -2,9 +2,24 @@
/** /**
* Visitor and page-view smoothed curves for a single shared time range. * Visitor and page-view smoothed curves for a single shared time range.
*/ */
import { computed } from 'vue' import { computed, onMounted, onUnmounted, ref } from 'vue'
import { makeSeries } from './analytics/time.js' import { makeSeries } from './analytics/time.js'
import { CHART_H, CHART_W, buildChart, fmtY } from './analytics/chart.js' import {
CHART_H,
CHART_W,
MARGIN_B,
MARGIN_L,
VIEW_H,
VIEW_W,
buildChart,
} from './analytics/chart.js'
const DAY_REFRESH_MS = 15000
// Keep the whole svg within page bounds: full width below the natural
// size, centered with equal side margins above it (max() clamps the
// centering margin to 0 at the breakpoint, so the rule is continuous).
const CHART_MARGIN = `max(0px, calc(50% - ${VIEW_W / 2}px))`
const props = defineProps({ const props = defineProps({
data: { type: Object, default: null }, data: { type: Object, default: null },
@@ -22,108 +37,118 @@ const allViews = computed(() => {
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range)) const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
const viewSeries = computed(() => makeSeries(allViews.value, props.range)) const viewSeries = computed(() => makeSeries(allViews.value, props.range))
const unit = computed(() => (props.range === 'week' ? 'h' : 'day'))
const visitChart = computed(() => buildChart(visitSeries.value)) function freqLabel(unit) {
const viewChart = computed(() => buildChart(viewSeries.value)) return unit === '5min' ? '5 min' : unit === 'hour' ? 'hourly' : 'daily'
}
/** Vertical axis caption: "visits / 5 min" on the day view, else "hourly visits" style. */
function axisLabel(unit, ylabel) {
return unit === '5min' ? `${ylabel} / 5 min` : `${freqLabel(unit)} ${ylabel}`
}
/** Legend label for the overlaid past weeks: "Week M" or "Week MN". */
function pastLabel(series) {
const oldest = series.at(-1).label.slice(5) // strip "Week "
return series.length > 2 ? `Week ${oldest}${series[1].label.slice(5)}` : `Week ${oldest}`
}
const now = ref(Date.now())
let refreshInterval = null
onMounted(() => {
refreshInterval = setInterval(() => { now.value = Date.now() }, DAY_REFRESH_MS)
})
onUnmounted(() => {
if (refreshInterval) clearInterval(refreshInterval)
})
const visitChart = computed(() => buildChart(visitSeries.value, now.value))
const viewChart = computed(() => buildChart(viewSeries.value, now.value))
</script> </script>
<template> <template>
<section v-for="c in [ <section v-for="c in [
{ ylabel: 'visitors', chart: visitChart, empty: 'no visits recorded yet' }, { ylabel: 'visits', chart: visitChart, legend: true, empty: 'no visits recorded yet' },
{ ylabel: 'views', chart: viewChart, empty: 'no views recorded yet' }, { ylabel: 'views', chart: viewChart, legend: false, empty: 'no views recorded yet' },
]" :key="c.ylabel"> ]" :key="c.ylabel">
<template v-if="c.chart"> <template v-if="c.chart">
<div class="chartwrap"> <svg class="chart" :viewBox="`${-MARGIN_L} 0 ${VIEW_W} ${VIEW_H}`"
<div class="plot"> :style="{ maxWidth: `${VIEW_W}px`, marginLeft: CHART_MARGIN }"
<div class="plotarea"> role="img" :aria-label="axisLabel(c.chart.unit, c.ylabel)">
<span class="yaxis-label">{{ c.ylabel }}/{{ unit }}</span> <line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
<svg class="chart" :viewBox="`0 0 ${CHART_W} ${CHART_H}`" :x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
preserveAspectRatio="none" role="img" :aria-label="`${c.ylabel} per ${unit}`"> <template v-for="t in c.chart.xticks" :key="'t' + t.x">
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value" <line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" /> class="minor vertical" />
<template v-for="t in c.chart.xticks" :key="'t' + t.x"> </template>
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H" <template v-if="c.chart.bars">
class="minor vertical" /> <rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
</template> :x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
<template v-for="(s, i) in c.chart.series" :key="i"> <path :d="c.chart.skyline" class="line" />
<path v-if="s.area" :d="s.area" class="area" /> </template>
<path :d="s.line" class="line" :style="{ opacity: s.opacity }" /> <template v-else>
</template> <!-- Oldest overlay weeks first so the current week paints on top. -->
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5" <template v-for="(s, i) in [...c.chart.series].reverse()" :key="i">
class="axis" /> <path v-if="s.area" :d="s.area" class="area" />
</svg> <path :d="s.line" class="line" :class="{ past: s.past }"
<span v-for="g in c.chart.majors" :key="g.value" class="ylab" :style="{ opacity: s.opacity }" />
:style="{ bottom: g.bottom + '%' }">{{ fmtY(g.value) }}</span> </template>
</div> </template>
<div class="xlabels"> <line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
<span v-for="t in c.chart.xticks" :key="t.x" class="xlab" class="axis" />
:style="{ left: t.left + '%' }">{{ t.label }}</span> <text v-for="g in c.chart.majors" :key="'y' + g.value" x="-5" :y="g.y"
</div> text-anchor="end" dominant-baseline="middle" class="ylab">{{ g.label }}</text>
</div> <text :x="-(MARGIN_L - 10)" :y="CHART_H / 2" text-anchor="middle"
</div> :transform="`rotate(-90 ${-(MARGIN_L - 10)} ${CHART_H / 2})`"
<div v-if="c.chart.series.length > 1" class="legend"> class="yaxis-label">{{ axisLabel(c.chart.unit, c.ylabel) }}</text>
<span v-for="(s, i) in c.chart.series" :key="i" :style="{ opacity: s.opacity }"> <text v-for="t in c.chart.xticks" :key="'x' + t.x" :x="t.x" :y="CHART_H + MARGIN_B - 8"
{{ s.label }} text-anchor="middle" class="xlab">{{ t.label }}</text>
</span> <!-- Week overlay legend, top right inside the plot: current week in
</div> accent, one muted specimen for the whole past range. -->
<g v-if="c.legend && c.chart.series.length > 1">
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="10" y2="10" class="line" />
<text :x="CHART_W - 72" y="10" dominant-baseline="middle"
class="leglab">{{ c.chart.series[0].label }}</text>
<line :x1="CHART_W - 98" :x2="CHART_W - 78" y1="25" y2="25"
class="line past" style="opacity: 0.6" />
<text :x="CHART_W - 72" y="25" dominant-baseline="middle"
class="leglab">{{ pastLabel(c.chart.series) }}</text>
</g>
</svg>
</template> </template>
<p v-else class="empty">{{ c.empty }}</p> <p v-else class="empty">{{ c.empty }}</p>
</section> </section>
</template> </template>
<style scoped> <style scoped>
/* The svg is stretched (preserveAspectRatio none), so all text lives in /* Each chart is a self-contained SVG: the viewBox includes the axis label
HTML overlays positioned by the same fractions the geometry uses. */ margins, so nothing is positioned with HTML overlays. Never upscale past
.chartwrap { the natural size (1 viewBox unit = 1 px, max-width set inline) — that
padding-left: 2.2rem; /* y labels */ would blow up the constant-size text; smaller panels still scale the
} chart down to fit. The margin-left (set inline) centers the chart above
its natural width; the svg always stays within page bounds.
.plot { overflow: visible lets wider fonts extend past the viewBox instead of
display: flex; clipping. */
flex-direction: column;
width: 100%;
}
.plotarea {
position: relative;
height: 8rem;
}
.xlabels {
position: relative;
height: 1.2rem;
}
.chart { .chart {
display: block; display: block;
width: 100%; width: 100%;
height: 100%; height: auto;
overflow: visible;
} }
.ylab { .chart .ylab,
position: absolute; .chart .xlab,
left: -2.2rem; .chart .yaxis-label,
width: 1.9rem; .chart .leglab {
text-align: right; font-family: system-ui, sans-serif; /* theme fonts can be overly styled */
transform: translateY(50%); font-size: 11px;
font-size: 0.7rem; fill: var(--muted);
color: var(--muted); }
.chart .ylab {
font-variant-numeric: tabular-nums; font-variant-numeric: tabular-nums;
} }
.xlab {
position: absolute;
top: 0.25rem;
transform: translateX(-50%);
font-size: 0.7rem;
color: var(--muted);
white-space: nowrap;
}
.xlabels .xlab:first-child { transform: none; }
.xlabels .xlab:last-child { transform: translateX(-100%); }
.chart .minor { .chart .minor {
stroke: var(--line); stroke: var(--line);
stroke-width: 1; stroke-width: 1;
@@ -154,6 +179,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
opacity: 0.15; opacity: 0.15;
} }
.chart .bar {
fill: var(--accent);
opacity: 0.15;
}
.chart .line { .chart .line {
fill: none; fill: none;
stroke: var(--accent); stroke: var(--accent);
@@ -163,27 +193,10 @@ const viewChart = computed(() => buildChart(viewSeries.value))
stroke-linecap: round; stroke-linecap: round;
} }
.legend { /* Past overlay weeks contrast with the current week's accent color. */
display: flex; .chart .line.past {
gap: 1.2rem; stroke: var(--muted);
margin-top: 0.4rem;
font-size: 0.75rem;
color: var(--muted);
} }
.legend span { color: var(--accent); }
.yaxis-label {
position: absolute;
top: 50%;
left: -2.2rem;
font-size: 0.7rem;
color: var(--muted);
writing-mode: vertical-rl;
transform: translateY(-50%) rotate(180deg);
}
section { margin-top: 1.8rem; }
.empty { color: var(--muted); } .empty { color: var(--muted); }
</style> </style>
+12 -6
View File
@@ -1,6 +1,9 @@
// Analytics page entry: mounts AnalyticsView inside the normal page layout. // Analytics page entry: mounts AnalyticsView inside the normal page layout.
// The backend renders #analytics-app inside #main and links this module for // In production the backend inlines this module into the /_a page (and
// the initial load; pagerite.js also imports it on fetch-navigation to /_a. // pagerite.js re-creates the script element after fetch-navigations there);
// in dev pagerite.js imports it from the Vite dev server on demand. Either
// way it auto-mounts on #analytics-app when it evaluates, and unmounts when
// pagerite.js announces a swap away from /_a.
import { createApp } from 'vue' import { createApp } from 'vue'
import AnalyticsView from './AnalyticsView.vue' import AnalyticsView from './AnalyticsView.vue'
@@ -8,9 +11,7 @@ let app = null
export function mount(container) { export function mount(container) {
if (app) return if (app) return
app = createApp(AnalyticsView, { app = createApp(AnalyticsView)
initialRange: new URLSearchParams(location.search).get('range') || 'week',
})
app.mount(container) app.mount(container)
} }
@@ -19,6 +20,11 @@ export function unmount() {
app = null app = null
} }
// Auto-mount on a normal (non-fetch) page load. // pagerite.js calls this before swapping away from /_a; each evaluation
// (the inlined production module evaluates fresh on every visit) replaces
// the handle.
window.__pageriteAnalyticsUnmount = unmount
// Auto-mount when the page holding #analytics-app is present.
const container = document.getElementById('analytics-app') const container = document.getElementById('analytics-app')
if (container) mount(container) if (container) mount(container)
+152 -113
View File
@@ -1,15 +1,21 @@
/** /**
* Chart geometry, smoothing, and SVG path generation for analytics charts. * Chart geometry, smoothing, and SVG path generation for analytics charts.
* *
* Fixed 720x180 viewBox, stretched to the panel width; values are per-unit * Fixed 720x180 plot area inside a larger viewBox that also holds the axis
* rates (hour on the week view, day on month+). * labels, so each chart SVG is self-contained; values are per-unit rates
* (hour on the week view, day on month+).
*/ */
import { DAY, HOUR, WEEK, mondayUTC } from './time.js' import { DAY, HOUR, MIN5, WEEK, mondayUTC } from './time.js'
import { formatCount } from './format.js'
export const CHART_W = 720 export const CHART_W = 1000
export const CHART_H = 180 export const CHART_H = 150
export const PAD_TOP = 14 // room above the highest point export const PAD_TOP = 14 // room above the highest point
export const MARGIN_L = 40 // y tick labels + vertical axis label
export const MARGIN_B = 24 // x tick labels
export const VIEW_W = MARGIN_L + CHART_W + 8
export const VIEW_H = CHART_H + MARGIN_B
/** /**
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at * Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
@@ -36,24 +42,19 @@ export function yScale(maxValue) {
} }
/** /**
* Edge-aware adaptive Gaussian smoothing. A change-point detector first * Edge-aware Gaussian smoothing with a fixed bandwidth. A change-point
* finds traffic-level shifts (two-unit totals compared on both sides of * detector first finds traffic-level shifts (two-unit totals compared on
* each bucket; strong ratio + significance marks a candidate, and each run * both sides of each bucket; strong ratio + significance marks a candidate,
* of candidates keeps only its best-scoring bucket as an edge). Each * and each run of candidates keeps only its best-scoring bucket as an
* edge-delimited segment is then smoothed independently: a broad two-unit * edge). Each edge-delimited segment is then smoothed independently: every
* pilot estimates the local traffic rate, which ramps the Gaussian sigma * bucket spreads its count with a fixed Gaussian sigma chosen so N events
* from ~0.4 units (isolated events stay narrow, peaking at ~1 event/unit) * in a single bucket peak at N events per unit, clipped to the segment and
* up to 1 unit (busy traffic gets full smoothing), and every bucket spreads * renormalized so total visitor count is preserved exactly. The unit is
* its count with its local sigma, clipped to the segment and renormalized * one hour on the week view and one day on the month+ views, so the
* so total visitor count is preserved exactly. The unit is one hour on the * smoothing time scale follows the range. The raw series is drawn faintly
* week view and one day on the month+ views, so the smoothing time scale * behind the curve for reference. Operates on raw counts.
* follows the range (month+ sigmas are 24x the hourly ones). The raw series
* is drawn faintly behind the curve for reference. Operates on raw counts.
*/ */
export function smooth(counts, binMinutes, unitMinutes, { export function smooth(counts, binMinutes, unitMinutes, {
minSigmaMinutes = unitMinutes / Math.sqrt(2 * Math.PI),
maxSigmaMinutes = unitMinutes,
pilotSigmaMinutes = 2 * unitMinutes,
detectorWindowMinutes = 2 * unitMinutes, detectorWindowMinutes = 2 * unitMinutes,
// Count thresholds are defined per hour and scale with the unit, so // Count thresholds are defined per hour and scale with the unit, so
// "low traffic" means the same thing on hourly and daily views // "low traffic" means the same thing on hourly and daily views
@@ -61,8 +62,6 @@ export function smooth(counts, binMinutes, unitMinutes, {
highTrafficEvents = 10 * unitMinutes / 60, highTrafficEvents = 10 * unitMinutes / 60,
minRatio = 2.5, minRatio = 2.5,
minSignificance = 4, minSignificance = 4,
sigmaRampStart = 5 * unitMinutes / 60,
sigmaRampEnd = 20 * unitMinutes / 60,
} = {}) { } = {}) {
const n = counts.length const n = counts.length
if (!n) return counts if (!n) return counts
@@ -104,68 +103,25 @@ export function smooth(counts, binMinutes, unitMinutes, {
i = j i = j
} }
const reflectIndex = (i, length) => { // Fixed sigma: N events in one bucket peak at N events per unit.
while (i < 0 || i >= length) { // sigma_bins * sqrt(2*pi) = rate = unitMinutes / binMinutes.
i = i < 0 ? -i - 1 : 2 * length - i - 1 const sigmaBins = unitMinutes / (binMinutes * Math.sqrt(2 * Math.PI))
} const radius = Math.ceil(4 * sigmaBins)
return i
}
const gaussianFilterReflect = (values, sigmaBins) => { // Process each discontinuity-delimited regime independently so the
const length = values.length // Gaussian cannot see through a detected boundary. Each input bin spreads
const radius = Math.ceil(4 * sigmaBins) // its count with the fixed sigma; the kernel is renormalized after
const kernel = new Float64Array(radius * 2 + 1) // clipping to the segment, preserving total visitor count apart from
let sum = 0 // floating-point error.
for (let k = -radius; k <= radius; k++) {
const w = Math.exp(-0.5 * (k / sigmaBins) ** 2)
kernel[k + radius] = w
sum += w
}
for (let i = 0; i < kernel.length; i++) kernel[i] /= sum
const out = new Float64Array(length)
for (let i = 0; i < length; i++) {
let value = 0
for (let k = -radius; k <= radius; k++) {
value += values[reflectIndex(i + k, length)] * kernel[k + radius]
}
out[i] = value
}
return out
}
// Process each discontinuity-delimited regime independently so neither
// the pilot nor the final Gaussian can see through a detected boundary.
const bounds = [0, ...edges, n] const bounds = [0, ...edges, n]
const smoothed = new Float64Array(n) const smoothed = new Float64Array(n)
for (let b = 0; b < bounds.length - 1; b++) { for (let b = 0; b < bounds.length - 1; b++) {
const lo = bounds[b] const lo = bounds[b]
const length = bounds[b + 1] - lo const length = bounds[b + 1] - lo
const segment = counts.slice(lo, lo + length) const segment = counts.slice(lo, lo + length)
// Broad pilot estimates only the generic local traffic level used for
// choosing sigma; it is not the final displayed curve.
const pilot = gaussianFilterReflect(segment, pilotSigmaMinutes / binMinutes)
// Keep isolated/sparse traffic at the minimum bandwidth through
// sigmaRampStart events, then ramp toward maxSigmaMinutes (thresholds
// are per-hour rates scaled to the unit: low traffic is low traffic
// on every range).
const sigmaMinutes = new Float64Array(length)
for (let i = 0; i < length; i++) {
const ratePerUnit = pilot[i] * unitMinutes / binMinutes
let mix = (ratePerUnit - sigmaRampStart) / (sigmaRampEnd - sigmaRampStart)
mix = Math.sqrt(Math.max(0, Math.min(1, mix)))
sigmaMinutes[i] = minSigmaMinutes + mix * (maxSigmaMinutes - minSigmaMinutes)
}
// Each input bin spreads its own count using its local sigma. The
// per-bin kernel is renormalized after clipping to the segment,
// preserving total visitor count apart from floating-point error.
for (let j = 0; j < length; j++) { for (let j = 0; j < length; j++) {
const count = segment[j] const count = segment[j]
if (!count) continue if (!count) continue
const sigmaBins = sigmaMinutes[j] / binMinutes
const radius = Math.ceil(4 * sigmaBins)
const start = Math.max(0, j - radius) const start = Math.max(0, j - radius)
const end = Math.min(length, j + radius + 1) const end = Math.min(length, j + radius + 1)
let weightSum = 0 let weightSum = 0
@@ -200,15 +156,16 @@ export function spline(pts) {
const c1y = clampY(p1.y + (p2.y - p0.y) / 6) const c1y = clampY(p1.y + (p2.y - p0.y) / 6)
const c2y = clampY(p2.y - (p3.y - p1.y) / 6) const c2y = clampY(p2.y - (p3.y - p1.y) / 6)
d += `C${p1.x + (p2.x - p0.x) / 6},${c1y} ` d += `C${p1.x + (p2.x - p0.x) / 6},${c1y} `
+ `${p2.x - (p3.x - p1.x) / 6},${c2y} ${p2.x},${p2.y}` + `${p2.x - (p3.x - p1.x) / 6},${c2y} ${p2.x},${p2.y}`
} }
return d return d
} }
/** Build a full chart model from a series descriptor produced by time.js. */ /** Build a full chart model from a series descriptor produced by time.js. */
export function buildChart(input) { export function buildChart(input, now = Date.now()) {
if (!input || !input.series.length) return null if (!input || !input.series.length) return null
const { series, t0, t1, rate, binMinutes, unitMinutes } = input if (input.unit === '5min') return buildDayChart(input, now)
const { series, t0, t1, rate, binMinutes, unitMinutes, unit } = input
// Values are per-unit rates (hour on the week view, day on month+); the // Values are per-unit rates (hour on the week view, day on month+); the
// y max is derived from the *smoothed* curves so random single-bucket // y max is derived from the *smoothed* curves so random single-bucket
// spikes don't blow up the scale. Smoothing works on raw counts (its edge // spikes don't blow up the scale. Smoothing works on raw counts (its edge
@@ -238,7 +195,7 @@ export function buildChart(input) {
const nMajor = Math.round(max / step) const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) { for (let k = 0; k <= nMajor; k++) {
const v = k * step const v = k * step
majors.push({ value: v, y: y(v), bottom: (1 - PAD_TOP / CHART_H) * (v / max) * 100 }) majors.push({ value: v, y: y(v), label: fmtY(v) })
} }
if (minor) { if (minor) {
for (let v = minor; v < max; v += minor) { for (let v = minor; v < max; v += minor) {
@@ -252,38 +209,120 @@ export function buildChart(input) {
// ranges: boundary lines at Mondays / months / years. // ranges: boundary lines at Mondays / months / years.
const isWeek = t1 - t0 === WEEK const isWeek = t1 - t0 === WEEK
const isMonth = !isWeek && t1 - t0 <= 31 * DAY const isMonth = !isWeek && t1 - t0 <= 31 * DAY
const xticks = isWeek let xticks
? Array.from({ length: 7 }, (_, d) => { if (isWeek) {
const t = t0 + d * DAY + 12 * HOUR xticks = Array.from({ length: 7 }, (_, d) => {
return { const t = t0 + d * DAY + 12 * HOUR
x: x(t), left: ((t - t0) / (t1 - t0)) * 100, return {
label: new Date(t).toLocaleDateString(undefined, { x: x(t),
weekday: 'short', timeZone: 'UTC', label: new Date(t).toLocaleDateString(undefined, {
}), weekday: 'short', timeZone: 'UTC',
line: false, }),
} line: false,
}
})
} else if (isMonth) {
// t0 is day-aligned; label every day whose noon falls inside the range.
xticks = []
for (let day = t0; day + 12 * HOUR < t1; day += DAY) {
const date = new Date(day)
const t = day + 12 * HOUR
xticks.push({
x: x(t),
label: date.getUTCDate() === 1
? date.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' })
: String(date.getUTCDate()),
line: false,
}) })
: isMonth }
? Array.from( } else {
{ length: Math.floor((t1 - Math.ceil(t0 / DAY) * DAY) / DAY) }, xticks = xticksFor(t0, t1).map((t) => ({
(_, d) => { x: x(t), label: fmtTick(t, t1 - t0), line: true,
const day = Math.ceil(t0 / DAY) * DAY + d * DAY }))
const date = new Date(day) }
const t = day + 12 * HOUR return { max, majors, minors, series: drawn, xticks, unit }
return { }
x: x(t), left: ((t - t0) / (t1 - t0)) * 100,
label: date.getUTCDate() === 1 /**
? date.toLocaleDateString(undefined, { month: 'short', timeZone: 'UTC' }) * Day view: 5-minute bars for the last 24 hours. Bars are drawn at raw
: String(date.getUTCDate()), * counts; the skyline uses a projected full-bucket value for the still-open
line: false, * final bucket. The y scale is derived from the projected skyline maximum.
} */
}, export function buildDayChart(input, now = Date.now()) {
) const { series, t0, t1 } = input
: xticksFor(t0, t1).map((t) => ({ const points = series[0]?.points || []
x: x(t), left: ((t - t0) / (t1 - t0)) * 100, const n = points.length
label: fmtTick(t, t1 - t0), line: true, if (!n) return null
})) const bucketMs = (t1 - t0) / n
return { max, majors, minors, series: drawn, xticks } const bucketWidth = CHART_W / n
const gap = 0.2
const barWidth = Math.max(0.2, bucketWidth - gap)
const x = (i) => i * bucketWidth + gap / 2
const prevRaw = n > 1 ? points[n - 2].count : 0
const projected = points.map((p, i) => {
if (i !== n - 1) return p.count
const bucketStart = t0 + i * bucketMs
const elapsed = Math.max(1, Math.min(bucketMs, now - bucketStart))
// Blend the observed partial bucket with the previous full bucket:
// the longer the current bucket has run, the less we borrow from it.
const share = elapsed / bucketMs
return p.count + prevRaw * (1 - share)
})
const highest = Math.max(0, ...projected)
const { max, step, minor } = yScale(highest)
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const bars = points.map((p, i) => {
const bx = x(i)
const by = y(p.count)
return {
x: bx,
y: by,
width: barWidth,
height: CHART_H - by,
raw: p.count,
projected: projected[i],
}
})
let skyline = ''
for (let i = 0; i < bars.length; i++) {
const b = bars[i]
const top = y(b.projected)
if (i === 0) {
skyline += `M${b.x},${top} H${b.x + b.width}`
} else {
skyline += ` V${top} H${b.x + b.width}`
}
}
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), label: fmtY(v) })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
const xticks = []
const tickStep = 3 * HOUR
const firstTick = Math.ceil(t0 / tickStep) * tickStep
for (let t = firstTick; t < t1; t += tickStep) {
if (t < t0) continue
const d = new Date(t)
xticks.push({
x: ((t - t0) / (t1 - t0)) * CHART_W,
label: `${String(d.getUTCHours()).padStart(2, '0')}:00`,
line: false,
})
}
return { bars, skyline: skyline.trim(), max, majors, minors, xticks, unit: '5min', series: [] }
} }
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month /** X ticks for year/all: Monday boundaries up to a quarter, UTC month
@@ -300,7 +339,7 @@ export function xticksFor(t0, t1) {
if (span <= 4 * 365 * DAY) { if (span <= 4 * 365 * DAY) {
const d = new Date(t0) const d = new Date(t0)
let t = Date.UTC(d.getUTCFullYear(), d.getUTCMonth() + 1, 1) let t = Date.UTC(d.getUTCFullYear(), d.getUTCMonth() + 1, 1)
for (; t <= t1; ) { for (; t <= t1;) {
ticks.push(t) ticks.push(t)
const m = new Date(t) const m = new Date(t)
t = Date.UTC(m.getUTCFullYear(), m.getUTCMonth() + 1, 1) t = Date.UTC(m.getUTCFullYear(), m.getUTCMonth() + 1, 1)
@@ -327,7 +366,7 @@ export function fmtTick(t, span) {
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' }) return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
} }
/** Y labels: integers when the step allows, one decimal for fractional steps. */ /** Y labels use the same compact formatter as text labels. */
export function fmtY(v) { export function fmtY(v) {
return Number.isInteger(v) ? String(v) : v.toFixed(1) return formatCount(v)
} }
+313 -53
View File
@@ -25,11 +25,44 @@ export const hostIP = (ip) => {
} }
} }
/** Copy the full IP to the clipboard, ignoring failures. */ function showCopiedFeedback(el) {
export async function copyIp(ip) { if (!el || typeof document === 'undefined') return
const popup = document.createElement('span')
popup.textContent = 'Copied!'
popup.className = 'copy-popup'
popup.style.cssText =
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
'pointer-events:none;z-index:10;'
el.classList.add('has-copy-popup')
el.appendChild(popup)
setTimeout(() => {
popup.remove()
el.classList.remove('has-copy-popup')
}, 1200)
}
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
export async function copyIp(ip, event) {
if (!ip) return if (!ip) return
const el = event?.currentTarget
try { try {
await navigator.clipboard.writeText(ip) await navigator.clipboard.writeText(ip)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
export async function copyList(text, event) {
if (!text) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(text)
showCopiedFeedback(el)
} catch { } catch {
/* ignore */ /* ignore */
} }
@@ -44,6 +77,44 @@ export function calcTotalViews(views) {
return n return n
} }
// Very short reads are navigation/skims, not real reading time.
export const MIN_READ_SECONDS = 10
/** Average minutes per visit and average of per-article median read minutes. */
export function calcReadStats(visits) {
const perArticle = {}
let totalVisitSeconds = 0
let visitCount = 0
for (const v of visits || []) {
const secs = Object.values(v.read || {}).filter((s) => s >= MIN_READ_SECONDS)
if (!secs.length) continue
visitCount++
totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
for (const [path, s] of Object.entries(v.read || {})) {
if (s >= MIN_READ_SECONDS) {
; (perArticle[path] || (perArticle[path] = [])).push(s)
}
}
}
const avgMinPerVisit = visitCount
? Math.max(1, Math.round(totalVisitSeconds / visitCount / 60))
: 0
let articleMedianSum = 0
const articleCount = Object.keys(perArticle).length
for (const arr of Object.values(perArticle)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median = arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
articleMedianSum += Math.max(MIN_READ_SECONDS, median)
}
const avgArticleMedianMin = articleCount
? Math.max(1, Math.round(articleMedianSum / articleCount / 60))
: 0
return { avgMinPerVisit, avgArticleMedianMin }
}
/** Build a path -> page title lookup from the site tree. */ /** Build a path -> page title lookup from the site tree. */
function buildTitleMap(pageTree) { function buildTitleMap(pageTree) {
const titles = new Map() const titles = new Map()
@@ -59,22 +130,22 @@ function buildTitleMap(pageTree) {
/** Last path segment for display; front page becomes a house icon. */ /** Last path segment for display; front page becomes a house icon. */
function slugOf(path) { function slugOf(path) {
return path === '/' ? '🏠' : path.split('/').pop() return path === '/' ? '🏠' : path.split('/').pop()
} }
/** Host name of an external https origin, with scheme stripped. */ /** Host name of an external https origin, with scheme and www. stripped. */
function externalSlug(origin) { function externalSlug(origin) {
try { try {
return new URL(origin).host return new URL(origin).host.replace(/^www\./, '')
} catch { } catch {
return origin.replace(/^https?:\/\//, '') return origin.replace(/^https?:\/\//, '').replace(/^www\./, '')
} }
} }
/** Format one trail step: an internal page or an external https origin. */ /** Format one trail step: an internal page or an external https origin. */
function stepOf(path, titles) { function stepOf(path, titles) {
if (path?.startsWith('/')) { if (path?.startsWith('/')) {
return { path, slug: slugOf(path), title: titles.get(path) || '', external: false } return { path, slug: slugOf(path), title: titles.get(path) || '', external: false, home: path === '/' }
} }
if (path?.startsWith('https://')) { if (path?.startsWith('https://')) {
return { return {
@@ -114,6 +185,8 @@ export function formatWhen(ts, now = Date.now()) {
if (adiff <= 86400000) { if (adiff <= 86400000) {
return formatter return formatter
.format(Math.round(diff / 3600000), 'hour') .format(Math.round(diff / 3600000), 'hour')
.replace('hours', 'h')
.replace('hour', 'h')
.replaceAll(' ', '\u202F') .replaceAll(' ', '\u202F')
} }
if (adiff <= 604800000) { if (adiff <= 604800000) {
@@ -140,6 +213,52 @@ export function formatWhenTooltip(ts) {
return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC') return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC')
} }
/** Full local timestamp for tooltips, e.g. "21 Aug 2026, 17:38:48". */
export function formatWhenLocal(ts) {
return new Date(ts).toLocaleString('en-ie', {
year: 'numeric',
month: 'short',
day: 'numeric',
hour: '2-digit',
minute: '2-digit',
second: '2-digit',
})
}
/** Preserve locale case with the region/country subtag upper-cased. */
export function formatLang(value) {
if (!value || value === '—') return value
const parts = value.split('-')
if (parts.length > 1) {
parts[parts.length - 1] = parts[parts.length - 1].toUpperCase()
}
return parts.join('-')
}
/** ISO 8601 UTC timestamp without subseconds, e.g. "2026-08-21T00:20:48Z". */
export function formatWhenIso(ts) {
return `${new Date(ts).toISOString().split('.')[0]}Z`
}
/**
* Compact read time for tooltips: "50s" under a minute, "1m23s" otherwise.
*/
export function formatReadTime(seconds) {
if (seconds < 60) return `${seconds}s`
return `${Math.floor(seconds / 60)}m${seconds % 60}s`
}
/**
* Compact visitor counts: plain below 1k, then 1.2k / 10k / 1.2M.
* Truncated, not rounded.
*/
export function formatCount(n) {
if (n < 1000) return String(n)
if (n < 10000) return `${Math.trunc(n / 1000)}.${Math.trunc((n % 1000) / 100)}k`
if (n < 1_000_000) return `${Math.trunc(n / 1000)}k`
return `${Math.trunc(n / 1_000_000)}.${Math.trunc((n % 1_000_000) / 100_000)}M`
}
/** /**
* Format recent visits for display, newest first. Each step is a linked slug * Format recent visits for display, newest first. Each step is a linked slug
* pointing to its article; external referers/origins are shown as their * pointing to its article; external referers/origins are shown as their
@@ -196,67 +315,188 @@ export function formatCounts(entries) {
/** /**
* Count distinct User-Agent strings among crawler hits, most common first. * Count distinct User-Agent strings among crawler hits, most common first.
* Returns an array of [ua, count] pairs. * Returns an array of [ua, count] pairs. ``clients`` maps client hashes to
* client records.
*/ */
export function countCrawlerUas(crawlers) { export function countCrawlerUas(crawlers, clients) {
const counts = {} const counts = {}
for (const c of crawlers || []) { for (const c of crawlers || []) {
const value = c.ua_pretty || c.ua || '(no UA)' const client = (clients || {})[c.client] || {}
const value = client.ua_pretty || client.ua || '(no UA)'
counts[value] = (counts[value] || 0) + 1 counts[value] = (counts[value] || 0) + 1
} }
return Object.entries(counts).sort((a, b) => b[1] - a[1]) return Object.entries(counts).sort((a, b) => b[1] - a[1])
} }
/** /**
* Group raw crawler hits by the same (ip, ua) pair we use to tell a real * Reduce a reverse-DNS hostname to its right-most components that fit
* visitor from a crawler, and format each group as a row showing every * within ``limit`` characters. This keeps the meaningful main domain
* internal page that crawler visited. Rows are sorted by total hits, * while avoiding absurdly long subdomains like ``xxx.yyy.zzz...provider.net``.
* most active crawler first, rather than by most recent hit.
*/ */
export function formatCrawlerRows(crawlers, pageTree, now = Date.now()) { export function mainDomain(host, limit = 24) {
if (!host) return host
const labels = host.split('.').filter(Boolean)
if (!labels.length) return host
const parts = [labels.pop()]
while (labels.length) {
const next = labels[labels.length - 1]
const candidate = `${next}.${parts.join('.')}`
if (candidate.length > limit) break
parts.unshift(labels.pop())
}
return parts.join('.')
}
/**
* Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by total hits,
* most active crawler first, rather than by most recent hit.
* ``clients`` maps client hashes to client records.
*/
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
const groups = new Map() const groups = new Map()
for (const c of crawlers || []) { for (const c of crawlers || []) {
const key = `${c.ip}\0${c.ua}` const client = (clients || {})[c.client] || {}
const g = groups.get(key) || { const g = groups.get(c.client) || {
ip: c.ip || '', clientHash: c.client,
ua: c.ua_pretty || c.ua || '—', client,
uaRaw: c.ua || '',
lastStart: 0, lastStart: 0,
pages: new Map(), pages: new Map(),
} }
const start = new Date(c.start).getTime() const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start if (start > g.lastStart) g.lastStart = start
if (c.entry?.startsWith('/')) { if (c.entry?.startsWith('/')) {
g.pages.set(c.entry, (g.pages.get(c.entry) || 0) + 1) const existing = g.pages.get(c.entry) || { count: 0, status: c.status || 200 }
existing.count += 1
if (c.status != null) existing.status = c.status
g.pages.set(c.entry, existing)
} }
groups.set(key, g) groups.set(c.client, g)
} }
const totalHits = (g) => { const totalHits = (g) => {
let n = 0 let n = 0
for (const c of g.pages.values()) n += c for (const p of g.pages.values()) n += p.count
return n return n
} }
return [...groups.values()] return [...groups.values()]
.sort((a, b) => totalHits(b) - totalHits(a) || b.lastStart - a.lastStart) .sort((a, b) => totalHits(b) - totalHits(a) || b.lastStart - a.lastStart)
.slice(0, 10) .slice(0, 10)
.map((g) => ({ .map((g) => {
when: formatWhen(g.lastStart, now), const client = g.client || {}
whenTooltip: formatWhenTooltip(g.lastStart), const host = client.host || ''
pages: [...g.pages.entries()] const isHost = !!host
.sort((a, b) => b[1] - a[1]) return {
.map(([path, count]) => ({ lastSeen: formatWhen(g.lastStart, now),
path, lastSeenIso: formatWhenIso(g.lastStart),
slug: slugOf(path), lastSeenLocal: formatWhenLocal(g.lastStart),
title: titles.get(path) || '', pages: [...g.pages.entries()]
count, .sort((a, b) => b[1].count - a[1].count)
.map(([path, info]) => ({ ...stepOf(path, titles), count: info.count, status: info.status })),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical paths are collapsed into one entry with their hit count.
* Flagged paths (the ones that triggered abuse classification) are lifted to
* the top, followed by other 404s, then document GETs from the abuser. Within
* each category paths are sorted by count descending, then earliest first.
* Rows are sorted by most recent hit first. Visitor metadata comes from the
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
* how many distinct client variations the IP produced. Paths are shown
* verbatim (query string included), not resolved against the page tree.
* ``clients`` maps client hashes to client records.
*/
export function formatAbuseRows(abuse, clients, now = Date.now()) {
const groups = new Map()
for (const a of abuse || []) {
const client = (clients || {})[a.client] || {}
const ip = client.ip || ''
const g = groups.get(ip) || {
ip,
pathCounts: new Map(),
clientHashes: new Set(),
lastStart: 0,
lastClient: a.client,
}
const start = new Date(a.start).getTime()
if (start > g.lastStart) {
g.lastStart = start
g.lastClient = a.client
}
const path = a.path || ''
const existing = g.pathCounts.get(path) || {
path,
count: 0,
firstStart: start,
flag: a.flag || false,
is_404: a.is_404 || false,
}
existing.count += 1
if (start < existing.firstStart) existing.firstStart = start
if (a.flag) existing.flag = true
if (!a.is_404) existing.is_404 = false
g.pathCounts.set(path, existing)
g.clientHashes.add(a.client)
groups.set(ip, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pathCounts.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart)
.slice(0, 10)
.map((g) => {
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2)
const paths = [...g.pathCounts.values()].sort(
(a, b) =>
pathCategory(a) - pathCategory(b) ||
b.count - a.count ||
a.firstStart - b.firstStart,
)
const client = (clients || {})[g.lastClient] || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
paths: paths.map((p) => ({
path: p.path,
count: p.count,
flag: p.flag,
is_404: p.is_404,
})), })),
ip: g.ip, allPaths: paths
ipDisplay: hostIP(g.ip) || g.ip || '—', .map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path))
ua: g.ua, .join('\n'),
uaRaw: g.uaRaw, clientCount: g.clientHashes.size,
total: totalHits(g), ip: client.ip || g.ip,
})) ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
} }
/** /**
@@ -264,31 +504,51 @@ export function formatCrawlerRows(crawlers, pageTree, now = Date.now()) {
* with display strings; missing values become "—". ``trail`` starts with the * with display strings; missing values become "—". ``trail`` starts with the
* external referer (when present), then the entry page and any further internal * external referer (when present), then the entry page and any further internal
* pages or external exit origins. Only the 20 most recent visits are shown. * pages or external exit origins. Only the 20 most recent visits are shown.
* ``clients`` maps client hashes to client records.
*/ */
export function formatVisitRows(visits, pageTree, now = Date.now()) { export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().slice(0, 20).map((v) => { return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const trail = [v.referer, v.entry, ...(v.trail || [])] const client = (clients || {})[v.client] || {}
.map((p) => stepOf(p, titles)) const read = v.read || {}
const statuses = v.statuses || {}
const trail = [v.entry, ...(v.trail || [])]
.map((p) => {
const step = stepOf(p, titles)
if (step) {
if (read[p]) step.readSeconds = read[p]
if (statuses[p]) step.status = statuses[p]
}
return step
})
.filter(Boolean) .filter(Boolean)
const utm = Object.entries(v.utm || {}) const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean)
const utm = utmValues.length ? utmValues.join(' · ') : ''
const utmTitle = Object.entries(v.utm || {})
.map(([k, value]) => `${k}=${value}`) .map(([k, value]) => `${k}=${value}`)
.join(', ') .join(', ')
const dash = (s) => (s || '—') const dash = (s) => (s || '—')
const host = client.host || ''
const isHost = !!host
return { return {
when: formatWhen(v.start, now), lastSeen: formatWhen(v.start, now),
whenTooltip: formatWhenTooltip(v.start), lastSeenIso: formatWhenIso(v.start),
lastSeenLocal: formatWhenLocal(v.start),
langDisplay: formatLang(client.lang),
trail, trail,
refererStep: stepOf(v.referer, titles),
referer: dash(v.referer), referer: dash(v.referer),
ip: v.ip || '', ip: client.ip || '',
ipDisplay: v.host || hostIP(v.ip) || v.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
host: dash(v.host), isHost,
lang: dash(v.lang), lang: dash(client.lang),
country: dash(v.country), country: dash(client.country),
city: dash(v.city), city: dash(client.city),
ua: v.ua_pretty || v.ua || '—', ua: client.ua_pretty || client.ua || '—',
uaRaw: v.ua || '', uaRaw: client.ua || '',
utm: utm || '—', utm: utm || '—',
utmTitle,
} }
}) })
} }
+117 -23
View File
@@ -13,6 +13,7 @@ export const DAY = 86400e3
export const WEEK = 7 * DAY export const WEEK = 7 * DAY
export const RANGES = { export const RANGES = {
day: { label: 'day', span: DAY, bucket: MIN5 },
week: { label: 'week' }, week: { label: 'week' },
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR }, month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
year: { label: 'year', span: 365 * DAY, bucket: DAY }, year: { label: 'year', span: 365 * DAY, bucket: DAY },
@@ -25,6 +26,15 @@ export function mondayUTC(t) {
return (d - ((d + 3) % 7)) * DAY return (d - ((d + 3) % 7)) * DAY
} }
/** ISO 8601 week number of the week containing t (via its Thursday). */
export function isoWeek(t) {
const d = new Date(t)
d.setUTCHours(0, 0, 0, 0)
d.setUTCDate(d.getUTCDate() + 4 - (d.getUTCDay() || 7))
const yearStart = Date.UTC(d.getUTCFullYear(), 0, 1)
return Math.ceil(((d - yearStart) / DAY + 1) / 7)
}
/** Parse sparse timestamp buckets into a { epochMs: count } map. */ /** Parse sparse timestamp buckets into a { epochMs: count } map. */
export function rawTimes(buckets) { export function rawTimes(buckets) {
const raw = {} const raw = {}
@@ -43,7 +53,9 @@ export function sumRange(raw, t0, t1) {
/** /**
* One series per overlaid week: [this week, 1 week ago, ...], at native * One series per overlaid week: [this week, 1 week ago, ...], at native
* 5-minute resolution, up to 8 weeks back (and only weeks that overlap the * 5-minute resolution, up to 8 weeks back (and only weeks that overlap the
* recorded data at all). The current week is truncated at the current bucket * recorded data at all). Each older week's timestamps are shifted forward
* onto the current week's axis so all curves overlay inside the plot.
* The current week is truncated at the current bucket
* — no fake zeroes drawn for the future. Counts are rates per hour * — no fake zeroes drawn for the future. Counts are rates per hour
* (bucket count * 12): a lone visit in a 5-minute bucket reads as "12/h". * (bucket count * 12): a lone visit in a 5-minute bucket reads as "12/h".
* The coarser ranges use per-day rates instead (unitMinutes = 24*60). * The coarser ranges use per-day rates instead (unitMinutes = 24*60).
@@ -66,12 +78,13 @@ export function weeklySeries(buckets) {
: start + WEEK : start + WEEK
const points = [] const points = []
for (let t = start; t < end; t += MIN5) { for (let t = start; t < end; t += MIN5) {
points.push({ t, count: raw[t] || 0 }) points.push({ t: t + back * WEEK, count: raw[t] || 0 })
} }
out.push({ out.push({
points, points,
label: back === 0 ? 'this week' : `${back}w ago`, label: `Week ${isoWeek(start)}`,
opacity: Math.max(0.15, 1 - back * 0.25), opacity: Math.max(0.15, 1 - back * 0.25),
past: back > 0,
area: back === 0, area: back === 0,
}) })
} }
@@ -91,54 +104,135 @@ export function weeklySeries(buckets) {
* to per-day rates (the unit the month+ charts are read in). * to per-day rates (the unit the month+ charts are read in).
* Ranges without a fixed span use the full data reach, but never less than * Ranges without a fixed span use the full data reach, but never less than
* their configured minSpan so the chart keeps a readable minimum x scale. * their configured minSpan so the chart keeps a readable minimum x scale.
* t0 is aligned to the UTC day so the x labels cover the whole range;
* t1 is now, so the scale never extends into the future. The bucket size
* follows the resulting window (6h up to 31 days, daily beyond), so ranges
* covering the same window — "all" at its 30-day minimum vs "month" —
* render the identical curve.
*/ */
export function rollingSeries(buckets, rangeKey) { export function rollingSeries(buckets, rangeKey) {
const raw = rawTimes(buckets) const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number) const times = Object.keys(raw).map(Number)
if (!times.length) return null if (!times.length) return null
const { span, bucket, minSpan = 0 } = RANGES[rangeKey] const { span, bucket, minSpan = 0 } = RANGES[rangeKey]
const t1 = Math.floor(Date.now() / bucket) * bucket + bucket const t1 = Date.now()
const earliest = Math.floor(Math.min(...times) / bucket) * bucket const earliest = Math.min(...times)
const t0 = span != null const t0 = Math.floor((span != null
? t1 - span ? t1 - span
: Math.min(earliest, t1 - minSpan) : Math.min(earliest, t1 - minSpan)) / DAY) * DAY
// The bucket follows the actual window length, not the range key: when
// "all" is capped to its 30-day minimum it covers the very window "month"
// does, and daily bins would draw a different curve over the same data
// (coarser edge detection, points a day apart plotted at bin starts, the
// last point stuck at today's midnight instead of reaching now).
const bucketMs = t1 - t0 <= 31 * DAY ? Math.min(bucket, 6 * HOUR) : bucket
const points = [] const points = []
for (let t = t0; t < t1; t += bucket) { for (let t = t0; t < t1; t += bucketMs) {
points.push({ t, count: sumRange(raw, t, t + bucket) }) points.push({ t, count: sumRange(raw, t, t + bucketMs) })
} }
return { return {
series: [{ points, label: '', opacity: 1, area: true }], series: [{ points, label: '', opacity: 1, area: true }],
t0, t0,
t1, t1,
rate: DAY / bucket, rate: DAY / bucketMs,
binMinutes: bucket / 60e3, binMinutes: bucketMs / 60e3,
unitMinutes: 24 * 60, unitMinutes: 24 * 60,
unit: 'day', unit: 'day',
} }
} }
/** Dispatch to weekly or rolling series based on the selected range. */ /**
* Day view: raw 5-minute bucket counts for the current 24-hour window.
* No smoothing or rate conversion is applied; counts are used as-is.
*/
export function daySeries(buckets) {
const raw = rawTimes(buckets)
const now = Date.now()
const { span, bucket } = RANGES.day
const t1 = Math.floor(now / bucket) * bucket + bucket
const t0 = t1 - span
const points = []
for (let t = t0; t < t1; t += bucket) {
points.push({ t, count: raw[t] || 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: false }],
t0,
t1,
rate: 1,
binMinutes: bucket / 60e3,
unitMinutes: bucket / 60e3,
unit: '5min',
}
}
/** Dispatch to daily, weekly or rolling series based on the selected range. */
export function makeSeries(buckets, rangeKey) { export function makeSeries(buckets, rangeKey) {
return rangeKey === 'week' if (rangeKey === 'day') return daySeries(buckets)
? weeklySeries(buckets) if (rangeKey === 'week') return weeklySeries(buckets)
: rollingSeries(buckets, rangeKey) return rollingSeries(buckets, rangeKey)
} }
/** /**
* Absolute UTC time window for a given range key. Used to filter visits, * Absolute UTC time window for a given range key. Used to filter visits,
* transitions and views to the same period the charts are showing. * transitions and views for the non-chart stats on the analytics page.
* Every bounded range is a rolling span ending at now; the charts instead
* align week to Monday 00:00 UTC (overlaying previous weeks) and month+
* to UTC day boundaries, so their x windows differ from the stats range
* on purpose.
* Returns { t0, t1 } where null means unbounded. * Returns { t0, t1 } where null means unbounded.
*/ */
export function rangeWindow(rangeKey) { export function rangeWindow(rangeKey) {
const now = Date.now() const now = Date.now()
if (rangeKey === 'week') {
const start = mondayUTC(now)
return { t0: start, t1: start + WEEK }
}
if (rangeKey === 'all') { if (rangeKey === 'all') {
return { t0: null, t1: null } return { t0: null, t1: null }
} }
const { span, bucket } = RANGES[rangeKey] const span = rangeKey === 'week' ? WEEK : RANGES[rangeKey].span
const t1 = Math.floor(now / bucket) * bucket + bucket return { t0: now - span, t1: now }
return { t0: t1 - span, t1 } }
/**
* Sum the bucketed transition matrix (from -> to -> bucket ISO -> count)
* into a plain from -> to -> count matrix for the window [t0, t1).
*/
export function filterTransitionsByRange(transitions, t0, t1) {
const out = {}
for (const [fr, tos] of Object.entries(transitions || {})) {
for (const [to, buckets] of Object.entries(tos)) {
let n = 0
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) n += c
}
if (n) {
out[fr] = out[fr] || {}
out[fr][to] = n
}
}
}
return out
}
/** Keep only the 5-minute view buckets that fall inside [t0, t1). */
export function filterViewsByRange(views, t0, t1) {
const filtered = {}
for (const [path, buckets] of Object.entries(views || {})) {
const out = {}
for (const [k, c] of Object.entries(buckets)) {
const t = Date.parse(k)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out[k] = c
}
if (Object.keys(out).length) filtered[path] = out
}
return filtered
}
/** Keep only records whose start time falls inside [t0, t1). */
export function filterRecordsByRange(records, t0, t1) {
const out = []
for (const r of records || []) {
const t = Date.parse(r.start)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out.push(r)
}
return out
} }
File diff suppressed because it is too large Load Diff
+48 -5
View File
@@ -65,6 +65,12 @@
box-sizing: border-box; box-sizing: border-box;
} }
/* Links never underline — including SVG link text, which the UA stylesheet
underlines by default. */
a {
text-decoration: none;
}
html { html {
scroll-behavior: smooth; scroll-behavior: smooth;
/* Native-scrollbar fallback styling (JS off or before pagerite.js runs): /* Native-scrollbar fallback styling (JS off or before pagerite.js runs):
@@ -186,6 +192,10 @@ body {
/* One line always: pagerite.js shrinks the font size to fit instead of /* One line always: pagerite.js shrinks the font size to fit instead of
wrapping (the themed size is the maximum). */ wrapping (the themed size is the maximum). */
white-space: nowrap; white-space: nowrap;
/* Shrink-wrap to the text: as a flex child of the column-direction
#banner it would otherwise stretch full-width, making the empty banner
area beside the text a link to the front page. */
align-self: flex-start;
margin: auto 1.25rem 0; margin: auto 1.25rem 0;
padding-top: 1.5rem; padding-top: 1.5rem;
color: var(--text); color: var(--text);
@@ -791,8 +801,11 @@ figure:has(img[width]) {
The rules below re-anchor the bleed for the layouts where the article The rules below re-anchor the bleed for the layouts where the article
is not viewport-centered; each just overrides width/margin-inline, and is not viewport-centered; each just overrides width/margin-inline, and
later rules win at equal specificity. */ later rules win at equal specificity. The analytics dashboard uses the
figure:has(.wide) { same breakout directly on its container (div.wide — it is the page's
whole content, not a figure). */
figure:has(.wide),
div.wide {
width: 100vw; width: 100vw;
max-width: none; max-width: none;
margin-inline: calc(50% - 50vw); margin-inline: calc(50% - 50vw);
@@ -898,6 +911,30 @@ article h2 {
(explicit img widths still shrink-wrap), while .wide keeps its full (explicit img widths still shrink-wrap), while .wide keeps its full
viewport bleed. */ viewport bleed. */
@media (max-width: 48rem) { @media (max-width: 48rem) {
/* Nav type shrinks fluidly as space runs out. The nav font-size is
em-based both in base and in every theme override, so scaling the
banner's font-size (nothing else in the banner is em-sized — brand and
gaps use rem) reaches the nav through all themes with a single rule.
2.6vw crosses 1rem at ≈38.5rem, so only genuinely narrow viewports
shrink. */
#banner {
font-size: clamp(0.65rem, 2.6vw, 1rem);
}
/* Tighter margins/padding/gaps: the 1.25rem side gutter is wasted space
on a phone. */
#brand {
margin-inline: 0.6rem;
}
#nav {
padding: 0.25rem 0.6rem;
}
#nav ul {
gap: 0.15rem 0.9rem;
}
#content { #content {
display: flex; display: flex;
flex-direction: column; flex-direction: column;
@@ -909,13 +946,19 @@ article h2 {
max-height: none; max-height: none;
overflow-y: visible; overflow-y: visible;
border-radius: 0; border-radius: 0;
padding: 0.5rem 1rem; padding: 0.4rem 0.8rem;
/* Smaller type: the horizontal link strip fits roughly a third more
items per line. The nested-list gaps below are em-based and shrink
along. */
font-size: 0.8rem;
} }
#sidebar ul { /* Only the main level becomes a horizontal wrapping strip; submenus stay
vertical blocks attached under their parent item. */
#sidebar > ul {
flex-direction: row; flex-direction: row;
flex-wrap: wrap; flex-wrap: wrap;
gap: 0.5rem 1.2rem; gap: 0.3rem 0.75rem;
} }
figure:has(.right), figure:has(.right),
+248 -52
View File
@@ -55,6 +55,19 @@ import "overlayscrollbars/overlayscrollbars.css";
let isAdmin = false; let isAdmin = false;
let editorMeta = null; let editorMeta = null;
// Asset URLs for the on-demand bundles. Dev renders them as
// pagerite:* meta tags (Vite dev-server URLs); production inlines all
// page assets and carries the on-demand URLs in a JSON script instead.
const assets = (() => {
const el = document.getElementById("pagerite-assets");
if (el) return JSON.parse(el.textContent);
const map = {};
for (const m of document.querySelectorAll('meta[name^="pagerite:"]')) {
map[m.name] = m.content;
}
return map;
})();
function makePen(mode) { function makePen(mode) {
const btn = document.createElement("button"); const btn = document.createElement("button");
btn.type = "button"; btn.type = "button";
@@ -124,11 +137,11 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
async function setupAuth() { async function setupAuth() {
const src = document.querySelector('meta[name="pagerite:editor-src"]')?.content; const src = assets["pagerite:editor-src"];
if (!src) { pingEntryOnce(); return; } if (!src) { pingEntryOnce(); return; }
editorMeta = { editorMeta = {
src, src,
css: document.querySelector('meta[name="pagerite:editor-css"]')?.content, css: assets["pagerite:editor-css"],
}; };
// Detect whether Paskia SSO is available on this site. // Detect whether Paskia SSO is available on this site.
@@ -147,6 +160,26 @@ import "overlayscrollbars/overlayscrollbars.css";
// No auth proxy / dev. // No auth proxy / dev.
} }
if (isAdmin) {
// Teach the backend the site's public origin (used for absolute
// social/canonical URLs): unlike request headers, location.origin
// reflects the real scheme and host even behind reverse proxies.
fetch("/_api/site-url", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ url: location.origin }),
}).catch(() => {});
// Warm the cache with the editor bundle: the hashed asset is
// immutable, so preloading costs nothing and the pens then open
// instantly. The analytics page has no editor.
if (currentPath !== "/_a" && !import.meta.env.DEV) {
const preload = document.createElement("link");
preload.rel = "modulepreload";
preload.href = src;
document.head.append(preload);
}
}
renderAuthUi(); renderAuthUi();
pingEntryOnce(); pingEntryOnce();
} }
@@ -230,6 +263,7 @@ import "overlayscrollbars/overlayscrollbars.css";
// buttons; re-add whichever auth UI is appropriate for this session. // buttons; re-add whichever auth UI is appropriate for this session.
renderAuthUi(); renderAuthUi();
placeEditPen(); placeEditPen();
fitNav();
// Multi-column layout only when there is enough text to justify it. // Multi-column layout only when there is enough text to justify it.
// Split the body into columned segments: h1s, h2s and wide figures are // Split the body into columned segments: h1s, h2s and wide figures are
// full-width separators and never go inside columns. // full-width separators and never go inside columns.
@@ -285,14 +319,17 @@ import "overlayscrollbars/overlayscrollbars.css";
// internal link is fetched exactly once, and navigation is served from // internal link is fetched exactly once, and navigation is served from
// memory with no fetch at all. Editor re-renders (swapdoc.loadPlain) // memory with no fetch at all. Editor re-renders (swapdoc.loadPlain)
// announce their fresh copies via pagerite:page-fetched, keeping the // announce their fresh copies via pagerite:page-fetched, keeping the
// cache in sync after edits. // cache in sync after edits. The current page is NOT preloaded: we just
// received it as the document (re-fetching would be redundant, and
// browser heuristics may send it without if-none-match, defeating the
// conditional request); it enters the cache when navigated to.
const pageCache = new Map(); // pathname -> HTML text const pageCache = new Map(); // pathname -> HTML text
addEventListener("pagerite:page-fetched", (ev) => { addEventListener("pagerite:page-fetched", (ev) => {
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html); pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html);
}); });
function preload() { function preload() {
const urls = new Set([location.pathname]); const urls = new Set();
for (const a of document.querySelectorAll( for (const a of document.querySelectorAll(
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]', '#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
)) { )) {
@@ -300,7 +337,10 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
for (const url of urls) { for (const url of urls) {
if (pageCache.has(url)) continue; if (pageCache.has(url)) continue;
fetch(url) // x-pagerite-preload: idle cache warm-up, not a page view — the
// server excludes these GETs from analytics (the ping sent on actual
// navigation does the counting).
fetch(url, { headers: { "x-pagerite-preload": "1" } })
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html") .then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
? r.text() : "")) ? r.text() : ""))
.then((html) => { if (html) pageCache.set(url, html); }) .then((html) => { if (html) pageCache.set(url, html); })
@@ -338,29 +378,109 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
// --- Analytics pings --------------------------------------------------- // --- Analytics pings ---------------------------------------------------
// Fire-and-forget POST /_a {fr, to}: on the initial page load (starts the // Fire-and-forget POST /_a {fr, to, read}: on the initial page load
// visit — the server counts nothing from the document GET alone), for // (starts the visit — the server counts nothing from the document GET
// internal fetch-navigations and for external https exits. Excluded: // alone), for internal fetch-navigations, for external https exits, and
// back/forward (popstate never pings) and everything while we know the // on window close. ``read`` is the active time (ms) spent on ``fr``.
// user is an admin — but only when SSO is actually in use; with no auth // Reading time pauses after 1 minute of inactivity and resumes on the
// (dev/test) "admin" is everyone's state and nothing would be recorded — // next mouse/touch/scroll/keyboard event.
// or has the editor open (admin noise, not visits). The analytics page // Excluded: back/forward (popstate never pings), everything while the
// itself (/_a) is also excluded even though fetch-navigation treats it like // editor is open (body.editing — admin noise, not visits), and
// a normal article. // navigations TO the analytics page (/_a — admin machinery, and the
// server rejects it as a ping target anyway). Navigations AWAY from /_a
// must ping: load() already fetched the target page without the preload
// header, and without the ping that GET would flush to the crawler list.
// Admins (when SSO is actually in use — with no auth proxy "admin" is
// everyone's state) ping normally but with hide=1: the server then
// records nothing and scrubs any session the same browser accumulated
// before logging in, so admins never show up as visits or crawlers.
// See docs/analytics.md. // See docs/analytics.md.
function ping(to, fr = currentPath) { function ping(to, fr = currentPath, read = 0) {
if ((ssoAvailable && isAdmin) || document.body.classList.contains("editing") if (document.body.classList.contains("editing")) return;
|| to === "/_a" || fr === "/_a") return; if (to && to === "/_a") return;
const hide = ssoAvailable && isAdmin ? 1 : 0;
const body = JSON.stringify({
fr, to, hide,
read: Math.max(0, Math.round(read / 1000)),
});
try { try {
fetch("/_a", { fetch("/_a", {
method: "POST", method: "POST",
keepalive: true, keepalive: true,
headers: { "content-type": "application/json" }, headers: { "content-type": "application/json" },
body: JSON.stringify({ fr, to }), body,
}); });
} catch { /* analytics must never break navigation */ } } catch { /* analytics must never break navigation */ }
} }
// Active reading time for the current page. The clock stops after 1 minute
// without activity and restarts on the next mouse/touch/scroll/keyboard
// event.
const INACTIVE_MS = 60_000;
let readStart = performance.now();
let readElapsed = 0;
let reading = true;
let readInactivityTimer = null;
let closePingedFor = null;
function markReadActivity() {
if (!reading) {
reading = true;
readStart = performance.now();
}
clearTimeout(readInactivityTimer);
readInactivityTimer = setTimeout(() => {
if (reading) {
readElapsed += performance.now() - readStart;
reading = false;
}
}, INACTIVE_MS);
}
function takeReadTime() {
if (reading) {
readElapsed += performance.now() - readStart;
readStart = performance.now();
}
const ms = Math.max(0, Math.round(readElapsed));
readElapsed = 0;
return ms;
}
function resetReadTime() {
readElapsed = 0;
reading = true;
readStart = performance.now();
clearTimeout(readInactivityTimer);
}
function sendClosePing() {
if (closePingedFor === currentPath) return;
const read = Math.max(0, Math.round(takeReadTime() / 1000));
if (read <= 0) return;
const hide = ssoAvailable && isAdmin ? 1 : 0;
const body = JSON.stringify({ fr: currentPath, hide, read });
const blob = new Blob([body], { type: "application/json" });
try {
if (navigator.sendBeacon) {
navigator.sendBeacon("/_a", blob);
} else {
fetch("/_a", {
method: "POST",
keepalive: true,
headers: { "content-type": "application/json" },
body,
});
}
} catch { /* analytics must never break navigation */ }
closePingedFor = currentPath;
}
for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) {
addEventListener(ev, markReadActivity, { passive: true });
}
addEventListener("pagehide", sendClosePing);
// The initial page load pings too — it is what starts the visit and // The initial page load pings too — it is what starts the visit and
// counts the entry page view (the document GET alone records nothing). // counts the entry page view (the document GET alone records nothing).
// Sent once per load, after the auth probes so the admin gate applies; // Sent once per load, after the auth probes so the admin gate applies;
@@ -377,29 +497,39 @@ import "overlayscrollbars/overlayscrollbars.css";
// --- Analytics page mount/unmount -------------------------------------- // --- Analytics page mount/unmount --------------------------------------
// The analytics page is a normal page whose body is rendered by the server // The analytics page is a normal page whose body is rendered by the server
// but whose content is a Vue app. We load the entry module on demand so the // but whose content is a Vue app. In dev the entry module is imported from
// analytics bundle is only fetched when visiting /_a, and unmount the app // the Vite dev server on demand; in production it is inlined into the /_a
// before swapping away so Vue teardown runs cleanly. // page as script#pagerite-js-analytics, which a fetch-navigation swap does
let analyticsUnmount = null; // not execute — re-create the element so the fresh module auto-mounts on
// #analytics-app (see analytics-main.js). The module exposes its unmount
// as window.__pageriteAnalyticsUnmount.
function teardownAnalytics() { function teardownAnalytics() {
analyticsUnmount?.(); // Remove even the server-rendered script element so a later return to
analyticsUnmount = null; // /_a re-mounts from a fresh copy (the module has torn itself down).
document.getElementById("pagerite-js-analytics")?.remove();
window.__pageriteAnalyticsUnmount?.();
window.__pageriteAnalyticsUnmount = null;
} }
async function mountAnalytics(doc) { async function mountAnalytics(doc) {
const src = doc.querySelector('meta[name="pagerite:analytics-src"]')?.content; if (!doc.getElementById("analytics-app")) return;
if (!src) { // Already mounted: on a full /_a load the inline script has run.
teardownAnalytics(); if (document.getElementById("pagerite-js-analytics")) return;
const inline = doc.getElementById("pagerite-js-analytics");
if (inline) {
const s = document.createElement("script");
for (const a of inline.attributes) s.setAttribute(a.name, a.value);
s.textContent = inline.textContent;
document.body.append(s);
return; return;
} }
try { try {
const mod = await import(/* @vite-ignore */ src); // Dev: the cached module auto-mounts only on its first evaluation,
// so call mount() explicitly for repeat visits (it no-ops when the
// app is already up).
const mod = await import(/* @vite-ignore */ assets["pagerite:analytics-src"]);
const container = document.getElementById("analytics-app"); const container = document.getElementById("analytics-app");
if (container) { if (container) mod.mount(container);
mod.mount(container);
analyticsUnmount = mod.unmount;
}
} catch (e) { } catch (e) {
console.error("analytics mount failed:", e); console.error("analytics mount failed:", e);
} }
@@ -426,8 +556,8 @@ import "overlayscrollbars/overlayscrollbars.css";
// Reflect any redirect the server issued. // Reflect any redirect the server issued.
if (res.redirected) finalUrl = res.url; if (res.redirected) finalUrl = res.url;
const html = await res.text(); const html = await res.text();
// Populate the cache too, or the post-swap preload (which includes // Populate the cache too, so returning here (back/forward, or a
// location.pathname) would fetch the very page we just loaded again. // self-link in the nav) is served from memory.
pageCache.set(new URL(finalUrl, location.href).pathname, html); pageCache.set(new URL(finalUrl, location.href).pathname, html);
doc = new DOMParser().parseFromString(html, "text/html"); doc = new DOMParser().parseFromString(html, "text/html");
} catch { } catch {
@@ -456,27 +586,47 @@ import "overlayscrollbars/overlayscrollbars.css";
} else if (oldSidebar) { } else if (oldSidebar) {
oldSidebar.remove(); oldSidebar.remove();
} }
// Site-wide custom CSS lives in <head id="pagerite-user"> and must be // Stylesheets live in <head> with stable ids — links in dev, inline
// kept in sync across fetch-navigations. It is kept last in <head>: // <style> elements in production — and must follow the swap: the
// in dev Vite injects the base stylesheet after the server-rendered // analytics sheet exists on /_a only, and theme/banner/custom CSS
// tag, and equal-specificity :root rules are decided by order. // may have changed since this page was loaded. Diff by id, keeping
const oldUserStyle = document.getElementById("pagerite-user"); // the fresh document's order; unchanged sheets keep their elements
const newUserStyle = doc.getElementById("pagerite-user"); // so their @keyframes are never torn down. Editor-injected sheets
if (oldUserStyle && newUserStyle) { // (data-pagerite, no id) and Vite's dev styles (no id) are left
oldUserStyle.textContent = newUserStyle.textContent; // alone. Mirrors the head sync in swapdoc.js.
document.head.appendChild(oldUserStyle); const sel = 'link[rel="stylesheet"][id], style[id]';
} else if (newUserStyle) { const fresh = [...doc.head.querySelectorAll(sel)];
document.head.appendChild(document.importNode(newUserStyle, true)); const freshIds = new Set(fresh.map((el) => el.id));
} else if (oldUserStyle) { for (const el of [...document.head.querySelectorAll(sel)]) {
oldUserStyle.remove(); if (!freshIds.has(el.id)) el.remove();
} }
let anchor = null;
for (const el of fresh) {
const cur = document.getElementById(el.id);
if (cur && cur.outerHTML === el.outerHTML) {
anchor = cur;
continue;
}
const imported = document.importNode(el, true);
if (cur) cur.replaceWith(imported);
else if (anchor) anchor.after(imported);
else {
const base = document.getElementById("pagerite-base");
if (base) base.after(imported);
else document.head.append(imported);
}
anchor = imported;
}
// Custom CSS must stay last: equal-specificity :root rules (font
// variables) are decided by order, and in dev Vite injects the base
// stylesheet after the server-rendered tag.
const userStyle = document.getElementById("pagerite-user");
if (userStyle) document.head.appendChild(userStyle);
document.title = doc.title; document.title = doc.title;
// Banners may contain scripts (canvas etc.), content pages may too. // Banners may contain scripts (canvas etc.), content pages may too.
runScripts(document.getElementById("page-banner")); runScripts(document.getElementById("page-banner"));
runScripts(document.getElementById("main")); runScripts(document.getElementById("main"));
applyEffects(); applyEffects();
// The fetched doc carries the analytics meta; the live document's
// <head> is never swapped, so querying it would never find the entry.
mountAnalytics(doc); mountAnalytics(doc);
}; };
// Rotating cube page transition (see the FRAGILE block in pagerite.css); // Rotating cube page transition (see the FRAGILE block in pagerite.css);
@@ -538,7 +688,10 @@ import "overlayscrollbars/overlayscrollbars.css";
if (url.origin !== location.origin) { if (url.origin !== location.origin) {
// External link: the browser navigates; record the full https URL so // External link: the browser navigates; record the full https URL so
// different links to the same domain stay distinct in analytics. // different links to the same domain stay distinct in analytics.
if (url.protocol === "https:") ping(url.href); if (url.protocol === "https:") {
closePingedFor = currentPath;
ping(url.href, currentPath, takeReadTime());
}
return; return;
} }
// Same-page anchor links (footnotes etc.): let the browser handle them // Same-page anchor links (footnotes etc.): let the browser handle them
@@ -550,7 +703,12 @@ import "overlayscrollbars/overlayscrollbars.css";
ev.preventDefault(); ev.preventDefault();
// Capture the source now: load() updates currentPath before pinging. // Capture the source now: load() updates currentPath before pinging.
const from = currentPath; const from = currentPath;
load(url).then((ok) => { if (ok) ping(url.pathname, from); }); load(url).then((ok) => {
if (!ok) return;
closePingedFor = null;
ping(url.pathname, from, takeReadTime());
resetReadTime();
});
}); });
addEventListener("popstate", () => { addEventListener("popstate", () => {
@@ -621,6 +779,44 @@ import "overlayscrollbars/overlayscrollbars.css";
fit(); fit();
} }
// --- Nav condense-to-fit -------------------------------------------------
// The top nav stays on one row even on too-narrow screens: first the link
// gaps shrink, then the nav's side padding, and only in extreme cases the
// font size. #nav is replaced on fetch-navigation swaps, so this re-runs
// from applyEffects (fresh elements each time); CSS keeps flex-wrap: wrap
// as the no-JS fallback.
function fitNav() {
const nav = document.getElementById("nav");
const ul = nav?.querySelector("ul");
if (!ul) return;
// Restore the themed defaults before measuring.
nav.style.fontSize = "";
nav.style.paddingInline = "";
ul.style.columnGap = "";
ul.style.flexWrap = "nowrap";
const overflow = () => ul.scrollWidth - ul.clientWidth;
if (overflow() <= 0) return;
// 1) shrink the gaps between items (down to a fifth of the themed gap)
const gap = parseFloat(getComputedStyle(ul).columnGap) || 0;
const joints = Math.max(ul.children.length - 1, 1);
if (gap > 0) {
ul.style.columnGap = `${Math.max(0.2 * gap, gap - overflow() / joints)}px`;
}
// 2) shrink the nav's side padding (down to 0.4x)
if (overflow() > 0) {
const pad = parseFloat(getComputedStyle(nav).paddingInlineStart) || 0;
nav.style.paddingInline = `${Math.max(0.4 * pad, pad - overflow() / 2)}px`;
}
// 3) shrink the font to fit what remains
if (overflow() > 0) {
const fs = parseFloat(getComputedStyle(nav).fontSize);
nav.style.fontSize = `${fs * ul.clientWidth / ul.scrollWidth}px`;
}
}
addEventListener("resize", fitNav);
document.fonts?.ready.then(fitNav);
setupAuth(); setupAuth();
applyEffects(); applyEffects();
mountAnalytics(document); mountAnalytics(document);
+26 -19
View File
@@ -59,32 +59,39 @@ function swapRegions(doc) {
curUserStyle.remove() curUserStyle.remove()
} }
// Theme and other public stylesheets live in <head>, rendered with stable // Theme and other public stylesheets live in <head>, rendered with stable
// ids by the backend; sync them positionally so the custom CSS (rendered // ids by the backend (links in dev, inline <style> elements in prod);
// last) always keeps winning by order. Diff-based: unchanged sheets keep // sync them positionally so the custom CSS (rendered last) always keeps
// their elements, so their @keyframes are never torn down (re-creating // winning by order. Diff-based: unchanged sheets keep their elements, so
// keyframes would replay the editor's slide-in animation). // their @keyframes are never torn down (re-creating keyframes would
const freshLinks = [...doc.head.querySelectorAll('link[rel="stylesheet"]')] // replay the editor's slide-in animation).
const freshIds = new Set(freshLinks.map((l) => l.id)) const sel = 'link[rel="stylesheet"][id], style[id]'
for (const link of [...document.head.querySelectorAll('link[rel="stylesheet"]')]) { const freshEls = [...doc.head.querySelectorAll(sel)]
if (!link.dataset.pagerite && !freshIds.has(link.id)) link.remove() const freshIds = new Set(freshEls.map((el) => el.id))
for (const el of [...document.head.querySelectorAll(sel)]) {
if (!freshIds.has(el.id)) el.remove()
} }
// Insert missing sheets in the fresh document's order, each right after // Insert missing sheets in the fresh document's order, each right after
// its predecessor's element. The first sheet rendered is always the base // its predecessor's element. The first sheet rendered is always the base
// CSS, so its link doubles as the fallback anchor when nothing matched yet // CSS, so its element doubles as the fallback anchor when nothing matched
// (e.g. no theme was selected before and the position is otherwise lost). // yet (e.g. no theme was selected before and the position is otherwise
// lost).
let anchor = null let anchor = null
for (const link of freshLinks) { for (const el of freshEls) {
const cur = link.id && document.getElementById(link.id) const cur = el.id && document.getElementById(el.id)
if (cur && cur.href === link.href) { if (cur && cur.outerHTML === el.outerHTML) {
anchor = cur anchor = cur
continue continue
} }
const el = document.importNode(link, true) const imported = document.importNode(el, true)
// Same id, new URL (theme switch): replace in place, keeping position. // Same id, new content (theme switch): replace in place, keeping position.
if (cur) cur.replaceWith(el) if (cur) cur.replaceWith(imported)
else if (anchor) anchor.after(el) else if (anchor) anchor.after(imported)
else document.getElementById('pagerite-base')?.after(el) ?? document.head.append(el) else {
anchor = el const base = document.getElementById('pagerite-base')
if (base) base.after(imported)
else document.head.append(imported)
}
anchor = imported
} }
// The editor keeps its own title while open; only inherit the server title // The editor keeps its own title while open; only inherit the server title
// when navigating outside the editor (e.g. fetch-navigation swaps). // when navigating outside the editor (e.g. fetch-navigation swaps).
+4 -5
View File
@@ -7,11 +7,10 @@ import vueDevTools from 'vite-plugin-vue-devtools'
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200' const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200'
// Proxy content pages (/slug, /path/to/slug) to the FastAPI backend in dev. // Proxy everything except Vite's own dev-time paths and the backend machinery
// Excludes Vite internals (/@..., /src, /node_modules, /__...) and the // to the FastAPI backend in dev. /_api, /_f, /_themes and /_a are handled by
// backend's /_ prefix. /_api, /_f, /_themes and the /_a analytics ping are // the fastapi-vue plugin, and /@..., /src, /node_modules, /__... stay with Vite.
// handled by the fastapi-vue plugin. const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
const CONTENT_PROXY = '^\\/(?!_|@|src|node_modules|__)(?:[^./?]+(?:\\/[^./?]+)*)?(?:\\?.*)?$'
// https://vite.dev/config/ // https://vite.dev/config/
export default defineConfig({ export default defineConfig({
+481 -108
View File
@@ -5,15 +5,27 @@ ping on page load starts a visit, later pings extend it, and pings with no
known session start a fresh one (missing data, not dropped). The document known session start a fresh one (missing data, not dropped). The document
GET handler stashes the entry referer (external https origin) and any GET handler stashes the entry referer (external https origin) and any
utm_* query parameters in in-memory IP tables, consumed when the ping utm_* query parameters in in-memory IP tables, consumed when the ping
starts the visit; nothing is counted without a ping (bots and admin starts the visit; nothing is counted without a ping (plain bots that only
browsing stay invisible). The session map is in-memory only. The visitor fetch documents end up in the crawler list). JS-running crawlers
IP and, when available, its reverse-DNS host name are stored on the visit (Googlebot, GoogleOther, Applebot, ...) do ping, but their UA gives them
record itself. away (``_is_bot_ua``) and their pings are ignored, so they land in the
crawler list too. Idle-time link preloads from pagerite.js carry an
``x-pagerite-preload`` header and are not tracked at all — the ping sent
when the user actually navigates does the counting.
Admin clients ping with ``hide=1``, which records nothing and removes any
visit the session accumulated before logging in. Scanner telltale 404s
(dotpaths, *.php) classify the source IP as abuse; its hits — including
earlier crawler hits — are moved to the abuse list, which the viewer
groups by IP with full request paths. Client metadata (IP, UA, language,
country/city, host) is stored once per unique client hash and referenced
from visits, crawler hits and abuse hits. The session map is in-memory
only.
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db), Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
rewritten atomically on every recorded event. rewritten atomically on every recorded event.
""" """
import ipaddress
import os import os
import re import re
import tempfile import tempfile
@@ -23,6 +35,7 @@ from datetime import UTC, datetime, timedelta
from pathlib import Path from pathlib import Path
from urllib.parse import parse_qs, urlparse from urllib.parse import parse_qs, urlparse
import blake3
import msgspec import msgspec
from ua_parser import parse from ua_parser import parse
@@ -41,7 +54,10 @@ def _compact_user_agent(ua: str) -> str:
dev = r.device.family if r.device else None dev = r.device.family if r.device else None
if browser in (None, "Other") and os_name in (None, "Other"): if browser in (None, "Other") and os_name in (None, "Other"):
return ua return ua
browser = browser if browser and browser != "Other" else "" if browser and browser != "Other":
browser = browser.split()[0]
else:
browser = ""
os_name = os_name if os_name and os_name != "Other" else "" os_name = os_name if os_name and os_name != "Other" else ""
if dev in (None, "Other") or dev == browser: if dev in (None, "Other") or dev == browser:
dev = "" dev = ""
@@ -49,51 +65,92 @@ def _compact_user_agent(ua: str) -> str:
return " ".join(p for p in parts if p).strip() return " ".join(p for p in parts if p).strip()
class Client(msgspec.Struct, omit_defaults=True):
"""Client metadata shared by visits, crawler hits and abuse hits.
Identified by a 6-byte blake3 hash of the IPv4 address or IPv6 /64
network, the full User-Agent string and the extracted language tag.
Country/city/host are filled in asynchronously after the first event.
"""
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
ip: str = ""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter country code from the DB-IP geoip lookup, or "".
country: str = ""
#: City name from the DB-IP geoip lookup, or "".
city: str = ""
#: Raw User-Agent header.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
class Visit(msgspec.Struct, omit_defaults=True): class Visit(msgspec.Struct, omit_defaults=True):
"""One visit: the initial-load data plus everything seen afterwards. """One visit: the initial-load data plus everything seen afterwards.
``trail`` holds page paths and external exit URLs in first-seen ``trail`` holds page paths and external exit URLs in first-seen
order; re-visiting an already seen page does not append. The entry order; re-visiting an already seen page does not append. The entry
page itself is in ``entry``, not in the trail. page itself is in ``entry``, not in the trail. Client metadata is
held in ``Analytics.clients`` keyed by ``client``.
""" """
start: datetime start: datetime
entry: str entry: str
#: External https origin of the initial load, "" for direct visits. #: External https origin of the initial load, "" for direct visits.
referer: str = "" referer: str = ""
#: Visitor IP address (first X-Forwarded-For hop or direct peer). #: 6-byte blake3 hash referencing ``Analytics.clients``.
ip: str = "" client: bytes = b""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
trail: list[str] = [] trail: list[str] = []
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter region subtag derived from ``lang`` (e.g. "US"), or "".
#: Overwritten by the DB-IP geoip lookup when a database is available.
country: str = ""
#: City name from the DB-IP geoip lookup, or "".
city: str = ""
#: Raw User-Agent header from the initial ping.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
#: UTM query parameters from the landing URL, keyed by parameter name. #: UTM query parameters from the landing URL, keyed by parameter name.
utm: dict[str, str] = {} utm: dict[str, str] = {}
#: Active reading time per path (seconds), keyed by path.
read: dict[str, int] = {}
#: HTTP status of the response when the path was first seen (200 or 404).
statuses: dict[str, int] = {}
class CrawlerHit(msgspec.Struct, omit_defaults=True): class CrawlerHit(msgspec.Struct, omit_defaults=True):
"""A document GET that was never followed by an analytics ping.""" """A document GET that was never followed by an analytics ping.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime start: datetime
entry: str entry: str
ip: str = "" #: 6-byte blake3 hash referencing ``Analytics.clients``.
ua: str = "" client: bytes = b""
#: Compact display form of ``ua`` when parsable.
ua_pretty: str = ""
#: External https origin of the initial load, "" for direct/none. #: External https origin of the initial load, "" for direct/none.
referer: str = "" referer: str = ""
#: Raw query string of the landing URL (UTM tags can be parsed from it). #: Raw query string of the landing URL (UTM tags can be parsed from it).
query: str = "" query: str = ""
#: HTTP status of the served response (200 or 404 for content pages).
status: int = 200
class AbuseHit(msgspec.Struct, omit_defaults=True):
"""A request from an IP classified as a scanner/abuser.
Unlike crawler hits the full request path (query string included) is
kept: the interesting part is exactly which paths were probed.
``flag`` marks the path that triggered classification; ``is_404``
distinguishes 404 responses from document GETs made by the abuser.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime
#: Full request path including the query string (e.g. "/.env?x=1").
path: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: True when this path triggered abuse classification (telltale path
#: or the 404 that crossed the threshold).
flag: bool = False
#: True for 404 responses; false for document GETs from the abuser.
is_404: bool = False
class Analytics(msgspec.Struct, omit_defaults=True): class Analytics(msgspec.Struct, omit_defaults=True):
@@ -103,6 +160,12 @@ class Analytics(msgspec.Struct, omit_defaults=True):
visits: list[Visit] = [] visits: list[Visit] = []
#: Document GETs that never produced a ping, treated as crawler/bot hits. #: Document GETs that never produced a ping, treated as crawler/bot hits.
crawlers: list[CrawlerHit] = [] crawlers: list[CrawlerHit] = []
#: Requests from abusive IPs (see AbuseHit), grouped by IP in the viewer.
abuse: list[AbuseHit] = []
#: Client metadata keyed by 6-byte blake3 hash.
clients: dict[bytes, Client] = {}
#: IPs classified as scanners/abusers (keys; values always True).
abuse_ips: dict[str, bool] = {}
#: Page transitions per 5-minute bucket (sparse): #: Page transitions per 5-minute bucket (sparse):
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or #: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
#: "(direct)" for initial loads, a page path for pings. #: "(direct)" for initial loads, a page path for pings.
@@ -186,9 +249,62 @@ def _utm_tags(query: str) -> dict[str, str]:
_CRAWLER_TIMEOUT = timedelta(seconds=10) _CRAWLER_TIMEOUT = timedelta(seconds=10)
#: UAs of JS-running crawlers, which would register as visitors on their
#: ping. Anything calling itself a "bot" or "spider" matches; known crawlers
#: without those tokens (GoogleOther) are listed as extra alternates. No
#: source verification: a spoofed bot UA just lands in the crawler list, and
#: scanners that probe telltale paths are caught by the abuse rules anyway.
_BOT_UA = re.compile(r"bot|spider|googleother", re.IGNORECASE)
def _is_bot_ua(ua: str) -> bool:
"""True when the UA claims a crawler identity (bot or spider)."""
return bool(_BOT_UA.search(ua))
#: Plain-404 count per IP that classifies it as abuse even without a
#: telltale path hit.
_ABUSE_404_THRESHOLD = 10
#: Paths that instantly classify an IP as abuse when they 404: any segment
#: starting with a dot ("/.env", "/.git/config") or ending in ".php".
_ABUSE_PATH = re.compile(r"(^|/)\.|\.php$", re.IGNORECASE)
def _is_abuse_path(path: str) -> bool:
"""Telltale scanner path: dot segment or *.php."""
return bool(_ABUSE_PATH.search(path.split("?")[0]))
def _network_ip(ip: str) -> str:
"""IPv4 address unchanged, IPv6 collapsed to its /64 network address.
We hash the network rather than the full address so that clients in the
same /64 (a typical end-user allocation) are treated as one visitor.
"""
if not ip:
return ip
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ip
if isinstance(addr, ipaddress.IPv6Address):
return str(ipaddress.IPv6Network(f"{ip}/64", strict=False).network_address)
return ip
def _client_hash(ip: str, ua: str, lang: str) -> bytes:
"""6-byte blake3 digest identifying a visitor/client tuple.
The key is the prettified IP (IPv6 /64), the raw UA string and the
extracted language tag, separated by null bytes.
"""
return blake3.blake3(
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
).digest()[:6]
class Store: class Store:
"""In-memory analytics data plus the (IP, UA) -> visit session map.""" """In-memory analytics data plus the client-hash -> visit session map."""
def __init__(self, path: Path) -> None: def __init__(self, path: Path) -> None:
self.path = path self.path = path
@@ -198,8 +314,12 @@ class Store:
self.data = msgspec.json.decode(path.read_bytes(), type=Analytics) self.data = msgspec.json.decode(path.read_bytes(), type=Analytics)
except msgspec.DecodeError, OSError: except msgspec.DecodeError, OSError:
pass # legacy schema / corrupt or unreadable file: start fresh pass # legacy schema / corrupt or unreadable file: start fresh
#: (ip, user-agent) -> index of the current visit in data.visits #: client hash -> index of the current visit in data.visits
self.sessions: dict[tuple[str, str], int] = {} self.sessions: dict[bytes, int] = {}
#: visit index -> count events recorded for that visit, so
#: ``_remove_visit`` can reverse all of them — not just the ones
#: from the visit's creation. In-memory only, like ``sessions``.
self._count_log: dict[int, list[tuple]] = {}
#: ip -> external https origin of the latest document GET carrying #: ip -> external https origin of the latest document GET carrying
#: one, stashed for the visit the client's initial ping starts. #: one, stashed for the visit the client's initial ping starts.
#: Internal or absent referers never touch the table. #: Internal or absent referers never touch the table.
@@ -212,6 +332,12 @@ class Store:
#: Document GETs that have not yet been matched by a ping. Kept #: Document GETs that have not yet been matched by a ping. Kept
#: in RAM only; expired entries are written to ``data.crawlers``. #: in RAM only; expired entries are written to ``data.crawlers``.
self.pending_crawlers: list[CrawlerHit] = [] self.pending_crawlers: list[CrawlerHit] = []
#: client hash -> {path: status} for recent document GETs, consumed
#: by the matching ping to record the status of each visited path.
self.pending_statuses: dict[bytes, dict[str, int]] = {}
#: ip -> number of plain (non-telltale) 404s seen, in RAM only;
#: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse.
self.not_found_counts: dict[str, int] = {}
#: Callables to notify when persisted data changes. Registered by the #: Callables to notify when persisted data changes. Registered by the
#: analytics WebSocket broadcaster. #: analytics WebSocket broadcaster.
self._on_change: list[Callable[[], None]] = [] self._on_change: list[Callable[[], None]] = []
@@ -244,20 +370,26 @@ class Store:
else: else:
self._notify() self._notify()
def _flush_crawlers(self, now: datetime | None = None) -> None: def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]:
"""Move expired pending crawler hits into persistent ``data.crawlers``.""" """Move expired pending crawler hits into persistent ``data.crawlers``.
Returns the client hashes of the newly flushed hits so callers can
schedule async enrichment.
"""
if not self.pending_crawlers: if not self.pending_crawlers:
return return []
now = now or datetime.now(UTC) now = now or datetime.now(UTC)
cutoff = now - _CRAWLER_TIMEOUT cutoff = now - _CRAWLER_TIMEOUT
expired: list[CrawlerHit] = [] expired: list[CrawlerHit] = []
remaining: list[CrawlerHit] = [] remaining: list[CrawlerHit] = []
for hit in self.pending_crawlers: for hit in self.pending_crawlers:
(expired if hit.start <= cutoff else remaining).append(hit) (expired if hit.start <= cutoff else remaining).append(hit)
if expired: if not expired:
self.pending_crawlers = remaining return []
self.data.crawlers.extend(expired) self.pending_crawlers = remaining
self._save() self.data.crawlers.extend(expired)
self._save()
return [hit.client for hit in expired]
def _count(self, table: dict[str, int], key: str) -> None: def _count(self, table: dict[str, int], key: str) -> None:
table[key] = table.get(key, 0) + 1 table[key] = table.get(key, 0) + 1
@@ -267,70 +399,235 @@ class Store:
buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {}) buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {})
self._count(buckets, _bucket(now)) self._count(buckets, _bucket(now))
def _uncount(self, table: dict[str, int], key: str) -> None:
"""Reverse one ``_count``: decrement and drop empty keys."""
if key in table:
table[key] -= 1
if table[key] <= 0:
del table[key]
def _remove_visit(self, index: int) -> None:
"""Delete a visit and reverse every count it recorded.
Used when a known visitor turns out to be an admin (hide=1 ping):
the session is scrubbed from the stats. The in-memory
``_count_log`` tracks each site-visit/view/transition count the
visit produced, so the scrub reverses all of them — including the
ones logged by later pings inside the visit.
"""
for event in self._count_log.pop(index, ()):
kind = event[0]
if kind == "site":
self._uncount(self.data.site_visits, event[1])
elif kind == "view":
views = self.data.views.get(event[1])
if views is not None:
self._uncount(views, event[2])
if not views:
del self.data.views[event[1]]
else: # transition
_, fr, to, bucket = event
fr_map = self.data.transitions.get(fr)
if fr_map is not None:
buckets = fr_map.get(to)
if buckets is not None:
self._uncount(buckets, bucket)
if not buckets:
del fr_map[to]
if not fr_map:
del self.data.transitions[fr]
del self.data.visits[index]
# Sessions and count logs store list indices; shift the ones past
# the removed visit.
for key, i in list(self.sessions.items()):
if i > index:
self.sessions[key] = i - 1
self._count_log = {
i - 1 if i > index else i: log for i, log in self._count_log.items()
}
def _client_ip(self, client_hash: bytes) -> str:
"""Return the IP stored for ``client_hash``, or "" if missing."""
client = self.data.clients.get(client_hash)
return client.ip if client else ""
def _ensure_client(
self,
ip: str,
ua: str,
lang: str,
*,
country: str = "",
) -> bytes:
"""Get or create a ``Client`` record; return its 6-byte hash."""
h = _client_hash(ip, ua, lang)
if h not in self.data.clients:
self.data.clients[h] = Client(
ip=ip,
ua=ua,
ua_pretty=_compact_user_agent(ua),
lang=lang,
country=country,
)
self._save()
return h
def enrich_client(
self,
client_hash: bytes,
*,
host: str = "",
country: str = "",
city: str = "",
) -> None:
"""Fill in host/geoip fields on a client record after async lookups."""
client = self.data.clients.get(client_hash)
if client is None:
return
changed = False
if host and not client.host:
client.host = host
changed = True
if country:
client.country = country
changed = True
if city:
client.city = city
changed = True
if changed:
self._save()
def _abuse_hit(
self,
client_hash: bytes,
path: str,
start: datetime | None = None,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Append one abuse hit referencing a client by hash."""
self.data.abuse.append(
AbuseHit(
start=start or datetime.now(UTC),
path=path,
client=client_hash,
flag=flag,
is_404=is_404,
)
)
def classify_abuse(
self,
ip: str,
client_hash: bytes,
path: str,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Classify an IP as a scanner/abuser and record the triggering hit.
All earlier crawler hits from the same IP (persisted and pending)
are moved to the abuse list — a random-UA scanner must not pollute
the crawler stats of the legitimate bots it impersonates.
"""
if ip not in self.data.abuse_ips:
self.data.abuse_ips[ip] = True
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
if moved:
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip]
for h in moved:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip]
if pending:
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip]
for h in pending:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
self._abuse_hit(client_hash, path, flag=flag, is_404=is_404)
self._save()
def track_404(
self,
ip: str,
ua: str,
path: str,
accept_language: str = "",
) -> bytes:
"""Record a 404 response for ``path`` (full path, query included).
A telltale path (dot segment or *.php) classifies the IP as abuse
immediately; enough plain 404s from one IP do too. Hits from
already-classified IPs go straight to the abuse list.
Returns the client hash so callers can schedule async enrichment.
"""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
self._abuse_hit(client_hash, path, flag=_is_abuse_path(path), is_404=True)
self._save()
return client_hash
if _is_abuse_path(path):
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
self.not_found_counts[ip] = self.not_found_counts.get(ip, 0) + 1
if self.not_found_counts[ip] >= _ABUSE_404_THRESHOLD:
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
return client_hash
def _new_visit( def _new_visit(
self, self,
entry: str, entry: str,
referer: str, referer: str,
key: tuple[str, str], client_hash: bytes,
ip: str = "",
lang: str = "",
country: str = "",
ua: str = "",
utm: dict[str, str] | None = None, utm: dict[str, str] | None = None,
status: int = 200,
) -> Visit: ) -> Visit:
now = datetime.now(UTC) now = datetime.now(UTC)
visit = Visit( visit = Visit(
start=now, start=now,
entry=entry, entry=entry,
referer=referer, referer=referer,
ip=ip, client=client_hash,
lang=lang,
country=country,
ua=ua,
ua_pretty=_compact_user_agent(ua),
utm=utm or {}, utm=utm or {},
) )
visit.statuses[entry] = status
self.data.visits.append(visit) self.data.visits.append(visit)
self.sessions[key] = len(self.data.visits) - 1 index = len(self.data.visits) - 1
self._count(self.data.site_visits, _bucket(now)) self.sessions[client_hash] = index
self._count(self.data.views.setdefault(entry, {}), _bucket(now)) bucket = _bucket(now)
self._count_transition(referer or "(direct)", entry, now) fr = referer or "(direct)"
self._count(self.data.site_visits, bucket)
self._count(self.data.views.setdefault(entry, {}), bucket)
self._count_transition(fr, entry, now)
self._count_log[index] = [
("site", bucket),
("view", entry, bucket),
("transition", fr, entry, bucket),
]
return visit return visit
def enrich_visit(
self,
index: int,
*,
host: str = "",
country: str = "",
city: str = "",
) -> None:
"""Fill in host/geoip fields on an existing visit after async lookups."""
if index < 0 or index >= len(self.data.visits):
return
visit = self.data.visits[index]
changed = False
if host and not visit.host:
visit.host = host
changed = True
if country:
visit.country = country
changed = True
if city:
visit.city = city
changed = True
if changed:
self._save()
def track_entry( def track_entry(
self, self,
referer: str, referer: str,
own_origin: str, own_origin: str,
ip: str, ip: str,
ua: str, ua: str,
entry: str, full_path: str,
query: str = "", accept_language: str = "",
) -> None: *,
status: int = 200,
) -> list[bytes]:
"""Stash the entry referer/UTM tags and queue a pending crawler hit. """Stash the entry referer/UTM tags and queue a pending crawler hit.
Nothing is counted here — the client's initial /_a ping starts the Nothing is counted here — the client's initial /_a ping starts the
@@ -341,11 +638,28 @@ class Store:
does not erase an earlier tagged landing. does not erase an earlier tagged landing.
Every document GET is also queued as a pending crawler hit. If a ping Every document GET is also queued as a pending crawler hit. If a ping
from the same (IP, UA) pair arrives within ``_CRAWLER_TIMEOUT``, the from the same client arrives within ``_CRAWLER_TIMEOUT``, the hit is
hit is discarded; otherwise it is flushed to ``data.crawlers``. discarded; otherwise it is flushed to ``data.crawlers``. The
Accept-Language header is stored on the client record immediately;
host/geoip are filled in later by async enrichment.
GETs from IPs already classified as abuse are recorded as abuse hits
with the full request path (query string included).
Returns the client hashes of any hits flushed to persistent storage,
so callers can schedule async enrichment.
""" """
entry = full_path.split("?")[0]
query = full_path.split("?", 1)[1] if "?" in full_path else ""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
flushed = self._flush_crawlers()
self._abuse_hit(client_hash, full_path, is_404=False, flag=False)
self._save()
return flushed
now = datetime.now(UTC) now = datetime.now(UTC)
self._flush_crawlers(now) flushed = self._flush_crawlers(now)
if referer: if referer:
origin = _origin(referer) origin = _origin(referer)
if origin is not None and origin != own_origin: if origin is not None and origin != own_origin:
@@ -357,71 +671,130 @@ class Store:
CrawlerHit( CrawlerHit(
start=now, start=now,
entry=entry, entry=entry,
ip=ip, client=client_hash,
ua=ua,
ua_pretty=_compact_user_agent(ua),
referer=self.pending_referers.get(ip, ""), referer=self.pending_referers.get(ip, ""),
query=query, query=query,
status=status,
) )
) )
self.pending_statuses.setdefault(client_hash, {})[entry] = status
return flushed
def _add_read(self, client_hash: bytes, path: str, seconds: int) -> None:
"""Add ``seconds`` of reading time for ``path`` to the current visit."""
if seconds <= 0:
return
index = self.sessions.get(client_hash)
if index is None or index >= len(self.data.visits):
return
visit = self.data.visits[index]
visit.read[path] = visit.read.get(path, 0) + seconds
def ping( def ping(
self, self,
from_: str, from_: str,
to: str, to: str | None,
ip: str, ip: str,
ua: str, ua: str,
accept_language: str = "", accept_language: str = "",
) -> int | None: hide: bool = False,
"""Record a client navigation ping ({from, to} from pagerite.js). read: int = 0,
) -> tuple[int | None, list[bytes]]:
"""Record a client navigation ping ({from, to, read} from pagerite.js).
``to`` is an internal path ("/...") or an https URL for exit links; a
missing/empty ``to`` means the page is being closed and only the
``read`` time should be recorded. The transition is always counted when
``to`` is present; the trail only grows on first sight of a page within
the visit. ``read`` is the active time (seconds) spent on ``from_``.
``to`` is an internal path ("/...") or an https URL for exit links;
anything else is ignored. The transition is always counted; the trail
only grows on first sight of a page within the visit.
A ping with no known session starts a fresh visit, consuming the A ping with no known session starts a fresh visit, consuming the
referer and UTM tags stashed by the document GET if there are any. referer and UTM tags stashed by the document GET if there are any.
Returns the index of the new visit when one is created, so callers ``hide`` is set by admin clients: the ping cancels pending crawler
can enrich it later with non-blocking lookups (host, geoip country). hits as usual, and any existing visit for this client session is
removed from the stats (the admin browsed anonymously before logging
in). Nothing new is recorded.
Pings from IPs classified as abuse, and pings whose User-Agent
claims a JS-running crawler identity (``_is_bot_ua``), are ignored
entirely — the crawler's pending hits stay queued and flush to
``data.crawlers`` normally.
Returns the index of the new visit when one is created (or None) and
the client hashes of any crawler hits flushed by this call, so callers
can schedule async enrichment (host, geoip country/city).
""" """
self._flush_crawlers() flushed = self._flush_crawlers()
# A real visitor ping cancels any pending crawler hits from this lang, country = _parse_accept_language(accept_language)
# (IP, UA) pair. client_hash = _client_hash(ip, ua, lang)
if hide:
# Admin ping: cancel pending crawler hits and scrub the session.
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash
]
self.pending_statuses.pop(client_hash, None)
index = self.sessions.pop(client_hash, None)
if index is not None and index < len(self.data.visits):
self._remove_visit(index)
self._save()
return None, flushed
if ip in self.data.abuse_ips:
return None, flushed
if _is_bot_ua(ua):
# A JS-running crawler (Googlebot, GoogleOther, Applebot execute
# JS and ping): never a visit. Its pending crawler hits are
# kept and flush to ``data.crawlers`` normally.
return None, flushed
# A real visitor ping cancels any pending crawler hits from this client.
self.pending_crawlers = [ self.pending_crawlers = [
hit for hit in self.pending_crawlers if not (hit.ip == ip and hit.ua == ua) hit for hit in self.pending_crawlers if hit.client != client_hash
] ]
fr_path = _internal_path(from_) if from_ else ""
if fr_path and read > 0:
self._add_read(client_hash, fr_path, read)
if not to:
if read > 0:
self._save()
return None, flushed
if to.startswith("/") and not to.startswith("//"): if to.startswith("/") and not to.startswith("//"):
target = _internal_path(to) or "" target = _internal_path(to) or ""
else: else:
target = _external_target(to) or "" target = _external_target(to) or ""
if not target: if not target:
return None return None, flushed
key = (ip, ua) index = self.sessions.get(client_hash)
index = self.sessions.get(key) fr = fr_path or "(direct)"
fr = (_internal_path(from_) or "(direct)") if from_ else "(direct)" statuses = self.pending_statuses.setdefault(client_hash, {})
target_status = statuses.pop(target, None) or 200
if not statuses:
self.pending_statuses.pop(client_hash, None)
if index is None or index >= len(self.data.visits): if index is None or index >= len(self.data.visits):
# No known session: the initial ping of a fresh page load (or # No known session: the initial ping of a fresh page load (or
# missing data after a server restart) — start a visit. # missing data after a server restart) — start a visit.
lang, country = _parse_accept_language(accept_language)
index = len(self.data.visits) index = len(self.data.visits)
self._ensure_client(ip, ua, lang, country=country)
self._new_visit( self._new_visit(
target, target,
self.pending_referers.pop(ip, ""), self.pending_referers.pop(ip, ""),
key, client_hash,
ip=ip,
lang=lang,
country=country,
ua=ua,
utm=self.pending_utms.pop(ip, {}), utm=self.pending_utms.pop(ip, {}),
status=target_status,
) )
else: else:
visit = self.data.visits[index] visit = self.data.visits[index]
now = datetime.now(UTC) now = datetime.now(UTC)
bucket = _bucket(now)
log = self._count_log.setdefault(index, [])
if target.startswith("/"): if target.startswith("/"):
self._count(self.data.views.setdefault(target, {}), _bucket(now)) self._count(self.data.views.setdefault(target, {}), bucket)
log.append(("view", target, bucket))
self._count_transition(fr, target, now) self._count_transition(fr, target, now)
log.append(("transition", fr, target, bucket))
# First-seen only: repeat pages and repeated exits don't append. # First-seen only: repeat pages and repeated exits don't append.
if visit.entry != target and target not in visit.trail: if visit.entry != target and target not in visit.trail:
visit.trail.append(target) visit.trail.append(target)
visit.statuses[target] = target_status
self._save() self._save()
return index if index is not None and index < len(self.data.visits) else None visit_index = index if index is not None and index < len(self.data.visits) else None
return visit_index, flushed
+279 -34
View File
@@ -27,14 +27,16 @@ from email.utils import format_datetime
from functools import lru_cache from functools import lru_cache
from pathlib import Path from pathlib import Path
from urllib.parse import urlparse from urllib.parse import urlparse
from xml.sax.saxutils import escape as xml_escape
import blake3 import blake3
import msgspec import msgspec
from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect
from fastapi.responses import HTMLResponse, RedirectResponse, Response from fastapi.responses import RedirectResponse, Response
from fastapi_vue import Frontend from fastapi_vue import Frontend
from kanta import Kanta from kanta import Kanta
from pydantic import BaseModel from pydantic import BaseModel
from zstandard import ZstdCompressor
from pagerite import analytics, seed, views from pagerite import analytics, seed, views
from pagerite.__main__ import DEVMODE from pagerite.__main__ import DEVMODE
@@ -126,13 +128,21 @@ class GeoIP:
return "" return ""
def city(self, ip: str) -> str: def city(self, ip: str) -> str:
"""City name for ``ip``, or "" when unavailable.""" """City name for ``ip``, or "" when unavailable.
GeoIP sometimes appends district names in parentheses (e.g.
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
the value is stored.
"""
if not ip or self._reader is None: if not ip or self._reader is None:
return "" return ""
try: try:
rec = self._reader.get(ip) rec = self._reader.get(ip)
if rec: if rec:
return (rec.get("city") or {}).get("names", {}).get("en", "") city = (rec.get("city") or {}).get("names", {}).get("en", "")
if city:
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
return city
except Exception: except Exception:
pass pass
return "" return ""
@@ -273,6 +283,79 @@ async def _headers(request: Request, call_next) -> Response:
return response return response
# Dynamic HTML is compressed per request at level 9 (static assets are
# already pre-compressed by fastapi-vue's Frontend).
_zstd = ZstdCompressor(9)
def _render_html(kind: str, path: str, base_url: str) -> str:
"""Render one of the generated pages (see _html_response)."""
if kind == "page":
return views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, base_url)
if kind == "category":
return views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
if kind == "not-found":
return views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
return views.render_analytics(data.menu, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
@lru_cache(maxsize=128)
def _cached_body(kind: str, path: str, base_url: str, version: int, zstd: bool) -> bytes:
"""Rendered page body. Every input the output depends on is in the key:
data.version bumps on any content/settings change, base_url feeds the
social meta URLs, and zstd selects the stored encoding (both variants
are cached rather than re-compressed).
"""
body = _render_html(kind, path, base_url).encode()
return _zstd.compress(body) if zstd else body
def _html_response(
request: Request,
kind: str,
path: str,
status_code: int = 200,
headers: dict | None = None,
etag: bool = False,
) -> Response:
"""Response for a generated page, zstd-compressed when the client
accepts it (no gzip fallback).
Done per handler rather than in middleware so that Frontend's
already-compressed asset responses are never touched. The ETag stays
identical across encodings (revalidation compares it before
compression); ``vary: accept-encoding`` keeps caches from mixing the
representations. In dev the cache is bypassed so theme/design edits on
disk apply immediately.
``etag=True`` derives the validator from a blake3 hash of the
(uncompressed) body — for pages like /_a that have no Node whose
modified timestamp could serve as one — and answers matching
if-none-match revalidations with a 304.
"""
zstd = "zstd" in request.headers.get("accept-encoding", "")
# Absolute social/canonical URLs use the learned public origin; until
# an admin visit teaches it, fall back to the request's own base URL.
base_url = data.site_url or str(request.base_url).rstrip("/")
if DEVMODE:
identity = _render_html(kind, path, base_url).encode()
body = _zstd.compress(identity) if zstd else identity
else:
identity = _cached_body(kind, path, base_url, data.version, False)
body = _cached_body(kind, path, base_url, data.version, True) if zstd else identity
h = dict(headers or {})
if zstd:
h["vary"] = "accept-encoding"
if etag:
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
h["etag"] = tag
if request.headers.get("if-none-match") == tag:
return Response(status_code=304, headers=h)
if zstd:
h["content-encoding"] = "zstd"
return Response(body, status_code, h, media_type="text/html")
class PageIn(BaseModel): class PageIn(BaseModel):
"""Payload for creating or replacing a page.""" """Payload for creating or replacing a page."""
@@ -426,6 +509,33 @@ async def put_settings(settings: SettingsIn) -> None:
data.version += 1 data.version += 1
class SiteUrlIn(BaseModel):
"""Payload for learning the site's public origin."""
url: str
@app.post("/_api/site-url", status_code=204)
async def learn_site_url(payload: SiteUrlIn) -> None:
"""Learn the site's public origin (scheme + host) from an admin browser.
pagerite.js reports location.origin once an admin session is detected:
unlike request Host headers it reflects the real public scheme and host
even behind reverse proxies, with zero manual configuration. Stored in
the database with a version bump so cached pages re-render with correct
absolute social/canonical URLs.
"""
url = payload.url.rstrip("/")
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or not parsed.netloc or parsed.path:
raise HTTPException(400, "not an origin")
if url == data.site_url:
return
with kanta.transaction("learn site url"):
data.site_url = url
data.version += 1
@app.put("/_api/settings/favicon") @app.put("/_api/settings/favicon")
async def put_favicon(request: Request) -> dict[str, str]: async def put_favicon(request: Request) -> dict[str, str]:
"""Upload a favicon into the content-addressed store and activate it. """Upload a favicon into the content-addressed store and activate it.
@@ -605,6 +715,12 @@ def _client_ip(request: Request) -> str:
return forwarded or (request.client.host if request.client else "") return forwarded or (request.client.host if request.client else "")
def _query_suffix(request: Request) -> str:
"""The request's query string as a "?..." suffix, or "" when absent."""
query = str(request.url.query)
return f"?{query}" if query else ""
@lru_cache(maxsize=4096) @lru_cache(maxsize=4096)
def _cached_ptr(ip: str) -> str: def _cached_ptr(ip: str) -> str:
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or "".""" """Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
@@ -638,14 +754,21 @@ async def _geoip_city(ip: str) -> str:
return await asyncio.to_thread(_geoip.city, ip) return await asyncio.to_thread(_geoip.city, ip)
async def _enrich_visit(index: int, ip: str) -> None: async def _enrich_client(client_hash: bytes) -> None:
"""Run non-blocking reverse-DNS and geoip enrichment for a new visit.""" """Run non-blocking reverse-DNS and geoip enrichment for a client."""
if not ip: client = analytics_store.data.clients.get(client_hash)
if not client or not client.ip:
return return
host = await _lookup_host(ip) host = await _lookup_host(client.ip)
country = await _geoip_country(ip) country = await _geoip_country(client.ip)
city = await _geoip_city(ip) city = await _geoip_city(client.ip)
analytics_store.enrich_visit(index, host=host, country=country, city=city) analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
"""Start background host/geoip enrichment for the given client hashes."""
for client_hash in client_hashes:
asyncio.create_task(_enrich_client(client_hash))
async def _broadcast_analytics() -> None: async def _broadcast_analytics() -> None:
@@ -683,22 +806,27 @@ class AnalyticsPing(BaseModel):
"""Navigation ping from pagerite.js (see docs/analytics.md).""" """Navigation ping from pagerite.js (see docs/analytics.md)."""
fr: str = "" fr: str = ""
to: str to: str | None = None
#: 1 from admin clients: scrub the session instead of recording it.
hide: int = 0
#: Active reading time on ``fr`` (ms), if any.
read: int = 0
@app.get("/_a", response_model=None) @app.get("/_a", response_model=None)
async def analytics_page(request: Request) -> HTMLResponse: async def analytics_page(request: Request) -> Response:
"""Render the analytics viewer as a normal site page at /_a. """Render the analytics viewer as a normal site page at /_a.
The page itself is public, but the data stream (/_api/ws/analytics) stays The page itself is public, but the data stream (/_api/ws/analytics) stays
admin-gated like the rest of /_api, so only authorized users see the admin-gated like the rest of /_api, so only authorized users see the
statistics; others get the viewer with a "could not be loaded" message. statistics; others get the viewer with a "could not be loaded" message.
""" """
return HTMLResponse( return _html_response(
views.render_analytics( request,
data.menu, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html "analytics",
), "",
headers={"cache-control": "no-cache"}, headers={"cache-control": "no-cache"},
etag=True,
) )
@@ -710,42 +838,59 @@ async def analytics_ping(ping: AnalyticsPing, request: Request) -> None:
the response is never delayed by slow DNS or the first MMDB decompress. the response is never delayed by slow DNS or the first MMDB decompress.
""" """
ip = _client_ip(request) ip = _client_ip(request)
index = analytics_store.ping( visit_index, flushed_clients = analytics_store.ping(
ping.fr, ping.fr,
ping.to, ping.to,
ip, ip,
request.headers.get("user-agent", ""), request.headers.get("user-agent", ""),
request.headers.get("accept-language", ""), request.headers.get("accept-language", ""),
hide=bool(ping.hide),
read=ping.read,
) )
if index is not None: if visit_index is not None:
asyncio.create_task(_enrich_visit(index, ip)) visit = analytics_store.data.visits[visit_index]
asyncio.create_task(_enrich_client(visit.client))
_schedule_client_enrichment(flushed_clients)
def _track_entry(path: str, request: Request) -> None: def _track_entry(path: str, request: Request, *, status: int = 200) -> list[bytes]:
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET. """Stash the referer/UTM tags and queue a pending crawler hit for the GET.
Nothing is counted on the GET itself — the client's /_a ping starts the Nothing is counted on the GET itself — the client's /_a ping starts the
visit, so bots and admin browsing never register as visits. visit, so bots never register as visits (JS-running crawlers ping too,
but the ping handler ignores known bot UAs). (Admin clients ping too,
but with hide=1, which scrubs their session instead of recording it.)
The devserver's health probe (``GET /?from=devserver.py`` from The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be ``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
logged as a crawler hit. The root-path and localhost checks prevent logged as a crawler hit. The root-path and localhost checks prevent
remote visitors from hiding traffic with the same query string. remote visitors from hiding traffic with the same query string.
Returns the client hashes of any pending crawler hits flushed to persistent
storage, so callers can schedule async geoip and reverse-DNS enrichment.
""" """
if request.headers.get("x-pagerite-preload"):
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
# ping sent when the user actually navigates does the counting.
# (Forging the header only hides a GET from the crawler stats; the
# path-based abuse classification is unaffected.)
return []
if ( if (
path == "" path == ""
and str(request.url.query) == "from=devserver.py" and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1" and _client_ip(request) == "127.0.0.1"
): ):
return return []
own_origin = f"https://{urlparse(str(request.base_url)).netloc}" own_origin = f"https://{urlparse(str(request.base_url)).netloc}"
analytics_store.track_entry( full_path = f"{request.url.path}{_query_suffix(request)}"
return analytics_store.track_entry(
request.headers.get("referer", ""), request.headers.get("referer", ""),
own_origin, own_origin,
_client_ip(request), _client_ip(request),
request.headers.get("user-agent", ""), request.headers.get("user-agent", ""),
"/" if path == "" else f"/{path}", full_path,
str(request.url.query), request.headers.get("accept-language", ""),
status=status,
) )
@@ -952,22 +1097,108 @@ async def front_page(request: Request) -> Response:
return await show_page(request, "") return await show_page(request, "")
@app.get("/sitemap.xml")
async def sitemap(request: Request) -> Response:
"""Dynamically generate a sitemap of all published article pages."""
base = str(request.base_url).rstrip("/")
entries: list[tuple[str, datetime, int]] = []
def walk(
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
) -> None:
first_content_slug = next(
(
slug
for slug, node in sorted_nodes(nodes)
if node.published and node.content is not None
),
None,
)
for slug, node in sorted_nodes(nodes):
path = f"{prefix}/{slug}" if prefix else slug
depth = path.count("/") if path else 0
if (
not parent_has_content
and slug == first_content_slug
and node.published
and node.content is not None
and depth > 0
):
depth -= 1
if node.published and node.content is not None:
entries.append((path, node.modified, depth))
if node.children:
walk(node.children, path, node.content is not None)
walk(data.menu, "")
def priority(depth: int) -> float:
return max(0.1, 1.0 - depth * 0.2)
lines = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
]
for path, modified, depth in entries:
loc = xml_escape(f"{base}/{path}" if path else base)
lastmod = (
modified.astimezone(UTC).replace(microsecond=0).isoformat().replace("+00:00", "Z")
)
lines.append(
f" <url>"
f"<loc>{loc}</loc>"
f"<lastmod>{lastmod}</lastmod>"
f"<priority>{priority(depth):.1f}</priority>"
f"</url>"
)
lines.append("</urlset>")
return Response(
"\n".join(lines),
media_type="application/xml",
headers={"cache-control": "no-cache"},
)
@app.get("/robots.txt")
async def robots_txt(request: Request) -> Response:
"""Allow all crawling and point crawlers at the sitemap."""
base = str(request.base_url).rstrip("/")
body = f"User-agent: *\nAllow: /\nSitemap: {base}/sitemap.xml\n"
return Response(
body,
media_type="text/plain",
headers={"cache-control": "no-cache"},
)
# Vue build asset routes are inserted at this position during load(): the # Vue build asset routes are inserted at this position during load(): the
# build mirrors the URL space (/_assets/*, /favicon.ico at the root). # build mirrors the URL space (/_assets/*, /favicon.ico at the root).
frontend.route(app, "/") frontend.route(app, "/")
@app.get("/{path:path}", response_model=None) @app.get("/{path:path}", response_model=None)
async def show_page(request: Request, path: str) -> HTMLResponse | Response: async def show_page(request: Request, path: str) -> Response:
"""Render the content page at a slug path, or 404. """Render the content page at a slug path, or 404.
A node without content is a category label: its URL renders a A node without content is a category label: its URL renders a
placeholder page (nav links point straight at its first child). placeholder page (nav links point straight at its first child).
""" """
path = path.strip("/") path = path.strip("/")
ua = request.headers.get("user-agent", "")
accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path): if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its # Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page. # built-in 404 instead of rendering an editable article page.
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
# as abuse in analytics.
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
raise HTTPException(404) raise HTTPException(404)
chain = resolve(data.menu, path) chain = resolve(data.menu, path)
node = chain[-1] if chain else None node = chain[-1] if chain else None
@@ -982,9 +1213,12 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
if request.headers.get("if-none-match") == etag: if request.headers.get("if-none-match") == etag:
return Response(status_code=304) return Response(status_code=304)
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) flushed = _track_entry(path, request)
return HTMLResponse( _schedule_client_enrichment(flushed)
views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, str(request.base_url).rstrip("/")), return _html_response(
request,
"page",
path,
headers={ headers={
"etag": etag, "etag": etag,
"last-modified": _http_date(node.modified), "last-modified": _http_date(node.modified),
@@ -995,9 +1229,12 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
# Category label without a landing page: placeholder with the pen # Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real). # to create it (404 — no page here, but the node is real).
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) flushed = _track_entry(path, request, status=404)
return HTMLResponse( _schedule_client_enrichment(flushed)
views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), return _html_response(
request,
"category",
path,
404, 404,
headers={ headers={
"last-modified": _http_date(node.modified), "last-modified": _http_date(node.modified),
@@ -1011,5 +1248,13 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
if item.published: if item.published:
return RedirectResponse(f"/{slug}") return RedirectResponse(f"/{slug}")
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) client_hash = analytics_store.track_404(
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), 404) _client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
flushed = _track_entry(path, request, status=404)
_schedule_client_enrichment(flushed)
return _html_response(request, "not-found", path, 404)
+6
View File
@@ -100,6 +100,12 @@ class Data(msgspec.Struct):
#: Favicon: name of a file in `files` (content-addressed), linked as #: Favicon: name of a file in `files` (content-addressed), linked as
#: <link rel="icon"> on every page. Empty = the build's /favicon.ico. #: <link rel="icon"> on every page. Empty = the build's /favicon.ico.
favicon: str = "" favicon: str = ""
#: Public origin (scheme + host) of the site, learned from admin
#: browsers (POST /_api/site-url — location.origin is correct even
#: behind reverse proxies, unlike request Host headers). Used for
#: absolute social/canonical URLs; empty = fall back to the request's
#: own base URL.
site_url: str = ""
#: Legacy flat page store (pre-tree databases); migrated into `menu` #: Legacy flat page store (pre-tree databases); migrated into `menu`
#: on startup, then cleared. Never written otherwise. #: on startup, then cleared. Never written otherwise.
pages: dict[str, Page] = {} pages: dict[str, Page] = {}
+19 -4
View File
@@ -27,15 +27,30 @@
let my = 0 let my = 0
let lastMove = 0 let lastMove = 0
addEventListener('mousemove', e => { // Mouse and touch tracked with separate listeners (pointer events arrive
// too late on some mobile browsers). Passive listeners: a drag on the
// banner still scrolls the page — on browsers that stop delivering
// touchmove once scrolling takes over, the gaze just follows until then.
const track = (x, y) => {
const r = c.getBoundingClientRect() const r = c.getBoundingClientRect()
// Convert viewport coordinates into the canvas' CSS-pixel coordinate // Convert viewport coordinates into the canvas' CSS-pixel coordinate
// system. This remains correct with browser zoom, CSS transforms, etc. // system. This remains correct with browser zoom, CSS transforms, etc.
mx = (e.clientX - r.left) * c.clientWidth / r.width mx = (x - r.left) * c.clientWidth / r.width
my = (e.clientY - r.top) * c.clientHeight / r.height my = (y - r.top) * c.clientHeight / r.height
lastMove = performance.now() lastMove = performance.now()
}) }
addEventListener('mousemove', e => track(e.clientX, e.clientY))
const trackTouch = e => {
const t = e.touches[0]
if (t) track(t.clientX, t.clientY)
}
addEventListener('touchstart', trackTouch, { passive: true })
addEventListener('touchmove', trackTouch, { passive: true })
let gx = 0.5 let gx = 0.5
let gy = 0.5 let gy = 0.5
+114 -30
View File
@@ -110,6 +110,35 @@ def _editor_css_url(vite_url: str | None) -> str | None:
return None return None
def _inline_asset(url: str) -> str:
"""Read a served asset's content for inlining into the page (prod only).
Handles build assets (``/_assets/...`` from the Vite build) and theme
files (``/_themes/{name}/...`` from pagerite/themes/).
"""
if url.startswith("/_themes/"):
name, _, file = url.removeprefix("/_themes/").partition("/")
if _valid_name(name) and _valid_name(file):
return (THEMES / name / file).read_text()
raise ValueError(f"not a theme asset: {url}")
return (BUILD / url.lstrip("/")).read_text()
def _inline_script(url: str) -> str:
"""Read a built JS bundle for inlining (prod only).
Inline modules resolve relative imports against the document URL, not
the bundle's directory, so rewrite the build's relative chunk
specifiers ("./chunk.js") to absolute /_assets/ paths.
"""
js = _inline_asset(url)
for chunk in _manifest().values():
file = chunk.get("file", "")
if file.endswith(".js"):
js = js.replace(f'"./{file.rsplit("/", 1)[-1]}"', f'"/{file}"')
return js
def _layout( def _layout(
modules: list[str] = (), modules: list[str] = (),
stylesheets: list[str] = (), stylesheets: list[str] = (),
@@ -118,16 +147,21 @@ def _layout(
banner_design: str = "", banner_design: str = "",
favicon: str = "", favicon: str = "",
social: dict[str, str] | None = None, social: dict[str, str] | None = None,
extra_meta: dict[str, str] | None = None,
) -> Template: ) -> Template:
"""Page layout template with standard asset URLs and ES-module scripts. """Page layout template with standard assets and ES-module scripts.
Stylesheets use ``blocking="render"`` so the browser waits for them before In dev (PAGERITE_VITE_URL set) assets are linked from the Vite dev
showing the page, avoiding a flash of unstyled content. Order matters and server and stylesheets use ``blocking="render"`` so the browser waits
is fixed: base (Vite build, absent in dev where Vite injects it from JS), for them before showing the page, avoiding a flash of unstyled content.
theme and banner design (backend-served from pagerite/themes/), entry- In production all page assets are inlined into the document: stylesheets
specific stylesheets (e.g. overlayscrollbars.css), then the user's custom become ``<style>`` elements and module scripts inline ``<script>``s, so
CSS last so it always wins. a page loads with no asset round trips. The on-demand bundles (editor,
analytics) stay external in both modes.
Order matters and is fixed: base (Vite build, absent in dev where Vite
injects it from JS), theme and banner design (from pagerite/themes/),
entry-specific stylesheets (e.g. overlayscrollbars.css), then the user's
custom CSS last so it always wins.
In dev, pagerite.js re-appends the backend-rendered theme/design links In dev, pagerite.js re-appends the backend-rendered theme/design links
(and the custom CSS) after the Vite-injected base styles, keeping this (and the custom CSS) after the Vite-injected base styles, keeping this
@@ -135,9 +169,6 @@ def _layout(
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as ``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
property attributes, everything else (description, twitter:*) as name. property attributes, everything else (description, twitter:*) as name.
``extra_meta`` is emitted as plain ``<meta name="..." content="...">``
tags after the editor meta tags; used for page-specific import hints.
""" """
doc = Document(E.Title, lang="en") doc = Document(E.Title, lang="en")
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs # Responsive layout (see the 48rem breakpoint in pagerite.css) needs
@@ -155,33 +186,63 @@ def _layout(
# one, browsers fall back to the build's /favicon.ico by convention. # one, browsers fall back to the build's /favicon.ico by convention.
if favicon: if favicon:
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon") doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
# Editor asset URLs for pagerite.js, which injects the 🖊️ edit pens # Asset URLs for the on-demand bundles (editor, analytics) for
# itself once it has validated the session (pages render identically # pagerite.js, which injects the 🖊️ edit pens itself once it has
# for everyone; editing is gated by the auth proxy in front of /_api). # validated the session (pages render identically for everyone; editing
script, editor_css = _editor_assets() # is gated by the auth proxy in front of /_api). Dev passes the Vite
doc.meta(name="pagerite:editor-src", content=script[-1]) # dev-server URLs as meta tags (Vite serves the modules and injects
if editor_css: # their CSS for hot reloads); production inlines all page assets and
doc.meta(name="pagerite:editor-css", content=editor_css) # carries the on-demand URLs in one JSON script instead.
for key, value in (extra_meta or {}).items():
doc.meta(name=key, content=value)
# Stylesheet links carry stable ids so the site editor's hot swap can
# keep each sheet at its rendered position (see swapRegions).
vite_url = os.environ.get("PAGERITE_VITE_URL") vite_url = os.environ.get("PAGERITE_VITE_URL")
editor_scripts, editor_css = _editor_assets()
config = {
"pagerite:editor-src": editor_scripts[-1],
"pagerite:analytics-src": _analytics_assets()[0][0],
}
if editor_css:
config["pagerite:editor-css"] = editor_css
if vite_url:
for key, value in config.items():
doc.meta(name=key, content=value)
else:
# Inert JSON script; URLs never contain "</", but stay safe.
doc.script(
HTML(json.dumps(config).replace("</", "<\\/")),
type="application/json",
id="pagerite-assets",
)
# Stylesheets carry stable ids so the fetch-navigation and the site
# editor's hot swap can sync <head> positionally (see swapdoc.js).
# Production inlines the CSS as <style> elements: one less round trip
# per sheet, and fetch-navigation can carry them across swaps whole.
sheets = [ sheets = [
("pagerite-base", _base_css_url(vite_url)), ("pagerite-base", _base_css_url(vite_url)),
("pagerite-theme", _theme_css_url(theme)), ("pagerite-theme", _theme_css_url(theme)),
("pagerite-banner", _banner_css_url(banner_design)), ("pagerite-banner", _banner_css_url(banner_design)),
] ]
for id_, url in sheets: for id_, url in sheets:
if url: if not url:
continue
if vite_url:
doc.link(rel="stylesheet", href=url, blocking="render", id=id_) doc.link(rel="stylesheet", href=url, blocking="render", id=id_)
else:
doc.style(HTML(_inline_asset(url)), id=id_)
for url in stylesheets: for url in stylesheets:
doc.link(rel="stylesheet", href=url, blocking="render") if vite_url:
doc.link(rel="stylesheet", href=url, blocking="render")
else:
# Id from the file stem minus the content hash, so the head
# sync can match sheets across pages (e.g. the analytics sheet
# exists on /_a only and is added/removed on swaps).
stem = url.rsplit("/", 1)[-1].removesuffix(".css")
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
doc.style(HTML(_inline_asset(url)), id=f"pagerite-css-{name}")
for src in modules: for src in modules:
doc.script(src=src, type="module") if vite_url:
doc.script(src=src, type="module")
if custom_css.strip(): if custom_css.strip():
doc.style(custom_css, id="pagerite-user") doc.style(custom_css, id="pagerite-user")
return Template( body = (
doc doc
.header( .header(
E.div(E.Banner, id="page-banner"), E.div(E.Banner, id="page-banner"),
@@ -194,8 +255,23 @@ def _layout(
E.main(E.Main, id="main"), E.main(E.Main, id="main"),
id="content", id="content",
) )
.footer(None), # kept empty for now; zero-height (see pagerite.css) .footer(None) # kept empty for now; zero-height (see pagerite.css)
) )
if not vite_url:
# Inline the bundles at the end of the body: module scripts are
# deferred anyway, and the page can render before they execute.
# Escape "</script" so it cannot terminate the element early (only
# ever occurs inside string literals, where the backslash escape is
# a no-op).
for src in modules:
js = re.sub(r"</script", r"<\\/script", _inline_script(src), flags=re.I)
# Stable id from the file stem minus the content hash; the
# analytics page's script (pagerite-js-analytics) is found and
# re-created by pagerite.js on fetch-navigations to /_a.
stem = src.rsplit("/", 1)[-1].removesuffix(".js")
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
body.script(HTML(js), type="module", id=f"pagerite-js-{name}")
return Template(body)
def _brand_link(brand: str, brand_html: str = "") -> HTML: def _brand_link(brand: str, brand_html: str = "") -> HTML:
@@ -682,14 +758,23 @@ def render_analytics(
favicon: str = "", favicon: str = "",
brand_html: str = "", brand_html: str = "",
) -> str: ) -> str:
"""Render the analytics viewer as a normal page at /_a.""" """Render the analytics viewer as a normal page at /_a.
The analytics entry is inlined into this page only (prod) or loaded
from the Vite dev server (dev); its stylesheet rides along in <head>
so fetch-navigations can sync it into the live document. The initial
range is not rendered in: the client takes it from the URL hash or
derives it from the analytics data itself.
"""
page_scripts, page_stylesheets = _page_assets() page_scripts, page_stylesheets = _page_assets()
analytics_scripts, analytics_stylesheets = _analytics_assets() analytics_scripts, analytics_stylesheets = _analytics_assets()
scripts = page_scripts + analytics_scripts scripts = page_scripts + analytics_scripts
stylesheets = page_stylesheets + analytics_stylesheets stylesheets = page_stylesheets + analytics_stylesheets
doc = E.article doc = E.article
with doc: with doc:
doc.div(id="analytics-app") # .wide: the dashboard breaks out of the article column to the full
# viewport width, like wide figures (see the .wide rules).
doc.div(id="analytics-app", class_="wide")
return str( return str(
_layout( _layout(
scripts, scripts,
@@ -698,7 +783,6 @@ def render_analytics(
theme, theme,
banner_design(menu, "_a", theme), banner_design(menu, "_a", theme),
favicon, favicon,
extra_meta={"pagerite:analytics-src": analytics_scripts[0]},
)( )(
Title=f"Analytics {brand}" if brand else "Analytics", Title=f"Analytics {brand}" if brand else "Analytics",
Brand=_brand_link(brand, brand_html), Brand=_brand_link(brand, brand_html),
+1
View File
@@ -27,6 +27,7 @@ dependencies = [
"mdit-py-plugins>=0.6.1", "mdit-py-plugins>=0.6.1",
"pygments>=2.20.0", "pygments>=2.20.0",
"ua-parser>=1.0.2", "ua-parser>=1.0.2",
"zstandard>=0.25.0",
] ]
[project.scripts] [project.scripts]
+594 -146
View File
@@ -6,23 +6,18 @@
# "playwright>=1.45.0", # "playwright>=1.45.0",
# ] # ]
# /// # ///
"""Generate fake browser visits and crawler hits for a Pagerite site. """Generate fake browser visits, crawler hits, and abuse scans for a Pagerite site.
The script drives a real Chromium browser with Playwright, clicking visible Browser sessions (ordinary users) come from realistic residential IPv4 and IPv6
internal links so the site's own analytics JavaScript records normal visits addresses and stay mostly stable; an IPv6 host part may rotate once mid-session,
(POST /_a). Most browser sessions enter the site with a cross-origin and an IPv4 session may switch to another residential address. Crawler hits come
``Referer: https://somedomain.com/`` header, and outbound links found on the from datacenter IPs, with each crawler profile paired to a matching provider IP
page are followed to real external sites (ending the session). Browser when possible. Abuse scanners fire bursts of vulnerability probes from pinned
sessions and crawler GETs send a small rotating pool of real public IPs in datacenter IPs.
X-Forwarded-For, so the backend can reverse-DNS and GeoIP them instead of seeing
every hit as 127.0.0.1.
Sessions start with a Poisson inter-arrival delay (``--arrival-rate``) to
spread traffic out a little, while still keeping the overall run fast.
Run against a local dev server, e.g.: Run against a local dev server, e.g.:
uv run scripts/fake_traffic.py http://localhost:3200 -b 8 -c 20 uv run scripts/fake_traffic.py http://localhost:3200
Repeat whenever you want more traffic; each run appends new events to the Repeat whenever you want more traffic; each run appends new events to the
site's analytics file. site's analytics file.
@@ -39,7 +34,7 @@ from collections.abc import Sequence
from dataclasses import dataclass from dataclasses import dataclass
from datetime import UTC, datetime from datetime import UTC, datetime
from typing import Any from typing import Any
from urllib.parse import urljoin, urlparse from urllib.parse import urlencode, urljoin, urlparse
import httpx import httpx
@@ -59,6 +54,7 @@ class BrowserProfile:
class CrawlerProfile: class CrawlerProfile:
name: str name: str
user_agent: str user_agent: str
ip: str
BROWSER_PROFILES: list[BrowserProfile] = [ BROWSER_PROFILES: list[BrowserProfile] = [
@@ -95,39 +91,397 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
CrawlerProfile( CrawlerProfile(
"googlebot", "googlebot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"66.249.64.66", # US, Google
), ),
CrawlerProfile( CrawlerProfile(
"bingbot", "bingbot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"40.77.167.0", # US, Microsoft
), ),
CrawlerProfile( CrawlerProfile(
"duckduckbot", "DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)" "duckduckbot",
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
"95.217.0.1", # Germany, Hetzner VPS
),
CrawlerProfile(
"curl",
"curl/8.5.0",
"139.162.0.1", # Singapore, Linode VPS
), ),
CrawlerProfile("curl", "curl/8.5.0"),
] ]
# Small pool of real public resolver IPs. They have real reverse-DNS and GeoIP # Residential IPv4 addresses and IPv6 /64 prefixes used for ordinary browser
# entries, and cycling through a handful avoids hammering DNS during traffic # sessions. IPv6 entries keep the network part stable and randomise only the
# generation. # host part; the host may rotate once mid-session.
SOURCE_IPS: list[str] = [ RESIDENTIAL_SOURCE_IPS: list[str] = [
"8.8.8.8", # Residential IPv4
"1.1.1.1", "91.154.140.209", # Finland, Elisa
"9.9.9.9", "84.143.145.207", # Germany, Deutsche Telekom
"208.67.222.222", "220.165.255.254", # China, Chinanet / China Telecom
"185.228.168.9", "84.235.83.162", # Saudi Arabia, SaudiNet / STC
"94.140.14.14", # Residential IPv6 /64 prefixes
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
"240e:45d:1e60:5b0::/64", # China, China Telecom
"2409:8904:6720:4123::/64", # China, China Unicom
] ]
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
# the whole scan burst.
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
# and index 2 uses a fixed crawler UA.
ABUSE_SOURCE_IPS: list[str] = [
"45.63.0.12", # US, Vultr VPS
"138.197.0.89", # US, DigitalOcean / Cloudways
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
]
# Paths commonly probed by attackers looking for exposed config, admin panels,
# version control, credentials, backups, or debug endpoints.
SUSPICIOUS_PATHS: list[str] = [
"/.env",
"/env",
"/.env.local",
"/env.development",
"/config",
"/config.json",
"/config.yaml",
"/config.yml",
"/configuration.json",
"/configuration.yaml",
"/configuration.yml",
"/settings.json",
"/settings.yaml",
"/settings.yml",
"/app.config",
"/appsettings.json",
"/appsettings.Development.json",
"/credentials",
"/credentials.json",
"/secrets",
"/secrets.json",
"/.aws/credentials",
"/.ssh/id_rsa",
"/id_rsa",
"/id_rsa.pub",
"/known_hosts",
"/sftp-config.json",
"/admin",
"/administrator",
"/adminer.php",
"/login",
"/signin",
"/auth/login",
"/api/login",
"/api/.env",
"/api/config",
"/api/v1/config",
"/api/v2/config",
"/webhook",
"/webhooks",
"/callback",
"/proxy",
"/image",
"/images",
"/preview",
"/download",
"/downloads",
"/log",
"/logs",
"/debug",
"/trace",
"/phpinfo.php",
"/info.php",
"/phpmyadmin",
"/pma",
"/myadmin",
"/phpMyAdmin",
"/wp-admin",
"/wp-login.php",
"/wp-config.php",
"/xmlrpc.php",
"/wp-json/wp/v2/users",
"/.git/config",
"/.git/HEAD",
"/git/config",
"/swagger-ui.html",
"/v2/api-docs",
"/actuator/env",
"/actuator/health",
"/actuator/configprops",
"/server-status",
"/.htaccess",
"/web.config",
"/package.json",
"/composer.json",
"/vendor/autoload.php",
"/docker-compose.yml",
"/Dockerfile",
"/manage",
"/console",
"/manager",
"/manager/html",
"/metrics",
"/prometheus",
"/healthz",
"/_api",
"/api",
"/api/v1/",
"/api/v2/",
"/graphql",
"/query",
"/feed",
"/rss",
"/_debug",
"/test",
"/testing",
"/tmp",
"/temp",
"/backup",
"/backups",
"/dump",
"/dumps",
"/sql",
"/db",
"/database",
"/dump.sql",
"/backup.sql",
"/db.sql",
"/backup.zip",
"/backup.tar.gz",
"/site.zip",
"/site.tar.gz",
"/source.zip",
"/src.zip",
"/upload",
"/uploads",
"/import",
"/export",
"/token",
"/tokens",
"/oauth",
"/oauth2",
"/openid",
"/jwks",
"/keys",
"/key",
"/private",
"/public",
]
# Realistic external referers. Most sessions arrive with a generic referer;
# a subset carries matching UTM tags on the landing URL.
PLAIN_REFERRERS: list[str] = [
"https://example.com/",
"https://somedomain.com/",
"https://another-site.org/",
"https://friend-site.net/",
]
# (referer origin, utm parameter dict) pairs used for tagged traffic.
TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
("https://chatgpt.com/", {"utm_source": "chatgpt.com"}),
("https://www.google.com/", {"utm_source": "google", "utm_medium": "organic"}),
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
]
# Fraction of referered sessions that also carry UTM tags.
UTM_RATE = 0.25
# Innocent-looking paths that do not exist on a Pagerite site. Hitting many of
# these from a single IP is itself a telltale of a spray-and-pray scanner.
NORMAL_404_PATHS: list[str] = [
"/about",
"/about-us",
"/services",
"/products",
"/contact",
"/contact-us",
"/team",
"/careers",
"/jobs",
"/pricing",
"/features",
"/demo",
"/trial",
"/docs",
"/documentation",
"/api-docs",
"/support",
"/help",
"/faq",
"/knowledge-base",
"/terms",
"/terms-of-service",
"/privacy",
"/privacy-policy",
"/legal",
"/blog",
"/news",
"/articles",
"/press",
"/events",
"/webinars",
"/podcast",
"/videos",
"/resources",
"/whitepapers",
"/case-studies",
"/customers",
"/clients",
"/testimonials",
"/reviews",
"/partners",
"/integrations",
"/api-reference",
"/developers",
"/status",
"/security",
"/trust",
"/compliance",
"/gdpr",
"/ccpa",
"/sitemap",
"/archive",
"/tags",
"/categories",
"/search",
"/users",
"/accounts",
"/dashboard",
"/profile",
"/settings",
"/preferences",
"/notifications",
"/messages",
"/inbox",
"/calendar",
"/reports",
"/analytics",
"/billing",
"/invoice",
"/orders",
"/cart",
"/checkout",
"/store",
"/shop",
"/home",
"/main",
"/start",
"/welcome",
"/intro",
"/overview",
"/summary",
"/portfolio",
"/projects",
"/work",
"/solutions",
]
ABUSE_USER_AGENTS: list[str] = [
# Desktop browsers
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_5 like Mac OS X) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Mobile/15E148 Safari/604.1",
# Well-known crawlers / bots
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; "
"+http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; "
"+http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html)",
"Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36 "
"(compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)",
"Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com)",
"Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html)",
"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)",
# Social / service fetchers
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
"Twitterbot/1.0",
"LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)",
"Slackbot-LinkExpanding 1.0 (+https://api.slack.com/robots)",
"WhatsApp/2.23.20.0",
# Command-line / library clients
"curl/8.5.0",
"Wget/1.21.4 (linux-gnu)",
"python-requests/2.32.3",
"Go-http-client/1.1",
"Node.js/20.5.1",
]
def _random_ipv6_host(prefix: str) -> str:
"""Return a concrete address within an IPv6 /64 prefix.
The host part is generated randomly, mimicking a fresh OS privacy address.
The input prefix must end in ``::/64`` (e.g. ``2a02:8109:ac82:6f0c::/64``).
"""
if "/" not in prefix:
return prefix
base, mask = prefix.split("/")
if mask != "64":
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
if base.endswith("::"):
base = base[:-2]
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
return f"{base}:{host}"
def _concretize_ip(entry: str) -> str:
"""Return a concrete IP address; randomise the host part for IPv6 /64 prefixes."""
if ":" in entry and "/" in entry:
return _random_ipv6_host(entry)
return entry
class _SessionIP:
"""Stable IP for a browser session, with one optional mid-session rotation.
IPv6 prefixes get a fresh random host part; IPv4 addresses are swapped for
another address from the residential pool.
"""
def __init__(self, entry: str, pool: Sequence[str]):
self.entry = entry
self.pool = pool
self._value = _concretize_ip(entry)
def current(self) -> str:
return self._value
def rotate(self) -> None:
if ":" in self.entry and "/" in self.entry:
self._value = _random_ipv6_host(self.entry)
return
# IPv4: switch to another IPv4 address from the residential pool.
for _ in range(20):
candidate_entry = random.choice(self.pool)
if ":" in candidate_entry and "/" in candidate_entry:
continue
candidate = _concretize_ip(candidate_entry)
if candidate != self._value:
self._value = candidate
return
def _sleep(base: float, jitter: float) -> None: def _sleep(base: float, jitter: float) -> None:
time.sleep(max(0.0, base + random.uniform(-jitter, jitter))) time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
def _source_ip(index: int) -> str:
"""Pick one of the small pool of real public IPs."""
return SOURCE_IPS[index % len(SOURCE_IPS)]
def _normalize_url(url: str) -> str: def _normalize_url(url: str) -> str:
"""Return a usable base URL, adding missing scheme/host/port parts. """Return a usable base URL, adding missing scheme/host/port parts.
@@ -232,43 +586,75 @@ def _run_browser_session(
paths: Sequence[str], paths: Sequence[str],
profile: BrowserProfile, profile: BrowserProfile,
session_index: int, session_index: int,
max_clicks: int, ip_entry: str,
stay: tuple[float, float],
headless: bool,
fake_ip: str,
referer_rate: float,
include_external: bool = True,
) -> dict[str, Any]: ) -> dict[str, Any]:
from playwright.sync_api import sync_playwright from playwright.sync_api import sync_playwright
MAX_CLICKS = 6
STAY = (2.0, 6.0)
HEADLESS = True
REFERER_RATE = 0.75
INCLUDE_EXTERNAL = True
ip_provider = _SessionIP(ip_entry, RESIDENTIAL_SOURCE_IPS)
ips_used: list[str] = [ip_provider.current()]
trail: list[str] = [] trail: list[str] = []
start_time = datetime.now(UTC) start_time = datetime.now(UTC)
try: try:
with sync_playwright() as p: with sync_playwright() as p:
browser = p.chromium.launch( browser = p.chromium.launch(
headless=headless, headless=HEADLESS,
args=["--no-sandbox", "--disable-dev-shm-usage"], args=["--no-sandbox", "--disable-dev-shm-usage"],
) )
extra_headers = { extra_headers = {
"X-Forwarded-For": fake_ip, "X-Forwarded-For": ip_provider.current(),
"Accept-Language": profile.accept_language, "Accept-Language": profile.accept_language,
} }
# Most sessions arrive from an external origin; some are direct. # Most sessions arrive from an external origin; some are direct.
if random.random() < referer_rate: # A subset of referered sessions carries realistic UTM tags on the
extra_headers["Referer"] = "https://somedomain.com/" # landing URL; the referer origin is paired with the UTM source.
tagged: dict[str, str] = {}
if random.random() < REFERER_RATE:
if random.random() < UTM_RATE:
referer, tagged = random.choice(TAGGED_REFERRERS)
else:
referer = random.choice(PLAIN_REFERRERS)
extra_headers["Referer"] = referer
context = browser.new_context( context = browser.new_context(
user_agent=profile.user_agent, user_agent=profile.user_agent,
viewport={"width": profile.viewport[0], "height": profile.viewport[1]}, viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
extra_http_headers=extra_headers, extra_http_headers=extra_headers,
) )
page = context.new_page() page = context.new_page()
# Update X-Forwarded-For per request; the value stays stable unless we
# explicitly rotate it once mid-session.
def _route_handler(route, request):
headers = dict(request.headers)
headers["X-Forwarded-For"] = ip_provider.current()
ips_used.append(headers["X-Forwarded-For"])
route.continue_(headers=headers)
page.route("**/*", _route_handler)
# Pick one point during the session to emulate an IP rotation.
rotate_at = random.randint(0, MAX_CLICKS - 1) if MAX_CLICKS > 0 else -1
entry = random.choice(paths) if paths else "/" entry = random.choice(paths) if paths else "/"
page.goto(urljoin(base, entry), wait_until="networkidle") landing = urljoin(base, entry)
if tagged:
sep = "&" if "?" in landing else "?"
landing += sep + urlencode(tagged)
page.goto(landing, wait_until="networkidle")
trail.append(page.url) trail.append(page.url)
for _ in range(max_clicks): for click_idx in range(MAX_CLICKS):
_sleep(random.uniform(*stay) / 2, 0.3) _sleep(random.uniform(*STAY) / 2, 0.3)
links = _collect_links(page, include_external) if click_idx == rotate_at:
ip_provider.rotate()
ips_used.append(ip_provider.current())
logger.debug("rotated session IP to %s", ip_provider.current())
links = _collect_links(page, INCLUDE_EXTERNAL)
visible = [item for item in links if item.get("visible")] visible = [item for item in links if item.get("visible")]
if not visible: if not visible:
visible = links visible = links
@@ -291,14 +677,15 @@ def _run_browser_session(
break break
page.wait_for_load_state("networkidle") page.wait_for_load_state("networkidle")
trail.append(page.url) trail.append(page.url)
_sleep(random.uniform(*stay), 0.5) _sleep(random.uniform(*STAY), 0.5)
browser.close() browser.close()
return { return {
"profile": profile.name, "profile": profile.name,
"entry": entry, "entry": entry,
"ip": fake_ip, "ip": ips_used[0],
"ips_seen": len(set(ips_used)),
"pages": len(trail), "pages": len(trail),
"trail": [urlparse(u).path or "/" for u in trail], "trail": [urlparse(u).path or "/" for u in trail],
"duration": (datetime.now(UTC) - start_time).total_seconds(), "duration": (datetime.now(UTC) - start_time).total_seconds(),
@@ -312,12 +699,10 @@ def _run_crawler_hit(
base: str, base: str,
paths: Sequence[str], paths: Sequence[str],
profile: CrawlerProfile, profile: CrawlerProfile,
profile_index: int,
session_index: int,
) -> dict[str, Any]: ) -> dict[str, Any]:
path = random.choice(paths) if paths else "/" path = random.choice(paths) if paths else "/"
url = urljoin(base, path) url = urljoin(base, path)
fake_ip = _source_ip(session_index) fake_ip = profile.ip
headers = { headers = {
"User-Agent": profile.user_agent, "User-Agent": profile.user_agent,
"X-Forwarded-For": fake_ip, "X-Forwarded-For": fake_ip,
@@ -337,6 +722,78 @@ def _run_crawler_hit(
return {"profile": profile.name, "path": path, "error": str(exc)} return {"profile": profile.name, "path": path, "error": str(exc)}
def _abuse_ua() -> str:
"""Return a randomized, syntactically valid user agent for an abuse scan."""
return random.choice(ABUSE_USER_AGENTS)
def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
"""Fire a burst of vulnerability probes from a single fake IP.
Scanner 0 randomises its user agent every request, scanner 1 uses a fixed
browser UA, and scanner 2 uses a fixed crawler UA.
"""
ip_entry = ABUSE_SOURCE_IPS[ip_index % len(ABUSE_SOURCE_IPS)]
if ":" in ip_entry and "/" in ip_entry:
fake_ip = _random_ipv6_host(ip_entry)
else:
fake_ip = ip_entry
MIN_HITS = 15
MAX_HITS = 25
total_hits = random.randint(MIN_HITS, MAX_HITS)
# Ensure the burst contains both telltales: suspicious paths and more
# than ten normal-looking 404 paths.
suspicious_count = max(5, total_hits // 3)
normal_count = total_hits - suspicious_count
if normal_count < 11:
normal_count = 11
suspicious_count = max(3, total_hits - normal_count)
paths = random.choices(SUSPICIOUS_PATHS, k=suspicious_count) + random.choices(
NORMAL_404_PATHS, k=normal_count
)
random.shuffle(paths)
ua_mode = ip_index % 3
if ua_mode == 0:
get_ua = _abuse_ua
elif ua_mode == 1:
def get_ua() -> str:
return BROWSER_PROFILES[0].user_agent
else:
def get_ua() -> str:
return CRAWLER_PROFILES[0].user_agent
scan_results: list[dict[str, Any]] = []
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
for path in paths:
headers = {
"User-Agent": get_ua(),
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": random.choice(
["en-US,en;q=0.9", "en-GB,en;q=0.8", "en;q=0.7"]
),
}
try:
r = client.get(urljoin(base, path), headers=headers)
scan_results.append(
{"path": path, "status": r.status_code, "ua": headers["User-Agent"]}
)
except Exception as exc: # noqa: BLE001
scan_results.append({"path": path, "error": str(exc)})
_sleep(0.15, 0.1)
return {
"scanner": ip_index + 1,
"ip": fake_ip,
"hits": len(scan_results),
"results": scan_results,
}
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace: def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
description="Generate fake traffic for a Pagerite site.", description="Generate fake traffic for a Pagerite site.",
@@ -351,54 +808,13 @@ def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
"missing scheme defaults to http://.", "missing scheme defaults to http://.",
) )
parser.add_argument( parser.add_argument(
"-b", "-t",
"--browsers", "--duration",
type=int,
default=5,
help="Number of simulated browser sessions",
)
parser.add_argument(
"-c", "--crawlers", type=int, default=10, help="Number of crawler HTTP GETs"
)
parser.add_argument(
"--max-clicks",
type=int,
default=6,
help="Max internal link clicks per browser session",
)
parser.add_argument(
"--stay",
type=float, type=float,
nargs=2, default=60.0,
default=[2.0, 6.0], metavar="SECONDS",
metavar=("MIN", "MAX"), help="Rough maximum time to generate traffic (0 runs one preset batch)",
help="Seconds to stay on a page before clicking again",
) )
parser.add_argument(
"--headless",
action=argparse.BooleanOptionalAction,
default=True,
help="Run browsers headlessly",
)
parser.add_argument(
"--arrival-rate",
type=float,
default=1.0,
help="Average arrivals per second (Poisson). 0 disables inter-arrival waits",
)
parser.add_argument(
"--referer-rate",
type=float,
default=0.75,
help="Share of browser sessions that arrive with a cross-origin Referer",
)
parser.add_argument(
"--external-links",
action=argparse.BooleanOptionalAction,
default=True,
help="Include real outbound links in random navigation",
)
parser.add_argument("--seed", type=int, default=None, help="Random seed")
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging") parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
return parser.parse_args(argv) return parser.parse_args(argv)
@@ -414,8 +830,6 @@ def main(argv: Sequence[str] | None = None) -> int:
logger.error("%s", exc) logger.error("%s", exc)
return 2 return 2
random.seed(args.seed)
# Discover content paths from the public page tree if we can. # Discover content paths from the public page tree if we can.
paths: list[str] = [] paths: list[str] = []
try: try:
@@ -428,61 +842,95 @@ def main(argv: Sequence[str] | None = None) -> int:
paths = ["/"] paths = ["/"]
logger.info( logger.info(
"Generating fake traffic against %s (%d content paths, %d browsers, %d crawlers)", "Generating fake traffic against %s (%d content paths, duration=%ss)",
base, base,
len(paths), len(paths),
args.browsers, args.duration,
args.crawlers,
) )
results: list[dict[str, Any]] = [] results: list[dict[str, Any]] = []
arrival_rate = 1.0
for i in range(args.browsers): def _wait() -> None:
if i > 0: wait = _poisson_wait(arrival_rate)
wait = _poisson_wait(args.arrival_rate) logger.debug("waiting %.2fs before next session", wait)
logger.debug("waiting %.2fs before next browser session", wait) time.sleep(wait)
time.sleep(wait)
profile = random.choice(BROWSER_PROFILES)
fake_ip = _source_ip(i)
logger.info(
"[%d/%d] browser session: %s (ip=%s)",
i + 1,
args.browsers,
profile.name,
fake_ip,
)
result = _run_browser_session(
base,
paths,
profile,
i,
args.max_clicks,
(args.stay[0], args.stay[1]),
args.headless,
fake_ip,
args.referer_rate,
args.external_links,
)
results.append(result)
logger.debug(" trail: %s", result.get("trail", []))
for i in range(args.crawlers): if args.duration <= 0:
if i > 0: # One preset batch.
wait = _poisson_wait(args.arrival_rate) for i in range(5):
logger.debug("waiting %.2fs before next crawler hit", wait) if i > 0:
time.sleep(wait) _wait()
profile_index = i % len(CRAWLER_PROFILES) profile = random.choice(BROWSER_PROFILES)
profile = CRAWLER_PROFILES[profile_index] ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
fake_ip = _source_ip(i) logger.info(
logger.info( "browser session: %s (ip=%s)",
"[%d/%d] crawler hit: %s (ip=%s)", profile.name,
i + 1, _concretize_ip(ip_entry),
args.crawlers, )
profile.name, result = _run_browser_session(base, paths, profile, i, ip_entry)
fake_ip, results.append(result)
) logger.debug(" trail: %s", result.get("trail", []))
result = _run_crawler_hit(base, paths, profile, profile_index, i)
results.append(result) for i in range(10):
if i > 0:
_wait()
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
results.append(result)
for i in range(3):
if i > 0:
_wait()
ip_entry = ABUSE_SOURCE_IPS[i % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, i)
results.append(result)
logger.debug(
" hits: %s", [r.get("path") for r in result.get("results", [])]
)
else:
deadline = time.time() + args.duration
session_index = 0
while time.time() < deadline:
if session_index > 0:
_wait()
phase = session_index % 3
if phase == 0:
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(
base, paths, profile, session_index, ip_entry
)
logger.debug(" trail: %s", result.get("trail", []))
elif phase == 1:
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
else:
ip_entry = ABUSE_SOURCE_IPS[session_index % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, session_index // 3)
logger.debug(
" hits: %s",
[r.get("path") for r in result.get("results", [])],
)
results.append(result)
session_index += 1
ok = sum(1 for r in results if "error" not in r) ok = sum(1 for r in results if "error" not in r)
logger.info("Done: %d/%d requests succeeded.", ok, len(results)) logger.info("Done: %d/%d requests succeeded.", ok, len(results))